In-vehicle virtual sound field playback method and device based on time domain simplification processing, electronic equipment and storage medium
Through the method based on time domain simplification processing, the virtual sound field playback technology on the vehicle embedded platform is optimized, which solves the problem that the existing technology cannot achieve real-time operation and poor sense of reality, and realizes efficient virtual sound field playback on the vehicle platform.
Patent Information
- Application Number
- CN202510135410.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-07
- Publication Date
- 2025-05-30
AI Technical Summary
The existing virtual sound field playback technology cannot be run in real time on the vehicle-mounted embedded platform, and the virtual sound field has poor realism and excessive computing power and memory consumption.
Using a time domain simplification process method, by obtaining the measured impulse response sequence of a specified scene, simplifying it according to the set threshold, the main energy sequence is obtained, and the input signal is delayed according to the maximum delay value of the main energy sequence, and a virtual sound field playback signal is generated by combining amplitude weighting summation processing.
Real-time playback of virtual sound field is realized on the existing vehicle-mounted embedded platform, and the playback effect of virtual scenes is good, reducing the consumption of computing power and memory resources.
Smart Images

Figure CN120075725A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method, device, electronic device and storage medium for in-vehicle virtual sound field reproduction based on time-domain simplified processing. Background Art
[0002] With the technological progress of the automotive manufacturing industry and the continuous improvement of people's living standards, the number and quality of speakers in the car cabin are also constantly improving, which has promoted the rapid development of in-vehicle sound field reproduction technology. In addition, with the rapid development of artificial intelligence technology, new technological products such as the metaverse, digital twin, VR (virtual reality), AR (augmented reality), and MR (mixed reality) have emerged continuously. During the reproduction of program sources such as movies, games, and music, the demand for virtual sound field reproduction is becoming increasingly strong. Especially in the narrow space inside the car, car manufacturers and consumers hope to experience the real effects of various playback scenarios such as opera houses, cinemas, football fields, and clubs through more speakers arranged.
[0003] In response to the application requirements of virtual sound field reproduction, many audio effect software production companies have developed audio effect software for virtual sound field reconstruction based on the PC platform. However, these audio effect software are all developed based on the computing power of the PC's CPU. These audio effect algorithms occupy a very large amount of computing power resources during actual operation. Because of the too large computing power consumption, these audio effect algorithms cannot be directly applied to the embedded chip platform. Although there are already some methods for realizing reverberant sound fields, such as using time-domain convolution operations to perform real-time convolution processing on input signals, however, when multiple signals are replayed with reverberation output, it will consume a very large amount of computing resources, resulting in the inability to complete real-time operation on the existing in-vehicle embedded platform.
[0004] Generally speaking, there are two existing ways to realize virtual sound field reproduction: Way 1 is to use an artificial reverberation model and simulate the reverberation characteristics of the indoor playback space through a multi-stage comb filter. The virtual sound field generated by this artificial reverberation method has a large gap from the real sound field in the real scene, and the realism of the scene is relatively poor from the perspective of the experience effect. Way 2 is to use the actually measured impulse response and perform impulse response convolution operation on the sound source signal to realize the reproduction of the virtual sound field. This implementation method based on the measured impulse response and time-domain convolution requires a very high computing power overhead and also a large memory overhead, and cannot be realized in real time on the in-vehicle embedded processing platform.
[0005] The above information disclosed in the background art section is only used to strengthen the understanding of the background of the present application. Therefore, it may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention
[0006] In view of the above problems, the present invention provides a method, device, electronic device and storage medium for in-vehicle virtual sound field reproduction based on time-domain simplification processing, which solves the problem of excessive computing power consumption of the reverberant sound field time-domain convolution algorithm, and ensures that real-time reproduction of the virtual sound field and good authenticity of the virtual scene reproduction effect can be achieved on the existing in-vehicle embedded platform.
[0007] The first aspect of the present invention provides a method for in-vehicle virtual sound field reproduction, which includes the following steps:
[0008] S1. Obtain the measured impulse response sequence of a specified scene;
[0009] S2. Simplify the impulse response sequence according to a set threshold level to obtain the main energy sequence of the impulse response sequence;
[0010] S3. Find the maximum delay value of the main energy sequence of the impulse response sequence, and delay the input signal according to the maximum delay value to obtain the delayed sequence of the input signal;
[0011] S4. Numerically extract the delayed sequence of the input signal in turn according to the delay value of the main energy sequence to obtain the input signal data frame delayed according to the delay value of the main energy sequence; then perform dot product processing on the input signal data frame of the delayed processing obtained through extraction in turn according to the amplitude value of the main energy sequence to obtain the delayed input signal frame after amplitude weighting processing; then sum up all the delayed input signals after amplitude weighting to obtain the virtual sound field reproduction signal.
[0012] In a preferred embodiment, in step S1, select the sound source emission position and the positions of the listener's two ears in the specified scene, generate an impulse sound source at the sound source emission position, and obtain the impulse response sequence at the positions of the listener's two ears as the measured impulse response sequence.
[0013] In a preferred embodiment, in step S2, set a screening threshold Th, and compare the measured impulse response sequences h L and h R at the positions of the listener's two ears with the screening threshold Th respectively, remove the impulse response sequences smaller than the threshold Th, and only retain the impulse response sequences greater than or equal to the threshold Th to obtain the main energy sequence of the impulse response sequence at the positions of the listener's two ears; among them, the main energy sequence of the impulse response sequence at the position of the listener's left ear includes the amplitude value A S,L (n) and the delay value T S,L (n), and the selected sequence length is N S,L ; the main energy sequence of the impulse response sequence at the position of the listener's right ear includes the amplitude value A S,R (n) and the delay value TS,R (n), the selected sequence length is N S,R .
[0014] In a preferred embodiment, in step S3,
[0015] Find the delay value T corresponding to the Nth sequence of the impulse response main energy sequence at the position of the listener's left ear S,L The delay value T corresponding to the Nth sequence of the impulse response main energy sequence at the position of the listener's left ear S,L (N S,L ) is denoted as the maximum delay value Tmax of this main energy sequence S,L , assuming that the number of sample points in a single data frame of the input signal is K, calculate the number of data frames M that need to be delayed when obtaining the maximum delay value of the input signal according to the following formula L :
[0016]
[0017] Find the delay value T corresponding to the Nth sequence of the impulse response main energy sequence at the position of the listener's right ear S,R The delay value T corresponding to the Nth sequence of the impulse response main energy sequence at the position of the listener's right ear S,R (N S,R ) is denoted as the maximum delay value Tmax of this main energy sequence S,R , assuming that the number of sample points in a single data frame of the input signal is K, calculate the number of data frames that need to be delayed when the input signal generates the maximum delay value as M R :
[0018]
[0019] In the above formula, is the ceiling operation
[0020] In a preferred embodiment, step S4 includes
[0021] Step S400: The left-channel input signal after being processed by the maximum delay value Tmax S,L is X T,L , according to the delay value T of the main energy sequence S,L (n), successively in the input signal X after the maximum delay T,L sequence, according to the time position where the delay value T S,L (n) is located, take out K sample points to generate the delay data frame X corresponding to the delay value T S,L (n) L (N - T S,L (n)); The right-channel input signal after being processed by the maximum delay value Tmax S,R is X T,R , according to the delay value T of the main energy sequence S,R (n), successively in the input signal X after the maximum delayT,R In the sequence, according to the time position where the delay value T S,R (n) is located, K sample points are taken to generate a delay data frame X corresponding to the delay value T S,R (n); R (N - T S,R (n));
[0022] Step S401: According to the amplitude value A S,L of the main energy sequence of the impulse response at the position of the listener's left ear, the input signal data frame X L after decimation and delay processing is successively subjected to dot product processing to obtain the amplitude-weighted delay data frame A S,L (n) × X S,L (N - T L (n)); Then, all the delayed input signals after amplitude weighting are summed to obtain the left channel signal frame X S,L for virtual sound field reproduction; L,R (n);
[0023] Step S402: According to the amplitude value A S,R of the main energy sequence of the impulse response at the position of the listener's right ear, the input signal data frame X R after decimation and delay processing is successively subjected to dot product processing to obtain the amplitude-weighted delay data frame A S,R (n) × X S,R (N - T R (n)); Then, all the delayed input signals after amplitude weighting are summed to obtain the left channel signal frame X S,R (n). R,R (n).
[0024] In a preferred embodiment, in step S4, according to the order of increasing delay time, the main energy sequence of the impulse response sequence at the position of the listener's left ear is divided into multiple components. For each component, step S401 is respectively executed to obtain multiple sound field parts after amplitude weighting and summation, and the timbre of each sound field part is adjusted using an equalizer;
[0025] In step S4, according to the order of increasing delay time, the main energy sequence of the impulse response sequence at the position of the listener's right ear is divided into multiple components. For each component, step S402 is respectively executed to obtain multiple sound field parts after amplitude weighting and summation, and the timbre of each sound field part is adjusted using an equalizer.
[0026] In a more preferred embodiment, in step S4, the multiple sound field portions include three portions: an early reflection sound field, a mid-term scattering sound field, and a late diffusion sound field;
[0027] Among them, according to the maximum delay value M1 of the early reflection sound field and the maximum delay value M2 of the mid-term scattering sound field, the main energy sequence of the impulse response at the position of the listener's left ear is divided from 0 to the maximum delay value Tmax S,L into three time periods. The first period is the early reflection sound field, which is composed of the first N1 impulse sequences of the main energy sequence of the impulse response, and step S401 is executed, so that the early reflection sound field portion of the left-channel virtual sound field output signal is generated by the amplitude weighted sum and summation processing of the delay signals; the second period is the mid-term scattering sound field, which is composed of the middle N2 impulse sequences of the main energy sequence of the impulse response, and step S401 is executed, so that the mid-term scattering sound field portion of the left-channel virtual sound field output signal is generated by the amplitude weighted sum and summation processing of the delay signals; the third period is the late diffusion sound field, which is composed of the last N3 impulse sequences, and step S401 is executed, so that the late diffusion sound field portion of the left-channel virtual sound field output signal is generated by the amplitude weighted sum and summation processing of the delay signals;
[0028] According to the maximum delay value M1 of the early reflection sound field and the maximum delay value M2 of the mid-term scattering sound field, the main energy sequence of the impulse response at the position of the listener's right ear is divided from 0 to the maximum delay value Tmax S,L into three time periods. The first period is the early reflection sound field, which is composed of the first N1 impulse sequences of the main energy sequence of the impulse response, and step S402 is executed, so that the early reflection sound field portion of the right-channel virtual sound field output signal is generated by the amplitude weighted sum and summation processing of the delay signals; the second period is the mid-term scattering sound field, which is composed of the middle N2 impulse sequences of the main energy sequence of the impulse response, and step S402 is executed, so that the mid-term scattering sound field portion of the right-channel virtual sound field output signal is generated by the amplitude weighted sum and summation processing of the delay signals; the third period is the late diffusion sound field, which is composed of the last N3 impulse sequences, and step S402 is executed, so that the late diffusion sound field portion of the right-channel virtual sound field output signal is generated by the amplitude weighted sum and summation processing of the delay signals.
[0029] The second aspect of the present invention provides an in-vehicle virtual sound field reproduction device, which includes:
[0030] An impulse response sequence reading module, which is used to obtain the measured impulse response sequence of a specified scenario;
[0031] An impulse response sequence selection module, which is used to simplify the impulse response sequence according to a set threshold level to obtain the main energy sequence of the impulse response sequence;
[0032] A delay processing module, which is configured to receive an input signal, find the maximum delay value of the main energy sequence of the impulse response sequence, and perform delay processing on the input signal according to the maximum delay value to obtain a delay sequence of the input signal;
[0033] A delay data frame extraction module, which is configured to sequentially extract numerical values from the delay sequence of the input signal according to the delay value of the main energy sequence to obtain an input signal data frame that is delay-processed according to the delay value of the main energy sequence;
[0034] A main energy sequence decomposition module, which is configured to divide the main energy sequence of the impulse response sequence into multiple components in the order of increasing delay time;
[0035] An amplitude weighting module, which is configured to perform dot product processing on the delay-processed input signal data frames obtained by extraction in sequence for each component according to the amplitude value of the response component of the main energy sequence to obtain a delay input signal frame that is amplitude-weighted;
[0036] A signal summation module, which is configured to sum up all the delay input signals that are amplitude-weighted for each component to obtain a virtual sound field reproduction signal for multiple sound field parts.
[0037] In a preferred embodiment, the impulse response sequence reading module includes an artificial head disposed at the position of a listener in a specified scenario and binaural microphones disposed at the two ears of the artificial head.
[0038] A third aspect of the present invention provides an electronic device, which includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the in-vehicle virtual sound field reproduction method described above is implemented.
[0039] In an embodiment, the electronic device may be an in-vehicle embedded platform or a component of an in-vehicle embedded platform.
[0040] A fourth aspect of the present invention provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the in-vehicle virtual sound field reproduction method described above is implemented.
[0041] The present invention adopts the above solutions and has the following advantages compared with the prior art:
[0042] (1) The present invention adopts a measured impulse response sequence, and can obtain the impulse response sequences of multiple scenarios, ensuring that the reproduction effect of the virtual scenario has a certain degree of authenticity, and the experience of the virtual scenario reproduction effect will exceed the artificial virtual scenario reproduction method.
[0043] (2) The present invention proposes to perform time-domain simplification processing on the measured impulse response sequences in multiple scenarios, reducing the number of time-domain impulse response sequences while retaining the main energy sequences of the time-domain impulse response sequences, thereby ensuring that the experience of virtual scene replay is basically the same as that of the original impulse response sequence.
[0044] (3) The virtual scene replay method proposed by the present invention occupies relatively small computing power resources and memory resources, and is very suitable for implementing the virtual sound field replay function in in-vehicle hosts and power amplifier products, while achieving basically the same replay effect as the real scene. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] In order to more clearly illustrate the technical solutions of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0046] Figure 1 FIG. is a flowchart of an in-vehicle virtual sound field replay method according to an embodiment of the present invention.
[0047] Figure 2 FIG. is a signal processing flowchart of in-vehicle virtual sound field replay according to an embodiment of the present invention.
[0048] Figure 3 FIG. is a schematic diagram of the division of three parts of a virtual sound field according to an embodiment of the present invention.
[0049] Figure 4 FIG. is a signal processing flowchart of three sound field parts of in-vehicle virtual sound field replay according to an embodiment of the present invention.
[0050] Figure 5 FIG. is a module diagram of an in-vehicle virtual sound field storage device according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0051] The following will elaborate on the preferred embodiments of the present invention in conjunction with the drawings, so that the advantages and features of the present invention can be more easily understood by those skilled in the art. It should be noted here that the description of these embodiments is for helping to understand the present invention, but does not constitute a limitation to the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0052] Regarding the chip platform used in in-vehicle power amplifiers and car radios, in view of the problem that the computing power overhead of the reverberant sound field time-domain convolution algorithm is too large, this embodiment proposes a method and device for in-vehicle virtual sound field replay based on time-domain simplification processing, ensuring that real-time replay of the virtual sound field can be achieved on the existing in-vehicle embedded platform.
[0053] Figure 1 A method for in-vehicle virtual sound field replay based on time-domain simplification processing is shown, and the specific process is as follows.
[0054] (1) Obtain the measured impulse response sequence of a specified scenario. Specifically, select the sound source emission position and the positions of the listener's two ears in the specified scenario. Generate an impulse sound source at the sound source emission position, and obtain the impulse response sequence at the positions of the listener's two ears as the measured impulse response sequence.
[0055] To produce the virtual sound field replay effect of a specified scenario, it is first necessary to measure the impulse response of the specified scenario to obtain the measured data of the impulse response. To obtain the measured data of the impulse response, it is necessary to select the position where the sound source emits in the specified scenario, use an air gun to generate an impulse sound source at the sound source emission position, and at the same time, it is also necessary to select the position where the listener is located in the audience seat area, and use an artificial head and a binaural microphone to obtain the impulse response sequence at the two ears of the artificial head. Assume that the impulse response sequence at the left ear of the artificial head is h L and the impulse response sequence at the right ear of the artificial head is h R .
[0056] It should also be noted that multiple measured impulse response sequences of specified scenarios can be obtained, so as to achieve virtual sound field replay of multiple scenarios.
[0057] (2) Simplify the time-domain impulse response sequence according to the set threshold level to obtain the main energy sequence of the impulse response sequence.
[0058] To obtain the main energy sequence of the measured impulse response sequence, it is first necessary to set a screening threshold Th, compare the measured impulse response sequences h L and h R with the screening threshold Th respectively, remove the impulse sequences smaller than the threshold Th, and only retain the impulse response sequences greater than or equal to the threshold Th, so as to obtain the main energy sequence of the measured impulse response sequence.
[0059] The main energy sequence of the impulse response sequence at the left ear of the artificial head is h S,L , and its expression is as follows:
[0060]
[0061] Similarly, the main energy sequence of the impulse response sequence at the right ear of the artificial head is hS,R , and its expression is as follows:
[0062]
[0063] The main energy sequence of the artificial head's left ear impulse response includes: amplitude value A S,L (n) and delay value T S,L (n). To reduce the computational overhead, consider truncating the main energy sequence, and the selected sequence length is N S,L . Similarly, the main energy sequence of the artificial head's right ear impulse response includes: amplitude value A S,R (n) and delay value T S,R (n), and the selected sequence length is N S,R .
[0064] (3) Find the maximum delay value of the main energy sequence of the measured impulse response sequence, and delay the input signal according to the maximum delay value to obtain the delayed sequence of the input signal.
[0065] Find the delay value T S,L corresponding to the N S,L -th sequence of the main energy sequence of the left ear impulse response S,L and denote it as the maximum delay value Tmax of this main energy sequence S,L . Assume that the number of sample points in a single data frame of the input signal is K. From Tmax S,L , the number of data frames M that need to be delayed when the input signal generates the maximum delay value can be obtained L , and its calculation formula is as follows:
[0066]
[0067] where is the ceiling operation;
[0068] Similarly, find the delay value T S,R corresponding to the N S,R -th sequence of the main energy sequence of the right ear impulse response S,R and denote it as the maximum delay value Tmax of this main energy sequence S,R . Assume that the number of sample points in a single data frame of the input signal is K. From Tmax S,R , the number of data frames M that need to be delayed when the input signal generates the maximum delay value can be obtained R , and its calculation formula is as follows:
[0069]
[0070] Assume that the signals of the left and right channels of the input stereo are X L and X R, according to the number M of data frames to be delayed when generating the maximum delay value L and M R , respectively perform maximum delay processing on the input stereo signals X L and X R according to the delay values Tmax S,L and Tmax S,R to obtain the input signals X T,L and X T,R after maximum delay.
[0071] (4) According to the delay values of the impulse response main energy sequence, numerically extract the delay sequence of the input signal in turn to obtain the input signal data frames processed by delay according to the delay values of the main energy sequence. Then, according to the amplitude values of the impulse response main energy sequence, perform dot product processing on the input signal data frames processed by delay obtained through extraction in turn to obtain the delay input signal frames processed by amplitude weighting. Finally, sum up all the delay input signals after amplitude weighting to obtain the signal frames for virtual sound field reproduction.
[0072] As Figure 2 shown, the left-channel input signal X L is divided into multiple data blocks according to the data frame length K, and the maximum delay value Tmax S,L is generated according to the left-ear impulse response main energy sequence, and the number of data frames to be delayed is M L . When implemented on an embedded platform, only one memory space area needs to be allocated to store all M L× K sample point data, which can effectively save the memory space overhead. The input signal after processing by the maximum delay value Tmax S,L is X T,L , and then according to the delay value T S,L (n) of the main energy sequence, in the input signal X T,L sequence after maximum delay in turn, according to the time position where the delay value T S,L (n) is located, take out K sample points to generate the delay data frame X S,L (n) corresponding to the delay value T L (N - T S,L (n)). Then, according to the amplitude value A S,L (n) of the impulse response main energy sequence, perform dot product processing on the input signal data frames X L (N - T S,L (n)) processed by delay obtained through extraction in turn to obtain the delay data frames A S,L (n) × X L (N - T S,L(n)). Finally, sum up all the delayed input signals after amplitude weighting to obtain the signal frame X for virtual sound field reproduction. L,R (n). Similarly, for the right-channel input signal X R it will also be delayed according to the delay values T S,R (n) of the main energy sequence of the right-ear impulse response to obtain the delayed data frame X R (N - T S,R (n)), and then weighted according to the amplitude value A S,R (n) to obtain the amplitude-weighted delayed data frame A S,R (n) × X R (N - T S,R (n)). Finally, sum up all the delayed input signals after amplitude weighting to obtain the signal frame X R,R (n).
[0073] Furthermore, divide the maximum delay value of the main energy sequence of the impulse response into three components according to the increasing delay time, so that the main energy sequence of the impulse response can be divided into three parts: the early reflection sound field part (Early), the mid-term scattering sound field part (Cluster), and the late diffusion sound field part (Diffuse).
[0074] As Figure 3 shown, according to the maximum delay value M1 of the early reflection sound field and the maximum delay value M2 of the mid-term scattering sound field, divide the main energy sequence of the left-ear impulse response from 0 to the maximum delay value Tmax S,L into three time periods. The first period is the early reflection sound field, which consists of the first N1 pulse sequences of the main energy sequence of the impulse response. The second period is the mid-term scattering sound field, which consists of the middle N2 pulse sequences of the main energy sequence of the impulse response. The third period is the late diffusion sound field, which consists of the last N3 pulse sequences.
[0075] As Figure 4 shown, according to this division method, the amplitude weighting and summation processes in step (4) will be divided into three components, respectively generating the three components of the early reflection, mid-term scattering, and late diffusion of the virtual sound field.
[0076] Similarly, the main energy sequence of the right-ear impulse response is also divided into three parts in this way, and the three sound field parts corresponding to the right-channel virtual sound field output signal are generated by the amplitude weighting and summation of the delayed signals.
[0077] Use an EQ equalizer to adjust the timbre effects of the early reflection, mid-term scattering, and late diffusion of the virtual sound field respectively.
[0078] Figure 5shows an in-vehicle virtual sound field reproduction device according to this embodiment. Refer to Figure 5 As shown, the in-vehicle virtual sound field reproduction device includes: a pulse response sequence reading module 101, a pulse response sequence selection module 102, a delay processing module 103, a delayed data frame extraction module 104, a main energy sequence decomposition module 105, an amplitude weighting module 106, and a signal accumulation and summation module 107.
[0079] The pulse response sequence reading module 101 is configured to read the pulse response sequences actually measured in various scenarios and send them to the pulse response sequence selection module 102. Specifically, the pulse response sequence reading module 101 includes an artificial head arranged at the position of the listener in a specified scenario and binaural microphones arranged at the two ears of the artificial head, and the pulse response sequences at the two ears of the artificial head are collected through the binaural microphones.
[0080] The input end of the pulse response sequence selection module 102 is electrically connected to the output end of the pulse response sequence reading module 101. The pulse response sequence selection module 102 is configured to simplify the read pulse response sequences according to a certain set threshold to obtain the main energy sequence of the pulse response sequences, and send it to the delay processing module 103.
[0081] The delay processing module 103 has a first input end and a second input end. The first input end is electrically connected to the output end of the pulse response sequence selection module 102 to receive the main energy sequence of the pulse response sequences; the second input end is used to access the input signal frame. The delay processing module 103 delays the input signal frame according to the maximum delay value of the pulse response main energy sequence to generate a delayed input signal frame, and sends it to the delayed data frame extraction module 104.
[0082] The input end of the delayed data frame extraction module 104 is electrically connected to the output end of the delay processing module 103. The delayed data frame extraction module 104 is configured to perform time positioning on the delayed input signal frame according to the delay value of the pulse response main energy sequence, and extract one frame of data at each delay position corresponding to the main energy sequence as the delayed data frame after being processed by this delay value, and send it to the main energy sequence decomposition module 105.
[0083] The main energy sequence decomposition module 105 is specifically an early, middle, and late main energy sequence decomposition module, and its input end is electrically connected to the output end of the delayed data frame extraction module 104. The main energy sequence decomposition module 105 is configured to divide the main energy sequence into three pulse sequences of early, middle, and late according to the delay time from short to long, and send them to the amplitude weighting module. Specifically, as Figure 3As shown, according to the order of the delay time from short to long, the delay time period is divided into three components. In this way, the main energy sequence of the impulse response can be divided into three parts: the early reflection sound field part (Early), the mid-term scattering sound field part (Cluster), and the late diffusion sound field part (Diffuse).
[0084] The input end of the amplitude weighting module 106 is electrically connected to the output end of the main energy sequence decomposition module 105. The amplitude weighting module 106 is used to perform amplitude weighting processing on the delay data frames respectively by using the early, mid-term, and late pulse sequences, generate the data frames after amplitude weighting processing, and send them to the signal accumulation and summation module 107.
[0085] The input end of the signal accumulation and summation module 107 is electrically connected to the output end of the amplitude weighting module 106. The signal accumulation and summation module 107 is used to accumulate and sum all the data frames processed by the amplitude weighting 106 to generate the virtual sound field output data frames.
[0086] This embodiment also provides an electronic device, which includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the in-vehicle virtual sound field replay method described above. This electronic device can be an in-vehicle embedded platform or a component of an in-vehicle embedded platform.
[0087] This embodiment also provides a computer-readable storage medium. A computer program is stored on the computer-readable storage medium. When this program is executed by a processor, it implements the constant beam shape design method described above. The computer-readable storage medium includes, but is not limited to, FLASH memory, read-only memory, magnetic disks, or optical discs.
[0088] In this embodiment, the measured impulse response sequences of multiple scenarios are adopted, which ensures that the replay effect of the virtual scenario has a certain degree of authenticity, and the experience of the virtual scenario replay effect will exceed the replay method of artificial virtual scenarios. It is proposed to perform time-domain simplification processing on the measured impulse response sequences of multiple scenarios, which reduces the number of time-domain impulse response sequences, while retaining the main energy sequence of the time-domain impulse response sequences. Thus, in the experience of virtual scenario replay, it is ensured that the experience is basically consistent with that of the original impulse response sequence. The occupied computing power resources and memory resources are relatively small, which is very suitable for implementing the virtual sound field replay function in in-vehicle host and power amplifier products, and at the same time, the replay effect is basically consistent with the real scenario.
[0089] Unless otherwise specifically defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art. If the definitions used herein are contradictory or inconsistent with the definitions contained in other published documents, the definitions used herein shall prevail.
[0090] As shown in this specification and the claims, the terms "comprising" and "including" merely indicate the inclusion of the expressly identified steps and elements, and these steps and elements do not constitute an exclusive listing. A method or apparatus may also include other steps or elements. The term "and / or" as used herein includes any combination of one or more of the related listed items.
[0091] It should be understood that various parts of the present application can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one of the following techniques known in the art or a combination thereof can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application-specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0092] Those of ordinary skill in the art of this technology can understand that all or part of the steps carried by the method of the above embodiments can be completed by instructing relevant hardware through a program. The said program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiments.
[0093] In addition, each functional unit in the embodiments can be integrated into a processing module, or each unit can exist physically alone, or two or more units can be integrated into one module. The above integrated module can be implemented in the form of hardware or in the form of a software functional module. When the above integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.
[0094] The above embodiments are only for illustrating the technical concept and features of the present invention and are a preferred embodiment. The purpose is to enable those who are familiar with this technology to understand the content of the present invention and implement it accordingly, and it is not intended to limit the protection scope of the present invention. Any equivalent transformation or modification made according to the spirit and essence of the present invention should be covered within the protection scope of the present invention.
Claims
1. A method for reproducing a virtual sound field in a vehicle, characterized in that: The steps include: S1, obtaining the measured impulse response sequence of the specified scene; S2, simplifying the impulse response sequence according to the set threshold level to obtain the main energy sequence of the impulse response sequence; S3, finding the maximum delay value of the main energy sequence of the impulse response sequence, performing delay processing on the input signal according to the maximum delay value, and obtaining a delay sequence of the input signal; S4. According to the delay value of the main energy sequence, the delay sequence of the input signal is numerically extracted in sequence to obtain the input signal data frame that is delayed according to the delay value of the main energy sequence; then, according to the amplitude value of the main energy sequence, the input signal data frame that is delayed after extraction is subjected to dot multiplication processing in sequence to obtain the delayed input signal frame that is subjected to amplitude weighting processing; then, all the delayed input signals that are amplitude weighted are added together to obtain the virtual sound field playback signal.
2. The in-vehicle virtual sound field playback method according to claim 1, characterized in that: In step S1, a sound source emission position and the positions of the listener's ears are selected in a specified scene, a pulse sound source is generated at the sound source emission position, and an impulse response sequence at the positions of the listener's ears is obtained as the measured impulse response sequence.
3. The in-vehicle virtual sound field playback method according to claim 1, characterized in that: In step S2, a screening threshold Th is set, and the measured impulse response sequence h at the binaural position of the listener is L and h R , respectively compared with the screening threshold Th, the impulse response sequence less than the threshold Th is removed, and only the impulse response sequence greater than or equal to the threshold Th is retained to obtain the main energy sequence of the impulse response sequence at the binaural position of the listener; wherein the main energy sequence of the impulse response sequence at the left ear position of the listener includes the amplitude value A S,L (n) and delay value T S,L (n), the selected sequence length is N S,L The main energy sequence of the impulse response sequence at the listener's right ear position includes the amplitude value A S,R (n) and delay value T S,R (n), the selected sequence length is N S,R .
4. The in-vehicle virtual sound field playback method according to any one of claims 1 to 3, characterized in that: In step S3, Find the Nth main energy sequence of the impulse response at the listener's left ear S,L The delay value T corresponding to the sequence S,L (N S,L ) is recorded as the maximum delay value Tmax of the main energy sequence S,L Assuming that the number of sample points of a single data frame of the input signal is K, the number of data frames required to delay when obtaining the maximum delay value of the input signal is calculated according to the following formula: L : Find the Nth main energy sequence of the impulse response at the listener's right ear S,R The delay value T corresponding to the sequence S,R (N S,R ) is recorded as the maximum delay value Tmax of the main energy sequence S,R Assuming that the number of sample points of a single data frame of the input signal is K, the number of data frames required to delay when the input signal produces the maximum delay value is calculated according to the following formula: R : In the above formula, It is a round-up operation.
5. The in-vehicle virtual sound field playback method according to any one of claims 1 to 3, characterized in that: Step S4 includes step S400: after a maximum delay value Tmax S,L The processed left channel input signal is X T,L , according to the main energy sequence delay value T S,L (n), respectively, the input signal X after the maximum delay T,L In the sequence, according to T S,L (n) The time position of the delay value, take out K sample points, and generate the corresponding delay value T S,L (n) delayed data frame X L (NT S,L (n)); after the maximum delay value Tmax S,R The processed right channel input signal is X T,R , according to the main energy sequence delay value T S,R (n), respectively, the input signal X after the maximum delay T,R In the sequence, according to T S,R (n) The time position of the delay value, take out K sample points, and generate the corresponding delay value T S,R (n) delayed data frame X R (NT S,R (n)); Step S401: according to the amplitude value A of the main energy sequence of the impulse response at the left ear position of the listener S,L (n), sequentially process the input signal data frame X after the delay processing obtained by extraction L (NT S,L (n)) is processed by point multiplication to obtain the delayed data frame A after amplitude weighting processing S,L (n)×X L (NT S,L (n)); then add all the delayed input signals after amplitude weighting to obtain the left channel signal frame X of the virtual sound field playback L,R (n); Step S402: according to the amplitude value A of the main energy sequence of the impulse response at the listener's right ear position S,R (n), sequentially process the input signal data frame X after the delay processing obtained by extraction R (NT S,R (n)) is processed by point multiplication to obtain the delayed data frame A after amplitude weighting processing S,R (n)×X R (NT S,R (n)); then add all the delayed input signals after amplitude weighting to obtain the left channel signal frame X of the virtual sound field playback R,R (n).
6. The in-vehicle virtual sound field playback method according to claim 5, characterized in that: In step S4, the main energy sequence of the impulse response sequence at the left ear position of the listener is divided into a plurality of components in the order of delay time from short to long, and for each component, step S401 is performed respectively to obtain a plurality of corresponding sound field parts after amplitude weighting and summing, and the timbre of each sound field part is adjusted by using an equalizer respectively; In step S4, the main energy sequence of the impulse response sequence at the right ear position of the listener is divided into multiple components in the order of delay time from short to long. For each component, step S402 is executed respectively to obtain the corresponding multiple sound field parts after amplitude weighting and addition, and the timbre of each sound field part is adjusted using an equalizer.
7. The in-vehicle virtual sound field playback method according to claim 6, characterized in that: In step S4, the multiple sound field parts include three parts: early reflection sound field, mid-scattering sound field and late diffusion sound field; Among them, according to the maximum delay value M1 of the early reflected sound field and the maximum delay value M2 of the mid-term scattered sound field, the main energy sequence of the impulse response at the left ear position of the listener is extended from 0 to the maximum delay value Tmax S,L The range of is divided into three time periods, the first section is an early reflection sound field, which is composed of the first N1 pulse sequences of the main energy sequence of the impulse response, and step S401 is executed, so that the early reflection sound field part of the left channel virtual sound field output signal is generated after the amplitude weighting and addition processing of the delayed signal; the second section is a mid-term scattered sound field, which is composed of the middle N2 pulse sequences of the main energy sequence of the impulse response, and step S401 is executed, so that the mid-term scattered sound field part of the left channel virtual sound field output signal is generated after the amplitude weighting and addition processing of the delayed signal; the third section is a late diffuse sound field, which is composed of the rear end N3 pulse sequences, and step S401 is executed, so that the late diffuse sound field part of the left channel virtual sound field output signal is generated after the amplitude weighting and addition processing of the delayed signal; According to the maximum delay value M1 of the early reflected sound field and the maximum delay value M2 of the mid-term scattered sound field, the main energy sequence of the impulse response at the listener's right ear position is extended from 0 to the maximum delay value Tmax. S,L The range of is divided into three time periods, the first section is the early reflection sound field, which is composed of the first N1 pulse sequences of the pulse response main energy sequence, and step S402 is executed, so that the early reflection sound field part of the right channel virtual sound field output signal is generated after the amplitude weighting and addition processing of the delayed signal; the second section is the mid-term scattered sound field, which is composed of the middle N2 pulse sequences of the pulse response main energy sequence, and step S402 is executed, so that the mid-term scattered sound field part of the right channel virtual sound field output signal is generated after the amplitude weighting and addition processing of the delayed signal; the third section is the late diffuse sound field, which is composed of the rear-end N3 pulse sequences, and step S402 is executed, so that the late diffuse sound field part of the right channel virtual sound field output signal is generated after the amplitude weighting and addition processing of the delayed signal.
8. A virtual sound field playback device in a car, characterized in that: include: An impulse response sequence reading module, which is used to obtain a measured impulse response sequence of a specified scene; An impulse response sequence selection module, which is used to simplify the impulse response sequence according to a set threshold level to obtain a main energy sequence of the impulse response sequence; A delay processing module, which is used to receive an input signal, find out the maximum delay value of the main energy sequence of the impulse response sequence, perform delay processing on the input signal according to the maximum delay value, and obtain a delay sequence of the input signal; A delayed data frame extraction module, which is used to extract the values of the delay sequence of the input signal in sequence according to the delay value of the main energy sequence, and obtain the input signal data frame that is delayed according to the delay value of the main energy sequence; A main energy sequence decomposition module, which is used to divide the main energy sequence of the impulse response sequence into multiple components in the order of delay time from short to long; An amplitude weighting module, which is used to perform point multiplication processing on the delayed input signal data frame obtained through extraction in sequence according to the amplitude value of the response component of the main energy sequence for each component, so as to obtain a delayed input signal frame after amplitude weighting processing; The signal accumulation and summing module is used to add up all delayed input signals of various components after amplitude weighting to obtain virtual sound field playback signals of multiple sound field parts.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the in-vehicle virtual sound field playback method as described in any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the in-vehicle virtual sound field playback method as claimed in any one of claims 1 to 7.