Directional gain method and apparatus for spatial sound field, device, medium, and product
By performing direction gain processing on the spatial sound field signal, the problem of the inability to adjust the volume of a specific direction in the prior art is solved, and more targeted volume control and better user experience are achieved.
Patent Information
- Application Number
- PCT/CN2024/140568
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-02-06
- Filing Date
- 2024-12-19
- Publication Date
- 2025-08-14
AI Technical Summary
Existing spatial audio technology cannot adjust the volume in a specific direction in the spatial sound field signal, affecting the user experience.
By obtaining the spatial sound field signal and applying the direction gain function for gain processing, including decoding, channel gain and encoding steps, the volume of a specific direction is enhanced or attenuated, and the pre-calculated direction gain function is stored to reduce the calculation amount, and the cross-in-out algorithm is used to smooth out the volume change.
It realizes accurate adjustment of the volume in the spatial sound field signal in a specific direction, improves the playback experience of the spatial sound field, reduces computing resource consumption and playback lag, and ensures the smoothness of volume changes.
Smart Images

Figure CN2024140568_14082025_PF_FP_ABST
Abstract
Description
Directional gain method, device, equipment, medium and product for spatial sound field
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application is based on the Chinese application with application number 202410172067.7 and application date February 6, 2024, and claims its priority. The disclosed content of the Chinese application is hereby introduced as a whole into this application. Technical Field
[0003] The present application relates to the field of sound processing technology, and in particular to a method, device, equipment, medium and product for directional gain of a spatial sound field. Background Art
[0004] Spatial audio technology is an audio technology that creates sounds that give the listener the perception of sound channels coming from different directions (including depth and height). This technology leverages human acoustic and psychoacoustic properties to simulate the natural auditory space. Spatial audio technologies include stereo, surround sound, multi-channel sound, and 3D audio. Spatial audio technology plays a vital role in modern audio processing and virtual reality technology. Summary of the Invention
[0005] A first aspect of the present application provides a directional gain method for a spatial sound field, comprising:
[0006] Acquiring spatial sound field signals;
[0007] The spatial sound field signal is subjected to gain processing according to a determined directional gain function to obtain a target sound field signal.
[0008] The second aspect of the present application provides a directional gain device for a spatial sound field, comprising:
[0009] An acquisition module, configured to acquire a spatial sound field signal;
[0010] The gain processing module is configured to perform gain processing on the spatial sound field signal according to a determined directional gain function to obtain a target sound field signal.
[0011] The third aspect of the present application proposes an electronic device, comprising a memory and a processor, wherein the memory stores a computer program, the computer program can be run on the processor, and the processor implements the method described in the first aspect when executing the computer program.
[0012] A fourth aspect of the present application provides a non-transitory computer-readable storage medium, wherein the non-transitory computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to execute the method described in the first aspect.
[0013] A fifth aspect of the present application provides a computer program product, comprising computer program instructions, wherein when the computer program instructions are executed on a computer, the computer is caused to execute the method described in the first aspect.
[0014] A sixth aspect of the present application proposes a computer program, comprising a program code, wherein when the computer program or program code is run on a computer, the computer is caused to execute the method described in the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] In order to more clearly illustrate the technical solutions in this application or related technologies, the following briefly introduces the drawings required for use in the embodiments or related technical descriptions. Obviously, the drawings described below are merely embodiments of this application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0016] FIG1 is a schematic diagram of an application scenario of an embodiment of the present application;
[0017] FIG2A is a flow chart of a method for directional gain of a spatial sound field according to an embodiment of the present application;
[0018] FIG2B is a schematic diagram of spherical harmonic basis functions for the combination (l=3, m=3) according to an embodiment of the present application;
[0019] FIG3 is a structural block diagram of a directional gain device for a spatial sound field according to an embodiment of the present application;
[0020] FIG4 is a schematic structural diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION
[0021] It is understandable that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) must comply with the requirements of relevant laws, regulations and relevant provisions.
[0022] The principles and spirit of the present application will be described below with reference to several exemplary embodiments. It should be understood that these embodiments are provided solely to enable those skilled in the art to better understand and implement the present application, and are not intended to limit the scope of the present application in any way. Rather, these embodiments are provided to make the present application more thorough and complete, and to fully convey the scope of the present application to those skilled in the art.
[0023] It is understandable that before using the technical solutions of each embodiment of the present disclosure, the type, scope of use, usage scenarios, etc. of the personal information involved will be informed to the user in an appropriate manner, and the user's authorization will be obtained.
[0024] For example, in response to a user's active request, a prompt message is sent to the user to clearly inform the user that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the electronic device, application, server, storage medium, or other software or hardware that performs the operation of the disclosed technical solution based on the prompt message.
[0025] As an optional but non-limiting implementation, in response to a user's active request, the prompt information may be sent to the user in the form of a pop-up window, in which the prompt information may be presented in text form. Furthermore, the pop-up window may also contain a selection control for the user to select "agree" or "disagree" to provide personal information to the electronic device.
[0026] It is understandable that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of the present disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of the present disclosure.
[0027] It should be understood herein that any number of elements in the drawings is for illustration only and not for limitation, and any naming is only for distinction and does not have any limiting meaning.
[0028] Based on the description of the above background technology, the following situations also exist in the related art:
[0029] Ambisonics is a three-degree-of-freedom representation of a spatial sound field. Once the listener's position, the positions of all sound channels, and the scene's shape and materials are determined, the three-degree-of-freedom spatial sound field at the listener's location can be encoded into a set of spherical harmonic basis functions, capturing and reproducing the omnidirectional information of the sound field at the listener's location.
[0030] However, currently, spatial sound field signals (e.g., ambisonic signals) generated using spatial audio technology are generally limited to full-scene playback and the adjustment of the orientation of individual channels to achieve full-scene playback. However, this current method of controlling the spatial sound field only adjusts the overall volume, not the volume in specific directions. In particular, when an ambisonic signal is already present, it is impossible to adjust the volume of sounds in different directions, which affects the user experience.
[0031] In view of this, the present application proposes an improved method, device, equipment, medium and product for directional gain of a spatial sound field.
[0032] Based on the above description, the principles and spirit of the present application are explained in detail below with reference to several representative implementations of the present application.
[0033] Refer to Figure 1, which is a schematic diagram of an application scenario for the directional gain method for a spatial sound field provided in an embodiment of the present application. This application scenario includes a playback terminal 101 and at least one sound channel 102. After the playback terminal 101 determines the corresponding spatial sound field signal, it adjusts the volume of the control sound field signal in the corresponding sound direction according to the determined directional gain function to obtain a target sound field signal. This allows the target sound field signal to be played back panoramically using each sound channel 102.
[0034] The following describes the directional gain method for a spatial sound field according to an exemplary embodiment of the present application in conjunction with the application scenario of FIG1 . It should be noted that the above application scenario is only provided to facilitate understanding of the spirit and principles of the present application, and the embodiments of the present application are not limited in this respect. On the contrary, the embodiments of the present application can be applied to any applicable scenario.
[0035] An embodiment of the present application provides a directional gain method for a spatial sound field, which is applied to a playback terminal.
[0036] As shown in FIG2A , the method includes:
[0037] Step 201: Acquire a spatial sound field signal.
[0038] In some embodiments, the spatial sound field signal is a sound signal (eg, an ambisonic signal) representing a three-degree-of-freedom control sound field. The spatial sound field signal can be obtained by performing spatial determination based on the listener's position, the position of the sound channel, and the required scene shape.
[0039] As an example, the spatial sound field signal may be a sound signal recorded by a microphone, an audio signal stored locally on a playback terminal, or an audio signal obtained from a server or other terminal via a network.
[0040] Step 202 : Perform gain processing on the spatial sound field signal based on a directional gain function to obtain a target sound field signal.
[0041] In some embodiments, the directional gain function is a gain matrix that enhances or attenuates sound in one or more sound field directions, determined based on playback requirements. The directional gain function can be manually adjusted by a user of the playback terminal, or determined by a server based on playback requirements of a spatial sound field signal.
[0042] In some embodiments, the gain processing may include multiplying the spatial sound field signal with a directional gain function, so that the volume of the spatial sound field signal can be adjusted in the direction corresponding to the required gain to obtain the target sound field signal, and then the target sound field signal can be played.
[0043] Through the above scheme, a directional gain function can be determined, so that the directional gain function can perform gain processing to enhance or attenuate one or several sound directions of the spatial sound field signal, thereby realizing volume gain adjustment for one or several sound directions, and enabling more targeted volume adjustment in specific directions, thereby meeting the diverse needs of sound playback and improving the playback experience of the spatial sound field.
[0044] In some embodiments, step 202 includes:
[0045] Step I: Decoding the spatial sound field signal using a decoding function to obtain a multi-channel signal.
[0046] In some embodiments, the spatial sound field signal includes: an Ambisonic signal Wherein, l is the order corresponding to the spatial sound field signal, m=-l,-l+1,…,0,1,…,l, and t is time.
[0047] The decoding function is a decoding array D(θ,φ,l) determined according to the angle (θ,φ) of each channel, where θ is the horizontal angle of the channel, is the pitch angle of the channel.
[0048] In some embodiments, a decoding function is used to decode the spatial sound field signal to obtain a spherically arranged multi-channel signal, which forms a dense and uniform spherical channel array.
[0049] Specifically, the spatial sound field signal is multiplied by the decoding function to obtain the multi-channel signal X c(θ,φ) (t): Where c is the sound field number.
[0050] Step II: determining a channel gain function, and performing gain processing on the multi-channel signal by using the channel gain function as a directional gain function to obtain a multi-channel signal after gain.
[0051] In some embodiments, the channel gain function is a gain array G(θ,φ) composed of gain data of channels in various directions. c(θ,φ) (t) is multiplied by the channel gain function G(θ,φ), and the multi-channel signal X after gain can be obtained. c (t), X c (t) = X c(θ,φ) (t)*G(θ,φ).
[0052] Step II: Processing the multi-channel signal after gaining by using a coding function to obtain the target sound field signal.
[0053] In some embodiments, the encoding function is an encoding matrix E(θ, φ, l) corresponding to the decoding function, which is the opposite of the decoding process of the decoding function. Specifically, the encoding matrix is the inverse matrix of the decoding function E(θ, φ, l) = D(θ, φ, l) -1 . In this way, the multi-channel signal X after the gain obtained above c (t) is multiplied by the encoding function E(θ,φ,l), thus obtaining the target sound field signal After obtaining the target sound field signal, it can be transmitted to other terminals or servers, or sent to the corresponding channel for playback. During playback, the listener can feel the sound effect after the corresponding direction gain.
[0054] In the above scheme, since the spatial sound field signal is a real-time changing signal, if the corresponding channel gain function also changes in real time, it is necessary to use the above steps I to III to calculate the multi-channel signal after gain in real time to determine the target sound field signal. This ensures that the target sound field signal is updated in real time according to the real-time changes in the spatial sound field signal and the real-time changes in the channel gain function, ensuring the real-time effect of the target sound field signal.
[0055] However, in order to ensure real-time effects, the amount of calculation corresponding to steps I to III is relatively large. If the corresponding channel gain function is fixed, the following embodiment can be adopted to reduce the amount of calculation:
[0056] In some embodiments, the method further comprises:
[0057] Step I': determining a decoding function, a channel gain function, and a coding function, and multiplying the decoding function, the channel gain function, and the coding function to obtain the directional gain function.
[0058] In some embodiments, the decoding function is a decoding array D(θ, φ, l) determined according to the angle (θ, φ) of each channel, the channel gain function is a gain array G(θ, φ) composed of gain data of the channels in each direction, and the encoding function is an encoding function E(θ, φ, l) corresponding to the decoding function.
[0059] In some embodiments, since the decoding function, the channel gain function and the encoding function are all fixed, if they are used for each spatial sound field signal that changes in real time, All of them need to be multiplied, which requires a lot of calculations, occupies more computing resources, and takes a long time to calculate. This will easily cause freezes and sound drops, affecting the playback experience.
[0060] Therefore, it is necessary to multiply the fixed decoding function, channel gain function and encoding function in sequence to obtain the directional gain function Mpp =D(θ,φ,l)*G(θ,φ)*D(θ,φ,l) -1 , the directional gain function M pp Store. In this way, there is no need to repeatedly perform multiplication calculations in the future. The above calculation direction gain function M pp It is only calculated once and can be retrieved directly when needed later.
[0061] Step 202 specifically includes:
[0062] Step II': retrieve the stored directional gain function, and multiply the spatial sound field signal by the directional gain function to obtain the target sound field signal.
[0063] In some embodiments, due to the spatial sound field signal It changes in real time, so we only need to call the stored directional gain function M pp , the spatial sound field signal With the directional gain function M pp The target sound field signal can be obtained by multiplying
[0064] Through the above scheme, since the gain signal M pp It has been pre-calculated and stored, so only one multiplication needs to be calculated for the real-time changing spatial sound field signal, which saves a lot of computing resources and shortens the calculation time, thereby quickly and accurately obtaining the target sound field signal.
[0065] In some embodiments, the process of determining the vocal channel gain function includes:
[0066] Step a1: Determine the gain corresponding to the angle of each channel.
[0067] Step a2: performing array combination on the channel gain amounts corresponding to the angles of the various channels to obtain the channel gain function.
[0068] In some embodiments, the gain amount is the gain size corresponding to each channel, and then the gain amounts corresponding to each specific channel are arrayed in an appropriate manner, for example, arrayed according to the angular order of each channel (for example, any appropriate angular order, from large to small), to obtain a gain matrix G(θ, φ) as the channel gain function.
[0069] Through the above solution, the channel gain function can be accurately obtained, and the channel gain function is an array, which makes it easier for the computer to perform multiplication calculation processing in step 202, making the calculation faster and more accurate.
[0070] In some embodiments, the vocal tract gain function obtained in step a2 may still be inaccurate. To further improve the accuracy of the vocal tract gain function, after step a2, the method further includes:
[0071] Step a3: determine the gain weight corresponding to the angle of each channel, and multiply the channel gain function by the gain weight corresponding to the angle of each channel to obtain an adjusted channel gain function.
[0072] In some embodiments, since the channel distribution of the panoramic playback of the spatial sound field forms a spherical state, the gain weight voronoi(θ, φ) corresponding to the angle of each channel is determined according to the density of the distribution of the sphere to ensure that the energy of the sound playback is evenly distributed on the sphere. Then, the channel gain function G(θ, φ) is multiplied by the gain weight voronoi(θ, φ) corresponding to the angle of each channel to obtain the adjusted channel gain function G v (θ,φ), that is, G v (θ,φ)=G(θ,φ)*coronoi(θ,φ).
[0073] Then, the adjusted channel gain function G v (θ, φ) replaces the vocal tract gain function G(θ, φ) to perform the subsequent process of step 202.
[0074] Through the above solution, the gain weight of the channel gain function is adjusted, and the adjusted channel gain function can ensure that the energy of the sound playback is evenly distributed on the sphere, thereby ensuring the panoramic effect of the sound playback.
[0075] In some embodiments, the process of determining the decoding function includes:
[0076] Step b1, construct spherical harmonic basis functions.
[0077] In some embodiments, the spherical harmonic basis functions are the basis for converting a spatial sound field signal into a spherically distributed sound field composed of multiple channels.
[0078] In some embodiments, step b1 includes:
[0079] Step b11: determining the order corresponding to the spatial sound field signal.
[0080] In some embodiments, after the Ambisonib order l is determined, a (l, m) combination can be given based on the order l, where m=-l, -l+1, ..., 0, 1, ..., l.
[0081] Step b12: determine the term function of Legendre polynomials, and use the term function of Legendre polynomials to perform calculations according to the order to obtain spherical harmonic basis functions.
[0082] In some embodiments, the corresponding spherical harmonic basis functions are formulated as follows:
[0083] in, is the term function of the Legendre polynomial.
[0084] Step b2: determining the spherical harmonic basis value array corresponding to the angle of each channel based on the spherical harmonic basis function, wherein one spatial sound field signal corresponds to multiple channels, and a spherical harmonic basis value array is obtained corresponding to the angle of each channel.
[0085] In some embodiments, a spherical harmonic basis function is determined based on the (l, m) combination (as shown in FIG2B , for the (l=3, m=3) combination). Thus, for each channel angle (θ, φ), a corresponding spherical harmonic basis value array SH(θ, φ, l) is obtained:
[0086] Step b3: Arrange the spherical harmonic basis value arrays in order of angle to obtain the decoding function.
[0087] In some embodiments, for each angle: Where n is the number of channels. The spherical harmonic value arrays SH(θ, φ, l) corresponding to each angle are combined in the order of the angles to obtain the decoding function D(θ, φ, l), that is: D(θ, φ, l) = [SH(θ1, φ1, l), SH(θ2, φ2, l), …, SH(θ n ,φ n ,l)]
[0088] Through the above scheme, it can be ensured that the determined decoding function has the characteristics of spherical harmonic distribution, and that the subsequent decoding using the decoding function also has more spherical harmonic distribution characteristics, thereby improving the spherical distribution effect of the sound field and making the spatial sound field playback stereo effect better.
[0089] In some embodiments, if the directional gain function (eg, G(θ,φ) or M pp ) will change over time, and the directional gain function corresponding to each moment can be determined for application. In view of this situation, in step 202, for each two adjacent moments: the first moment and the second moment, the execution process includes:
[0090] Step I'' determines a first directional gain function at a first moment and a second directional gain function at a second moment, wherein the first moment and the second moment are adjacent moments.
[0091] In step II, gain processing is performed on the spatial sound field signal according to the first directional gain function to obtain a first target sound field signal corresponding to the first moment.
[0092] In step III, gain processing is performed on the spatial sound field signal according to the second directional gain function to obtain a second target sound field signal corresponding to the second moment.
[0093] In step IV, the first target sound field signal and the second target sound field signal are used as the target sound field signal.
[0094] The above steps can be divided into two cases:
[0095] One case: the directional gain function is the channel gain function G(θ, φ), which is converted into the channel gain function G(θ, φ, t) that changes with time. In this way, the first directional gain function corresponding to the first moment t1 is the channel gain function G(θ, φ, t1), and the second directional gain function corresponding to the second moment t2 is the channel gain function G(θ, φ, t2). Then, for the spatial sound field signal at the first moment t1, Decoding using decoding function Then perform gain processing Finally, the encoding process is performed to obtain the first target sound field signal Similarly, for the spatial sound field signal at the second moment t2 Decoding using decoding function Then perform gain processing Finally, the second target sound field signal is obtained by encoding and processing.
[0096] Another case: the directional gain function is M pp , so the first direction gain function corresponding to the first moment t1 can be calculated as M pp (t1) = D(θ, φ, l) * G(θ, φ, t1) * E(θ, φ, l); the second directional gain function corresponding to the second moment t2 is M pp (t2) = D(θ, φ, l)*G(θ, φ, t2)*E(θ, φ, l). Then, gain processing is performed to obtain the first target sound field signal and the second target sound field signal
[0097] The first target sound field signal obtained above and the second target sound field signal It can be directly used as the target sound field signal and played at the corresponding time, that is, the first target sound field signal and the second target sound field signal It can indicate the target sound field signal at different times.
[0098] Through the above solution, any two adjacent moments can be implemented according to the above solution, which can adapt to the real-time changing directional gain function and ensure the accuracy of the obtained target sound field signal.
[0099] However, directly playing the first target sound field signal and the second target sound field signal may cause the two sounds to be connected unsmoothly, resulting in auditory discomfort. To avoid this situation, the following embodiment is adopted:
[0100] In some embodiments, step IV" comprises:
[0101] In a time period corresponding to the first moment to the second moment, the first target sound field signal is smoothly transitioned to the second target sound field signal by using a cross-fade algorithm.
[0102] In some embodiments, during the time period from the first moment to the second moment, the first target sound field signal may be processed using a gradual-out curve, such that the first target sound field signal does not disappear abruptly but gradually disappears following the gradual-out amplitude of the gradual-out curve. During the time period from the first moment to the second moment, the second target sound field signal may be processed using a gradual-in curve, such that the first target sound field signal does not appear abruptly but gradually appears following the gradual-in amplitude of the gradual-in curve.
[0103] Through the above solution, the jump-like syllable change directly from the first target sound field signal to the second target sound field signal is avoided, and the smooth transition from the first target sound field signal to the second target sound field signal is ensured, making the syllable change heard by the listener softer and smoother.
[0104] In some embodiments, directional gain functions corresponding to a plurality of time periods are pre-calculated and stored, wherein one time period corresponds to one directional gain function; and the process of determining the directional gain function includes:
[0105] A target time period corresponding to the spatial sound field signal is determined, and a directional gain function corresponding to the target time period is selected from a plurality of directional gain functions.
[0106] In some embodiments, if the directional gain function for each time period can be pre-calculated and stored, then when receiving the spatial sound field signal in real time, it is only necessary to call the directional gain function of the corresponding target time period, and multiply the spatial sound field signal with the corresponding directional gain function to obtain the target sound field signal for the sound channel to play.
[0107] Through the above solution, the amount of calculation required for real-time calculation of the directional gain function can be reduced, ensuring that accurate target sound field signals can be obtained for playback in subsequent time periods.
[0108] It should be noted that the method of the embodiment of the present application can be performed by a single device, such as a computer or server. The method of this embodiment can also be applied in a distributed scenario and performed by multiple devices working together. In such a distributed scenario, one of the multiple devices may only perform one or more steps of the method of the embodiment of the present application, and the multiple devices will interact with each other to complete the method.
[0109] It should be noted that the above description is limited to some embodiments of the present application. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in an order different from that described in the above embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0110] Based on the same concept, corresponding to the directional gain method of the spatial sound field in any of the above embodiments, the present application also provides a directional gain device for the spatial sound field.
[0111] Referring to FIG3 , the apparatus comprises:
[0112] An acquisition module 301 is configured to acquire a spatial sound field signal;
[0113] The gain processing module 302 is configured to perform gain processing on the spatial sound field signal according to a determined directional gain function to obtain a target sound field signal.
[0114] In some embodiments, the gain processing module 302 is specifically configured to:
[0115] The method comprises the steps of: decoding the spatial sound field signal using a decoding function to obtain a multi-channel signal; determining a channel gain function, and performing gain processing on the multi-channel signal using the channel gain function as a directional gain function to obtain a gained multi-channel signal; and encoding the gained multi-channel signal using an encoding function to obtain the target sound field signal.
[0116] In some embodiments, the apparatus includes: a gain calculation storage module configured to:
[0117] Before performing gain processing on the spatial sound field signal according to the determined directional gain function, determining a decoding function, a channel gain function, and a coding function, multiplying the decoding function, the channel gain function, and the coding function to obtain and store the directional gain function;
[0118] The gain processing module 302 is specifically configured to:
[0119] The stored directional gain function is retrieved, and the spatial sound field signal is multiplied by the directional gain function to obtain the target sound field signal.
[0120] In some embodiments, the apparatus further includes a vocal channel gain function determination module configured to:
[0121] Determine the gain corresponding to the angle of each channel; and perform array combination on the channel gain corresponding to the angle of each channel to obtain the channel gain function.
[0122] In some embodiments, the vocal channel gain function determination module is further configured to:
[0123] After obtaining the channel gain function, the gain weight corresponding to the angle of each channel is determined, and the channel gain function is multiplied by the gain weight corresponding to the angle of each channel to obtain an adjusted channel gain function.
[0124] In some embodiments, the apparatus further includes a decoding function determination module configured to:
[0125] Construct spherical harmonic basis functions; based on the spherical harmonic basis functions, determine the spherical harmonic basis value arrays corresponding to the angles of each channel, wherein a spatial sound field signal corresponds to multiple channels, and the angle of each channel corresponds to a spherical harmonic basis value array; arrange the spherical harmonic basis value arrays in order of angles to obtain the decoding function.
[0126] In some embodiments, the decoding function determination module is further configured to:
[0127] Determine the order corresponding to the spatial sound field signal; determine the term function of the Legendre polynomial, and use the term function of the Legendre polynomial to perform calculation processing according to the order to obtain the spherical harmonic basis function.
[0128] In some embodiments, the gain processing module 302 is further configured to:
[0129] Determine a first directional gain function at a first moment and a second directional gain function at a second moment, wherein the first moment and the second moment are adjacent moments; perform gain processing on the spatial sound field signal according to the first directional gain function to obtain a first target sound field signal corresponding to the first moment; perform gain processing on the spatial sound field signal according to the second directional gain function to obtain a second target sound field signal corresponding to the second moment; and use the first target sound field signal and the second target sound field signal as the target sound field signal.
[0130] In some embodiments, the gain processing module 302 is further configured to:
[0131] In a time period corresponding to the first moment to the second moment, the first target sound field signal is smoothly transitioned to the second target sound field signal by using a cross-fade algorithm.
[0132] In some embodiments, the gain calculation and storage module is further configured to: pre-calculate and store directional gain functions corresponding to a plurality of time periods, wherein one time period corresponds to one directional gain function;
[0133] The process of determining the directional gain function in the gain processing module 302 includes:
[0134] A target time period corresponding to the spatial sound field signal is determined, and a directional gain function corresponding to the target time period is selected from a plurality of directional gain functions.
[0135] For the convenience of description, the above devices are described as being divided into various modules according to their functions. Of course, when implementing this application, the functions of each module can be implemented in the same or multiple software and / or hardware.
[0136] The apparatus of the above embodiment is used to implement the corresponding method in any of the above embodiments, and has the beneficial effects of the corresponding method embodiment, which will not be described in detail here.
[0137] Based on the same concept, corresponding to the method of any of the above embodiments, the present application also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and runnable on the processor, wherein the processor implements the method described in any of the above embodiments when executing the program.
[0138] FIG4 shows a more specific schematic diagram of the hardware structure of an electronic device provided in this embodiment. The device may include: a processor 410, a memory 420, an input / output interface 430, a communication interface 440, and a bus 450. The processor 410, the memory 420, the input / output interface 430, and the communication interface 440 are connected to each other within the device via the bus 450.
[0139] The processor 410 can be implemented using a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this specification.
[0140] The memory 420 can be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage devices, dynamic storage devices, etc. The memory 420 can store an operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 420 and is called and executed by the processor 410.
[0141] The input / output interface 430 is used to connect input / output modules to implement information input and output. The input / output modules can be configured as components within the device (not shown in the figure) or can be externally connected to the device to provide corresponding functions. Input devices may include a keyboard, mouse, touch screen, microphone, various sensors, etc., and output devices may include a display, speaker, vibrator, indicator light, etc.
[0142] The communication interface 440 is used to connect to a communication module (not shown) to enable communication between the device and other devices. The communication module can communicate via a wired method (such as USB, network cable, etc.) or a wireless method (such as mobile network, WiFi, Bluetooth, etc.).
[0143] The bus 450 comprises a pathway for transmitting information between the various components of the device, such as the processor 410 , the memory 420 , the input / output interface 430 , and the communication interface 440 .
[0144] It should be noted that although the above device only shows the processor 410, the memory 420, the input / output interface 430, the communication interface 440, and the bus 450, in a specific implementation, the device may also include other components necessary for normal operation. In addition, it will be understood by those skilled in the art that the above device may only include the components necessary to implement the embodiments of this specification, and does not necessarily include all the components shown in the figure.
[0145] The electronic device of the above embodiment is used to implement the corresponding spatial sound field directional gain method in any of the above embodiments, and has the beneficial effects of the corresponding spatial sound field directional gain method embodiment, which will not be described in detail here.
[0146] Based on the same concept, corresponding to any of the above-mentioned embodiments, the present application also provides a non-transitory computer-readable storage medium, which stores computer instructions, and the computer instructions are used to enable the computer to execute the method described in any of the above embodiments.
[0147] The computer-readable media of this embodiment include permanent and non-permanent, removable and non-removable media that can be used to store information by any method or technology. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, read-only compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device.
[0148] The computer instructions stored in the storage medium of the above embodiment are used to enable the computer to execute the method described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.
[0149] Based on the same concept, corresponding to any of the above-mentioned embodiment methods, the present application also provides a computer program product, including computer program instructions. When the computer program instructions are run on a computer, the computer executes the method described in any of the above embodiments, which has the beneficial effects of the corresponding method embodiments and will not be repeated here.
[0150] Those skilled in the art should understand that the discussion of any of the above embodiments is merely illustrative and is not intended to imply that the scope of the present application (including the claims) is limited to these examples. Within the scope of the present application, the technical features in the above embodiments or different embodiments may be combined, the steps may be implemented in any order, and there are many other variations of the different aspects of the embodiments of the present application as described above, which are not provided in detail for the sake of simplicity.
[0151] In addition, for simplicity of description and discussion, and in order not to make the embodiment of the application difficult to understand, the known power supply / ground connection with integrated circuit (IC) chip and other components may or may not be shown in the accompanying drawings provided. In addition, the device can be shown in the form of a block diagram to avoid making the embodiment of the application difficult to understand, and this also takes into account the following fact, that is, the details of the embodiment of these block diagram devices are highly dependent on the platform to be implemented in the embodiment of the application (that is, these details should be fully within the scope of understanding of those skilled in the art). When specific details (for example, circuit) are set forth to describe exemplary embodiments of the application, it will be apparent to those skilled in the art that the embodiment of the application can be implemented without these specific details or when these specific details are changed. Therefore, these descriptions should be considered to be illustrative rather than restrictive.
[0152] Although the present invention has been described in conjunction with specific embodiments thereof, many alternatives, modifications, and variations of these embodiments will be apparent to those skilled in the art based on the foregoing description. For example, other memory architectures (e.g., dynamic RAM (DRAM)) may utilize the embodiments discussed.
[0153] [Corrected 08.01.2025 in accordance with Rule 26] The embodiments of this application are intended to cover all such substitutions, modifications, and variations that fall within the broad scope of the appended claims. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the embodiments of this application shall be included within the scope of protection of this application.
Claims
1. A method for directional gain of a spatial sound field, comprising: Acquiring spatial sound field signals; Gain processing is performed on the spatial sound field signal based on a directional gain function to obtain a target sound field signal.
2. The method according to claim 1, wherein The performing gain processing on the spatial sound field signal based on the directional gain function to obtain a target sound field signal includes: Decoding the spatial sound field signal using a decoding function to obtain a multi-channel signal; Determine the vocal tract gain function, Performing gain processing on the multi-channel signal using the channel gain function as a directional gain function to obtain a gained multi-channel signal; The multi-channel signal after gaining is coded using a coding function to obtain the target sound field signal.
3. The method according to claim 1, wherein Before performing gain processing on the spatial sound field signal based on the directional gain function, the method further includes: Determine the decoding function, channel gain function and encoding function, Multiplying the decoding function, the channel gain function and the encoding function to obtain the directional gain function; The step of performing gain processing on the spatial sound field signal according to the determined directional gain function to obtain a target sound field signal includes: The spatial sound field signal is multiplied by the directional gain function to obtain the target sound field signal.
4. The method according to claim 2 or 3, wherein: The process of determining the vocal tract gain function includes: Determine the amount of gain corresponding to the angle of each channel; The channel gain amounts corresponding to the angles of the various channels are combined in an array to obtain the channel gain function.
5. The method according to claim 4, wherein After obtaining the vocal channel gain function, the method further includes: Determine the gain weight corresponding to the angle of each channel, The channel gain function is multiplied by the gain weight corresponding to the angle of each channel to obtain an adjusted channel gain function.
6. The method according to claim 2 or 3, wherein: The process of determining the decoding function includes: Construct spherical harmonic basis functions; Based on the spherical harmonic basis functions, determining a spherical harmonic basis value array corresponding to the angle of each channel, wherein a spatial sound field signal corresponds to multiple channels, and a spherical harmonic basis value array is obtained corresponding to the angle of each channel; The decoding function is obtained by arranging the spherical harmonic basis value arrays in the order of angles.
7. The method according to claim 6, wherein: The constructing of spherical harmonic basis functions comprises: Determine the order corresponding to the spatial sound field signal; Determine the term functions of the Legendre polynomials, The term functions of the Legendre polynomials are used to perform calculations according to the order to obtain spherical harmonic basis functions.
8. The method according to claim 1, wherein The performing gain processing on the spatial sound field signal based on the directional gain function to obtain a target sound field signal includes: Determine a first directional gain function at a first moment and a second directional gain function at a second moment, wherein the first moment and the second moment are adjacent moments; performing gain processing on the spatial sound field signal according to the first directional gain function to obtain a first target sound field signal corresponding to the first moment; performing gain processing on the spatial sound field signal according to the second directional gain function to obtain a second target sound field signal corresponding to the second moment; At least one of the first target sound field signal and the second target sound field signal is used as the target sound field signal.
9. The method according to claim 8, wherein The taking at least one of the first target sound field signal and the second target sound field signal as the target sound field signal includes: In a time period corresponding to the first moment to the second moment, the first target sound field signal is smoothly transitioned to the second target sound field signal by using a cross-fade algorithm.
10. The method according to claim 1, wherein Precalculating and storing directional gain functions corresponding to a plurality of time periods, wherein one time period corresponds to one directional gain function; The process of determining the directional gain function includes: A target time period corresponding to the spatial sound field signal is determined, and a directional gain function corresponding to the target time period is selected from a plurality of directional gain functions.
11. A directional gain device for a spatial sound field, wherein: include: An acquisition module, configured to acquire a spatial sound field signal; The gain processing module is configured to perform gain processing on the spatial sound field signal based on a directional gain function to obtain a target sound field signal.
12. An electronic device comprising a memory and a processor, wherein a computer program is stored in the memory, wherein: The processor executes the program to implement the method according to any one of claims 1 to 10.
13. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are executed by a computer to implement the method according to any one of claims 1 to 10.
14. A computer program product comprising computer program instructions, wherein: When the computer program instructions are executed on a computer, the computer is caused to implement the method according to any one of claims 1 to 10.
15. A computer program, comprising program codes, which, when executed by a processor, enable the processor to implement the method according to any one of claims 1 to 10.
Citation Information
Patent Citations
A signal processing apparatus for enhancing a voice component within a multi-channel audio signal
CN107004427A
Speech enhancement method and device, computer equipment and storage medium
CN114550743A
Virtual stereo generation method and electronic equipment
CN116320908A
Audio signal processing method and system, audio signal playing method and system and electronic equipment
CN117354707A
Direction gain method, device and equipment of space sound field, medium and product
CN117880734A