Audio processing method, terminal, storage medium and program product
By applying high-order background sound HOA coefficients and spherical harmonic beamforming technology on portable devices, the problem of limited number of microphones is solved, the encoding of high-order audio signals and sound field reconstruction are realized, the signal fidelity of audio signals and the accuracy of sound field reconstruction are improved, and the immersive listening experience is enhanced.
Patent Information
- Application Number
- CN202580000768.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-06
- Publication Date
- 2025-09-19
AI Technical Summary
In the existing technology, when the HOA coding method is applied on portable devices, it is limited by the number of microphones and array regularity, and cannot effectively realize the coding of high-order audio signals and sound field reconstruction, resulting in insufficient signal fidelity of audio signals and insufficient accuracy of sound field reconstruction.
Through terminal-based microphone array configuration, the audio signal is up-coded using the high-order background sound HOA coefficients, and combined with spherical harmonic beamforming and time domain alignment strategies to generate a second audio signal, breaking through the limitation on the number of microphones and improving the signal fidelity of the audio signal and the accuracy of sound field reconstruction.
It achieves the processing of higher-order audio signals without increasing the number of microphone devices, improves the flexibility of audio processing and the immersive listening experience, and enhances the signal fidelity of audio signals and the accuracy of sound field reconstruction.
Smart Images

Figure CN120677718A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of audio signal processing, and in particular to an audio processing method, terminal, storage medium, and program product. Background Art
[0002] Higher-order ambisonics (HOA) technology is a mainstream spatial audio technology. Its advantages lie in its device independence and ability to compensate for head motion. Among related technologies, HOA encoding relies on a regular spherical microphone array (SMA). However, further research is needed to apply HOA encoding methods to portable devices such as terminals. Summary of the Invention
[0003] In order to improve the output quality of audio signals, embodiments of the present disclosure provide an audio processing method, a terminal, a storage medium, and a program product.
[0004] According to a first aspect of an embodiment of the present disclosure, an audio processing method is proposed, which is executed by a terminal. The method includes:
[0005] Acquire a first audio signal collected by a microphone array of the terminal, where the microphone array includes at least one microphone device;
[0006] The first audio signal is up-order encoded according to the first high-order background sound (HOA) coefficient to generate a second audio signal.
[0007] According to a second aspect of an embodiment of the present disclosure, an audio processing device is provided, which is applied to a terminal. The device includes:
[0008] a transceiver module, configured to obtain a first audio signal collected by a microphone array of the terminal, the microphone array comprising at least one microphone device;
[0009] The processing module is configured to perform up-order coding on the first audio signal according to the first high-order background sound effect (HOA) coefficient to generate a second audio signal.
[0010] According to a third aspect of an embodiment of the present disclosure, a terminal is provided. The terminal is configured to execute the audio processing method described in any one of the first aspects of the present disclosure.
[0011] According to a fourth aspect of an embodiment of the present disclosure, a storage medium is proposed, which stores instructions. When the instructions are executed on a terminal, the terminal executes the audio processing method as described in any one of the first aspects of the present disclosure.
[0012] According to the fifth aspect of the embodiment of the present disclosure, a program product is proposed, including at least one of a program and an instruction, which, when executed by a terminal, implements the steps of the audio processing method described in any one of the first aspects of the present disclosure.
[0013] By adopting the above technical solution, at least the following beneficial technical effects can be achieved:
[0014] The terminal receives a first audio signal collected by a microphone array comprising at least one microphone device and performs up-order encoding on the first audio signal based on a first high-order background sound (HOA) coefficient to generate a second audio signal. This enables HOA encoding of the audio signal based on the terminal's microphone array configuration, improving the signal fidelity of the output audio and significantly enhancing the accuracy of sound field reconstruction. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure, the following drawings required for describing the embodiments are introduced. The following drawings are merely some embodiments of the present disclosure and do not impose specific limitations on the protection scope of the present disclosure.
[0016] Figure 1 It is a schematic diagram of a microphone array configuration of a terminal according to an embodiment of the present disclosure.
[0017] Figure 2A It is a flowchart of an audio processing method according to an embodiment of the present disclosure.
[0018] Figure 2B It is a flowchart of a method for determining an HOA coefficient according to an embodiment of the present disclosure.
[0019] Figure 2C It is a schematic diagram showing the audio simulation experiment results according to an embodiment of the present disclosure.
[0020] Figure 2D It is a schematic diagram showing the actual audio experiment results according to an embodiment of the present disclosure.
[0021] Figure 3A It is a flowchart of an audio processing method according to an embodiment of the present disclosure.
[0022] Figure 3B It is a flowchart of an audio processing method according to an embodiment of the present disclosure.
[0023] Figure 4 It is a flowchart of an audio processing method according to an embodiment of the present disclosure.
[0024] Figure 5 It is a schematic structural diagram of a terminal according to an embodiment of the present disclosure.
[0025] Figure 6 FIG6 is a structural diagram of an electronic device 600 according to an embodiment of the present disclosure.
[0026] Figure 7 7 is a schematic structural diagram of a chip 700 proposed according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0027] The embodiments of the present disclosure provide an audio processing method, a terminal, a storage medium, and a program product.
[0028] In a first aspect, an embodiment of the present disclosure provides an audio processing method, which is executed by a terminal. The method includes:
[0029] Acquire a first audio signal collected by a microphone array of the terminal, where the microphone array includes at least one microphone device;
[0030] The first audio signal is up-order encoded according to the first high-order background sound (HOA) coefficient to generate a second audio signal.
[0031] In the above embodiment, based on the microphone array configuration of the terminal, HOA encoding of the audio signal is implemented, the signal fidelity of the output audio is improved, and the accuracy of sound field reconstruction is significantly improved.
[0032] In combination with some embodiments of the first aspect, in some embodiments, the number of the at least one microphone device and the audio order of the second audio signal do not need to meet the first restriction condition.
[0033] In the above embodiment, by limiting the conditions for up-order encoding of the audio signal in this embodiment, it is possible to process higher-order audio signals without increasing the number of microphone devices, thereby improving the flexibility of audio processing.
[0034] In conjunction with some embodiments of the first aspect, in some embodiments, the method further includes:
[0035] Determining, according to the geometric coordinates of the microphone array, a frequency response of the microphone array to a plane wave sound source in K discrete directions, where K is a positive integer;
[0036] generating an array manifold matrix of the microphone array according to the frequency response;
[0037] Determine a plurality of spherical harmonic function values corresponding to the K discrete directions respectively;
[0038] generating a discrete spherical harmonic transform (DSHT) matrix according to the plurality of spherical harmonic function values, the array manifold matrix, and a first audio order, wherein the first audio order is a set order of the terminal;
[0039] The first HOA coefficient is determined according to the array manifold matrix and the DSHT matrix.
[0040] In the above embodiment, by converting the microphone sampling point constraint into the sound source direction constraint, the number of microphones is exceeded, and the HOA encoding of the first audio signal by the terminal is achieved.
[0041] In conjunction with some embodiments of the first aspect, in some embodiments, determining the first HOA coefficient according to the array manifold matrix and the DSHT matrix includes:
[0042] Determining a constraint condition of the first HOA coefficient according to the array manifold matrix and the DSHT matrix;
[0043] The first HOA coefficient is determined according to the constraint condition.
[0044] In the above embodiment, the constraint conditions can minimize the error between the reconstructed sound field and the original sound field, improve the accuracy of the sound field reconstruction, and ensure that the omnidirectional component of the sound field complies with physical rules.
[0045] In conjunction with some embodiments of the first aspect, in some embodiments, determining the multiple spherical harmonic function values corresponding to the K discrete directions includes:
[0046] The spherical harmonics values are determined by the following formula:
[0047]
[0048] Among them, the is the spherical harmonic function value in any discrete direction, θ is the polar angle, φ is the azimuth angle, n is a natural number less than or equal to the first audio order, m is an integer greater than or equal to -n and less than or equal to n, is the associated Legendre polynomial.
[0049] In the above embodiment, by decomposing the sound field into coefficients of spherical harmonic functions, encoding, reconstruction, sound source localization and beamforming of the sound field are achieved, thereby improving the immersive auditory experience.
[0050] In conjunction with some embodiments of the first aspect, in some embodiments, generating a discrete spherical harmonic transform (DSHT) matrix according to the multiple spherical harmonic function values and the array manifold matrix includes:
[0051] The DSHT matrix is determined by the following formula:
[0052]
[0053] Among them, the is the spherical harmonic function value in any discrete direction, and D(ω) is the array manifold matrix.
[0054] In the above embodiment, the discrete signals on the sphere can be converted into spherical harmonic coefficients through the DSHT matrix, thereby realizing the encoding, reconstruction, analysis and processing of the sound field, providing data support for creating an immersive audio experience.
[0055] With reference to some embodiments of the first aspect, in some embodiments, the constraint condition includes:
[0056]
[0057] Among them, the is the first HOA coefficient, D(ω) is the array manifold matrix, S is the DSHT matrix, and e n,m is a unit vector.
[0058] In the above embodiment, by optimizing the HOA coefficients according to the constraint conditions, the error between the reconstructed sound field and the original sound field can be minimized, and the accuracy of the sound field reconstruction can be improved to meet specific sound field reconstruction requirements.
[0059] In conjunction with some embodiments of the first aspect, in some embodiments, determining the first HOA coefficient by regularized least squares method according to the constraint condition includes:
[0060] The first HOA coefficient is determined by the following formula:
[0061]
[0062] Among them, the is the first HOA coefficient, D(ω) is the array manifold matrix, S is the DSHT matrix, and e n,m is a unit vector, and λ is a regularization coefficient.
[0063] In the above embodiment, by determining the HOA coefficients from discrete sampling points on the sphere, accurate description and processing of the three-dimensional sound field environment is achieved, thereby improving the accuracy of the output audio.
[0064] In conjunction with some embodiments of the first aspect, in some embodiments, the method further includes:
[0065] If it is determined that the signal simulation degree of the second audio signal is lower than a set threshold, obtaining a second audio order, where the second audio order is greater than the first audio order, and the first audio order is an initially set order of the terminal;
[0066] determining a second HOA coefficient according to the second audio order and the first HOA coefficient;
[0067] The first audio signal is encoded according to the second HOA coefficient to generate a third audio signal.
[0068] In the above embodiment, there is no limit on the number of upscaling steps of the audio signal. The number of upscaling steps of the audio signal is determined by measuring the signal fidelity of the audio signal generated after upscaling. This prevents incorrect upscaling that would result in a decrease in the signal fidelity of the audio signal, thereby enhancing the authenticity of the output audio.
[0069] In combination with some embodiments of the first aspect, in some embodiments, encoding the first audio signal according to the first high-order background sound effect (HOA) coefficient to generate the second audio signal includes:
[0070] determining, based on the first HOA coefficient, a third HOA coefficient of a fourth audio signal, where the fourth audio signal is an audio signal having an audio frequency in the first audio signal higher than a set frequency threshold;
[0071] encoding the fourth audio signal according to the third HOA coefficient to generate a fifth audio signal;
[0072] encoding the sixth audio signal according to the first HOA coefficient to generate a seventh audio signal, wherein the sixth audio signal is an audio signal having an audio frequency lower than or equal to the set frequency threshold in the first audio signal;
[0073] The second audio signal is generated according to the fifth audio signal and the seventh audio signal.
[0074] In the above embodiment, by separating phase and amplitude processing, the directionality information is retained in the aliasing frequency band, thereby improving the high-frequency robustness of the audio processing.
[0075] In a second aspect, an embodiment of the present disclosure provides an audio processing device, applied to a terminal, the device comprising:
[0076] a transceiver module, configured to obtain a first audio signal collected by a microphone array of the terminal, the microphone array comprising at least one microphone device;
[0077] The processing module is configured to perform up-order coding on the first audio signal according to the first high-order background sound effect (HOA) coefficient to generate a second audio signal.
[0078] In a third aspect, an embodiment of the present disclosure provides a terminal, which is used to execute the audio processing method described in any one of the first aspects of the present disclosure.
[0079] In a fourth aspect, an embodiment of the present disclosure proposes a storage medium, which stores instructions. When the instructions are executed on a terminal, the terminal executes the audio processing method as described in any one of the first aspects of the present disclosure.
[0080] In a fifth aspect, an embodiment of the present disclosure proposes a program product, comprising at least one of a program and an instruction, wherein when the at least one of the program and the instruction is executed by a terminal, the steps of the audio processing method described in any one of the first aspects of the present disclosure are implemented.
[0081] In a sixth aspect, an embodiment of the present disclosure provides a chip or a chip system, which includes a processing circuit configured to execute the method described in the optional implementation of the first aspect.
[0082] It is understood that the above-mentioned audio processing device, terminal, storage medium, program product, etc. are all used to execute the method proposed in the embodiment of the present disclosure. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the corresponding method and will not be repeated here.
[0083] The embodiments of the present disclosure are not exhaustive, but are merely illustrations of some embodiments, and are not intended to be a specific limitation on the scope of protection of the present disclosure. In the absence of contradiction, each step in a certain embodiment can be implemented as an independent embodiment, and the steps can be arbitrarily combined. For example, a solution after removing some steps in a certain embodiment can also be implemented as an independent embodiment, and the order of the steps in a certain embodiment can be arbitrarily exchanged. In addition, the optional implementation methods in a certain embodiment can be arbitrarily combined; in addition, the embodiments can be arbitrarily combined. For example, some or all steps of different embodiments can be arbitrarily combined, and a certain embodiment can be arbitrarily combined with the optional implementation methods of other embodiments. In each embodiment of the present disclosure, if there is no special explanation and logical conflict, the terms and / or descriptions between the embodiments are consistent and can be referenced to each other. The technical features in different embodiments can be combined to form new embodiments according to their internal logical relationships.
[0084] The terms used in the embodiments of the present disclosure are only for the purpose of describing specific embodiments and are not intended to limit the present disclosure.
[0085] In the embodiments of the present disclosure, unless otherwise specified, elements expressed in the singular, such as "a", "an", "the", "above", "said", "the", "the", etc., may mean "one and only one", or "one or more", "at least one", etc. For example, when using articles such as "a", "an", "the" in English in translation, the noun following the article may be understood as a singular expression or a plural expression.
[0086] In the embodiments of the present disclosure, “plurality” refers to two or more.
[0087] In some embodiments, the terms "at least one of A or B, at least one of A and B", "one or more", "a plurality of", "multiple" and the like can be used interchangeably.
[0088] In some embodiments, descriptions such as "at least one of A and B," "A and / or B," "A in one case, B in another case," and "in response to one case A, in response to another case B" may include the following technical solutions depending on the circumstances: in some embodiments, A (A is executed regardless of whether there is a branch B); in some embodiments, B (B is executed regardless of whether there is a branch A); in some embodiments, execution is selected from A and B (A and B are selectively executed); in some embodiments, A and B (both A and B are executed). The above is also applicable when there are more branches such as A, B, and C.
[0089] In some embodiments, "A or B" and other notations may include the following technical solutions, depending on the circumstances: in some embodiments, A (A is executed regardless of whether B branch exists); in some embodiments, B (B is executed regardless of whether A branch exists); in some embodiments, execution is selected from A and B (A and B are selectively executed). The above is also applicable when there are more branches such as A, B, and C.
[0090] The prefixes such as "first" and "second" in the embodiments of the present disclosure are only used to distinguish different description objects and do not constitute any restriction on the position, order, priority, quantity or content of the description objects. For the statement of the description object, please refer to the description in the context of the claims or embodiments, and no unnecessary restriction should be constituted due to the use of prefixes. For example, if the description object is a "field", the ordinal number before the "field" in the "first field" and the "second field" does not limit the position or order between the "fields". "First" and "second" do not limit whether the "fields" they modify are in the same message, nor do they limit the order of the "first field" and the "second field". For another example, if the description object is a "level", the ordinal number before the "level" in the "first level" and the "second level" does not limit the priority between the "levels". For another example, the number of description objects is not limited by the ordinal number and can be one or more. Taking "first device" as an example, the number of "devices" can be one or more. In addition, the objects modified by different prefixes can be the same or different. For example, if the description object is "device", then the "first device" and the "second device" can be the same device or different devices, and their types can be the same or different; for another example, if the description object is "information", then the "first information" and the "second information" can be the same information or different information, and their contents can be the same or different.
[0091] In some embodiments, “including A,” “comprising A,” “used to indicate A,” and “carrying A” can be interpreted as directly carrying A or indirectly indicating A.
[0092] In some embodiments, terms such as "time / frequency" and "time / frequency domain" refer to the time domain and / or the frequency domain.
[0093] In some embodiments, terms such as "in response to...", "in response to determining...", "in the case of...", "at...", "when...", "if...", "if...", etc. can be used interchangeably. These descriptions all mean that the device will make corresponding processing under certain objective circumstances. It is not necessary to limit the time, nor is it required that the device must perform a judgment action when implemented, nor does it mean that there must be other limitations.
[0094] In some embodiments, terms such as "greater than", "greater than or equal to", "not less than", "more than", "more than or equal to", "not less than", "higher than", "higher than or equal to", "not less than", and "above" can be replaced with each other, and terms such as "less than", "less than or equal to", "not greater than", "less than", "less than or equal to", "not more than", "lower than", "lower than or equal to", "not higher than", and "below" can be replaced with each other.
[0095] In some embodiments, devices and the like can be interpreted as physical or virtual, and their names are not limited to those described in the embodiments. Terms such as "device," "equipment," "device," "circuit," "network element," "network function," "network device," "function," "node," "unit," "section," "system," "network," "chip," "chip system," "entity," and "subject" can be used interchangeably.
[0096] In some embodiments, the terms "terminal", "terminal device", "user equipment (UE)", "user terminal", "mobile station (MS)", "mobile terminal (MT)", subscriber station, mobile unit, subscriber unit, wireless unit, remote unit, mobile device, wireless device, wireless communication device, remote device, mobile subscriber station, access terminal, mobile terminal, wireless terminal, remote terminal, handset, user agent, mobile client, client, etc. can be used interchangeably.
[0097] In some embodiments, obtaining data, information, etc. may comply with the laws and regulations of the country where the data is obtained.
[0098] In some embodiments, data, information, etc. may be obtained with the user's consent.
[0099] In addition, each element, each row, or each column in the table of the embodiment of the present disclosure can be implemented as an independent embodiment, and the combination of any elements, any rows, and any columns can also be implemented as an independent embodiment.
[0100] In some embodiments, neural HOA coding is implemented through a data-driven approach, which requires a large number of real microphone signals to train the model and has limited generalization capabilities. Among them, the model-driven approach includes: sound field decomposition method and least squares coding method. Among them, the sound field decomposition method relies on direction estimation (DOA, Direction-of-Arrival) and sound source separation, and has high accuracy requirements; the least squares coding method directly constructs a linear equation between the HOA coefficients and the array manifold, but is limited by the number of microphones and cannot break through the order limit of HOA coding.
[0101] In some embodiments, the least squares-based HOA coding method and the beamforming-based HOA coding method need to satisfy Q≥(N+1) 2 , it cannot be applied to terminals that implement HOA coding with four microphones. Furthermore, for high-frequency audio signals, insufficient microphone spacing can easily lead to degraded HOA coding performance. Methods that rely on DOA estimation or sound field model assumptions are highly complex and have limited generalization capabilities.
[0102] For example, construct a linear system of equations through the array manifold matrix: AB = p, and solve the pseudo-inverse matrix This method uses the 2.5D spherical harmonic beam pattern as the coding matrix. However, due to the limited number of microphones, it only supports low-order HOA coding. Underdetermined equations lead to failure of high-order coding and deterioration of the condition number in high-frequency bands. HOA coefficients are directly output using a 2.5D spherical harmonic beam pattern. However, this relies on a regular array and does not address the irregular arrangement of terminal microphones.
[0103] In some embodiments, a method for HOA encoding and order raising based on spherical harmonic beamforming for a terminal microphone array is disclosed. This method uses spherical harmonic beamforming to implement fourth-order HOA encoding using four microphones. A time-domain alignment strategy is introduced to optimize the high-frequency phase response, and a beamformer is designed directly based on the array manifold, avoiding the complex preprocessing required for HOA encoding.
[0104] For example, this embodiment is applied to the HOA audio collection and processing scenario of a terminal in a complex sound field environment. Figure 1 FIG is a schematic diagram of a microphone array configuration of a terminal according to an embodiment of the present disclosure. Figure 1 As shown, the terminal is equipped with four irregular microphones: Mic A, Mic B, Mic C, and Mic D. Among them, Mic A is the main microphone located at the bottom of the terminal, Mic B is the secondary microphone located at the bottom of the terminal, Mic C is the microphone located at the top of the terminal, and Mic D is the microphone located at the back of the terminal. The maximum distance between each microphone in the terminal is d max=8cm. The terminal includes: an audio processing module for running the beamforming algorithm disclosed herein to generate HOA signals up to the fourth order; an output interface for directly using the processed HOA signals for local audio recording on the terminal, real-time call noise reduction, or transmitting them to an external headset via Bluetooth. This embodiment breaks through the limitation on the number of microphones. Through spherical harmonic beamforming, four microphones are used to implement fourth-order HOA coding, suppress high-frequency aliasing, introduce a time-domain alignment strategy, optimize high-frequency phase response, and eliminate the need for DOA estimation. The beamformer can be designed directly based on the array manifold, avoiding the complex preprocessing of the audio.
[0105] In some embodiments, the signal transmission process of the present disclosure includes:
[0106] Step 1: The microphone array collects the sound pressure signals in the sound field environment currently in which the terminal is located: p1(ω), p2(ω), p3(ω), and p4(ω). For example, p1(ω) is the sound pressure signal collected by Mic A, p2(ω) is the sound pressure signal collected by Mic B, p3(ω) is the sound pressure signal collected by Mic C, and p4(ω) is the sound pressure signal collected by Mic D.
[0107] Step 2: Generate HOA coefficients through beamforming algorithm Where n≤4, |m|≤n;
[0108] Step 3: The HOA signal is used for end-to-end storage (e.g., spatial audio storage in video recording) or real-time communication (e.g., directional noise reduction).
[0109] In some embodiments, this embodiment is applicable to a variety of application scenarios. For example, in a video recording scenario, when a user records a video through a terminal, the terminal collects the audio signal of the ambient sound field through a microphone array and encodes it into an HOA signal, which can then be restored through headphones in 360° spatial audio. In a voice call scenario, in a noisy environment, the direction of the target speaker is separated through HOA encoding to improve call clarity. In a multimedia playback scenario, when the terminal plays music, the HOA signal is combined with the headphone virtual surround sound algorithm to enhance the stereo experience.
[0110] Figure 2A FIG. 1 is a flow chart of an audio processing method according to an embodiment of the present disclosure. Figure 2A As shown, the embodiment of the present disclosure relates to an audio processing method, which is executed by a terminal. The method includes:
[0111] Step S2101: Acquire a first audio signal collected by a microphone array of a terminal.
[0112] For example, this embodiment is executed by a terminal, which is provided with a microphone array. The terminal collects a first audio signal of the sound field environment currently located by the terminal based on the microphone array. The terminal is used to collect ambient audio of the sound field environment. After obtaining the first audio signal, the terminal encodes the first audio signal to generate a second audio signal for output. The audio effect of the second audio signal is the same as or similar to the audio effect of the first audio signal, thereby achieving high-precision restoration of the first audio signal and improving the accuracy of sound field reconstruction.
[0113] In some embodiments, the first audio signal is an audio signal of a sound field environment collected by the terminal through a microphone array.
[0114] For example, in this embodiment, the microphone array includes Q microphone devices, where Q is a positive integer. The microphone devices are set at different positions of the terminal, and the setting positions of the microphone devices are different, so the audio signals collected by different microphone devices are different. That is, the first audio signal is used to represent the environmental audio signal collected by each microphone device. For example, the microphone array includes 4 microphone devices, then the first audio signal is p(ω) = [p1(ω), p2(ω), p3(ω), p4(ω)] T Among them, p1(ω) is the audio signal collected by Mic A, p2(ω) is the audio signal collected by Mic B, p3(ω) is the audio signal collected by Mic C, and p4(ω) is the audio signal collected by Mic D.
[0115] In some embodiments, the name of the first audio signal is not limited, and may be, for example, “microphone acquisition signal”, “ambient sound field signal”, “ambient audio signal”, “background audio signal”, etc.
[0116] In some embodiments, the terminal achieves high-precision spatial audio capture by restoring the collected audio signal. For example, the terminal's usage scenarios may include at least one of the following: virtual reality (VR) application scenarios, augmented reality (AR) application scenarios, smart home application scenarios, remote conferencing system application scenarios, and drone acoustic monitoring application scenarios.
[0117] In some embodiments, the output end of the second audio signal can be a terminal. For example, in a scenario where video recording is performed based on a terminal, the ambient sound field is collected through the microphone array of the terminal to obtain a first audio signal. The first audio signal is encoded to generate a second audio signal which is associated with the recorded video data and saved. When the video data is played later through the speaker device of the terminal, the speaker device is the output end of the second audio signal, and the second audio signal is played through the speaker device, thereby restoring 360° spatial audio.
[0118] Optionally, in some embodiments, the output end of the second audio signal can also be another electronic device. For example, in a voice call environment, the terminal's microphone array collects audio signals of the ambient sound field (including user voice and environmental noise) to obtain a first audio signal. After the terminal encodes the first audio signal to generate a second audio signal, it transmits the second audio signal to another electronic device via wireless network communication. The second audio signal is played based on the speaker device of the other electronic device. This achieves a high degree of restoration of the call voice and improves the authenticity of the call voice.
[0119] In some embodiments, the microphone array includes at least one microphone device.
[0120] For example, in this embodiment, the arrangement positions of the microphone devices in the terminal can be regular or irregular. This embodiment does not limit this. The terminal can set the number and arrangement positions of the microphone devices based on assembly requirements.
[0121] Optionally, in some embodiments, the number of the at least one microphone device and the audio order of the second audio signal do not need to satisfy the first restriction condition.
[0122] For example, in this embodiment, the microphone array includes Q microphone devices, and the first constraint condition is: Q ≥ (N+1) 2 Where N is the audio order of the second audio signal. That is, in this embodiment, when implementing high-order audio coding, there is no restriction on the number of microphone devices included in the microphone array. For example, by configuring four microphone devices, fourth-order audio coding of an audio signal can be implemented; by configuring two microphone devices, third-order audio coding of an audio signal can be implemented; and by configuring nine microphone devices, fourth-order audio coding of an audio signal can be implemented.
[0123] It should be noted that in this embodiment, the terminal uses HOA technology to perform up-order encoding on the first audio signal to generate a second audio signal. The goal of HOA technology is to reproduce the sound scene in a three-dimensional space. Compared with stereo or surround sound, HOA technology can more accurately describe the direction, distance and spatial distribution of the sound, allowing the sound to move freely in the three-dimensional space to achieve an immersive sound experience. Among them, the order in HOA technology refers to the highest order of the spherical harmonic function used to describe the three-dimensional sound field. The spherical harmonic function is used to represent the sound signal distributed in the three-dimensional space. The HOA order determines the details and accuracy of the sound scene. The higher the order, the finer the description of the sound scene, and the more sound details that can be captured and reproduced.
[0124] For example, in related art, to implement fourth-order HOA coding, at least 25 microphone devices must be configured in the terminal. In this embodiment, higher-order HOA coding can be implemented with fewer microphone devices. For example, a microphone array includes four microphone devices. In related art, these four microphone devices must meet the first restriction, meaning that they can only implement the highest-order coding processing. However, in this embodiment, fourth-order HOA coding can be implemented based on these four microphone devices.
[0125] Step S2102: Up-order coding is performed on the first audio signal according to the first HOA coefficient to generate a second audio signal.
[0126] For example, in this embodiment, the first audio signal collected by the terminal through the microphone array is: Q (ω)=[p1(ω),p2(ω),p3(ω),…,p Q (ω)] T , performing HOA encoding on the first audio signal based on the first HOA coefficient to obtain a second audio signal. During the terminal's test phase, the terminal measures a plane wave sound source in a sound field simulation or anechoic chamber based on the arrangement positions of each microphone device in the microphone array, converts the microphone sampling point constraint into a sound source direction constraint, and converts the microphone number restriction into a measurement direction restriction, thereby generating the first HOA coefficient. During the terminal's application phase, the first HOA coefficient is multiplied by the collected first audio signal to obtain a second audio signal for output.
[0127] Figure 2B FIG. 1 is a flow chart of a method for determining an HOA coefficient according to an embodiment of the present disclosure. Figure 2B As shown, the embodiment of the present disclosure relates to a method for determining an HOA coefficient, which is executed by a terminal. The method includes:
[0128] Step S2201 : determining the frequency response of the microphone array to a plane wave sound source in K discrete directions according to the geometric coordinates of the microphone array.
[0129] For example, in this embodiment, a three-dimensional coordinate system is constructed based on the terminal, and the geometric coordinates of the microphone array are determined according to the position of each microphone device in the microphone array relative to the origin of the three-dimensional coordinate system. For example, the center of gravity of the terminal is used as the origin of the three-dimensional coordinate system, the extension direction of the long side of the display device corresponding to the terminal is the X-axis direction of the three-dimensional coordinate system, the extension direction of the short side of the display device perpendicular to the long side is the Z-axis direction of the three-dimensional coordinate system, and the direction perpendicular to the plane formed by the X-axis and the Z-axis is the Y-axis direction of the three-dimensional coordinate system. After determining the three-dimensional coordinate system, the geometric coordinates (x Q ,y Q , z Q ).
[0130] In some embodiments, K is a positive integer. For example, this embodiment is applied to the test phase of the terminal, in K discrete directions (θ k ,φ k ) is used to obtain the frequency response d(ω,θ k ,φ k ). Where θ is the polar angle (the angle between the vector from the positive Z axis and the point), which ranges from [0, π]; φ is the directional angle (the angle between the vector from the positive X axis and the point on the XY plane), which ranges from [0, 2π]. By measuring the frequency response of the microphone array in K discrete directions, the response characteristics of the microphone array to different frequencies and incident angles are determined, reflecting the sensitivity of the microphone array in different directions and the time delay of the sound wave reaching the microphone array.
[0131] It should be noted that to ensure the accuracy and regularity of the frequency response, the K discrete directions in this embodiment are uniformly sampled across the sphere. This allows the microphone array to fully capture all directions of the sound field, avoiding the loss of sound field information in certain directions, ensuring that the sound field is smoothly reconstructed in all directions, reducing spectral leakage, and improving the accuracy of the spherical harmonic coefficients. For example, in this embodiment, directions can be selected at equal intervals on the sphere, or directions can be selected on the sphere so that the spherical area corresponding to each direction is the same, thereby achieving uniform sampling across the sphere.
[0132] Step S2202: Generate an array manifold matrix of the microphone array according to the frequency response.
[0133] For example, in this embodiment, the frequency responses of Q microphone devices in the microphone array in K discrete directions are measured separately. The frequency response d of each microphone device in each discrete direction is obtained. Q (ω,θ K,φ k ), construct the array manifold matrix based on the frequency response, and obtain the array manifold matrix as follows:
[0134]
[0135] Wherein, ω is the frequency response, Q is the serial number of the microphone device, and K is the direction serial number corresponding to the discrete direction.
[0136] Step S2203: Determine a plurality of spherical harmonic function values corresponding to K discrete directions.
[0137] For example, in this embodiment, the spherical harmonic function values in K discrete directions are determined to construct a spherical harmonic basis matrix, and the spherical harmonic basis matrix is applied to achieve encoding and reconstruction of the sound field.
[0138] Optionally, in some embodiments, the step of determining a plurality of spherical harmonic function values corresponding to the K discrete directions includes:
[0139] The spherical harmonics values are determined by the following formula:
[0140]
[0141] in, is the spherical harmonic function value in the discrete direction (θ, φ), θ is the polar angle, φ is the azimuth angle, n is the order of the spherical harmonic function, n is a natural number less than or equal to the first audio order, m is an integer greater than or equal to -n and less than or equal to n, is the associated Legendre polynomial.
[0142] Step S2204 : generating a discrete spherical harmonic transform (DSHT) matrix according to the multiple spherical harmonic function values, the array manifold matrix, and the first audio order.
[0143] For example, the first audio order is the initial audio order set in the terminal. Usually, the first audio order is set to 3 in the terminal. Based on the first HOA coefficient, the terminal up-codes the first audio signal into a second audio signal, and the audio order of the second audio signal is 3. In this embodiment, the spherical harmonic function values at all discrete sampling points are constructed into a matrix to obtain a DSHT (Discrete Spherical Harmonic Transform) matrix S. Each column of the matrix S corresponds to a spherical harmonic function value. Each row corresponds to the discrete direction (θ, φ) of a discrete sampling point. The matrix S can be expressed as:
[0144]
[0145] Where l = n 2+n+m+1, 0≤n≤N, -n≤m≤n, the matrix dimension is K×(N+1) 2 . Calculate the matrix in Represents a pseudo-inverse operation.
[0146] Optionally, in some embodiments, the step of “generating a discrete spherical harmonic transform (DSHT) matrix according to the multiple spherical harmonic function values and the array manifold matrix” includes:
[0147] The DSHT matrix is determined by the following formula:
[0148]
[0149] in, is the spherical harmonic function value in any discrete direction, and D(ω) is the array manifold matrix.
[0150] For example, this embodiment introduces the DSHT matrix as an intermediate mapping to project the irregular array response into spherical harmonic space, avoiding reliance on the geometric symmetry of the microphone array and improving the flexibility and applicability of sound field reconstruction. By determining the mapping relationship between the corresponding response vector of the irregular microphone array and the spherical harmonic coefficient vector, the conversion of irregular sampling points to spherical harmonic space is achieved. In this embodiment, the DSHT matrix S can be solved using the least squares method or other optimization methods.
[0151] Step S2205: Determine the first HOA coefficient according to the array manifold matrix and the DSHT matrix.
[0152] For example, in this embodiment, the response of the microphone array can be converted into spherical harmonic coefficients by combining the array manifold matrix D(ω) and the DSHT matrix S. The HOA coefficients are the elements of the spherical harmonic coefficient vector and can be extracted using the following formula:
[0153]
[0154] Among them, e1 is a unit vector, which represents a vector with 1 at the first position and 0 at other positions. The response of the microphone array is converted into spherical harmonic coefficients through the above extraction formula, and the HOA coefficients are determined based on the spherical harmonic coefficients.
[0155] Optionally, in some embodiments, the above step S2206 includes:
[0156] Determine the constraints of the first HOA coefficient based on the array manifold matrix and the DSHT matrix;
[0157] According to the constraint conditions, the first HOA coefficient is determined.
[0158] For example, in this embodiment, constraints are constructed for each HOA coefficient of the microphone array. These constraints ensure that the reconstruction and processing of the sound field meet application requirements. Based on the constraints, the HOA coefficients are optimized, thereby achieving more accurate, efficient, and robust sound field processing.
[0159] Optionally, in some embodiments, the constraint condition includes:
[0160]
[0161] in, is the first HOA coefficient, D(ω) is the array manifold matrix, S is the DSHT matrix, e n,m is a unit vector.
[0162] For example, in this embodiment Yes (N+1) 2 dimensional unit vector, the (n,m)th element is 1 and the other elements are 0. Each HOA coefficient is converted into a unit vector through the array manifold matrix and the DSHT matrix. This means that at frequency ω, the (n,m)th HOA coefficient is accurately extracted, while the other coefficients are suppressed.
[0163] For example, in this embodiment, the regularized least squares method is used to minimize the error between the reconstructed sound field and the original sound field while satisfying the constraints, thereby improving the stability and accuracy of the coefficients and ensuring the accuracy and stability of the first HOA coefficients.
[0164] Optionally, in some embodiments, the step of “determining the first HOA coefficient by a regularized least squares method according to the constraint condition” includes:
[0165] The first HOA coefficient is determined by the following formula:
[0166]
[0167] in, is the first HOA coefficient, D(ω) is the array manifold matrix, S is the DSHT matrix, e n,m is a unit vector, and λ is the regularization coefficient.
[0168] Optionally, in some embodiments, the audio processing method may further include:
[0169] When it is determined that the signal simulation degree of the second audio signal is lower than the set threshold, obtaining a second audio order, where the second audio order is greater than the first audio order, and the first audio order is an initially set order of the terminal;
[0170] determining a second HOA coefficient based on the second audio order and the first HOA coefficient;
[0171] The first audio signal is encoded according to the second HOA coefficient to generate a third audio signal.
[0172] For example, in the present embodiment, during the initial HOA encoding process of the first audio signal, the input first audio order is the audio order preset by the terminal. The terminal generates a first HOA coefficient corresponding to the first audio order according to the first audio order through the above-mentioned determination method of the HOA coefficient, and encodes the first audio signal based on the first HOA coefficient to obtain a second audio signal after the up-order processing. The second audio signal obtained after encoding is compared with the first audio signal to determine the signal simulation degree of the second audio signal. The signal simulation degree is used to indicate the degree of difference between the second audio signal and the real sound field audio received by the microphone. If the signal simulation degree is lower than the set threshold, it is determined that the currently used first audio order does not match. The terminal can obtain the second HOA coefficient based on the second audio order and the above-mentioned determination method of the HOA coefficient. The first audio signal is then encoded according to the second HOA coefficient to generate a third audio signal.
[0173] It should be noted that, in this embodiment, the terminal will measure the audio signal generated after each encoding to determine the signal simulation degree of the audio signal. If the signal simulation degree is lower than the set threshold, the order of the spherical harmonic function will be increased, and new HOA coefficients will be regenerated. Then, based on the newly generated HOA coefficients, the collected audio signal will be re-encoded until the spherical harmonic function reaches the maximum order set by the terminal; if the signal simulation degree is higher than or equal to the set threshold, the terminal will use the audio signal generated after encoding as the output audio signal.
[0174] Optionally, in some embodiments, the above step S2102 includes:
[0175] determining a third HOA coefficient of a fourth audio signal according to the first HOA coefficient, where the fourth audio signal is an audio signal having an audio frequency in the first audio signal higher than a set frequency threshold;
[0176] encoding the fourth audio signal according to the third HOA coefficient to generate a fifth audio signal;
[0177] encoding the sixth audio signal according to the first HOA coefficient to generate a seventh audio signal, where the sixth audio signal is an audio signal having an audio frequency lower than or equal to a set frequency threshold in the first audio signal;
[0178] A second audio signal is generated based on the fifth audio signal and the seventh audio signal.
[0179] For example, in this embodiment, by separating phase and amplitude processing, the directional information is retained in the aliasing frequency band, thereby improving the high-frequency processing robustness of the high-frequency audio signal. The terminal performs frequency measurement on the collected first audio signal, determines the frequency distribution of each type of audio in the first audio signal, and processes the low-frequency audio signal and the high-frequency audio signal separately. Among them, the frequency threshold is set to the spatial aliasing frequency value, and the set frequency threshold can be set to: ω th =2πc / (2d max ). For example, d max =8cm, f th ≈2125 Hz), that is, the audio signal above 2125 Hz in the first audio signal is the fourth audio signal, and the audio signal below 2125 Hz in the first audio signal is the sixth audio signal. For the fourth audio signal, using the above-described method for determining the HOA coefficients, the array manifold matrix D(ω) is replaced with the amplitude response matrix |D(ω)|, thereby ignoring phase information to avoid aliasing errors, and obtaining third HOA coefficients. The fourth audio signal is encoded based on the third HOA coefficients to generate a fifth audio signal. The sixth audio signal is encoded based on the first HOA coefficients to generate a seventh audio signal. The fifth and seventh audio signals are mixed as the second audio signal for output.
[0180] Figure 2C It is a schematic diagram showing the audio simulation experiment results according to an embodiment of the present disclosure. Figure 2D FIG is a schematic diagram showing the audio real experiment results according to an embodiment of the present disclosure. Figure 2C and Figure 2D As shown, in the related art, HOA coding is limited to Q≥(N+1) 2 In this embodiment, the limitation on the number of microphones is converted into a limitation on the measurement direction vector. Experimental verification confirms that even with only four irregularly arranged microphones, 4th-order HOA encoding and up-spreading can be achieved.
[0181] In some embodiments, in simulation experiments, as the SNR (Signal to Noise ratio) decreases from 30 dB to 0 dB, the present method consistently outperforms the baseline method in terms of the εerror metric. In particular, under low SNR conditions (0 dB), the error of the present method (6.69) is significantly lower than that of the baseline method (10.90), indicating that it has excellent noise resistance.
[0182] In some embodiments, the present method outperforms the baseline method at all SNR levels in terms of SDR metrics, indicating higher signal fidelity.
[0183] In some embodiments, simulation results show that system performance degrades as RT60 increases, mainly manifested by a significant increase in the error metric εerror. However, compared with the baseline method, the proposed method exhibits better robustness to reverberation.
[0184] In some embodiments, under mild reverberation conditions (RT60 = 0.2s), the present method reduces the error from 3.36 to 1.89, while under high reverberation environments (RT60 = 2.0s), the present method maintains the error at 300.26, demonstrating a significant performance advantage over the baseline method of 1020.52.
[0185] In some embodiments, in actual experimental results, the method significantly improves spatial correlation in the 2-5kHz range while significantly reducing reconstruction error. In the high-frequency band, the frequency partition curve shows relatively stable performance, while the baseline method shows significant fluctuations.
[0186] In some embodiments, the above experiments demonstrate the effectiveness of the frequency partitioning strategy of this method above 2kHz. Maintaining frequency consistency when no phase information is provided is more effective than providing erroneous phase information. The spherical harmonic transform (SHT) also contributes to this high-frequency stability, which is fully reflected in the SDR performance.
[0187] In some embodiments, upscaling to the fourth order significantly improves the accuracy of sound field reconstruction, while further upscaling yields negligible improvements. The SDR metric shows that the baseline method not only significantly underperforms the present method in terms of overall average performance, but also exhibits a more pronounced downward trend in the mid- and high-frequency ranges, indicating a deviation from the theoretical HOA coefficients.
[0188] The communication method involved in the embodiment of the present disclosure may include at least one of steps S2101 to S2102. For example, step S2102 may be implemented as an independent embodiment, but is not limited thereto.
[0189] In some embodiments, step S2102 is optional, and one or more of these steps may be omitted or replaced in different embodiments.
[0190] In some embodiments, reference may be made to the steps and optional implementation methods of other embodiments recorded before or after the description corresponding to this embodiment, as well as other related parts in the description, which will not be repeated here.
[0191] Figure 3A FIG. 1 is a flow chart of an audio processing method according to an embodiment of the present disclosure. Figure 3A As shown, the embodiment of the present disclosure relates to an audio processing method, which includes:
[0192] Step S3101 : determining the frequency response of the microphone array to a plane wave sound source in K discrete directions according to the geometric coordinates of the microphone array.
[0193] Optional implementations of step S3101 can be found in Figure 2B Optional implementation of step S2201, and Figure 2B Other related parts in the embodiments involved will not be described in detail here.
[0194] Step S3102: Generate an array manifold matrix of the microphone array according to the frequency response.
[0195] Optional implementations of step S3102 can be found in Figure 2B Optional implementation of step S2202, and Figure 2B Other related parts in the embodiments involved will not be described in detail here.
[0196] Step S3103: Determine a plurality of spherical harmonic function values corresponding to K discrete directions.
[0197] Optional implementations of step S3103 can be found in Figure 2B Optional implementation of step S2203, and Figure 2B Other related parts in the embodiments involved will not be described in detail here.
[0198] Step S3104: Generate a discrete spherical harmonic transform (DSHT) matrix according to the multiple spherical harmonic function values, the array manifold matrix, and the first audio order.
[0199] In some embodiments, the first audio order is an initially set order of the terminal.
[0200] Optional implementations of step S3104 can be found in Figure 2B Optional implementation of step S2204, and Figure 2B Other related parts in the embodiments involved will not be described in detail here.
[0201] Step S3105: Determine the first HOA coefficient according to the array manifold matrix and the DSHT matrix.
[0202] Optional implementations of step S3105 can be found in Figure 2B Optional implementation of step S2205, and Figure 2B Other related parts in the embodiments involved will not be described in detail here.
[0203] Step S3106: Acquire a first audio signal collected by the microphone array of the terminal.
[0204] In some embodiments, the microphone array includes at least one microphone device.
[0205] Optional implementations of step S3106 can be found in Figure 2A Optional implementation of step S2101, and Figure 2A Other related parts in the embodiments involved will not be described in detail here.
[0206] Step S3107 : performing up-order coding on the first audio signal according to the first high-order background sound effect (HOA) coefficient to generate a second audio signal.
[0207] Optional implementations of step S3107 can be found in Figure 2A Optional implementation of step S2102, and Figure 2A Other related parts in the embodiments involved will not be described in detail here.
[0208] The communication method involved in the embodiment of the present disclosure may include at least one of steps S3101 to S3107. For example, steps S3106 and S3107 may be implemented as independent embodiments, but are not limited thereto.
[0209] In some embodiments, steps S3101-S3105 are optional, and one or more of these steps may be omitted or replaced in different embodiments.
[0210] In some embodiments, reference may be made to the steps and optional implementation methods of other embodiments recorded before or after the description corresponding to this embodiment, as well as other related parts in the description, which will not be repeated here.
[0211] Figure 3B FIG. 1 is a flow chart of an audio processing method according to an embodiment of the present disclosure. Figure 3B As shown, the embodiment of the present disclosure relates to an audio processing method, which includes:
[0212] Step S3201: Acquire a first audio signal collected by a microphone array of the terminal.
[0213] In some embodiments, the microphone array includes at least one microphone device.
[0214] Optional implementations of step S3201 can be found in Figure 2A Optional implementation of step S2101, and Figure 2A Other related parts in the embodiments involved will not be described in detail here.
[0215] Step S3202: Determine a third HOA coefficient of the fourth audio signal according to the first HOA coefficient.
[0216] In some embodiments, the fourth audio signal is an audio signal in the first audio signal whose audio frequency is higher than a set frequency threshold.
[0217] Optional implementations of step S3202 can be found in Figure 2A Optional implementation of step S2102, and Figure 2A Other related parts in the embodiments involved will not be described in detail here.
[0218] Step S3203: Encode the fourth audio signal according to the third HOA coefficient to generate a fifth audio signal.
[0219] Optional implementations of step S3203 can be found in Figure 2A Optional implementation of step S2102, and Figure 2A Other related parts in the embodiments involved will not be described in detail here.
[0220] Step S3204: Encode the sixth audio signal according to the first HOA coefficient to generate a seventh audio signal.
[0221] In some embodiments, the sixth audio signal is an audio signal in the first audio signal whose audio frequency is lower than or equal to a set frequency threshold.
[0222] Optional implementations of step S3204 can be found in Figure 2A Optional implementation of step S2102, and Figure 2A Other related parts in the embodiments involved will not be described in detail here.
[0223] Step S3205: Generate a second audio signal according to the fifth audio signal and the seventh audio signal.
[0224] Optional implementations of step S3205 can be found in Figure 2A Optional implementation of step S2102, and Figure 2A Other related parts in the embodiments involved will not be described in detail here.
[0225] The communication method involved in the embodiment of the present disclosure may include at least one of steps S3201 to S3205. For example, steps S3202+S3203+S3204+S3205 may be implemented as an independent embodiment, but are not limited thereto.
[0226] In some embodiments, step S3201 is optional, and one or more of these steps may be omitted or replaced in different embodiments.
[0227] In some embodiments, reference may be made to the steps and optional implementation methods of other embodiments recorded before or after the description corresponding to this embodiment, as well as other related parts in the description, which will not be repeated here.
[0228] Figure 4 FIG. 1 is a flow chart of an audio processing method according to an embodiment of the present disclosure. Figure 4 As shown, the embodiment of the present disclosure relates to an audio processing method, which is executed by a terminal. The method includes:
[0229] Step S4101: input the geometric coordinates of the microphone array and construct an array manifold matrix.
[0230] For example, the geometric coordinates of the microphone array are (x Q ,y Q , z Q ), where Q = 1, 2, 3, 4…K, in K discrete directions (θ k ,φ k ) and obtain the frequency response d(ω,θ k ,φ k ).
[0231] Construct the array manifold matrix D(ω) with dimension 4×K, that is:
[0232]
[0233] Among them, d q (ω,θ k ,φ k ) indicates the qth microphone in the direction (θ k ,φ k )'s complex response.
[0234] Step S4102: Input the target HOA order and direction sampling points to generate a discrete spherical harmonic transformation matrix.
[0235] For example, the input target HOA order N = 4, and the direction sampling point (θ k ,φ k ), calculate the spherical harmonic function value of each sampling point and construct the spherical harmonic basis matrix Its elements are:
[0236]
[0237] Where 0≤n≤N,-n≤m≤n, the matrix dimension is K×(N+1) 2 ; Calculate the DSHT (Discrete Spherical Harmonic Transform) matrix in Represents a pseudo-inverse operation.
[0238] Step S4103: Input the array manifold matrix and the DSHT matrix, and solve the beamforming weights using the regularized least squares method.
[0239] For example, for each target HOA coefficient Construct the constraint equation:
[0240]
[0241] in, is a unit vector (l=n 2 +n+m+1 bits are 1).
[0242] Solving for beamforming weights using regularized least squares
[0243]
[0244] Among them, λ is the Tikhonov regularization coefficient, which is used to suppress noise sensitivity.
[0245] Step S4104: input a spatial aliasing frequency threshold and perform aliasing coding on the audio signal.
[0246] For example, the spatial aliasing frequency threshold is: ω th =2πc / (2d max ), for example, d max =8cm, f th ≈2125Hz.
[0247] In some embodiments, for ω>ω th The high-band audio signal is replaced by the array manifold D(ω) with the amplitude response matrix |D(ω)|, thereby ignoring the phase information to avoid aliasing errors. Recalculate the beamforming weights for the high-band audio signal Used to optimize amplitude matching.
[0248] Step S4105: Perform real-time HOA encoding on the audio signal based on the HOA coefficient to obtain an HOA audio signal for output.
[0249] For example, the input microphone signal p(ω)=[p1(ω),p2(ω),p3(ω),p4(ω)] T , multiply the microphone signal by the HOA coefficient to obtain the HOA audio signal, that is, the HOA audio signal is calculated and determined by the following formula:
[0250]
[0251] In some embodiments, by converting microphone sampling point constraints into sound source direction constraints, the limitation on the number of microphones can be overcome and fourth-order encoding can be achieved. Furthermore, the DSHT matrix S is introduced as an intermediate mapping to project the irregular array response into spherical harmonic space, avoiding reliance on array geometric symmetry.
[0252] In some embodiments, by separating phase and amplitude processing, directional information is retained in the aliasing frequency band, thereby improving high-frequency robustness.
[0253] For example, the spherical harmonics in this embodiment are:
[0254]
[0255] in, is the Legendre polynomial.
[0256] In some embodiments, reference may be made to the steps and optional implementation methods of other embodiments recorded before or after the description corresponding to this embodiment, as well as other related parts in the description, which will not be repeated here.
[0257] The present disclosure also provides an apparatus (also referred to as an electronic device, etc.) for implementing any of the above audio processing methods. For example, a device is provided, comprising units or modules for implementing each step performed by the electronic device in any of the above audio processing methods.
[0258] It should be understood that the division of the various units or modules in the above device is only a division of logical functions. In actual implementation, they can be fully or partially integrated into one physical entity, or they can be physically separated. In addition, the units or modules in the device can be implemented in the form of a processor calling software: for example, the device includes a processor, the processor is connected to a memory, and the memory stores instructions. The processor calls the instructions stored in the memory to implement any of the above audio processing methods or implement the functions of the various units or modules of the above device, wherein the processor is, for example, a general-purpose processor, such as a central processing unit (CPU) or a microprocessor, and the memory is a memory within the device or a memory outside the device. Alternatively, the units or modules in the device can be implemented in the form of hardware circuits, and the functions of some or all of the units or modules can be realized by designing the hardware circuits. The above-mentioned hardware circuits can be understood as one or more processors; for example, in one implementation, the above-mentioned hardware circuit is an application-specific integrated circuit (ASIC), which realizes the functions of some or all of the above units or modules by designing the logical relationship of the components in the circuit; for example, in another implementation, the above-mentioned hardware circuit can be realized by a programmable logic device (PLD). Taking a field programmable gate array (FPGA) as an example, it can include a large number of logic gate circuits, and the connection relationship between the logic gate circuits is configured by configuring the configuration file, thereby realizing the functions of some or all of the above units or modules. All units or modules of the above devices can be realized in the form of software called by the processor, or in the form of hardware circuits, or in part by the form of software called by the processor, and the rest by hardware circuits.
[0259] In the embodiments of the present disclosure, the processor is a circuit with signal processing capabilities. In one implementation, the processor can be a circuit with instruction reading and execution capabilities, such as a central processing unit (CPU), a microprocessor, a graphics processing unit (GPU) (which can be understood as a microprocessor), or a digital signal processor (DSP). In another implementation, the processor can implement certain functions through the logical relationship of the hardware circuit. The logical relationship of the above-mentioned hardware circuit is fixed or reconfigurable. For example, the processor is a hardware circuit implemented by an application-specific integrated circuit (ASIC) or a programmable logic device (PLD), such as an FPGA. In a reconfigurable hardware circuit, the process of the processor loading a configuration document and implementing the hardware circuit configuration can be understood as the process of the processor loading instructions to implement the functions of some or all of the above units or modules. In addition, it can also be a hardware circuit designed for artificial intelligence, which can be understood as an ASIC, such as a neural network processing unit (NPU), a tensor processing unit (TPU), a deep learning processing unit (DPU), etc.
[0260] It should be noted that the various implementation methods / embodiments involved in the embodiments of the present disclosure can be used in conjunction with the aforementioned and / or later-described embodiments, or can be used independently. Whether used alone or in conjunction with the aforementioned and / or later-described embodiments, the implementation principles are similar. In the implementation of the present disclosure, some embodiments are described by using implementation methods used together. Of course, such examples are not intended to limit the embodiments of the present disclosure.
[0261] Figure 5 is a schematic diagram of a terminal structure according to an embodiment of the present disclosure. Terminal 500 is used to execute any of the above methods. In some embodiments, Figure 5 As shown, the terminal 500 may include at least one of: a transceiver module 501, a processing module 502, etc.
[0262] In some embodiments, the transceiver module 501 is configured to obtain a first audio signal collected by a microphone array of a terminal, where the microphone array includes at least one microphone device. The processing module 502 is configured to perform up-order coding on the first audio signal based on a first high-order background sound (HOA) coefficient to generate a second audio signal.
[0263] In some embodiments, the transceiver module may include a transmitting module and / or a receiving module, and the transmitting module and the receiving module may be separate or integrated. Optionally, the transceiver module may be interchangeable with the transceiver.
[0264] In some embodiments, the processing module may be a single module or may include multiple submodules. Optionally, the multiple submodules respectively execute all or part of the steps required to be executed by the processing module.
[0265] In some embodiments, the processing module may be interchangeable with the processor, and the transceiver module may be interchangeable with the transceiver.
[0266] Figure 6 6 is a schematic diagram of the structure of an electronic device 600 according to an embodiment of the present disclosure. For example, the electronic device 600 may be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.
[0267] Reference Figure 6 The electronic device 600 may include one or more of the following components: a processing component 602 , a memory 604 , a power component 606 , a multimedia component 608 , an audio component 610 , an input / output interface 612 , a sensor component 614 , and a communication component 616 .
[0268] The processing component 602 generally controls the overall operation of the electronic device 600, such as operations associated with display, phone calls, data communications, camera operation, and recording operations. The processing component 602 may include one or more processors 620 to execute instructions to perform all or part of the steps of the above-described method. In addition, the processing component 602 may include one or more modules to facilitate interaction between the processing component 602 and other components. For example, the processing component 602 may include a multimedia module to facilitate interaction between the multimedia component 608 and the processing component 602.
[0269] In some embodiments, processor 620 performs at least one of the processing steps.
[0270] The memory 604 is configured to store various types of data to support operations on the electronic device 600. Examples of such data include instructions for any application or method operating on the electronic device 600, contact data, phone book data, messages, pictures, videos, etc. The memory 604 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk.
[0271] The power supply assembly 606 provides power to the various components of the electronic device 600. The power supply assembly 606 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the electronic device 600.
[0272] The multimedia component 608 includes a screen that provides an output interface between the electronic device 600 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, slides, and gestures on the touch panel. The touch sensor can not only sense the boundaries of the touch or slide action, but also detect the duration and pressure associated with the touch or slide operation. In some embodiments, the multimedia component 608 includes a front camera and / or a rear camera. When the electronic device 600 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each front camera and rear camera can be a fixed optical lens system or have a focal length and optical zoom capability.
[0273] The audio component 610 is configured to output and / or input audio signals. For example, the audio component 610 includes a microphone (MIC), which is configured to receive external audio signals when the electronic device 600 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signal can be further stored in the memory 604 or transmitted via the communication component 616. In some embodiments, the audio component 610 also includes a speaker for outputting audio signals.
[0274] The input / output interface 612 provides an interface between the processing component 602 and peripheral interface modules, such as a keyboard, a click wheel, buttons, etc. These buttons may include but are not limited to: a home button, a volume button, a start button, and a lock button.
[0275] The sensor assembly 614 includes one or more sensors for providing various aspects of status assessment for the electronic device 600. For example, the sensor assembly 614 can detect the open / closed state of the electronic device 600, the relative positioning of components, such as the display and keypad of the electronic device 600. The sensor assembly 614 can also detect changes in the position of the electronic device 600 or a component of the electronic device 600, the presence or absence of user contact with the electronic device 600, the orientation or acceleration / deceleration of the electronic device 600, and temperature changes of the electronic device 600. The sensor assembly 614 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor assembly 614 may also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor assembly 614 may also include an accelerometer, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.
[0276] The communication component 616 is configured to facilitate wired or wireless communication between the electronic device 600 and other devices. The electronic device 600 can access a wireless network based on a communication standard, such as WiFi, 2G or 3G, or a combination thereof. In an exemplary embodiment, the communication component 616 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 616 also includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology and other technologies.
[0277] In an exemplary embodiment, the electronic device 600 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the above methods.
[0278] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 604 including instructions, and the instructions can be executed by the processor 620 of the electronic device 600 to perform the above method. For example, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.
[0279] Figure 7700 according to an embodiment of the present disclosure. Figure 7 The structure of the chip 700 is shown, but is not limited thereto.
[0280] The chip 700 includes one or more processors 701. The chip 700 is configured to execute any of the above methods.
[0281] In some embodiments, the chip 700 further includes one or more interface circuits 702. Optionally, terms such as interface circuit, interface, and transceiver pins may be used interchangeably. In some embodiments, the chip 700 further includes one or more memories 703 for storing data and / or instructions. Optionally, all or part of the memories 703 may be located outside the chip 700. Optionally, the interface circuit 702 is connected to the memory 703, and the interface circuit 702 may be used to receive data and / or instructions from the memory 703 or other devices, or to send data and / or instructions to the memory 703 or other devices. For example, the interface circuit 702 may read data and / or instructions stored in the memory 703 and send the data and / or instructions to the processor 701.
[0282] In some embodiments, processor 701 performs at least one of the processing steps.
[0283] The modules and / or devices described in various embodiments, such as virtual devices, physical devices, and chips, can be arbitrarily combined or separated according to circumstances. Optionally, some or all steps can also be performed collaboratively by multiple modules and / or devices, which is not limited here.
[0284] The present disclosure also proposes a storage medium having instructions stored thereon, which, when the instructions are executed on an electronic device, causes the electronic device to execute any of the above methods. Optionally, the storage medium is an electronic storage medium. Optionally, the storage medium is a computer-readable storage medium, but is not limited thereto, and may also be a storage medium readable by other devices. Optionally, the storage medium may be a non-transitory storage medium, but is not limited thereto, and may also be a temporary storage medium.
[0285] The present disclosure also provides a program product, including a program and / or instructions, which, when executed by an electronic device, causes the electronic device to perform any of the above methods. Optionally, the program product is a computer program product. Optionally, the program product is stored on the above storage medium.
[0286] The present disclosure also proposes a computer program, which, when executed on a computer, causes the computer to perform any one of the above methods.
Claims
1. An audio processing method, executed by a terminal, characterized in that: The method comprises: Acquire a first audio signal collected by a microphone array of the terminal, where the microphone array includes at least one microphone device; The first audio signal is up-order encoded according to the first high-order background sound (HOA) coefficient to generate a second audio signal.
2. The method according to claim 1, characterized in that The number of the at least one microphone device and the audio order of the second audio signal do not need to satisfy a first restriction condition.
3. The method according to claim 1 or 2, characterized in that The method further comprises: Determining, according to the geometric coordinates of the microphone array, a frequency response of the microphone array to a plane wave sound source in K discrete directions, where K is a positive integer; generating an array manifold matrix of the microphone array according to the frequency response; Determine a plurality of spherical harmonic function values corresponding to the K discrete directions respectively; generating a discrete spherical harmonic transform (DSHT) matrix according to the plurality of spherical harmonic function values, the array manifold matrix, and a first audio order, wherein the first audio order is a set order of the terminal; The first HOA coefficient is determined according to the array manifold matrix and the DSHT matrix.
4. The method according to claim 3, wherein determining the first HOA coefficient according to the array manifold matrix and the DSHT matrix comprises: Determining a constraint condition of the first HOA coefficient according to the array manifold matrix and the DSHT matrix; The first HOA coefficient is determined according to the constraint condition.
5. The method according to claim 3 or 4, characterized in that Determining the multiple spherical harmonic function values corresponding to the K discrete directions includes: The spherical harmonics values are determined by the following formula: Among them, the is the spherical harmonic function value in any discrete direction, θ is the polar angle, φ is the azimuth angle, n is a natural number less than or equal to the first audio order, m is an integer greater than or equal to -n and less than or equal to n, is the associated Legendre polynomial.
6. The method according to any one of claims 3 to 5, characterized in that Generating a discrete spherical harmonic transform (DSHT) matrix according to the plurality of spherical harmonic function values, the array manifold matrix, and the first audio order includes: The DSHT matrix is determined by the following formula: Among them, the is the spherical harmonic function value in any discrete direction, and D(ω) is the array manifold matrix.
7. The method according to claim 4, characterized in that The constraints include: Among them, the is the first HOA coefficient, D(ω) is the array manifold matrix, S is the DSHT matrix, and e n,m is a unit vector.
8. The method according to claim 4 or 7, characterized in that The determining the first HOA coefficient according to the constraint condition includes: The first HOA coefficient is determined by the following formula: Among them, the is the first HOA coefficient, D(ω) is the array manifold matrix, S is the DSHT matrix, and e n,m is a unit vector, and λ is a regularization coefficient.
9. The method according to any one of claims 1 to 8, characterized in that The method further comprises: When the signal simulation degree of the second audio signal is lower than a set threshold, obtaining a second audio order, where the second audio order is greater than the first audio order, and the first audio order is a set order of the terminal; determining a second HOA coefficient according to the second audio order and the first HOA coefficient; The first audio signal is encoded according to the second HOA coefficient to generate a third audio signal.
10. The method according to any one of claims 1 to 9, characterized in that The encoding of the first audio signal according to the first high-order background sound effect HOA coefficient to generate a second audio signal includes: determining, based on the first HOA coefficient, a third HOA coefficient of a fourth audio signal, where the fourth audio signal is an audio signal having an audio frequency in the first audio signal higher than a set frequency threshold; encoding the fourth audio signal according to the third HOA coefficient to generate a fifth audio signal; encoding the sixth audio signal according to the first HOA coefficient to generate a seventh audio signal, wherein the sixth audio signal is an audio signal having an audio frequency lower than or equal to the set frequency threshold in the first audio signal; The second audio signal is generated according to the fifth audio signal and the seventh audio signal.
11. An audio processing device, characterized in that: Applied to a terminal, the device includes: a transceiver module, configured to obtain a first audio signal collected by a microphone array of the terminal, the microphone array comprising at least one microphone device; The processing module is configured to perform up-order coding on the first audio signal according to the first high-order background sound effect (HOA) coefficient to generate a second audio signal.
12. A terminal, characterized in that: The terminal is used to execute the audio processing method according to any one of claims 1 to 10.
13. A storage medium storing instructions, characterized in that: When the instruction is executed on a terminal, the terminal is enabled to execute the audio processing method according to any one of claims 1 to 10.
14. A program product, comprising at least one of a program and instructions, characterized in that: When at least one of the program and the instruction is executed by the terminal, the steps of the audio processing method according to any one of claims 1 to 10 are implemented.
Citation Information
Cited By
Hearing aid speech enhancement method and device based on artificial intelligence, and medium
CN121531284A