Sound signal processing methods, devices and headphones
By setting up dual microphones on the headphones and using cross-correlation analysis and head transfer function to determine the location of the sound source, and configuring filters to process the headphone signal, the problem of decreased external sound perception when wearing headphones is solved, achieving the same external sound perception effect as when not wearing headphones, thus improving user safety and experience.
Patent Information
- Application Number
- CN202411984708.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-30
- Publication Date
- 2025-11-14
- Estimated Expiration
- 2044-12-30
AI Technical Summary
When wearing headphones, users' ability to perceive external sounds decreases, resulting in an inability to respond to external sounds in a timely manner, affecting safety and communication efficiency. Existing ambient sound modes cannot fully simulate the natural auditory experience when not wearing headphones.
By setting microphones in the left and right earpieces of the headphones, symmetrical external sound signals are acquired from both ears. The location of the sound source is determined by cross-correlation analysis and head transfer function, and a filter is configured for filtering to improve the perception of external sound signals.
It achieves the same effect of external sound signals when wearing headphones as when not wearing headphones, improves the user's ability to perceive external sounds, ensures timely response, and enhances user experience and security.
Smart Images

Figure CN119893368B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of headphone technology, and more specifically, to a sound signal processing method, apparatus, and headphone. Background Technology
[0002] With the rapid development of modern technology, headphones have become an indispensable entertainment and communication tool in people's daily lives. Whether commuting, working, or relaxing, headphones provide users with a convenient audio experience, making the enjoyment of various audio content such as music, calls, and games more personalized and private.
[0003] However, despite the numerous conveniences headphones offer, they also present some challenges. When users wear headphones, the direct contact between the headphones and the ears partially or completely isolates them from external sounds. This isolation makes users less sensitive to external sounds, potentially leading to delayed responses to external noises in many situations, thus impacting user safety and communication efficiency. For example, in public places, users may miss urgent announcements or important calls from others because they are wearing headphones.
[0004] In the existing technology, although some attempts have been made to solve this problem by adding an ambient sound mode or a transparency mode, these solutions are often limited in their effectiveness and cannot fully simulate the natural auditory experience when not wearing headphones. Summary of the Invention
[0005] One objective of this application is to provide a novel sound signal processing method that can improve a user's ability to perceive external sounds while wearing headphones, ensuring that the user can respond to external sounds in a timely and accurate manner, thereby enhancing user experience and security.
[0006] According to a first aspect of the present invention, a sound signal processing method is provided, the sound signal processing method being applied to an earphone, the earphone including a left earphone and a right earphone, the left earphone and the right earphone each being provided with at least one microphone, the method comprising:
[0007] The first sound signal is obtained by acquiring external sound signals from two microphones symmetrically positioned in each ear;
[0008] Based on the first sound signal, determine the location of the sound source of the external sound signal relative to the first sound source of the headphones;
[0009] Based on the location of the first sound source, a first filter corresponding to the location of the first sound source is determined; the first sound signal is filtered using the filter parameters configured for the first filter to obtain the target sound signal.
[0010] Optionally, the first sound signal includes a left-side sound signal collected by a microphone located at the left earphone position and a right-side sound signal collected by a microphone located at the right earphone position. Determining the location of the external sound source relative to the earphones based on the first sound signal includes:
[0011] Perform cross-correlation analysis on the left and right sound signals to determine the first time difference between them;
[0012] Based on the first time difference, candidate coordinate values of the sound source of the external sound signal relative to the headphone coordinate system are determined; wherein, the headphone coordinate system is a coordinate system constructed based on the left and right headphones;
[0013] Based on the root mean square value of the amplitude of the left-side sound signal and / or the root mean square value of the amplitude of the right-side sound signal, a target coordinate value is selected from the candidate coordinate values as the first sound source location.
[0014] Optionally, the step of performing cross-correlation analysis on the left and right sound signals to determine the first time difference between the left and right sound signals includes:
[0015] A first bandpass filter is applied to the left-side sound signal and the right-side sound signal within a first set frequency range to obtain a first left-side filtered sound signal and a first right-side filtered sound signal.
[0016] Perform cross-correlation analysis on the first left-side filtered audio signal and the first right-side filtered audio signal to determine the first time difference between the first left-side filtered audio signal and the first right-side filtered audio signal;
[0017] The step of selecting a target coordinate value from the candidate coordinate values as the first sound source location based on the root mean square value of the amplitude of the left-side sound signal and / or the root mean square value of the amplitude of the right-side sound signal includes:
[0018] A second bandpass filter is applied to the left-side sound signal and the right-side sound signal within a second set frequency range to obtain a second left-side filtered sound signal and a second right-side filtered sound signal.
[0019] Based on the root mean square value of the amplitude of the second left-side filtered sound signal and / or the root mean square value of the amplitude of the second right-side filtered sound signal, a target coordinate value is selected from the candidate coordinate values as the first sound source location; wherein, the lower limit of the second set frequency range is greater than the upper limit of the first set frequency range.
[0020] Optionally, the step of performing cross-correlation analysis on the left and right sound signals to determine the first time difference between the left and right sound signals includes:
[0021] Perform cross-correlation analysis on the left and right sound signals, and take the time difference between the left and right sound signals corresponding to the maximum value in the cross-correlation analysis results as the first time difference.
[0022] Optionally, determining the candidate coordinate values of the sound source of the external sound signal relative to the headphone coordinate system based on the first time difference includes:
[0023] The second time difference is determined based on the first time difference and the sampling rate of the first sound signal;
[0024] The candidate coordinate values are determined based on the second time difference.
[0025] Optionally, selecting the target coordinate value from the candidate coordinate values based on the root mean square value of the amplitude of the left-side sound signal and / or the root mean square value of the amplitude of the right-side sound signal includes:
[0026] Obtain the head transfer function corresponding to each preset sound source location;
[0027] Based on the head transfer function corresponding to each preset sound source location, determine the candidate head transfer function corresponding to the candidate coordinate value;
[0028] The reference signal energy corresponding to the candidate coordinate value is determined based on the candidate head transfer function;
[0029] The target coordinate value is selected from the candidate coordinate values based on the root mean square value of the amplitude of the left sound signal and / or the root mean square value of the amplitude of the right sound signal, and the reference signal energy of the candidate coordinate values.
[0030] Optionally, determining the first filter corresponding to the first sound source location based on the first sound source location includes:
[0031] Based on the first sound source location and the first correspondence, a first filter corresponding to the first sound source location is determined; wherein, the first correspondence is a preset correspondence between the sound source location and the filter.
[0032] Optionally, the step of determining the first correspondence includes:
[0033] Obtain the head transfer function corresponding to each preset sound source location;
[0034] For any preset sound source location, the filter parameters corresponding to the preset sound source location are determined according to the head transfer function corresponding to the preset sound source location, thus obtaining the filter parameters corresponding to each preset sound source location.
[0035] Configure the filter to be configured in the filter array according to the filter parameters corresponding to each preset sound source location, and establish the correspondence between each preset sound source location and the configured filter in the filter array to obtain the first correspondence.
[0036] Optionally, determining the first filter corresponding to the first sound source location based on the first sound source location includes:
[0037] Based on the first sound source location and the second correspondence, the first filter parameters corresponding to the first sound source location are determined; wherein, the second correspondence is a preset correspondence between the sound source location and the filter parameters;
[0038] Configure the filter to be configured according to the first filter parameters to obtain the first filter corresponding to the first sound source location.
[0039] Optionally, the step of determining the second correspondence includes:
[0040] Obtain the head transfer function corresponding to each preset sound source location;
[0041] For any preset sound source location, the filter parameters corresponding to the preset sound source location are determined according to the head transfer function corresponding to the preset sound source location, thus obtaining the filter parameters corresponding to each preset sound source location.
[0042] The second correspondence is established based on the filter parameters corresponding to each preset sound source location.
[0043] Optionally, determining the filter parameters corresponding to the preset sound source location based on the head transfer function corresponding to the preset sound source location includes:
[0044] Based on the head transfer function corresponding to the preset sound source location and the preset relational expression, the corresponding filter parameters are determined; wherein, the preset relational expression reflects the relationship between the head transfer function and the filter parameters.
[0045] According to a second aspect of the present invention, a sound signal processing apparatus is also provided, comprising a memory and a processor, the memory being configured to store executable instructions; the processor being configured to operate under the control of the instructions to perform the method as described in the first aspect.
[0046] According to a third aspect of the present invention, an earphone is also provided, including a left earphone and a right earphone, and a sound signal processing device as described in the second aspect;
[0047] The left and right earpieces are each equipped with at least one microphone. External sound signals are collected through one microphone in each earpiece to obtain a first sound signal, which is then sent to the sound signal processing device.
[0048] One beneficial effect of this invention is that by acquiring a first sound signal from external sound signals collected by two symmetrical microphones in each ear, determining the location of the sound source of the external sound signal relative to the headphones based on the first sound signal, determining a first filter corresponding to the first sound source location based on the first sound source location, and filtering the first sound signal using the filter parameters configured for the first filter to obtain the target sound signal, this invention achieves the same effect as the sound heard by the human ear when wearing headphones as when not wearing headphones. This improves the user's ability to perceive external sounds while wearing headphones, ensuring that the user can react to external sounds promptly and accurately, thereby enhancing user experience and safety. Attached Figure Description
[0049] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments of the invention and, together with their description, serve to explain the principles of the invention.
[0050] Figure 1 This is a schematic diagram of the hardware structure of the headphones according to an embodiment of this application;
[0051] Figure 2 This is a schematic flowchart of a sound signal processing method according to an embodiment of the present invention;
[0052] Figure 3 This is a schematic diagram of the sound propagation path from an external sound signal to the human ear according to an embodiment of the present invention;
[0053] Figure 4 This is a schematic diagram illustrating the transmission of external sound signals corresponding to different sound sources according to an example of the present invention;
[0054] Figure 5 This is a schematic diagram showing the arrangement of each preset sound source location according to an example of the present invention;
[0055] Figure 6 This is a schematic diagram of the head transfer function corresponding to each preset sound source location according to an embodiment of the present invention;
[0056] Figure 7 This is a schematic diagram of the structure of a sound signal processing device according to an embodiment of the present invention;
[0057] Figure 8This is a schematic diagram of the structure of an earphone according to an embodiment of the present invention. Detailed Implementation
[0058] Various exemplary embodiments of the present invention will now be described in detail with reference to the accompanying drawings. It should be noted that, unless otherwise specifically stated, the relative arrangement, numerical expressions, and values of the components and steps set forth in these embodiments do not limit the scope of the invention.
[0059] The following description of at least one exemplary embodiment is merely illustrative and is in no way intended to limit the invention or its application or use.
[0060] Techniques, methods, and equipment known to those skilled in the art may not be discussed in detail, but where appropriate, such techniques, methods, and equipment should be considered part of the specification.
[0061] In all the examples shown and discussed herein, any specific values should be interpreted as merely exemplary and not as limitations. Therefore, other examples of exemplary embodiments may have different values.
[0062] It should be noted that similar labels and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be discussed further in subsequent figures.
[0063] <Hardware Configuration>
[0064] Figure 1 This is a block diagram of the hardware configuration of the earphone 100 according to an embodiment of this application.
[0065] like Figure 1 As shown, the earphone 100 includes a left earphone 1000, a right earphone 2000, and a sound signal processing device 3000.
[0066] Headphones can be in-ear headphones or over-ear headphones.
[0067] The left earphone 1000 and the right earphone 2000 are respectively connected to the sound signal processing device 3000. The left earphone 1000 and the right earphone 2000 are each equipped with at least one microphone to collect external sound signals and obtain a first sound signal.
[0068] In this embodiment, refer to Figure 1 As shown, the left earphone 1000 may include a processor 1100, a memory 1200, a microphone 1300, a communication device 1400, etc.
[0069] Processor 1100 may be a mobile processor. Memory 1200 includes, for example, ROM (Read-Only Memory), RAM (Random Access Memory), and non-volatile memory such as a hard disk. Microphone 1300 can collect external sound signals. Communication device 1400 is capable of wired or wireless communication, and may include short-range communication devices, such as any device that performs short-range wireless communication based on short-range wireless communication protocols such as Hilink, WiFi (IEEE 802.11), Mesh, Bluetooth, ZigBee, Thread, Z-Wave, NFC, UWB, and LiFi. Communication device 1400 may also include long-range communication devices, such as any device that performs WLAN, GPRS, or 2G / 3G / 4G / 5G long-range communication.
[0070] In this embodiment, the memory 1200 of the left earphone 1000 is used to store instructions for controlling the processor 1100 to operate in order to at least execute the method performed by the left earphone according to any embodiment of the present invention. Those skilled in the art can design instructions based on the disclosed scheme of the present invention. How the instructions control the processor to operate is well known in the art and will not be described in detail here.
[0071] In this embodiment, refer to Figure 1 As shown, the right earphone 2000 may include a processor 2100, a memory 2200, a microphone 2300, a communication device 2400, etc.
[0072] Processor 2100 may be a mobile processor. Memory 2200 includes, for example, ROM (Read-Only Memory), RAM (Random Access Memory), and non-volatile memory such as a hard disk. Microphone 2300 can collect external sound signals. Communication device 2400 is capable of wired or wireless communication, for example. Communication device 1400 may include short-range communication devices, such as any device that performs short-range wireless communication based on short-range wireless communication protocols such as Hilink, WiFi (IEEE 802.11), Mesh, Bluetooth, ZigBee, Thread, Z-Wave, NFC, UWB, and LiFi. Communication device 2400 may also include long-range communication devices, such as any device that performs WLAN, GPRS, or 2G / 3G / 4G / 5G long-range communication.
[0073] In this embodiment, the memory 2200 of the right earphone 2000 is used to store instructions for controlling the processor 2100 to operate in order to at least execute the method performed by the right earphone according to any embodiment of the present invention. Those skilled in the art can design instructions according to the disclosed scheme of the present invention. How the instructions control the processor to operate is well known in the art and will not be described in detail here.
[0074] The sound signal processing device 3000 is used to process the left sound signal obtained by the microphone 1300 of the left earphone and the right sound signal obtained by the microphone 2300 of the right earphone to obtain the target sound signal.
[0075] In this embodiment, refer to Figure 1 As shown, the audio signal processing device 3000 may include a processor 3100, a memory 3200, a communication device 3300, a filter 3400, etc.
[0076] Processor 3100 may be a mobile processor. Memory 3200 includes, for example, ROM (Read-Only Memory), RAM (Random Access Memory), and non-volatile memory such as a hard disk. Communication device 3300 may be capable of wired or wireless communication. Communication device 3300 may include short-range communication devices, such as any device that performs short-range wireless communication based on short-range wireless communication protocols such as Hilink, WiFi (IEEE 802.11), Mesh, Bluetooth, ZigBee, Thread, Z-Wave, NFC, UWB, and LiFi. Communication device 3300 may also include long-range communication devices, such as any device that performs WLAN, GPRS, or 2G / 3G / 4G / 5G long-range communication. Filter 3400 is used to filter the first sound signal obtained by microphones 1300 and 2300 from external sound signals to obtain the target sound signal.
[0077] In this embodiment, the memory 3200 of the sound signal processing apparatus 3000 is used to store instructions for controlling the processor 3100 to operate to at least execute the method performed by the sound signal processing apparatus according to any embodiment of the present invention. Those skilled in the art can design instructions according to the disclosed scheme of the present invention. How the instructions control the processor to operate is well known in the art and will not be described in detail here.
[0078] Despite Figure 1 The present invention illustrates multiple devices of the audio signal processing apparatus 3000, but may refer to only some of these devices. For example, the audio signal processing apparatus 3000 may refer only to the memory 3200 and the processor 3100.
[0079] <Method Implementation>
[0080] Figure 2 This is a schematic flowchart of a sound signal processing method according to an embodiment of this application. The method is applied to a sound signal processing device for headphones, i.e. Figure 1The sound signal processing device 3000 is included. The earphone 100 includes a left earphone 1000 and a right earphone 2000, with at least one microphone 1300 in the left earphone and at least one microphone 2300 in the right earphone.
[0081] according to Figure 2 As shown, the sound signal processing method of this embodiment may include the following steps S2100 to S2400:
[0082] Step S2100: Obtain the first sound signal obtained by collecting external sound signals from one microphone on each ear.
[0083] In this embodiment, the two microphones symmetrically positioned in each ear refer to the two microphones in the left and right earpieces that are symmetrically placed relative to the center of the wearer's head in terms of spatial layout. Both of these microphones can be referred to as feedforward microphones.
[0084] like Figure 3 As shown, it can represent the sound propagation path of the external sound signal to the human ear position when wearing headphones. The external sound signal is x(t), and the first sound signal obtained by the two symmetrical microphones of each ear can be represented as x(t)*R(s), where R(s) represents the total transfer function of the external sound signal from the environment to the two symmetrical microphone positions of each ear.
[0085] In some embodiments, the first sound signal includes a left-side sound signal collected by a microphone located at the left earphone position and a right-side sound signal collected by a microphone located at the right earphone position.
[0086] In this embodiment, R(s) can be the sum of the first transfer function R1(s) of the external sound signal being transmitted from the environment to the microphone position located at the left earphone position and the second transfer function R2(s) of the external sound signal being transmitted from the environment to the microphone position located at the right earphone position. Therefore, the sound signal on the left can be represented as x(t)*R1(s), and the sound signal on the right can be represented as x(t)*R2(s).
[0087] Step S2200: Determine the location of the sound source of the external sound signal relative to the first sound source of the headphones based on the first sound signal.
[0088] In existing technology, the phase difference can be determined based on the interaural time difference (ITD) between the left and right sound signals, and then the location of the external sound source relative to the first sound source in the headphones can be determined based on this phase difference. The interaural time difference refers to the time difference between the arrival of the external sound signal at each of the two symmetrical microphones in each ear. The interaural time difference causes a wavelength-dependent phase difference.
[0089] However, in some cases, the accuracy of determining the location of the first sound source using the time difference between the ears is not high. For example, as... Figure 4 As shown, the time it takes for an external sound signal to travel from the front of the wearer to each of the two symmetrical microphones in each ear is T1 and T2. The time it takes for the external sound signal to travel from the back of the wearer to each of the two symmetrical microphones in each ear is T1' and T2'. Furthermore, T1 and T1' are equal, and T2 and T2' are equal. Therefore, simply using the time difference between the external sound signal's arrival at each of the two symmetrical microphones in each ear is insufficient to distinguish whether the external sound signal's source is in front of or behind the wearer, resulting in low accuracy.
[0090] Therefore, the inventors proposed a new method for locating the source of external sound signals to solve the problem of low accuracy in existing location methods. Specifically:
[0091] In some embodiments, determining the location of the sound source of the external sound signal relative to the first sound source of the headphones based on the first sound signal in step S2200 includes: steps S2200.1 to S2200.3.
[0092] Step S2200.1: Perform cross-correlation analysis on the left-side sound signal and the right-side sound signal to determine the first time difference between the left-side sound signal and the right-side sound signal.
[0093] In this embodiment, cross-correlation can be used to measure the similarity between the left and right sound signals in order to determine the time delay between them, i.e., the first time difference.
[0094] In some embodiments, step S2200.1 performs cross-correlation analysis on the left-side sound signal and the right-side sound signal to determine a first time difference between the left-side sound signal and the right-side sound signal, including:
[0095] Perform cross-correlation analysis on the left and right sound signals, and take the time difference between the left and right sound signals corresponding to the maximum value in the cross-correlation analysis results as the first time difference.
[0096] In this embodiment, x(n) is the left-side sound signal collected by the microphone set in the left earphone, n is the sampling point index of the left-side sound signal, n = 1, 2...N, and N is the total number of sampling points of x(n). y(m) is the right-side sound signal collected by the microphone set in the right earphone, m is the sampling point index of the right-side sound signal, m = 1, 2...M, and M is the total number of sampling points of y(m). The cross-correlation function can be expressed as shown in formula (1):
[0097]
[0098] Where c(iN) represents the value of the cross-correlation function, i.e., the result of the cross-correlation analysis, and iN is the time offset. i is the index variable for summation, used to traverse the sampling points of the left and right audio signals when calculating the cross-correlation function.
[0099] The cross-correlation function c(iN) reaches its maximum value when the left and right audio signals are perfectly aligned in time. The maximum value of the cross-correlation function indicates the strongest correlation between the two signals. The time difference between the left and right audio signals at the point where the cross-correlation function reaches its maximum value is taken as the first time difference. This first time difference is due to the difference in distance along the different paths taken by the sound source to reach the microphones set on the left and right earphones.
[0100] According to an embodiment of this application, by using the time difference between the left and right sound signals corresponding to the maximum value in the cross-correlation analysis results as the first time difference, the accuracy of determining the time difference can be improved, thereby improving the accuracy of positioning.
[0101] To avoid affecting the accuracy of determining the first time difference due to the fact that the wavelengths of the sound waves in the left and right sound signals are smaller than the distance between the two symmetrical microphones, in some embodiments, step S2200.1 performs cross-correlation analysis on the left and right sound signals to determine the first time difference between the left and right sound signals, including steps S2200.11 and S2200.12.
[0102] Step S2200.11: Perform a first bandpass filter on the left-side sound signal and the right-side sound signal within a first set frequency range to obtain a first left-side filtered sound signal and a first right-side filtered sound signal.
[0103] In this embodiment, the inventors discovered that because the wavelength of sound waves at frequencies above 1.5kHz is smaller than the distance between the two symmetrical microphones in each ear, it affects the determination of the first time difference. Therefore, the first set frequency range can be, for example, 100Hz to 1500Hz, or other frequencies below 1500Hz, and is not limited here.
[0104] By performing a first bandpass filter on the left and right sound signals within a first set frequency range, only the low-frequency components of the left and right sound signals can be retained, while the high-frequency components are filtered out, resulting in the first left-filtered sound signal and the first right-filtered sound signal.
[0105] Step S2200.12: Perform cross-correlation analysis on the first left-side filtered audio signal and the first right-side filtered audio signal to determine the first time difference between the first left-side filtered audio signal and the first right-side filtered audio signal.
[0106] In this embodiment, the first left-side filtered audio signal and the first right-side filtered audio signal are substituted into the above formula (1) to calculate the cross-correlation function, that is, x(n)n=1,2…N is the first left-side filtered audio signal, and y(m)m=1,2…M is the first right-side filtered audio signal. Then, the time difference between the first left-side filtered audio signal and the first right-side filtered audio signal when the two signals reach the maximum cross-correlation value is taken as the first time difference.
[0107] Thus, by applying a first bandpass filter to the left and right sound signals within a first set frequency range, only the low-frequency components in the left and right sound signals can be retained for cross-correlation analysis, thereby improving the accuracy of determining the first time difference between the two signals.
[0108] Step S2200.2: Based on the first time difference, determine the candidate coordinate values of the sound source of the external sound signal relative to the headphone coordinate system.
[0109] In this embodiment, the first time difference can be converted into the corresponding phase difference, and then the candidate coordinate value can be determined based on the phase difference. Alternatively, the candidate coordinate value can be determined directly based on the first time difference. No limitation is made here.
[0110] The headphone coordinate system is a coordinate system constructed based on the left and right headphones. The headphone coordinate system can be constructed with the midpoint of the line connecting the left and right headphones as the center and the line connecting the left and right headphones as the reference line (i.e., the directions of the line connecting the left and right headphones are 0° and 180° respectively), or it can be constructed with the midpoint of the line connecting the left and right headphones as the center and the line perpendicular to the line connecting the left and right headphones as the reference line (i.e., the directions perpendicular to the line connecting the left and right headphones are 0° and 180° respectively). There is no limitation here.
[0111] For example, such as Figure 5 As shown, the center of the headphone coordinate system can be defined as the midpoint of the line connecting the left and right earphones. Starting from this center, the directions extending perpendicular to the line connecting the left and right earphones are defined as 0° and 180°, respectively, thus constructing the headphone coordinate system.
[0112] Candidate coordinate values can be the angle values of the sound source of the external sound signal in the headphone coordinate system. In other words, candidate coordinate values can be the angle values between the line connecting the sound source of the external sound signal to the center of the circle and the reference line of the headphone coordinate system.
[0113] In one example, the process of determining candidate coordinate values based on the first time difference can be as follows: the preset sound source location may include, for example... Figure 5 The five preset sound source orientations shown (0°, 45°, 90°, 135°, and 180°) were pre-tested on a simulated human head to obtain the head transfer function corresponding to each preset sound source orientation. The head transfer function includes the amplitude and phase response corresponding to each preset sound source orientation. The head transfer function is shown below. Figure 6 As shown. According to Figure 6 The data shows that the time difference is largest at 90° (estimated to be approximately 0.7ms based on the diameter of a human head), followed by 45° / 135° (approximately 0.35ms), while the time difference at 0° / 180° is close to 0. If the first time difference is 0.35ms, then the candidate coordinates are 45° and 135°. If the first time difference is 0, then the candidate coordinates are 0° and 180°.
[0114] In some embodiments, step S2200.2, which determines the candidate coordinate values of the sound source of the external sound signal relative to the headphone coordinate system based on the first time difference, includes steps SB1 to SB2.
[0115] Step SB1: Determine the second time difference based on the first time difference and the sampling rate of the first sound signal.
[0116] In this embodiment, the "data point difference" obtained from the cross-correlation analysis actually refers to the time difference between the two signals at the sampling point in the cross-correlation function, rather than the actual time difference.
[0117] By dividing the first time difference by the sampling rate of the first audio signal, the second time difference is obtained. This allows the time difference between the two signals at the sampling points to be converted into the actual time difference, i.e., the second time difference. The sampling rate refers to the number of times the microphone samples per unit time, which determines the actual time interval between sampling points.
[0118] Step SB2: Determine the candidate coordinate values based on the second time difference.
[0119] Since the second time difference is more accurate than the first time difference, determining candidate coordinate values based on the second time difference can improve positioning accuracy.
[0120] Step S2200.3: Select the target coordinate value from the candidate coordinate values as the first sound source location based on the root mean square value of the amplitude of the left sound signal and / or the root mean square value of the amplitude of the right sound signal.
[0121] In this embodiment, because the auricle amplifies sound wave reflection, the sound signal energy transmitted to the ear from a sound source located in front of the wearer is greater than that from a sound source located behind the wearer. Furthermore, the sound signal energy can be characterized by the root mean square (RMS) value of the sound signal amplitude. Based on this, the target coordinate value can be selected from candidate coordinate values according to the RMS value of the sound signal amplitude on the left or right side.
[0122] In some embodiments, step S2200.3, selecting target coordinate values from candidate coordinate values based on the root mean square value of the amplitude of the left sound signal and / or the root mean square value of the amplitude of the right sound signal, includes steps SC1 to SC3.
[0123] Step SC1: Obtain the head transfer function corresponding to each preset sound source location.
[0124] In this embodiment, the head transfer function corresponding to each preset sound source location is obtained by testing on a simulated human head beforehand. The preset sound source locations can be as follows: Figure 5 The indicated sound source location can also be set to other sound source locations according to accuracy requirements; no limitation is made here.
[0125] The head transfer function includes the amplitude and phase response corresponding to each preset sound source location. Furthermore, for any preset sound source location, the head transfer function for that location includes a first head transfer function corresponding to the left earphone and a second head transfer function corresponding to the right earphone.
[0126] Step SC2: Determine the candidate head transfer function corresponding to the candidate coordinate value based on the head transfer function corresponding to each preset sound source location.
[0127] In this embodiment, from multiple head transfer functions corresponding to multiple preset sound source locations, the head transfer function whose preset sound source location matches the candidate coordinate value is determined as the candidate head transfer function corresponding to the candidate coordinate value. The candidate head transfer function includes a first head transfer function corresponding to the left earphone and a second head transfer function corresponding to the right earphone.
[0128] Step SC3: Determine the reference signal energy corresponding to the candidate coordinate value based on the candidate head transfer function.
[0129] In this embodiment, the reference signal energy is the root mean square value of the amplitude response of the candidate head transfer function.
[0130] If the candidate coordinates are 45° and 135°, the reference signal energy corresponding to each candidate coordinate value can be determined based on the root mean square value of the amplitude response of the head transfer function when the preset sound source orientation is 45° and 135°. Since the auricle amplifies sound wave reflection, the signal energy of a sound source located in front of the wearer is greater than that located behind the wearer. Therefore, the calculated reference signal energy corresponding to the candidate coordinate value of 45° is greater than that corresponding to the candidate coordinate value of 135°.
[0131] The reference signal energy includes a first reference signal energy corresponding to the left earphone and a second reference signal energy corresponding to the right earphone. For example, the reference signal energy corresponding to the candidate coordinate value 45° includes the first reference signal energy corresponding to 45° and the second reference signal energy corresponding to 45°.
[0132] Step SC4: Select the target coordinate value from the candidate coordinate values based on the root mean square value of the amplitude of the left sound signal and / or the root mean square value of the amplitude of the right sound signal, and the reference signal energy of the candidate coordinate values.
[0133] In one example, if the candidate coordinates are 45° and 135°, it's impossible to determine whether the sound source is in front of or behind the wearer. Through steps SC1-SC3, the reference signal energy corresponding to 45° and 135° is determined. Then, the root mean square (RMS) value of the left-side sound signal amplitude is compared with the first reference signal energy corresponding to 45° and 135°. If the RMS value of the left-side sound signal amplitude is greater than the first reference signal energy corresponding to 135° but less than the first reference signal energy corresponding to 45°, the target coordinate value is determined to be 45°. If the RMS value of the left-side sound signal amplitude is greater than the first reference signal energy corresponding to 45°, the target coordinate value is determined to be 45°. If the RMS value of the left-side sound signal amplitude is less than the first reference signal energy corresponding to 135°, the target coordinate value is determined to be 135°.
[0134] In another example, after determining the reference signal energy corresponding to 45° and 135°, the root mean square (RMS) values of the amplitudes of the left and right sound signals can be compared with the reference signal energies. If the RMS value of the left sound signal amplitude is greater than the first reference signal energy corresponding to 45°, and the RMS value of the right sound signal amplitude is greater than the second reference signal energy corresponding to 45°, then the target coordinate value is determined to be 45°. If the RMS value of the left sound signal amplitude is less than the first reference signal energy corresponding to 135°, and the RMS value of the right sound signal amplitude is less than the second reference signal energy corresponding to 135°, then the target coordinate value is determined to be 135°.
[0135] By comparing at least one of the root mean square values of the amplitudes of the actual acquired left-side and right-side sound signals with the reference signal energy corresponding to the candidate coordinate values, the specific sound source location can be selected from the candidate coordinate values, thereby improving the accuracy of sound source location.
[0136] The inventors discovered that the reflection of sound waves by the auricle primarily affects high-frequency sounds. To improve the accuracy of determining the target coordinate value from candidate coordinate values based on the energy difference of sound signals generated by the reflection of sound waves by the auricle, low-frequency sounds can be filtered out and high-frequency sounds retained when calculating the root mean square (RMS) values of the amplitudes of the left and / or right sound signals. Then, the target coordinate value is determined from the candidate coordinate values by calculating the RMS value of the amplitude of the high-frequency sound signals.
[0137] Based on this, step S2200.3 selects a target coordinate value from the candidate coordinate values as the first sound source location according to the root mean square value of the amplitude of the left sound signal and / or the root mean square value of the amplitude of the right sound signal, including steps S2200.31 and S2200.32.
[0138] Step S2200.31: Perform a second bandpass filter on the left and right sound signals within a second set frequency range to obtain a second left-filtered sound signal and a second right-filtered sound signal.
[0139] In this embodiment, the second set frequency range can be, for example, 3000 to 6000 Hz, or other frequency ranges greater than 2000 Hz, and is not limited here.
[0140] In one embodiment, the lower limit of the second set frequency range is greater than the upper limit of the first set frequency range.
[0141] Step S2200.32: Select the target coordinate value from the candidate coordinate values as the first sound source location based on the root mean square value of the amplitude of the second left-side filtered sound signal and / or the root mean square value of the amplitude of the second right-side filtered sound signal.
[0142] By applying a second bandpass filter to the left and right sound signals, only the high-frequency components in the left and right sound signals can be retained for signal energy calculation, thereby improving the accuracy of determining the target coordinate value from the candidate coordinate value.
[0143] After determining the location of the first sound source in step S2200, step S2300 is executed to determine the first filter corresponding to the location of the first sound source based on the location of the first sound source.
[0144] In one example, a filter corresponding to each preset sound source location can be pre-configured, with each filter having different filter parameters. Thus, after determining the first sound source location, a first filter corresponding to that first sound source location can be determined based on the first sound source location.
[0145] In another example, filter parameters can be left unconfigured initially. After determining the location of the first sound source, the filter parameters corresponding to that location can be determined based on the location of the first sound source. Then, the filter to be configured can be configured based on these filter parameters to obtain the first filter corresponding to the location of the first sound source.
[0146] In some embodiments, step S2300, determining the first filter corresponding to the first sound source location based on the first sound source location, includes:
[0147] Based on the first sound source location and the first correspondence, determine the first filter corresponding to the first sound source location.
[0148] In this embodiment, the first correspondence is the correspondence between the preset sound source location and the filter. That is, each preset sound source location corresponds to a filter, and each filter is configured with filter parameters corresponding to its preset sound source location.
[0149] In some embodiments, the step of determining the first correspondence includes: steps S3100 to S3300.
[0150] Step S3100: Obtain the head transfer function corresponding to each preset sound source location.
[0151] In this embodiment, the head transfer function is obtained by collecting the transfer function of sound sources from different preset sound source locations to the simulated ear on the simulated head.
[0152] Specifically, the head transfer function can be obtained through the same steps as step SC1 above, which will not be repeated here.
[0153] Step S3200: For any preset sound source location, determine the filter parameters corresponding to the preset sound source location based on the head transfer function corresponding to the preset sound source location, and obtain the filter parameters corresponding to each preset sound source location.
[0154] In some embodiments, the step S3200 of determining the filter parameters corresponding to the preset sound source location based on the head transfer function corresponding to the preset sound source location includes: determining the corresponding filter parameters based on the head transfer function corresponding to the preset sound source location and a preset relation.
[0155] In this embodiment, the preset relation reflects the relationship between the head transfer function and the filter parameters.
[0156] The predefined relation can be obtained through the following derivation process, such as... Figure 3 As shown, this can represent the sound propagation path of external sound signals to the ear when headphones are worn. The sound signal at the eardrum includes the sound signal remaining after passive sound insulation and the sound signal transmitted to the ear after being collected by the microphone, filtered, and then transmitted through the speaker. The sound signal remaining after passive sound insulation is x(t)*P(s), and the sound signal transmitted to the ear after being collected by the microphone, filtered, and then transmitted through the speaker is x(t)*R(s)*T(s)*G. v (s). Based on this, the sound signal at the eardrum can be represented as: e v (t)=x(t)*P(s)+x(t)*R(s)*T(s)*G v (s). Where x(t) is the external sound signal, P(s) is the transfer function of the external sound signal after passive sound insulation to the ear position, and R(s) represents the total transfer function of the external sound signal from the environment to the positions of two microphones symmetrically positioned at both ears. T(s) are the filter parameters, and G... v (s) is the transfer function of sound from the loudspeaker to the ear.
[0157] It should be noted that P(s), R(s), and G v (s) can all be obtained through actual measurement in the test environment. The test environment is a pressure field environment, and there are no special restrictions on the simulated human head used in the test.
[0158] For the sound signal heard when wearing headphones to be the same as the external sound signal heard when not wearing headphones, formula (2) must be satisfied:
[0159] X(s)P(s)+X(s)R(s)T(s)G v (s)=E v (s)=X(s)P HRTF (s) (2)
[0160] The external sound signal heard without wearing headphones can be obtained by multiplying the external sound signal by the head transfer function P. HRTF (s) is obtained, and this process is actually simulating the process of sound traveling from the sound source to each ear.
[0161] According to the above formula (2), the relationship between the filter parameters and the head transfer function can be calculated, that is, the following formula (3), which is the preset relationship.
[0162]
[0163] By substituting the head transfer function of any preset sound source location into formula (3), the filter parameters corresponding to the preset sound source location can be obtained.
[0164] Step S3300: Configure the filter to be configured in the filter array according to the filter parameters corresponding to each preset sound source location, and establish the correspondence between each preset sound source location and the configured filter in the filter array to obtain the first correspondence.
[0165] In this embodiment, a corresponding number of filters are set for each preset sound source location to form a filter array. If there are five preset sound source locations, five filters are set, and these five filters are in a state of pending configuration, that is, the filter array includes five filters to be configured. Then, for any preset sound source location, any filter to be configured in the filter array is configured according to the filter parameters calculated by the above formula (3). After configuration, the filter to be configured becomes a configured filter, and finally, the correspondence between the preset sound source location and the configured filter is established to obtain the first correspondence.
[0166] In some other embodiments, step S2300 involves determining the first filter corresponding to the first sound source location based on the first sound source location, including steps S4100 and S4200.
[0167] Step S4100: Determine the first filter parameters corresponding to the first sound source location based on the first sound source location and the second correspondence.
[0168] In this embodiment, the second correspondence is the correspondence between the preset sound source location and the filter parameters.
[0169] In some embodiments, the step of determining the second correspondence includes: steps S5100 to S5300.
[0170] Step S5100: Obtain the head transfer function corresponding to each preset sound source location.
[0171] This step is the same as step S3100 above, and will not be repeated here.
[0172] Step S5200: For any preset sound source location, determine the filter parameters corresponding to the preset sound source location based on the head transfer function corresponding to the preset sound source location, and obtain the filter parameters corresponding to each preset sound source location.
[0173] This step is the same as step S3200 above, and will not be repeated here.
[0174] Step S5300: Establish the second correspondence based on the filter parameters corresponding to each preset sound source location.
[0175] In this embodiment, each preset sound source location corresponds to a filter parameter.
[0176] Step S4200: Configure the filter to be configured according to the first filter parameters to obtain the first filter corresponding to the first sound source location.
[0177] In this embodiment, the filter to be configured is configured according to the first filter parameters corresponding to the first sound source location to obtain the first filter corresponding to the first sound source location.
[0178] Step S2400: Filter the first sound signal using the filter parameters configured in the first filter to obtain the target sound signal.
[0179] In this embodiment, the filter parameters of the filter corresponding to any preset sound source location are calculated under the condition that the sound signal heard by a person wearing headphones is the same as the external sound signal heard by a person not wearing headphones. Therefore, by filtering the first sound signal using the filter parameters configured in the first filter to obtain the target sound signal, and then playing the target sound signal through the speaker, the effect of the external sound signal heard by the human ear when wearing headphones is the same as the sound heard by the human ear when not wearing headphones can be achieved.
[0180] According to an embodiment of this application, a first sound signal is obtained by acquiring external sound signals from two microphones symmetrically positioned in each ear. Based on the first sound signal, the location of the sound source of the external sound signal relative to the headphones is determined. Based on the location of the first sound source, a first filter corresponding to the location of the first sound source is determined. The first sound signal is then filtered using the filter parameters configured for the first filter to obtain a target sound signal. This allows the user to hear the same external sound signals when wearing headphones as when not wearing headphones, improving the user's ability to perceive external sounds while wearing headphones, ensuring that the user can react to external sounds promptly and accurately, thereby enhancing user experience and safety.
[0181] In addition, by determining candidate coordinate values based on the first time difference between the left and right sound signals, and then selecting the target coordinate value from the candidate coordinate values based on the root mean square value of the amplitude of the left sound signal and / or the root mean square value of the amplitude of the right sound signal, the accuracy of sound source localization can be improved, further enhancing the user's ability to perceive external sounds when wearing headphones.
[0182] <Device Embodiment>
[0183] The figure is a schematic diagram of the structure of a sound signal processing apparatus according to an embodiment of the present application.
[0184] like Figure 7 As shown, the audio signal processing apparatus 700 includes a memory 710 and a processor 720. The memory 710 is used to store executable instructions; the processor 720 is used to operate according to the control of the instructions to execute the method described in any of the above method embodiments.
[0185] How instructions control the processor to perform operations is well known in the art, and therefore will not be described in detail here.
[0186] Figure 8 This is a schematic diagram of the structure of an earphone according to an embodiment of this application.
[0187] like Figure 8 As shown, the earphone 800 includes a left earphone 810, a right earphone 820, and a sound signal processing device 830;
[0188] The left earphone 810 and the right earphone 820 are each equipped with at least one microphone. External sound signals are collected through one microphone in the left earphone 810 and one microphone in the right earphone 820 to obtain a first sound signal, and the first sound signal is sent to the sound signal processing device 830.
[0189] The sound signal processing device 830 can be as follows: Figure 7 The device shown, or, as Figure 1 The apparatus shown.
[0190] In some embodiments, the headphones 800 are in-ear headphones.
[0191] In other embodiments, the earphone 800 is a headset.
[0192] This invention can be a system, method, and / or computer program product. A computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for causing a processor to implement various aspects of the invention.
[0193] Computer-readable storage media can be tangible devices capable of holding and storing instructions for use by an instruction execution device. Computer-readable storage media can be, for example—but not limited to—electrical storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital multifunction disc (DVD), memory sticks, floppy disks, mechanical encoding devices, such as punch cards or recessed protrusions storing instructions thereon, and any suitable combination thereof. The computer-readable storage media used herein are not to be construed as transient signals themselves, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses through fiber optic cables), or electrical signals transmitted through wires.
[0194] The computer-readable program instructions described herein can be downloaded from computer-readable storage media to various computing / processing devices, or downloaded via a network, such as the Internet, local area network, wide area network, and / or wireless network, to an external computer or external storage device. The network may include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to the computer-readable storage media in the respective computing / processing device.
[0195] The computer program instructions used to perform the operations of this invention may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, electronic circuitry, such as programmable logic circuitry, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), is personalized by utilizing state information from the computer-readable program instructions. This electronic circuitry can execute the computer-readable program instructions to implement various aspects of the invention.
[0196] Various aspects of the present invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0197] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processor of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner; thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0198] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.
[0199] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction containing one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions. It will be known to those skilled in the art that implementation in hardware, implementation in software, and implementation using a combination of software and hardware are equivalent.
[0200] The various embodiments of the present invention have been described above. These descriptions are exemplary and not exhaustive, and are not limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles, practical application, or technical improvements to the embodiments in the market, or to enable others skilled in the art to understand the embodiments disclosed herein. The scope of the invention is defined by the appended claims.
Claims
1. A sound signal processing method, characterized in that, The sound signal processing method is applied to an earphone, the earphone including a left earphone and a right earphone, the left earphone and the right earphone each being provided with at least one microphone, the method including: The first sound signal is obtained by acquiring external sound signals from two microphones symmetrically positioned in each ear; Based on the first sound signal, determine the location of the sound source of the external sound signal relative to the first sound source of the headphones; Based on the location of the first sound source, a first filter corresponding to the location of the first sound source is determined; the first sound signal is filtered using the filter parameters configured for the first filter to obtain the target sound signal. The first sound signal includes a left-side sound signal collected by a microphone located at the left earphone position and a right-side sound signal collected by a microphone located at the right earphone position. Determining the location of the external sound source relative to the earphones based on the first sound signal includes: Perform cross-correlation analysis on the left and right sound signals to determine the first time difference between them; Based on the first time difference, candidate coordinate values of the sound source of the external sound signal relative to the headphone coordinate system are determined; wherein, the headphone coordinate system is a coordinate system constructed based on the left and right headphones; Based on the root mean square value of the amplitude of the left-side sound signal and / or the root mean square value of the amplitude of the right-side sound signal, a target coordinate value is selected from the candidate coordinate values as the first sound source location.
2. The method according to claim 1, characterized in that, The step of performing cross-correlation analysis on the left and right sound signals to determine the first time difference between them includes: A first bandpass filter is applied to the left-side sound signal and the right-side sound signal within a first set frequency range to obtain a first left-side filtered sound signal and a first right-side filtered sound signal. Perform cross-correlation analysis on the first left-side filtered audio signal and the first right-side filtered audio signal to determine the first time difference between the first left-side filtered audio signal and the first right-side filtered audio signal; The step of selecting a target coordinate value from the candidate coordinate values as the first sound source location based on the root mean square value of the amplitude of the left-side sound signal and / or the root mean square value of the amplitude of the right-side sound signal includes: A second bandpass filter is applied to the left-side sound signal and the right-side sound signal within a second set frequency range to obtain a second left-side filtered sound signal and a second right-side filtered sound signal. Based on the root mean square value of the amplitude of the second left-side filtered sound signal and / or the root mean square value of the amplitude of the second right-side filtered sound signal, a target coordinate value is selected from the candidate coordinate values as the first sound source location; wherein, the lower limit of the second set frequency range is greater than the upper limit of the first set frequency range.
3. The method according to claim 1, characterized in that, The step of performing cross-correlation analysis on the left and right sound signals to determine the first time difference between them includes: Perform cross-correlation analysis on the left and right sound signals, and take the time difference between the left and right sound signals corresponding to the maximum value in the cross-correlation analysis results as the first time difference.
4. The method according to claim 1, characterized in that, Determining the candidate coordinate values of the sound source of the external sound signal relative to the headphone coordinate system based on the first time difference includes: The second time difference is determined based on the first time difference and the sampling rate of the first sound signal; The candidate coordinate values are determined based on the second time difference.
5. The method according to claim 1, characterized in that, The step of selecting target coordinate values from candidate coordinate values based on the root mean square value of the amplitude of the left-side sound signal and / or the root mean square value of the amplitude of the right-side sound signal includes: Obtain the head transfer function corresponding to each preset sound source location; Based on the head transfer function corresponding to each preset sound source location, determine the candidate head transfer function corresponding to the candidate coordinate value; The reference signal energy corresponding to the candidate coordinate value is determined based on the candidate head transfer function; The target coordinate value is selected from the candidate coordinate values based on the root mean square value of the amplitude of the left sound signal and / or the root mean square value of the amplitude of the right sound signal, and the reference signal energy of the candidate coordinate values.
6. The method according to claim 1, characterized in that, The step of determining the first filter corresponding to the first sound source location based on the first sound source location includes: Based on the first sound source location and the first correspondence, a first filter corresponding to the first sound source location is determined; wherein, the first correspondence is a preset correspondence between the sound source location and the filter.
7. The method according to claim 6, characterized in that, The steps for determining the first correspondence include: Obtain the head transfer function corresponding to each preset sound source location; For any preset sound source location, the filter parameters corresponding to the preset sound source location are determined according to the head transfer function corresponding to the preset sound source location, thus obtaining the filter parameters corresponding to each preset sound source location. Configure the filter to be configured in the filter array according to the filter parameters corresponding to each preset sound source location, and establish the correspondence between each preset sound source location and the configured filter in the filter array to obtain the first correspondence.
8. The method according to claim 1, characterized in that, The step of determining the first filter corresponding to the first sound source location based on the first sound source location includes: Based on the first sound source location and the second correspondence, the first filter parameters corresponding to the first sound source location are determined; wherein, the second correspondence is a preset correspondence between the sound source location and the filter parameters; Configure the filter to be configured according to the first filter parameters to obtain the first filter corresponding to the first sound source location.
9. The method according to claim 8, characterized in that, The steps for determining the second correspondence include: Obtain the head transfer function corresponding to each preset sound source location; For any preset sound source location, the filter parameters corresponding to the preset sound source location are determined according to the head transfer function corresponding to the preset sound source location, thus obtaining the filter parameters corresponding to each preset sound source location. The second correspondence is established based on the filter parameters corresponding to each preset sound source location.
10. The method according to claim 7 or 9, characterized in that, The step of determining the filter parameters corresponding to the preset sound source location based on the head transfer function corresponding to the preset sound source location includes: Based on the head transfer function corresponding to the preset sound source location and the preset relational expression, the corresponding filter parameters are determined; wherein, the preset relational expression reflects the relationship between the head transfer function and the filter parameters.
11. A sound signal processing device, characterized in that, It includes a memory and a processor, the memory being used to store executable instructions; the processor being used to operate under the control of the instructions to perform the method as described in any one of claims 1 to 10.
12. An earphone, characterized in that, It includes a left earphone and a right earphone, and the sound signal processing device as described in claim 11; The left and right earpieces are each equipped with at least one microphone. External sound signals are collected through one microphone in each earpiece to obtain a first sound signal, which is then sent to the sound signal processing device.
Citation Information
Patent Citations
Filtering method and device in transparent mode, earphone and readable storage medium
CN114363770A