Method for processing a sound signal and device using the same

CN122802839APending Publication Date: 2026-09-22DEEP HEARING CO
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610321978.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2025-05-29
Filing Date
2026-03-17
Publication Date
2026-09-22

AI Technical Summary

Benefits of technology

[0027]根据本发明实施例的方法和装置,仅利用配置于相隔开的位置的两个声传感器即可设定三维形状的声音处理区域,因此,能够以简单且小型化的结构有效地设定声音信号的三维接收区域。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122802839A_ABST
    Figure CN122802839A_ABST
Patent Text Reader

Abstract

A sound signal processing method according to an embodiment of the present application includes the steps of: receiving training sound signals from the same sound source through first and second sound sensors respectively disposed at spaced apart positions; training an artificial neural network model for setting a three-dimensional sound processing region based on the received training sound signals; and processing a received processing target sound signal using the trained artificial neural network model and based on the three-dimensional sound processing region.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for processing sound signals and an apparatus thereof, and more specifically, to a method for processing sound signals and an apparatus thereof, wherein an artificial neural network model for setting a three-dimensional shape of a sound processing region is trained based on training sound signals received by two sound sensors disposed at spaced apart, and the trained artificial neural network model is used to process the sound signal of the object to be processed. Background Technology

[0002] The technology of using multiple acoustic sensors to selectively receive sound signals that are transmitted within a specific distance range or in a specific direction is applicable to a variety of applications.

[0003] For example, selective audio signal reception technology can be applied to situations where, in conference systems or remote video conferencing equipment, communication quality can be improved by clearly receiving only the voices of specific speakers.

[0004] As another application example, the technology of selectively receiving sound signals can be applied to situations such as: in hearing aids, reducing hearing burden by emphasizing only sounds from a specific direction, or in small wearable devices such as wireless headphones or smart glasses, improving the user experience by suppressing ambient noise and selectively receiving only the desired sounds.

[0005] Recently, selective sound reception technology has been applied to various fields, such as voice recognition-based artificial intelligence assistants, external sound recognition for autonomous vehicles, and monitoring systems. Summary of the Invention

[0006] The problem that the invention aims to solve

[0007] The technical problem to be solved by the present invention is to provide a method for processing sound signals and an apparatus thereof, which can train an artificial neural network model for setting a three-dimensional shape of a sound processing region based on training sound signals received by two sound sensors respectively arranged at spaced apart positions, and use the trained artificial neural network model to process the sound signal of the processing object.

[0008] Solution for solving the problem

[0009] According to an embodiment of the present invention, a method for processing sound signals using two sound sensors may include the following steps: receiving training sound signals from the same sound source using a first sound sensor and a second sound sensor disposed at spaced apart; training an artificial neural network model based on the received training sound signals, wherein the artificial neural network model defines a bounded three-dimensional sound processing region according to its relative positional relationship with the first sound sensor and the second sound sensor; and processing the received sound signals of the target object using the trained artificial neural network model and based on the three-dimensional sound processing region.

[0010] According to an embodiment, the aforementioned three-dimensional sound processing area can be a region that defines the range of sound to be received, or a region that defines the range of sound not to be received.

[0011] According to an embodiment, the aforementioned three-dimensional sound processing region can be defined as at least a portion of a conical region centered on the axis passing through the first sound sensor and the second sound sensor.

[0012] According to the embodiment, the sound processing area of ​​the above-mentioned three-dimensional shape can be determined based on a preset angle range formed with the above-mentioned axis as a reference and a preset distance range from the above-mentioned first sound sensor and the above-mentioned second sound sensor.

[0013] According to an embodiment, the range of distances between the first acoustic sensor and the second acoustic sensor can be determined by the distance on the shaft from the midpoint between the first acoustic sensor and the second acoustic sensor.

[0014] According to the embodiments, the above-mentioned artificial neural network model can be an artificial neural network model with a U-NET structure or an artificial neural network model with a recurrent neural network (RNN) structure.

[0015] According to an embodiment, in the step of training the artificial neural network model, the artificial neural network model can be trained based on the time difference between the first training sound signal received by the first sound sensor and the second training sound signal received by the second sound sensor, as well as the signal strength difference between the first training sound signal and the second training sound signal.

[0016] According to an embodiment, the steps of training the artificial neural network model may include the following steps: performing a Short Time Fourier Transform (STFT) on the first training sound signal and the second training sound signal respectively; and using the first training sound signal and the second training sound signal after STFT transformation, calculating the time difference between the first training sound signal and the second training sound signal, as well as the signal strength difference between the first training sound signal and the second training sound signal.

[0017] According to the embodiment, in the step of training the above-mentioned artificial neural network model, the time difference between the first training sound signal and the second training sound signal, as well as the signal strength difference between the first training sound signal and the second training sound signal, can be used as input features of the above-mentioned artificial neural network model to train the above-mentioned artificial neural network model.

[0018] According to an embodiment, the time difference between the first training sound signal and the second training sound signal can be calculated based on the phase difference between the transformed first training sound signal and the transformed second training sound signal.

[0019] According to an embodiment, the time difference between the first training sound signal and the second training sound signal can be calculated based on the difference between the time point in the time domain where the magnitude of the first training sound signal has a maximum value and the time point in the time domain where the magnitude of the second training sound signal has a maximum value.

[0020] According to an embodiment, the signal strength difference between the first training sound signal and the second training sound signal can be calculated based on the magnitude difference between the transformed first training sound signal and the transformed second training sound signal, or the magnitude ratio between the transformed first training sound signal and the transformed second training sound signal.

[0021] According to an embodiment, the signal strength difference between the first training sound signal and the second training sound signal can be calculated based on the signal energy difference between the first training sound signal in the time domain and the second training sound signal in the time domain.

[0022] According to an embodiment, the step of processing the received sound signal of the processing object based on the sound processing area of ​​the above-mentioned three-dimensional shape may include the following steps: removing the sound signal of the processing object that exceeds the sound processing area of ​​the above-mentioned three-dimensional shape, or reducing the size of the sound signal of the processing object that exceeds the sound processing area of ​​the above-mentioned three-dimensional shape.

[0023] According to an embodiment, the step of processing the received object sound signal based on the sound processing region of the three-dimensional shape may include the following steps: removing the object sound signal within the sound processing region of the three-dimensional shape, or reducing the size of the object sound signal within the sound processing region of the three-dimensional shape.

[0024] According to an embodiment, in the step of training the above-mentioned artificial neural network model, the artificial neural network model can be trained based on the phase difference matrix, which includes the time difference information between the first training sound signal received by the first sound sensor and the second training sound signal received by the second sound sensor, and the signal strength difference between the first training sound signal and the second training sound signal.

[0025] A sound signal processing apparatus according to an embodiment of the present invention, which processes sound signals using two sound sensors, may include: a first sound sensor; a second sound sensor disposed at a position spaced apart from the first sound sensor; and a processor that trains an artificial neural network model based on training sound signals received from the same sound source through the first sound sensor and the second sound sensor, respectively. The artificial neural network model sets a three-dimensional sound processing region according to the relative positional relationship with the first sound sensor and the second sound sensor, and processes the received sound signal of the target object using the trained artificial neural network model and based on the three-dimensional sound processing region.

[0026] Invention Effects

[0027] According to the method and apparatus of the present invention, a three-dimensional sound processing area can be set using only two sound sensors arranged at spaced-away positions. Therefore, a three-dimensional sound signal receiving area can be effectively set with a simple and miniaturized structure.

[0028] The methods and apparatus according to embodiments of the present invention can also be applied to small devices or wearable devices, thereby minimizing the size and power consumption of the device. Attached Figure Description

[0029] To provide a fuller understanding of the accompanying drawings referenced in the detailed description of the invention, a brief description of each drawing is provided.

[0030] Figure 1 This is a block diagram of a sound signal processing apparatus according to an embodiment of the present invention.

[0031] Figure 2 This is a flowchart of a sound signal processing method according to an embodiment of the present invention.

[0032] Figure 3 It is used for explanation Figure 2A diagram illustrating the process of determining the three-dimensional shape of the sound processing region in a sound signal processing method.

[0033] Figure 4 It is based on Figure 3 An example of a three-dimensional shape sound processing area defined by the process.

[0034] Figure 5 It shows the basis Figure 2 The diagram shows a detailed process of one embodiment of the steps involved in training an artificial neural network model.

[0035] Figure 6 It is used to explain the relevant calculations. Figure 5 A diagram illustrating the time difference between the first and second training sound signals used in the process of training an artificial neural network model.

[0036] Figure 7 It is used to explain the relevant calculations. Figure 5 A diagram illustrating the signal strength difference between the first and second training sound signals used in the process of training an artificial neural network model.

[0037] Figure 8 This is a graph illustrating the sound signal processing effect of the sound signal processing apparatus according to an embodiment of the present invention.

[0038] Figure 9 yes Figure 1 The example shown is an audio signal processing device implemented in a form combined with a wearable device.

[0039] Figure 10 yes Figure 1 The example shown is an audio signal processing device implemented in the form of a wireless microphone.

[0040] Figure 11 yes Figure 1 The example shown is an audio signal processing device implemented in combination with wireless headphones.

[0041] Explanation of reference numerals in the attached figures:

[0042] 100: Sound signal processing device

[0043] 110, 120: Sound sensors

[0044] 130: Processor

[0045] 140: Memory. Detailed Implementation

[0046] The technical concept of this invention can be modified in many ways and can have many embodiments. Specific embodiments are illustrated in the accompanying drawings, and these specific embodiments are described in detail. However, this does not mean that the technical concept of this invention is limited to specific implementations, but should be understood as including all modifications, equivalents, or substitutions included in the scope of the technical concept of this invention.

[0047] When describing the technical concept of the present invention, detailed descriptions of relevant prior art will be omitted if they are deemed unnecessary to obscure the spirit of the invention. Furthermore, the numbers used in this specification (e.g., first, second, etc.) are merely identification numbers used to distinguish one component from another.

[0048] Furthermore, when this specification refers to a component being “connected” or “linked” to another component, the aforementioned component may also be directly connected or directly linked to the aforementioned other component. Unless otherwise stated, it should be understood that the connection or link may also be achieved by intervening other components in the middle.

[0049] Furthermore, the terms "~unit", "~device", "~sub-unit", and "~module" used in this specification refer to a unit used to process at least one function or action, which can be implemented by hardware, software, or a combination of hardware and software such as processor, microprocessor, micro controller, central processing unit (CPU), graphics processing unit (GPU), accelerated processor unit (APU), drive signal processor (DSP), application specific integrated circuit (ASIC), and field programmable gate array (FPGA), or can also be implemented in combination with memory for storing data required to process at least one function or action.

[0050] Furthermore, it should be clarified that the division of structural units in this specification is based solely on the primary functions each structural unit is responsible for. That is, it can be configured such that two or more structural units described below can be combined into one structural unit, or a single structural unit can be divided into two or more units based on more specific functions. Moreover, in addition to their own primary functions, each structural unit described below can also perform some or all of the functions of other structural units. Of course, some of the primary functions of each structural unit may also be performed by other structural units.

[0051] Figure 1 This is a block diagram of a sound signal processing apparatus according to an embodiment of the present invention.

[0052] Reference Figure 1 The sound signal processing device 100 may include a first sound sensor 110, a second sound sensor 120, a processor 130, and a memory 140.

[0053] The first sound sensor 110 and the second sound sensor 120 are respectively positioned at a distance from each other, so that they can receive sound signals transmitted from the same sound source.

[0054] According to the embodiments, the sound signal processing device of the present invention only requires two sound sensors, and the received sound signals can be processed through a simple structure without additional sound sensors.

[0055] The first sound sensor 110 and the second sound sensor 120 can respectively receive training sound signals or processing object sound signals transmitted from the same sound source.

[0056] The first acoustic sensor 110 and the second acoustic sensor 120 are positioned at a distance from each other, resulting in a time difference and a signal strength difference in the sound signals received by the first acoustic sensor 110 and the second acoustic sensor 120, respectively. The resulting time difference and signal strength difference in the sound signals can be used in the training process of an artificial neural network model for subsequent processing of the sound signals.

[0057] According to the embodiment, the time difference between the sound signals received by the first sound sensor 110 and the second sound sensor 120 respectively can be called the interaural time difference (ITD), but is not limited thereto.

[0058] According to the embodiment, the signal strength difference of the sound signals received by the first sound sensor 110 and the second sound sensor 120 respectively can be called the interaural level difference (ILD), but is not limited to this.

[0059] According to an embodiment, the sound signal processing apparatus of this invention may also use a phase difference matrix including the aforementioned time difference information (rather than the time difference between the sound signals received by the first sound sensor 110 and the second sound sensor 120, respectively) for training an artificial neural network model to process the sound signal. The details of the phase difference matrix will be explained later.

[0060] According to the embodiments, the first acoustic sensor 110 and the second acoustic sensor 120 can be implemented in the form of various types of sensors that can receive sound signals.

[0061] According to the embodiments, the first acoustic sensor 110 and the second acoustic sensor 120 can be implemented as microphones of various forms, such as dynamic microphones that convert signals into electrical signals, electret condenser microphones (ECM), micro-electro-mechanical systems (MEMS) microphones, etc.

[0062] The processor 130 can train an artificial neural network model using training sound signals received by the first sound sensor 110 and the second sound sensor 120. The artificial neural network model defines a bounded three-dimensional sound processing region based on its relative positional relationship with the first sound sensor 110 and the second sound sensor 120.

[0063] According to the embodiments, the bounded region can also be expressed in different ways, such as a limited region, a finite spatial region, or a closed region.

[0064] According to the embodiment, the relative positional relationship with the first acoustic sensor 110 and the second acoustic sensor 120 can refer to, for example, the distance from the base point or baseline set by the first acoustic sensor 110 and the second acoustic sensor 120, and the positional relationship relative to the base point or baseline set by the first acoustic sensor 110 and the second acoustic sensor 120.

[0065] The processor 130 can process the sound signal of the object to be processed using a trained artificial neural network model and based on a sound processing region of three-dimensional shape.

[0066] According to the embodiments, artificial neural network models can be implemented in various forms, such as artificial neural network models with U-NET structures or artificial neural network models with recurrent neural network (RNN) structures.

[0067] According to an embodiment, the three-dimensional sound processing area can be a region that defines the range of sound to be received or a region that defines the range of sound not to be received.

[0068] The following content will refer to Figures 2 to 7 The detailed operation of processor 130 is explained.

[0069] The memory 140 is connected to the processor 130 and can store data such as training sound signals and processing target sound signals received by the first sound sensor 110 and the second sound sensor 120 respectively, various data required for processing the sound signals, data generated according to the processing process or processing result of the sound signals, and trained artificial neural network models.

[0070] Figure 2 This is a flowchart of a sound signal processing method according to an embodiment of the present invention. Figure 3 It is used for explanation Figure 2 A diagram illustrating the process of determining the three-dimensional shape of the sound processing region in a sound signal processing method. Figure 4 It is based on Figure 3 An example of a three-dimensional shape sound processing area defined by the process. Figure 5 It shows the basis Figure 2 The diagram shows a detailed process of one embodiment of the steps involved in training an artificial neural network model. Figure 6 It is used to explain the relevant calculations. Figure 5 A diagram illustrating the time difference between the first and second training sound signals used in the process of training an artificial neural network model. Figure 7 It is used to explain the relevant calculations. Figure 5 A diagram illustrating the signal strength difference between the first and second training sound signals used in the process of training an artificial neural network model.

[0071] Reference Figure 2 According to an embodiment of the present invention, the sound signal processing apparatus 100 can receive training sound signals from the same sound source through a first sound sensor 110 and a second sound sensor 120 arranged at separate positions (S10).

[0072] According to an embodiment, the first acoustic sensor 110 can receive a first training sound signal from the same sound source, and the second acoustic sensor 120 can receive a second training sound signal from the same sound source. The first acoustic sensor 110 and the second acoustic sensor 120 are configured separately from each other. Therefore, there may be a time difference (time difference or phase difference between the receiving time points) and a signal strength difference (signal magnitude difference) between the first training sound signal and the second training sound signal.

[0073] The sound signal processing device 100 can train an artificial neural network model (S20) based on the training sound signal received in step S10. The artificial neural network model sets a bounded three-dimensional sound processing area according to the relative positional relationship with the first sound sensor 110 and the second sound sensor 120.

[0074] According to the embodiments, artificial neural network models can be implemented in various forms, such as artificial neural network models with U-NET structure or artificial neural network models with recurrent neural network (RNN) structure.

[0075] According to an embodiment, a bounded three-dimensional sound processing region can be defined as at least a portion of a conical region.

[0076] Combined with reference Figure 3 A bounded three-dimensional sound processing region A-RC can be defined as at least a portion of a conical region centered on the axis AX that passes through the first sound sensor 110 and the second sound sensor 120 together.

[0077] According to an embodiment, at least a portion of the conical shape can be determined based on a preset angle range (e.g., 0~θ) formed with respect to the axis AX passing through the first acoustic sensor 110 and the second acoustic sensor 120, and a preset distance range (e.g., R~L) from the first acoustic sensor 110 and the second acoustic sensor 120. In this case, the preset distance range (e.g., R~L) can be determined based on the distance on the axis AX from the midpoint P-ct of the first acoustic sensor 110 and the second acoustic sensor 120.

[0078] In this case, the axis AX passing through the first acoustic sensor 110 and the second acoustic sensor 120 becomes the central axis of the cone shape, and the midpoint P-ct of the first acoustic sensor 110 and the second acoustic sensor 120 can be equivalent to the vertex of the cone shape. Furthermore, the upper limit θ of the preset angle range corresponds to the opening angle of the cone shape, and the base of the cone shape can be formed at the upper limit of the distance range (e.g., L). In the cone shape formed as described above, only the region within the preset distance range (e.g., R~L) can be defined as the bounded three-dimensional sound processing region A-RC. In this case, in the cone shape with height L, the shape region other than the cone shape with height R can be defined as the bounded three-dimensional sound processing region A-RC.

[0079] According to an embodiment, the lower limit of the preset angle range can have any angle value other than 0 degrees. In this case, the sound processing area A-RC can be a shape where a large cone shape (e.g., a cone shape with height L) lacks a region inside that corresponds to a small cone shape (e.g., a cone shape with height R). At this time, the vertices of the large cone shape and the small cone shape can be the same.

[0080] Reference Figure 4 At least a portion of the conical shape can be defined as follows: Figure 4 The shape of the sound processing area A-RC shown in the figure.

[0081] According to an embodiment, Figure 2 The steps for training an artificial neural network model can be as follows: Figure 5 The detailed steps are as follows.

[0082] Reference Figure 5 The processor 130 of the sound signal processing device 100 can perform short-time fourier transform (STFT) on the first training sound signal S1(t) received by the first sound sensor 110 and the second training sound signal S2(t) received by the second sound sensor 120 (S201, S202).

[0083] In step S201, the transformed first training sound signal X1(w) can be output from the result of performing a short-time Fourier transform on the first training sound signal S1(t).

[0084] In step S202, the transformed second training sound signal X2(w) can be output from the result of the short-time Fourier transform of the second training sound signal S2(t).

[0085] The processor 130 of the sound signal processing device 100 can use the transformed first training sound signal X1(w) and the transformed second training sound signal X2(w) to calculate the time difference ITD between the first training sound signal S1(t) and the second sound sensor 120 (S203).

[0086] According to an embodiment, the time difference ITD between the first training sound signal S1(t) and the second sound sensor 120 can be calculated based on the phase difference between the transformed first training sound signal X1(w) and the transformed second training sound signal X2(w). In this case, the processor 130 can calculate the time difference ITD based on the following mathematical formula 1.

[0087] Mathematical formula 1:

[0088] ITD=ang(X1 * (w)X2(w))

[0089] According to another embodiment, the time difference ITD between the first training sound signal S1(t) and the second training sound signal S2(t) can be calculated based on the difference between the time point when the magnitude of the first training sound signal S1(t) reaches its maximum and the time point when the magnitude of the second training sound signal S2(t) reaches its maximum. In this case, the processor 130 can calculate the time difference ITD based on the following mathematical formula 2.

[0090] Mathematical formula 2:

[0091] ITD=argmax(S1(t))-argmax(S2(t))

[0092] Combined with reference Figure 6 This illustrates the process of using the time difference ITD between the first training sound signal S1(t) and the second training sound signal S2(t) to obtain information about the angle θ formed by the axis AX of the sound source relative to the axis passing through the first sound sensor 110 and the second sound sensor 120.

[0093] The first acoustic sensor 110 and the second acoustic sensor 120 are arranged apart by a predetermined distance d, and there is an axis AX that passes through the first acoustic sensor 110 and the second acoustic sensor 120 together.

[0094] The sound signal generated from the sound source can be transmitted to the first sound sensor 110 at a first distance r1 and to the second sound sensor 120 at a second distance r2.

[0095] When the angle formed by the line segment connecting the sound source and the first sound sensor 110 with the axis AX is set as θ, the arrival time difference of the sound signal between the first sound sensor 110 and the second sound sensor 120 is ( t) can be approximately represented by the following mathematical expression 3, and θ can be calculated using mathematical expression 3.

[0096] Mathematical formula 3:

[0097] t≒dcosθ / v

[0098] (The above v is the speed of sound)

[0099] Furthermore, the phase difference of the sound signal between the first sound sensor 110 and the second sound sensor 120 can be represented by the following mathematical formula 4 ( ITD).

[0100] Mathematical formula 4:

[0101]

[0102] (where f is the frequency of the sound signal)

[0103] According to another embodiment, the processor 130 can train an artificial neural network model using a phase difference matrix (PDM) that includes the time difference (ITD) information between a first training sound signal S1(t) and a second training sound signal S2(t). The phase difference matrix (PDM) can express the phase difference between the first training sound signal S1(t) and the second training sound signal S2(t) in matrix form in the time-frequency domain.

[0104] First, the processor 130 can perform Fourier transforms (e.g., short-time Fourier transform (STFT)) on the first training audio signal S1(t) and the second training audio signal S2(t) respectively to generate the time-frequency domain spectrum. , (The above q is the time frame index, and the above w is the frequency index).

[0105] Subsequently, processor 130 can utilize the spectrum in the time-frequency domain. , And based on the following mathematical formula 5, a phase difference matrix including phase difference information at each time-frequency position (tq, wt) is obtained.

[0106] Mathematical formula 5:

[0107]

[0108] (The above PDM represents the phase difference matrix, and Im(·) represents the imaginary part.)

[0109] According to the relationship in the following mathematical formula 6, the phase difference matrix can include time difference (ITD) information.

[0110] Mathematical formula 6:

[0111]

[0112] (The above nw is a natural number)

[0113] According to an embodiment, the processor 130 inputs the phase difference matrix into the artificial neural network model in the form of an image and trains it, thereby greatly improving the confusion caused by the phase repeating every 2π cycles.

[0114] The processor 130 of the sound signal processing device 100 can use the transformed first training sound signal X1(w) and the transformed second training sound signal X2(w) to calculate the signal strength difference ILD between the first training sound signal S1(t) and the second sound sensor 120 (S204).

[0115] According to an embodiment, the signal strength difference ILD between the first training sound signal S1(t) and the second training sound signal S2(t) is the difference in magnitude between the transformed first training sound signal X1(w) and the transformed second training sound signal X2(w). In this case, the processor 130 can calculate the signal strength difference ILD based on the following mathematical formula 7.

[0116] Mathematical expression 7:

[0117]

[0118] According to another embodiment, the signal strength difference ILD between the first training audio signal S1(t) and the second training audio signal S2(t) can be calculated based on the magnitude ratio of the transformed first training audio signal X1(w) to the transformed second training audio signal X2(w). In this case, the processor 130 can calculate the signal strength difference ILD based on the following mathematical formula 8.

[0119] Mathematical formula 8:

[0120]

[0121] According to another embodiment, the signal strength difference ILD between the first training sound signal S1(t) and the second training sound signal S2(t) can be calculated based on the signal energy difference between the first training sound signal S1(t) in the time domain and the second training sound signal S2(t) in the time domain. In this case, the processor 130 can calculate the signal strength difference ILD based on the following mathematical formula 9.

[0122] Mathematical formula 9:

[0123] ILD=

[0124] Combined with reference Figure 7 This illustrates the process of obtaining information related to the distance between the sound source and the first sound sensor 110 and the second sound sensor 120 by using the signal intensity difference ILD between the first training sound signal S1(t) and the second training sound signal S2(t).

[0125] The first acoustic sensor 110 and the second acoustic sensor 120 are arranged apart by a predetermined distance d, and there is an axis AX that passes through the first acoustic sensor 110 and the second acoustic sensor 120 together.

[0126] The signal strength of the sound signal generated at the sound source is inversely proportional to the square of the distance r from the sound source.

[0127] That is, the signal strength I1 of the sound signal obtained by the first sound sensor 110 and the signal strength I2 of the sound signal obtained by the second sound sensor 120 have the following mathematical formula 10 relationship.

[0128] Mathematical formula 10:

[0129]

[0130] At this time, the ratio of the signal strength I1 of the sound signal acquired by the first sound sensor 110 to the signal strength I2 of the sound signal acquired by the second sound sensor 120 is ( ILD can be calculated according to the following mathematical formula 11.

[0131] Mathematical formula 11:

[0132]

[0133] The processor 130 of the sound signal processing device 100 can use the time difference ITD between the first training sound signal S1(t) and the second training sound signal S2(t) and the signal strength difference ILD between the first training sound signal S1(t) and the second training sound signal S2(t) as input features of the artificial neural network model, and train the artificial neural network model (S205).

[0134] According to an embodiment, the processor 130 of the sound signal processing device 100 can concatenate the input features corresponding to the time difference ITD of the first training sound signal S1(t) and the second training sound signal S2(t) and the signal strength difference ILD of the first training sound signal S1(t) and the second training sound signal S2(t) into a vector, and use the concatenated vector to train an artificial neural network model.

[0135] According to an embodiment, the processor 130 of the sound signal processing apparatus 100 can train an artificial neural network model by inputting a phase difference matrix (rather than the time difference ITD between the first training sound signal S1(t) and the second training sound signal S2(t)) in the form of an image. In this case, the artificial neural network model can be trained by inputting the phase difference matrix and the signal intensity difference ILD.

[0136] When the output signal Xout(w) of the artificial neural network model is output, the processor 130 of the sound signal processing device 100 can output the time-domain output signal Sout(t) by performing an inverse short-time fourier transform (ISTFT) on the output signal Xout(w).

[0137] When the sound processing region, which has a three-dimensional shape, defines the range of the sound to be received, the processor 130 of the sound signal processing device 100 directly outputs or amplifies the sound signal corresponding to the sound source as an output signal Sout(t) when the sound source is within the sound processing region (e.g., A-RC). When the sound source is outside the sound processing region A-RC, the output signal Sout(t) can be attenuated and then output. According to an embodiment, the processor 130 of the sound signal processing device 100 can train an artificial neural network model by removing the processed sound signal that exceeds the three-dimensional sound processing region (e.g., A-RC) and outputting an output signal Sout(t) corresponding to 0. According to another embodiment, the processor 130 of the sound signal processing device 100 can train an artificial neural network model by outputting an output signal Sout(t) that reduces the magnitude of the processed sound signal that exceeds the three-dimensional sound processing region (e.g., A-RC). For example, the processor 130 can train an artificial neural network in such a way that the magnitude of the processed object sound signal beyond a three-dimensional sound processing region (e.g., A-RC) decreases in a manner inversely proportional to the distance beyond the sound processing region (e.g., A-RC) or inversely proportional to the square of the distance beyond the sound processing region (e.g., A-RC).

[0138] When the sound processing area of ​​a three-dimensional shape is a region that is defined as a range in which no sound is received, the processor 130 of the sound signal processing device 100 attenuates the sound signal corresponding to the sound source when the sound source is present in the sound processing area (e.g., A-RC) and outputs it as an output signal Sout(t). When the sound source is outside the sound processing area A-RC, the processor 130 can directly output or amplify the output signal Sout(t).

[0139] According to an embodiment, the processor 130 of the sound signal processing device 100 can train an artificial neural network model by removing the sound signal of the object to be processed within a three-dimensional sound processing region (e.g., A-RC) and outputting an output signal Sout(t) corresponding to 0.

[0140] According to another embodiment, the processor 130 of the sound signal processing apparatus 100 can train an artificial neural network model by outputting an output signal Sout(t) relating to the reduction of the magnitude of the processed object's sound signal within a three-dimensional sound processing region (e.g., A-RC). For example, the processor 130 can train the artificial neural network model such that the magnitude of the processed object's sound signal within the three-dimensional sound processing region (e.g., A-RC) decreases at a uniform scale within the sound processing region (e.g., A-RC).

[0141] According to an embodiment, the processor 130 of the sound signal processing device 100 can train an artificial neural network model using supervised learning.

[0142] Figure 8 This is a graph illustrating the sound signal processing effect of the sound signal processing apparatus according to an embodiment of the present invention.

[0143] Reference Figure 8 For a three-dimensional sound processing area, in the sound signal processing apparatus according to an embodiment of the present invention, the area to be received sound is set, that is, the sound processing area is set in the following manner: the distance from the base point (e.g., the midpoint between the first sound sensor 110 and the second sound sensor 120) is 0 to 20 cm, and the angle formed relative to the baseline (e.g., the axis AX that passes through the first sound sensor 110 and the second sound sensor 120 together) is 0 to 60 degrees.

[0144] Figure 8 The horizontal axis of the graph shown represents the distance from the base point, and the vertical axis represents the magnitude of the received sound (e.g., the root mean square (RMS) value of the sound signal).

[0145] Referring to the graph, sound within the 0-60 degree angle range (0 degrees, 30 degrees, 60 degrees) of the three-dimensional sound processing area can be received normally, while sound at an angle exceeding 90 degrees cannot be received. Furthermore, it can be confirmed that sound within the 0-20cm distance range of the three-dimensional sound processing area can be received, but the intensity of sound signals exceeding the set distance range is received as a value approaching 0.

[0146] Figure 9 yes Figure 1 The example shown is an audio signal processing device implemented in a form combined with a wearable device.

[0147] Reference Figure 9 An embodiment is shown in which a sound signal processing device 100 is combined with a wearable device and used only to extract the voice of a user wearing the wearable device.

[0148] In this case, the first acoustic sensor 110 and the second acoustic sensor 120 included in the sound signal processing device 100 can be configured such that the axis AX passing through the first acoustic sensor 110 and the second acoustic sensor 120 together passes around the user's mouth of the wearable device.

[0149] For example, in the case of a wearable device implemented in the form of glasses, the first acoustic sensor 110 and the second acoustic sensor 120 may be configured at mutually spaced positions on the frame of the glasses.

[0150] In this case, the sound processing area A-RC can be positioned around the user's mouth so that only the user's voice is received.

[0151] Figure 10 yes Figure 1 The example shown is an audio signal processing device implemented in the form of a wireless microphone.

[0152] Reference Figure 1 and Figure 10 The sound signal processing device 100 can be implemented in the form of a wireless microphone (e.g., a Bluetooth microphone).

[0153] The first acoustic sensor 110 and the second acoustic sensor 120 can be configured at mutually spaced positions on the wireless microphone device.

[0154] In this case, the sound processing area A-RC can be positioned around the user's face so that only the user's voice is received.

[0155] Figure 11 yes Figure 1 The example shown is an audio signal processing device implemented in combination with wireless headphones.

[0156] Figure 11 yes Figure 1 The example shown illustrates the audio signal processing device implemented in conjunction with wireless headphones, such as True Wireless Stereo (TWS) headphones.

[0157] According to an embodiment, the sound signal processing device 100 can be implemented in combination with one of the two separate wireless earphone units.

[0158] Reference Figure 11An embodiment is shown in which a sound signal processing device 100 is combined with a wireless headset and used only to extract the voice of a user wearing the wireless headset.

[0159] In this case, the first acoustic sensor 110 and the second acoustic sensor 120 included in the sound signal processing device 100 can be configured such that the axis AX passing through the first acoustic sensor 110 and the second acoustic sensor 120 together passes around the mouth of the wireless earphone user.

[0160] For example, the first acoustic sensor 110 and the second acoustic sensor 120 can be configured at positions spaced apart from each other on one side of the wireless earphone.

[0161] In this case, the sound processing area A-RC can be positioned around the user's mouth so that only the user's voice is received.

[0162] According to yet another embodiment, when the sound signal processing device is implemented in conjunction with a hearing assistance device (rather than a wireless headset), the three-dimensional sound processing area according to an embodiment of the present invention can be set as an area defining a range in which sound is not received.

[0163] In this case, by setting a three-dimensional sound processing area around the hearing aid wearer's mouth, the hearing aid wearer's voice can be excluded, thereby solving the problem of the hearing aid wearer's voice being amplified and heard.

[0164] The present invention has been described above through preferred embodiments, but the present invention is not limited to the above embodiments. Those skilled in the art can make various modifications and variations without departing from the technical concept and scope of the present invention.

Claims

1. A method for processing sound signals, wherein, The method, which uses two acoustic sensors to process sound signals, includes the following steps: Training sound signals are received from the same sound source through a first sound sensor and a second sound sensor located at separate positions, respectively. The artificial neural network model is trained based on the received training sound signal, and the artificial neural network model defines a bounded three-dimensional sound processing region according to its relative positional relationship with the first sound sensor and the second sound sensor. as well as The received sound signal of the processing object is processed using the trained artificial neural network model and based on the sound processing region of the three-dimensional shape.

2. The method for processing sound signals according to claim 1, wherein, The three-dimensional sound processing area is a region that defines the range of sound to be received, or a region that defines the range of sound not to be received.

3. The method for processing sound signals according to claim 1, wherein, The three-dimensional sound processing region is defined as at least a portion of a conical region centered on the axis passing through the first sound sensor and the second sound sensor.

4. The method for processing sound signals according to claim 3, wherein, The sound processing area of ​​the three-dimensional shape is determined based on a preset angle range formed with the axis as a reference and a preset distance range from the first sound sensor and the second sound sensor.

5. The method for processing sound signals according to claim 1, wherein, The range of distances between the first acoustic sensor and the second acoustic sensor is determined by the distance on the axis from the midpoint between the first acoustic sensor and the second acoustic sensor.

6. The method for processing sound signals according to claim 1, wherein, The artificial neural network model is either a U-NET structure artificial neural network model or a recurrent neural network structure artificial neural network model.

7. The method for processing sound signals according to claim 1, wherein, In the step of training the artificial neural network model, the artificial neural network model is trained based on the time difference between the first training sound signal received by the first sound sensor and the second training sound signal received by the second sound sensor, as well as the signal strength difference between the first training sound signal and the second training sound signal.

8. The method for processing sound signals according to claim 7, wherein, The steps for training the artificial neural network model include the following: Perform short-time Fourier transforms on the first training audio signal and the second training audio signal respectively; and Using the first training sound signal after short-time Fourier transform and the second training sound signal after short-time Fourier transform, the time difference between the first training sound signal and the second training sound signal, as well as the signal strength difference between the first training sound signal and the second training sound signal, are calculated.

9. The method for processing sound signals according to claim 8, wherein, In the step of training the artificial neural network model, the time difference between the first training sound signal and the second training sound signal, as well as the signal strength difference between the first training sound signal and the second training sound signal, are used as input features of the artificial neural network model to train the artificial neural network model.

10. The method for processing sound signals according to claim 9, wherein, The time difference between the first training sound signal and the second training sound signal is calculated based on the phase difference between the transformed first training sound signal and the transformed second training sound signal.

11. The method for processing sound signals according to claim 9, wherein, The time difference between the first training sound signal and the second training sound signal is calculated based on the difference between the time point in the time domain where the magnitude of the first training sound signal has the maximum value and the time point in the time domain where the magnitude of the second training sound signal has the maximum value.

12. The method for processing sound signals according to claim 9, wherein, The signal strength difference between the first training sound signal and the second training sound signal is calculated based on the magnitude difference between the transformed first training sound signal and the transformed second training sound signal, or the magnitude ratio between the transformed first training sound signal and the transformed second training sound signal.

13. The method for processing sound signals according to claim 9, wherein, The signal strength difference between the first training sound signal and the second training sound signal is calculated based on the signal energy difference between the first training sound signal in the time domain and the second training sound signal in the time domain.

14. The method for processing sound signals according to claim 1, wherein, The steps for processing the received sound signal of the processing object based on the sound processing region of the three-dimensional shape include the following steps: Remove the sound signal of the object being processed that exceeds the sound processing area of ​​the three-dimensional shape, or reduce the size of the sound signal of the object being processed that exceeds the sound processing area of ​​the three-dimensional shape.

15. The method for processing sound signals according to claim 1, wherein, The steps for processing the received sound signal of the processing object based on the sound processing region of the three-dimensional shape include the following steps: Remove the object sound signal within the sound processing area of ​​the three-dimensional shape, or reduce the size of the object sound signal within the sound processing area of ​​the three-dimensional shape.

16. The method for processing sound signals according to claim 1, wherein, In the step of training the artificial neural network model, the artificial neural network model is trained based on a phase difference matrix including time difference information between a first training sound signal received by the first sound sensor and a second training sound signal received by the second sound sensor, and the signal strength difference between the first training sound signal and the second training sound signal.

17. A sound signal processing device, wherein, The sound signal processing device, which uses two sound sensors to process sound signals, includes: First sound sensor; A second acoustic sensor is disposed at a position spaced apart from the first acoustic sensor; and The processor trains an artificial neural network model based on training sound signals received from the same sound source via the first sound sensor and the second sound sensor, respectively. The artificial neural network model sets a three-dimensional sound processing region according to its relative positional relationship with the first sound sensor and the second sound sensor, and processes the received sound signal of the processing target using the trained artificial neural network model and based on the three-dimensional sound processing region.