Improved sound source location tracking system using artificial intelligence and method therefor
A cost-effective and efficient sound source tracking system using a small number of microphones and AI correction accurately locates sound sources in real-time, suitable for voice recognition, sound-based surveillance, and augmented reality.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- PAICHAI UNIV IND ACADEMIC COOPERATION FOUND
- Filing Date
- 2024-11-19
- Publication Date
- 2026-04-23
AI Technical Summary
Existing sound source tracking systems using multiple microphones face high costs and complex installation, and struggle with real-time processing due to high data volume.
A system utilizing a small number of microphones, combined with an artificial intelligence model, to estimate and correct sound source location through a sound signal collection device, an AI model generation device, and a mobile device, employing TDOA calculation and LSTM-based AI model for real-time correction.
Accurately tracks sound source location in real-time with reduced costs and simplified installation, enabling applications in voice recognition, sound-based surveillance, and augmented reality systems.
Smart Images

Figure KR2024018278_23042026_PF_FP_ABST
Abstract
Description
Enhanced sound source location tracking system and method using artificial intelligence
[0001] The present invention relates to an enhanced sound source location tracking system and method using artificial intelligence, and more specifically, to a technology that tracks the location of a sound source in real time using a smaller number of microphones and corrects the location estimation value using artificial intelligence.
[0002] Existing sound source tracking systems have performed location estimation using multiple microphones, but these systems suffer from high costs and complex installation. Furthermore, systems composed of multiple microphones face difficulties in real-time processing due to the high volume of data.
[0003] Accordingly, there is a demand for technology that can accurately track the location of sound sources by utilizing artificial intelligence while reducing costs through the use of a small number of microphones.
[0004] The present invention aims to provide an enhanced sound source location tracking system and method using artificial intelligence.
[0005] As a technical means for achieving the above-mentioned objective, one embodiment may provide a system for estimating the location of a sound source, comprising a sound signal collection device, an artificial intelligence model generation device, and a mobile device. The sound signal collection device comprises a square array of microphones composed of four microphones, a signal processing module for processing sound signals input from the microphones, a TDOA calculation module for calculating a Time Difference of Arrival (TDOA) value for each pair of microphones, a location estimation module for estimating the location of a sound source using the TDOA value, and a data transmission module for transmitting non-real-time data to the artificial intelligence model generation device and transmitting real-time data to the mobile device. The artificial intelligence model generation device generates an artificial intelligence model for correcting the location of the sound source and provides the artificial intelligence model to the mobile device, and the mobile device corrects the estimated location of the sound source using the artificial intelligence model and outputs the corrected location of the sound source in real time.
[0006] The present invention has the technical effect of accurately tracking the location of a sound source at low cost using a small number of microphones. In particular, it can indicate an accurate location through position correction using artificial intelligence.
[0007] In addition, it has the advantage of simple system installation and low data throughput, allowing for the accurate display of sound source locations in real time. Accordingly, it can be utilized in various application fields, such as voice recognition systems in conference rooms, sound-based surveillance systems, robot navigation, and augmented reality systems.
[0008] FIG. 1 is a diagram schematically illustrating a system for estimating the location of a sound source according to one embodiment.
[0009] FIGS. 2 and 3 are flowcharts for explaining a method for estimating the location of a sound source according to one embodiment.
[0010] FIGS. 4a and 4b are drawings illustrating an example of estimating the location of a sound source using a Time Difference of Arrival (TDOA) value according to one embodiment.
[0011] FIG. 5 is a diagram showing the configuration of an artificial intelligence model according to one embodiment.
[0012] FIGS. 6a and 6b are drawings showing examples of sound source location output screens according to one embodiment.
[0013] FIG. 7 is a block diagram illustrating the configuration of a system for estimating the location of a sound source according to one embodiment.
[0014] A first aspect of the present invention may provide a system for estimating the location of a sound source, comprising a sound signal collection device, an artificial intelligence model generation device, and a mobile device. The sound signal collection device comprises a square array of microphones composed of four microphones, a signal processing module for processing sound signals input from the microphones, a TDOA calculation module for calculating a Time Difference of Arrival (TDOA) value for each pair of microphones, a location estimation module for estimating the location of a sound source using the TDOA value, and a data transmission module for transmitting non-real-time data to an artificial intelligence model generation device and transmitting real-time data to a mobile device. The artificial intelligence model generation device generates an artificial intelligence model for correcting the location of a sound source and provides the artificial intelligence model to the mobile device. The mobile device corrects the estimated location of the sound source using the artificial intelligence model and outputs the corrected location of the sound source in real time.
[0015] In addition, the position estimation module can estimate the location of a sound source using non-linear equations and the least squares method with the TDOA value.
[0016] In addition, the artificial intelligence model generation device can train the artificial intelligence model using non-real-time data received from the sound signal collection device.
[0017] In addition, the artificial intelligence model may be a model based on LSTM (Long Short-Term Memory).
[0018] In addition, corrected location coordinates can be output by inputting sequence-type location estimation values into the artificial intelligence model.
[0019] In addition, a background image layer and a sound source location display layer can be mixed and output on the screen of a mobile device.
[0020] A second aspect of the present invention may provide a method for estimating the location of a sound source, comprising the steps of: processing sound signals input from four microphones in a square array; calculating a TDOA value for each pair of microphones; estimating the location of a sound source using the TDOA value; generating an artificial intelligence model for correcting the location of the sound source; correcting the estimated location of the sound source using the artificial intelligence model; and outputting the corrected location of the sound source in real time.
[0021] The invention will be described in detail below by way of exemplary embodiments with reference to the attached drawings. The following embodiments are intended only to embody the invention and do not limit or restrict the scope of the invention. Anything that can be easily inferred by a person skilled in the art to which the invention pertains from the detailed description and embodiments shall be interpreted as falling within the scope of the invention.
[0022] Terms such as 'composed' or 'comprising' as used in this specification should not be interpreted as necessarily including all of the various components or steps, and should be interpreted as some of the components or steps may not be included, or additional components or steps may be included.
[0023] The terms used in this specification are described in their current, general form considering the functions mentioned herein; however, they may refer to various other terms depending on the intent of those skilled in the art, case law, the emergence of new technologies, etc. Therefore, the terms used in this specification should not be interpreted solely by their names, but should be interpreted based on the meaning of the terms and the overall content of this specification. Furthermore, singular expressions include a plural meaning unless the context clearly indicates a singular meaning.
[0024] The terms “above” and similar designations used in this specification (particularly in the claims) may indicate both singular and plural forms. Furthermore, unless there is a description explicitly specifying the order of the steps describing the method according to this specification, the described steps may be performed in a suitable order. The present invention is not limited by the order in which the described steps are described.
[0025] Phrases such as "in one embodiment" appearing in various places in this specification do not necessarily refer to the same embodiment.
[0026] Some embodiments of this specification may be represented by functional block configurations and various processing steps. Some or all of these functional blocks may be implemented by various numbers of hardware and / or software configurations that execute specific functions.
[0027] The embodiments relate to an enhanced sound source location tracking system and method using artificial intelligence, and detailed descriptions of matters widely known to those skilled in the art to which the following embodiments belong are omitted. The present invention will be described in detail below with reference to the attached drawings.
[0028] FIG. 1 is a diagram schematically illustrating a system for estimating the location of a sound source according to one embodiment. A system for estimating the location of a sound source (100) according to one embodiment may include a sound signal collection device (110), an AI model generation device (120), and a mobile device (130).
[0029] The sound signal collection device (110) can collect sound signals. For example, the collected sound signals may include non-real-time data and real-time data. Additionally, the sound signal collection device can receive sound signals through a microphone array. Furthermore, the signal received through the microphone array can be processed to estimate the location of the sound source using the time difference between each microphone. The AI model generation device (120) can train and verify an artificial intelligence model for correcting the location of the sound source. The artificial intelligence model can be trained using the non-real-time data received from the sound signal collection device (110). Additionally, the performance of the trained model can be verified, and if the performance is poor, it can be retrained. The trained artificial intelligence model can then be transmitted to a mobile device (130). For example, the trained or retrained AI model can be provided to the mobile device (130) for real-time prediction.
[0030] Accordingly, the user's mobile device (130) receives a sound signal in real time from the sound signal collection device (110) and corrects the location of the sound source through an AI model to predict the location of the sound source more accurately. In this specification, the location of the sound source may refer to the source location where the sound originates. For example, in a conference room, it may refer to the location of the speaker. The mobile device (130) can immediately provide the prediction result to the user. For example, the mobile device (130) can display the location of the sound source in real time on the screen of the mobile device (130).
[0031] In this specification, the learning and prediction performance of sound source location estimation can be maximized by utilizing an LSTM model as an AI model. Since the LSTM model has excellent ability to effectively process and remember time sequence data, it can predict the location more accurately by learning from past location data.
[0032] This sound source location tracking technology can be utilized in various application fields, such as voice recognition systems in conference rooms, sound-based surveillance systems, robot navigation, and augmented reality systems.
[0033] In one embodiment, the location of a sound source can be tracked and displayed on the video in a conference system. The location of the speaker can be tracked in real time by installing a four-microphone array in the conference room. By tracking the location of the sound source, the location of the conference speaker can be predicted and displayed on the real-time video, which can then be output to a monitor. Through this, conference participants can easily identify the location of the speaker, and an automatic camera system can be linked to follow the speaker.
[0034] In another embodiment, the location of a sound source in an inter-floor noise measurement system can be tracked and an image displayed. By installing an array of four microphones within a building, the location of the noise generation can be tracked in real time. That is, the location of the noise within the building can be calculated and displayed in real time through a display. Through this, the location of the noise generation can be identified and actively responded to.
[0035] FIGS. 2 and 3 are flowcharts for explaining a method for estimating the location of a sound source according to one embodiment.
[0036] In step S210, sound signals input from four microphones in a square array can be processed. For example, four microphones in a square array are illustrated in FIG. 4a. Since each of the four microphones is positioned at a vertex of the square, the relative positions between each microphone can be easily established. This square array is used to estimate the location of a sound source through the difference in signal arrival times between the microphones.
[0037] In one embodiment, analog sound signals input through microphones can be collected. The analog signals can then be converted into digital signals and saved in audio file formats (e.g., WAV and MP3 files). As illustrated in FIG. 3, preprocessing can be performed to collect signals in real time using microphones, digitize them, and remove noise, or preprocessing can be performed to filter and convert previously recorded audio files.
[0038] In step S220, the TDOA value can be calculated for each pair of microphones. That is, the TDOA value is obtained by calculating the difference in signal arrival times between each microphone. An example of calculating the TDOA value will be examined in more detail with reference to Fig. 4b.
[0039] In step S230, the location of the sound source can be estimated using the TDOA value. In one embodiment, the location of the sound source can be estimated using the TDOA value through a non-linear equation and the least squares method. An example of estimating the location of the sound source using the TDOA value will be examined in more detail with reference to FIG. 4b. As shown in FIG. 3, the initial location of the sound source can be estimated using the TODA value, and then the initial location can be corrected through an artificial intelligence model to output the final location.
[0040] In step S240, an artificial intelligence model for correcting the location of a sound source can be created. As illustrated in FIG. 3, the artificial intelligence model can be trained using previously collected and preprocessed non-real-time sound signal data. Additionally, the non-real-time data can be divided into a training dataset for training the artificial intelligence model and a test set for evaluating the model.
[0041] In one embodiment, the artificial intelligence model may be a model based on LSTM (Long Short-Term Memory). An example of the configuration of the artificial intelligence model will be examined with reference to FIG. 5.
[0042] As shown in Fig. 3, after creating an artificial intelligence model, it can be trained and its performance evaluated again. If the evaluation result is good, it can be saved, and if it is determined that the model's performance is poor, it can be adjusted and retrained.
[0043] In step S250, the estimated location of the sound source can be corrected using an artificial intelligence model. As illustrated in FIG. 3, the location estimate values in the form of a sequence can be input into the artificial intelligence model to output corrected location coordinates. That is, if real-time data is processed and input according to the input format of the artificial intelligence model, the artificial intelligence model can calculate the corrected location values.
[0044] In step S260, the location of the corrected sound source can be output in real time. In one embodiment, a background image layer and a sound source location display layer can be mixed and output on the screen of a user's mobile device. Examples of output will be examined in more detail with reference to FIGS. 6a and 6b.
[0045] FIGS. 4a and 4b are drawings illustrating an example of estimating the location of a sound source using a Time Difference of Arrival (TDOA) value according to one embodiment.
[0046] As shown in FIGS. 4a and 4b, four microphones are arranged in a square array. Each microphone is positioned at a vertex of the square, making it easy to understand the relative positions between the microphones due to this arrangement. The distance between each microphone is d, and the positions of each microphone are precisely positioned to form a square with all sides of equal length. That is, the position coordinates of microphone 1 are (0,0), the position coordinates of microphone 2 are (d, 0), the position coordinates of microphone 3 are (d, d), and the position coordinates of microphone 4 are (0, d).
[0047] This square array can be used to estimate the location of a sound source through the Time Difference of Arrival (TDOA) between microphones. Since the relative positions between each microphone are accurately known, the location of the sound source can be efficiently estimated through TDOA calculation.
[0048] After converting the angle to radians, calculate the distance between a pair of adjacent microphones using the following mathematical formula.
[0049]
[0050] And the signals input from each microphone can be digitally filtered and converted and transmitted to the TDOA calculation module.
[0051] The time difference (TDOA) between two microphones i and j is calculated using the following formula.
[0052]
[0053] In mathematical formula 2 is the time difference (TDOA) between two microphones i and j, and is the maximum correlation delay between the two microphone signals, and is the sampling rate. For example, the sampling rate can be 16,000 Hz.
[0054] Also, when the speed of sound is v, the difference in distance between microphones i and j It is calculated as follows. For example, the speed of sound v is 343 m / s.
[0055]
[0056] The positions of microphones i and j ( , )and ( , Let ) and let the source location of the sound be (x, y). Then the nonlinear equation is Equation 4.
[0057]
[0058] For example, the distance between microphone 1 and microphone 2 can be calculated using the above nonlinear equation.
[0059] For multiple pairs of microphones, the above nonlinear equation is solved using the least squares method to estimate the location (x, y) of the sound source as follows. For example, the same nonlinear equation is set up for other pairs of microphones and solved using the least squares method. For example, it could be between microphone 1 and microphone 3, between microphone 1 and microphone 4, etc.
[0060]
[0061] FIG. 5 is a diagram showing the configuration of an artificial intelligence model according to one embodiment. As the artificial intelligence model, a Long Short-Term Memory (LSTM) model may be used to improve the accuracy of sound source location estimation through AI correction. Since the LSTM model has excellent ability to effectively process and remember temporal sequence data, it enables the prediction of a more accurate location by learning past location data. Non-real-time data may be collected to create the artificial intelligence model and perform model training, and the location values of real-time data can be corrected through the model generated by training.
[0062] First, sound signal data collected in real time through a microphone array is processed to estimate the location value of the sound source, and the estimated location value of the sound source is converted to fit the input format of the LSTM model.
[0063] The structure of the LSTM model consists of an input layer, an LSTM layer, an LSTM layer, and an output layer, and the data input to the model is time steps and features.
[0064] For example, time step T is 10 historical data points, and feature F may include a sequence of 6 TDOA values and 2 initial position estimates. Through this, the output data of the model may be the corrected sound source position values (e.g., xy coordinates).
[0065] For example, when creating an AI model, multiple data are collected and trained, and when using the AI model in real-time to correct the location of a sound source, 10 data points can be input into the AI model.
[0066] For example, the 10 data points in sequence form input into an artificial intelligence model are as follows.
[0067] [0.0001, 0.0002, 0.0003, 0.0001, 0.0025, 0.0001, 1.5, 2.5]
[0068] [0.0002, 0.0003, 0.0015, 0.001, 0.0025, 0.001, 1.6, 2.4]
[0069] …
[0070] [0.0002, 0.0003, 0.0015, 0.001, 0.0025, 0.001, 1.6, 2.4]
[0071] For example, when the above data is input into an artificial intelligence model, the output location coordinates of the sound source are as follows.
[0072] [1.65, 2.35]
[0073] FIGS. 6a and 6b are drawings illustrating examples of a sound source location output screen according to one embodiment. As shown in FIG. 6a, the configuration of the screen displaying the location of a sound source in real time can be intuitively designed so that the user can easily identify the location of the sound source. For example, it can be configured with a structure that displays the location of the sound source by distinguishing two layers on top of a real-time video. Looking at FIG. 6a, the display of a mobile device may include coordinate information and status indicators of the sound source, settings and control buttons, and an output screen composed of a background screen layer and a sound source location display layer.
[0074] In one embodiment, the coordinate information and status indicator may provide text information that is updated in real time. That is, it allows the user to check the current coordinates of the sound source and the system status in real time.
[0075] As illustrated in FIG. 6b, the output screen may display a combination of a background layer and a sound source location display layer. The background layer can display a video captured in real time to serve as a background that allows for the visual recognition of the sound source location. The sound source location display layer can display the location of the sound source as a graphic element by overlaying it on the real-time video. This allows the user to intuitively identify the location of the sound source.
[0076] FIG. 7 is a block diagram illustrating the configuration of a system for estimating the location of a sound source according to one embodiment.
[0077] As illustrated in FIG. 7, a system (100) for estimating the location of a sound source according to one embodiment may include a sound signal collection device (110), an AI model generation device (120), and a mobile device (130). However, not all of the illustrated components are essential components. The system (100) may be implemented with more components than those illustrated, or with fewer components. The above components will be examined in turn below.
[0078] The sound signal collection device (110) can collect sound signal data through a microphone array (111). Then, the signal processing module (113) processes the sound signals input from the microphones, converts the analog signals into digital signals, and saves the digital signals in an audio file format (e.g., WAV and MP3 files).
[0079] The TDOA calculation module (115) can calculate the Time Difference of Arrival (TDOA) value for each pair (2) of four microphones. The location estimation module (117) can estimate the location of the sound source using the TDOA value. The data transmission module (119) can transmit non-real-time data to the AI model generation device (120) and transmit real-time data to the mobile device (130).
[0080] The AI model generation device (120) can train and verify an artificial intelligence model for correcting the location of a sound source. First, non-real-time sound signal data received from the sound signal collection device (110) is stored in the data storage (121), and the artificial intelligence model can be trained in the model training module (123) using this data. After evaluating the performance of the artificial intelligence model trained with verification data, if the performance is poor, it can be retrained in the model training module (123). Additionally, the model provision module (125) can transmit the trained model to a mobile device (130).
[0081] The data receiving module (131) of the user's mobile device (130) can receive an artificial intelligence model from the AI model generation device (120) and receive real-time sound signal data from the sound signal collection device (110). The AI model application module (133) and the data processing module (135) can apply the real-time sound signal data received from the sound signal collection device (110) to the artificial intelligence model to correct the location of the sound source. The display module (137) can immediately provide the user with the location of the sound source that has been finally determined. For example, a background screen layer and a sound source location display layer can be mixed and output on the screen of the mobile device (130).
[0082] Meanwhile, the embodiments of the present invention described above can be written as a program executable on a computer and can be implemented on a general-purpose digital computer that operates the program using a computer-readable recording medium.
[0083] The above computer-readable recording media includes storage media such as magnetic storage media (e.g., ROM, floppy disk, hard disk, etc.) and optical reading media (e.g., CD-ROM, DVD, etc.).
[0084] Although embodiments of the present invention have been described above with reference to the attached drawings, those skilled in the art will understand that the present invention may be implemented in other specific forms without changing its technical concept or essential features. Therefore, the embodiments described above should be understood as illustrative in all respects and not restrictive.
Claims
1. In a system for estimating the location of a sound source, Includes a sound signal collection device, an artificial intelligence model generation device, and a mobile device, The above sound signal collection device is, A square array of microphones consisting of four microphones; A signal processing module that processes sound signals input from the above microphones; A TDOA calculation module that calculates the Time Difference of Arrival (TDOA) value for each pair of the above microphones; A position estimation module that estimates the location of a sound source using the above TDOA value; and A data transmission module that transmits non-real-time data to the artificial intelligence model generation device and transmits real-time data to the mobile device; Includes, The above artificial intelligence model generation device generates an artificial intelligence model for correcting the location of the sound source and provides the artificial intelligence model to the mobile device, and A system in which the mobile device uses the artificial intelligence model to correct the location of the estimated sound source and outputs the corrected location of the sound source in real time.
2. In Paragraph 1, A system in which the above-mentioned position estimation module estimates the position of a sound source through a non-linear equation and the least squares method using the above-mentioned TDOA value.
3. In Paragraph 1, The above artificial intelligence model generation device is a system that trains an artificial intelligence model using non-real-time data received from the sound signal collection device.
4. In Paragraph 1, The above artificial intelligence model is a system that is a model based on LSTM (Long Short-Term Memory).
5. In Paragraph 4, A system that inputs sequence-type position estimation values into the above artificial intelligence model and outputs corrected position coordinates.
6. In Paragraph 1, A system that outputs a combination of a background image layer and a sound source location display layer on the screen of the mobile device.
7. In a method for estimating the location of a sound source, A step of processing sound signals input from four microphones in a square array; A step of calculating a TDOA value for each pair of the above microphones; A step of estimating the location of a sound source using the above TDOA value; A step of generating an artificial intelligence model to correct the location of the above sound source; A step of correcting the location of the estimated sound source using the artificial intelligence model; and A step of outputting the location of the corrected sound source in real time; A method including
Citation Information
Patent Citations
Upper structure of liquified gas storage tank and liquefied gas storage tank including the same
KR1020230014441A
Apparatus for non-vehicle fuel cell system
KR1020230123190A
High density lipoprotein mimicking solid lipid nanoparticles for drug delivery and uses thereof
KR102402620B1
Control method of intelligent for sound-position tracking and system thereof
KR102497914B1
Method and apparatus for identifying aircraft
KR102672778B1