Information processing device, information processing method, and program
The audio system uses spread spectrum modulation and machine learning to determine the two-dimensional position of a microphone in a stereo speaker setup, overcoming synchronization limitations and enabling precise audio adjustments.
Patent Information
- Application Number
- JP2023525383
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-06-03
- Filing Date
- 2022-02-09
- Publication Date
- 2025-10-22
- Estimated Expiration
- 2042-02-09
AI Technical Summary
Existing audio systems with stereo speakers and a microphone cannot accurately determine the two-dimensional position of the microphone without time-synchronization with multiple transmitting devices.
An audio system using stereo speakers and a microphone that calculates the position of the microphone based on arrival time difference and peak power ratio of spread spectrum modulated audio signals received from known positions, employing spread codes and machine learning for precise positioning.
Enables accurate two-dimensional positioning of the microphone relative to stereo speakers, allowing for realistic audio adjustments and sound field corrections based on the user's movement.
Smart Images

Figure 0007758036000006 
Figure 0007758036000007 
Figure 0007758036000008
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to an information processing device, an information processing method, and a program, and more particularly to an information processing device, an information processing method, and a program that are capable of measuring the position of a microphone using stereo speakers and a microphone. [Background technology]
[0002] A technology has been proposed in which a transmitting device modulates a data code with a code sequence to generate a modulated signal, and emits the modulated signal as sound, while a receiving device receives the emitted sound, calculates the correlation between the modulated signal, which is the received sound signal, and the code sequence, and measures the distance from the transmitting device based on the peak of the correlation (see Patent Document 1). [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2014-220741 Summary of the Invention [Problem to be solved by the invention]
[0004] However, when using the technology described in Patent Document 1, although the receiving device can measure the distance to the transmitting device, in order to determine the two-dimensional position of the receiving device, if the receiving device and the transmitting device are not time-synchronized, it is necessary to use at least three transmitting devices.
[0005] That is, in a typical audio system consisting of a stereo speaker (transmitter) made up of two speakers and a microphone (receiver), it is not possible to obtain the two-dimensional position of the microphone.
[0006] The present disclosure has been made in view of the above circumstances, and particularly aims to enable measurement of the two-dimensional position of a microphone in an audio system including stereo speakers and a microphone. [Means for solving the problem]
[0007] An information processing device and a program according to one aspect of the present disclosure include an audio receiving unit that receives audio signals output from two audio output blocks located at known positions, the audio signals consisting of spread code signals whose spread codes have been spread spectrum modulated, and a position calculation unit that calculates the position of the audio receiving unit based on an arrival time difference distance, which is the difference in distance determined from the arrival time, which is the time it takes for the audio signals from the two audio output blocks to arrive at the audio receiving unit and be received.
[0008] An information processing method according to one aspect of the present disclosure is an information processing method for an information processing device having an audio receiving unit that receives audio signals output from two audio output blocks located at known positions, the audio signals consisting of spread code signals whose spread codes have been spread spectrum modulated, and includes a step of calculating the position of the audio receiving unit based on an arrival time difference distance, which is the difference in distance determined from the arrival time, which is the time it takes for the audio signals from the two audio output blocks to arrive at the audio receiving unit and be received.
[0009] In one aspect of the present disclosure, an audio receiving unit receives audio signals output from two audio output blocks located at known positions, the audio signals consisting of spread code signals whose spread codes are spread spectrum modulated, and the position of the audio receiving unit is calculated based on an arrival time difference distance, which is the difference in distance determined from the arrival time, which is the time it takes for the audio signals from the two audio output blocks to arrive at the audio receiving unit and be received. [Brief explanation of the drawings]
[0010] [Figure 1] 1 is a diagram illustrating an example configuration of a home audio system according to the present disclosure. [Figure 2]2 is a diagram illustrating an example of the configuration of the audio output block in FIG. 1. FIG. [Figure 3] FIG. 2 is a diagram illustrating an example of the configuration of the electronic device in FIG. [Figure 4] 4 is a diagram illustrating an example of the configuration of a position calculation unit in FIG. 3. FIG. [Figure 5] FIG. 1 is a diagram illustrating communication using a spreading code. [Figure 6] 1A and 1B are diagrams illustrating autocorrelation and cross-correlation of spreading codes; [Figure 7] FIG. 10 is a diagram illustrating the arrival time of a spreading code using cross-correlation. [Figure 8] FIG. 10 is a diagram illustrating an example of the configuration of an arrival time calculation unit. [Figure 9] FIG. 1 is a diagram illustrating human hearing. [Figure 10] FIG. 10 is a diagram illustrating a frequency shift of a spreading code. [Figure 11] FIG. 10 is a diagram illustrating a procedure for frequency shifting a spreading code. [Figure 12] FIG. 10 is a diagram illustrating an example in which multipath is taken into consideration. [Figure 13] FIG. 10 is a diagram illustrating a time difference of arrival distance. [Figure 14] FIG. 10 is a diagram illustrating a peak power ratio. [Figure 15] FIG. 2 is a diagram illustrating a configuration example of a position calculation unit according to a first embodiment. [Figure 16] 10 is a flowchart illustrating a sound output process by the sound output block. [Figure 17] 4 is a flowchart illustrating a sound collection process performed by the electronic device of FIG. 3. [Figure 18] 10A and 10B are diagrams illustrating an application example of the position calculation unit according to the first embodiment. [Figure 19] FIG. 10 is a diagram illustrating the audible range when the TV is positioned at the center between the audio output blocks. [Figure 20] FIG. 10 is a diagram illustrating the audible range when the TV is not positioned at the center between the audio output blocks. [Figure 21] FIG. 10 is a diagram illustrating a peak power frequency component ratio. [Figure 22] FIG. 10 is a diagram illustrating a configuration example of an electronic device according to a second embodiment. [Figure 23] 23 is a diagram illustrating an example of the configuration of a position calculation unit in FIG. 22. FIG. [Figure 24] FIG. 10 is a diagram illustrating a configuration example of a position calculation unit according to a second embodiment. [Figure 25] 23 is a flowchart illustrating a sound collection process performed by the electronic device of FIG. 22. [Figure 26] This is an example of the configuration of a home audio system in which an IMU is provided in an electronic device. [Figure 27] FIG. 10 is a diagram illustrating a configuration example of an electronic device according to a third embodiment. [Figure 28] FIG. 28 is a diagram illustrating an example of the configuration of a position calculation unit in FIG. 27. [Figure 29] 28 is a flowchart illustrating a sound collection process performed by the electronic device of FIG. 27. [Figure 30] FIG. 10 is a diagram illustrating an application example of the third embodiment. [Figure 31] 1 shows an example of the configuration of a general-purpose computer. DETAILED DESCRIPTION OF THE INVENTION
[0011] Preferred embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings. In this specification and drawings, components having substantially the same functional configurations are designated by the same reference numerals, and redundant description will be omitted.
[0012] Hereinafter, embodiments of the present technology will be described in the following order. 1. First embodiment 2. Application example of the first embodiment 3. Second Embodiment 4. Third Embodiment 5. Application Example of the Third Embodiment 6. Software implementation example
[0013] <<1. First Embodiment>> <Home audio system configuration> The present disclosure is particularly directed to an audio system that includes a stereo speaker system with two speakers and a microphone, and enables the position of the microphone to be measured.
[0014] FIG. 1 shows an example of the configuration of a home audio system to which the present disclosure is applied.
[0015] 1 includes a display device 30 such as a TV (television receiver), audio output blocks 31-1 and 31-2, and an electronic device 32. Hereinafter, when there is no need to distinguish between the audio output blocks 31-1 and 31-2, they will be simply referred to as audio output block 31, and the same will be used for other components. Furthermore, the display device 30 will also be simply referred to as TV 30.
[0016] The audio output blocks 31-1 and 31-2 each have a speaker and emit audio from music content, games, etc., including audio consisting of a modulated signal obtained by spectrum spread modulation of a data code for identifying the position of the electronic device 32 with a spread code.
[0017] The electronic device 32 is carried or worn by a user, and is, for example, a smartphone used as a game controller or an HMD (Head Mounted Display).
[0018] The electronic device 32 includes an audio input block 41 having an audio input unit 51 such as a microphone that receives audio emitted from each of the audio output blocks 31-1 and 31-2, and a position detection unit 52 that detects its own position relative to the audio output blocks 31-1 and 31-2.
[0019] The audio input block 41 recognizes in advance the spatial positions of the display device 30 and the audio output blocks 31-1 and 31-2 as known position information, and uses the audio input unit 51 to pick up the audio emitted from the audio output block 31.The position detection unit 52 calculates the distance to each of the audio output blocks 31-1 and 31-2 based on the modulated signal contained in the picked up audio, and detects its own two-dimensional position (x, y) relative to the audio output blocks 31-1 and 31-2.
[0020] This allows the position of the electronic device 32 relative to the audio output blocks 31-1 and 31-2 to be identified, and the audio output from the audio output blocks 31-1 and 31-2 can be corrected for sound field positioning according to the identified position, allowing the user to hear and see realistic audio that corresponds to their own movements.
[0021] <Example of audio output block configuration> Next, an example of the configuration of the audio output block 31 will be described with reference to FIG.
[0022] The audio output block 31 includes a spreading code generation unit 71 , a known music sound source generation unit 72 , an audio generation unit 73 , an audio output unit 74 , and a communication unit 75 .
[0023] The spreading code generator 71 generates a spreading code and outputs it to the voice generator 73 .
[0024] The known music sound source generating unit 72 stores known music pieces, generates known music sound sources based on the stored known music pieces, and outputs the generated sound sources to the sound generating unit 73 .
[0025] The sound generating unit 73 applies spread spectrum modulation using a spreading code to the known music sound source to generate sound consisting of a spread spectrum signal, and outputs the sound to the sound output unit 74 .
[0026] More specifically, the sound generation unit 73 includes a diffusion unit 81 , a frequency shift processing unit 82 , and a sound field control unit 83 .
[0027] The spreading unit 81 applies spread spectrum modulation using a spreading code to the known music sound source to generate a spread spectrum signal.
[0028] The frequency shift processing unit 82 shifts the frequency of the spreading code in the spread spectrum signal to a frequency band that is difficult for the human ear to hear.
[0029] The sound field control unit 83 reproduces a sound field according to the positional relationship with the electronic device 32 itself, based on the information on the position of the electronic device 32 supplied from the electronic device 32.
[0030] The audio output unit 74 is, for example, a speaker, and outputs the known music sound source supplied from the audio generation unit 73 and audio based on the spread spectrum signal.
[0031] The communication unit 75 communicates with the electronic device 32 by wireless communication such as Bluetooth (registered trademark), and sends and receives various data and commands.
[0032] <Example of electronic device configuration> Next, an example of the configuration of the electronic device 32 will be described with reference to FIG.
[0033] The electronic device 32 includes a voice input block 41, a control unit 42, an output unit 43, and a communication unit 44.
[0034] The audio input block 41 receives audio inputs emitted from the audio output blocks 31-1 and 31-2, and calculates the arrival times and peak powers of the received audio based on the correlation between the received audio, the spread spectrum signal, and the spread code.The audio input block 41 then calculates its own two-dimensional position (x, y) based on the arrival time difference distance based on the calculated arrival times and the peak power ratio, which is the ratio of the peak powers of the audio output blocks 31-1 and 31-2, and outputs the calculated position to the control unit 42.
[0035] The control unit 42 controls, for example, the communication unit 44 based on the position of the electronic device 32 supplied from the audio input block 41, and when it acquires information notified by the audio output blocks 31-1 and 31-2, it presents the information to the user via the output unit 43, which is made up of a display, a speaker, etc. The control unit 42 also controls the communication unit 44 to send commands to the audio output blocks 31-1 and 31-2 to set a sound field based on the two-dimensional position of the electronic device 32.
[0036] The communication unit 44 communicates with the audio output block 31 by wireless communication such as Bluetooth (registered trademark), and sends and receives various data and commands.
[0037] More specifically, the voice input block 41 includes a voice input unit 51 and a position detection unit 52 .
[0038] The audio input unit 51 is, for example, a microphone, which collects the audio emitted from the audio output blocks 31-1 and 31-2 and outputs the audio to the position detection unit 52.
[0039] The position detection unit 52 determines the position of the electronic device 32 based on the sounds picked up by the sound input unit 51 and emitted from the sound output blocks 31-1 and 31-2.
[0040] The position detection unit 52 includes a known music sound source removal unit 91 , a space transfer characteristic calculation unit 92 , an arrival time calculation unit 93 , a peak power detection unit 94 , and a position calculation unit 95 .
[0041] The spatial transfer characteristic calculation unit 92 calculates spatial transfer characteristics based on the audio information supplied from the audio input unit 51, the characteristics of the microphone that constitutes the audio input unit 51, and the characteristics of the speaker that constitutes the audio output unit 74 of the audio output block 31, and outputs the calculated spatial transfer characteristics to the known music sound source removal unit 91.
[0042] The known music sound source remover 91 stores the music sound source stored in advance in the known music sound source generator 72 in the audio output block 31 as the known music sound source.
[0043] Then, the known music sound source removal unit 91 takes into consideration the spatial transfer characteristics supplied from the spatial transfer characteristic calculation unit 92, removes the components of the known music sound source from the audio supplied from the audio input unit 51, and outputs the result to an arrival time calculation unit 93 and a peak power detection unit 94.
[0044] That is, the known music sound source removal unit 91 removes the known music sound source components from the sound collected by the sound input unit 51 and outputs only the spread spectrum signal components to the arrival time calculation unit 93 and the peak power detection unit 94.
[0045] The arrival time calculation unit 93 calculates the arrival time from when the sound is emitted from each of the sound output blocks 31-1 and 31-2 to when the sound is picked up, based on the spectrum spread signal component contained in the sound picked up by the sound input unit 51, and outputs the calculated arrival time to the peak power detection unit 94 and the position calculation unit 95.
[0046] The method for calculating the arrival time will be described in detail later.
[0047] The peak power detector 94 detects the power of the spread spectrum signal component at the peak detected by the arrival time calculator 93 and outputs it to the position calculator 95 .
[0048] The position calculation unit 95 calculates the arrival time difference distance and peak power ratio based on the arrival times of each of the audio output blocks 31-1 and 31-2 supplied from the arrival time calculation unit 93 and the peak power supplied from the peak power detection unit 94, and then calculates the position (two-dimensional position) of the electronic device 32 based on the calculated arrival time difference distance and peak power ratio, and outputs the position to the control unit 42.
[0049] The detailed configuration of the position calculation unit 95 will be described later with reference to FIG.
[0050] <Configuration example of position calculation unit> Next, an example of the configuration of the position calculation unit 95 will be described with reference to FIG.
[0051] The position calculation unit 95 includes an arrival time difference distance calculation unit 111 , a peak power ratio calculation unit 112 , and a position calculation unit 113 .
[0052] The arrival time difference distance calculation unit 111 calculates the difference between the distances to the audio output blocks 31-1 and 31-2, which is determined based on the arrival times of the audio output blocks 31-1 and 31-2, as the arrival time difference distance, and outputs the calculated distance to the position calculation unit 113.
[0053] The peak power ratio calculation unit 112 calculates the ratio between the peak powers of the sounds output from the sound output blocks 31 - 1 and 31 - 2 as a peak power ratio and outputs the ratio to the position calculation unit 113 .
[0054] The position calculation unit 113 calculates the position of the electronic device 32 relative to the audio output blocks 31-1 and 31-2 by machine learning using a neural network based on the arrival time difference distance of the audio output blocks 31-1 and 31-2 supplied from the arrival time difference distance calculation unit 111 and the peak power ratio of the audio output blocks 31-1 and 31-2 supplied from the peak power ratio calculation unit 112, and outputs the calculated position to the control unit 42.
[0055] <Principles of communication using spreading codes> Next, the principle of communication using spreading codes will be described with reference to FIG.
[0056] On the transmitting side on the left side of the figure, a spreading unit 81 performs spread spectrum modulation by multiplying an input signal Di to be transmitted, which has a pulse width Td, by a spreading code Ex, to generate a transmission signal De with a pulse width Tc, which is then transmitted to the receiving side on the right side of the figure.
[0057] In this case, if the frequency band Dif of the input signal Di is, for example, represented by the frequency band -1 / Td to 1 / Td, the frequency band Exf of the transmission signal De is widened by multiplying it by the spreading code Ex to the frequency band -1 / Tc to 1 / Tc (1 / Tc>1 / Td), and the energy is spread on the frequency axis.
[0058] FIG. 5 shows an example in which the transmission signal De is interfered with by the interference wave IF.
[0059] At the receiving side, the signal that has been interfered with by the interfering wave IF from the transmission signal De is received as a reception signal De'.
[0060] The arrival time calculation unit 93 restores the received signal Do by despreading the received signal De' with the same spreading code Ex.
[0061] At this time, the frequency band Exf' of the received signal De' contains the interference wave component IFEx, but in the frequency band Dof of the despread received signal Do, the interference wave component IFEx is restored as the spread frequency band IFD, thereby spreading the energy, thereby making it possible to reduce the influence of the interference wave IF on the received signal Do.
[0062] That is, as described above, in communications using spreading codes, it is possible to reduce the influence of interference waves IF occurring on the transmission path of the transmission signal De, and to improve noise resistance.
[0063] Furthermore, the spreading code has an impulse-like autocorrelation, as shown in the waveform diagram in the upper part of Fig. 6, and a cross-correlation of 0, as shown in the waveform diagram in the lower part of Fig. 6. Fig. 6 shows the change in correlation value when a Gold sequence is used as the spreading code, with the horizontal axis representing the coded sequence and the vertical axis representing the correlation value.
[0064] In other words, by setting highly random spreading codes for each of the audio output blocks 31-1 and 31-2, the audio input block 41 can appropriately distinguish and recognize the spectrum signals contained in the audio for each of the audio output blocks 31-1 and 31-2.
[0065] The spreading code is not limited to the Gold sequence, but may be an M sequence or PN (Pseudorandom Noise), for example.
[0066] <Method of calculating arrival time by arrival time calculation unit> The timing at which the peak of the observed cross-correlation is observed in the audio input block 41 is the timing at which the audio emitted in the audio output block 31 is picked up in the audio input block 41, and therefore varies depending on the distance between the audio input block 41 and the audio output block 31.
[0067] That is, for example, when the distance between the audio input block 41 and the audio output block 31 is a first distance, a peak is detected at time T1 as shown in the left part of Figure 7, and when the distance between the audio input block 41 and the audio output block 31 is a second distance that is farther than the first distance, the peak is observed at time T2 (>T1) as shown in the right part of Figure 7.
[0068] In FIG. 7, the horizontal axis indicates the time that has elapsed since the sound was output from the sound output block 31, and the vertical axis indicates the strength of the cross-correlation.
[0069] In other words, the distance between the audio input block 41 and the audio output block 31 can be calculated by multiplying the time from when audio is emitted from the audio output block 31 until a peak is observed in the cross-correlation, i.e., the arrival time of the audio emitted from the audio output block 31 until it is picked up by the audio input block 41, by the speed of sound.
[0070] <Configuration example of arrival time calculation unit> Next, an example of the configuration of the arrival time calculation unit 93 will be described with reference to FIG.
[0071] The arrival time calculation unit 93 includes an inverse shift processing unit 130 , a cross-correlation calculation unit 131 , and a peak detection unit 132 .
[0072] The inverse shift processing unit 130 restores the spread spectrum modulated spread code signal, which has been frequency shifted by upsampling in the frequency shift processing unit 82 of the audio output block 31 in the audio signal collected by the audio input unit 51, to the original frequency band by downsampling, and outputs the restored signal to the cross-correlation calculation unit 131.
[0073] The shifting of the frequency band by the frequency shifter 82 and the restoration of the frequency band by the inverse shifter 130 will be described in detail later with reference to FIG.
[0074] The cross-correlation calculation unit 131 calculates the cross-correlation between the spreading code and the received signal from which the known music sound source has been removed in the audio signal collected by the audio input unit 51 of the audio input block 41, and outputs the result to the peak detection unit 132.
[0075] The peak detection unit 132 detects the time at which the cross-correlation calculated by the cross-correlation calculation unit 131 reaches a peak, and outputs the time as the arrival time.
[0076] Here, since it is known that the calculation of the cross-correlation performed by the cross-correlation calculation unit 131 generally requires a very large amount of calculation, it is realized by an equivalent calculation requiring a smaller amount of calculation.
[0077] Specifically, the cross-correlation calculation unit 131 performs Fourier transforms on the transmission signal output as audio by the audio output unit 74 of the audio output block 31 and the received signal from which the known music sound source has been removed in the audio signal received by the audio input unit 51 of the audio input block 41, as shown in the following equations (1) and (2).
[0078]
number
number
[0079] Here, g is a received signal from which the known music sound source has been removed in the audio signal received by the audio input unit 51 of the audio input block 41, and G is a result of Fourier transform of the received signal g from which the known music sound source has been removed in the audio signal received by the audio input unit 51 of the audio input block 41.
[0080] Furthermore, h is the transmission signal output as sound by the sound output unit 74 of the sound output block 31, and H is the result of the Fourier transform of the transmission signal output as sound by the sound output unit 74 of the sound output block 31.
[0081] Furthermore, V is the speed of sound, v is the speed of the electronic device 32 (the audio input unit 51 thereof), t is time, and f is frequency.
[0082] Next, cross-correlation calculation section 131 obtains a cross spectrum by multiplying the Fourier transform results G and H together, as shown in the following equation (3).
[0083]
number
[0084] Here, P is the cross spectrum obtained by multiplying the Fourier transform results G and H together.
[0085] Then, the cross-correlation calculation unit 131 performs an inverse Fourier transform on the cross spectrum P as shown in the following equation (4), thereby finding the cross-correlation between the transmission signal h that is output as audio by the audio output unit 74 of the audio output block 31 and the reception signal g from which the known music sound source has been removed in the audio signal received by the audio input unit 51 of the audio input block 41.
[0086]
number
[0087] Here, p is the cross-correlation between the transmission signal h output as audio by the audio output unit 74 of the audio output block 31 and the received signal g from which the known music sound source has been removed in the audio signal received by the audio input unit 51 of the audio input block 41.
[0088] Then, peak detection unit 132 detects the peak of cross-correlation p, detects the arrival time T based on the detected peak of cross-correlation p, and outputs it to position calculation unit 95. Arrival time difference distance calculation unit 111 of position calculation unit 95 calculates the distance between audio input block 41 and audio output block 31 by calculating the following equation (5) based on the detected peak of cross-correlation p.
[0089]
number
[0090] Here, D is the distance (arrival time distance) between the voice input block 41 (voice input unit 51) and the voice output block 31 (voice output unit 74), T is the arrival time, and V is the speed of sound. The speed of sound V is, for example, 331.5 + 0.6 × Q (m / s) (Q is temperature ° C).
[0091] Then, arrival time difference distance calculation section 111 calculates the difference between the distances between voice input block 41 and voice output blocks 31-1 and 31-2 obtained as described above as the arrival time difference distance, and outputs the calculated distance to position calculation section 113.
[0092] The peak power detection unit 94 detects the power of each of the sounds picked up at the timing when the cross-correlation p between the sound input block 41 and the sound output blocks 31-1 and 31-2 detected by the peak detection unit 132 of the arrival time calculation unit 93 reaches its peak, and outputs the detected power to the peak power ratio calculation unit 112 of the position calculation unit 95.
[0093] The peak power ratio calculation unit 112 calculates the ratio of the peak powers supplied from the peak power detection unit 94 and outputs it to the position calculation unit 113 .
[0094] The cross-correlation calculation unit 131 may further calculate the velocity v of the electronic device 32 (the voice input unit 51 thereof) by calculating the cross-correlation p.
[0095] More specifically, the cross-correlation calculation unit 131 calculates the cross-correlation p while changing the velocity v in a predetermined step (e.g., 0.01 m / s step) within a predetermined range (e.g., −1.00 m / s to 1.00 m / s), and calculates the velocity v showing the maximum peak of the cross-correlation p as the velocity v of the electronic device 32 (of its voice input unit 51).
[0096] It is also possible to find the absolute velocity of the electronic device 32 (the audio input block 41 thereof) based on the velocity v found for each of the audio output blocks 31-1 to 31-4.
[0097] <Frequency shift> The frequency band of the spreading code signal is the Nyquist frequency Fs, which is half the sampling frequency. For example, if the Nyquist frequency Fs is 8 kHz, the frequency band is 0 to 8 kHz, which is lower than the Nyquist frequency Fs.
[0098] As shown in Figure 9, it is known that human hearing is highly sensitive to sounds in the frequency band around 3 kHz, regardless of the loudness level, decreases from around 10 kHz, and is almost inaudible above 20 kHz.
[0099] Figure 9 shows the change in sound pressure level for each frequency at loudness levels of 0, 20, 40, 60, 80, and 100 phon, where the horizontal axis is frequency and the vertical axis is sound pressure level. The thick dashed line indicates the sound pressure level at the microphone, which is constant regardless of the loudness level.
[0100] Therefore, when the frequency band of the spread spectrum signal is 0 to 8 kHz, if the sound of the spread code signal is emitted together with the sound of a known music source, it may be heard as noise by the human ear.
[0101] For example, assuming that a piece of music is played at -50 dB, the range below the sensitivity curve L in Figure 10 is considered to be the range Z1 that is inaudible to humans (a range that is difficult to recognize with the human hearing), and the range above the sensitivity curve L is considered to be the range Z2 that is audible to humans (a range that is easy to recognize with the human hearing).
[0102] In FIG. 10, the horizontal axis indicates the frequency band, and the vertical axis indicates the sound pressure level.
[0103] Therefore, for example, when the range in which the audio of the known music source being played and the audio of the spread code signal can be separated is within -30 dB, if the audio of the spread code signal is output in the range of 16 kHz to 24 kHz shown as range Z3 within range Z1, it can be made inaudible to humans (hard to recognize with the human hearing).
[0104] Therefore, the frequency shift processing unit 82 upsamples the spreading code signal Fs including the spreading code as shown in the upper left part of Figure 11 by m times as shown in the middle left part, to generate spreading code signals Fs, 2Fs, ... mFs.
[0105] Then, as shown in the lower left part of Figure 11, the frequency shift processing unit 82 applies band limitation to the spreading code signal uFs of 16 to 24 kHz, which is a frequency band that is inaudible to humans and which was described with reference to Figure 10, thereby frequency shifting the spreading code signal Fs including the spreading code signal and emitting it from the audio output unit 74 together with the known music sound source.
[0106] As shown in the lower right part of FIG. 11 , the reverse shift processing unit 130 extracts a spreading code signal uFs as shown in the middle right part of FIG. 11 by limiting the band of the sound collected by the sound input unit 51, from which the known music sound source has been removed by the known music sound source removal unit 91, to a range of 16 to 24 kHz.
[0107] Then, as shown in the upper right part of FIG. 10, the inverse shift processing unit 130 downsamples to 1 / m to generate a spreading code signal Fs including a spreading code, thereby restoring the frequency band to the original band.
[0108] By performing frequency shift in this manner, even if audio containing a spread code signal is emitted while audio of a known music source is being emitted, the audio containing the spread code signal can be made difficult to hear (difficult to recognize by the human ear).
[0109] In the above, we have explained an example of making voice including a spread code signal difficult to hear (difficult to recognize by the human hearing) by using frequency shift. However, since high-frequency sound has a tendency to travel in a straight line and is easily affected by multipath propagation due to reflections from walls and other objects, and sound blocking by obstacles, it is desirable to also use voice in a lower frequency band, including a low-frequency band below 10 kHz, for example, around 3 kHz, which is prone to diffraction, as shown in Figure 12. Figure 12 shows an example in which voice of a spread code signal is output even in a range including a low-frequency band below 10 kHz, as shown in range Z3'. Therefore, in the case of Figure 12, voice of a spread code signal is also easily recognized by the human hearing in range Z11.
[0110] In such a case, for example, the sound pressure level of the known music sound source may be set to -50 dB, and the range required for separation may be set to up to -30 dB. Then, using an auditory compression technique such as that used in ATRAC (registered trademark) or MP3 (registered trademark), the spread code signal may be auditorily masked by the known music, and the spread code signal may be emitted in an inaudible manner.
[0111] More specifically, the frequency components of the music being played back may be analyzed for each predetermined playback time unit (for example, 20 ms), and the sound pressure level of the audio of the spread code signal for each critical band (24 bark) may be dynamically increased or decreased according to the analysis results so as to cause auditory masking.
[0112] <About arrival time difference distance> Next, the arrival time difference distance will be described with reference to Fig. 13. Fig. 13 is a plot diagram of the arrival time difference distance according to the position relative to the audio output blocks 31-1 and 31-2. Note that in Fig. 13, the y-axis represents the position relative to the front direction in which the audio output blocks 31-1 and 31-2 emit audio, and the x-axis represents the position perpendicular to the direction in which the audio output blocks 31-1 and 31-2 emit audio, and the distribution shown is when the arrival time difference distance obtained when x and y are expressed in standardized units is plotted in standardized units.
[0113] That is, in FIG. 13, the audio output blocks 31-1 and 31-2 are separated by three normalized units in the x-axis direction, and the plot results of the arrival time difference distance in the range of three units from the audio output blocks 31-1 and 31-2 in the y-axis direction are shown.
[0114] As shown in FIG. 13, a correlation with the time difference of arrival distance is observed in the x-axis direction, and therefore it is considered possible to obtain the x-axis direction with a certain degree of accuracy or higher.
[0115] However, in the y-axis direction, no correlation is observed, particularly near the position of 1.5 units in the x-axis direction, i.e., near the center between the audio output blocks 31-1 and 31-2, and therefore it is thought that it is not possible to obtain the value with a precision higher than a certain level.
[0116] As a result, it is thought that only the position in the x-axis direction can be determined with a certain degree of accuracy or higher by using the time difference of arrival distance alone.
[0117] <About peak power ratio> Next, the peak power ratio will be described with reference to Fig. 14. Fig. 14 is a plot diagram of the peak power ratio as a function of position for the audio output blocks 31-1 and 31-2.
[0118] In Figure 14, the y-axis represents the position of the audio output blocks 31-1 and 31-2 relative to the front direction in which they emit audio, and the x-axis represents the position of the audio output blocks 31-1 and 31-2 in the direction perpendicular to the direction in which they emit audio.The distribution shown is when the peak power ratios obtained when x and y are expressed in standardized units are plotted in standardized units.
[0119] That is, in FIG. 14, the audio output blocks 31-1 and 31-2 are separated by three normalized units in the x-axis direction, and the plot results of peak power in the range of three units from the audio output blocks 31-1 and 31-2 in the y-axis direction are shown.
[0120] As shown in FIG. 14, a correlation with the peak power ratio is observed in both the x- and y-axis directions, and therefore it is believed that it is possible to obtain the peak power ratio with a certain degree of accuracy or higher in both the x- and y-axis directions.
[0121] <Configuration Example of First Embodiment of Position Calculation Unit> From the above, it is believed that machine learning using the arrival time difference distance and peak power ratio can determine the position (x, y) of the electronic device 32 (audio input block 41) relative to the audio output blocks 31-1 and 31-2 with a certain level of accuracy or higher.
[0122] However, it is assumed here that the positions of the audio output blocks 31-1 and 31-2 are fixed to known positions, or that the position of one of the audio output blocks 31-1 and 31-2 is known and the distance between them is also known.
[0123] Therefore, the position calculation unit 113 constructs a predetermined hidden layer 152 by machine learning for an input layer 151 consisting of a predetermined number of data items, which are the arrival time difference distance D and the peak power ratio PR, as shown in FIG. 15, and obtains an output layer 153 consisting of the position (x, y) of the electronic device 32 (voice input block 41).
[0124] More specifically, for example, for 32,760 samples, as shown in the upper right part of Figure 15, 10 pieces of data are used as information for one input layer, with peak calculations performed by sliding the time ts for each set of 3,276 samples C1, C2, C3, etc.
[0125] That is, it is assumed that the input layer 151, which is made up of the arrival time difference distance D and the peak power ratio PR for which peak calculation has been performed, is made up of 10 pieces of data.
[0126] The hidden layer 152 is composed of n layers, for example, a first layer 152a to an n-th layer 152n, with the first layer 152a being a layer with a function to mask data that satisfies a predetermined condition with respect to the data of the input layer 151, and is composed of, for example, a layer with 1280 channels. The second layer 152b to the n-th layer 152n are each composed of a layer with 128 channels.
[0127] The first layer 152a masks data from the input layer 151 that does not satisfy certain conditions, such as a peak signal-to-noise ratio of 8 times or more or an arrival time difference distance of 3 m or less, so that it is not used for processing in subsequent layers.
[0128] As a result, the second layer 152b to the nth layer 152n of the hidden layer 152 use only the data from the input layer 151 that satisfies certain conditions to determine the two-dimensional position (x, y) of the electronic device 32, which becomes the output layer 153.
[0129] The position calculation unit 113 configured as described above outputs the position (x, y) of the electronic device 32 (voice input block 41), which is the output layer 153, from the hidden layer 152 configured by machine learning, for the input layer 151 consisting of the time difference of arrival distance D and the peak power ratio PR.
[0130] <Audio sound processing> Next, the audio output process by the audio output block 31 will be described with reference to the flowchart of FIG.
[0131] In step S 11 , the spreading code generation unit 71 generates a spreading code and outputs it to the voice generation unit 73 .
[0132] In step S12 , the known music sound source generating unit 72 generates the stored known music sound source and outputs it to the sound generating unit 73 .
[0133] In step S13, the voice generating unit 73 controls the spreading unit 81 to multiply a predetermined data code by a spreading code to perform spread spectrum modulation, thereby generating a spreading code signal.
[0134] In step S14, the voice generating unit 73 controls the frequency shift processing unit 82 to frequency shift the spread code signal as described with reference to the left part of FIG.
[0135] In step S15, the sound generating unit 73 outputs the known music sound source and the frequency-shifted spread code signal to the sound output unit 74 made up of a speaker, and emits (outputs) them as sound at a predetermined sound output.
[0136] By performing the above processing in each of the audio output blocks 31-1 and 31-2, it becomes possible for the user who owns the electronic device 32 to hear the audio of the known music sound source being output.
[0137] Furthermore, since it is possible to shift the spread code signal to a frequency band that is inaudible to the human user and output it as sound, the electronic device 32 can measure the distance to the sound output block 31 based on the sound that is made up of the emitted spread code signal and shifted to a frequency band that is inaudible to humans, without making the user hear unpleasant sounds.
[0138] In step S16, the voice generation unit 73 controls the communication unit 75 to determine, through processing described below, whether or not a notification that a peak cannot be detected has been received from the electronic device 32. If a notification that a peak cannot be detected has been received in step S16, the processing proceeds to step S17.
[0139] In step S17, the sound generating unit 73 controls the communication unit 75 to determine whether or not a command to adjust the sound output has been transmitted from the electronic device 32.
[0140] If it is determined in step S17 that a command to adjust the sound output has been transmitted, the process proceeds to step S18.
[0141] In step S18, the sound generation unit 73 adjusts the sound output of the sound output unit 74, and the process returns to step S15. Note that, if a command to adjust the sound output has not been transmitted in step S17, the process of step S18 is skipped.
[0142] That is, if no peak is detected from the sound output, the process of emitting sound is repeated until a peak is detected. At this time, when a command instructing adjustment of the sound output is transmitted from the electronic device 32, the sound generation unit 73 controls the sound output unit 74 based on this command to adjust the sound output so that a peak can be detected from the sound emitted by the electronic device 32.
[0143] The command sent from the electronic device 32 to instruct adjustment of the sound output will be described in detail later.
[0144] <Sound collection processing by the electronic device 32 in FIG. 3> Next, the sound pickup process by the electronic device 32 of FIG. 3 will be described with reference to the flowchart of FIG.
[0145] In step S31, the audio input unit 51, which is made up of a microphone, collects audio and outputs the collected audio to the known music sound source removal unit 91 and the spatial transfer characteristic calculation unit 92.
[0146] In step S32, the spatial transfer characteristic calculation unit 92 calculates spatial transfer characteristics based on the audio supplied from the audio input unit 51, the characteristics of the audio input unit 51, and the characteristics of the audio output unit 74 of the audio output block 31, and outputs the calculated spatial transfer characteristics to the known music sound source removal unit 91.
[0147] In step S33, the known music sound source removal unit 91 generates an inverted phase signal of the known music sound source in consideration of the spatial transfer characteristic supplied from the spatial transfer characteristic calculation unit 92, removes the component of the known music sound source from the audio supplied from the audio input unit 51, and outputs the result to the arrival time calculation unit 93 and the peak power detection unit 94.
[0148] In step S34, the inverse shift processing unit 130 of the arrival time calculation unit 93 inversely shifts the frequency band of the spreading code signal obtained by removing the known music sound source from the sound input by the sound input unit 51 and supplied from the known music sound source removal unit 91, as described with reference to the right part of FIG. 11 .
[0149] In step S35, the cross-correlation calculation unit 131 calculates the cross-correlation between the spreading code signal, in which the frequency band has been reverse-shifted and the known music sound source has been removed from the audio input by the audio input unit 51, and the spreading code signal of the audio output from the audio output block 31, by calculation using the above-mentioned equations (1) to (4).
[0150] In step S36, the peak detection unit 132 detects peaks in the calculated cross-correlation.
[0151] In step S37, the peak power detection unit 94 detects the power of the frequency band component of the spread code signal at the timing when the cross-correlation corresponding to the distance to each of the detected audio output blocks 31-1 and 31-2 reaches its peak as the peak power, and outputs it to the position calculation unit 95.
[0152] In step S38, the control unit 42 determines whether or not the arrival time calculation unit 93 has detected peaks of the cross-correlation corresponding to the distances to the audio output blocks 31-1 and 31-2.
[0153] If it is determined in step S38 that a cross-correlation peak has not been detected, the process proceeds to step S39.
[0154] In step S39, the control unit 42 controls the communication unit 44 to notify the electronic device 32 that a cross-correlation peak has not been detected.
[0155] In step S40, the control unit 42 determines whether or not either of the peak powers of the audio output blocks 31-1 and 31-2 is greater than a predetermined threshold value, i.e., whether or not either of the peak powers corresponding to the distances to the audio output blocks 31-1 and 31-2 is an extremely large value.
[0156] If it is determined in step S40 that any of the peak powers is greater than the predetermined threshold, the process proceeds to step S41.
[0157] In step S41, the control unit 42 controls the communication unit 44 to send a command to the electronic device 32 to adjust the sound output, and the process returns to step S31. Note that, if it is not determined in step S40 that any of the peak powers is greater than the predetermined threshold, the process of step S41 is skipped.
[0158] That is, sound emission is repeated until a cross-correlation peak is detected, and if either peak power is greater than a predetermined threshold, the sound emission output is adjusted. At this time, the sound emission output levels of both the sound output blocks 31-1 and 31-2 are adjusted equally so as not to affect the peak power ratio.
[0159] If it is determined in step S38 that a cross-correlation peak has been detected, the process proceeds to step S342.
[0160] In step S42, the peak power ratio calculation unit 112 of the position calculation unit 95 calculates the ratio of peak powers according to the distances to the audio output blocks 31-1 and 31-2 as a peak power ratio, and outputs the calculated ratio to the position calculation unit 113.
[0161] In step S43, the peak detection unit 132 of the arrival time calculation unit 93 outputs the time detected as the peak in the cross-correlation to the position calculation unit 95 as the arrival time.
[0162] The cross-correlation between the voices output from the voice output blocks 31-1 and 31-2 and the spreading code signals is calculated, and the arrival times corresponding to the voice output blocks 31-1 and 31-2 are found.
[0163] In step S44, the arrival time difference distance calculation unit 111 of the position calculation unit 95 calculates the arrival time difference distance based on the arrival time corresponding to the distance to each of the audio output blocks 31-1 and 31-2, and outputs the calculated distance to the position calculation unit 113.
[0164] In step S45, the position calculation unit 113 configures the input layer 151 described with reference to FIG. 15 based on the arrival time difference distance supplied from the arrival time difference distance calculation unit 111 and the peak power ratio supplied from the peak power ratio calculation unit 112, and masks data that does not satisfy a predetermined condition using the first layer 152a at the top of the hidden layer 152.
[0165] In step S46, the position calculation unit 113 sequentially uses the second layer 152b to the nth layer 152n of the hidden layer 152 described with reference to Figure 15 to calculate the two-dimensional position of the electronic device 32 as the output layer 153 and output it to the control unit 42.
[0166] In step S47, the control unit 42 executes processing based on the obtained two-dimensional position of the electronic device 32, and ends the processing.
[0167] For example, the control unit 42 controls the communication unit 44 to send commands to the audio output blocks 31-1 and 31-2 to control the level and timing of the audio output from the audio output units 74 of the audio output blocks 31-1 and 31-2 so as to realize a sound field based on the desired position of the electronic device 32.
[0168] As a result, in the audio output blocks 31-1 and 31-2, the sound field control unit 83 controls the level and timing of the audio output from the audio output unit 74 based on the command sent from the electronic device 32 so as to realize a sound field corresponding to the position of the user holding the electronic device 32.
[0169] By performing such processing, the user wearing the electronic device 32 can listen to music output from the audio output blocks 31-1 and 31-2 in an appropriate sound field that corresponds to the user's movements in real time.
[0170] As described above, it is possible to find the position of the electronic device 32 relative to the audio output blocks 31-1 and 31-2 using only two audio output blocks 31-1 and 31-2 that constitute a typical stereo speaker and the electronic device 32 (audio input block 41).
[0171] In addition, at this time, a spread code signal is emitted using a sound band that is difficult for humans to hear, and the distance between the sound output block 31 and the electronic device 32 equipped with the sound input block 41 is measured, thereby making it possible to determine the position of the electronic device 32 in real time.
[0172] Furthermore, the speaker constituting the audio output section 74 of the audio output block 31 of the present disclosure and the microphone constituting the audio input section 51 of the electronic device 32 can be made from existing audio equipment, for example, one with two speakers, so they can be implemented at low cost and the effort involved in installation can be simplified.
[0173] Furthermore, since it uses sound and can utilize existing audio equipment, no certification or other licenses that are required when using radio waves, etc., are required, which also simplifies the costs and effort involved in use.
[0174] Furthermore, it is possible to measure the position of a user who carries or wears electronic device 32 in real time while allowing the user to enjoy music or the like by playing a known music sound source without having to listen to or see sounds that the user finds unpleasant.
[0175] In the above, an example has been described in which the two-dimensional position of the electronic device 32 is determined by a learning device formed by machine learning based on the time difference of arrival distance and the peak power ratio, but the two-dimensional position may also be determined based on an input consisting of only the time difference of arrival distance or only the peak power ratio.
[0176] However, when two-dimensional position is calculated based on input consisting of only the time difference of arrival distance or only the peak power ratio, accuracy decreases, so the method of use may be devised or restricted.
[0177] For example, if the input is formed only from the time difference of arrival distance, it is expected that the accuracy of the position in the y direction will be low, so only the position in the x direction may be used.
[0178] Furthermore, for example, when the input is formed only from the peak power ratio, it is expected that accuracy will decrease if the input is moved away from the audio output block 31 by more than a predetermined distance, so it may be possible to limit the use of only two-dimensional positions within a range within a predetermined distance from the audio output block 31.
[0179] Furthermore, in the above, information when a cross-correlation peak is detected is used as the input layer, but information when a cross-correlation peak cannot be detected may also be used as the input layer. By doing so, for example, when a position is very close to one audio output block 31, the audio signal from the other audio output block 31 cannot be received and a peak may not be detected, and even in such cases, it is possible to appropriately identify the two-dimensional position.
[0180] Furthermore, although the above has described an example of determining the two-dimensional position of the electronic device 32 (audio input block 41) relative to the audio output blocks 31-1 and 31-2, if the position of only the electronic device 32 or the audio output blocks 31-1 and 31-2 is known, it is also possible to determine the position of the audio output block 31, whose position is unknown, by using similar processing.
[0181] <<2. Application Example of the First Embodiment>> <Example of using the distance between audio output blocks as the input layer> The above has described an example in which the technology of the present disclosure is applied to a home audio system consisting of two audio output blocks 31-1, 31-2 and an electronic device 32, in which the position of the electronic device 32 equipped with an audio input block 41 is determined in real time, and the audio output from the audio output block 31 is controlled based on the position of the electronic device 32, thereby realizing an appropriate sound field.
[0182] However, in the above description, an example has been described in which the positions of the audio output blocks 31-1 and 31-2 are known, or the position of at least one of them is known, and the distance between them is also known.
[0183] For example, in addition to the information of input layer 151 consisting of time difference of arrival distance D and peak power ratio PR, input layer 151α consisting of mutual distance SD may be configured and input to hidden layer 152.
[0184] For example, as shown in FIG. 18, in the position calculation unit 113, in the hidden layer 152′, masked data that does not satisfy a predetermined condition between the time difference of arrival distance D and the peak power ratio PR constituting the first layer 152a and an input layer 151α consisting of the mutual distance SD are input to the second layer 152′b, and the position of the electronic device 32 may be obtained as the output layer 153′ by processing the second layer 152′b to the n-th layer 152′n.
[0185] With this configuration, even if the distance between the audio output blocks 31-1 and 31-2 changes in various ways, it is possible to find the position of the electronic device 32 from the two audio output blocks 31-1 and 31-2 and one electronic device 32 (microphone).
[0186] <<3. Second Embodiment>> The above has described an example in which, using audio output blocks 31-1 and 31-2 located at known positions and electronic device 32, the position of electronic device 32 is identified using the arrival time difference distance D and peak power ratio PR when audio emitted from audio output blocks 31-1 and 31-2 is picked up by electronic device 32.
[0187] Incidentally, the positions in the x direction, which is perpendicular to the sound output direction of the sound output blocks 31-1 and 31-2, can be determined with a relatively high degree of accuracy using the above-described method, as described with reference to FIG.
[0188] On the other hand, as described with reference to FIG. 14, the position in the y direction, which is the sound output direction of the audio output blocks 31-1 and 31-2, is slightly less accurate than the position in the x direction.
[0189] Here, as shown in Figure 1, when audio output blocks 31-1 and 31-2 are provided and TV 30 is provided at approximately the center of audio output blocks 31-1 and 31-2, it is generally expected that a user will listen to the audio emitted from audio output blocks 31-1 and 31-2 while watching TV 30.
[0190] At this time, the sounds emitted from the sound output blocks 31-1 and 31-2 that the user listens to are evaluated for their listening ability within a range defined by an angle set with the center position of the TV 30 as the reference.
[0191] For example, as shown in FIG. 19, consider a case where audio output blocks 31-1 and 31-2 are provided and TV 30 is set at approximately the center of audio output blocks 31-1 and 31-2 with its display surface parallel to the line connecting audio output blocks 31-1 and 31-2.
[0192] In this case, when user H1 is positioned directly opposite TV 30, the range of angle α based on the center position of TV 30 in the figure is considered to be the audible range, and the closer the position within this range is to the dotted line in the figure, the better the audible range.
[0193] Furthermore, when user H2 is present in front of TV 30, the range of angle β based on the center position of TV 30 in the figure is considered to be the audible range, and the closer the position within this range is to the dotted line in the figure, the better the audible range.
[0194] In other words, when audio output blocks 31-1 and 31-2 are provided and TV 30 is provided at approximately the center of audio output blocks 31-1 and 31-2, the user's viewing position can be determined with a certain degree of accuracy in the x-direction, which is perpendicular to the sound emission direction of the above-mentioned audio output blocks 31-1 and 31-2, and thereby better viewing and listening than a predetermined level can be achieved.
[0195] In other words, when audio output blocks 31-1 and 31-2 are provided and TV 30 is provided at approximately the center of audio output blocks 31-1 and 31-2, even if the user's viewing position in the y direction, which is the sound emission direction of the above-mentioned audio output blocks 31-1 and 31-2, is not determined with a specified accuracy, as long as the position in the x direction is determined with a specified accuracy, viewing and listening quality better than a specified level can be achieved.
[0196] As a result, as described above, even if the positional accuracy in the y direction is lower than a predetermined level, as long as the positional accuracy in the x direction is at or above a predetermined level, good visual and hearing quality at or above a predetermined level can be achieved.
[0197] However, as shown in Figure 20, when audio output blocks 31-1 and 31-2 are provided and TV 30 is provided at a position shifted from approximately the center of audio output blocks 31-1 and 31-2, there is a risk that an appropriate sound field will not be realized due to the audio being emitted for user H11 who is positioned directly opposite approximately the center of audio output blocks 31-1 and 31-2, because the viewing position for TV 30 is shifted from the center of audio output blocks 31-1 and 31-2.
[0198] Therefore, in this case, in addition to the position of the electronic device 32 in the x direction relative to the audio output blocks 31-1 and 31-2, the position in the y direction also needs to be determined with a precision higher than a predetermined precision.
[0199] <Peak power frequency component ratio> Therefore, in addition to the arrival time difference distance D and the peak power ratio PR when the sound emitted from the sound output blocks 31-1 and 31-2 is picked up by the electronic device 32, a hidden layer may be formed by machine learning using the frequency component ratio between the peak power of the cross-correlation of the high frequency components of each of the sound output blocks 31-1 and 31-2 and the peak power of the cross-correlation of the low frequency components, thereby obtaining an output layer consisting of the position of the electronic device 32.
[0200] For example, the peak power frequency component ratio FR (=HP / LP), obtained from the peak power LP, which is the power at the peak of the cross-correlation of the low-frequency components (e.g., 18 to 21 kHz) of the sound emitted from the sound output block 31, and the peak power HP, which is the power at the peak of the cross-correlation of the high-frequency components (e.g., 21 to 24 kHz), is distributed in the x and y directions when the sound output block 31 is positioned at the center, as shown in FIG.
[0201] That is, as shown by the distribution of peak power frequency component ratio FR (=HP / LP) in Figure 21, with respect to the y direction, which is the sound emission direction of the audio output block 31, the correlation with the distance in the y direction is higher in the range directly facing the audio output block 31, so it is possible to identify the position in the y direction with high accuracy.
[0202] However, as the distance from the audio output block 31 in the x-axis direction and the y-axis direction increases or the angle increases, i.e., as the distance from the audio output block 31 increases or the angle increases relative to the sound emission direction, the peak power HP of the high-frequency component attenuates, and the peak power frequency component ratio FR (=HP / LP) decreases, for example, as shown in the ranges Z1 and Z2 in Figure 21.
[0203] Therefore, the accuracy of the position in the y direction needs to be determined based on the distance from the audio output block 31, so for example, the value for the position in the y direction up to a predetermined distance from the audio output block 31 may be used.
[0204] Note that Figure 21 shows an example of the component ratio FR of one audio output block 31, so when two audio output blocks 31-1 and 31-2 are used, the respective peak power frequency component ratios FRL and FRR are used in addition to the arrival time difference distance D and the peak power ratio PR for processing in the hidden layer, making it possible to determine the x-direction and y-direction positions of the electronic device 32 with high accuracy.
[0205] <Example of the configuration of an electronic device when the peak power frequency component ratio is used as the input layer> Next, referring to FIG. 22, an example configuration of an electronic device 32 will be described in which a peak power frequency component ratio FR (=HP / LP) obtained from the peak power LP of the cross-correlation of the low frequency components (e.g., 18 to 21 kHz) of the sound emitted from the sound output block 31 and the peak power HP of the cross-correlation of the high frequency components (e.g., 21 to 24 kHz) is newly added and used in the input layer.
[0206] In the electronic device 32 of FIG. 22, components having the same functions as those of the electronic device 32 of FIG. 3 are denoted by the same reference numerals, and the description thereof will be omitted as appropriate.
[0207] The electronic device 32 in FIG. 22 differs from the electronic device 32 in FIG. 3 in that a peak power frequency component ratio calculation unit 201 is newly provided, and a position calculation unit 202 is provided instead of the position calculation unit 95.
[0208] The peak power frequency component ratio calculation unit 201 performs the same processing as that performed by the arrival time calculation unit 93 to find the cross-correlation peak in each of the low frequency bands (e.g., 18 to 21 kHz) and high frequency bands (e.g., 21 to 24 kHz) of the audio emitted from each of the audio output blocks 31-1 and 31-2, and finds the peak power LP of each low frequency band and the peak power HP of each high frequency band.
[0209] Then, the peak power frequency component ratio calculation unit 201 calculates peak power frequency component ratios FRR, FRL using the peak power HP of the high frequency band of each of the audio output blocks 31-1, 31-2 as the numerator and the peak power LP of the low frequency band as the denominator, and outputs these to the position calculation unit 202. That is, the peak power frequency component ratio calculation unit 201 calculates the ratio of the peak power HP of the high frequency band to the peak power LP of the low frequency band of each of the audio output blocks 31-1, 31-2 as the peak power frequency component ratios FRR, FRL, and outputs these to the position calculation unit 202.
[0210] The position calculation unit 202 calculates the position (x, y) of the electronic device 32 using a neural network formed by machine learning based on the arrival time, peak power, and peak power frequency component ratios of the audio output blocks 31-1 and 31-2.
[0211] <Configuration example of the position calculation unit in Fig. 22> Next, a configuration example of the position calculation unit 202 of the electronic device 32 in Fig. 22 will be described with reference to Fig. 23. In the position calculation unit 202 in Fig. 23, components having the same functions as those of the position calculation unit 95 in Fig. 4 are denoted by the same reference numerals, and descriptions thereof will be omitted as appropriate.
[0212] The position calculation unit 202 in FIG. 23 differs from the position calculation unit 95 in FIG. 4 in that a position calculation unit 211 is provided instead of the position calculation unit 113.
[0213] The basic function of the position calculation unit 211 is the same as that of the position calculation unit 113, but the position calculation unit 113 functions as a hidden layer with the arrival time difference distance and the peak power ratio as the input layer, and obtains the position of the electronic device 32 as the output layer.
[0214] In contrast, the position calculation unit 211 functions as a hidden layer for an input layer that includes the arrival time difference distance D, the peak power ratio PR, the peak power frequency component ratios FRR, FRL of the audio output blocks 31-1, 31-2, and the mutual distance DS between the audio output blocks 31-1, 31-2, and calculates the position of the electronic device 32 as an output layer.
[0215] More specifically, as shown in FIG. 24, the input layer 221 in the position calculation unit 211 is composed of the arrival time difference distance D, the peak power ratio PR, and the peak power frequency component ratios FRR and FRL of the audio output blocks 31-1 and 31-2, respectively, and further, the input layer 221α is composed of the mutual distance DS between the audio output blocks 31-1 and 31-2.
[0216] Then, the position calculation unit 211 functions as a hidden layer 222 configured from a neural network consisting of a first layer 222a to an n-th layer 222n, and an output layer 223 consisting of the position (xy) of the electronic device 32 is obtained.
[0217] 15. The input layer 221, hidden layer 222, and output layer 223 in FIG. 24 correspond to the input layer 151, hidden layer 152, and output layer 153 in FIG.
[0218] <Sound collection processing by the electronic device 32 in FIG. 3> Next, the sound pickup process by the electronic device 32 of Fig. 22 will be described with reference to the flowchart of Fig. 25. Note that the sound emission process is the same as the process of Fig. 16, and therefore description thereof will be omitted.
[0219] The processes in steps S101 to S112, S114 to S116, and S118 in the flowchart of FIG. 25 are similar to the processes in steps S31 to S45 and S47 in the flowchart of FIG. 17, and therefore will not be described again.
[0220] That is, once the cross-correlation peak is detected and the peak power ratio is calculated in steps S101 to S112, the process proceeds to step S113.
[0221] In step S113, the peak power frequency component ratio calculation unit 201 finds peaks based on cross-correlation in the low frequency band (e.g., 18 to 21 kHz) and the high frequency band (e.g., 21 to 24 kHz) of the audio output from each of the audio output blocks 31-1 and 31-2.
[0222] Then, the peak power frequency component ratio calculation unit 201 obtains the peak power LP of the low frequency band and the peak power HP of the high frequency band, calculates the ratio of the peak power LP of the low frequency band to the peak power HP of the high frequency band as peak power frequency component ratios FRR, FRL, and outputs them to the position calculation unit 202.
[0223] In steps S114 to S116, the arrival times are calculated, the arrival time difference distances are calculated, and data that do not satisfy the predetermined conditions is masked.
[0224] In step S117, the position calculation unit 202 sequentially executes processing using the second layer 222b to the n-th layer 222n of the hidden layer 222 on the input layer 221 consisting of the arrival time difference distance D, the peak power ratio PR, and the peak power frequency component ratios FRR and FRL of the audio output blocks 31-1 and 31-2 described with reference to FIG. 24, and the input layer 221α consisting of the mutual distance DS, to calculate the position (two-dimensional position (x, y)) of the electronic device 32 as the output layer 223, and outputs the calculated position to the control unit 42.
[0225] In step S118, the control unit 42 executes processing based on the determined position of the electronic device 32, and then ends the processing.
[0226] As described above, with only two audio output blocks 31-1 and 31-2 that constitute a typical stereo speaker and the electronic device 32 (audio input block 41), and even if the TV 30 is positioned away from the center position between the audio output blocks 31-1 and 31-2, it is possible to determine the position of the electronic device 32 relative to the audio output blocks 31-1 and 31-2 with high accuracy in the x and y directions.
[0227] <<4. Third Embodiment>> In the above, an example has been described in which an input layer 221 is formed from the arrival time difference distance D, the peak power ratio PR, and the peak power frequency component ratios FRR and FRL of the audio output blocks 31-1 and 31-2, and processing is performed by a hidden layer 222 consisting of a neural network formed by machine learning, thereby determining the position (x, y) of the electronic device 32 (audio input block 41) as the output layer 223.
[0228] However, even if the input layer is configured with two audio output blocks 31 and electronic device 32 (audio input block 41) using the arrival time difference distance D and peak power ratio PR, the x-direction position of electronic device 32 relative to audio output blocks 31-1 and 31-2, which is perpendicular to the sound emission direction of audio output block 31, can be determined with relatively high accuracy, but the accuracy of the y-direction position is somewhat inferior.
[0229] Therefore, an IMU may be provided in the electronic device 32 to detect the posture when the electronic device 32 is tilted toward each of the voice output blocks 31-1 and 31-2, and the angle θ between the voice output blocks 31-1 and 31-2 with respect to the electronic device 32 may be obtained to accurately obtain the position in the y direction.
[0230] That is, an IMU is mounted on the electronic device 32. As shown in FIG. 26, for example, the posture when the upper end portion of the electronic device 32 in the figure is directed toward each of the voice output blocks 31-1 and 31-2 is detected, and the angle θ formed by the direction to each of the voice output blocks 31-1 and 31-2 from the detected posture change is obtained.
[0231] At this time, for example, the known positions of the voice output blocks 31-1 and 31-2 are represented by (a1, b1) and (a2, b2), and the position of the electronic device 32 is represented by (x, y). Note that x is a known value obtained by constructing the input layer with the arrival time difference distance D and the peak power ratio PR.
[0232] The vector A1 to the voice output block 31-1 with respect to the position of the electronic device 32 is represented by (a1 - x, b1 - y). Similarly, the vector A2 to the voice output block 31-2 is represented by (x - a2, y - b2).
[0233] Here, the inner product (A1, A2) of the vectors A1 and A2 is expressed by the relational expression (A1, A2)=|A1|·|A2|cosθ. As described above, since the vectors A1 and A2 are known values except for y, the value of y may be obtained by solving for y from this relational expression of the inner product.
[0234] <Configuration example of an electronic device provided with an IMU> Next, referring to FIG. 27, a configuration example of the electronic device 32 in the case of providing an IMU to obtain the angle θ formed by the voice output blocks 31-1 and 31-2 with respect to the electronic device 32 and obtaining the value of y using the relational expression of the inner product from the angle θ will be described.
[0235] In the electronic device 32 of FIG. 27, components having the same functions as those of the electronic device 32 of FIG. 3 are denoted by the same reference numerals, and the description thereof will be omitted as appropriate.
[0236] The electronic device 32 in FIG. 27 differs from the electronic device 32 in FIG. 3 in that an IMU (Inertial Measurement Unit) 230 and an attitude calculation unit 231 are newly provided, and a position calculation unit 232 is provided instead of the position calculation unit 95.
[0237] The IMU 230 detects angular velocity and acceleration and outputs them to the attitude calculation unit 231 .
[0238] The attitude calculation unit 231 calculates the attitude of the electronic device 32 based on the angular velocity and acceleration supplied from the IMU 230, and outputs the calculated attitude to the position calculation unit 232. Note that, since roll and pitch can always be obtained from the direction of gravity, only yaw on the xy plane will be considered here among roll, pitch, and yaw that can be obtained as attitude.
[0239] The position calculation unit 232 basically has the same functions as the position calculation unit 95, and determines the position of the electronic device 32 in the x direction, acquires orientation information supplied from the orientation calculation unit 231 described above, determines the angle θ between the electronic device 32 and the audio output blocks 31-1 and 31-2 based on the electronic device 32 described with reference to Figure 26, and determines the position of the electronic device 32 in the y direction from the relational equation for the dot product.
[0240] At this time, the control unit 42 controls the output unit 43 consisting of a speaker and a display to instruct the user to point a specific part of the electronic device 32 toward each of the audio output blocks 31-1 and 31-2 as necessary, and the position calculation unit 232 calculates the angle θ based on the posture (direction) at that time.
[0241] <Configuration example of the position calculation unit in Fig. 27> Next, a configuration example of the position calculation unit 232 of the electronic device 32 in Fig. 22 will be described with reference to Fig. 28. In the position calculation unit 232 in Fig. 28, components having the same functions as those of the position calculation unit 95 in Fig. 4 are denoted by the same reference numerals, and descriptions thereof will be omitted as appropriate.
[0242] The position calculation unit 232 in FIG. 28 differs from the position calculation unit 95 in FIG. 4 in that a position calculation unit 241 is provided instead of the position calculation unit 113.
[0243] The basic function of the position calculation unit 241 is the same as that of the position calculation unit 113, but the position calculation unit 113 functions as a hidden layer with the arrival time difference distance and the peak power ratio as the input layer, and obtains the position of the electronic device 32 as the output layer.
[0244] In contrast, the position calculation unit 241 uses only the position in the x direction as the position of the electronic device 32, which is the output layer for which the arrival time difference distance and the peak power ratio are calculated as the input layer. The position calculation unit 241 also acquires the information on the attitude supplied from the attitude calculation unit 231, calculates the angle θ described with reference to Fig. 26, and calculates the position in the y direction of the electronic device 32 from the relational expression between the information on the position in the x direction and the dot product.
[0245] <Sound collection processing by the electronic device 32 in FIG. 27> Next, the sound pickup process by the electronic device 32 of Fig. 27 will be described with reference to the flowchart of Fig. 29. Note that the sound emission process is the same as the process of Fig. 16, and therefore description thereof will be omitted.
[0246] Here, the processing of step S151 in the flowchart of Fig. 25 is the processing of steps S31 to S46 in the flowchart of Fig. 17, and of the position information of the electronic device 32 obtained here, only the position information in the x direction is used, and the description thereof will be omitted. Also, the audio output blocks 31-1 and 31-2 will also be referred to as the left audio output block 31-1 and the right audio output block 31-2, respectively.
[0247] Once the position in the x direction is determined by the processing of step S151, in step S152 the control unit 42 controls the output unit 43 consisting of a touch panel to display an image requesting a tap operation, for example, with the top end of the electronic device 32 facing the left audio output block 31-1.
[0248] In step S153, the control unit 42 controls the output unit 43 to determine whether or not a tap has been performed, and repeats the same process until it is determined that a tap operation has been performed.
[0249] In step S153, for example, if the user performs a tap operation while pointing the top end of the electronic device 32 toward the left audio output block 31-1, it is deemed that a tap operation has been performed, and the process proceeds to step S154.
[0250] In step S154, upon acquiring the acceleration and angular velocity information supplied from the IMU 230, the attitude calculation unit 231 converts it into attitude information and outputs it to the position calculation unit 232. In response to this, the position calculation unit 232 stores the attitude (direction) in a state facing the direction of the left audio output block 31-1.
[0251] In step S155, the control unit 42 controls the output unit 43, which is made up of a touch panel, to display, for example, an image requesting a tap operation with the top end of the electronic device 32 facing the right audio output block 31-2.
[0252] In step S156, the control unit 42 controls the output unit 43 to determine whether or not a tap has been performed, and repeats the same process until it is determined that a tap operation has been performed.
[0253] In step S156, for example, if the user performs a tap operation while pointing the top end of the electronic device 32 toward the right audio output block 31-2, it is deemed that a tap operation has been performed, and the process proceeds to step S157.
[0254] In step S157, upon acquiring the acceleration and angular velocity information supplied from the IMU 230, the attitude calculation unit 231 converts it into attitude information and outputs it to the position calculation unit 232. In response to this, the position calculation unit 232 stores the attitude (direction) in a state facing the direction of the right audio output block 31-2.
[0255] In step S158, the position calculation unit 241 of the position calculation unit 232 calculates the angle θ between the left and right audio output blocks 31-1 and 31-2 based on the electronic device 32 from the stored information on the posture (direction) of the audio output blocks 31-1 and 31-2 as they face each direction.
[0256] In step S159, the position calculation unit 241 calculates the y-direction position of the electronic device 32 from the inner product relational equation based on the known positions of the audio output blocks 31-1 and 31-2, the x-direction position of the electronic device 32, and the angle θ between the left and right audio output blocks 31-1 and 31-2 relative to the electronic device 32.
[0257] In step S160, the position calculation unit 241 determines whether the position in the y direction of the electronic device 32 has been appropriately determined based on whether the determined value of the y direction position is an extremely large value or an extremely small value, such as larger or smaller than a predetermined value.
[0258] If it is determined in step S160 that the position in the y direction has not been properly determined, the process returns to step S152.
[0259] That is, the processes of steps S152 to S160 are repeated until the position of the electronic device 32 in the y direction is properly determined.
[0260] Then, if it is determined in step S160 that the position in the y direction has been properly determined, the process proceeds to step S161.
[0261] In step S161, the control unit 42 controls the output unit 43, which is made up of a touch panel, to display, for example, an image requesting a tap operation with the top end of the electronic device 32 facing the TV 30.
[0262] In step S162, the control unit 42 controls the output unit 43 to determine whether or not a tap has been performed, and repeats the same process until it is determined that a tap operation has been performed.
[0263] In step S162, for example, if the user performs a tap operation while pointing the top end of the electronic device 32 toward the TV 30, it is deemed that a tap operation has been performed, and the process proceeds to step S163.
[0264] In step S163, upon acquiring the acceleration and angular velocity information supplied from the IMU 230, the attitude calculation unit 231 converts the information into attitude information and outputs it to the position calculation unit 232. In response to this, the position calculation unit 232 stores the attitude (direction) in a state facing the direction of the TV 30.
[0265] In step S164, the control unit 42 executes processing based on the obtained position of the electronic device 32, the attitude (direction) of the TV 30 from the electronic device 32, and the known positions of the audio output blocks 31-1 and 31-2, and then ends the processing.
[0266] For example, the control unit 42 controls the communication unit 44 to send commands to the audio output blocks 31-1 and 31-2 to control the level and timing of the audio output from the audio output units 74 of the audio output blocks 31-1 and 31-2 so as to realize an appropriate sound field based on the determined position of the electronic device 32, the direction of the TV 30, and the known positions of the audio output blocks 31-1 and 31-2.
[0267] By the above processing, an IMU is provided in the electronic device 32, the attitude of the electronic device 32 when tilted toward each of the audio output blocks 31-1 and 31-2 is detected, and the angle θ between the audio output blocks 31-1 and 31-2 with the electronic device 32 as the reference is calculated, thereby making it possible to calculate the position in the y direction with high accuracy from the relational expression for the inner product. As a result, it becomes possible to measure the position of the electronic device 32 with high accuracy using the audio output blocks 31-1 and 31-2, which are composed of two speakers or the like, and the electronic device 32 (audio input block 41).
[0268] <<5. Application Example of the Third Embodiment>> In the above, an example has been described in which an IMU is provided in the electronic device 32, and the input layer is used with two audio output blocks 31 and the electronic device 32 to determine the x-direction position of the electronic device 32 relative to the audio output blocks 31-1 and 31-2, which is perpendicular to the sound emission direction of the audio output block 31, using the arrival time difference distance D and peak power ratio PR, and further the attitude of the electronic device 32 when tilted toward each of the audio output blocks 31-1 and 31-2 is detected, and the angle θ between the audio output blocks 31-1 and 31-2 with the electronic device 32 as the reference is determined, thereby determining the y-direction position from the inner product relational equation.
[0269] However, other configurations are possible as long as the positional relationship and direction between the electronic device 32 and the audio output blocks 31-1 and 31-2 are known. For example, in addition to the IMU, two audio input units 51 each consisting of a microphone or the like may be provided, and the two microphones may be used to recognize the positional relationship and direction between the electronic device 32 and the audio output blocks 31-1 and 31-2.
[0270] That is, as shown in FIG. 30, an audio input unit 51-1 may be provided at the upper end of the electronic device 32, and an audio input unit 51-2 may be provided at the lower end.
[0271] Here, it is assumed that the audio input units 51-1 and 51-2 are located at the center of the electronic device 32 as indicated by the dashed-dotted line, and that the distance between them is L. In this case, it is assumed that the TV 30 is located on the center line of the electronic device 32 indicated by the dashed-dotted line in the drawing, on the straight line connecting the audio input units 51-1 and 51-2, and that the angle formed with the sound output direction of the audio output blocks 31-1 and 31-2 is θ.
[0272] In addition, in the case of FIG. 30, if the position of the voice input unit 51-1 is expressed as (x, y), the position of the voice input unit 51-2 is expressed as (x+L sin θ, y+L cos θ).
[0273] Here, assuming that the mutual distance L between the audio input units 51-1 and 51-2 is known and the three parameters of the position (x, y) and angle θ of the electronic device 32 (of its audio input unit 51-1) are unknown, it is possible to determine the position (x, y) and angle θ of the electronic device 32 (of its audio input unit 51-1) using simultaneous equations consisting of the arrival time distance between the audio output block 31-1 and the audio input unit 51-1, the arrival time distance between the audio output block 31-1 and the audio input unit 51-2, the arrival time distance between the audio output block 31-2 and the audio input unit 51-1, and the arrival time distance between the audio output block 31-2 and the audio input unit 51-2, as well as the coordinates of the known positions of the audio input unit 51-1 and the audio output blocks 31-1 and 31-2.
[0274] Furthermore, when the four parameters of the distance L, the position (x, y) of the electronic device 32 (of its voice input unit 51-1), and the angle θ are unknown, the position of the electronic device 32 in the x direction can be determined as the output layer by configuring the input layer with the arrival time difference distance D and the peak power ratio PR, as described above, and processing is performed by a hidden layer consisting of a neural network formed by machine learning.
[0275] Furthermore, since the x-direction position of the audio input unit 51-1 of the electronic device 32 is known, the unknown distance L, the y-direction position of the electronic device 32 (of its audio input unit 51-1), and the angle θ can be determined using simultaneous equations consisting of the arrival time distance between the audio output block 31-1 and the audio input unit 51-1, the arrival time distance between the audio output block 31-1 and the audio input unit 51-2, the arrival time distance between the audio output block 31-2 and the audio input unit 51-1, and the arrival time distance between the audio output block 31-2 and the audio input unit 51-2, as well as the x-direction position of the audio input unit 51-1 and the coordinates of the known positions of the audio output blocks 31-1 and 31-2.
[0276] That is, as shown in FIG. 30, when two audio input units 51-1 and 51-2 are provided, the distance L is known, and three parameters, namely the position (x, y) and angle θ of the electronic device 32 (of its audio input unit 51-1), are unknown, it is possible to determine the unknown parameters by an analytical method using simultaneous equations.
[0277] On the other hand, when the four parameters of the distance L, the position (x, y) of the electronic device 32 (of its voice input unit 51-1), and the angle θ are unknown, the position in the x direction can be determined using a method that uses a neural network formed by machine learning, and then the remaining parameters can be determined analytically using simultaneous equations.
[0278] This makes it possible to determine the two-dimensional position of the electronic device 32 (audio input block 41) using two speakers (audio output blocks 31-1, 31-2) and one electronic device 32 (audio input block 41) in various types of electronic devices 32, regardless of whether the distance L is known or not, thereby enabling the setting of an appropriate sound field.
[0279] <<6. Example of execution by software>> The above-described series of processes can be executed by hardware, but can also be executed by software. When the series of processes are executed by software, the programs constituting the software are installed from a recording medium into a computer incorporated in dedicated hardware, or into, for example, a general-purpose computer that can execute various functions by installing various programs.
[0280] 31 shows an example of the configuration of a general-purpose computer. This computer has a built-in CPU (Central Processing Unit) 1001. An input / output interface 1005 is connected to the CPU 1001 via a bus 1004. A ROM (Read Only Memory) 1002 and a RAM (Random Access Memory) 1003 are connected to the bus 1004.
[0281] Connected to the input / output interface 1005 are an input unit 1006 including input devices such as a keyboard and a mouse through which a user inputs operation commands, an output unit 1007 that outputs a processing operation screen and images of processing results to a display device, a storage unit 1008 including a hard disk drive or the like that stores programs and various data, and a communication unit 1009 including a LAN (Local Area Network) adapter or the like that executes communication processing via a network typified by the Internet. Also connected to the input / output interface 1005 is a drive 1010 that reads and writes data from / to removable storage media 1011 such as a magnetic disk (including a flexible disk), an optical disk (including a CD-ROM (Compact Disc-Read Only Memory) and a DVD (Digital Versatile Disc)), a magneto-optical disk (including an MD (Mini Disc)), or a semiconductor memory.
[0282] The CPU 1001 executes various processes in accordance with a program stored in a ROM 1002 or a program read from a removable storage medium 1011 such as a magnetic disk, optical disk, magneto-optical disk, or semiconductor memory, installed in a storage unit 1008, and loaded from the storage unit 1008 into a RAM 1003. The RAM 1003 also stores data necessary for the CPU 1001 to execute various processes as appropriate.
[0283] In a computer configured as described above, the CPU 1001 performs the above-described series of processes by, for example, loading a program stored in the memory unit 1008 into the RAM 1003 via the input / output interface 1005 and the bus 1004 and executing it.
[0284] The program executed by the computer (CPU 1001) can be provided by being recorded on a removable storage medium 1011 such as a package medium, for example. The program can also be provided via a wired or wireless transmission medium such as a local area network, the Internet, or digital satellite broadcasting.
[0285] In a computer, a program can be installed in the storage unit 1008 via the input / output interface 1005 by inserting a removable storage medium 1011 into the drive 1010. The program can also be received by the communication unit 1009 via a wired or wireless transmission medium and installed in the storage unit 1008. Alternatively, the program can be installed in the ROM 1002 or the storage unit 1008 in advance.
[0286] The program executed by the computer may be a program that processes in chronological order according to the order described in this specification, or may be a program that processes in parallel or at the required timing, such as when called.
[0287] 31 realizes the functions of the audio output block 31 and the audio input block 41 in FIG.
[0288] In this specification, a system refers to a collection of multiple components (devices, modules (components), etc.), regardless of whether all the components are contained in the same housing. Therefore, multiple devices housed in separate housings and connected via a network, and a single device housed in a single housing with multiple modules, are both systems.
[0289] The embodiments of the present disclosure are not limited to the above-described embodiments, and various modifications are possible within the scope of the gist of the present disclosure.
[0290] For example, the present disclosure can be configured as a cloud computing system in which a single function is shared and processed collaboratively by multiple devices via a network.
[0291] Furthermore, each step described in the above flowchart can be executed by one device, or can be shared and executed by multiple devices.
[0292] Furthermore, when one step includes multiple processes, the multiple processes included in that one step can be executed by one device or can be shared and executed by multiple devices.
[0293] The present disclosure can also be configured as follows. <1> an audio receiving unit for receiving audio signals output from two audio output blocks located at known positions, the audio signals being spread spectrum modulated spread codes; a position calculation unit that calculates a position of the audio receiving unit based on an arrival time difference distance that is a difference in distance specified from an arrival time that is the time it takes for the audio signal of the two audio output blocks to arrive at the audio receiving unit and be received; An information processing device comprising: <2> an arrival time calculation unit that calculates the arrival time of each of the audio signals from the two audio output blocks to the audio receiving unit; and an arrival time difference distance calculation unit that calculates the difference in distance between the two audio output blocks and itself as the arrival time difference distance based on the arrival times of the audio signals of the two audio output blocks and the known positions of the two audio output blocks, wherein the position calculation unit calculates the position of the audio receiving unit as an output layer by processing an input layer consisting of the arrival time difference distances using a hidden layer consisting of a neural network formed by machine learning. <1> The information processing device described in <3> The arrival time calculation unit a cross-correlation calculation unit that calculates a cross-correlation between a spreading code signal in the voice signal received by the voice receiving unit and a spreading code signal of the voice signal output from the two voice output blocks; a peak detection unit that detects a time at which a peak occurs in the cross-correlation as the arrival time, The time difference of arrival distance calculation unit Calculating the difference in distance between the two audio output blocks and the target audio output block based on the arrival times detected by the peak detection unit as the arrival time difference distance. <2> The information processing device described in <4> The position calculation unit calculates the position of the voice receiving unit as an output layer by performing processing using a hidden layer made of a neural network formed by the machine learning on an input layer made of the arrival time difference distance and a peak power ratio, which is a ratio of powers at the timings at which the cross-correlation of the voice signals output from each of the two voice output blocks becomes a peak. <3> The information processing device described in <5> a peak power detection unit that detects, as a peak power, the power of the audio signals output from each of the two audio output blocks at the peak when received by the audio receiving unit; and a peak power ratio calculation unit that calculates a peak power ratio from the ratio of the peak powers of the audio signals output from the two audio output blocks, detected by the peak power detection unit. <4> The information processing device described in <6> The position calculation unit calculates the position of the voice receiving unit as the output layer by performing processing using the hidden layer on an input layer made up of the arrival time difference distance calculated by the arrival time difference distance calculation unit and the peak power ratio calculated by the peak power ratio calculation unit. <5> The information processing device described in <7> The position calculation unit calculates the position of the voice receiving unit as an output layer by performing processing using a hidden layer made of a neural network formed by the machine learning on an input layer made of the arrival time difference distance, the peak power ratio of the voice signals output from each of the two voice output blocks, and a peak power frequency component ratio, which is the ratio of peak power of low frequency components to peak power of high frequency components of the voice signals output from each of the two voice output blocks. <4> The information processing device described in <8> The audio signal processing device further includes a peak power frequency component ratio calculation unit that detects peak power of the low frequency component and peak power of the high frequency component of the audio signal output from each of the two audio output blocks at the peak, and calculates the ratio of the peak power of the high frequency component to the peak power of the low frequency component as the peak power frequency component ratio. <7> The information processing device described in <9> The position calculation unit calculating, by the machine learning, positions in a direction perpendicular to a sound output direction of the audio signals in the audio output blocks of the audio receiving unit based on the arrival time difference distances of the audio signals in the two audio output blocks received by the audio receiving unit; Calculating the position of the sound output direction of the audio signal in the audio output block of the audio receiving unit based on the angle formed by the directions of the two audio output blocks with the audio receiving unit as a reference <2> The information processing device described in <10> an IMU (Inertial Measurement Unit) that detects angular velocity and acceleration of the audio receiving unit; a posture detection unit that detects the posture of the robot itself based on the angular velocity and the acceleration, The position calculation unit calculating an angle formed by the directions of the two audio output blocks with the audio receiving unit as a reference based on the orientation of the device itself detected by the orientation detection unit; Calculating the position of the sound output direction of the audio signal of the audio output block of the audio receiving unit based on the calculated angle formed by the directions of the two audio output blocks with the audio receiving unit as a reference <9> The information processing device described in <11> The position calculation unit calculates an angle formed by the directions of the two audio output blocks with the audio receiving unit as a reference, based on the orientation of the device when facing the two audio output blocks, detected by the orientation detection unit. <10> The information processing device described in <12> The position calculation unit calculates the position of the audio output block of the audio receiving unit in the sound output direction of the audio signal from an inner product relational expression based on the angle formed by the directions of the two audio output blocks with the audio receiving unit as a reference. <11> The information processing device described in <13> further including another audio receiving unit different from the audio receiving unit, The position calculation unit calculating, by machine learning, positions of the audio output blocks of the audio receiving unit in a direction perpendicular to a sound output direction of the audio signals based on the arrival time difference distances of the audio signals of the two audio output blocks received by the audio receiving unit; Based on the arrival time difference distance of the audio signals of the two audio output blocks received by the audio receiving unit and the other audio receiving unit, respectively, simultaneous equations are constructed and solved to calculate the position of the audio receiving unit in the sound output direction of the audio signal of the audio output block, the angle formed by the direction connecting the audio receiving unit and the other audio receiving unit with the sound output direction of the audio signal of the audio output block, and the distance between the audio receiving unit and the other audio receiving unit. <2> The information processing device described in <14> further including another audio receiving unit different from the audio receiving unit, When the distance between the audio receiving unit and the other audio receiving unit is known, the position calculation unit calculates the two-dimensional position of the audio receiving unit and the angle formed by the direction connecting the audio receiving unit and the other audio receiving unit with the emitting direction of the audio signal of the audio output block by forming and solving simultaneous equations based on the arrival time difference distance of the audio signals of the two audio output blocks received by the audio receiving unit and the other audio receiving unit, respectively. <2> The information processing device described in <15> The information processing device is a smartphone or a head mounted display (HMD). <1> ~ <14> 10. An information processing device according to claim 9, wherein: <16> 1. An information processing method for an information processing device having an audio receiving unit that receives audio signals output from two audio output blocks located at known positions, the audio signals being spread spectrum modulated spread code signals, the spread code of which is Calculating the position of the audio receiving unit based on an arrival time difference distance, which is a difference in distance specified from an arrival time, which is the time it takes for the audio signal from the second audio output block to arrive at the audio receiving unit and be received. An information processing method comprising the steps. <17> an audio receiving unit for receiving audio signals output from two audio output blocks located at known positions, the audio signals being spread spectrum modulated spread codes; a position calculation unit that calculates a position of the audio receiving unit based on an arrival time difference distance that is a difference in distance specified from an arrival time that is the time it takes for the audio signal of the two audio output blocks to arrive at the audio receiving unit and be received; programs that make a computer function. [Explanation of symbols]
[0294] 11 Home audio system, 31, 31-1, 31-2 Audio output block, 32 Electronic device, 41 Audio input block, 42 Control unit, 43 Output unit, 44 Communication unit, 51, 51-1, 51-2 Audio input unit, 71 Spreading code generator, 72 Known music sound source generator, 73 Audio generator, 74 Audio output unit, 81 Diffusion unit, 82 Frequency shift processor, 83 Sound field controller, 91 Known music sound source eliminator, 92 Spatial transfer characteristic calculator, 93 Arrival time calculator, 94 Peak power detector, 95 Position calculator, 111 Arrival time difference distance calculator, 112 Peak power ratio calculator, 113 Position calculator, 130 Reverse shift processor, 131 Cross-correlation calculator, 132 Peak detector, 201, peak power frequency component ratio calculator, 202, position calculator, 211, position calculator, 230, IMU, 231, attitude calculator, 232, position calculator, 241, position calculator
Claims
1. an audio receiving unit for receiving audio signals output from two audio output blocks located at known positions, the audio signals being spread spectrum modulated with spread codes; an arrival time calculation unit that calculates the arrival time of each of the audio signals from the two audio output blocks to the audio receiving unit; an arrival time difference distance calculation unit that calculates a difference in distance between each of the two audio output blocks and itself as an arrival time difference distance based on the arrival times of the audio signals of the two audio output blocks and known positions of the two audio output blocks; a position calculation unit that calculates a position of the voice receiving unit based on the time difference of arrival distance, The arrival time calculation unit a cross-correlation calculation unit that calculates the cross-correlation between a spreading code signal in the voice signal received by the voice receiving unit and a spreading code signal in the voice signal output from the two voice output blocks; a peak detection unit that detects a time at which a peak occurs in the cross-correlation as the arrival time, The position calculation unit calculates the position of the voice receiving unit as an output layer by performing processing using a hidden layer made of a neural network formed by machine learning on an input layer made of the arrival time difference distance and a peak power ratio which is a ratio of powers at the timings at which the peaks occur in the cross-correlation of the voice signals output from the two voice output blocks. Information processing device.
2. a peak power detection unit that detects, as a peak power, the power of the audio signals output from each of the two audio output blocks at the peak when received by the audio receiving unit; a peak power ratio calculation unit that calculates a peak power ratio as a ratio between the peak powers of the audio signals output from the two audio output blocks, the ratio being detected by the peak power detection unit. The information processing device according to claim 1 .
3. The position calculation unit calculates the position of the voice receiving unit as the output layer by performing processing using the hidden layer on an input layer made up of the arrival time difference distance calculated by the arrival time difference distance calculation unit and the peak power ratio calculated by the peak power ratio calculation unit. The information processing device according to claim 2 .
4. The position calculation unit calculates the position of the voice receiving unit as an output layer by performing processing using a hidden layer made of a neural network formed by the machine learning on an input layer made of the arrival time difference distance, the peak power ratio of the voice signals output from each of the two voice output blocks, and a peak power frequency component ratio, which is the ratio of peak power of low frequency components to peak power of high frequency components of the voice signals output from each of the two voice output blocks. The information processing device according to claim 1 .
5. a peak power frequency component ratio calculation unit that detects peak power of the low frequency component and peak power of the high frequency component of the audio signal output from each of the two audio output blocks at the peak, and calculates the ratio of the peak power of the high frequency component to the peak power of the low frequency component as the peak power frequency component ratio; The information processing device according to claim 4 .
6. The position calculation unit calculating, by the machine learning, positions of the audio signals in the audio output blocks of the audio receiving unit in a direction perpendicular to the sound emission direction of the audio signals based on the arrival time difference distances of the audio signals in the two audio output blocks received by the audio receiving unit; Calculating the position of the sound output direction of the audio signal in the audio output block of the audio receiving unit based on the angle formed by the directions of the two audio output blocks with the audio receiving unit as a reference The information processing device according to claim 1 .
7. an IMU (Inertial Measurement Unit) that detects the angular velocity and acceleration of the audio receiving unit; a posture detection unit that detects the posture of the robot itself based on the angular velocity and the acceleration, The position calculation unit calculating an angle formed by the directions of the two audio output blocks with the audio receiving unit as a reference based on the orientation of the device itself detected by the orientation detection unit; Calculating the position of the sound output direction of the audio signal of the audio output block of the audio receiving unit based on the calculated angle formed by the directions of the two audio output blocks with the audio receiving unit as a reference. The information processing device according to claim 6 .
8. The position calculation unit calculates an angle formed by the directions of the two audio output blocks with the audio receiving unit as a reference, based on the orientation of the device when facing the two audio output blocks, detected by the orientation detection unit. The information processing device according to claim 7 .
9. The position calculation unit calculates the position of the audio output block of the audio receiving unit in the sound output direction of the audio signal from an inner product relational expression based on the angle formed by the directions of the two audio output blocks with the audio receiving unit as a reference. The information processing device according to claim 8 .
10. further including another audio receiving unit different from the audio receiving unit, The position calculation unit calculating, by machine learning, positions of the audio output blocks of the audio receiving unit in a direction perpendicular to a sound output direction of the audio signals, based on the arrival time difference distances of the audio signals of the two audio output blocks received by the audio receiving unit; Based on the arrival time difference distances of the audio signals of the two audio output blocks received by the audio receiving unit and the other audio receiving unit, respectively, simultaneous equations are constructed and solved to calculate the position of the audio receiving unit in the sound output direction of the audio signal of the audio output block, the angle formed by the direction connecting the audio receiving unit and the other audio receiving unit with the sound output direction of the audio signal of the audio output block, and the distance between the audio receiving unit and the other audio receiving unit. The information processing device according to claim 1 .
11. further including another audio receiving unit different from the audio receiving unit, When the distance between the audio receiving unit and the other audio receiving unit is known, the position calculation unit calculates the two-dimensional position of the audio receiving unit and the angle formed by the direction connecting the audio receiving unit and the other audio receiving unit with the emitting direction of the audio signal of the audio output block by forming and solving simultaneous equations based on the arrival time difference distance of the audio signals of the two audio output blocks received by the audio receiving unit and the other audio receiving unit, respectively. The information processing device according to claim 1 .
12. The information processing device is a smartphone or a head mounted display (HMD). The information processing device according to claim 1 .
13. 1. An information processing method for an information processing device having an audio receiving unit that receives audio signals output from two audio output blocks located at known positions, the audio signals being spread spectrum modulated spread code signals, the spread code of which is an arrival time calculation process for calculating an arrival time of each of the audio signals from the two audio output blocks to reach the audio receiving unit; an arrival time difference distance calculation process for calculating a difference in distance between each of the two audio output blocks and itself as an arrival time difference distance based on the arrival times of the audio signals of the two audio output blocks and known positions of the two audio output blocks; performing a position calculation process to calculate a position of the voice receiving unit based on the time difference of arrival distance; The arrival time calculation process includes: a cross-correlation calculation process for calculating a cross-correlation between a spreading code signal in the voice signal received by the voice receiving unit and a spreading code signal of the voice signal output from the two voice output blocks; a peak detection process for detecting a time when a peak occurs in the cross-correlation as the arrival time, The position calculation process calculates the position of the voice receiving unit as an output layer by performing processing using a hidden layer made of a neural network formed by machine learning on an input layer made of the arrival time difference distance and a peak power ratio, which is a ratio of powers at the timings at which the peaks occur in the cross-correlation of the voice signals output from the two voice output blocks. Information processing methods.
14. an audio receiving unit for receiving audio signals output from two audio output blocks located at known positions, the audio signals being spread spectrum modulated with spread codes; an arrival time calculation unit that calculates the arrival time of each of the audio signals from the two audio output blocks to the audio receiving unit; an arrival time difference distance calculation unit that calculates a difference in distance between each of the two audio output blocks and itself as an arrival time difference distance based on the arrival times of the audio signals of the two audio output blocks and known positions of the two audio output blocks; causing a computer to function as a position calculation unit that calculates a position of the voice receiving unit based on the time difference of arrival distance; The arrival time calculation unit a cross-correlation calculation unit that calculates the cross-correlation between a spreading code signal in the voice signal received by the voice receiving unit and a spreading code signal in the voice signal output from the two voice output blocks; a peak detection unit that detects a time at which a peak occurs in the cross-correlation as the arrival time, The position calculation unit calculates the position of the voice receiving unit as an output layer by performing processing using a hidden layer made of a neural network formed by machine learning on an input layer made of the arrival time difference distance and a peak power ratio which is a ratio of powers at the timings at which the peaks occur in the cross-correlation of the voice signals output from the two voice output blocks. program.
Citation Information
Patent Citations
Indoor positioning method and positioning system based on spread spectrum sound waves
CN112379331A
Local positioning system using ultrasonic wave
JP2001337157A
Position estimation method, device and program
JP2013195405A
Communication system, demodulation device, and modulation signal generating device
JP2014220741A
Audio processing for an acoustical environment
US20170278519A1