Program, information processing device, and information processing method
By receiving spreading code signals and using machine learning to infer device position, and combining simulated pseudo-random noise to generate datasets, the problem of low position measurement accuracy in existing technologies is solved, and the accuracy and robustness of position measurement are improved at low cost.
Patent Information
- Application Number
- CN202480024577.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-04-19
- Filing Date
- 2024-04-02
- Publication Date
- 2025-11-04
AI Technical Summary
In existing technologies, when using machine learning for location measurement, it is necessary to collect a large amount of real-world environmental data that meets various conditions, and the complexity of the simulated communication environment makes it difficult to improve accuracy.
By receiving and analyzing spread spectrum code signals, machine learning is used to infer the device's location. Simulated pseudo-random noise is combined to generate a simulated dataset, reducing reliance on actual environmental data and improving the accuracy of location measurement.
It achieves improved accuracy in position measurement of distance between devices at low cost, reduces the time and effort required to collect data from the actual environment, and enhances measurement accuracy in complex communication environments.
Smart Images

Figure CN120898149A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to a program, an information processing apparatus, and an information processing method, and more particularly, to a program, an information processing apparatus, and an information processing method capable of achieving position measurement using a modulated signal with high accuracy by machine learning. BACKGROUND
[0002] A technique has been proposed that achieves a receiving device position measurement based on a distance between devices obtained from a transmission / reception time at which a modulated signal transmitted from a plurality of transmission devices is received by a receiving device by using machine learning of actual environment data (see Non-Patent Literature 1).
[0003] LIST OF CITATIONS
[0004] NON-PATENT LITERATURE
[0005] Non-Patent Literature 1: Improving GNSS Positioning Using Neural-Network-Based Corrections SUMMARY
[0006] Technical problem to be solved by the invention
[0007] In the technique described in Non-Patent Literature 1, in order to improve the accuracy of position measurement in machine learning, a large amount of actual environment data satisfying various conditions needs to be collected, and then machine learning is performed using the collected actual environment data.
[0008] However, it takes time and effort to collect a large amount of actual environment data satisfying various conditions, and there are actual environment data that are difficult to collect depending on the conditions. Therefore, the position measurement using machine learning cannot easily improve the accuracy.
[0009] Furthermore, as a data set instead of actual environment data, for example, it can be thought that machine learning is achieved by creating an inter-device distance obtained from a transmission / reception time of a modulated signal by simulation and generating a data set combined with the inter-device distance.
[0010] However, in order to achieve an accurate simulation of a communication environment according to a modulated signal, for example, a complex and laborious data set needs to take into account all influences such as communication shielding and reflection of objects present in the communication environment, and therefore, the accuracy cannot be easily improved.
[0011] The present disclosure is made in view of this situation, and specifically, an object of the present disclosure is to easily improve accuracy of position measurement by machine learning using a distance between devices obtained from transmission / reception times at which modulated signals transmitted from a plurality of transmission devices are received by a receiving device.
[0012] Technical solution to technical problem
[0013] An information processing apparatus and program according to an aspect of the present disclosure is an information processing apparatus and program including: a ranging signal receiving unit that receives a ranging signal including a spread code signal obtained by performing spread modulation on a spread code and output from a plurality of ranging signal output blocks present at known positions; and a position inferring unit that infers a position of the ranging signal receiving unit by machine learning based on pseudo distances to the plurality of ranging signal output blocks and the known positions of the plurality of ranging signal output blocks, the pseudo distances being determined from transmission times until ranging signals of the plurality of ranging signal output blocks are transmitted to and received by the ranging signal receiving unit.
[0014] An information processing method according to an aspect of the present disclosure is an information processing method of an information processing apparatus including: a ranging signal receiving unit that receives a ranging signal including a spread code signal obtained by performing spread modulation on a spread code and output from a plurality of ranging signal output blocks present at known positions; and a position inferring unit that infers a position of the ranging signal receiving unit by machine learning based on pseudo distances to the plurality of ranging signal output blocks and the known positions of the plurality of ranging signal output blocks, the pseudo distances being determined from transmission times until ranging signals of the plurality of ranging signal output blocks are transmitted to and received by the ranging signal receiving unit, the information processing method including the step of inferring, by the position inferring unit, the position of the ranging signal receiving unit by machine learning based on the pseudo distances and the known positions of the plurality of ranging signal output blocks.
[0015] In one aspect of the present disclosure, a ranging signal including a spread code signal obtained by performing spread modulation on a spread code and output from a plurality of ranging signal output blocks present at known positions is received by a ranging signal receiving unit; and a position of the ranging signal receiving unit is inferred by machine learning based on pseudo distances to the plurality of ranging signal output blocks and the known positions of the plurality of ranging signal output blocks, the pseudo distances being determined from transmission times until ranging signals of the plurality of ranging signal output blocks are transmitted to and received by the ranging signal receiving unit. BRIEF DESCRIPTION OF DRAWINGS
[0016] Figure 1 is a diagram for explaining a configuration example of an acoustic positioning system of the present disclosure.
[0017] Figure 2 is a diagram for explaining a function implemented by an audio output block of the present disclosure. Figure 1
[0018] Figure 3 is a diagram for explaining a function implemented by an electronic device of the present disclosure. Figure 1
[0019] Figure 4 is a diagram for explaining communication using a spread code.
[0020] Figure 5 is a diagram for explaining autocorrelation and cross-correlation of a spread code.
[0021] Figure 6 is a diagram for explaining a transmission time of a spread code using cross-correlation.
[0022] Figure 7 is a diagram for explaining a configuration example of a transmission time calculation unit.
[0023] Figure 8 is a diagram for explaining a method of obtaining a position of an electronic device using an analysis method.
[0024] Figure 9 is a diagram for explaining how to obtain a position of an electronic device of the present disclosure.
[0025] Figure 10 is a diagram for explaining a function implemented by a position inference unit of the present disclosure. Figure 1
[0026] is a flowchart showing a position inference unit learning process by a learning device of the present disclosure. Figure 11 Figure 10 is a flowchart showing a position measurement process by an electronic device.
[0027] Figure 12 is a flowchart showing a position measurement process by an audio output block.
[0028] Figure 13 is a flowchart showing a transmission time calculation process.
[0029] Figure 14 shows a configuration example of a general-purpose computer.
[0030] Figure 15 DETAILED DESCRIPTION
[0031] Hereinafter, a preferred embodiment of the present disclosure will be described in detail with reference to the accompanying drawings. Note that, in this specification and the drawings, constituent elements having substantially the same function are denoted with the same reference numerals, and redundant description is omitted.
[0032] Hereinafter, a mode for carrying out the present technology will be described. The description will be given in the following order.
[0033] 1. Preferred Embodiment
[0034] 2. Example Performed by Software
[0035] <<1. Preferred Embodiment>>
[0036] <Configuration Example of Acoustic Positioning System>
[0037] Specifically, the present disclosure makes it possible to easily improve the accuracy of position measurement by machine learning based on the distance between devices obtained from the transmission / reception time when a modulated signal transmitted from a plurality of transmission devices is received by a receiving device.
[0038] Figure 1 A configuration example of an acoustic positioning system to which the technology of the present disclosure is applied is shown.
[0039] Figure 1 The acoustic positioning system 11 of the present embodiment includes audio output blocks 31-1 to 31-4 and an electronic device 32. Note that, hereinafter, in a case where it is not necessary to particularly distinguish the audio output blocks 31-1 to 31-4 from each other, they are simply referred to as audio output blocks 31, and other configurations are also mentioned in a similar manner.
[0040] Each of the audio output blocks 31-1 to 31-4 includes a speaker, and emits sound by including an audio signal that is a ranging signal including a modulated signal obtained by performing spread spectrum modulation on a data code for determining the position of the electronic device 32 by using a spread code in sound such as music content, a game, or known music.
[0041] The electronic device 32 is carried or worn by a user, and is, for example, a smartphone or a head-mounted display (HMD) used as a game controller.
[0042] The electronic device 32 includes an audio input block 41 including an audio input unit 51 such as a microphone that receives audio including an audio signal that is a ranging signal emitted from each of the audio output blocks 31-1 to 31-4, and a position detection unit 52.
[0043] The audio input block 41 identifies in advance the position in the space of each of the audio output blocks 31-1 to 31-4 as known position information through communication (ranging signal or other communication means) between the audio output block 31 and the electronic device 32. The audio input block 41 causes the audio input unit 51 to receive an audio signal that is a ranging signal including a modulated signal included in the audio emitted from the audio output block 31, and outputs the audio signal to the position detection unit 52. The position detection unit 52 obtains the distance to each of the audio output blocks 31-1 to 31-4 based on the audio signal that is the ranging signal including the modulated signal supplied from the audio input block 41, and detects the position of itself with respect to the audio output blocks 31-1 to 31-4 based on the obtained distance.
[0044] With this configuration, for example, in a case where the electronic device 32 is an HMD or a smartphone including a see-through display unit, it is possible to track the movement of the head of the user wearing the HMD or the smartphone (i.e., the electronic device 32).
[0045] Further, because the position of the HMD or the smartphone that is the electronic device 32 with respect to the audio output blocks 31-1 to 31-4 is determined, it is possible to output the audio output from the audio output blocks 31-1 to 31-4 in a case where the sound field positioning is corrected in accordance with the determined position. This configuration allows the user wearing the HMD that is the electronic device 32 to experience a sound with a sense of reality in accordance with the movement of the head.
[0046] <Function configuration example of audio output block>
[0047] Next, the function implemented by the audio output block 31 will be described with reference to Figure 2 The function implemented by the audio output block 31 will be described with reference to
[0048] The audio output block 31 includes a spread code generation unit 71, a known music sound source generation unit 72, an audio generation unit 73, an audio output unit 74, and a communication unit 75.
[0049] The spread code generation unit 71 generates a spread code and outputs the spread code to the audio generation unit 73.
[0050] The known music sound source generation unit 72 stores known music, generates a known music sound source based on the stored known music, and outputs the known music sound source to the audio generation unit 73.
[0051] The audio generation unit 73 applies spread modulation to the known music sound source using the spread code to generate audio including a spread signal, and outputs the audio to the audio output unit 74.
[0052] More specifically, the audio generation unit 73 includes a spread spectrum unit 81, a frequency shift processing unit 82, and a sound field control unit 83.
[0053] The spread spectrum unit 81 applies spread spectrum modulation using a spread spectrum code to the known music sound source to generate a spread spectrum signal.
[0054] The frequency shift processing unit 82 shifts the frequency of the spread spectrum code in the spread spectrum signal to a frequency band according to the frequency band characteristics of the audio output unit 74.
[0055] The sound field control unit 83 reproduces a sound field according to a positional relationship with its own position based on information about the position of the electronic device 32 provided from the electronic device 32.
[0056] The audio output unit 74 is, for example, a speaker, and outputs the known music sound source and the audio based on the spread spectrum signal provided from the audio generation unit 73.
[0057] The communication unit 75 is controlled by the audio generation unit 73, communicates with the electronic device 32 via Wi-Fi communication, Bluetooth (registered trademark) communication, or the like, and receives a request for sound emission to measure the position provided from the electronic device 32. Further, the communication unit 75 transmits its own position information to the electronic device 32 at a time before emitting sound.
[0058] <Configuration example of electronic device>
[0059] Next, a configuration example of the electronic device 32 will be described with reference to Figure 3
[0060] The electronic device 32 includes an audio input block 41, a control unit 42, and a communication unit 43.
[0061] The audio input block 41 receives input of audio emitted from each of the audio output blocks 31-1 to 31-4, obtains a distance to each of the audio output blocks 31-1 to 31-4 based on a correlation between a spread spectrum code of the received audio and a spread spectrum code, obtains its own position based on the obtained distance, and outputs the own position to the control unit 42.
[0062] In a case where the electronic device 32 is a smartphone used as a game controller, the control unit 42 controls, for example, the communication unit 43 based on the position of the electronic device 32 provided from the audio input block 41 to transmit a command to set a sound field based on the position of the electronic device 32 to the audio output blocks 31-1 to 31-4.
[0063] In this case, the sound field control unit 83 of each of the audio output blocks 31-1 to 31-4 adjusts the audio output from the audio output unit 74 based on a command for setting a sound field transmitted by the electronic device 32 to achieve the best sound field for the user who possesses the electronic device 32.
[0064] The audio input unit 51 is, for example, a microphone, and collects the audio emitted from each of the audio output blocks 31-1 to 31-4 and outputs the audio to the position detection unit 52.
[0065] The position detection unit 52 detects its own position based on the audio emitted from each of the audio output blocks 31-1 to 31-4 and outputs its own position to the control unit 42.
[0066] More specifically, the position detection unit 52 includes a known music sound source removal unit 91, a space transmission characteristic calculation unit 92, a transmission time calculation unit 93, a pseudo distance calculation unit 94, and a position inference unit 95.
[0067] The space transmission characteristic calculation unit 92 calculates a space transmission characteristic based on information about the audio provided from the audio input unit 51, characteristics of the microphones constituting the audio input unit 51, and characteristics of the speakers constituting the audio output units 74 of the audio output blocks 31, and outputs the space transmission characteristic to the known music sound source removal unit 91.
[0068] The known music sound source removal unit 91 stores the music sound source pre-stored in the known music sound source generation unit 72 of the audio output block 31 as a known music sound source.
[0069] Then, the known music sound source removal unit 91 removes the component of the known music sound source from the audio provided from the audio input unit 51 in consideration of the space transmission characteristic provided from the space transmission characteristic calculation unit 92, and outputs the result to the transmission time calculation unit 93.
[0070] That is, the known music sound source removal unit 91 removes the component of the known music sound source from the audio collected by the audio input unit 51, and outputs only the spread spectrum signal component to the transmission time calculation unit 93.
[0071] The transmission time calculation unit 93 calculates the transmission time from the emission of the audio from each of the audio output blocks 31-1 to 31-4 to the collection of the audio based on the spread spectrum signal component included in the audio collected by the audio input unit 51, and outputs the transmission time to the pseudo distance calculation unit 94.
[0072] Note that the method for calculating the transmission time will be described in detail later.
[0073] The pseudo distance calculation unit 94 calculates the pseudo distance between the electronic device 32 and each of the audio output blocks 31-1 to 31-4 based on the transmission time of each of the audio output blocks 31-1 to 31-4 provided by the transmission time calculation unit 93, and outputs the pseudo distance to the position inference unit 95.
[0074] The position inference unit 95 is, for example, an inference unit that includes a deep neural network (DNN) on which machine learning, described later, is performed, and infers the position of the electronic device 32 based on the pseudo distances between the electronic device 32 and the audio output blocks 31-1 to 31-4 provided by the pseudo distance calculation unit 94 and the pre-acquired positions of the audio output blocks 31-1 to 31-4, and outputs the position to the control unit 42.
[0075] Note: See later. Figure 10 Detailed description of the learning device 201 that enables the position inference unit 95 to perform machine learning ( Figure 10 ).
[0076] When it begins its own position measurement process, the control unit 42 communicates with the audio output block 31 to obtain the corresponding position information and outputs the position information to the position inference unit 95.
[0077] <Communication Principles Using Spread Codes>
[0078] Next, we will refer to Figure 4 Describe the principle of communication using spreading codes.
[0079] On the transmission side of the left portion of the figure, the spreading unit 81 performs spread spectrum modulation by multiplying the input signal Di with pulse width Td to be transmitted with the spreading code Ex to generate a transmission signal De with pulse width Tc, and transmits the transmission signal De to the receiving side of the right portion of the figure.
[0080] At this time, when the frequency band Dif of the input signal Di is indicated by, for example, the frequency band from -1 / Td to 1 / Td, the frequency band Exf of the transmitted signal De is widened to the frequency band from -1 / Tc to 1 / Tc (1 / Tc > 1 / Td) by multiplying it with the spreading code Ex, thereby spreading energy across the frequency axis.
[0081] Notice, Figure 4 An example is shown where the transmitted signal De is interfered with by the interference wave IF.
[0082] On the receiving side, the transmitted signal De, which is interfered with by the interference wave IF, is received as the received signal De'.
[0083] Transmission time calculation unit 93 (cross-correlation calculation unit 131) Figure 7The received signal Do is recovered by applying despreading to the received signal De' by the same spreading code Ex.
[0084] At this time, the frequency band Exf' of the received signal De' includes the component IFEx of the interference wave, but in the frequency band Dof of the despread received signal Do, the energy is spread by recovering the component IFEx of the interference wave to the spread frequency band IFD, so that the influence of the interference wave IF on the received signal Do can be reduced.
[0085] That is, as described above, in the communication using the spreading code, the influence of the interference wave IF generated on the transmission path of the transmission signal De can be reduced, and the noise resistance can be improved.
[0086] Further, in the spreading code, for example, the autocorrelation is in the form of a pulse as shown in the upper waveform chart of Figure 5 , and the cross-correlation is 0 as shown in the lower waveform of Figure 5 . Note that Figure 5 shows the change in the correlation value in the case where the Gold sequence is used as the spreading code, in which the horizontal axis indicates the code sequence, and the vertical axis indicates the correlation value.
[0087] That is, with the highly random spreading code set to each of the audio output blocks 31-1 to 31-4, the audio input block 41 can appropriately distinguish and recognize the spectrum signal included in the audio for each of the audio output blocks 31-1 to 31-4.
[0088] The spreading code can not only be the Gold sequence, but also an M sequence, a pseudo-random noise (PN), or the like.
[0089] <Method of calculating transmission time by transmission time calculation unit>
[0090] The time at which the peak of the observed cross-correlation is observed in the audio input block 41 is the time at which the audio emitted by the audio output block 31 is collected in the audio input block 41, and thus differs depending on the distance between the audio input block 41 and the audio output block 31.
[0091] That is, for example, as shown in the left portion of Figure 6 , when the distance between the audio input block 41 and the audio output block 31 is a first distance, the peak is detected at the time T1; as shown in the right portion of Figure 6 , when the distance between the audio input block 41 and the audio output block 31 is a second distance longer than the first distance, the peak is observed at the time T2 (> T1).
[0092] Note that in Figure 6 , the horizontal axis indicates the time elapsed from the output of the audio from the audio output block 31, and the vertical axis indicates the intensity of the cross-correlation.
[0093] That is, by multiplying the time from when the audio is emitted from the audio output block 31 to when the peak in the cross-correlation is observed (i.e., the transmission time from when the audio is emitted from the audio output block 31 to when the audio is collected in the audio input block 41) by the speed of sound, the distance between the audio input block 41 and the audio output block 31 can be obtained.
[0094] <Configuration example of transmission time calculation unit>
[0095] Next, a configuration example of the transmission time calculation unit 93 will be described with reference to Figure 7
[0096] The transmission time calculation unit 93 includes an inverse shift processing unit 130, a cross-correlation calculation unit 131, and a peak detection unit 132.
[0097] The inverse shift processing unit 130 restores a spread code signal that has been spread-modulated, which has been frequency-shifted by the up-sampling in the frequency shift processing unit 82 of the audio output block 31, to the original frequency band in the audio signal collected by the audio input unit 51 by down-sampling, and outputs the restored signal to the cross-correlation calculation unit 131.
[0098] Note that, regarding the shifting of the frequency band by the frequency shift processing unit 82 and the restoration of the frequency band by the inverse shift processing unit 130, refer to International Publication No. 2022 / 163801 filed by the present applicant.
[0099] The cross-correlation calculation unit 131 calculates the cross-correlation between the spread code and a reception signal obtained by removing the known musical sound source from the audio signal collected by the audio input unit 51 of the audio input block 41, and outputs the cross-correlation to the peak detection unit 132.
[0100] The peak detection unit 132 detects the peak time in the cross-correlation calculated by the cross-correlation calculation unit 131 and outputs the peak time as the transmission time.
[0101] Here, since it is well known that the calculation of the cross-correlation performed in the cross-correlation calculation unit 131 has particularly large computational complexity, the calculation is implemented by an equivalent calculation that has less computational complexity.
[0102] Specifically, as expressed by the following Equations (1) and (2), the cross-correlation calculation unit 131 performs a Fourier transform on each of the transmission signal outputted as the audio output by the audio output unit 74 of the audio output block 31 and the reception signal obtained by removing the known musical sound source from the audio signal received by the audio input unit 51 of the audio input block 41.
[0103] [Mathematical Expression 1]
[0104]
[0105] [Math. 2]
[0106]
[0107] Here, g represents a received signal obtained by removing a known music sound source from an audio signal received by the audio input unit 51 of the audio input block 41, and G represents a result of performing a Fourier transform on the received signal g obtained by removing the known music sound source from the audio signal received by the audio input unit 51 of the audio input block 41.
[0108] Further, h represents a transmission signal to be output as audio by the audio output unit 74 of the audio output block 31, and H represents a result of performing a Fourier transform on the transmission signal to be output as audio by the audio output unit 74 of the audio output block 31.
[0109] Further, V represents a sound velocity, v represents a velocity of the electronic device 32 (of the audio input unit 51), t represents time, and f represents a frequency.
[0110] Next, the cross-correlation calculation unit 131 obtains a cross spectrum by multiplying the results G and H of the Fourier transform with each other as represented by the following Equation (3).
[0111] [Math. 3]
[0112]
[0113] Here, P represents a cross spectrum obtained by multiplying the results G and H of the Fourier transform with each other.
[0114] Then, the cross-correlation calculation unit 131 performs an inverse Fourier transform on the cross spectrum P as represented by the following Equation (4) to obtain a cross-correlation between the transmission signal h output as audio by the audio output unit 74 of the audio output block 31 and the received signal g obtained by removing the known music sound source from the audio signal received by the audio input unit 51 of the audio input block 41.
[0115] [Math. 4]
[0116]
[0117] Here, p represents a cross-correlation between the transmission signal h output as audio by the audio output unit 74 of the audio output block 31 and the received signal g obtained by removing the known music sound source from the audio signal received by the audio input unit 51 of the audio input block 41.
[0118] The pseudo-distance calculation unit 94 calculates the following Equation (5) by the transmission time T obtained based on the peak of the cross-correlation p, thereby calculating the pseudo-distance between the audio input block 41 and the audio output block 31.
[0119] [Equation 5]
[0120] ...(5)
[0122] Here, D is the pseudo-distance between the audio input block 41 (of the audio input unit 51) and the audio output block 31 (of the audio output unit 74), T is the transmission time, and V is the speed of sound. Further, the speed of sound V is, for example, 331.5 + 0.6 x Q (m / s) (Q is the temperature °C).
[0123] Note that the cross-correlation calculation unit 131 can further obtain the speed v of the audio input unit 51 of the electronic device 32 by obtaining the cross-correlation p.
[0124] More specifically, the cross-correlation calculation unit 131 obtains the cross-correlation p while changing the speed v in a predetermined range (for example, -1.00 m / s to 1.00 m / s) at a predetermined step (for example, 0.01 m / s step), and obtains the speed v indicating the maximum peak of the cross-correlation p as the speed v of the audio input unit 51 of the electronic device 32.
[0125] The absolute speed of the audio input block 41 of the electronic device 32 can also be obtained based on the speed v obtained for each of the audio output blocks 31-1 to 31-4.
[0126] <Method of obtaining position of electronic device using analysis method>
[0127] (TOA positioning)
[0128] Next, a method of obtaining the position of the audio input unit 51-i of the electronic device 32 using an analysis method based on the distance Dik between the audio input unit 51-i of the electronic device 32 and the audio output block 31-k will be described.
[0129] Note that in the present disclosure, the position inference unit 95 including the DNN infers the position of the electronic device 32 based on the pseudo-distance between the electronic device 32 and each of the audio output blocks 31-1 to 31-4 provided from the pseudo-distance calculation unit 94 and the position of each of the audio output blocks 31-1 to 31-4, and does not use the analysis method described below to obtain the position.
[0130] However, the description will be made because of information required for a data set used for machine learning that describes the position estimation unit 95.
[0131] There are various methods for obtaining the position of the audio input unit 51-i of the electronic device 32 using an analysis method, and time of arrival (TOA) positioning and time difference of arrival (TDOA) positioning are well known.
[0132] The time of arrival (TOA) positioning is a method of obtaining the position of a reception-side device based on a distance relationship between the reception-side device and a plurality of transmission-side devices whose positions are known.
[0133] On the other hand, the time difference of arrival (TDOA) positioning is a method of obtaining the position of a reception-side device based on a distance difference relationship between the reception-side device and a plurality of transmission-side devices whose positions are known.
[0134] First, the time of arrival (TOA) positioning will be described. Here, it is assumed that the positions of the audio output units 74 of the audio output blocks 31 serving as transmission-side devices are known.
[0135] For example, as shown in FIG. 6, it is assumed that the position of the audio output unit 74-1 of the audio output block 31-1 is (X1, Y1, Z1), the position of the audio output unit 74-2 of the audio output block 31-2 is (X2, Y2, Z2), the position of the audio output unit 74-3 of the audio output block 31-3 is (X3, Y3, Z3), and the position of the audio output unit 74-4 of the audio output block 31-4 is (X4, Y4, Z4). Figure 8 Further, it is assumed that the position of the audio input unit 51-1 of the audio input block 41-1 of the electronic device 32-1 is (x1, y1, z1) and the position of the audio input unit 51-2 of the audio input block 41-2 of the electronic device 32-2 is (x2, y2, z2).
[0136] These are generalized so that the position of the audio output unit 74-k of the audio output block 31-k is (Xk, Yk, Zk), and the position of the audio input unit 51-i of the audio input block 41-i of the electronic device 32-i is (xi, yi, zi).
[0137] In this case, the distance Dik between the audio output unit 74-k of the audio output block 31-k and the audio input unit 51-i of the audio input block 41-i of the electronic device 32-i is represented by the following equation (6).
[0138] [Equation 6]
[0139]
[0140]
[0141] Here, Ds represents a distance offset corresponding to a system delay between the audio output block 31 and the audio input block 41.
[0142] Therefore, in a case where distances Di1 to Di4 between the audio input unit 51-i (of the audio input block 41-i of the electronic device 32-i) and the audio output block 31-1 to 31-4 (of the audio output unit 74) are obtained, the position (xi, yi, zi) of the audio input unit 51-i (of the audio input block 41-i of the electronic device 32-i) can be obtained by solving simultaneous equations represented by the following equation (7).
[0143] [Equation 7]
[0144]
[0145] Note that, in the above, in a case where the distance offset Ds corresponding to the time offset caused by the operation delay is known, the number of unknowns is three, and having three simultaneous equations is enough; therefore, if the positions of the three audio output units 74 are known, it can be solved.
[0146] (TDOA positioning)
[0147] As described above, in a case of a method of obtaining the position of the electronic device based on TOA positioning, in order to measure the transmission time described above, the transmission time and the reception time need to be strictly managed, and complete synchronization of the clocks used in the audio input unit 51 (of the audio input block 41 of the electronic device 32) and the audio output block 31 is a mandatory requirement.
[0148] However, it is actually difficult to completely synchronize the clocks of the audio input unit 51 (of the audio input block 41 of the electronic device 32) and the audio output block 31.
[0149] Therefore, in TDOA positioning, the use of the difference in the distance between the electronic device 32 and the audio output block 31-1 to 31-i compensates for the error caused by the asynchronous clocks used to measure the distance, and eliminates the need for synchronization of the clocks. Note that synchronization is not required for the clocks used in the electronic device 32 and the audio output block 31, but the times at which the audio is emitted from the audio output block 31-1 to 31-4 need to be synchronized.
[0150] More specifically, it can be obtained by solving simultaneous equations as represented by the following equation (8).
[0151] [Equation 8]
[0152]
[0153] <Inference of position of electronic device in the present disclosure>
[0154] As described above, when the distances Di1 to Di4 between the audio input unit 51-i (of the audio input block 41-i of the electronic device 32-i) and the audio output unit 74 (of the audio output block 31-1 to 31-4) are obtained, even in a state where the clocks of the electronic device 32 and the audio output block 31 are not synchronized, the position (xi, yi, zi) of the audio input unit 51-i (of the audio input block 41-i of the electronic device 32-i) can be obtained by solving the simultaneous equations represented by the above equation (8).
[0155] Meanwhile, the simultaneous equations represented by the above equation (7) or (8) are based on the assumption that the audio emitted from the audio output unit 74 of the audio output block 31 is transmitted linearly and received by the audio input unit 51 of the electronic device 32.
[0156] However, in reality, the audio emitted from the audio output unit 74 of the audio output block 31 can be transmitted linearly but not collected by the audio input unit 51-i of the audio input block 41-i of the electronic device 32-i.
[0157] For example, as shown in FIG. 15, in a case where a shielding object 151 exists between the audio output block 31-2 and the audio input block 41-1 of the electronic device 32-1, the audio emitted from the audio output block 31-2 is shielded by the shielding object 151, and thus, there is a possibility that the audio input block 41-1 cannot collect the audio at a sufficient level and cannot appropriately obtain the distance therebetween. Figure 9 Further, as shown in FIG. 16, in a case where a wall surface 152 exists in the upper portion of the figure, there is a possibility that the audio emitted from the audio output block 31-2 is reflected by the wall surface 152 and collected by the audio input block 41-1, as indicated by the alternate long and short dashed arrow.
[0158] Figure 9 As described above, in a case where the audio emitted from the audio output block 31-2 is shielded by the shielding object 151, reflected by the wall surface 152, and collected by the audio input block 41-1, the emitted audio is transmitted linearly, while the transmission path becomes long.
[0159] As a result, compared to the actual situation, the transmission of the emitted audio is delayed, and the time at which the emitted audio is collected is shifted, so that the transmission time becomes long, and it is possible to obtain a distance farther than the actual distance as the distance therebetween.
[0160] As described above, in a case where the audio emitted from the audio output block 31-2 is shielded by the shielding object 151, reflected by the wall surface 152, and collected by the audio input block 41-1, the emitted audio is transmitted linearly, while the transmission path becomes long.
[0161] In the above TDOA, as a method of suppressing the influence of such as shadowing and reflection, it is conceivable to more robustly obtain the position of the electronic device 32 by constructing four or more equations employing an iteratively reweighted least squares method (IRLS).
[0162] However, even with IRLS, any influence of shadowing and reflection cannot be robustly handled.
[0163] Therefore, in the present disclosure, the position of the electronic device 32 and the position of the audio output block 31 are randomly set by simulation within a settable range, and pseudo-random noise that simulates noise occurring on a transmission path is imparted to a mutual distance (hereinafter, a simulation distance), thereby generating a simulation pseudo-distance.
[0164] Then, by a data set including the simulation position of the electronic device 32, the simulation position of the audio output block 31, and the simulation pseudo-distance, a position inference unit 95 that infers the position of the electronic device 32 from the pseudo-distance and the position of the audio output block 31 performs machine learning.
[0165] More specifically, at the simulation position of the electronic device 32 and the simulation position of the audio output block 31 that are randomly set by simulation, pseudo-random noise that assumes a time-of-arrival offset due to shadowing or reflection is imparted to a simulation distance obtained by collecting audio emitted from the audio output block 31 in an ideal state in the electronic device 32, thereby generating a simulation pseudo-distance.
[0166] Then, by machine learning using a data set including the generated simulation position of the electronic device 32, the simulation position of the audio output block 31, and the simulation pseudo-distance, a position inference unit 95 that infers the position of the electronic device 32 based on the pseudo-distances of the electronic device 32 and the audio output blocks 31-1 to 31-4 provided from the pseudo-distance calculation unit 94 and the positions of the audio output blocks 31-1 to 31-4 that are known is generated.
[0167] In order to improve the inference accuracy of the position inference unit 95, it is necessary to generate and learn more simulation pseudo-distances in which various pseudo-random noises are added to the simulation distance as true values of a data set.
[0168] In the present disclosure, pseudo-random noise is imparted to the simulation distance as a true value, so that a data set that assumes various shadowing, reflection, and the like to which noise is imparted is generated at low cost.
[0169] As a result, machine learning of the position inference unit 95 using many data sets is promoted, and the accuracy of the position of the electronic device 32 inferred by the position inference unit 95 from the pseudo-distance and the position of the audio output block 31 is improved.
[0170] <Configuration example of learning device>
[0171] Next, a specific configuration example of a learning device 201 that causes the position estimation unit 95 of the present disclosure to perform machine learning will be described with reference to Figure 10
[0172] The learning device 201 includes a position simulator 211, a noise imparting unit 212, and a position estimation unit 213 (95). Figure 10 The position simulator 211 generates, as a simulation position, a position in which the electronic device 32 can exist and a position in which the audio output block 31 can exist, which are required for machine learning of the position estimation unit 213 (95), and outputs, as a simulation distance, a mutual distance between the electronic device 32 and the audio output block 31 at each simulation position.
[0173] More specifically, the position simulator 211 includes an electronic device position output unit 221, an audio output block position output unit 222, and a simulation distance calculation unit 223.
[0174] The electronic device position output unit 221 randomly generates, as a simulation position, a three-dimensional position in which the audio input unit 51 of the audio input block 41 of the electronic device 32 can exist in a predetermined space, and outputs the simulation position to the simulation distance calculation unit 223 and the position estimation unit 213 (95).
[0175] The audio output block position output unit 222 randomly generates, as a simulation position, a three-dimensional position in which the audio output unit 74 of the audio output block 31 can exist in a predetermined space, and outputs the simulation position to the simulation distance calculation unit 223 and the position estimation unit 213 (95).
[0176] The simulation distance calculation unit 223 calculates, as a simulation distance, a mutual distance between the simulation position of the electronic device 32 provided by the electronic device position output unit 221 and the simulation position of the audio output block 31 provided from the audio output block position output unit 222, and outputs the simulation distance to the noise imparting unit 212.
[0177] The simulation distance, which is a mutual distance between the simulation position of the electronic device 32 and the simulation position of the audio output block 31 calculated by the simulation distance calculation unit 223, is a mutual distance calculated in an ideal state that can be considered as a true value obtained by transmitting audio emitted from the audio output block 31 to the electronic device 32 and collecting the audio without being affected by noise such as shielding or reflection.
[0178]
[0179] The noise-imposing unit 212 imposes pseudo-random noise on the simulated distance provided from the simulated distance calculation unit 223 of the position simulator 211 to convert the simulated distance into a simulated pseudo-range, and outputs the simulated distance to the position inferring unit 213.
[0180] By imposing pseudo-random noise on the simulated distance, the audio emitted from the audio output blocks 31 is transmitted to the electronic device 32 and collected in a state affected by noise caused by various factors such as shielding and reflection, and is converted into a mutual distance (i.e., a simulated pseudo-range).
[0181] More specifically, for example, the noise-imposing unit 212 imposes noise of an influence of at least one of a hypothetical observation error, a multipath error, an NLOS error, and an audio emission time error of the plurality of audio output blocks 31 on the simulated distance.
[0182] The observation error is a random error of a time at which a peak of cross-correlation between the emitted audio detected by collecting the emitted audio and the collected audio is observed. Figure 6 and Figure 7 The peak of the cross-correlation between the emitted audio detected by collecting the emitted audio and the collected audio is observed.
[0183] The noise-imposing unit 212 imposes noise of an influence of the hypothetical observation error by imposing Gaussian noise or noise according to a uniform distribution (hereinafter, it is also referred to as uniform distribution noise) as a slight error for several frames on the simulated distance, which is a true value.
[0184] The multipath error is an error simulating a case in which a reflected wave of the emitted audio is erroneously collected in a case in which a direct wave of the emitted audio is to be collected, or an error simulating a case in which a direct wave of the emitted audio is erroneously collected in a case in which a reflected wave of the emitted audio is to be collected, and has a lower occurrence probability than the observation error.
[0185] In order to apply noise having an error simulating a case in which a reflected wave of the emitted audio is erroneously collected when a direct wave of the emitted audio is to be collected, the noise-imposing unit 212 imposes, as noise of the simulated multipath error, uniform distribution noise having a distance longer than the simulated distance on the simulated distance (which is a true value) observed by collecting the direct wave of the emitted audio, due to a collection delay caused by collecting the reflected wave of the emitted audio.
[0186] In addition, in order to apply noise having an error simulating a case in which a reflected wave of the emitted audio is erroneously detected when a reflected wave of the emitted audio is to be collected, the noise-imposing unit 212 imposes, as noise of the simulated multipath error, uniform distribution noise having a distance shorter than the simulated distance on the simulated distance (which is a true value) observed by collecting the reflected wave of the emitted audio, due to early collection of the direct wave of the emitted audio.
[0187] A non-line-of-sight (NLOS) error is an error in a case where transmitted audio is shielded by a shielding object or the like and cannot be collected, and has a large error amount maximum value, but has a lower occurrence probability than an observation error.
[0188] The noise-imposing unit 212 imposes uniform distribution noise of a simulated non-line-of-sight (NLOS) error on the simulated distance that is a true value.
[0189] The audio transmission time error of the plurality of audio output blocks 31-1 to 31-4 is an error caused by the fact that clocks in the plurality of audio output blocks 31-1 to 31-4 cannot be completely synchronized, and is an error caused by the fact that audio transmitted from each audio output unit 74 cannot be synchronized.
[0190] The noise-imposing unit 212 imposes Gaussian noise or uniform distribution noise that simulates the audio transmission time error of the plurality of audio output blocks 31-1 to 31-4 on the simulated distance that is a true value.
[0191] As described above, the noise-imposing unit 212 generates a simulated pseudo distance by applying pseudo random noise that simulates noise of at least one of the plurality of different types of errors described above to the simulated distance, and outputs the simulated pseudo distance to the position inference unit 213 (95).
[0192] The position inference unit 213 has a configuration corresponding to the position inference unit 95 and includes a deep neural network (DNN). In the position inference unit 213, machine learning is performed based on the simulated position of the electronic device 32 provided from the electronic device position output unit 221, the simulated position of the audio output block 31 provided from the audio output block position output unit 222, and the simulated pseudo distance provided from the noise-imposing unit 212.
[0193] By performing such machine learning, the position inference unit 213 (95) functions as an inference unit that infers the position of the electronic device 32 based on the pseudo distance calculated by the pseudo distance calculation unit 94 of the position detection unit in the electronic device 32 and the respective positions of the audio output blocks 31-1 to 31-4 that are known.
[0194] With such a configuration, the learning device 201 of the present disclosure can generate a large amount of data set to which noise simulating various types of errors is imposed at a low cost, and cause the position inference unit 213 (95) to perform machine learning.
[0195] As a result, since it is possible to generate a large number of data sets at low cost, it is easy to implement machine learning of the position inference unit 95, and it is possible to highly accurately infer the position of the electronic device 32 from the pseudo distances calculated by the pseudo distance calculation unit 94 of the position detection unit in the electronic device 32 and the known positions of the audio output blocks 31-1 to 31-4.
[0196] <Position Inference Unit Learning Processing>
[0197] Next, the learning processing of the learning device 201 for the position inference unit 213 (95) will be described with reference to the flowchart of Figure 11 Figure 10
[0198] In step Sll, the electronic device position output unit 221 of the position simulator 211 randomly generates a three-dimensional position in which the electronic device 32 (the audio input unit 51 of the audio input block 41) can exist in a predetermined space by simulation, and outputs the three-dimensional position as a simulation position to the simulation distance calculation unit 223 and the position inference unit 213 (95).
[0199] In step S12, the audio output block position output unit 222 of the position simulator 211 randomly generates a three-dimensional position in which the audio output block 31 (the audio output unit 74) can exist in a predetermined space by simulation, and outputs the three-dimensional position as a simulation position to the simulation distance calculation unit 223 and the position inference unit 213 (95).
[0200] In step S13, the simulation distance calculation unit 223 calculates the mutual distance as a simulation distance (which is a true value) from the simulation position of the electronic device 32 provided from the electronic device position output unit 221 and the simulation position of the audio output block 31 provided from the audio output block position output unit 222, and outputs the simulation distance to the noise imparting unit 212.
[0201] In step S14, the noise imparting unit 212 imparts pseudo random noise that simulates at least one of the above-described plurality of different types of errors to the simulation distance provided from the simulation distance calculation unit 223 of the position simulator 211 to convert the simulation distance into a simulation pseudo distance and output to the position inference unit 213.
[0202] In step S15, machine learning of the position inference unit 213 (95) that infers the position of the electronic device 32 based on the pseudorange calculated from the position detection unit in the electronic device 32 and the known position of the audio output block 31 is performed based on the data set including the simulated position of the electronic device 32 provided from the electronic device position output unit 221, the simulated position of the audio output block 31 provided from the audio output block position output unit 222, and the simulated pseudorange provided from the noise imparting unit 212.
[0203] With the above processing, it is possible to generate a large number of data sets to which noise that simulates various types of errors is added at low cost (easily), and to cause the position inference unit 213 to perform machine learning.
[0204] As a result, it becomes possible to easily implement machine learning of the position inference unit 213 (95) using a large number of data sets, and the electronic device 32 can highly accurately infer the position of the electronic device 32 (which is the own position) by collecting audio emitted from the audio output block 31.
[0205] Note that in the above, an example in which the electronic device position output unit 221 and the audio output block position output unit 222 randomly set the positions from all positions in which the electronic device 32 and the audio output block 31 can exist in the predetermined space has been described.
[0206] However, in a case where the range in which the electronic device 32 and the audio output block 31 can exist in the predetermined space is determined within a predetermined range, it is desirable to randomly set the positions within the determined predetermined range.
[0207] That is, in a state in which the positions of the electronic device 32 and the audio output block 31 output from the electronic device position output unit 221 and the audio output block position output unit 222 narrow within a predetermined range in the predetermined space, machine learning based on the generated data sets becomes possible.
[0208] As a result, compared to machine learning using data sets generated for all positions in which the electronic device 32 and the audio output block 31 can exist in the predetermined space, it is possible to generate the position inference unit 95 with equivalent inference accuracy by machine learning of a smaller number of data sets within the predetermined range, and it is possible to improve the processing efficiency of the learning processing.
[0209] Furthermore, by machine learning of data sets within a determined predetermined range in the predetermined space (the number of which is the same as data sets generated for all positions in which the electronic device 32 and the audio output block 31 can exist in the predetermined space), it is possible to train a position inference unit 95 with higher inference accuracy.
[0210] Therefore, regarding the position inference unit 95, the position inference unit that has already undergone learning processing using a dataset generated for all possible positions in the predetermined space can be installed and used as is. However, when the learning processing is performed again under specified conditions after determining at least one of the arrangement of the audio output block position output unit 222 and the range of movement of the electronic device 32, the efficiency related to the learning processing and inference accuracy can be improved.
[0211] <Location Measurement Processing>
[0212] Next, we will refer to Figure 12 to Figure 14 The flowchart in the document describes the position measurement and processing of the audio output block 31 and the electronic device 32.
[0213] Note that, under the assumption that position inference unit learning processing has already been performed, refer to Figure 11 Flowchart description Figure 12 to Figure 14 The processing.
[0214] also, Figure 12 This is a flowchart illustrating the processing of electronic device 32, and Figure 13 This is a flowchart illustrating the processing of audio output block 31. Furthermore, Figure 14 It is used to illustrate as in Figure 12 The flowchart for the transmission time calculation process in step S38 is shown.
[0215] In step S31 ( Figure 12 In the process, the control unit 42 of the electronic device 32 determines whether a command for starting position measurement processing has been given by the user through operation of the operation unit (not shown), and repeats similar processing until a command is given.
[0216] Then, in step S31, if the start position measurement process is indicated, the process proceeds to step S32.
[0217] In step S32, the control unit 42 controls the communication unit 43 to request the audio output block 31 to start position measurement processing.
[0218] In step S51 ( Figure 13 In the audio output block 31, the audio generation unit 73 controls the communication unit 75 to determine whether the electronic device 32 has requested to start position measurement processing, and repeats similar processing until a request is made.
[0219] Then, in step S51, if the electronic device 32 requests the start of position measurement processing, the process proceeds to step S52.
[0220] In step S52, the audio generation unit 73 controls the communication unit 75 to transmit its own position information to the electronic device 32 together with the information indicating the start of the position measurement processing. By this processing, the position information about each of the audio output blocks 31-1 to 31-4 is transmitted from the audio output blocks 31-1 to 31-4 to the electronic device 32.
[0221] In step S33 ( Figure 12 ), the control unit 42 of the electronic device 32 controls the communication unit 43 to acquire the information indicating the start of the position measurement processing and the position information provided from the audio output block 31, and provides the acquired information and the position information to the position inference unit 95.
[0222] In step S34, the control unit 42 controls the communication unit 43 to request the audio output block 31 to emit the audio.
[0223] In step S53 ( Figure 13 ), the audio generation unit 73 of the audio output block 31 controls the communication unit 75 to determine whether or not the emission of the audio is requested, and repeats similar processing until the emission is requested.
[0224] Then, in step S53, when the emission of the audio is requested, the processing proceeds to step S54.
[0225] In step S54, the audio generation unit 73 controls the spread code generation unit 71 to generate and acquire the spread code.
[0226] In step S55, the audio generation unit 73 controls the known music sound source generation unit 72 to generate and acquire the stored known music sound source.
[0227] In step S56, the audio generation unit 73 controls the spread code unit 81 to multiply the predetermined data code with the spread code and perform spread modulation to generate a spread code signal.
[0228] In step S57, the audio generation unit 73 controls the frequency shift processing unit 82 to frequency shift the spread code signal according to each frequency characteristic of the audio output unit 74.
[0229] In step S58, the audio generation unit 73 outputs the known music sound source and the frequency-shifted spread code signal to the audio output unit 74 including a speaker, and emits (outputs) the signal as the audio.
[0230] By executing the above-described processing in each of the audio output blocks 31-1 to 31-4, the audio can be emitted and the user who possesses the electronic device 32 is allowed to listen to the audio as the known music sound source.
[0231] Further, since the spread code signal can be shifted to a frequency band including inaudible to a person and output as audio, the electronic device 32 can measure the distance to the audio output block 31 based on the emission audio including the spread code signal shifted to the inaudible frequency band to the person without making the person hear an unpleasant sound.
[0232] In step S35 ( Figure 12 ), the audio input unit 51 including a microphone collects audio and outputs the collected audio to the known music sound source removal unit 91 and the spatial transmission characteristic calculation unit 92 of the position detection unit 52.
[0233] In step S36, the spatial transmission characteristic calculation unit 92 calculates the spatial transmission characteristic based on the audio provided from the audio input unit 51, the characteristics of the audio input unit 51, and the characteristics of the audio output unit 74 of the audio output block 31, and outputs the spatial transmission characteristic to the known music sound source removal unit 91.
[0234] In step S37, the known music sound source removal unit 91 generates an inverse signal of the known music sound source in consideration of the spatial transmission characteristic provided from the spatial transmission characteristic calculation unit 92, removes the component of the known music sound source from the audio provided from the audio input unit 51, and outputs the audio to the transmission time calculation unit 93.
[0235] In step S38, the transmission time calculation unit 93 performs the transmission time calculation processing, calculates the transmission time until the audio output from the audio output block 31 is transmitted to the audio input unit 51, and outputs the transmission time to the pseudo distance calculation unit 94.
[0236] <Transmission time calculation processing>
[0237] Here, the transmission time calculation processing by the transmission time calculation unit 93 will be described with reference to the flowchart of Figure 14 .
[0238] In step S71, the inverse shift processing unit 130 inversely shifts the frequency band of the spread code signal from which the known music sound source has been removed from the audio input from the audio input unit 51 provided from the known music sound source removal unit 91.
[0239] In step S72, the cross-correlation calculation unit 131 calculates the cross-correlation between the spread code signal obtained by inversely shifting the frequency band from the audio input from the audio input unit 51 and removing the known music sound source and the spread code signal of the audio output from the audio output block 31 by the calculation using the above-described Equations (1) to (4).
[0240] In step S73, the peak detection unit 132 detects a peak in the calculated cross-correlation.
[0241] In step S74, the peak detection unit 132 outputs the time detected as a peak in the cross-correlation to the pseudo distance calculation unit 94 as a transmission time.
[0242] Note that the transmission time corresponding to each of the plurality of audio output blocks 31 is obtained by calculating the cross-correlation with the spread code signal of the audio output from each of the plurality of audio output blocks 31.
[0243] Here, the description returns to Figure 12 the flowchart.
[0244] In step S39, the pseudo distance calculation unit 94 multiplies the transmission time corresponding to each of the plurality of audio output blocks 31 by the speed of sound to calculate the pseudo distance between each of the plurality of audio output blocks 31 and itself (the electronic device 32), and outputs the pseudo distance to the position inference unit 95.
[0245] In step S40, the position inference unit 95 infers the position of itself (the electronic device 32) from the distance between each of the plurality of audio output blocks 31 and itself (the electronic device 32) and the information on the positions of the plurality of audio output blocks 31, and outputs information on the position of itself (self position) as an inference result to the control unit 42.
[0246] In step S41, the control unit 42 performs processing based on the inferred position of the electronic device 32 (the position of itself), and ends the processing.
[0247] For example, the control unit 42 controls the communication unit 43 to transmit a command for controlling the level and time of the audio output from each of the audio output units 74 of the audio output blocks 31-1 to 31-4 to the audio output blocks 31-1 to 31-4, so that a sound field based on the obtained position of the electronic device 32 can be realized.
[0248] As a result, in the audio output blocks 31-1 to 31-4, the sound field control unit 83 controls the level and time of the audio output from the audio output units 74 based on the command transmitted from the electronic device 32 to realize a sound field corresponding to the position of the user who owns the electronic device 32.
[0249] With the above series of processing, when the position of the electronic device 32 is measured, the position of the electronic device 32 is inferred as a self position by the position inference unit 95 based on the pseudo distance and the positions of the audio output blocks 31, on which machine learning based on the pseudo distance obtained from the audio emitted from the audio output blocks 31, the positions of the audio output blocks, and the position of the electronic device 32 is performed, so that the accuracy of the position of the electronic device 32 can be improved.
[0250] Note that, in the above, an example in which an audio signal based on sound (sound wave) is used as a transmission medium of a modulation signal used as a ranging signal has been described; however, a different transmission medium, such as a radio wave or light, can be used. For example, in a smart factory or the like, a user position tracking system or the like that determines the position of a smart phone possessed by a user by transmitting and receiving a ranging signal including a modulation signal using a radio wave such as ultra wide band (UWB) as a transmission medium can be implemented.
[0251] As described above, in a case in which a ranging signal including a modulation signal is transmitted and received using a radio wave such as UWB as a transmission medium, the audio output block 31 that transmits (outputs) a ranging signal including an audio signal to emit audio can be replaced by, for example, a radio wave output block 31 that transmits (outputs) a ranging signal including a radio wave signal, and can be configured to be installed at a known position of a smart factory. Further, in this case, the audio input block 41 that collects and receives (inputs) a ranging signal including an audio signal can be replaced by, for example, a radio wave input block 41 or the like that receives (inputs) a ranging signal including a radio wave signal, and can be configured to be incorporated into a smart phone possessed by a user.
[0252] Further, in a case in which a transmission medium other than a sound wave and a radio wave, such as light, is used, the audio output block 31 and the audio input block 41 can be replaced by, for example, a light output block 31 and a light input block 41 that transmit and receive a light signal used as a ranging signal including a modulation signal using light as a transmission medium.
[0253] Further, the audio output block 31 and the audio input block 41 can be replaced by a ranging signal output block 31 that outputs (transmits) a ranging signal and a ranging signal input block 41 that receives (inputs) a ranging signal, respectively, regardless of the type of a transmission medium of a ranging signal.
[0254] Further, regarding the modulation signal, an example in which a signal modulated by a spread code is transmitted and received has been described; however, other signals can be used, specifically, a transmission signal such as frequency modulated continuous wave (FMCW) or a BLE beacon can be used.
[0255] <<2. Examples of execution by software>>
[0256] Incidentally, the series of processes described above can be executed by hardware, but can also be executed by software. In a case in which the series of processes is executed by software, a program forming the software is installed from a recording medium into, for example, a computer built into dedicated hardware or a general-purpose computer capable of executing various functions by installing various programs, or the like.
[0257] Figure 15A configuration example of a general-purpose computer is shown. The computer includes a central processing unit (CPU) 1001. An input / output interface 1005 is connected to the CPU 1001 via a bus 1004. A read only memory (ROM) 1002 and a random access memory (RAM) 1003 are connected to the bus 1004.
[0258] The input / output interface 1005 is connected to an input unit 1006 including an input device such as a keyboard or mouse by which a user inputs operation commands, an output unit that outputs an image of a processing operation screen and a processing result to a display device, a storage unit 1008 including a hard disk drive or the like that stores programs and various types of data, and a communication unit 1009 including a local area network (LAN) adapter or the like and performs communication processing via a network represented by the Internet. Further, a drive 1010 that reads and writes data from and to a removable storage medium 1011 such as a magnetic disk (including a floppy® disk), an optical disk (including a compact disc read only memory (CD-ROM) and a digital versatile disc (DVD)), a magneto-optical disk (including a mini disk (MD)), or a semiconductor memory is connected.
[0259] The CPU 1001 executes various types of processing in accordance with a program stored in the ROM 1002 or a program read from a removable storage medium 1011 such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory that is installed in the storage unit 1008 and loaded from the storage unit 1008 to the RAM 1003. The RAM 1003 also appropriately stores data required for the CPU 1001 to execute various types of processing and the like.
[0260] In the computer configured as described above, for example, the CPU 1001 loads a program stored in the storage unit 1008 into the RAM 1003 via the input / output interface 1005 and the bus 1004 and executes the program, thereby executing the series of processing described above.
[0261] The program executed by the computer (CPU 1001), for example, can be provided by being recorded in the removable storage medium 1011 as a package medium or the like. Further, the program can be provided via a wired or wireless transmission medium such as a local area network, the Internet, or digital satellite broadcasting.
[0262] In the computer, the program can be installed in the storage unit 1008 by attaching the removable storage medium 1011 to the drive 1010 via the input / output interface 1005. Further, the program can be received by the communication unit 1009 via a wired or wireless transmission medium to be installed on the storage unit 1008. Alternatively, the program can be installed in the ROM 1002 or the storage unit 1008 in advance.
[0263] Note that a program executed by a computer can be a program that executes processing in the time series order described in this specification, or can be a program that executes processing in parallel or at a necessary time such as when called.
[0264] Note that the CPU 1001 in Figure 15 implements the functions of the audio output block 31 and the audio input block 41 of Figure 1 , and the learning device 201 of Figure 10 .
[0265] Further, in this specification, a system means a group of a plurality of constituent elements (apparatuses, modules (parts), and the like), and whether all constituent elements are in the same housing is irrelevant. Therefore, a plurality of apparatuses housed in separate housings and connected to each other via a network and one apparatus including a plurality of modules housed in one housing are all systems.
[0266] Note that the embodiments of the present disclosure are not limited to the above-described embodiments, and various modifications can be made without departing from the spirit of the present disclosure.
[0267] For example, the present disclosure can have a configuration of cloud computing in which one function is shared by a plurality of devices via a network and processing is executed in cooperation.
[0268] Further, each step described in the above-described flowcharts can be executed by a single device, or can be executed in a shared manner by a plurality of devices.
[0269] Further, in a case where a plurality of processes are included in one step, the plurality of processes included in one step can be executed by one device or in a shared manner by a plurality of devices.
[0270] Note that the present disclosure can also have the following configuration.
[0271] <1> A program for causing a computer to function as:
[0272] a ranging signal receiving unit that receives a ranging signal, the ranging signal including a spread code signal obtained by performing spread spectrum modulation on a spread code and being output from a plurality of ranging signal output blocks present at known positions; and
[0273] a position inference unit that infers a position of the ranging signal receiving unit by machine learning based on pseudo distances to the plurality of ranging signal output blocks and the known positions of the plurality of ranging signal output blocks, the pseudo distances being determined from transmission times, the transmission times being times until ranging signals of the plurality of ranging signal output blocks are transmitted to the ranging signal receiving unit and received by the ranging signal receiving unit.
[0274] <2> The program according to <1>, wherein
[0275] The position inferring unit performs machine learning using a data set based on the simulated distances, which are true values of distances between the ranging signal receiving unit and each of the plurality of ranging signal output blocks.
[0276] <3> The program according to <2>, wherein
[0277] The simulated distances are true values of distances between a simulated position of the ranging signal receiving unit and simulated positions of each of the plurality of ranging signal output blocks, the simulated positions being generated by simulation and randomly set in a predetermined space.
[0278] <4> The program according to <3>, wherein
[0279] The position inferring unit performs machine learning using a data set including simulated pseudo-distances generated by assigning noise assumed in transmission and reception of the ranging signal to the simulated distances.
[0280] <5> The program according to <4>, wherein
[0281] The position inferring unit includes a DNN (Deep Neural Network), and performs machine learning to infer the position of the ranging signal receiving unit based on pseudo-distances and known positions of the plurality of ranging signal output blocks, using a data set including the simulated pseudo-distances generated based on the simulated distances, simulated positions of the ranging signal receiving unit corresponding to the pseudo-distances, and simulated positions of each of the plurality of ranging signal output blocks.
[0282] <6> The program according to <4>, wherein
[0283] The simulated pseudo-distances are generated by assigning pseudo-random noise corresponding to noise assumed in transmission and reception of the ranging signal to the simulated distances.
[0284] <7> The program according to <6>, wherein
[0285] The simulated pseudo-distances are generated by assigning pseudo-random noise simulating noise of at least one of an observation error, a multipath error, an NLOS (Non Line of Sight) error, and an audio transmission time error of the plurality of ranging signal output blocks assumed in transmission and reception of the ranging signal to the simulated distances.
[0286] <8> The program according to <7>, wherein
[0287] The simulated pseudo-distances are generated by assigning Gaussian noise or uniform distribution noise simulating an observation error assumed in transmission and reception of the ranging signal to the simulated distances.
[0288] <9> The program according to <7>, in which,
[0289] The pseudo distance is generated by assigning a uniform distribution noise simulating a multipath error assumed in transmission and reception of the ranging signal to the simulated distance.
[0290] <10> The program according to <7>, in which,
[0291] The pseudo distance is generated by assigning a uniform distribution noise simulating an NLOS error assumed in transmission and reception of the ranging signal to the simulated distance.
[0292] <11> The program according to <7>, in which,
[0293] The pseudo distance is generated by assigning a Gaussian noise or a uniform distribution noise simulating an audio transmission time error of a plurality of ranging signal output blocks assumed in transmission and reception of the ranging signal to the simulated distance.
[0294] <12> The program according to <1>, further comprising:
[0295] a transmission time calculating unit that calculates a transmission time until each ranging signal of the plurality of ranging signal output blocks is transmitted to the ranging signal receiving unit; and
[0296] a pseudo distance calculating unit that calculates a pseudo distance of each of the plurality of ranging signal output blocks from the transmission time of each ranging signal of the plurality of ranging signal output blocks, wherein,
[0297] a position inferring unit infers a position of the ranging signal receiving unit by machine learning based on the known positions of the plurality of ranging signal output blocks and the pseudo distances of the plurality of ranging signal output blocks.
[0298] <13> The program according to <12>, in which,
[0299] the transmission time calculating unit includes:
[0300] a cross-correlation calculating unit that calculates a cross-correlation between a spread code signal in the ranging signal received by the ranging signal receiving unit and a spread code signal of the ranging signal output from the plurality of ranging signal output blocks; and
[0301] a peak detecting unit that detects a peak time in the cross-correlation as the transmission time;
[0302] the pseudo distance calculating unit calculates the pseudo distance of each of the plurality of ranging signal output blocks from the transmission time of each ranging signal of the plurality of ranging signal output blocks, and
[0303] The position inference unit infers the position of the ranging signal receiving unit by machine learning based on pseudo distances to the plurality of ranging signal output blocks and known positions of the plurality of ranging signal output blocks.
[0304] <14> The program according to any one of <1> to <13>, in which,
[0305] The ranging signal receiving unit is provided on a smartphone or an HMD (Head Mounted Display).
[0306] <15> The program according to any one of <1> to <14>, in which,
[0307] The ranging signal includes an audio signal, a radio wave signal, and a light signal.
[0308] <16> An information processing apparatus comprising:
[0309] a ranging signal receiving unit that receives a ranging signal, the ranging signal including a spread code signal obtained by performing spread spectrum modulation on a spread code and being output from a plurality of ranging signal output blocks present at known positions; and
[0310] a position inference unit that infers the position of the ranging signal receiving unit by machine learning based on pseudo distances to the plurality of ranging signal output blocks and the known positions of the plurality of ranging signal output blocks, the pseudo distances being determined from transmission times, the transmission times being times until ranging signals of the plurality of ranging signal output blocks are transmitted to and received by the ranging signal receiving unit.
[0311] <17> An information processing method of an information processing apparatus including:
[0312] a ranging signal receiving unit that receives a ranging signal, the ranging signal including a spread code signal obtained by performing spread spectrum modulation on a spread code and being output from a plurality of ranging signal output blocks present at known positions; and
[0313] a position inference unit that infers the position of the ranging signal receiving unit by machine learning based on pseudo distances to the plurality of ranging signal output blocks and the known positions of the plurality of ranging signal output blocks, the pseudo distances being determined from transmission times, the transmission times being times until ranging signals of the plurality of ranging signal output blocks are transmitted to and received by the ranging signal receiving unit,
[0314] the information processing method including the step of inferring, by the position inference unit, the position of the ranging signal receiving unit by machine learning based on the pseudo distances and the known positions of the plurality of ranging signal output blocks.
[0315] List of reference symbols
[0316] 11 acoustic positioning system, 31, 31-1 to 31-4 audio output block, 32 electronic device, 41 audio input block, 42 control unit, 43 communication unit, 51, 51-1 to 51-4 audio input unit, 71 spread code generating unit, 72 known music sound source generating unit, 73 audio generating unit, 74 audio output unit, 81 spread unit, 82 frequency shift processing unit, 83 sound field control unit, 91 known music sound source removing unit, 92 spatial transmission characteristic calculating unit, 93 transmission time calculating unit, 94 pseudo distance calculating unit, 95 position inference unit, 130 inverse shift processing unit, 131 cross correlation calculating unit, 132 peak value detecting unit, 201 learning device, 211 position simulator, 212 noise imparting unit, 213 position inference unit, 221 electronic device position output unit, 222 audio output block position output unit, 223 simulated distance calculating unit.
Claims
1. A program for causing a computer to function as: a ranging signal receiving unit that receives a ranging signal including a spread code signal obtained by performing spread modulation on a spread code and output from a plurality of ranging signal output blocks present at known positions; and a position inference unit that infers a position of the ranging signal receiving unit based on pseudo distances to the plurality of ranging signal output blocks and the known positions of the plurality of ranging signal output blocks, the pseudo distances being determined from transmission times until ranging signals of the plurality of ranging signal output blocks are transmitted to and received by the ranging signal receiving unit, by machine learning.
2. The program according to claim 1, wherein the position inference unit performs the machine learning using a data set based on simulated distances, the simulated distances being true values of distances between the ranging signal receiving unit and each of the plurality of ranging signal output blocks.
3. The program according to claim 2, wherein the simulated distances are true values of distances between simulated positions of the ranging signal receiving unit and simulated positions of each of the plurality of ranging signal output blocks, the simulated positions being generated by simulation and randomly set in a predetermined space.
4. The program according to claim 3, wherein the position inference unit performs the machine learning using a data set including simulated pseudo distances, the simulated pseudo distances being generated by assigning noise assumed in transmission and reception of the ranging signals to the simulated distances.
5. The program according to claim 4, wherein the position inference unit includes a DNN (Deep Neural Network) and performs the machine learning to infer the position of the ranging signal receiving unit based on the pseudo distances and the known positions of the plurality of ranging signal output blocks using a data set including the simulated pseudo distances generated based on the simulated distances, the simulated positions of the ranging signal receiving unit, and the simulated positions of each of the plurality of ranging signal output blocks.
6. The program according to claim 4, wherein the simulated pseudo distances are generated by assigning pseudo-random noise corresponding to noise assumed in transmission and reception of the ranging signals to the simulated distances.
7. The program according to claim 6, wherein the simulated pseudo distances are generated by assigning the pseudo-random noise to the simulated distances, the pseudo-random noise simulating noise of at least one of an observation error, a multipath error, an NLOS (Non Line of Sight) error, and an audio transmission time error of the plurality of ranging signal output blocks assumed in transmission and reception of the ranging signals.
8. The program according to claim 7, wherein the simulated pseudo distances are generated by assigning Gaussian noise or uniform distribution noise simulating the observation error assumed in transmission and reception of the ranging signals to the simulated distances.
9. The program according to claim 7, wherein The simulated pseudo-range is generated by simulating the uniform distribution noise of the multipath error assumed in the transmission and reception of the ranging signal.
10. The procedure according to claim 7, wherein, The simulated pseudo-range is generated by simulating a uniformly distributed noise of the NLOS error assumed in the transmission and reception of the ranging signal.
11. The procedure according to claim 7, wherein, The simulated pseudo-distance is generated by imposing Gaussian noise or uniformly distributed noise, simulating the audio transmission time error of the plurality of ranging signal output blocks assumed in the transmission and reception of the ranging signals, onto the simulated distance.
12. The procedure according to claim 1, further comprising: A transmission time calculation unit calculates the transmission time until each ranging signal of the plurality of ranging signal output blocks is transmitted to the ranging signal receiving unit; as well as A pseudo-distance calculation unit calculates the pseudo-distance of each of the plurality of ranging signal output blocks based on the transmission time of each ranging signal from the plurality of ranging signal output blocks. The location inference unit infers the location of the ranging signal receiving unit based on the known locations of the plurality of ranging signal output blocks and the pseudo-distances of the plurality of ranging signal output blocks through machine learning.
13. The procedure according to claim 12, wherein, The transmission time calculation unit includes: A cross-correlation calculation unit calculates the cross-correlation between the spreading code signal in the ranging signal received by the ranging signal receiving unit and the spreading code signal of the ranging signal output from the plurality of ranging signal output blocks; and A peak detection unit detects the peak time in the cross-correlation as the transmission time; The pseudo-distance calculation unit calculates the pseudo-distance of each of the plurality of ranging signal output blocks based on the transmission time of each ranging signal from the plurality of ranging signal output blocks, and The location inference unit infers the location of the ranging signal receiving unit based on the pseudo distance to the plurality of ranging signal output blocks and the known locations of the plurality of ranging signal output blocks through machine learning.
14. The procedure according to claim 1, wherein, The ranging signal receiving unit is installed on a smartphone or HMD (head-mounted display).
15. The procedure according to claim 1, wherein, The ranging signal includes audio signals, radio wave signals, and optical signals.
16. An information processing apparatus, comprising: A ranging signal receiving unit receives a ranging signal, the ranging signal including a spreading code signal obtained by performing spreading modulation on the spreading code and output from a plurality of ranging signal output blocks existing at a known location; and A position inference unit, which infers the position of the ranging signal receiving unit by machine learning based on the pseudo distance to the plurality of ranging signal output blocks and the known positions of the plurality of ranging signal output blocks. The pseudo distance is determined based on the transmission time, which is the time until the ranging signals of the plurality of ranging signal output blocks are transmitted to and received by the ranging signal receiving unit.
17. An information processing method of an information processing apparatus, the information processing apparatus comprising: A ranging signal receiving unit receives a ranging signal, the ranging signal including a spreading code signal obtained by performing spreading modulation on the spreading code and output from a plurality of ranging signal output blocks existing at a known location; and A position inference unit infers the position of the ranging signal receiving unit based on pseudo-distances to the plurality of ranging signal output blocks and the known positions of the plurality of ranging signal output blocks, using machine learning. The pseudo-distances are determined based on transmission time, which is the time until the ranging signals from the plurality of ranging signal output blocks are transmitted to and received by the ranging signal receiving unit. The information processing method includes the following steps: the position inference unit infers the position of the ranging signal receiving unit based on the pseudo distance and the known positions of the plurality of ranging signal output blocks through machine learning.
Citation Information
Patent Citations
Information processing apparatus, information processing method, and program
WO2022163801A1