Control method, control device, and program
The control method synchronizes presentation sound data with real sounds by considering environmental factors, addressing interference issues and enhancing user experience in acoustic XR systems.
Patent Information
- Application Number
- PCT/JP2024/020851
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-06-07
- Publication Date
- 2025-12-11
AI Technical Summary
Existing acoustic XR systems fail to provide optimal audio experiences for users due to variations in how real sounds are heard based on distance, direction, and environmental factors, leading to interference with real sounds and inability to synchronize presentation sound data with real sounds effectively.
A control method that senses the location of real sounds and user environmental information, controlling sound emission based on the relationship between a microphone and the user to synchronize presentation sound data with real sounds, considering factors like distance and transmission characteristics.
Provides optimal acoustic XR experiences tailored to each user's environment by synchronizing presentation sound data with real sounds, enhancing user understanding and immersion.
Smart Images

Figure JP2024020851_11122025_PF_FP_ABST
Abstract
Description
Control method, control device, and program
[0001] The present invention relates to acoustic XR technology that combines real sounds with sounds emitted from a sound output device.
[0002] With the spread of open-ear acoustic devices, attention is being focused on acoustic XR (Extended Reality or Cross Reality), which allows users to hear real sounds from their surroundings while playing virtual sounds from the acoustic device depending on their location, providing an experience that goes beyond hearing real sounds.
[0003] In exhibition venues such as art museums and aquariums, audio XR is used as audio guidance to introduce exhibits. Sensors installed in the venue communicate with devices such as smartphones owned by users to detect when the user approaches an exhibit, and audio guidance about the exhibit is played on an open audio device via the smartphone or other device (see Non-Patent Documents 1 and 2). This allows users to enjoy the realistic sounds of the exhibits themselves, such as digital art and the sounds of living creatures, while the audio guidance is played in their ears, deepening their understanding of the exhibits.
[0004] Additionally, efforts are being made to deepen users' understanding of classical music by having them listen to commentary on the music being played through open-ear audio devices while watching an orchestra perform at a concert venue.
[0005] In existing systems, the combination of the area where the sound will be played and the presentation sound data is set in advance, and when a user is detected within the area, the presentation sound data is delivered to a device such as a smartphone owned by the user, and when the presentation sound data is received, the presentation sound is played from the audio device.
[0006] "ABOUT WORK OUR PROJECT Project Introduction: Development of a Location-Linked Audio Guide App for Art Exhibitions - An Audio Guide App Providing a New Viewing Experience," [online], Techfirm Inc., [Retrieved May 29, 2024], Internet <URL: https: / / career.techfirm.co.jp / project / museum / > "The teamLab Guide App displays location-linked exhibit explanations and videos on smartphones, also for infection prevention measures and audio guides," [online], Techfirm Inc., [Retrieved May 29, 2024], Internet <URL: https: / / creatorzine.jp / news / detail / 1127>
[0007] However, the time when real sounds from exhibits, performers, etc. reach the user's ears and how they are heard varies depending on the distance between the user and the location where the real sounds are generated, the direction of the user's head, and the surrounding environment (noise, reverberation characteristics of structures, etc.). Therefore, simply delivering and playing presentation sound data based on whether the user is present in the playback area does not allow for timing control of the time when real sounds reach the user or for correcting differences in how they are heard due to differences in the surrounding environment.
[0008] In addition, since the combination of the area in which the sound will be played and the presentation sound data is set in advance, there may be cases where the user is unable to request the desired presentation sound, or where the presentation sound gets in the way and prevents the user from hearing the real sounds occurring in the real space.
[0009] As such, existing systems may not be able to provide the optimal acoustic XR experience for each user's environment, in line with user requests and real sounds occurring in the real space.
[0010] The purpose of this invention is to provide optimal acoustic XR for each user's environment by sensing the location of real sounds occurring in real space and the user's environmental information (position, head direction, surrounding environment, etc.), and then sending playback information for each user calculated from this information together with presentation sound data in response to the user's request.
[0011] In order to solve the above problem, according to one aspect of the present invention, a control method makes external sounds audible and controls an acoustic signal heard by a user wearing a sound emitting device worn on the user's head. The control method acquires a relationship between a predetermined microphone and the user, and controls the sound emitting device to emit a predetermined acoustic signal based on the relationship, the relationship including at least one of distance and transmission characteristics, and the predetermined acoustic signal takes into account the time until the direct sound of the acoustic signal emitted from a predetermined sound source reaches the user and the time for the predetermined acoustic signal to be emitted.
[0012] The present invention has the effect of being able to provide optimal acoustic XR for each user's environment.
[0013] FIG. 1 is a diagram showing an example of the configuration of an audio distribution system according to a first embodiment. FIG. 2 is a diagram showing an example of the processing flow of the audio distribution system according to the first embodiment. FIG. 3 is a functional block diagram of a server according to the first embodiment. FIG. 4 is a diagram showing an example of data stored in a storage unit. FIG. 5 is a diagram showing an example of the processing flow of an audio distribution system according to a second embodiment. FIG. 6 is a diagram showing an example of the processing flow of an audio distribution system according to the second embodiment. FIG. 7 is a functional block diagram of a server according to the second embodiment. FIG. 8 is a diagram showing an example of the configuration of a computer to which the present technique is applied.
[0014] Hereinafter, an embodiment of the present invention will be described. In the drawings used in the following description, components having the same functions and steps performing the same processes are denoted by the same reference numerals, and redundant description will be omitted.
[0015] <Audio Distribution System According to First Embodiment> In the first embodiment, presentation sound data requested by a user is distributed while the location of real sounds generated in real space and the user's location are fixed. Fig. 1 shows an example of the configuration of an audio distribution system according to this embodiment, and Fig. 2 shows an example of the processing flow. The presentation sound data is obtained by applying predetermined acoustic processing to the output signal of a sound pickup microphone, and it can be said that the user is requesting the output signal of the sound pickup microphone.
[0016] The sound distribution system includes N user control terminals 110-n, N acoustic devices 120-n, M sound pickup microphones 130-m, a server 140, and a position measurement sensor 150. N and M are each an integer equal to or greater than 1, where n=1, 2, ..., N and m=1, 2, ..., M. M is the number of sound pickup microphones installed within the target area and corresponds to the total number of anticipated real sound generation positions, and m is the identifier of the sound pickup microphone. N is the number of users present within the target area, and n indicates the user's identifier.
[0017] The user control terminal 110-n and the acoustic device 120-n are connected to each other by wire or wirelessly so as to be able to communicate with each other. The server 140 is connected to the N user control terminals 110-n, the M sound pickup microphones 130-m, and the position measurement sensor 150 so as to be able to communicate with each other, either directly or via a communication line.
[0018] <User control terminal 110-n and acoustic device 120-n> Each user n has one set of user control terminal 110-n and open-ear acoustic device 120-n, and is assumed to be stationary at any position within the target area. Furthermore, the position 10-m where real sounds are generated in the real space is fixed, and a sound pickup microphone 130-m is installed at that position.
[0019] The user control terminal 110-n' transmits a user request to the server 140 in accordance with the operation of the user n', receives playback information and presentation sound data from the server 140, and outputs the presentation sound data, which has undergone predetermined processing using the playback information, to the acoustic device 120-n' at a predetermined timing. For example, the user control terminal 110-n' may be a smartphone or the like, and functions as the user control terminal 110-n' of this embodiment by installing an application that performs predetermined processing. The user request includes the identifier n' of the user control terminal and the identifier m' of the sound pickup microphone. n' is the identifier of the user control terminal included in the user request and is one of 1, 2, ..., N, and m' is the identifier of the sound pickup microphone included in the user request and is one of 1, 2, ..., M.
[0020] The acoustic device 120-n makes external sounds audible, includes a sound emitting device worn on the head of the user n, and plays back the received presentation sound data. For example, the acoustic device 120-n is formed of an open-ear earphone or the like.
[0021] In other words, user n' specifies the sound he or she wants to hear (a real sound, or a sound pickup microphone 130-m' installed at the location 10-m' where the sound is being generated) via the user control terminal 110-n' at any time, and requests the server 140 to play that sound. Upon receiving the request, the server 140 transmits playback information calculated from the environmental information of user n' and presentation sound data to the user control terminal 110-n'. User n' can hear the presentation sound data corresponding to that environment via the acoustic device 120-n'.
[0022] <Position Measurement Sensor 150> The position measurement sensor 150 acquires and outputs information (hereinafter also referred to as "sensor information") used to measure the positions of N user control terminals 110-n and M sound pickup microphones 130-m present within a target area. For example, the position measurement sensor 150 measures distance based on differences in the strength and arrival time of the signal emitted by the position measurement sensor 150, and uses WiFi, RFID, Bluetooth, sound waves, UWB, GPS, etc. The position measurement sensor 150 is installed at a location where sensor information can be acquired, for example, within the environment of the target area. For example, any position measurement method may be used as long as the sound pickup microphone 130-m or user n to be measured is equipped with or wears a sensor that receives signals from the position measurement sensor 150. Note that the sensor information may not only be information used to measure the positions of the N user control terminals 110-n and M sound pickup microphones 130-m, but may also be information indicating the positions themselves.
[0023] <Server 140> FIG. 3 shows a functional block diagram of the server 140.
[0024] The server 140 includes a position measurement unit 141, a playback time calculation unit 142, a transfer characteristic calculation unit 143, a sound acquisition unit 144, a target sound setting unit 145, an acoustic processing control unit 146, a presentation sound generation unit 147, a distribution control unit 148, and a memory unit 149.
[0025] In this embodiment, the server 140 performs an initial process prior to the distribution process.
[0026] In the initial processing, the server 140 receives the sensor information and the output signal of the microphone of the acoustic device as input, obtains playback information, and stores it in the storage unit 149 .
[0027] In the distribution process, the server 140 receives the output signal from the sound pickup microphone and a user request, performs predetermined acoustic processing, and outputs the signal to the user control terminal 110-n. The output signal from the sound pickup microphone is assumed to have a time code attached to it indicating the time it was acquired. However, if no time code is attached, the server 140 adds a time code to the output signal from the sound pickup microphone indicating the time the server 140 received the output signal from the sound pickup microphone.
[0028] The server 140 is a special device configured by loading a special program into a publicly known or dedicated computer having, for example, a central processing unit (CPU) and a main memory (RAM). The server 140 executes each process under the control of the central processing unit. Data input to the server 140 and data obtained from each process are stored in, for example, the main memory, and the data stored in the main memory is read by the central processing unit as needed and used for other processes. At least a portion of each processing unit of the server 140 may be configured with hardware such as an integrated circuit. Each storage unit of the server 140 may be configured with, for example, a main storage unit such as RAM (Random Access Memory) or middleware such as a relational database or key-value store. However, each storage unit does not necessarily have to be internal to the server 140; it may be configured with an auxiliary storage unit configured with a hard disk, optical disk, or semiconductor memory element such as flash memory, and be configured external to the server 140.
[0029] Each part will be explained below.
[0030] First, the initial processing will be described.
[0031] <Position measurement unit 141> The position measurement unit 141 receives sensor information from the position measurement sensor 150, and uses the sensor information to measure the positions of the sound pickup microphone 130-m and user n within the target area (S141), and outputs the position of the sound pickup microphone 130-m within the target area. m =(x m ,y m ,z m ), m∈M, the user's position is user n =(x n ,y n ,z n ),n∈N.
[0032] <Transfer characteristic calculation unit 143> The transfer characteristic calculation unit 143 calculates the transfer characteristic of the sound that is generated in the real space and that is caused by the surrounding environment until it reaches the user n (S143), and stores the calculated transfer characteristic in the storage unit 149. For example, using existing techniques such as the M-sequence method or the TSP method, a measurement signal is output from the speaker 20-m placed near the sound collection microphone 130-m in order from the sound collection microphones 130-1 to 130-M, and a transfer function TF is calculated from the output signal (observation signal) of the microphone of the acoustic device 120-n worn by the user n. nm Therefore, when calculating the transfer function using this method, the acoustic device 120-n includes a microphone. Note that if the acoustic device 120-n does not include a microphone but the user control terminal 110-n does include a microphone, similar processing may be performed using the output signal from the microphone of the user control terminal 110-n.
[0033] <Playback time calculation unit 142> The playback time calculation unit 142 calculates the playback time at the position mic of the sound pickup microphone 130-m. m and the user's location n In order to control the timing of playing the presentation sound data at the user control terminal 110-n, the arrival time AT nm (S142), and the calculated arrival time AT nm is stored in the storage unit 149. Arrival time AT nmis calculated from the distance from the sound pickup microphone 130-m to the user n as follows: c is the speed of sound and t is the gas temperature.
[0034] The user control terminal 110-n inputs the arrival time AT nm By playing the presentation sound data at the time added together, it is possible to play it at the same timing as the real sound generated in the real space.
[0035] It is assumed that the times of the sound pickup microphone 130-m and the user control terminal 110-n are completely synchronized.
[0036] <Storage Unit 149> The storage unit 149 stores a transfer function TF nm and arrival time AT nm 4 shows an example of data stored in the storage unit 149. The arrival time AT nm , transfer function TF nm is calculated only the first time and stored in the database. This allows the arrival time AT n'm' , transfer function TF n'm' Therefore, it is not necessary to calculate the above, and the amount of processing can be reduced.
[0037] When the initial processing is completed, the acquisition of sound data from the sound pickup microphone 130-m (S144) and the reception of a request (user request) for a sound to be heard from the user control terminal 110-n (S145) are started.
[0038] <Sound Acquisition Unit 144> The sound acquisition unit 144 always acquires the output signal (sound data) sound_orginal of each sound pickup microphone 130-n, regardless of whether or not there is a request from the user control terminal 110-n. m (S144) and transmits it to the acoustic processing control unit 146.
[0039] <Objective sound setting unit 145> Upon receiving a request from user n' (S145), the objective sound setting unit 145 transmits the user's identifier n' and the identifier m' of the sound pickup microphone corresponding to the request from user n' to the acoustic processing control unit 146. At this time, the sound that the user wants to hear (real sound) and the sound pickup microphone 130-m (the sound pickup microphone installed at the generation position 10-m of the real sound) can be matched within the server 140, and the identifier m' of the sound pickup microphone may be specified directly by user n' via the user control terminal 110-n', or may be specified by referring to a table that associates the UI (User Interface) of the user control terminal 110-n with the identifier m of the sound pickup microphone 130-m, which is held within the server 140.
[0040] <Sound processing control unit 146> The sound processing control unit 146 always receives M pieces of sound data sound_orginal from the sound acquisition unit 144. m Furthermore, the sound processing control unit 146 receives the user identifier n' and the sound collection microphone identifier m' from the objective sound setting unit 145, and transmits the user identifier n', the sound collection microphone identifier m', and sound data sound_orginal collected by the sound collection microphone 130-m' corresponding to the received identifier m' to the presentation sound generation unit 147 in the order of reception. m' and requests the generation of presentation sound data (S146). The acoustic processing control unit 146 continues to request the presentation sound generation unit 147 to perform processing until a request to change the identifier m′ of the sound pickup microphone or a request to stop distribution is received from the objective sound setting unit 145.
[0041] <Presentation Sound Generator 147> The presentation sound generator 147 receives the user identifier n', the sound pickup microphone identifier m', and sound data sound_orginal m' Receive the sound data sound_orginal m'The system performs predetermined acoustic processing on the input signal n', m' to generate a sound (presentation sound) to be presented to the user (S147), and outputs the identifiers n', m' and the generated presentation sound data. In this system, the acoustic processing used to generate the presentation sound is not limited. Acoustic processing may be set in advance for each sound pickup microphone, and acoustic processing may be performed according to the sound pickup microphone requested by the user, acoustic processing specified by the user's request, or acoustic processing may be performed according to the current state. Note that this also includes the option of not performing acoustic processing.
[0042] For example, possible acoustic processing includes voice enhancement, noise reduction, echo reduction, dereverberation, and suppression of specific frequency components. By performing such processing, the user can hear real sounds that include noise, etc., and also hear clear sounds that have been removed from the acoustic device 120-n. The user can easily recognize voices, etc., while experiencing real sounds that include noise and echoes.
[0043] Other possible acoustic processing techniques include adding noise, adding reverberation, adding reverberation, and emphasizing specific frequency components. By performing such processing, the user can hear real sounds as well as sounds reproduced by the acoustic device 120-n that have been added with noise or the like that occurs in a specific environment. While listening to real sounds, the user can experience being in a specific virtual environment. For example, by reproducing sounds that have been added with reverberation that occurs in a cave or the like, the user can be given the sensation of being inside a cave.
[0044] Another possible acoustic processing method is to separate sound sources and generate and play a signal that is out of phase with the sound emitted from a specific sound source. By performing such processing, a user can cancel out the sound emitted from a specific sound source from real sounds. For example, if the loud noise of a campaign car driving outside bothers the user while working at home, the user can play a sound that cancels out the noise of the campaign car, thereby making the unpleasant sound inaudible from real sounds.
[0045] <Distribution control unit 148> The distribution control unit 148 acquires the relationship between a predetermined microphone and a user. This relationship includes at least one of distance and transmission characteristics. In this embodiment, the relationship is determined by the arrival time AT n'm' and the transfer function TF n'm' Since the arrival time can be calculated from the distance using equation (1), and conversely, the distance can be calculated from the arrival time, the relationship in this embodiment can be said to include the distance and the transfer characteristic.
[0046] For example, the distribution control unit 148 receives the identifiers n′ and m′ and the presentation sound data, and based on the identifiers n′ and m′, extracts the arrival time AT n'm' and the transfer function TF n'm' (S148), and transmits the presentation sound data and reproduction information (arrival time, transfer function) received from the presentation sound generation unit 147 to the user control terminal 110-n' of user n'.
[0047] The user control terminal 110-n' receives the presentation sound data and the reproduction information (arrival time, transfer function), and applies the transfer function TF n'm' The user control terminal 110-n applies a transfer function TF n'm' The presentation sound data to which the time code of the presentation sound data is applied is calculated as the arrival time AT n'm' and reproduces the audio via the audio device 120-n when the time on the user control terminal 110-n matches the time on the user control terminal 110-n.
[0048] With this configuration, the server 140 can control the sound output device included in the acoustic device to output a predetermined acoustic signal (an acoustic signal that has been subjected to acoustic processing by the presentation sound generation unit 147) based on the relationship (arrival time and transfer function) between the user n' and a predetermined microphone (a sound collection microphone corresponding to a request from the user n'). The predetermined acoustic signal is output over a time period (arrival time AT n'm' ) and the time period during which this predetermined acoustic signal is emitted are taken into consideration.
[0049] <Effects> In the first embodiment, while the position of the real sound generated in the real space and the position of the user are fixed, a presentation sound that has been acoustically processed from the sound requested by the user is played back at the time the real sound generated in the real space reaches the user, making it possible to provide optimal acoustic XR for each user's environment.
[0050] The server 140 controls the audio signals (presentation sound data) listened to by the user wearing the audio device, and is therefore also called a control device. The audio distribution system of this embodiment is also called a control system.
[0051] <Modification> In this embodiment, the arrival time AT n'm' and the transfer function TF n'm' For example, when the acoustic processing performed by the presentation sound generation unit 147 is echoing or reverberation and a transfer function is added to the presentation sound data itself, the transfer function TF n'm' may not be applied.
[0052] Second Embodiment The following description will focus on the differences from the first embodiment.
[0053] In the second embodiment, a configuration will be described in which a sound requested by a user is delivered while the position of real sounds occurring in the real space and the position of the user change.
[0054] The overall configuration of this embodiment is similar to that of the first embodiment, but differs in that the position of the sound pickup microphone 130-m changes in accordance with the generation position 10-m of the real sound generated in the real space, and that user n can move to any position within the target area.
[0055] <Server 240> FIGS. 5 and 6 show an example of a processing flow of the sound distribution system according to this embodiment, and FIG. 7 is a functional block diagram of the server 240 according to this embodiment.
[0056] The server 240 includes a position measurement unit 241, a playback time calculation unit 142, a transfer characteristic calculation unit 243, a sound acquisition unit 144, a target sound setting unit 145, an acoustic processing control unit 146, a presentation sound generation unit 147, a distribution control unit 148, and a memory unit 249.
[0057] In the second embodiment, the position of the real sound generated in the real space and the position of the user move, and so the playback time and transfer function are calculated when either of these changes, which is a difference from the first embodiment. The differences from the first embodiment mainly lie in the timing of processing in the position measurement unit, playback time calculation unit, and transfer characteristic calculation unit, and the timing of storing data in the memory unit 249. This part will be explained below.
[0058] Fig. 5 shows an example of the processing flow of the position measurement unit 241, the playback time calculation unit 142, and the transfer characteristic calculation unit 243. Fig. 6 shows an example of the processing flow of the sound acquisition unit 144, the target sound setting unit 145, the acoustic processing control unit 146, the presentation sound generation unit 147, and the distribution control unit 148. The processing (distribution processing) is the same as in the first embodiment, so a description thereof will be omitted.
[0059] <Position measurement unit 241> The position measurement unit 241 receives sensor information from the position measurement sensor 150. The position measurement unit 241 also receives the user identifier n′ and the sound collection microphone identifier m′ from the objective sound setting unit 145 (YES in S145 of FIG. 5 ), and measures the positions of the sound collection microphone 130-m′ and the user n′ within the target area at regular intervals w using the sensor information in the order received (S241-1).
[0060] The position measurement unit 241 calculates the amount of change between the position of the sound pickup microphone 130-m' and the position of the user n' (the difference from the position calculated in the previous cycle), and determines whether either amount of change is equal to or greater than the allowable error (S241-2).
[0061] If the amount of change is greater than or equal to the allowable error (YES in S241-2), the position measurement unit 241 sends a control signal to the transfer characteristic calculation unit 243 to instruct it to execute the transfer characteristic calculation process S243, and outputs a control signal to the playback time calculation unit 142 to instruct it to execute the playback time calculation process S142, as well as the positions of the sound pickup microphone 130-m' and user n'.
[0062] The transfer characteristic calculation unit 243 executes the transfer characteristic calculation process S243, the reproduction time calculation unit 142 executes the reproduction time calculation process S142, and the transfer function TF stored in the storage unit 249 using the newly obtained transfer characteristic and arrival time is calculated.n'm' and arrival time AT n'm' The playback time calculation process S142 in the playback time calculation unit 142 is the same as in the first embodiment, so a description thereof will be omitted. In the first embodiment, the arrival time AT nm , transfer function TF nm is calculated initially and stored in a database. However, in this embodiment, it is assumed that the position of the real sound generated in the real space and the position of the user change. Therefore, when a sound requested by a user is delivered, the arrival time AT n'm' , transfer function TF n'm' is calculated at regular intervals.
[0063] If the amount of change is less than the allowable error (NO in S241-2), the position measurement unit 241 does not perform the transfer characteristic calculation process S243 and the reproduction time calculation process S142, and therefore does not transmit a control signal for executing each process.
[0064] In the case of the first position measurement process S241-1, since there is no position calculated in the previous cycle, the position measurement unit 241 performs the same process as when the amount of change is equal to or greater than the allowable error (YES in S241-2).
[0065] As described above, the position measurement unit 241 performs the position measurement process S241-1 at regular intervals w (S241-3).
[0066] When the position measurement unit 241 receives a request to change the identifier of the sound pickup microphone 130-m from the objective sound setting unit 145 (S241-4), it performs a position measurement process S241-1.
[0067] When the position measurement unit 241 receives a stop request from the objective sound setting unit 145 (S241-5), it ends the position measurement process S241-1.
[0068] <Transfer characteristic calculation unit 243> The transfer characteristic calculation unit 243 calculates the transfer characteristic of sound caused by the surrounding environment until a real sound generated in a real space reaches the user n′ (S243), and stores the calculated transfer characteristic in the memory unit 249.
[0069] In this embodiment, it may be difficult to install the speaker 20-m' because the position of the sound pickup microphone 130-m' changes. In such cases, a transfer function is calculated from the sound data observed by the sound pickup microphone 130-m' using a method of estimating a transfer function from only the observed signal (see, for example, Reference 1).
[0070] (Reference 1) YOICHI SAT0, "A Method of Self-Recovering Equalization for Multilevel Amplitude-Modulation Systems", [online], IEEE TRANSACTIONS ON COMMUNICATIONS, [Retrieved June 3, 2024], Internet<URL:https: / / faratarjome.ir / u / media / shopping_files / store-EN-1484205766-4349.pdf> <Effects> With this configuration, it is possible to obtain the same effects as in the first embodiment. Furthermore, in this embodiment, it is possible to flexibly respond to cases where the position of a real sound generated in the real space and the position of the user change.
[0071] Third Embodiment The following description will focus on the differences from the first and second embodiments.
[0072] In the first and second embodiments, the user control terminal 110-n plays the presentation sound via the acoustic device 120-n at the same time that the real sound generated in the real space reaches the user, but the configuration may be such that the presentation sound is played before the real sound generated in the real space reaches the user (in other words, the configuration is such that the time at which the real sound (direct sound) reaches the user is controlled to be later than the time at which a predetermined acoustic signal (presentation sound data) is emitted). In this case, the playback timing control parameter α is used to control the arrival time to AT n'm' From AT n'm' -α. This change process may be performed by the playback time calculation unit 142, the distribution control unit 148, or the user control terminal 110-n. α is 0 or more and AT n'm' It can be set to the following values:
[0073] This configuration can achieve the same effect as in the first embodiment. Furthermore, the configuration of this embodiment allows the user to hear the reproduced sound before the real sound reaches the user, making it possible to provide more flexible audio XR. For example, if the real sound is one that alerts the user to danger (such as a scream, gunshot, or explosion), the user can be notified of the danger before the real sound reaches the user.
[0074] <Modification> A configuration may be adopted in which the presentation sound is played after a real sound generated in the real space reaches the user. In this case, the playback timing control parameter β is used to determine the arrival time relative to the AT n'm' From AT n'm' +β. α can be set to a value greater than or equal to 0.
[0075] <Fourth embodiment> In the first to third embodiments, sound requested by a user is subjected to acoustic processing and then distributed, but in this embodiment, the system determines the acoustic processing to be performed by the acoustic processing control unit 146 depending on the surrounding environment and the user's situation.
[0076] When an incident occurs in the city, even if a witness warns people to run away, people a short distance away may not hear due to noise or may not be paying attention. In this embodiment, such situations are automatically detected and a warning sound is played regardless of whether the user requests it.
[0077] In this embodiment, the acoustic processing control unit 146 executes the acoustic processing control process S146 and also performs the following processes.
[0078] The sound processing control unit 146 receives the M pieces of sound data sound_orginal m For example, sound signals expressing danger (words informing danger, screams, gunshots, explosions, etc.) are stored in advance in a storage unit (not shown), and M sound data sound_orginal are stored. m It is determined whether or not the audio signal representing a danger that is stored in the audio signal list is included in the audio signal list.
[0079] The sound processing control unit 146 generates M pieces of sound data sound_orginalm If it is determined that the sound data expressing danger is included in the sound data sound_orginal, sound processing is performed to emphasize the sound data expressing danger, regardless of whether or not the user requests it, and the sound data sound_orginal that includes the sound data expressing danger is m The sound will be broadcast to all users around the corresponding microphone.
[0080] This configuration can encourage the user to take action to avoid danger.
[0081] Furthermore, by utilizing sensors other than the sound pickup microphone and acoustic devices within the target area, the system may determine the acoustic processing depending on the surrounding environment and the user's reaction. For example, if a negative reaction such as tilting one's head or crossing one's arms is detected from the camera image within the target area when the user hears the sound presentation, it can be determined that the sound presented to the user is not appropriate, and the acoustic processing can be adjusted to improve the situation.
[0082] <Other Variations> Furthermore, a device (terminal) for using the device of the present invention, the system of the present invention, or the method of the present invention via a network (telecommunications line) may also be included. The "device (terminal) for use" may be provided with functions (e.g., control function, decoding function, restoration function, input / output function, etc.) necessary to obtain the effects of implementing the device of the present invention, the system of the present invention, or the method of the present invention. Note that a configuration including a device (terminal) for using the device of the present invention or the method of the present invention via a network (telecommunications line) is also referred to as an audio distribution system.
[0083] <Hardware, Programs, and Recording Media> The functions realized by the components described in this specification may be implemented in circuitry or processing circuitry, including general-purpose processors, application-specific processors, integrated circuits, ASICs (Application Specific Integrated Circuits), a CPU (a Central Processing Unit), conventional circuits, and / or combinations thereof, programmed to realize the described functions. A processor includes transistors and other circuits and is considered to be circuitry or processing circuitry. A processor may also be a programmed processor that executes a program stored in a memory.
[0084] In this specification, a circuitry, unit, or means is hardware that is programmed to realize or performs the described functions, which may be any hardware disclosed herein or any hardware known to be programmed to realize or perform the described functions.
[0085] If the hardware is a processor considered to be a type of circuitry, the circuitry, means, or unit is a combination of the hardware and software used to configure the hardware and / or processor.
[0086] The various processes described above can be implemented by loading a program that executes each step of the above method into the recording unit 2020 of the computer 2000 shown in Figure 8, and operating the control unit 2010, input unit 2030, output unit 2040, display unit 2050, etc.
[0087] The program describing the processing contents can be recorded on a computer-readable recording medium, which may be, for example, a magnetic recording device, an optical disk, a magneto-optical recording medium, a semiconductor memory, or any other suitable recording medium.
[0088] The program may be distributed by, for example, selling, transferring, lending, etc. portable recording media such as DVDs and CD-ROMs on which the program is recorded. Furthermore, the program may be stored in a storage device of a server computer, and then transferred from the server computer to other computers via a network, thereby distributing the program.
[0089] A computer that executes such a program may first temporarily store the program recorded on a portable recording medium or transferred from a server computer in its own storage device. Then, when executing a process, the computer reads the program stored on its own recording medium and executes the process in accordance with the read program. Alternatively, the computer may read the program directly from a portable recording medium and execute the process in accordance with the program. Furthermore, the computer may execute the process in accordance with the program each time a program is transferred from a server computer to the computer. Alternatively, the server computer may not transfer the program to the computer, but may instead execute the process through a so-called ASP (Application Service Provider) service, which realizes the processing function by issuing an execution instruction and obtaining the results. Furthermore, the server computer may execute the process at the terminal using a so-called SaaS (Software as a Service) service, which allows users to use part of a server computer along with the program. In this embodiment, the program includes information used for processing by an electronic computer that is equivalent to a program (such as data that is not a direct instruction to a computer but has properties that dictate computer processing).
[0090] Furthermore, in this embodiment, the device is configured by executing a predetermined program on a computer, but at least a part of the processing contents may be realized by hardware.
[0091] <Other Modifications> The present invention is not limited to the above-described embodiments and modifications. For example, the various processes described above may not only be executed in chronological order as described, but may also be executed in parallel or individually depending on the processing capacity of the device that executes the processes or as needed. Other modifications are possible as long as they do not deviate from the spirit of the present invention.
Claims
1. A control method for making external sounds audible and controlling an acoustic signal heard by a user wearing a sound emission device worn on the user's head, the control method comprising: obtaining a relationship between a predetermined microphone and the user; and controlling the sound emission device to emit a predetermined acoustic signal based on the relationship, the relationship including at least one of distance and transmission characteristics; and the predetermined acoustic signal taking into account the time it takes for the direct sound of the acoustic signal emitted from a predetermined sound source to reach the user and the time at which the predetermined acoustic signal is emitted.
2. A control method according to claim 1, wherein the time at which the direct sound reaches the user is controlled to be later than the time at which the predetermined acoustic signal is emitted.
3. A control method according to claim 1, comprising: a transfer characteristic calculation step for calculating the transfer characteristic from the specified microphone to the user; an arrival time calculation step for calculating the arrival time for sound generated at the position of the specified microphone to reach the user; a presentation sound generation step for performing specified acoustic processing on the output signal of the specified microphone to generate the specified acoustic signal; and a reproduction step for applying a transfer function representing the transfer characteristic to the specified acoustic signal, and controlling the sound emission device to emit the specified acoustic signal at the time obtained by adding the arrival time to the acquisition time of the direct sound.
4. A control method according to claim 3, wherein the position of the specified microphone and the position of the user are fixed, and the transfer characteristic calculation step and the arrival time calculation step are performed only once before the presentation sound generation step and the playback step are executed.
5. A control method according to claim 3, wherein the position of the predetermined microphone and the position of the user change, and the control method includes a position measurement step of measuring the position of the predetermined microphone and the position of the user at regular intervals, and when the amount of change in the position of the predetermined microphone or the amount of change in the position of the user is equal to or greater than an allowable error, the control method performs the transfer characteristic calculation step and the arrival time calculation step, and updates the transfer characteristic and the arrival time.
6. A control device that makes external sounds audible and controls an acoustic signal heard by a user wearing a sound emission device worn on the user's head, the control device acquiring a relationship between a predetermined microphone and the user, and controlling the sound emission device to emit a predetermined acoustic signal based on the relationship, the relationship including at least one of distance and transmission characteristics, and the predetermined acoustic signal taking into account the time it takes for the direct sound of the acoustic signal emitted from a predetermined sound source to reach the user and the time at which the predetermined acoustic signal is emitted.
7. A program for causing a computer to execute the control method of any one of claims 1 to 5.
Citation Information
Patent Citations
Sound source management device, sound source management method, and sound source management system
JP2014236259A
Sound reproduction device and sound reproduction method
WO2012042905A1