Information processing device, information processing method, and program

By employing a central sound source and peripheral sound sources with HRTF-based processing, the technology effectively addresses the challenge of representing distance in spatial audio, providing a realistic experience.

JP7786451B2Active Publication Date: 2025-12-16SONY GROUP CORP
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2023503608
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-03-05
Filing Date
2022-01-13
Publication Date
2025-12-16
Estimated Expiration
2042-01-13

AI Technical Summary

Technical Problem

Conventional methods struggle to accurately represent the sense of distance from a user to a virtual sound source using head-related transfer functions (HRTF), making it difficult to create a realistic spatial audio experience.

Method used

The technology employs a central sound source and multiple peripheral sound sources positioned according to the size of the sound image, using HRTF information for convolution processing to output sound data that adjusts the perceived distance by controlling the size of the sound image.

Benefits of technology

This approach allows users to perceive distance to virtual sound sources realistically through spatial audio, enhancing the sense of presence and immersion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007786451000001
    Figure 0007786451000001
  • Figure 0007786451000002
    Figure 0007786451000002
  • Figure 0007786451000003
    Figure 0007786451000003
Patent Text Reader

Abstract

The present invention relates to an information processing device, an information processing method, and a program that make it possible to suitably reproduce the feeling of distance from a user to a virtual sound source and the apparent size of a virtual sound source in spatial acoustic expression. This information processing device comprises: a first sound source; a sound source setting unit for setting a plurality of second sound sources in positions corresponding to the size of a sound image of first sound which is sound from the first sound source; and an output control unit that causes first sound data obtained by convolution processing using HRTF information corresponding to the position of the first sound source and a plurality of pieces of second sound data obtained by convolution processing using HRTF processing corresponding to the positions of the each of the second sound sources to be output. Each of the second sound sources is set so as to be positioned around the first sound source. The present invention can be applied to devices that cause sound to be output from playback devices such as headphones.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] In particular, the present technology relates to an information processing device, an information processing method, and a program that are capable of appropriately reproducing the sense of distance from a user to a virtual sound source and the apparent size of the virtual sound source in a spatial audio representation. [Background technology]

[0002] As a method for using sound to allow a user to recognize a space, there is a method for expressing the direction, distance, movement, etc. of a virtual sound source by calculation using a head-related transfer function (HRTF). [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2010-004512 Summary of the Invention [Problem to be solved by the invention]

[0004] In order to allow the user to perceive space using sound, it is important to represent the direction and distance of a virtual sound source. Although the direction of a virtual sound source can be expressed by calculation using HRTF, it is difficult to adequately represent the sense of distance from the user to the virtual sound source using conventional methods.

[0005] The present technology has been made in consideration of such circumstances, and makes it possible to appropriately reproduce the sense of distance from the user to a virtual sound source and the apparent size of the virtual sound source. [Means for solving the problem]

[0006] An information processing device according to one aspect of the present technology includes: a first sound source; a sound source setting unit that sets a plurality of second sound sources at positions according to the size of a sound image of the first sound, which is the sound of the first sound source; and an output control unit that outputs first sound data obtained by convolution processing using HRTF information according to the position of the first sound source, and a plurality of second sound data obtained by convolution processing using HRTF information according to the positions of each of the second sound sources, wherein each of the second sound sources is set to be located in the vicinity of the first sound source.

[0007] In one aspect of the present technology, a first sound source and a plurality of second sound sources are set at positions corresponding to the size of a sound image of the first sound, which is the sound of the first sound source, and first sound data obtained by a convolution process using HRTF information corresponding to the position of the first sound source and a plurality of second sound data obtained by a convolution process using HRTF information corresponding to the positions of each of the second sound sources are output. Each of the second sound sources is set to be located in the periphery of the first sound source. [Brief explanation of the drawings]

[0008] [Figure 1] FIG. 1 is a diagram illustrating an example of how a listener perceives sound. [Figure 2] 1A and 1B are diagrams illustrating examples of distance representation in the present technology. [Figure 3] FIG. 1 is a diagram showing the positional relationship between a central sound source and a user. [Figure 4] FIG. 1 is a diagram showing the positional relationship between a central sound source and peripheral sound sources. [Figure 5] FIG. 10 is another diagram showing the positional relationship between the central sound source and the peripheral sound sources. [Figure 6] 10A and 10B are diagrams showing other examples of distance representation in the present technology. [Figure 7] 1A and 1B are diagrams illustrating the shape of a sound image in the present technology. [Figure 8] 1 is a diagram illustrating an example of the configuration of an audio reproduction system to which the present technology is applied. [Figure 9] 1 is a block diagram illustrating an example of the hardware configuration of an information processing device 10. FIG. [Figure 10] 1 is a block diagram showing an example of a functional configuration of an information processing device 10. FIG. [Figure 11] 10 is a flowchart illustrating processing of the information processing device 10. [Figure 12] FIG. 10 is a diagram illustrating another example configuration of an audio reproduction system to which the present technology is applied. [Figure 13] FIG. 10 is a diagram illustrating an example of a method for notifying an obstacle to which the present technology is applied. [Figure 14] FIG. 10 is another diagram showing an example of an obstacle notification method to which the present technology is applied. [Figure 15] 10A and 10B are diagrams illustrating an example of a method for notifying a distance to a destination to which the present technology is applied. [Figure 16] 1 is a diagram illustrating an example of a notification method for an alarm sound of a home appliance to which the present technology is applied. [Figure 17] FIG. 1 illustrates an example of the configuration of a remote conference system. [Figure 18] 10A and 10B are diagrams illustrating examples of screen displays that serve as user interfaces during a remote conference. [Figure 19] FIG. 10 is a diagram illustrating an example of the size of a sound image of each user's voice. [Figure 20] FIG. 10 is a diagram illustrating an example of a method for notifying a user of a pseudo engine sound of a vehicle. [Figure 21] FIG. 1 is a diagram illustrating an example of a playback device. [Figure 22] FIG. 10 is a diagram illustrating another example of a playback device. DETAILED DESCRIPTION OF THE INVENTION

[0009] Hereinafter, embodiments of the present technology will be described in the following order. 1.Explanation of how we perceive sound 2. Distance expression using multiple sound sources 3. Example of configuration of sound reproduction system and information processing device 4. Explanation of the operation of the information processing device 5. Modifications (Application Examples) 6. Other Examples

[0010] <1. Explaining how we perceive sound> FIG. 1 is a diagram showing an example of how a listener perceives sound.

[0011] In Figure 1, a car is shown as the object that is the sound source. The car is moving while emitting sounds such as engine noise and running noise. The way the user, who is the listener, perceives the sound changes depending on the distance from the car.

[0012] In the example of A in Figure 1, a car is located far away from the user. In this case, the user perceives the sound from the car as a point sound source. In the example of A in Figure 1, the point sound source perceived by the user is represented by a colored small circle #1.

[0013] On the other hand, in the example of B in Figure 1, a car is located close to the user. In this case, the user perceives the sound from the car as having a loudness represented by the colored circle #2 surrounding the car. In this specification, the apparent loudness of a sound perceived by the user is referred to as the loudness of the sound image.

[0014] In this way, the user perceives the distance to the sound source by sensing the size of the sound image.

[0015] <2. Distance expression using multiple sound sources> FIG. 2 is a diagram showing an example of distance expression in the present technology.

[0016] In this technology, the distance from the user to the object that serves as the virtual sound source is expressed by controlling the size of the sound image. By changing the size of the sound image that the user hears, it becomes possible for the user to perceive the sense of distance from the virtual sound source.

[0017] 2, in this technology, a user U wears an output device such as headphones 1 and listens to sound from a car, which is a virtual sound source. The sound from the virtual sound source is played back by, for example, a smartphone carried by the user U, and output from the headphones 1.

[0018] In the example of Fig. 2, the sound of a car, which is an object corresponding to a virtual sound source, is composed of sounds from a central sound source C and four peripheral sound sources U, namely, peripheral sound sources LU, RU, LD, and RD. Here, the central sound source C and the peripheral sound source U are each virtual sound sources expressed by calculation using HRTFs. In Fig. 2, the central sound source C and the peripheral sound sources LU, RU, LD, and RD are illustrated as speakers. This is the same in other figures described later.

[0019] In this technology, sound is presented, for example, by converting the sounds from each sound source generated by calculation using head-related transfer functions (HRTFs) corresponding to the positions of the central sound source and each peripheral sound source into two-channel sounds (L / R) and outputting them from headphones 1.

[0020] The sound from the central sound source is the central sound that represents the sound of the object that is the virtual sound source, and is referred to as the central sound in this specification. The sound from the peripheral sound sources is the sound that represents the size of the sound image of the central sound, and is referred to as the peripheral sound in this specification.

[0021] As shown in Figure 2, this technology allows the user to perceive the sense of distance to an object that is a virtual sound source by changing the size of the sound image of the central sound. In this technology, the size of the sound image of the central sound is controlled by changing the positions of the peripheral sound sources.

[0022] In the example of Fig. 2, a car is shown as a virtual sound source object near the user, but the virtual sound source object may or may not be near the user. Also, the virtual sound source object may or may not have a physical body.

[0023] This technology makes it possible to make objects around the user appear as if they are the source of sound, and also makes it possible to make sounds appear as if they are coming from empty space around the user.

[0024] By listening to the central sound and multiple peripheral sounds, the user perceives the sound image of the central sound, which is the center representing the sound from the virtual sound source, as having a size as shown by the colored circle #11. As explained with reference to Figure 1, the perceived size of the sound image determines the sense of distance to the object that is the virtual sound source. Therefore, when a large sound image is presented as shown in Figure 2, the user perceives the car that is the virtual sound source as if it were nearby.

[0025] In this way, the user can perceive the sense of distance from the user to the object that is the virtual sound source in the spatial audio, and can experience a realistic spatial audio.

[0026] FIG. 3 is a diagram showing the positional relationship between the central sound source and the user.

[0027] As shown in Fig. 3, a central sound source C, which is a virtual sound source, is set at position P1, which is the center position of the sound image that the user is intended to experience. Position P1 is a position shifted by a predetermined horizontal angle Azim (d: degrees) and a vertical angle Elev (d) from the front direction of the user, for example. The distance from the user to position P1 is a predetermined distance L (m).

[0028] The central sound, which is the sound of the central sound source C, is the central sound that expresses the sound of the object that is the virtual sound source. The central sound is also used as a reference sound that allows the user to perceive the distance from the user to the virtual sound source.

[0029] A plurality of peripheral sound sources are set around the central sound source C set in this way. For example, the plurality of peripheral sound sources are arranged at equal intervals on a circumference with the central sound source C at the center.

[0030] FIG. 4 is a diagram showing the positional relationship between the central sound source and the peripheral sound sources.

[0031] As shown in FIG. 4, four peripheral sound sources LU, RU, LD, and RD are arranged around a central sound source C.

[0032] The ambient sounds, which are the sounds of the ambient sound sources LU, RU, LD, and RD, are sounds that express the size of the sound image of the central sound. By listening to the central sound and the ambient sounds, the user perceives the sound image of the central sound as having a larger size. This allows the user to perceive the distance to the object, which is the virtual sound source.

[0033] For example, the peripheral sound source RU is placed at position P11, which is a position that is a horizontal angle rAzim(d) and a vertical angle rElev(d) away from position P1 where the central sound source C is placed, with the user U as the reference. Similarly, the remaining peripheral sound sources LU, RD, and LD are placed at positions P12, P13, and P14, respectively, that are set with position P1 as the reference.

[0034] Position P12 where the ambient sound source LU is located is a position distant from position P1 by a horizontal angle of -rAzim(d) and a vertical angle of rElev(d). Position P13 where the ambient sound source RD is located is a position distant from position P1 by a horizontal angle of rAzim(d) and a vertical angle of -rElev(d), and position P14 where the ambient sound source LD is located is a position distant from position P1 by a horizontal angle of -rAzim(d) and a vertical angle of -rElev(d).

[0035] For example, the distance from the central sound source C to each of the peripheral sound sources is the same. In this way, the four peripheral sound sources LU, RU, LD, and RD are arranged radially with respect to the central sound source C.

[0036] FIG. 5 is another diagram showing the positional relationship between the central sound source and the peripheral sound sources.

[0037] For example, when the central sound source and the peripheral sound source are viewed from diagonally above, the positional relationship between the central sound source and the peripheral sound source is as shown in A of Fig. 5. When the central sound source and the peripheral sound source are viewed from the side, the positional relationship between the central sound source and the peripheral sound source is as shown in B of Fig. 5.

[0038] As described above, the positions of the multiple peripheral sound sources set around the central sound source C vary depending on the size of the sound image of the central sound that is intended to be perceived by the user.

[0039] Although the example in which four peripheral sound sources are set has been described so far as a typical example, the number of peripheral sound sources is not limited to this.

[0040] FIG. 6 is another diagram showing an example of distance expression in the present technology.

[0041] A in Fig. 6 shows the positions of the peripheral sound sources when the virtual sound source is far away from the user U wearing the headphones 1. By arranging each peripheral sound source close to the central sound source as shown in A in Fig. 6 and representing the sound image of the central sound as small, the user perceives the virtual sound source as being far away. As described above, the smaller the perceived sound image, the farther the user perceives the virtual sound source to be.

[0042] B in Fig. 6 shows the positions of the peripheral sound sources when the virtual sound source is close to the user U wearing the headphones 1. As shown in B in Fig. 6, by arranging each peripheral sound source at a position away from the central sound source and enlarging the sound image of the central sound, the user perceives the virtual sound source as being closer. As described above, the larger the perceived sound image, the closer the user perceives the virtual sound source to be.

[0043] According to the present technology, the positions of the peripheral sound sources arranged around the central sound source are controlled, thereby allowing the user to perceive different distances to the virtual sound source.

[0044] FIG. 7 is a diagram showing the shape of a sound image in the present technology.

[0045] Figure 7A shows the shape of the sound source when the absolute value of the horizontal angle between the central sound source and the peripheral sound sources is greater than the absolute value of the vertical angle. In this case, the shape of the sound image of the central sound perceived by the user is horizontally elongated, as shown by the colored oval.

[0046] Figure 7B shows the shape of the sound source when the absolute value of the vertical angle between the central sound source and the peripheral sound sources is greater than the absolute value of the horizontal angle. In this case, the shape of the sound image of the central sound perceived by the user is vertically elongated, as shown by the colored oval.

[0047] In this way, by changing the position of the ambient sound to any position, it is possible to express the distance even for virtual sound sources that have characteristic shapes such as vertically long or horizontally long.

[0048] 3. Examples of configurations of sound reproduction systems and information processing devices Next, the configurations of an audio reproduction system and an information processing device to which the present technology is applied will be described.

[0049] 8 is a diagram showing an example of the configuration of an audio reproduction system to which the present technology is applied. The audio reproduction system is configured by connecting an information processing device 10 and headphones 1.

[0050] In the present technology, for example, a user wears headphones 1 and carries an information processing device 10. The user can experience the spatial audio of the present technology by listening to sound corresponding to sound data processed by the information processing device 10 through the headphones 1 connected to the information processing device 10.

[0051] The information processing device 10 is, for example, a smartphone, a mobile phone, a PC, a television, a tablet, or the like that is owned by a user.

[0052] The headphones 1 are also called a playback device, and other devices such as earphones are also envisioned in addition to the headphones 1. The headphones 1 are worn on the user's head, more specifically, on the user's ears, and are connected to the information processing device 10 via a wire or wirelessly.

[0053] FIG. 9 is a block diagram illustrating an example of the hardware configuration of the information processing device 10. As shown in FIG.

[0054] As shown in FIG. 9, the information processing device 10 includes a CPU (Central Processing Unit) 11, a ROM (Read Only Memory) 12, and a RAM (Random Access Memory) 13, which are interconnected by a bus .

[0055] The information processing device 10 also includes an input / output interface 15, an input unit 16 configured with various buttons and a touch panel, and an output unit 17 configured with a display, a speaker, etc. The bus 14 is connected to the input / output interface 15, and the input unit 16 and the output unit 17 are connected to the input / output interface 15.

[0056] The information processing device 10 further includes a storage unit 18 including a hard disk or nonvolatile memory, a communication unit 19 including a network interface, and a drive 20 that drives removable media 21. The storage unit 18, communication unit 19, and drive 20 are connected to the input / output interface 15.

[0057] The information processing device 10 functions as an information processing device that processes sound data reproduced by a reproduction device such as headphones 1 worn by a user.

[0058] When the information processing device 10 and the playback device are wirelessly connected, the communication unit 19 functions as an output unit that supplies audio data.

[0059] The communication unit 19 may also function as an acquisition unit that acquires virtual sound source data and HRTF information via a network.

[0060] FIG. 10 is a block diagram showing an example of the functional configuration of the information processing device 10. As shown in FIG.

[0061] 10, the information processing unit 30 has a sound source setting unit 31, a spatial sound generating unit 32, and an output control unit 33. Each component shown in FIG. 10 is realized by the CPU 11 in FIG. 9 executing a predetermined program.

[0062] The sound source setting unit 31 sets a virtual sound source at a predetermined position to express a sense of distance. The sound source setting unit 31 also sets a central sound source according to the position of the virtual sound source, and sets peripheral sound sources at positions according to the distance to the virtual sound source.

[0063] The spatial sound generating unit 32 generates sound data of sounds from the central sound source and peripheral sound sources set by the sound source setting unit 31 .

[0064] For example, the spatial sound generation unit 32 performs convolution processing on the virtual sound source data based on HRTF information corresponding to the position of the central sound source to generate sound data for the central sound. Also, the spatial sound generation unit 32 performs convolution processing on the virtual sound source data based on HRTF information corresponding to the positions of each peripheral sound source to generate sound data for each peripheral sound.

[0065] The virtual sound source data that is the subject of convolution processing based on HRTF information corresponding to the position of the central sound source and the virtual sound source data that is the subject of convolution processing based on HRTF information corresponding to the positions of the peripheral sound sources may be the same data, or may be different data.

[0066] The output control unit 33 converts the sound data of the central sound and the sound data of each peripheral sound generated by the spatial sound generation unit 32 into L / R sound data. The output control unit 33 controls the output unit 17 or the communication unit 19 to output the converted sound data from a playback device worn by the user.

[0067] The output control unit 33 also adjusts the volume of the central sound and the volume of each peripheral sound as appropriate. For example, it is possible to reduce the volume of the peripheral sound to reduce the size of the sound image of the central sound, or to increase the volume of the peripheral sound to increase the size of the sound image of the central sound. The volume values ​​of the respective peripheral sounds can be set to either the same value or different values.

[0068] In this way, the information processing unit 30 sets a virtual sound source, as well as a central sound source and peripheral sound sources. The information processing unit 30 also performs convolution processing based on HRTF information according to the positions of the central sound source and each peripheral sound source, thereby generating sound data for the central sound and peripheral sound, and outputs the generated data to the playback device.

[0069] The HRTF data corresponding to the position of the central sound source and the HRTF data corresponding to the positions of the respective peripheral sound sources may be synthesized by multiplying them on the frequency axis, for example, and the synthesized HRTF data may be used to realize processing equivalent to the above-described processing. The synthesized HRTF data is HRTF data for expressing the area, which is the apparent size of the virtual sound source.

[0070] When the central sound source and the peripheral sound source are equal, there is an advantage that the amount of calculation is reduced.

[0071] <4. Operation of the information processing device> The processing of the information processing device 10 will be described with reference to the flowchart of FIG.

[0072] In step S101, the sound source setting unit 31 sets a virtual sound source at a predetermined position.

[0073] In step S102, the sound source setting unit 31 sets a central sound source according to the position of the virtual sound source.

[0074] In step S103, the sound source setting unit 31 sets the peripheral sound source in accordance with the distance from the user to the virtual sound source. In steps S101 to S103, the volume of the sound of each sound source is set appropriately.

[0075] In step S104, the spatial audio generation unit 32 generates sound data for a central sound, which is the sound of a central sound source, and a peripheral sound, which is the sound of a peripheral sound source, by performing convolution processing based on the HRTF information. The sound data for the central sound and the sound data for the peripheral sound generated by the convolution processing based on the HRTF information are supplied to a playback device and used to output the central sound and the peripheral sound, respectively.

[0076] In step S105, the sound source setting unit 31 determines whether or not the distance from the user to the virtual sound source changes.

[0077] If it is determined in step S105 that the distance from the virtual sound source to the user has changed, in step S106 the sound source setting unit 31 controls the positions of the peripheral sound sources according to the changed distance. For example, to express that the virtual sound source is approaching, the sound source setting unit 31 controls the positions of the respective peripheral sound sources so as to move away from the central sound source. On the other hand, to express that the virtual sound source is moving away, the sound source setting unit 31 controls the positions of the respective peripheral sound sources so as to move closer to the central sound source.

[0078] In step S107, the spatial sound generation unit 32 performs convolution processing based on the HRTF information to generate data on the central sound and the peripheral sound that are reset according to the distance to the virtual sound source. After the central sound and the peripheral sound are output using the sound data generated by the convolution processing based on the HRTF information, the processing ends.

[0079] On the other hand, if it is determined in step S105 that the distance from the user to the virtual sound source has not changed, the process also ends. The above process is repeated while the user is listening to the sound of the virtual sound source.

[0080] Through the above processing, the information processing device 10 can appropriately express the sense of distance from the user to the virtual sound source.

[0081] The user can perceive the distance to the virtual sound source through a realistic spatial audio experience.

[0082] FIG. 12 is a diagram showing another example configuration of a sound reproduction system to which the present technology is applied.

[0083] As shown in Fig. 12, an audio reproduction system to which the present technology is applied may include an information processing device 10, a reproduction device 50, a virtual sound source data providing server 60, and an HRTF server 70. In the example of Fig. 12, the reproduction device 50 is shown instead of the headphones 1. The reproduction device 50 is a general term for devices such as the headphones 1 and earphones that a user wears to listen to sound.

[0084] As shown in FIG. 12, it is also possible that the information processing device 10 and the playback device 50 function by receiving data from a virtual sound source data providing server 60 and an HRTF server 70 connected via a network such as the Internet.

[0085] For example, the information processing device 10 communicates with the virtual sound source data providing server 60 and acquires the virtual sound source data provided by the virtual sound source data providing server 60 .

[0086] Furthermore, the information processing device 10 communicates with the HRTF server 70 and acquires HRTF information provided by the HRTF server 70. The HRTF information is data for adding transfer characteristics from a virtual sound source to the user's ear (eardrum), that is, data in which head-related transfer functions for localizing a sound image at the position of the virtual sound source are recorded for each direction of the virtual sound source as seen from the user.

[0087] The HRTF information acquired from the HRTF server 70 may be recorded in the information processing device 10, or may be acquired from the HRTF server 70 each time the sound of the virtual sound source is output.

[0088] As the head-related transfer function, information recorded in the form of HRIR (Head Related Impulse Response), which is information in the time domain, or information recorded in the form of HRTF, which is information in the frequency domain, may be used. In this specification, the description will be made assuming that HRTF information is used.

[0089] Furthermore, the HRTF information may be personalized according to the physical characteristics of each individual user, or may be shared by multiple users.

[0090] For example, the personalized HRTF information may be information obtained by placing the subject in a test environment and measuring them, or information calculated from an image of the subject's ear. Information calculated based on the size information of the subject's head and ears may also be used as personalized HRTF information.

[0091] The commonly used HRTF information may be information obtained by measurement using a dummy head, or information obtained by averaging the HRTF information of multiple people. The user may be asked to compare sounds reproduced using multiple pieces of HRTF information, and the HRTF information that the user determines to be most suitable for them may be used as the commonly used HRTF information.

[0092] 12 has a communication unit 51, a control unit 52, and an output unit 53. In this case, it is also possible for the playback device 50 to take on at least some of the above-mentioned functions of the information processing device 10, and for the processing for generating the sound of the virtual sound source to be performed by the playback device 50. The control unit 52 of the playback device 50 acquires virtual sound source data and HRTF information through communication in the communication unit 51, and performs the above-mentioned processing for generating the sound of the virtual sound source.

[0093] In FIG. 12, the virtual sound source data providing server 60 and the HRTF server 70 are each configured as one device, but they may also be configured as multiple devices on the cloud.

[0094] Furthermore, the virtual sound source data providing server 60 and the HRTF server 70 may be realized by one device.

[0095] <5. Modifications (Application Examples)> -Notifying visually impaired people of obstacles using spatial audio while they are walking

[0096] FIG. 13 is a diagram showing an example of a method for notifying an obstacle to which the present technology is applied.

[0097] 13 shows a user U walking with a white cane W. The user U is wearing headphones 1. The white cane W held by the user U includes an ultrasonic speaker unit that emits ultrasonic waves, a microphone unit that receives reflected ultrasonic waves, and a communication unit that communicates with the headphones 1 (none of which are shown).

[0098] The white cane W also includes a processing control unit that controls the output of ultrasonic waves from the ultrasonic speaker unit and processes sounds detected by the microphone unit. These components are provided, for example, in a housing formed at the top end of the white cane W.

[0099] The ultrasonic speaker and microphone attached to the white cane W function as sensors, and notify the user U of information about obstacles in the vicinity. The notification to the user U is performed using sounds from a virtual sound source that allow the user U to perceive distance based on the size of the sound image.

[0100] As shown in Figure 14, ultrasonic waves output from the ultrasonic speaker unit of the white cane W are reflected by a wall X, which is a surrounding obstacle. The ultrasonic waves reflected by the wall X are detected by the microphone unit of the white cane W. As a result, the processing control unit of the white cane W detects the distance to the wall X, which is a surrounding obstacle, and the direction of the wall X as spatial information.

[0101] When the processing control unit of the white cane W detects the distance to the wall X and the direction of the wall X, it sets the wall X, which is an obstacle, as an object corresponding to the virtual sound source.

[0102] Furthermore, the processing control unit sets a central sound source and peripheral sound sources that represent the distance to wall X and the direction of wall X. For example, the central sound source is set in the direction of wall X, and peripheral sound sources are set at positions according to the size of the sound image that represents the distance to wall X.

[0103] The processing control unit generates sound data of the central sound and the peripheral sound by treating data such as an alarm sound as virtual sound source data and performing convolution processing on the virtual sound source data based on HRTF information corresponding to the respective positions of the central sound source and the peripheral sound sources. The processing control unit transmits the sound data obtained by performing the convolution processing to headphones 1 worn by the user U, and causes the headphones 1 to output the central sound and the peripheral sound.

[0104] When walking with a regular white cane (a white cane without an ultrasonic speaker or microphone), a visually impaired user, for example, can only obtain information about an area around them of about one meter, and is unable to obtain information about obstacles such as walls, steps, and cars several meters ahead, which puts them in danger.

[0105] In this way, by expressing the distance and direction of an obstacle detected by the white cane W as spatial sound, the user U can recognize not only the direction of the obstacle in the vicinity but also the distance to the obstacle by sound alone.In addition to information about the obstacle, spatial information is also acquired about the situation, such as whether there is space below in front, which indicates the edge of the platform, etc.

[0106] In this application example, the white cane W acquires distance information to nearby obstacles by using the ultrasonic speaker unit and microphone unit as sensors, and based on the acquired distance information, expresses the distance to the obstacle using spatial audio.

[0107] For example, by repeating such processing at short intervals, such as 50 ms, the user can instantly obtain information about surrounding obstacles while walking.

[0108] 13 and 14, all of the components, including the ultrasonic speaker unit, microphone unit, processing control unit, and output control unit, are provided in the white cane W, but at least one of these components may be provided as a separate device from the white cane. The functions of the white cane described above are realized by the communication between each component.

[0109] Furthermore, since there are individual differences in how people perceive distance from sound, the relationship between how users perceive distance and the size of a sound image may be learned in advance, and the size of a sound image may be adjusted to match the user's recognition pattern.

[0110] Furthermore, the size of the sound image may be adjusted depending on whether the user is walking or standing still, thereby allowing the user to easily perceive the sense of distance.

[0111] -Sound-based map information presentation

[0112] FIG. 15 is a diagram showing an example of a method for notifying a distance to a destination to which the present technology is applied.

[0113] In FIG. 15, it is assumed that a user U carries an information processing device 10 (not shown) and is walking to a destination D where a store or the like is located.

[0114] The information processing device 10 carried by the user U includes a position detection unit that detects the current position of the user U, and a surrounding information acquisition unit that acquires information on surrounding stations and the like.

[0115] In this application example, the information processing device 10 acquires the position of the user U by a position detection unit and acquires surrounding information by a surrounding information acquisition unit. Furthermore, the information processing device 10 controls the size of the sound image presented to the user U in accordance with the distance to the destination D, thereby allowing the user U to intuitively perceive the distance to the destination D.

[0116] For example, the information processing device 10 increases the volume of the sound image of the sound representing the destination D as the user U approaches the destination D. This allows the user U to perceive that the distance to the destination D is short.

[0117] 15A is a diagram showing an example of a sound image when the distance to destination D is long. In this case, the sound representing destination D is presented as a small sound image, as indicated by a small colored circle #51.

[0118] 15B is a diagram showing an example of a sound image when the distance to destination D is short. In this case, the sound representing destination D is presented as a sound with a large sound image, as indicated by colored circle #52.

[0119] In this way, map information using sounds to guide the user to their destination can be presented in an easy-to-understand manner using spatial audio.

[0120] Furthermore, by changing the size of the sound image depending on the amount of surrounding noise, it is possible to make the sound more easily understandable.

[0121] Example of alarm sound

[0122] FIG. 16 is a diagram illustrating an example of a notification method of an alarm sound of a home appliance to which the present technology is applied.

[0123] FIG. 16 shows how the notification sound of a kettle pot is presented to the user U, for example.

[0124] The information processing device 10 possessed by the user U includes a detection unit that detects the urgency and importance of the contents of notifications in cooperation with other devices such as household electrical appliances (home appliances).

[0125] In this application example, the information processing device 10 changes the volume of the sound image of the alarm sound of the home appliance according to the urgency and importance detected by the detection unit, thereby intuitively conveying the urgency and importance of the alarm sound to the user U.

[0126] According to this application example, even if the user U does not notice the monotonous buzzer sound from the speaker attached to the home appliance, it is possible to make the user U aware of the alarm sound of the home appliance by increasing the size of the sound image and presenting the alarm sound.

[0127] The urgency and importance of the alarm sounds of home appliances are set according to the danger, for example. When the water boils, it is dangerous to leave the alarm sound unnoticed. In this case, the alarm is set to a high level of urgency and importance.

[0128] Although the home appliance has been described as a kettle pot, the present invention can also be applied to presenting alarm sounds for other home appliances. Applicable home appliances include refrigerators, microwave ovens, rice cookers, dishwashers, washing machines, kettle pots, vacuum cleaners, etc. The examples given here are general and are not limited to the illustrated examples.

[0129] Furthermore, if you want to draw the user's attention to a specific part of the device, you can gradually reduce the area of ​​the warning sound to guide the user's gaze. The specific part of the device could be, for example, a switch, button, or touch panel provided on the device.

[0130] In this way, according to the present technology, it is possible not only to allow the user to perceive the distance to the virtual sound source, but also to present the user with the importance and urgency of the device's alarm sound and to guide the user's gaze.

[0131] Example of a remote conference system

[0132] FIG. 17 is a diagram illustrating an example of the configuration of a remote conference system.

[0133] 17 shows how, for example, remote users A to D are holding a conference via a network 101 such as the Internet. A communication management server 100 is connected to the network 101.

[0134] The communication management server 100 controls the transmission and reception of voice data between users. Voice data transmitted from the information processing devices 10 used by each user is mixed in the communication management server 100 and distributed to all information processing devices 10.

[0135] The communication management server 100 also manages the position of each user on a spatial map and outputs the voice of each user as a sound with a sound image whose volume corresponds to the distance between each user on the spatial map. The communication management server 100 has the same functions as the information processing device 10 described above.

[0136] Each of users A to D wears headphones 1 and participates in the remote conference using information processing devices 10A to 10D. Each information processing device 10 has a built-in microphone or is connected to it, and a program for using the remote conference system is installed on it.

[0137] FIG. 18 is a diagram showing an example of a screen display that serves as a user interface during a remote conference.

[0138] The example in Fig. 18 is a screen of a remote conference system, and each user is represented by a circular icon I1, I2, or I3. Icons I1 to I3 represent, for example, users A to C, respectively. The user who views the screen in Fig. 18 and participates in the remote conference is, for example, user D.

[0139] User D can set the distance to a desired user by moving the icon and controlling the position of each user on the spatial map. In the example of Fig. 18, for example, the position of user B represented by icon I2 is set to be close, and the position of user A represented by icon I1 is set to be farther away.

[0140] 19 is a diagram showing an example of the size of the sound image of each user's voice. User U facing the screen is, for example, user D.

[0141] As shown by colored circle #61, the voice of user B, who is set to a nearby position on the spatial map, is output as a sound image with a larger volume depending on the distance. As shown by circles #62 and #63, the voices of users A and C are output as sound images with a volume depending on the respective distances.

[0142] If the voices of all users were mixed as monaural audio and output from headphones 1, the position of the speakers would be concentrated in one point, making it difficult to create the cocktail party effect and preventing users from focusing on the voice of a specific speaker. This also makes it difficult to hold group discussions in multiple groups.

[0143] In this way, by controlling the volume of the sound image of each speaker's voice according to the position of each speaker, it is possible to express the sense of distance between the user and each speaker.

[0144] By expressing the distance between the users and each speaker present at the conference, the user can have a conversation while feeling a sense of distance.

[0145] The voices of the speakers to be grouped may be output as a large sound image, as if they were localized at a nearby position, such as next to the ear. This makes it possible to express the feeling of being a group of speakers.

[0146] An HMD, a camera, etc. may be built into or connected to each information processing device 10. The direction of the user's face is detected using the HMD or the camera, and when it is detected that the user is paying attention to a specific speaker, the volume of the sound image of the speaker's voice that the user is paying attention to is increased, making it possible to make the user feel as if the specific speaker is speaking closer to the user.

[0147] In this example, each user can control the position of other users (speakers), but this is not limiting. For example, it is also possible that each participant in a conference controls their own or other participants' positions on the spatial map, and a position set by someone is shared among all participants.

[0148] -Example of a car's simulated engine sound

[0149] FIG. 20 is a diagram showing an example of a method for notifying a user of a pseudo engine sound of a vehicle.

[0150] Pedestrians are thought to recognize moving vehicles primarily based on visual and auditory information, but the engine noise of modern electric vehicles is quiet and difficult for pedestrians to notice. Furthermore, even if they can hear the sound of a car, it is difficult to notice that a car is approaching if other noises are heard along with it.

[0151] In this application example, a user U, who is a pedestrian, is made to hear a pseudo engine sound emitted by the car 110, thereby making the user U aware of the moving car 110. The car 110 is equipped with a device having functions similar to those of the information processing device 10. The user U, who is walking while wearing headphones 1, hears the pseudo engine sound output from the headphones 1 in accordance with control by the car 110.

[0152] In this application example, the car 110 is equipped with a camera that detects a user U who is a pedestrian, and a communication unit that transmits a pseudo engine sound as approach information to a user U walking nearby.

[0153] When the car 110 detects the user U, it generates a pseudo engine sound having a sound image whose size corresponds to the distance to the user U. The pseudo engine sound generated based on the central sound and the peripheral sound is transmitted to the headphones 1 and presented to the user U.

[0154] 20A is a diagram showing an example of a sound image when the distance between the car 110 and the user U is great. In this case, the pseudo engine sound is presented as a small sound image as indicated by the small colored circle #71.

[0155] 20B is a diagram showing an example of a sound image when the distance between the car 110 and the user U is short. In this case, the pseudo engine sound is presented as a loud sound image, as indicated by the colored circle #72.

[0156] The pseudo engine sound based on the central sound and the peripheral sound may be generated not in the car 110 but in the information processing device 10 owned by the user U.

[0157] According to this technology, the user U can perceive the direction from which the vehicle 110 is coming as well as the sense of distance to the vehicle 110, thereby improving the accuracy of risk avoidance.

[0158] The notification using the pseudo engine sound described above can be applied not only to cars with quiet engine noise, but also to conventional cars. By exaggerating the sense of distance by playing a pseudo engine sound with a sound image that varies in volume depending on the distance, the user can be made to perceive an approaching car and improve the accuracy of danger avoidance.

[0159] Example of a car obstacle warning sound

[0160] Although there are already systems that emit sounds to warn drivers when a car is approaching a wall, for example when parking, there are times when it is difficult to judge the distance between the car and the wall.

[0161] In this application example, the car is equipped with a camera for detecting the approach of a wall. In this case, the car is also equipped with a device having the same function as the information processing device 10.

[0162] The device installed in the vehicle detects the distance between the vehicle body and the wall based on images captured by the camera and controls the volume of the sound image of the warning sound. The closer the vehicle body is to the wall, the louder the sound image of the warning sound is output. By perceiving the distance to the wall based on the volume of the sound image of the warning sound, it is possible to improve the accuracy of crisis avoidance.

[0163] Example of fish school prediction detection

[0164] This technology can also be applied to the display of schools of fish by a fish detection and prediction device. For example, the larger the area of ​​a school of fish, the louder the sound image and the warning sound that is displayed. This allows the user to intuitively determine the predicted size of the school of fish.

[0165] Example of sound space expression

[0166] This technology allows the user to perceive the distance from the virtual sound source, and furthermore, by changing the area of ​​the reverberant sound (size of the sound image) relative to the direct sound, it is possible to express the expanse of space. In other words, by applying this technology to reverberant sound, it is possible to express a sense of depth.

[0167] Furthermore, by expressing the area of ​​the reverberation sound by reducing the amount of change as the user becomes more accustomed to it, the stimulating burden on the user can be reduced.

[0168] The perception of sound differs depending on whether the sound comes from the front, side, or back of the face. By setting parameters appropriate for each direction as parameters for area representation, it becomes possible to express the sound appropriately according to the direction from which it is presented.

[0169] Examples of video content and movies

[0170] This technology can be applied to the presentation of sounds from various types of content, including video content such as movies, audio content, and game content. By setting an object in the content as a virtual sound source and controlling the central and peripheral sounds, it is possible to create an experience in which the virtual sound source appears to move closer to or farther away from the user.

[0171] <6. Other examples> ·Playback device configuration

[0172] FIG. 21 is a diagram illustrating an example of a playback device.

[0173] The playback device used to output the sound of the virtual sound source may be a closed-type headphone (over-ear headphone) as shown in A of Fig. 21, or a shoulder-mounted neckband speaker as shown in B of Fig. 21. The left and right units constituting the neckband speaker are provided with speakers, and sound is output toward the user's ears.

[0174] FIG. 22 is a diagram illustrating another example of a playback device.

[0175] The playback device shown in FIG. 22 is an open-type earphone.

[0176] The open-type earphone shown in Fig. 22 is composed of a right unit 120R and a left unit 120L (not shown). As shown in an enlarged view in the speech bubble in Fig. 22, right unit 120R is composed of driver unit 121 and ring-shaped attachment part 123 joined via U-shaped sound guide tube 122. Right unit 120R is worn by pressing attachment part 123 against the area around the ear canal, with attachment part 123 and driver unit 121 sandwiching the right ear.

[0177] The left unit 120L has the same configuration as the right unit 120R. The left unit 120L and the right unit 120R are connected by wire or wirelessly.

[0178] Driver unit 121 of right unit 120R receives an audio signal transmitted from information processing device 10, and outputs sound corresponding to the audio signal from the tip of sound conduit 122, as indicated by arrow A1. A hole is formed at the joint between sound conduit 122 and attachment part 123, which outputs sound toward the external ear canal.

[0179] Wearing part 123 has a ring shape. In addition to the sound output from the tip of sound conduit 122, ambient sound also reaches the external ear canal, as indicated by arrow A2.

[0180] In this way, it is possible to use open-type earphones that do not seal the ear canal.

[0181] These playback devices may be provided with a detection unit that detects the orientation of the user's head. When the detection unit that detects the orientation of the user's head is provided, the HRTF information used in the convolution process is adjusted so that the position of the virtual sound source is fixed even if the orientation of the user's head changes.

[0182] About the program

[0183] The above-described series of processes can be executed by hardware or by software. When the series of processes is executed by software, the program that constitutes the software is installed from a program recording medium into a computer incorporated in dedicated hardware or a general-purpose personal computer.

[0184] The program to be installed is provided by being recorded on removable media such as an optical disc (CD-ROM (Compact Disc-Read Only Memory), DVD (Digital Versatile Disc), etc.) or semiconductor memory. It may also be provided via wired or wireless transmission media such as a local area network, the Internet, or digital broadcasting. The program can be pre-installed in a ROM or memory unit.

[0185] The program executed by the computer may be a program that processes in chronological order according to the order described in this specification, or may be a program that processes in parallel or at the required timing, such as when called.

[0186] In this specification, a system refers to a collection of multiple components (devices, modules (components), etc.), regardless of whether all the components are contained in the same housing. Therefore, multiple devices housed in separate housings and connected via a network, and a single device housed in a single housing with multiple modules, are both systems.

[0187] The effects described in this specification are merely examples and are not limiting, and other effects may also be present.

[0188] The embodiments of the present technology are not limited to the above-described embodiments, and various modifications are possible without departing from the spirit of the present technology.

[0189] For example, this technology can be configured as cloud computing, in which a single function is shared and processed collaboratively by multiple devices via a network.

[0190] Furthermore, each step described in the above flowchart can be executed by one device, or can be shared and executed by multiple devices.

[0191] Furthermore, when one step includes multiple processes, the multiple processes included in that one step can be executed by one device or can be shared and executed by multiple devices.

[0192] Configuration combination examples The present technology can also be configured as follows.

[0193] (1) a first sound source and a sound source setting unit that sets a plurality of second sound sources at positions according to the size of a sound image of the first sound, which is the sound of the first sound source; an output control unit that outputs first sound data obtained by convolution processing using HRTF information corresponding to the position of the first sound source and a plurality of second sound data obtained by convolution processing using HRTF information corresponding to the positions of the second sound sources; Equipped with Each of the second sound sources is set to be located around the first sound source. Information processing device. (2) The sound source setting unit sets the second sound sources with the first sound source as the center. The information processing device according to (1) above. (3) The sound source setting unit sets the second sound source at a position farther from the first sound source as the size of the sound image of the first sound increases. The information processing device according to (1) or (2). (4) The second sound source is four sound sources set around the first sound source. The information processing device according to any one of (1) to (3). (5) The sound source setting unit sets each of the second sound sources at a position according to the shape of the sound image of the first sound. The information processing device according to any one of (1) to (4). (6) The output control unit controls a playback device worn by a user to output two-channel audio data representing the first sound and a second sound that is a sound from the second sound source. The information processing device according to any one of (1) to (5). (7) The output control unit adjusts the volume of each of the first sound and the second sound according to the size of the sound image of the first sound. The information processing device according to (6) above. (8) The sound source setting unit determines that the size of the sound image of the first sound changes, and controls the position of the second sound source according to the size of the sound image of the first sound. The information processing device according to any one of (2) to (7). (9) The second sound, which is the first sound and the sounds of the second sound sources, is a sound for expressing a virtual sound source corresponding to an object. The information processing device according to any one of (2) to (5). (10) A detection unit is further provided to detect current location information of the user and destination information of the user, The sound source setting unit sets the position of the first sound source based on the current location information, and sets the position of the second sound source using the destination information. The information processing device according to any one of (2) to (9). (11) The information processing device a first sound source and a plurality of second sound sources are set at positions corresponding to the size of a sound image of the first sound, which is the sound of the first sound source; outputting first audio data obtained by performing convolution processing using HRTF data corresponding to the position of the first sound source, and a plurality of second audio data obtained by performing convolution processing using HRTF data corresponding to the positions of the second sound sources, each of which is set to be located in the vicinity of the first sound source; Information processing methods. (12) On the computer, a first sound source and a plurality of second sound sources are set at positions corresponding to the size of a sound image of the first sound, which is the sound of the first sound source; outputting first audio data obtained by performing convolution processing using HRTF data corresponding to the position of the first sound source, and a plurality of second audio data obtained by performing convolution processing using HRTF data corresponding to the positions of the second sound sources, each of which is set to be located in the vicinity of the first sound source; A program for executing a process. [Explanation of symbols]

[0194] 1 headphones, 10 information processing device, 30 information processing unit, 31 sound source setting unit, 32 spatial sound generation unit, 33 output control unit, 50 playback device, 60 virtual sound source data providing server, 70 HRTF server, 100 communication management server, 101 network, U user, C central sound source, LU, RU, LD, RD peripheral sound sources

Claims

1. a sound source setting unit that sets a first sound source and a plurality of second sound sources as sound sources around the first sound source; an output control unit that outputs first sound data obtained by convolution processing using HRTF information corresponding to the position of the first sound source and a plurality of second sound data obtained by convolution processing using HRTF information corresponding to the positions of the second sound sources; Equipped with When a virtual sound source represented by a first sound of the first sound source and a second sound of the second sound source approaches the user, the sound source setting unit controls the second sound source to move away from the first sound source from its previous position, and when the virtual sound source moves away from the user, the sound source setting unit controls the second sound source to move closer to the first sound source from its previous position. Information processing device.

2. The information processing device according to claim 1 , wherein the sound source setting unit sets the second sound sources with the first sound source at the center.

3. The sound source setting unit sets the second sound source at a position farther from the first sound source as the size of the sound image of the first sound increases. The information processing device according to claim 1 .

4. The second sound source is four sound sources set around the first sound source. The information processing device according to claim 1 .

5. The sound source setting unit sets the second sound sources at positions corresponding to the shapes of the sound images of the first sounds. The information processing device according to claim 1 .

6. The output control unit outputs two-channel audio data representing the first sound and the second sound from a playback device worn by the user. The information processing device according to claim 1 .

7. The output control unit adjusts the volume of each of the first sound and the second sound in accordance with the size of a sound image of the first sound. The information processing device according to claim 6 .

8. Further comprising a detection unit that detects current location information of the user and destination information of the user, The sound source setting unit sets the position of the first sound source based on the current location information, and sets the position of the second sound source using the destination information. The information processing device according to claim 2 .

9. The information processing device A first sound source and a plurality of second sound sources are set as sound sources around the first sound source; outputting first audio data obtained by performing convolution processing using HRTF data corresponding to the position of the first sound source, and a plurality of second audio data obtained by performing convolution processing using HRTF data corresponding to the positions of the second sound sources, each of which is set to be located in the vicinity of the first sound source; When a virtual sound source represented by a first sound of the first sound source and a second sound of the second sound source approaches the user, the second sound source is controlled to move away from the first sound source from its previous position, and when the virtual sound source moves away from the user, the second sound source is controlled to move closer to the first sound source from its previous position. Information processing methods.

10. On the computer, A first sound source and a plurality of second sound sources are set as sound sources around the first sound source; outputting first audio data obtained by performing a convolution process using HRTF data corresponding to the position of the first sound source, and a plurality of second audio data obtained by performing a convolution process using HRTF data corresponding to the positions of the second sound sources, each of which is set to be located in the periphery of the first sound source; When a virtual sound source represented by a first sound of the first sound source and a second sound of the second sound source approaches the user, the second sound source is controlled to move away from the first sound source from its previous position, and when the virtual sound source moves away from the user, the second sound source is controlled to move closer to the first sound source from its previous position. A program for executing a process.

Citation Information

Patent Citations

  • Stereophonic system

    JP1993119770A

  • Method of processing audio signal

    JP2010004512A

  • Audio data reproduction method and audio data reproduction apparatus

    JP2013038511A

  • Information processing device, information processing method, and program

    JP2019087973A