Sound source generation device, sound source generation method, and sound source generation program
The sound source generation device addresses the challenge of hearing difficulties by estimating and generating playback information distinct from the sound source information, improving the audibility of the target signal.
Patent Information
- Application Number
- JP2024029169
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-02-28
- Publication Date
- 2025-09-09
AI Technical Summary
The sound generated by a wearable device in the presence of a sound source becomes difficult to hear when the signal from the sound source arrives at the user.
A sound source generation device that estimates sound source information using sensors, generates playback information different from the sound source information, and presents the playback signal to the user, where the sound source information is the arrival direction of the sound source signal and the playback information is the arrival direction of the playback signal.
The device makes it easier for the user to hear the target signal by manipulating playback information to enhance the distinguishability of the sound source.
Smart Images

Figure 2025131430000001_ABST
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to a sound source generation device, a sound source generation method, and a sound source generation program. [Background technology]
[0002] Patent Document 1 describes a wearable terminal that generates a presentation sound related to a virtual sound source according to the positional relationship between the user and the virtual sound source. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2016-021169 Summary of the Invention [Problem to be solved by the invention]
[0004] However, when a sound source is actually present around the user and a signal from the sound source arrives at the user, the presented sound generated by the wearable device described in Patent Document 1 becomes difficult to hear.
[0005] The present disclosure aims to provide a sound source generating device that solves the above-mentioned problems. [Means for solving the problem]
[0006] According to one aspect of the present disclosure, there is provided a sound source generation device including: an estimation means for estimating sound source information of a sound source signal acquired by a sensor; a generation means for generating playback information different from the sound source information and adding the playback information to the playback signal; and a playback means for presenting the playback signal to a user, wherein the sound source information is the arrival direction of the sound source signal, and the playback information is the arrival direction of the playback signal.
[0007] According to one aspect of the present disclosure, there is provided a sound source generation method, comprising: an estimation step of estimating sound source information of a sound source signal acquired by a sensor; a generation step of generating playback information different from the sound source information and adding the playback information to the playback signal; and a playback step of presenting the playback signal to a user, wherein the sound source information is the arrival direction of the sound source signal, and the playback information is the arrival direction of the playback signal.
[0008] According to one aspect of the present disclosure, there is provided a sound source generation program including: an estimation step of estimating sound source information of a sound source signal acquired by a sensor; a generation step of generating playback information different from the sound source information and adding the playback information to the playback signal; and a playback step of presenting the playback signal to a user, wherein the sound source information is the arrival direction of the sound source signal, and the playback information is the arrival direction of the playback signal. [Effects of the Invention]
[0009] According to the present disclosure, it is possible to provide a sound source generating device that makes it easier for a user to hear a target signal. [Brief explanation of the drawings]
[0010] [Figure 1] FIG. 1 is a schematic diagram illustrating an example of mixed reality according to an embodiment of the present disclosure. [Figure 2] 1 is a block diagram showing a hardware configuration of a sound source generating device according to an embodiment of the present disclosure. [Figure 3] FIG. 2 is a block diagram showing a hardware configuration of a terminal according to an embodiment of the present disclosure. [Figure 4] 1 is a block diagram illustrating functional units of a sound source generating device according to an embodiment of the present disclosure. [Figure 5] 2 is a schematic diagram illustrating the directions of arrival of observed signals and the directions of arrival of a target signal according to an embodiment of the present disclosure; FIG. [Figure 6] 1 is a schematic diagram illustrating a time waveform of an observation signal and a time waveform of a target signal according to an embodiment of the present disclosure. [Figure 7]1 is a schematic diagram illustrating a time waveform of an observation signal and a time waveform of a target signal according to an embodiment of the present disclosure. [Figure 8] 1 is a schematic diagram illustrating a frequency distribution of an observed signal and a frequency distribution of a target signal according to an embodiment of the present disclosure. [Figure 9] 1 is a schematic diagram illustrating a time waveform of an observation signal and a time waveform of a target signal according to an embodiment of the present disclosure. [Figure 10] FIG. 1 is a schematic diagram illustrating an observed sound source, a target sound source, and user attention information according to an embodiment of the present disclosure. [Figure 11] FIG. 1 is a schematic diagram illustrating an observed sound source, a target sound source, and user attention information according to an embodiment of the present disclosure. [Figure 12] 1 is a sequence chart of a sound source generation method according to an embodiment of the present disclosure. [Figure 13] 1 is a sequence chart of a sound source generation method according to an embodiment of the present disclosure. [Figure 14] 1 is a block diagram illustrating a sound source generating device according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0011] A sound source generating device according to an embodiment of the present disclosure will be described below with reference to the drawings. In all the drawings, the same or corresponding components are designated by the same reference numerals, and common descriptions will be omitted.
[0012] [First embodiment] FIG. 1 is a schematic diagram showing a sound source generation device according to an embodiment of the present disclosure. The sound source generation device 1 is an information processing device worn by a user 3. The sound source generation device 1 may be, for example, a head-mounted display, smart glasses, a headset, a mobile terminal, or the like. The sound source generation device 1 desirably includes a sensor capable of acquiring sound and a sensor capable of acquiring images or videos. The sound source generation device 1 is directly or indirectly connected to a network 9. When the sound source generation device 1 is indirectly connected to the network 9, the sound source generation device 1 may be connected to the network 9 via a mobile terminal, smartphone, tablet, or the like carried by the user 3. The sound source generation device 1 can communicate with a terminal 81 or a server 82 via the network 9.
[0013] The terminal 81 is an information processing device operated by a user different from the user 3, and may be, for example, a personal computer, a laptop computer, a tablet computer, a smartphone, a mobile phone, etc. The terminal 81 can communicate with the sound source generating device 1 or the server 82 via the network 9.
[0014] The server 82 is an information processing device, and may be, for example, a personal computer, a workstation, a hardware server, a software server, or the like. If the server 3 is a hardware server, the server 3 may be, for example, a network server or a cloud server. If the server 3 is a software server, the server 3 may be, for example, server software or a server program. The server 82 can communicate with the sound source generation device 1 or the terminal 81 via the network 9. The server 82 may receive an audio signal, an image signal, and a video signal from the sound source generation device 1 and perform audio signal processing, image signal processing, and video signal processing. For example, the sound source generation device 1 acquires an audio signal, an image signal, and a video signal and transmits them to the server 82. The server 82 may estimate sound source information of the audio signal based on the image signal and the video signal using a model trained in advance using machine learning or the like. The sound source generation device 1 may generate a target sound source 5 and a target signal 51 using the sound source information estimated by the server 82.
[0015] The target sound source 5 is a sound source that generates a sound signal that the user 3 wants to hear. The target sound source 5 may be a sound source that the user 3 pays attention to. The target sound source 5 may be a sound source other than a sound source that generates a sound signal that the user 3 intentionally does not want to hear. That is, the target sound source 5 may include multiple sound sources. The target sound source 5 may be a sound database stored in a storage device or the like provided in the sound source generation device 1, or a sound source around the user 3. The target sound source 5 may be, for example, a virtual sound source in virtual reality, a virtual sound source in augmented reality, a virtual sound source in mixed reality, a terminal 81, a person, an animal, a motorcycle, a car, broadcasting equipment such as a store, etc. When the target sound source 5 is a sound database, the sound source generation device 1 generates a target signal 51 based on information in the sound database, and the sound source generation device 1 plays the target signal 51. When the target sound source 5 is a sound source around the user 3, the sound source generation device 1 acquires the target signal 51 via a sensor or the like, and the sound source generation device 1 plays the target signal 51.
[0016] The observed sound source 7 is a sound source that exists around the user 3. The observed sound source 7 can be, for example, a person, an animal, a motorcycle, a car, an airplane, broadcasting equipment in a store or the like, a construction site, etc. The observed sound source 7 is not limited to one that is clearly recognized by the user 3, and the observed sound source 7 may include the natural environment, etc. The observed sound source 7 generates an observed signal 71 that can be observed by the sound source generation device 1. Note that the observed signal 71 observed by the sound source generation device 1 does not have to be a direct sound source signal from the observed sound source 7. For example, the observed signal 71 generated by the observed sound source 7 may be observed by the sound source generation device 1 via reflection by an outer wall of a building or the like. Like the target sound source 5, the observed sound source 7 may also include multiple sound sources.
[0017] 2 is a block diagram showing a sound source generation device 1 according to an embodiment of the present disclosure. The sound source generation device 1 includes a CPU 11, a ROM 12, a RAM 13, a storage device 14, an input / output IF (Interface) 15, a communication IF 16, and a sensor 17. The CPU 11, the ROM 12, the RAM 13, the storage device 14, the input / output IF 15, the communication IF 16, and the sensor 17 are connected via a bus 19 so as to be able to communicate with each other.
[0018] CPU 11 is a central processing unit. CPU 11 controls each part of sound source generation device 1 using an application program. ROM 12 is a read only memory. ROM 12 is made up of non-volatile memory and stores application programs for controlling each part of sound source generation device 1. RAM 13 is a random access memory. RAM 13 provides a memory area necessary for the operation of CPU 11. Storage device 14 is a large-capacity storage device such as a hard disk drive.
[0019] The input / output IF 15 is an input / output interface that receives voice input from the user 3, outputs voice to the user 3, and transmits and receives data between the sound source generation device 1 and other devices. The input / output IF 15 may include a mouse, a touch panel, a trackball, a keyboard, earphones, headphones, speakers, a display, etc. The input / output IF 15 may include multiple displays and multiple speakers. The user 3 can operate the sound source generation device 1 via the input / output IF 15. The communication IF 16 performs mutual voice or data communication between the sound source generation device 1 and the terminal 81 or the server 82 via wired communication and / or wireless communication. The communication IF 16 has a GPS (Global Positioning System) receiver, and the sound source generation device 1 can obtain location information of the user 3 via the communication IF 16.
[0020] The sensor 17 is a sensor that acquires audio signals, image signals, and video signals of the surroundings of the user 3, including the user 3, and may be, for example, one or more microphones, one or more cameras, an infrared sensor, a gyro sensor, a speedometer, an accelerometer, a blood pressure monitor, a pulse rate monitor, an optical sensor, or the like. When the sensor 17 has multiple microphones, the sensor 17 may combine the multiple microphones to form a microphone array. When the sensor 17 forms a microphone array, directivity may be formed by a delay and sum array, or blind spots may be formed by an adaptive microphone array for noise reduction (AMNOR). The microphone array may be configured in two dimensions or three dimensions. When the sensor 17 has multiple microphones, the geometric arrangement of the multiple microphones is preferably an arrangement that allows estimation of the azimuth angle and elevation angle of the arrival direction of the target sound source 5 and the observation sound source 7.
[0021] If the sensor 17 has multiple cameras, the sensor 17 may combine the multiple cameras to acquire the distance from the sound source generating device 1 to an object present around the user 3. The sensor 17 may acquire the distance to an object present around the user 3 by irradiating near-infrared light, visible light, or ultraviolet light. That is, the sensor 17 may constitute a LiDAR (Light Detection and Ranging). If the sensor 17 has multiple cameras, it may acquire a face image and line of sight of the user 3. For example, a person or object that the user 3 is paying attention to may be acquired based on the face image of the user 3 and line of sight information of the user 3.
[0022] 3 is a block diagram showing a terminal 81 according to an embodiment of the present disclosure. The terminal 81 includes a CPU 811, a ROM 812, a RAM 813, a storage device 814, an input / output IF (Interface) 815, and a communication IF 816. The CPU 811, the ROM 812, the RAM 813, the storage device 814, the input / output IF 815, and the communication IF 816 are connected via a bus 819 so as to be able to communicate with each other.
[0023] The CPU 811 is a central processing unit. The CPU 811 controls each part of the terminal 81 using an application program. The ROM 812 is a read-only memory. The ROM 812 is made up of non-volatile memory, and stores application programs for controlling each part of the terminal 81. The RAM 813 is a random access memory. The RAM 813 provides a memory area necessary for the operation of the CPU 811. The storage device 814 is a large-capacity storage device such as a hard disk drive.
[0024] The input / output IF 815 is an input / output interface that receives voice input from the user, outputs voice to the user, and transmits and receives data between the terminal 81 and another device (for example, the terminal 81 owned by the user 3, the sound source generating device 1, or the server 82). The input / output IF 815 may include a mouse, a trackball, a touch panel, a keyboard, a speaker, etc. The user can operate the terminal 81 via the input / output IF 815. The communication IF 816 performs mutual data communication between the terminal 81, the server 82, and the sound source generating device 1 via wired communication and / or wireless communication. Note that the server 82 may have the same hardware configuration as the terminal 81. In this embodiment, for the sake of simplicity, it is assumed that the terminal 81 and the server 82 have the same hardware configuration.
[0025] 4 is a block diagram showing a functional unit 100 of the sound source generation device 1 according to an embodiment of the present disclosure. The functions of the functional unit 100 can be realized by the CPU 11 executing a program stored in the storage device 14 or the like. The functional unit 100 includes a control unit 101, an acquisition unit 102, an estimation unit 103, a generation unit 105, a playback unit 108, and a database 109. The control unit 101 controls each unit of the functional unit 100 and manages data transmission and reception between each unit of the functional unit 100, the processing order, notifications from the sound source generation device 1 to the user 3, data transmission and reception between the terminal 81 and the server 82 and the functional unit 100, etc.
[0026] The acquisition unit 102 acquires audio signals, image signals, video signals, biosignals of the user 3, etc. from the sensor 17. If the sensor 17 includes an infrared sensor, a speedometer, an accelerometer, etc., the acquisition unit 102 further acquires infrared images, infrared videos, speed, acceleration, etc. The acquisition unit 102 acquires signals from the sensor 17 and outputs them to the estimation unit 103, the generation unit 105, and the playback unit 108. The acquisition unit 102 may also directly acquire a target sound source 5 from the Internet, a brick-and-mortar store, a news site, etc. via the network 9 and the communication IF 16. For example, if the acquisition unit 102 acquires a warning sound such as an emergency earthquake alert, the target sound source 5 may be the warning sound. For example, if the acquisition unit 102 acquires advertising sound from a store in a shopping mall, the target sound source 5 may be the store, and the target signal 51 may be the advertising sound. If the user 3 is making a call with a remote terminal 81, the target sound source 5 may be the terminal 81, and the target signal 51 may be the call voice.
[0027] The estimation unit 103 acquires a sound signal, an image signal, a video signal, etc. from the acquisition unit 102, and estimates sound source information and statistical information of the observed sound source 7 and the observed signal 71. The sound source information of the observed sound source 7 and the observed signal 71 may be, for example, the position of the observed sound source 7, the arrival direction of the observed signal 71 from the observed sound source 7 to the sound source generating device 1, the fundamental frequency of the observed signal 71, the rise time of the observed signal 71, and the fall time of the observed signal 71. The statistical information of the observed sound source 7 and the observed signal 71 may be, for example, the average of the observed signal 71, the variance of the observed signal 71, the distribution function of the observed signal 71, the frequency distribution of the observed signal 71, the time distribution of the observed signal 71, and the average amplitude of the observed signal 71. The estimation unit 103 may estimate the observed signal 71 that may be generated by the observed sound source 7 at a future time based on the statistical information of the observed sound source 7 and the observed signal 71. When the target sound source 5 is a sound source around the user 3, the estimation unit 103 may estimate statistical information of the target sound source 5 and the target signal 51. For example, when the target sound source 5 is the terminal 81 and the target signal 51 is a call voice, the estimation unit 103 may analyze the call voice and estimate statistical information of the call voice.
[0028] The estimation unit 103 may perform image recognition of the observed sound source 7 based on the image signal or video signal. For example, if the observed sound source 7 is an ambulance, the estimation unit 103 can estimate that the observed signal 71 includes a siren sound, a running sound, and a human voice. That is, the estimation unit 103 can estimate sound source information and statistical information of the target sound source 5, the observed sound source 7, and the observed signal 71 based on multiple signals acquired by the sensor 17.
[0029] The generation unit 105 generates a target signal 51 having reproduction information different from the sound source information, according to sound source information and statistical information of the observed sound source 7 and the observed signal 71 in the estimation unit 103. When the estimation unit 103 performs image recognition or the like to further estimate the sound source information of the observed sound source 7 and the observed signal 71, the generation unit 105 may generate the target signal 51 according to the sound source information estimated from the image signal and video signal. For example, when the user 3 listens to the target signal 51 using earphones or headphones, and the generation unit 105 generates the target signal 51 having an arrival direction different from that of the observed signal 71, the generation unit 105 adds an interaural time difference and an interaural sound pressure difference according to the arrival direction of the target sound source 51 to the target signal 51 presented to the left ear and the right ear of the user 3. The sound source generation device 1 may acquire the head-related transfer characteristics of the user 3 in advance and store them in the database 109. The generation unit 105 may convolve the target signal 51 to be presented to the left ear of the user 3 and the target signal 51 to be presented to the right ear of the user 3 with the head-related transfer characteristic of the user 3 according to the direction from which the target sound source 5 arrives.
[0030] The generation unit 105 reads out a voice database or the like stored in the database 109 and generates the target signal 51. For example, if the estimation unit 103 estimates that the observed signal 71 is a male voice, the generation unit 105 may read out a female voice from the database 109 and generate the target signal 51. The generation unit 105 may generate a synthetic voice using the voice database stored in the database 109. The generation unit 105 may acquire a voice signal from an external device or the like other than the sound source generation device 1 and generate the target signal 51. The external device may be, for example, the terminal 81, the server 82, a cloud server via a network such as the Internet, or broadcasting equipment owned by a store visited by the user 3.
[0031] The reproduction unit 108 reproduces the target signal 51 and the observation signal 71 and presents the target signal 51 and the observation signal 71 to the user 3. When the user 3 directly hears sounds from sound sources around the user 3, the reproduction unit 108 may reproduce only the target signal 51 and present only the target signal 51 to the user 3. Depending on the manner in which the user 3 uses the sound source generating device 1, the reproduction unit 108 can select the signal to reproduce.
[0032] The database 109 stores a speech database, a labeled speech corpus, and head-related transfer functions of the user 3 in a part of the area of the storage device 14. The database 109 may store image patterns and the like used for image recognition in the estimation unit 103. When the estimation unit 103 estimates the observed sound source 7 or the target sound source 5 based on machine learning or the like, the database 109 may store coefficients and the like used for machine learning.
[0033] The functional unit 100 may further include a determination unit. The determination unit may compare the sound source information and statistical information estimated by the estimation unit 103 with predetermined determination conditions to determine the sound source information and reproduction information that the generation unit 105 can use. For example, if the arrival direction of the observed signal 71 is to the left of the user 3, the determination unit may determine that the arrival direction to be added to the target signal 51 is to the right of the user 3. Alternatively, if the average amplitude of the observed signal 71 is greater than a predetermined threshold, the determination unit may determine that the amplitude of the target signal 51 is not used in the reproduction information. The determination unit may also set priorities of multiple parameters included in the reproduction information. For example, the priorities may be the arrival direction, average amplitude, time distribution, center frequency, and frequency distribution, from highest to lowest. The estimation unit 103 may estimate the sound source information of the observed sound source 7 and the observed signal 71 according to the priorities. The determination unit may also determine a combination of reproduction information that can be used according to the available sound source information estimated by the estimation unit 103. For example, if the observed signal 71 and the target signal 51 are sounds or the like distributed in the same frequency band, the determination unit can determine that it is desirable to add reproduction information that combines the arrival direction of the target signal 51 and the average amplitude of the target signal 51 to the target signal 51. Combinations of reproduction information according to sound source information may be defined in advance and stored in the database 109, or may be determined based on the output of a machine learning model or the like. The combination of reproduction information may be, for example, a correspondence table between sound source information and reproduction information. The combination of reproduction information is not limited to two different pieces of reproduction information, but may also be a combination of three or more pieces of reproduction information.
[0034] The predetermined judgment condition may be set in advance by the user 3. For example, if the user 3 wants the arrival direction of the target signal 51 to be to the right of the user 3, the user 3 can limit the arrival direction of the target signal 51 to the right. If the user 3 does not want to increase the amplitude of the target signal 51, the user 3 can delete the judgment on the amplitude of the target signal 51 from the predetermined judgment conditions. In this embodiment, for simplicity of explanation, the estimation unit 103 and the generation unit 105 are assumed to realize some or all of the functions of the judgment unit.
[0035] With reference to FIG. 5-9, the sound source information and statistical information of the observed signal 71 estimated by the estimation unit 103 and the reproduction information and statistical information of the target signal 51 generated by the generation unit 105 will be described. In the embodiment of FIG. 5-9, the target sound source 5 is assumed to be a virtual sound source in virtual reality, augmented reality, or mixed reality. The function unit 100 can control the position, arrival direction, amplitude and phase, fundamental frequency, center frequency, and frequency distribution of the target sound source 5 and the target signal 51. The parameters of the target sound source 5 and the target signal 51 that can be controlled by the function unit 100 are not limited to those described above.
[0036] 5 is a schematic diagram showing the direction of arrival of an observation signal 71 and the direction of arrival of a target signal 51 according to an embodiment of the present disclosure, and is a top view of the head of a user 3. In FIG. 5, the observation signal 71 is directed in the direction of arrival D, with the front of the user 3 as the reference. O The acquisition unit 102 acquires a sound signal, an image signal, a video signal, etc., and the estimation unit 103 estimates the arrival direction D of the observed signal 71. O The generation unit 105 estimates the estimated arrival direction D O Depending on the arrival direction D O Arrival direction D is different from T That is, the sound source information is generated based on the arrival direction D O and the reproduced information is the arrival direction D T The generation unit 105 may calculate the arrival direction D of the observed signal 71. O Arrival direction D is different from T By adding this information to the target signal 51, it may become easier for the user 3 to distinguish between the observed signal 71 and the target signal 51.
[0037] When the observed sound source 7 moves around the user 3, the arrival direction D of the observed signal 71 O For example, if the observed sound source 7 moves around the user 3 in the direction M O When moving to the direction of arrival D O Also direction M O In such a case, the estimation unit 103 estimates the arrival direction DO and direction M O and the arrival direction D estimated by the generation unit 105 is O and direction M O Depending on the arrival direction D O and direction M O Arrival direction D is different from T and direction M T That is, the sound source information may be expressed as the arrival direction D O and direction M O and the reproduced information is the arrival direction D T and direction M T For example, in the direction M T is direction M O (i.e., direction M T is direction M O Alternatively, the arrival direction D O is direction M O , the generation unit 105 determines the arrival direction D T direction M so that T may be fixed in a predetermined direction.
[0038] In this embodiment, the target sound source 5 and the target signal 51 are not necessarily sound sources that exist in reality. T In this embodiment, the target signal 51 does not necessarily arrive at the user 3 from the arrival direction D T The target signal 51 having the direction of arrival D T The user 3 can perceive the target sound source 5 to be present in the arrival direction D T This means that the target signal 51 is perceived as coming from the direction M. T The same is true for the target signal 51 having
[0039] 6 is a schematic diagram showing the time waveform of the observed signal 71 and the time waveform of the target signal 51 according to an embodiment of the present disclosure, which shows the time waveform of the observed signal 71 and the time waveform of the target signal 51 from time t1 to time t2. The acquisition unit 102 acquires an audio signal, an image signal, a video signal, etc., and the estimation unit 103 estimates the amplitude A O The generation unit 105 estimates the estimated amplitude A O Depending on the amplitude A O and amplitude A T That is, the sound source information is generated as a target signal 51 having an amplitude A O and the reproduced information has amplitude A T For example, when the target signal 51 is reproduced in the same time interval as the observed signal 71, the generating unit 105 generates an amplitude A O Amplitude A is greater than T The target signal 51 may have an amplitude A T is the amplitude A of the observed signal 71. O When the amplitude A of the observed signal 71 estimated by the estimation unit 103 is larger than O may be the maximum amplitude of the observed signal 71 or the average amplitude of the observed signal 71. Similarly, the amplitude A T may be the maximum amplitude of the target signal 51 or the average amplitude of the target signal 51 .
[0040] FIG. 7 is a schematic diagram illustrating the time waveform of the observed signal 71 and the time waveform of the target signal 51 according to an embodiment of the present disclosure. The time waveform of the observed signal 71 is from time t1 to time t2, and the time waveform of the target signal 51 is from time t3 to time t4. The acquisition unit 102 acquires an audio signal, an image signal, a video signal, etc., and the estimation unit 103 estimates the rise time t1 and fall time t2 of the time signal of the observed signal 71. The generation unit 105 generates the target signal 51 having a time interval [t3, t4] different from the time interval [t1, t2] based on the estimated rise time t1 and fall time t2. That is, the sound source information may be the time interval [t1, t2], and the reproduction information may be the time interval [t3, t4]. For example, if the target signal 51 can be reproduced in a time interval different from that of the observed signal 71, the generation unit 105 may generate the target signal 51 in a time interval [t3, t4] different from the time interval [t1, t2] in which the time signal of the observed signal 71 exists. The time intervals [t1, t2] and [t3, t4] may partially overlap each other. It is preferable that the time intervals [t1, t2] and [t3, t4] are separated from each other on the time axis to such an extent that temporal masking does not occur between the observed signal 71 and the target signal 51.
[0041] FIG. 8 is a schematic diagram illustrating the frequency distribution of an observed signal 71 and the frequency distribution of a target signal 51 according to an embodiment of the present disclosure, where the observed signal 71 has a center frequency f O and bandwidth W O The acquisition unit 102 acquires an audio signal, an image signal, a video signal, etc., and the estimation unit 103 estimates the center frequency f of the observation signal 71. O and bandwidth W O The generating unit 105 estimates the estimated center frequency f O and bandwidth W O Depending on the center frequency f T and bandwidth W T That is, the sound source information is generated as a target signal 51 having a center frequency f O and bandwidth W O and the reproduced information has a center frequency f T and bandwidth WT For example, when the target signal 51 is reproduced in the same time interval as the observed signal 71, the generator 105 may generate a signal having a center frequency f O and a different center frequency f T The target signal 51 may have a center frequency f T is the center frequency f of the observed signal 71. O , the user 3 can easily distinguish between the observed signal 71 and the target signal 51. The frequency distribution of the observed signal 71 and the frequency distribution of the target signal 51 may partially overlap. The center frequency f on the frequency axis is set to a value that does not cause simultaneous masking (frequency masking) between the observed signal 71 and the target signal 51. O and center frequency f T and the bandwidth W O and bandwidth W T It is preferable that the and do not overlap.
[0042] The estimation unit 103 may further estimate the fundamental frequency of the observed signal 71. Here, the fundamental frequency is the lowest frequency contained in the signal. For example, if the observed signal 71 is a human voice and the target signal 51 is a human voice, it is desirable that the fundamental frequency of the observed signal 71 and the fundamental frequency of the target signal 51 are far enough apart that simultaneous masking between the observed signal 71 and the target signal 51 does not occur. For example, if the observed signal 71 is a male voice, the generation unit 105 may select a female voice and generate it as the target signal 51.
[0043] 9 is a schematic diagram showing the time waveform of an observed signal 71 and the time waveform of a target signal 51 according to an embodiment of the present disclosure, which is the time waveform of the periodically occurring observed signal 71. The acquisition unit 102 acquires an audio signal, an image signal, a video signal, etc., and the estimation unit 103 estimates the time interval P O The generation unit 105 estimates the estimated time interval P O and the period, the time interval P during which the observed signal 71 does not occur. TFor example, if the observation signal 71 is an intermittent sound generated by a press machine, the generation unit 105 generates the target signal 51 in a time interval when the observation signal 71 is not generated. If the target signal 51 is a voice, the generation unit 105 may divide the target signal 51 into syllables and generate syllables 511, 512, and 513. For example, if the target signal 51 is a voice reading "The next test is a pressure resistance test of item 3-1-1," the syllable 511 may be "The next test is," the syllable 512 may be "of item 3-1-1," and the voice 513 may be "It is a pressure resistance test." The number of syllables included in the syllables 511, 512, and 513 is not limited to one syllable, and may be generated over a time interval P T The number of syllables may be in accordance with the number of syllables in syllable 511. Note that some of the syllables in syllable 511, syllables in syllable 512, and syllables in syllable 513 may overlap. For example, syllable 511 may be "The next test is item 3-1-1's," and syllable 512 may be "The pressure resistance test of item 3-1-1."
[0044] 5 to 9 may be combined so that the generator 105 generates the target signal 51. For example, the generator 105 may generate the target signal 51 by detecting the arrival direction D of the observed signal 71. O Arrival direction D is different from T and the amplitude A of the observed signal 71 O and amplitude A T and the center frequency f of the observation signal 71 O and a different center frequency f T As long as the user 3 can distinguish between the target signal 51 and the observed signal 71 and can hear the target signal 51, the combination of playback information that the generation unit 105 adds to the target signal 51 is not limited to the above.
[0045] 5-9, the configuration in which the generator 105 controls the reproduction information of the target signal 51 has been described. However, when the reproducer 108 presents the target signal 51 and the observed signal 71 to the user 3, the generator 105 may control the sound source information of the observed signal 71. For example, the generator 105 may control the amplitude A O The amplitude A of the target signal 51 TThe generator 105 frequency-modulates the observed signal 71 to obtain a center frequency f O and the center frequency f of the target signal 51 T and may be at different frequencies.
[0046] FIG. 10 is a schematic diagram showing an observed sound source 71, a target sound source 51, and attention information of a user 3 according to an embodiment of the present disclosure. In FIG. 10, an ambulance (target sound source 5), a sports car (observed sound source 7A), and a large truck (observed sound source 7B) are traveling on a roadway. When the ambulance (target sound source 5) emits a siren sound (target signal 51), the user 3 hears the siren sound (target signal 51) directly or via the playback unit 108. At this time, when the user 3 directs his / her line of sight 31 toward the ambulance (target sound source 5), the acquisition unit 102 acquires gaze information of the user 3 via the sensor 17. Based on the gaze information of the user 3, the estimation unit 103 estimates an object or the like that the user 3 is paying attention to. In the example of FIG. 10, the estimation unit 103 estimates that the user 3 is paying attention to the ambulance (target sound source 5). When a sports car (observed sound source 7A) emits an engine sound (observed signal 71) that is louder than a siren sound (target signal 51), and a truck (observed sound source 7B) is emitting an engine sound (observed signal 71) near the user 3, it may be difficult for the user 3 to hear the siren sound (target signal 51). In such a case, for example, the generation unit 105 may increase the amplitude of the target signal 51, and the reproduction unit 108 may present the target signal 51 to the user 3.
[0047] When the user 3 is using earphones or headphones, the control unit 101 and the acquisition unit 102 may perform signal processing to form the directivity of the microphone array based on the line-of-sight information of the user 3. For example, the control unit 101 and the acquisition unit 102 may form the directivity of the microphone array toward an ambulance (target sound source 5), or may form blind spots of the microphone array toward a sports car (observed sound source 7A) and a large truck (observed sound source 7B). The reproduction unit 108 may present to the user 3 the target signal 51 acquired by forming directivity toward the target sound source 5, or may present to the user 3 the target signal 51 acquired by forming blind spots toward the observed sound source 7A and the observed sound source 7B. The generation unit 105 may add the reproduction information exemplified in FIG. 5-9 to the target signal 51 acquired by forming directivity toward the target sound source 5 or the target signal 51 acquired by forming blind spots toward the observed sound source 7A and the observed sound source 7B, and the reproduction unit 108 may present the target signal 51 to the user 3.
[0048] 11 is a schematic diagram showing an observed sound source 71, a target sound source 51, and attention information of a user 3 according to an embodiment of the present disclosure. At the work site in FIG. 11, the user 3 connects a mobile terminal (corresponding to a terminal 81 carried by the user 3) to the sound source generating device 1 and is making a voice call with another user. In parallel with the call with the other user, the user 3 is also talking with a technician at the work site. At the work site, when a work location (observed sound source 7) that is emitting a loud work sound (observed signal 71) is located in front of the right hand side of the user 3, the estimation unit 103 estimates the arrival direction D of the observed signal 71 arriving from the observed sound source 7. O The estimated arrival direction D O Based on this, the generation unit 105 calculates the arrival direction D of the other user's call voice (target sound source 5 or target signal 51). O Arrival direction D is different from T The generation unit 105 adds the reproduction information of the arrival direction D O and arrival direction D T The angle difference between the arrival direction D and TIt is desirable to add the reproduction information of the target signal 51 to the call voice (target sound source 5 or target signal 51). The predetermined angle can be, for example, 30 degrees, 45 degrees, 90 degrees, etc. The predetermined angle is not limited to the above, and may be any angle that makes it easier for the user 3 to hear the target signal 51 than the observed signal 71. The reproduction unit 108 detects the arrival direction D T The target signal 51 including the playback information is presented to the user 3. This can make it easier for the user 3 to hear the target signal 51 near the work location (observed sound source 7) that is emitting a loud work sound.
[0049] When the user 3 is prompted by another user to talk to the engineer about the work content, the user 3 directs his / her line of sight 31A toward the engineer (target sound source 5). The estimation unit 103 estimates that the user 3 is paying attention to the engineer (target sound source 5), and the acquisition unit 102 acquires the voice (target signal 51) of the engineer (target sound source 5). For example, the estimation unit 103 estimates the amplitude A of the observed signal 71. O The generation unit 105 estimates the amplitude A O Amplitude A is greater than T may be added to the target signal 51. The reproducing unit 108 adds the amplitude A T The target signal 51 including the playback information is presented to the user 3. This can make it easier for the user 3 to hear the target signal 51 near the work location (observed sound source 7) that is emitting a loud work sound.
[0050] After user 3 has spoken with the engineer, when user 3 is again engaged in a voice call with another user, the work location (observed sound source 7) emits a periodic work sound (observed signal 71) as shown in FIG. 9. In response to the work sound (observed signal 71), user 3 directs his / her gaze 31B toward the work location (observed sound source 7). At this time, for example, user 3 may frown or show a feeling of discomfort, directing his / her gaze 31B toward the work location (observed sound source 7). Based on the gaze 31B of user 3 and biometric information such as the facial expression of user 3, the estimation unit 103 estimates that user 3 is not paying attention to the work location (observed sound source 7) or that user 3 does not want to hear the observed signal 71 from the work location (observed sound source 7). Here, the amplitude A of the work sound (observed signal 71) Ois greater than the predetermined threshold, the generation unit 105 increases the amplitude A of the target signal 51 T with amplitude A O If the reproduction unit 108 presents the target signal 51 to the user 3, the ears of the user 3 may be damaged. In such a case, the generation unit 105 T The amplitude A of the observed signal 71 is O , and the reproduction unit 108 reduces the amplitude A O Alternatively, the generation unit 105 may divide a call voice from another user into a plurality of syllables, and the reproduction unit 108 may present the target signal 51 including the divided syllables to the user 3. That is, based on the gaze information of the user 3 and the biometric information of the user 3 (for example, a face image, blood pressure, pulse, speed, acceleration, etc.), the estimation unit 103 can estimate the sound source around the user 3 as the target sound source 5 or the observed sound source 7. The information used by the estimation unit 103 when estimating the sound source around the user 3 as the target sound source 5 or the observed sound source 7 is not limited to the above. For example, it may be an image signal or a video signal acquired by a camera or the like, or notification information via the communication IF 16. The notification information may be, for example, a character string indicating an advertising message or a character string indicating a warning message. The estimation unit 103 may analyze the character string or the like and estimate the sound source around the user 3 as the target sound source 5 or the observed sound source 7 according to the meaning or the like indicated by the character string or the like.
[0051] The combination of biometric information of user 3 used to estimate an object that user 3 is paying attention to or is not paying attention to is not limited to the above. For example, it may be gaze information of user 3 and the pulse rate of user 3, or a facial image of user 3 and the average sound pressure level of observed signals 71 around user 3. The sound source information of observed signals 71 controlled by generation unit 105 is not limited to the amplitude of observed signals 71. It may also be possible to control the spatial position of observed signals 71, the direction of arrival of observed signals 71, the frequency distribution of observed signals 71, etc.
[0052] 12 is a sequence chart of a sound source generation method according to an embodiment of the present disclosure. The embodiment in FIG. 12 illustrates a scene in which a user 3 hears an observation signal 71 coming from around the user 3 and a target signal 51 of a virtual sound source in virtual reality, augmented reality, or mixed reality. A sound source generation device 1 acquires the observation signal 71 from an observation sound source 7 around the user 3, generates a virtual sound source (target sound source 5) that differs from the sound source information of the observation signal 71, and presents the target signal 51 to the user 3.
[0053] The acquisition unit 102 acquires an observation signal 71 arriving from around the user 3 via the sensor 17 of the sound source generation device 1 (step S100). If the sensor 17 has multiple microphones, the acquisition unit 102 stores the observation signals 71 acquired by the multiple microphones as separate signals in the storage device 14 or the like. For example, if the sensor 17 has three microphones, A, B, and C, the acquisition unit 102 stores an observation signal 71A acquired by microphone A, an observation signal 71B acquired by microphone B, and an observation signal 71C acquired by microphone C in the storage device 14 or the like. When the acquisition unit 102 stores the observation signals 71 in the storage device 14 or the like, the respective positions of the multiple microphones or the relative positional relationship of the multiple microphones may be stored in the storage device 14 or the like. The positions of the multiple microphones may be stored in the database 109 in advance.
[0054] The acquisition unit 102 transmits the observed signal 71 acquired in step S100 to the estimation unit 103 (step S101). If the sensor 17 has multiple microphones, the acquisition unit 102 transmits multiple observed signals 71 acquired by the multiple microphones to the estimation unit 103.
[0055] The estimation unit 103 receives the observed signal 71 from the acquisition unit 102 and estimates sound source information of the observed signal 71 (step S102). The estimation unit 103 estimates at least one of the arrival direction of the observed signal 71, the movement direction of the observed signal 71, the amplitude of the observed signal 71, the time interval in which the observed signal 71 exists, and the center frequency, bandwidth, and fundamental frequency of the observed signal 71. The estimation unit 103 may further estimate statistical information of the observed signal 71, the period in which the observed signal 71 occurs, etc. The estimation unit 103 may request the acquisition unit 102 to acquire an image signal, and may estimate the position of the observed sound source 7 that generates the observed signal 71 based on the image signal.
[0056] For example, the estimation unit 103 estimates the average amplitude of the observed signal 71 and compares the estimated average amplitude with a predetermined threshold. If the estimated average amplitude is greater than the predetermined threshold, increasing the amplitude of the target signal 51 sufficiently beyond the estimated average amplitude may damage the ears of the user 3. The estimation unit 103 estimates sound source information other than the estimated average amplitude. The estimation unit 103 may estimate the direction of arrival of the observed signal 71. Here, if the direction of arrival of the observed signal 71 cannot be estimated or if there are multiple observed signals 71, it is not possible to add a direction of arrival different from the direction of arrival of the observed signal 71 to the target signal 51. In such cases, the estimation unit 103 may further estimate the center frequency, bandwidth, fundamental frequency, etc. of the observed signal. For example, if the observed signal 71 is the voices of multiple men, the estimation unit 103 may send a notification to the generation unit 105 to select a female voice as the target signal 51. That is, the estimation unit 103 may transmit to the generation unit 105 a notification or the like to select the target signal 51 that is distributed in a frequency band different from the frequency band in which the observed signal 71 is distributed.
[0057] When a plurality of observed sound sources 7 are present around the user 3 and the acquisition unit 102 acquires an observed signal 71 in which a plurality of different signals are mixed, the estimation unit 103 may select an observed signal 71 generated from a dominant observed sound source 7 among the plurality of observed sound sources 7, and estimate sound source information. When the sensor 17 has a plurality of microphones, for example, the estimation unit 103 may select the observed signal 71 of the dominant observed sound source 7 using principal component analysis or the like. The estimation unit 103 may select the observed signal 71 based on an eigenvalue, a contribution rate, a cumulative contribution rate, or the like. Alternatively, the estimation unit 103 may select the observed signal 71 of the dominant observed sound source 7 using a sound source separation method such as independent component analysis. After estimating the sound source information of the observed signal 71, the estimation unit 103 transmits the sound source information estimated in step S102 to the generation unit 105 (step S103).
[0058] The generation unit 105 receives the sound source information from the estimation unit 103 and generates the target signal 51 according to the sound source information (step S104). For example, when the estimation unit 103 selects the average amplitude of the observed signal 71 as the sound source information, the generation unit 105 generates the target signal 51 having an average amplitude that is sufficiently larger than the average amplitude of the observed signal 71. When the estimation unit 103 selects the arrival direction of the observed signal 71 as the sound source information, the generation unit 105 generates the target signal 51 having an arrival direction different from the arrival direction of the observed signal 71. When the generation unit 105 generates the target signal 51, it may convolve a head-related transfer characteristic of the user 3 according to the arrival direction to be added to the target signal 51. By convolving the head-related transfer characteristic of the user 3 with the target signal 51, it is possible to present the target signal 51 to the user 3, which has an interaural time difference, an interaural sound pressure difference, and a difference in frequency characteristics between the ears according to the arrival direction. When the observed signal 71 is a plurality of male and female voices, the target signal 51 is a read-out voice, and the estimation unit 103 selects sound source information for the plurality of observed signals 71, it may be difficult to differentiate between the observed signal 71 and the target signal 51. In such a case, the generation unit 105 may select a synthetic voice as the target signal 51. It is desirable that the generation unit 105 selects and generates the target signal 51 so that the user 3 can sufficiently distinguish between the observed signal 71 and the target signal 51.
[0059] The reproduction information used by the generation unit 105 when generating the target signal 51 may be determined by the generation unit 105 or by the estimation unit 103. When the estimation unit 103 determines the reproduction information, the estimation unit 103 can end estimation of the sound source information of the observed signal 71 when predetermined reproduction information is obtained. When the generation unit 105 determines the reproduction information, the estimation unit 103 estimates all necessary reproduction information and transmits it to the generation unit 105. The user 3 can modify the presented target signal 51 so that the reproduction information desired by the user 3 is added.
[0060] The generation unit 105 transmits the target signal 51 to the reproduction unit 108 (step S105), and the reproduction unit 108 presents the target signal 51 to the user 3 (step S106). When the user 3 wears a reproduction device that presents sound directly to the left and right ears, such as headphones or earphones, the reproduction unit 108 presents the target signal 51 to the user 3 via the headphones or earphones. When the user 3 listens to the target signal 51 via a speaker, the reproduction unit 108 presents the target signal 51 to the user 3 via the speaker.
[0061] Note that when the user 3 listens to the target signal 51 and the observed signal 71 using earphones or headphones, the estimation unit 103 may transmit the observed signal 71 received from the acquisition unit 102 in step S101 or the observed signal 71 processed in step S102 to the generation unit 105. For example, when the estimation unit 103 separates the observed signal 71 into separate sound source signals using a sound source separation technique, the generation unit 105 can individually control the sound source information of the separate sound source signals included in the observed signal 71 and present the observed signal 71 to the user 3 via the playback unit 108. In augmented reality or mixed reality, the user 3 can listen to the target signal 51 while listening to surrounding environmental sounds, etc.
[0062] In the sound source generating device 1, the estimation unit 103 estimates the sound source information of the observed signal 71, and the generation unit 105 generates the target signal 51 having playback information corresponding to the sound source information, and presents it to the user 3, thereby enabling the user 3 to distinguish between the observed signal 71 and the target signal 51.
[0063] Fig. 13 is a sequence chart of a sound source generation method according to an embodiment of the present disclosure. Similar to the embodiment of Fig. 11, the embodiment of Fig. 13 illustrates a scene in which a user 3 is having a voice call with another user in augmented reality or mixed reality, in which the user 3 hears an observed signal 71 arriving from around the user 3 and the call voice (target sound source 5 or target signal 51) with the other user. The process of acquiring the observed signal 71 (step S200) and the process of transmitting the observed signal 71 to the estimation unit 103 (step S201) in Fig. 13 are similar to the process of acquiring the observed signal 71 (step S100) and the process of transmitting the observed signal 71 to the estimation unit 103 (step S101) in Fig. 12. 13 is the same as the process of transmitting the target signal 51 to the playback unit 108 (step S105) and the process of presenting the target signal 51 to the user 3 (step S106) in Fig. 12. Therefore, the description of the process of step S200, the process of step S201, the process of step S210, and the process of step S211 will be omitted.
[0064] The acquisition unit 102 acquires an observation signal 71 arriving from the surroundings of the user 3 via a microphone or the like of the sensor 17 included in the sound source generation device 1 (step S200) and transmits the observation signal 71 to the estimation unit 103 (step S201). The acquisition unit 102 further captures an image signal via a camera or the like (step S202) and transmits the image signal to the estimation unit 103 (step S203). The acquisition unit 102 acquires an image signal or a video signal captured around the user 3 by a camera or the like. When the sensor 17 has one camera, it is preferable that the camera captures an image in front of the user 3 or in the same direction as the user 3's line of sight. The acquisition unit 102 further acquires biometric information of the user 3 via the sensor 17 (step S204) and transmits the biometric information to the estimation unit 103 (step S205). For example, the camera may capture both eyes and a facial image of the user 3, and the acquisition unit 102 may acquire the gaze information and facial expression of the user 3. The acquisition unit 102 may further acquire the pulse, blood pressure, etc. of the user 3. It is desirable that the acquisition of the observed signal 71 in step S200, the acquisition of the image signal or video signal in step S202, and the acquisition of the biometric information of the user 3 in step S204 are performed simultaneously.
[0065] The estimation unit 103 receives the observation signal 71, the image signal or video signal, and biometric information of the user 3 from the acquisition unit 102, and estimates sound source information of the observation signal 71 and attention information of the user 3 (step S206). The estimation unit 103 associates statistical information of the observation signal 71, image information of the observed sound source 7 or video information of the observed sound source 7 included in the image signal or video signal, and gaze information of the user 3, etc. For example, as in the embodiment of FIG. 11 , when the image signal includes a work location (observed sound source 7) and the gaze of the user 3 is directed toward the work location (observed sound source 7), the estimation unit 103 estimates whether the target corresponding to the gaze information of the user 3 is the target sound source 5 or the observed sound source 7 based on biometric information such as facial expression included in the facial image of the user 3. In the embodiment of FIG. 11 , the estimation unit 103 estimates that the target corresponding to the gaze information of the user 3 is the work location and the observed sound source 7. Since user 3 is engaged in a voice call with another user, estimation unit 103 may consider all voice signals other than the voice signal from terminal 81 to be observed signals 71 .
[0066] The estimation unit 103 further estimates sound source information of the observed signal 71 using information on the observed sound source 7 acquired using statistical information on the observed signal 71, image information on the observed sound source 7 or video information on the observed sound source 7 included in the image signal or video signal, and biometric information of the user 3. The estimation unit 103 estimates at least one of the arrival direction of the observed signal 71, the movement direction of the observed signal 71, the amplitude of the observed signal 71, the time interval in which the observed signal 71 exists, and the center frequency, bandwidth, and fundamental frequency of the observed signal 71. The estimation unit 103 may further estimate the statistical information of the observed signal 71, the period in which the observed signal 71 occurs, etc. The estimation unit 103 may transmit the observed signal 71, the image signal and video signal, and the biometric information of the user 3 to the server 82, and the server 82 may estimate the sound source information of the observed sound source 7 and the observed signal 71 using a model based on machine learning. In addition, if the estimation unit 103 can estimate the arrival direction or position of the observed sound source 7 based on the image signal or video signal, the estimation unit 103 may omit direction estimation using the observed signal 71.
[0067] The estimation unit 103 transmits the sound source information estimated in step S206 to the generation unit 105 (step S207). If the sensor 17 has multiple microphones and can form a microphone array, the estimation unit 103 may transmit not only the estimated sound source information but also the position of the observed sound source 7 or the arrival direction of the observed signal 71 and an instruction to form the directivity or blind spot of the microphone array to the generation unit 105. Furthermore, the generation unit 105 receives a voice signal from the terminal 81 (step S208) and generates a target signal 51 using the voice signal (step S209). The generation unit 105 adds reproduction information to the voice signal from the terminal 81 according to the sound source information. For example, the generation unit 105 may add a direction of arrival different from the arrival direction of the observed signal 71 to the voice signal of the terminal 81 to generate the target signal 51. 9, when the observed signal 71 is a loud noise that occurs periodically, the generation unit 105 may temporarily store the audio signal from the terminal 81 in the storage device 14, generate the stored audio signal from the terminal 81 during a time period when the observed signal 71 does not occur, and transmit the target signal 51 to the reproduction unit 108 (step S210). During a time period when the observed signal 71 does not occur, the reproduction unit 108 may present the target signal to the user 3 (step S211).
[0068] The estimation unit 103 estimates sound source information of the observation signal 71 based on signals from the multiple sensors 17, the generation unit 105 generates the target signal 51 having reproduction information corresponding to the sound source information, and the sound source generation device 1 can present the target signal 51 to the user 3. By controlling not only the reproduction information of the target signal 51 but also the sound source information of the observation signal 71 based on the biometric information of the user 3, the user 3 can distinguish between the observation signal 71 and the target signal 51.
[0069] [Other embodiments] 14 is a block diagram showing a sound source generation device according to an embodiment of the present disclosure. The sound source generation device 1000 includes an estimation unit 1001, a generation unit 1002, and a reproduction unit 1003. The estimation unit 1001 estimates sound source information of a sound source signal acquired by one or more sensors. The generation unit 1002 generates a reproduction signal having reproduction information different from the sound source information. The reproduction unit 1003 can present the reproduction signal to a user.
[0070] Furthermore, the scope of each embodiment also includes a processing method in which a program that operates the configuration of each embodiment to realize the functions of the above-described embodiments is recorded on a recording medium, the program recorded on the recording medium is read as code, and the program is executed on a computer. In other words, a computer-readable recording medium is also included in the scope of each embodiment. Furthermore, each embodiment includes not only a recording medium on which the above-described computer program is recorded, but also the computer program itself.
[0071] Examples of the recording medium that can be used include a floppy disk, a hard disk, an optical disk, a magneto-optical disk, a CD-ROM (Compact Disc-Read Only Memory), a magnetic tape, a non-volatile memory card, and a ROM. In addition, the scope of each embodiment is not limited to programs that execute processing by themselves recorded on the recording medium, but also includes programs that execute processing by operating on an OS (Operating System) in cooperation with other software and functions of an expansion board.
[0072] Although the present disclosure has been described above with reference to the embodiments, the present disclosure is not limited to the above-described embodiments. Various modifications that can be understood by those skilled in the art can be made to the configuration and detailed description of the present disclosure within the scope of the present disclosure. Furthermore, each embodiment can be combined with other embodiments as appropriate.
[0073] Some or all of the above-described embodiments can be described as, but are not limited to, the following supplementary notes.
[0074] (Appendix 1) an estimation means for estimating sound source information of a sound source signal acquired by a sensor; a generating means for generating reproduction information different from the sound source information and adding the reproduction information to a reproduction signal; a playback means for presenting the playback signal to a user; The sound source generating device, wherein the sound source information is the direction of arrival of the sound source signal, and the reproduction information is the direction of arrival of the reproduction signal.
[0075] (Appendix 2) the sound source information further includes at least one of a position of the sound source signal, a fundamental frequency of the sound source signal, a frequency distribution of the sound source signal, a rise time of the sound source signal, a fall time of the sound source signal, a time distribution of the sound source signal, an average amplitude of the sound source signal, and a generation period of the sound source signal; the reproduction information further includes at least one of a position of the reproduction signal, a fundamental frequency of the reproduction signal, a frequency distribution of the reproduction signal, a rise time of the reproduction signal, a fall time of the reproduction signal, a time distribution of the reproduction signal, an average amplitude of the sound source signal, and a generation period of the reproduction signal; 2. The sound source generating device according to claim 1, wherein the generating means selects the reproduction information based on the sound source information.
[0076] (Appendix 3) the reproduction information in which the average amplitude of the reproduction signal is different from the average amplitude of the sound source signal; the reproduction information in which the fundamental frequency of the reproduction signal is different from the fundamental frequency of the sound source signal; the reproduction information in which the frequency distribution of the reproduction signal and the frequency distribution of the sound source signal do not at least partially overlap; The reproduction information in which the rising time of the reproduction signal and the rising time of the sound source signal do not at least partially overlap, or the reproduction information in which the time distribution of the reproduction signal and the time distribution of the sound source signal do not at least partially overlap; 3. The sound source generating device according to claim 2, wherein the generating means selects:
[0077] (Appendix 4) 3. The sound source generating device according to claim 2, wherein the generating means divides the playback signal into a plurality of sections in accordance with the generation period of the sound source signal.
[0078] (Appendix 5) 5. The sound source generating device according to claim 4, wherein the generating means divides the playback signal into syllables.
[0079] (Appendix 6) The sensor further acquires a video signal of the user's surroundings and biological information of the user; the estimation means estimates attention information of the user based on the biometric information or the video signal; The sound source generating device according to claim 1, wherein the generating means selects the reproduction information based on the attention information.
[0080] (Appendix 7) the biometric information includes gaze information of the user, a face image of the user, and a pulse rate of the user; The sound source generating device according to claim 6, wherein the generating means changes the playback information or the sound source information based on the attention information.
[0081] (Appendix 8) the generating means selects a plurality of pieces of playback information according to the sound source information, The sound source generating device according to claim 2, wherein the plurality of pieces of playback information are selected based on a user selection or a predefined correspondence between the sound source information and the playback information.
[0082] (Appendix 9) an estimation step of estimating sound source information of a sound source signal acquired by a sensor; a generating step of generating reproduction information different from the sound source information and adding the reproduction information to a reproduction signal; a reproduction step of presenting the reproduction signal to a user; The sound source generating method, wherein the sound source information is the direction of arrival of the sound source signal, and the reproduction information is the direction of arrival of the reproduction signal.
[0083] (Appendix 10) the sound source information further includes at least one of a position of the sound source signal, a fundamental frequency of the sound source signal, a frequency distribution of the sound source signal, a rise time of the sound source signal, a fall time of the sound source signal, a time distribution of the sound source signal, an average amplitude of the sound source signal, and a generation period of the sound source signal; the reproduction information further includes at least one of a position of the reproduction signal, a fundamental frequency of the reproduction signal, a frequency distribution of the reproduction signal, a rise time of the reproduction signal, a fall time of the reproduction signal, a time distribution of the reproduction signal, an average amplitude of the sound source signal, and a generation period of the reproduction signal; 10. The sound source generating method according to claim 9, wherein the generating step selects the reproduction information based on the sound source information.
[0084] (Appendix 11) the reproduction information in which the average amplitude of the reproduction signal is different from the average amplitude of the sound source signal; the reproduction information in which the fundamental frequency of the reproduction signal is different from the fundamental frequency of the sound source signal; the reproduction information in which the frequency distribution of the reproduction signal and the frequency distribution of the sound source signal do not at least partially overlap; The reproduction information in which the rising time of the reproduction signal and the rising time of the sound source signal do not at least partially overlap, or the reproduction information in which the time distribution of the reproduction signal and the time distribution of the sound source signal do not at least partially overlap; 11. The sound source generation method according to claim 10, wherein the generating step selects:
[0085] (Appendix 12) 11. The sound source generating method according to claim 10, wherein the generating step divides the playback signal into a plurality of sections according to the generation period of the sound source signal.
[0086] (Appendix 13) 13. The sound source generation method according to claim 12, wherein the generating step divides the playback signal into syllables.
[0087] (Appendix 14) The sensor further acquires a video signal of the user's surroundings and biological information of the user; the estimating step estimates attention information of the user based on the biometric information or the video signal; 10. The sound source generation method according to claim 9, wherein the generating step selects the playback information based on the attention information.
[0088] (Appendix 15) the biometric information includes gaze information of the user, a face image of the user, and a pulse rate of the user; 15. The sound source generating method according to claim 14, wherein the generating step changes the playback information or the sound source information based on the attention information.
[0089] (Appendix 16) the generating step selects a plurality of pieces of playback information according to the sound source information; The sound source generation method according to claim 10, wherein the plurality of pieces of playback information are selected based on a user selection or a predefined correspondence between the sound source information and the playback information.
[0090] (Appendix 17) an estimation step of estimating sound source information of a sound source signal acquired by a sensor; a generating step of generating reproduction information different from the sound source information and adding the reproduction information to a reproduction signal; a reproduction step of presenting the reproduction signal to a user; The sound source generating program, wherein the sound source information is the arrival direction of the sound source signal, and the reproduction information is the arrival direction of the reproduction signal.
[0091] (Appendix 18) the sound source information further includes at least one of a position of the sound source signal, a fundamental frequency of the sound source signal, a frequency distribution of the sound source signal, a rise time of the sound source signal, a fall time of the sound source signal, a time distribution of the sound source signal, an average amplitude of the sound source signal, and a generation period of the sound source signal; the reproduction information further includes at least one of a position of the reproduction signal, a fundamental frequency of the reproduction signal, a frequency distribution of the reproduction signal, a rise time of the reproduction signal, a fall time of the reproduction signal, a time distribution of the reproduction signal, an average amplitude of the sound source signal, and a generation period of the reproduction signal; 18. The sound source generation program according to claim 17, wherein the generating step selects the playback information based on the sound source information.
[0092] (Appendix 19) the reproduction information in which the average amplitude of the reproduction signal is different from the average amplitude of the sound source signal; the reproduction information in which the fundamental frequency of the reproduction signal is different from the fundamental frequency of the sound source signal; the reproduction information in which the frequency distribution of the reproduction signal and the frequency distribution of the sound source signal do not at least partially overlap; The reproduction information in which the rising time of the reproduction signal and the rising time of the sound source signal do not at least partially overlap, or the reproduction information in which the time distribution of the reproduction signal and the time distribution of the sound source signal do not at least partially overlap; 19. The sound source generation program according to claim 18, wherein the generating step selects:
[0093] (Appendix 20) 19. The sound source generation program according to claim 18, wherein the generating step divides the playback signal into a plurality of sections according to the generation period of the sound source signal.
[0094] (Appendix 21) 21. The sound source generation program according to claim 20, wherein the generating step divides the playback signal into syllables.
[0095] (Appendix 22) The sensor further acquires a video signal of the user's surroundings and biological information of the user; the estimating step estimates attention information of the user based on the biometric information or the video signal; 18. The sound source generation program according to claim 17, wherein the generating step selects the playback information based on the attention information.
[0096] (Appendix 23) the biometric information includes gaze information of the user, a face image of the user, and a pulse rate of the user; 23. The sound source generating program according to claim 22, wherein the generating step changes the playback information or the sound source information based on the attention information.
[0097] (Appendix 24) the generating step selects a plurality of pieces of playback information according to the sound source information; 19. The sound source generating program according to claim 18, wherein the plurality of pieces of playback information are selected based on a user selection or a predefined correspondence between the sound source information and the playback information. [Explanation of symbols]
[0098] 1: Equipment 101: Control unit 102: Acquisition Department 103: Estimation part 105: Generation part 108: Playback section 109: Database 3: User 31: Gaze information 5: Target sound source 51: Target signal 7: Observed sound source 71: Observation signal 81: Terminal 82: Server 9: Network
Claims
1. an estimation means for estimating sound source information of a sound source signal acquired by a sensor; a generating means for generating reproduction information different from the sound source information and adding the reproduction information to a reproduction signal; a playback means for presenting the playback signal to a user; The sound source generating device, wherein the sound source information is the direction of arrival of the sound source signal, and the reproduction information is the direction of arrival of the reproduction signal.
2. the sound source information further includes at least one of a position of the sound source signal, a fundamental frequency of the sound source signal, a frequency distribution of the sound source signal, a rise time of the sound source signal, a fall time of the sound source signal, a time distribution of the sound source signal, an average amplitude of the sound source signal, and a generation period of the sound source signal; the reproduction information further includes at least one of a position of the reproduction signal, a fundamental frequency of the reproduction signal, a frequency distribution of the reproduction signal, a rise time of the reproduction signal, a fall time of the reproduction signal, a time distribution of the reproduction signal, an average amplitude of the sound source signal, and a generation period of the reproduction signal; The sound source generating device according to claim 1 , wherein the generating means selects the reproduction information based on the sound source information.
3. the reproduction information in which the average amplitude of the reproduction signal is different from the average amplitude of the sound source signal; the reproduction information in which the fundamental frequency of the reproduction signal is different from the fundamental frequency of the sound source signal; the reproduction information in which the frequency distribution of the reproduction signal and the frequency distribution of the sound source signal do not at least partially overlap; The reproduction information in which the rising time of the reproduction signal and the rising time of the sound source signal do not at least partially overlap, or the reproduction information in which the time distribution of the reproduction signal and the time distribution of the sound source signal do not at least partially overlap; 3. The sound source generating device according to claim 2, wherein said generating means selects:
4. 3. The sound source generating device according to claim 2, wherein said generating means divides said reproduction signal into a plurality of sections in accordance with said generation period of said sound source signal.
5. 5. The sound source generating device according to claim 4, wherein said generating means divides said reproduction signal into syllables.
6. The sensor further acquires a video signal of the user's surroundings and biological information of the user; the estimation means estimates attention information of the user based on the biometric information or the video signal; The sound source generating device according to claim 1 , wherein the generating means selects the reproduction information based on the attention information.
7. the biometric information includes gaze information of the user, a face image of the user, and a pulse rate of the user; The sound source generating device according to claim 6 , wherein the generating means changes the reproduction information or the sound source information based on the attention information.
8. the generating means selects a plurality of pieces of playback information according to the sound source information, The sound source generating device according to claim 1 , wherein the plurality of pieces of reproduction information are selected based on a user selection or a predefined correspondence between the sound source information and the reproduction information.
9. an estimation step of estimating sound source information of a sound source signal acquired by a sensor; a generating step of generating reproduction information different from the sound source information and adding the reproduction information to a reproduction signal; a reproduction step of presenting the reproduction signal to a user; The sound source generating method, wherein the sound source information is the direction of arrival of the sound source signal, and the reproduction information is the direction of arrival of the reproduction signal.
10. an estimation step of estimating sound source information of a sound source signal acquired by a sensor; a generating step of generating reproduction information different from the sound source information and adding the reproduction information to a reproduction signal; a reproduction step of presenting the reproduction signal to a user; The sound source generating program, wherein the sound source information is the arrival direction of the sound source signal, and the reproduction information is the arrival direction of the reproduction signal.
Citation Information
Patent Citations
Voice output timing controller
JP2002287783A
Sound providing system, sound providing method, program for this method, and recording medium
JP2006114942A
Navigation device, navigation device control method, program for the navigation device control method, and recoding medium with the program for navigation device control method stored thereon
JP2007333603A
Playback device, headphone, and playback method
JP2011097268A
Hearing support device and system, sound source localization device, input device, computer program, and distance detection device-integrated microphone array
JP2022122533A
Cited By
Process monitor and process monitoring method
US12424469B2