Target sound processing device, target sound processing method, and target sound processing program
The system addresses the lack of location-based sound processing by associating target sounds with location images and correcting for device characteristics to generate accurate sample sounds, replicating the actual sound experience.
Patent Information
- Application Number
- JP2022056007
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-03-30
- Publication Date
- 2025-09-02
- Estimated Expiration
- 2042-03-30
AI Technical Summary
Existing sound processing technologies fail to generate accurate preview sounds based on location information, as they do not account for the location where the target sound was collected.
A system that associates target sounds with location images, processes the sounds using device characteristics to correct for device-specific variations, and generates sample sounds that replicate the actual sound heard at a different location based on environmental conditions.
Enables the generation of accurate sample sounds that replicate the actual sound experience at a desired location, overcoming device-specific errors and environmental variations.
Smart Images

Figure 0007732938000002 
Figure 0007732938000003 
Figure 0007732938000004
Abstract
Description
[Technical Field]
[0001] The present invention relates to a target sound processing device, a target sound processing method, and a target sound processing program. [Background technology]
[0002] The reverberation and sound insulation level are usually expressed numerically, making it difficult for the average person, who is unfamiliar with such numerical values, to visualize how a sound will sound. Therefore, for example, methods are used to calculate the reverberation and sound insulation performance based on the design specifications of a building, and generate a sample sound by incorporating the predicted calculation results into a target sound, such as noise. For example, Patent Document 1 discloses a method for predicting the attenuation of environmental noise (target sound) for each propagation path into a sound receiving room, such as a direct transmission path through a partition wall, a roundabout propagation path from an opening, and a solid propagation path through a side wall, and then convolving the impulse response waveform obtained from the predicted attenuation with the source waveform of the environmental noise to generate an evaluation sound (sample sound) (see Claim 1, etc.). [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2003-156388 [Patent Document 2] Patent No. 4307622 [Patent Document 3] Patent No. 4234257 Summary of the Invention [Problem to be solved by the invention]
[0004] However, in the technology described in Patent Document 1, the target sound to be processed is collected at the location where the user wants to hear the preview sound, and then processed to generate the preview sound. However, since the technology does not have information about the location where the sound was collected, it is not possible to generate the preview sound based on the location information. [Means for solving the problem]
[0005] In order to solve the above problems, the target sound processing device according to the present invention comprises: a storage unit that stores a target sound collected by a sound collection unit at a predetermined location and a predetermined location image obtained by capturing the predetermined location in association with each other; an image receiving unit that receives a listening location image, the listening location image being an image of a location different from the predetermined location, where a user wishes to listen to a listening sound that reproduces absolute sound, which is a sound that is actually heard when the target sound is listened to under a predetermined environment at the predetermined location; a search unit that searches the storage unit for the predetermined location image similar to the received listening location image; a target sound acquisition unit that acquires a target sound associated with the searched predetermined location image; a first characteristic acquisition unit that acquires a first characteristic of the sound collection unit; a processed target sound generation unit that acquires a first processing correction value for processing the target sound based on the first characteristic, and processes the target sound using the acquired first processing correction value to generate a processed target sound; a sample sound generation unit that processes the generated processed target sound to generate the sample sound that reproduces an absolute sound that is a sound that is actually heard when the target sound is listened to in a predetermined environment at a location different from the predetermined location; a second characteristic acquisition unit that acquires a second characteristic of an output unit for outputting the generated sample sound; a processed sample sound generating unit that obtains a second processing correction value for processing the sample sound based on the second characteristic, and processes the sample sound using the obtained second processing correction value to generate a processed sample sound; an output control unit that controls the generated processed sample sound and outputs it from the output unit; Equipped with:
[0006] In order to solve the above problems, the target sound processing method according to the present invention includes: a storing step of storing a target sound collected by a sound collecting unit at a predetermined location and a predetermined location image obtained by capturing the predetermined location in association with each other; an image receiving step of receiving a listening location image of a location different from the predetermined location where a user wishes to listen to a listening sound that reproduces absolute sound, which is a sound that is actually heard when the user listens to the target sound under a predetermined environment at the predetermined location; a searching step of searching a storage unit for an image of the predetermined location similar to the received image of the listening location; a target sound acquisition step of acquiring a target sound associated with the searched predetermined location image; a first characteristic acquisition step of acquiring a first characteristic of the sound collection unit; a processed target sound generation step of acquiring a first processing correction value for processing the target sound based on the first characteristic, and processing the target sound using the acquired first processing correction value to generate a processed target sound; a sample sound generating step of processing the generated processed target sound to generate the sample sound that reproduces an absolute sound that is a sound that is actually heard when the target sound is listened to in a predetermined environment at a location different from the predetermined location; a second characteristic acquisition step of acquiring a second characteristic of an output unit for outputting the generated sample sound; a processed sample sound generating step of acquiring second processing correction values for processing the sample sound based on the second characteristic, and processing the sample sound using the acquired second processing correction values to generate a processed sample sound; an output control step of controlling the generated processed sample sound to output it from the output unit; Includes:
[0007] Furthermore, in order to solve the above problem, the target sound processing program according to the present invention comprises: a storing step of storing a target sound collected by a sound collecting unit at a predetermined location and a predetermined location image obtained by capturing the predetermined location in association with each other; an image receiving step of receiving a listening location image of a location different from the predetermined location where a user wishes to listen to a listening sound that reproduces absolute sound, which is a sound that is actually heard when the user listens to the target sound under a predetermined environment at the predetermined location; a searching step of searching a storage unit for an image of the predetermined location similar to the received image of the listening location; a target sound acquisition step of acquiring a target sound associated with the searched predetermined location image; a first characteristic acquisition step of acquiring a first characteristic of the sound collection unit; a processed target sound generation step of acquiring a first processing correction value for processing the target sound based on the first characteristic, and processing the target sound using the acquired first processing correction value to generate a processed target sound; a sample sound generating step of processing the generated processed target sound to generate the sample sound that reproduces an absolute sound that is a sound that is actually heard when the target sound is listened to in a predetermined environment at a location different from the predetermined location; a second characteristic acquisition step of acquiring a second characteristic of an output unit for outputting the generated sample sound; a processed sample sound generating step of acquiring second processing correction values for processing the sample sound based on the second characteristic, and processing the sample sound using the acquired second processing correction values to generate a processed sample sound; an output control step of controlling the generated processed sample sound to output it from the output unit; to be executed by the computer. [Effects of the Invention]
[0008] According to the present invention, the target sound collected to generate the sample sound is stored in association with information about the sound collection location, so that the sample sound can be generated based on the location information. [Brief explanation of the drawings]
[0009] [Figure 1A] 1 is a diagram for explaining an overview of a target sound processing system according to a preferred embodiment of the present invention; [Figure 1B] 1 is a sequence diagram for explaining an outline of the operation of a target sound processing system according to a preferred embodiment of the present invention; [Figure 2]1 is a diagram for explaining the configuration of a target sound processing device included in a target sound processing system according to a preferred embodiment of the present invention. FIG. [Figure 3A] 1 is a diagram showing an example of a microphone correction value table possessed by a target sound processing device included in a target sound processing system according to a preferred embodiment of the present invention; [Figure 3B] FIG. 2 is a diagram showing an example of a speaker correction value table held by a target sound processing device included in a target sound processing system according to a preferred embodiment of the present invention. [Figure 3C] FIG. 2 is a diagram showing an example of a target sound image table possessed by a target sound processing device included in a target sound processing system according to a preferred embodiment of the present invention. [Figure 4] 1 is a diagram illustrating a hardware configuration of a target sound processing device included in a target sound processing system according to a preferred embodiment of the present invention. [Figure 5A] 1 is a flowchart illustrating a processing procedure of a target sound processing device included in a target sound processing system according to a preferred embodiment of the present invention. [Figure 5B] 10 is a flowchart illustrating a processing procedure of a mobile terminal included in the target sound processing system according to a preferred embodiment of the present invention. [Figure 6] 10 is a graph showing frequency characteristics of collected sound data. [Figure 7] 10 is a graph showing the microphone's pickup level and linearity. [Figure 8] 10 is a graph showing frequency characteristics of sound reproduced through headphones. [Figure 9] 1 is a graph showing the reproducible level and linearity of headphones. [Figure 10] FIG. 1 is a diagram showing an outline of experimental conditions. [Figure 11] 10 is a graph showing the difference in sound pressure level between rooms. [Figure 12] 10 is a graph showing a sound pressure waveform on the conference room (sound source room) side. [Figure 13] 10 is a graph showing octave band levels on the conference room (sound source room) side. [Figure 14]10 is a graph showing a sound pressure waveform on the office (sound receiving room) side. [Figure 15] This is a graph showing the octave band levels on the office (receiving room) side. DETAILED DESCRIPTION OF THE INVENTION
[0010] Hereinafter, embodiments of the present invention will be described in detail by way of example with reference to the drawings. However, the configurations, numerical values, processing flows, functional elements, etc. described in the following embodiments are merely examples, and are open to modification and alteration, and are not intended to limit the technical scope of the present invention to the following description.
[0011] A target sound processing system 100 as a preferred embodiment of the present invention will be described with reference to Figures 1A to 5B. The target sound processing system 100 is used, for example, to evaluate how a target sound collected at a predetermined location sounds after undergoing performance determined by architectural specifications, etc. Figure 1A is a diagram for explaining an overview of the target sound processing system 100 according to this embodiment.
[0012] The target sound processing system 100 includes a target sound processing device 110 and a mobile terminal 120. A sound collection unit 130 (microphone) and an output unit 140 (speaker) are wired to the mobile terminal 120 using a cable from the outside of the mobile terminal 120. The sound collection unit 130 and the output unit 140 may be wirelessly connected to the mobile terminal 120, or the sound collection unit 130 and the output unit 140 may be built into the mobile terminal 120.
[0013] For example, a worker or user who owns the mobile terminal 120 collects a target sound at a predetermined location using the sound collection unit 130 (microphone) connected to the mobile terminal 120. Here, the target sound includes, for example, indoor sounds, outdoor sounds, etc., but is not limited to these. The predetermined location is, for example, a planned construction site for a detached house, an apartment building, etc., or an existing building, and is a location where stakeholders of the predetermined location or potential buyers want to know the sound environment of the predetermined location.
[0014] First, a sound collection worker at a predetermined location collects the target sound 131 at the predetermined location using the sound collection unit 130, and saves the collected target sound 131 in the mobile terminal 120. Then, the sound collection worker (or the owner of the mobile terminal 120, etc.) transmits the target sound data saved in the mobile terminal 120 to the target sound processing device 110. Note that the mobile terminal 120 and the target sound processing device 110 are connected by wireless connection. Also, the target sound processing device 110 may be a cloud server installed on the cloud, etc.
[0015] In addition, when transmitting the collected target sound 131 to the target sound processing device 110, the mobile terminal 120 also transmits a predetermined location image 160 that is an image of the predetermined location where the target sound 131 was collected, captured using an imaging device such as a camera.
[0016] The target sound processing device 110 associates the received target sound 131 with the predetermined location image 160 and stores them. That is, the target sound processing device 110 associates the target sound 131 with the predetermined location image 160 based on, for example, GPS (Global Positioning System) position data of the predetermined location, and stores them. In this way, the target sound processing device 110 stocks a pair (combination) of the target sound 131 and the predetermined location image 160, which is an image of the location where the target sound 131 was collected, in a predetermined storage or the like in preparation for future use.
[0017] Next, consider a case where, for example, a user listens to target sound 131 in a predetermined environment at a location different from the above-mentioned predetermined location and wants to listen to a sample sound that reproduces absolute pitch, which is the sound that is actually heard. Here, the location different from the above-mentioned predetermined location is a location (listening location) where the user wants to listen to the sample sound. In this case, the user first captures an image of the listening location using a camera 150 of the user's smartphone or the like, and transmits the captured listening location image 151 to the target sound processing device 110.
[0018] The target sound processing device 110 searches for a predetermined location image similar to the received listening location image from among the stored predetermined location images. Then, when the target sound processing device 110 finds a predetermined location image 160 similar to the listening location image 151, it acquires the target sound 131 associated with this predetermined location image 160. Then, the target sound processing device 110 generates a listening sound or the like for the acquired target sound 131, for example, according to the following procedure.
[0019] That is, the target sound processing device 110 processes the target sound 131 collected at a predetermined location based on the acquired target sound data of the target sound 131, to generate an evaluation sound (processed preview sound 141). When transmitting the target sound data to the target sound processing device 110, the mobile terminal 120 preferably also transmits data relating to the characteristics of the sound collection unit 130 (microphone) and output unit 140 (speaker) connected to the mobile terminal 120, but may transmit such data in response to a request from the target sound processing device 110.
[0020] This is because some frequency components of the picked-up target sound 131 are cut due to the sound collection characteristics of the sound collection unit 130, and the output sample sound has different frequency components from the sound that is actually heard due to the output characteristics of the output unit 140. For this reason, the target sound processing system 100 processes the target sound 131 while also taking into account the characteristics of the microphone and speaker to generate the sample sound (processed sample sound 141).
[0021] For example, built-in microphones and speakers built into mobile terminal 120 are smaller than external microphones and speakers, and may have limited performance or reduced functionality, in order to keep the price of mobile terminal 120 down or to conserve space within the housing of mobile terminal 120. Furthermore, even with external microphones and speakers, the available functions may be limited to certain functions, certain performance may be restricted, or conversely, certain functions may be enhanced, depending on the intended use and price. For this reason, there are many microphones and speakers with a variety of functions and performance capabilities.
[0022] In this way, if a microphone dedicated to collecting target sound 131 or a speaker dedicated to outputting a sample sound is used, there is no need to adjust for variations between microphones or speakers. However, if a dedicated microphone or speaker must be used, the dedicated microphone must be carried to a predetermined location each time target sound 131 is collected, and when the sample sound is to be listened to, the dedicated microphone must be transported to the location where the dedicated speaker is installed, making it difficult to respond flexibly and quickly.
[0023] Therefore, in the target sound processing system 100, in order to eliminate errors that depend on the characteristics of each device, at each stage of processing the picked-up target sound 131, absolute sounds (processed target sound, trial sound, processed trial sound 141) are generated that eliminate errors caused by the devices, and by processing these absolute sounds, sounds that are no different from real sounds can be reproduced so that the listener (user) can experience them.
[0024] Therefore, in the target sound processing system 100, the characteristics of the sound collection unit 130 (microphone) and the characteristics of the output unit 140 (speaker) are transmitted to the target sound processing device 110 along with the target sound data of the collected target sound 131 and the predetermined place image 160. Alternatively, the characteristics of the sound collection unit 130 and the output unit 140 may be transmitted to the target sound processing device 110 separately from the target sound 131 and the predetermined place image 160. For example, they may be transmitted before or after the transmission of the target sound 131 and the predetermined place image 160. The target sound processing device 110 processes the acquired target sound data using a correction value according to the characteristics of the sound collection unit 130 that collected the sound, thereby generating a processed target sound. Next, the target sound processing device 110 generates a trial sound from the generated processed target sound, reproducing the sound that would actually be heard when the target sound 131 is listened to at a predetermined place under a predetermined environment.
[0025] Here, the predetermined environment includes, for example, buildings such as apartment buildings or detached houses that are scheduled to be constructed at the predetermined location where the target sound 131 is collected, as well as the interiors of these buildings and their outdoor areas such as verandas. Furthermore, the predetermined environment may include various factors that realize the living environment of the building, such as the position and area of the building's walls, sound insulation performance (sound transmission loss), the position, area, and sound insulation performance (sound transmission loss) of windows attached to the building, the area and number of air intakes, sound insulation performance (normalized sound transmission loss), indoor surface area, and sound absorption performance (sound absorption power).
[0026] The target sound processing device 110 then recreates a predetermined environment (such as a building) in a virtual space, and generates a trial sound by predicting the acoustic effects in the predetermined environment using data on the picked-up target sound 131. The target sound processing device 110 processes the generated trial sound using a correction value according to the characteristics of the output unit 140 to generate a processed trial sound. By processing the trial sound using a correction value based on the characteristics of the output unit 140 in this way, it is possible to reliably reproduce the sound that is actually heard in the output unit 140 without using a dedicated speaker. Note that the trial sound may be regenerated based on the results of listening to the processed trial sound. By regenerating the trial sound in this way, trial sounds in various environments can be reproduced. Therefore, it is possible to simulate trial sounds in various environments by, for example, changing the window area or the grade of sound insulation performance.
[0027] 1B, an overview of the operation of the target sound processing system 100 will be described. In step S101, the sound collection unit 130 collects the target sound 131. In step S103, the collected target sound 131 is transmitted from the sound collection unit 130 to the mobile terminal 120. The target sound 131 collected by the sound collection unit 130 is, for example, analog data.
[0028] Then, in step S105, a camera or the like attached to the mobile terminal 120 is used to capture an image of the predetermined location where the target sound 131 was collected, thereby acquiring a predetermined location image 160. In step S107, the mobile terminal 120 transmits target sound data obtained by digitally converting the acquired target sound 131 and the predetermined location image data to the target sound processing device 110. Note that the predetermined location image 160 is previously digital data obtained by capturing an image with a digital camera or the like. In step S109, the target sound processing device 110 stores the received target sound data and predetermined location image data.
[0029] In step S111, for example, the mobile terminal 120 transmits a listening location image 151 obtained by capturing an image of a listening location, which is a location different from the predetermined location and where the user wants to listen to a listening sound, with the camera 150 to the target sound processing device 110. In step S113, the target sound processing device 110 searches for a predetermined location image 160 similar to the received listening location image 151, and acquires the target sound 131 associated with the predetermined location image 160.
[0030] In step S115, the target sound processing device 110 processes the acquired target sound 131 in accordance with predetermined conditions to generate a processed sample sound. In step S117, the target sound processing device 110 transmits the generated processed sample sound to the mobile terminal 120 (output unit 140). In step S119, the output unit 140 outputs the received processed sample sound. Here, the sound collection unit 130 and the output unit 140 may be built into the mobile terminal 120. Furthermore, digital conversion of the target sound 131 may be performed in the target sound processing device 110.
[0031] <Configuration of target sound processing device 110> The configuration of the target sound processing device 110 will be described with reference to Fig. 2. The target sound processing device 110 has a storage unit 211, an image receiving unit 212, a search unit 213, a target sound acquisition unit 214, a first characteristic acquisition unit 215, and a processed target sound generation unit 216. Furthermore, the target sound processing device 110 has a preview sound generation unit 217, a second characteristic acquisition unit 218, a processed preview sound generation unit 219, and an output control unit 220.
[0032] The storage unit 211 associates and stores target sound 131 collected by the sound collection unit 130 at a predetermined location with a predetermined location image 160 obtained by capturing an image of the predetermined location. The target sound 131 is collected, for example, by a microphone (sound collection unit 130). The predetermined location image 160 is captured by a camera or the like. The camera may be one built into the mobile terminal 120 such as a smartphone or tablet terminal, or one that exists independently as a camera. The captured predetermined location image 160 is saved as digital data in the camera's internal storage or external storage.
[0033] Furthermore, the predetermined location image 160 may be an image captured according to a predetermined imaging method. That is, when capturing an image of a predetermined location using a camera, imaging conditions such as the distance from a sound source or subject present at the predetermined location, exposure time (shutter speed), angle of view, and exposure are consistent. By consistent imaging conditions in this way, it is possible to shorten the search time by the search unit 213 (described later) and improve the search accuracy.
[0034] Furthermore, images may be taken under the same imaging conditions, such as the height of the camera lens from the ground, the angle of incidence of sunlight, etc. Note that there may be cases where it is not possible to uniform the imaging conditions due to various constraints, and in such cases, images may be taken under imaging conditions close to the predetermined imaging conditions.
[0035] The image receiving unit 212 receives a listening location image 151 captured at a location different from the predetermined location where the user wishes to listen to a sample sound that reproduces absolute pitch, which is a sound that is actually heard when listening to the target sound 131 in a predetermined environment at the predetermined location. The listening location image 151 is captured using a camera 150, for example.
[0036] Furthermore, listening location image 151 may be an image captured according to a predetermined imaging method, similar to the above-described predetermined location image 160. In this way, by unifying the imaging method between listening location image 151 and predetermined location image 160, it becomes possible to shorten the time required to search for predetermined location image 160 similar to received listening location image 151 and improve the accuracy of the search.
[0037] The search unit 213 searches the storage unit 211 for a predetermined location image 160 similar to the received listening location image 151. The search unit 213 searches the storage unit 211 for a predetermined location image 160 similar to the listening location image 151, for example, based on the degree of match between the listening location image 151 and the predetermined location image 160. The search unit 213 may, for example, extract feature points of both images and search for a predetermined location image 160 similar to the listening location image 151 based on the number of matching feature points among the extracted feature points.
[0038] Furthermore, the search unit 213 may input the predetermined location image 160, which is an image of the predetermined location, into artificial intelligence (AI) to perform machine learning. When the machine learning by the AI is completed, the search unit 213 generates a trained predetermined location image model. Note that the search unit 213 may store the generated trained predetermined location image model in a predetermined storage or the like. In this case, the stored trained predetermined location image model may be updated every time a new learning image is acquired, machine learning is performed, and a trained predetermined location image model is generated.
[0039] Machine learning using artificial intelligence is performed using a known algorithm. In machine learning, a loss function specifies weights and uses the inverse of the number of events. Furthermore, the search unit 213 inflates the number of images of a predetermined location that the artificial intelligence learns in order to improve the accuracy of the machine learning using the artificial intelligence and generate a more accurate model for similar image search. The search unit 213 obtains the inflated data by, for example, flipping the images horizontally. Furthermore, the search unit 213 may use transfer learning to improve the accuracy of the machine learning using the artificial intelligence. Here, transfer learning is a technique that aims to improve the performance of a model by reusing a trained model using a different dataset for a different problem and performing partial learning. This technique is particularly expected to improve inference performance and reduce learning time when training data is insufficient.
[0040] The target sound acquisition unit 214 acquires the target sound 131 associated with the searched predetermined place image 160. Since the storage unit 211 stores combinations of the target sound 131 and the predetermined place image 160, the target sound acquisition unit 214 acquires the target sound 131 associated with the searched predetermined place image 160 based on the searched predetermined place image 160.
[0041] The first characteristic acquisition unit 215 acquires the first characteristic of the sound collection unit 130. The first characteristic of the sound collection unit 130 is various conditions when the target sound 131 is collected, and includes, for example, mechanical characteristics such as the frequency characteristics of the sound collection unit 130 (microphone), characteristics of software incorporated in the sound collection unit 130, and environmental characteristics such as the temperature, humidity, wind direction, and wind volume of a predetermined location, but is not limited to these.
[0042] The processed target sound generation unit 216 acquires a first processing correction value for processing the target sound 131 based on the first characteristic, and processes the target sound 131 using the acquired first processing correction value to generate the processed target sound. The processed target sound generation unit 216 acquires the first processing correction value stored in, for example, internal storage or external storage. The first processing correction value is a correction value for processing the target sound 131 in accordance with the acquired first characteristic.
[0043] Then, the processed target sound generation unit 216 processes the target sound 131 (target sound data) using the acquired first processing correction value to generate the processed target sound. The processed target sound generation unit 216 processes the target sound data by, for example, canceling a specific frequency component, to generate the processed target sound.
[0044] The sample sound generation unit 217 processes the generated processed target sound to generate a sample sound that reproduces the absolute sound, which is the sound that will actually be heard when the target sound 131 is listened to in a location different from the predetermined location under a predetermined environment. Here, the predetermined environment is, for example, a building such as an apartment building or detached house that is scheduled to be constructed at the predetermined location where the target sound 131 is picked up, and includes the interior of these buildings and outdoor areas such as balconies. Furthermore, the predetermined environment may include various factors that realize the living environment of the building, such as the position and area of the building's walls, sound insulation performance (sound transmission loss), the position, area, and sound insulation performance (sound transmission loss) of windows attached to the building, the area and number of air intakes, and sound insulation performance (normalized sound transmission loss), the surface area of the room, sound absorption performance (sound absorption capacity), and sound reflection from surrounding buildings.
[0045] The generated test sounds include, for example, (1) indoor noise caused by external noise, (2) spatial sound insulation performance (noise propagating from adjacent rooms), (3) noise propagating from inside and outside the building to the site boundary, (4) indoor quieting performance caused by air conditioning equipment, (5) floor impact sound insulation performance (floor impact sound), and (6) indoor reverberation time (sound reverberation).
[0046] Specifically, (1) refers to how noise from an external noise source, such as a train, sounds indoors. (2) refers to how a room with a TV or conference sound sounds in the next room. (3) refers to how noise sources such as equipment and machinery on the building premises (inside or outside the building) and operational noise sound to the boundary or neighbors. (4) refers to how the quietness of the room changes or how the sound of the air conditioning equipment sounds indoors when the noise source is an air conditioner or total heat exchanger. (5) refers to how the sound of jumping or running from the room above sounds in the room below. (6) refers to how the sound of talking or audio equipment resonates in the room during a meeting or lecture.
[0047] The second characteristic acquisition unit 218 acquires the second characteristic of the output unit 140 for outputting the generated preview sound. The second characteristic is the output condition of the output unit 140 (speaker) for outputting the generated preview sound, and includes, for example, the mechanical characteristics of the output unit 140, the features of the software incorporated in the output unit 140, and the environmental characteristics of the output location, but is not limited to these.
[0048] The processed preview sound generation unit 219 acquires a second processing correction value for processing the preview sound based on the second characteristic, and processes the preview sound using the acquired second processing correction value to generate a processed preview sound.
[0049] When a sample sound is output from the output unit 140, the reproducibility of the sample sound depends on the characteristics of the output unit 140. Therefore, even if the same sample sound data is played back, if it is output from multiple output units 140 with different characteristics, the sound may sound different to the listener. When this happens, additional construction may be required after the building or the like is actually constructed. Therefore, in order to reproduce the sound that will actually be heard in advance, the generated sample sound is processed according to the output characteristics of the output unit 140. A processing correction value (second processing correction value) is used to process the sample sound.
[0050] The second processing correction values are determined after outputting a reference experimental sound from each of the output units 140 (speakers) and checking the output characteristics, such as the frequency characteristics, of each of the output units 140. Then, the processed sample sound generation unit 219 processes the sample sound using the second processing correction values determined as described above, and generates a processed sample sound by processing the generated sample sound in accordance with the characteristics of the output units 140.
[0051] The output control unit 220 controls the generated processed preview sound 141 to output it from the output unit 140. For example, the output control unit 220 transmits the generated processed preview sound 141 to the mobile terminal 120, and outputs the processed preview sound 141 from the output unit 140. Alternatively, the output control unit 220 may transmit the generated processed preview sound 141 directly to the output unit 140, thereby outputting the processed preview sound 141 from the output unit 140.
[0052] <Configuration of mobile terminal 120> Next, the configuration of the mobile terminal 120 will be described with reference to Fig. 2. The mobile terminal 120 has an acquisition unit 221, a transmission unit 222, a reception unit 223, and an output control unit 224. Note that the sound collection unit 130 and the output unit 140 may be built into the mobile terminal 120.
[0053] The acquisition unit 221 acquires the target sound 131 collected by the sound collection unit 130 (microphone) at a predetermined location. Here, the target sound 131 includes at least one of indoor sound and outdoor sound. The data of the acquired target sound 131 is analog data, but if the sound collection unit 130 has an AD converter, for example, the analog data of the collected target sound 131 may be converted into digital data in the sound collection unit 130. If the sound collection unit 130 does not have an AD converter, the analog data may be converted into digital data using an AD converter included in the mobile terminal 120. If neither the sound collection unit 130 nor the mobile terminal 120 has an AD converter, the analog data of the collected target sound 131 may be converted into digital data in an AD converter included in the target sound processing device 110.
[0054] The transmitting unit 222 transmits the acquired data of the target sound 131 to the target sound processing device 110. The data of the target sound 131 transmitted to the target sound processing device 110 may be analog data or digital data.
[0055] The receiving unit 223 receives the processed sample sound 141 transmitted from the target sound processing device 110. The data of the received processed sample sound 141 is digital data.
[0056] The output control unit 224 controls the received processed sample sound 141 to output it from the output unit 140. The output control unit 224 converts the received digital data of the processed sample sound 141 into analog data, sends it to the output unit 140, and controls the output unit 140 to play the analog data of the processed sample sound 141. As described above, the processed target sound may be sent directly from the output control unit 220 of the target sound processing device 110 to the output unit 140 to be output.
[0057] Next, an example of a microphone correction value table 301 included in the target sound processing device 110 will be described with reference to Fig. 3A. The microphone correction value table 301 stores characteristics 312 and correction values 313 in association with microphone IDs (Identifiers) 311. The microphone IDs 311 are identifiers for identifying microphones capable of collecting the target sound 131. The characteristics 312 indicate the characteristics of the microphone, including frequency characteristics and a sound collection level. The correction values 313 include a frequency correction value and a level correction value. The processed target sound generation unit 216 of the target sound processing device 110 then refers to the microphone correction value table 301, extracts a correction value according to the characteristics of the microphone (sound collection unit 130), and processes the received target sound 131.
[0058] 3B, an example of the speaker correction value table 302 included in the target sound processing device 110 will be described. The speaker correction value table 302 stores characteristics 322 and correction values 323 in association with speaker IDs 321. The speaker IDs 321 are identifiers for identifying speakers capable of outputting the processed sample sound 141 transmitted from the target sound processing device 110. The characteristics 322 indicate speaker characteristics, including frequency characteristics and available output levels. The correction values 323 include frequency correction values and level correction values. The processed sample sound generation unit 219 of the target sound processing device 110 then refers to the speaker correction value table 302, extracts correction values according to the characteristics of the speaker (output unit 140), and processes the generated sample sound.
[0059] 3C , an example of a target sound image table 303 held by the target sound processing device 110 will be described. The target sound image table 303 stores an image 332, a location 333, and sound collection / imaging conditions 334 in association with a target sound ID 331. The target sound ID 331 is an identifier for identifying each target sound 131 collected by the sound collection unit 130. The image 332 is an image stored in association with the collected target sound 131, and is an image of the location where the target sound 131 was collected. The location 333 is coordinate data indicating the location where the target sound 131 was collected, and multiple target sounds 131 and images may correspond to one location. The sound collection / imaging conditions 334 are the sound collection conditions when the target sound 131 was collected and the imaging conditions when an image of the sound collection location was captured, and include the distance from the subject, exposure time, weather, temperature, etc.
[0060] The hardware configuration of the target sound processing device 110 will be described with reference to FIG. 4. The CPU (Central Processing Unit) 410 is a processor for arithmetic and control, and executes programs to realize the various functional components of the target sound processing device 110 shown in FIG. 2. The CPU 410 may have multiple processors and execute different programs, modules, tasks, threads, etc. in parallel. The ROM (Read Only Memory) 420 stores fixed data such as initial data and programs, as well as other programs. The network interface 430 communicates with other devices via a network. The CPU 410 is not limited to a single CPU, and may include multiple CPUs or a GPU (Graphics Processing Unit) for image processing. The network interface 430 preferably has a CPU independent of the VPU 410 and writes and reads transmitted and received data to and from an area in the RAM (Random Access Memory) 440. It is also preferable to provide a DMAC (Direct Memory Access Controller) (not shown) for transferring data between the RAM 440 and the storage 450. The CPU 410 recognizes that data has been received or transferred to the RAM 440 and processes the data accordingly. The CPU 410 also prepares the processing results in the RAM 440, and leaves the subsequent transmission or transfer to the network interface 430 or DMAC.
[0061] The RAM 440 is a random access memory used by the CPU 410 as a work area for temporary storage. A storage area for storing data necessary for implementing this embodiment is secured in the RAM 440. The target sound data 441 is data of a target sound collected at a predetermined location using the sound collection unit 130, and is digitally converted data. The microphone correction value 442 is a correction value used to calibrate the microphone sensitivity (the difference between the actual sound and the collected sound) and the like that depend on the sound collection unit 130 (microphone) used to collect the target sound 131.
[0062] The speaker correction value 443 is a correction value used to eliminate speaker-specific acoustic characteristics (difference between the generated sample sound and the reproduced sample sound) that depend on the output unit 140 (speaker) that reproduces the generated processed sample sound. The processed target sound data 444 is sound data obtained by processing the collected target sound 131 using the first processing correction value.
[0063] The trial sound data 445 is data on the trial sound obtained by processing the processed target sound, and reproduces the sound that is actually heard when the target sound 131 is listened to at a predetermined location under a predetermined environment. The processed trial sound data 446 is trial sound that has been processed using the second processing correction value based on the output conditions of the output unit 140 for outputting the generated trial sound. The image data 447 is data on an image captured at the location where the target sound 131 was picked up.
[0064] The transmitted / received data 448 is data transmitted and received via the network interface 430. The RAM 440 also has an application execution area 449 for executing various application modules.
[0065] The storage 450 stores a database, various parameters, or the following data or programs required to implement this embodiment. The storage 450 stores a microphone correction value table 301, a speaker correction value table 302, and a target sound image table 303. The microphone correction value table 301 is a table that manages the relationship between the microphone ID 311 and the correction value 313, etc., shown in FIG. 3A. The speaker correction value table 302 is a table that manages the relationship between the speaker ID 321 and the correction value 323, etc., shown in FIG. 3B. The target sound image table 303 is a table that manages the relationship between the target sound ID 331 and the image 332, etc., shown in FIG. 3C.
[0066] The storage 450 further stores an image reception module 451, a search module 452, a target sound acquisition module 453, a first characteristic acquisition module 454, a processed target sound generation module 455, a preview sound generation module 456, a second characteristic acquisition module 457, a processed target sound generation module 458, and an output control module 459.
[0067] The image receiving module 451 is a module that captures an image of a location different from the predetermined location where a sample sound that reproduces absolute sound, which is the sound that is actually heard when the target sound 131 is listened to in a predetermined environment at the predetermined location, and receives a listening location image 151. The search module 452 is a module that searches the storage unit 211 for a predetermined location image 160 that is similar to the received listening location image 151. The target sound acquisition module 453 is a module that acquires the target sound 131 associated with the searched predetermined location image 160.
[0068] The first characteristic acquisition module 454 is a module that acquires the first characteristic of the sound collection unit 130. The processed target sound generation module 455 is a module that acquires a first processing correction value for processing the target sound 131 based on the first characteristic, and processes the target sound 131 using the acquired first processing correction value to generate the processed target sound. The preview sound generation module 456 is a module that processes the generated processed target sound to generate a preview sound.
[0069] The second characteristic acquisition module 457 is a module that acquires the second characteristic of the output unit 140 for outputting the generated preview sound. The processed target sound generation module 458 is a module that acquires a second processing correction value for processing the preview sound based on the second characteristic, and processes the preview sound using the acquired second processing correction value to generate a processed preview sound. The output control module 459 is a module that controls the generated processed preview sound and outputs it from the output unit 140.
[0070] These modules 451 to 459 are read into an application execution area 449 of the RAM 440 and executed by the CPU 410. The control program 470 is a program for controlling the target sound processing device 110 as a whole.
[0071] The input / output interface 460 interfaces input / output data with input / output devices. A display unit 461 and an operation unit 462 are connected to the input / output interface 460. A storage medium 464 may also be connected to the input / output interface 460. A speaker 463 serving as an audio output unit, a microphone (not shown) serving as an audio input unit, or a GPS position determination unit may also be connected. Note that the RAM 440 and storage 450 shown in FIG. 4 do not include programs or data related to the general-purpose functions of the target sound processing device 110 or other feasible functions.
[0072] Next, a processing procedure of the target sound processing device 110 will be described with reference to the flowchart shown in Fig. 5A. This flowchart is executed by the CPU 410 in Fig. 4 using the RAM 440, and realizes each functional configuration of the target sound processing device 110 in Fig. 2.
[0073] In step S501, the storage unit 211 associates and stores the target sound 131 picked up at a predetermined location with a predetermined location image capturing the predetermined location where the target sound 131 was picked up. In step S503, the image receiving unit 212 receives a listening location image 151 capturing an image of a location where the sample sound was listened to, which is a location different from the predetermined location where the target sound 131 was picked up.
[0074] In step S505, the search unit 213 searches the storage unit 211 for a predetermined location image similar to the received listening location image 151. The search is performed, for example, based on the degree of match between the predetermined location image 160 and the listening location image 151. In step S507, the search unit 213 determines whether or not a predetermined location image 160 similar to the listening location image 151 has been found. If it is determined that the search has not been successful (NO in step S507), the target sound processing device 110 ends the processing. If it is determined that the search has been successful (YES in step S507), the target sound processing device 110 proceeds to the next step.
[0075] In step S509, the target sound acquisition unit 214 acquires the target sound 131 associated with the searched predetermined location image 160. In step S511, the first characteristic acquisition unit 215 acquires the first characteristic of the sound collection unit 130. In step S513, the processed target sound generation unit 216 acquires a first processing correction value for processing the target sound 131 based on the acquired first characteristic, and processes the target sound 131 using the acquired first processing correction value to generate the processed target sound.
[0076] In step S515, the preview sound generation unit 217 processes the generated processed target sound to generate a preview sound that reproduces the absolute pitch, which is the sound that would actually be heard when the target sound 131 is listened to in a predetermined environment at a location different from the predetermined location. In step S517, the second characteristic acquisition unit 218 acquires the second characteristic of the output unit 140 for outputting the generated preview sound. In step S519, the processed preview sound generation unit 219 acquires a second processing correction value for processing the preview sound based on the acquired second characteristic, and processes the preview sound using the acquired second processing correction value to generate the processed preview sound. In step S521, the generated processed preview sound is controlled to be output from the output unit 140.
[0077] Next, a processing procedure of mobile terminal 120 will be described with reference to the flowchart shown in Fig. 5B. This flowchart is executed by a CPU (not shown) using a RAM (not shown), and realizes each functional configuration of mobile terminal 120 shown in Fig. 2.
[0078] In step S531, the mobile terminal 120 acquires the target sound 131 collected by the sound collection unit 130. Here, the mobile terminal 120 may digitally convert the acquired target sound 131. In step S533, the mobile terminal 120 transmits data of the acquired target sound 131 to the target sound processing device 110. In step S535, the mobile terminal 120 receives the processed sample sound 141 from the target sound processing device 110. The received processed sample sound 141 is output from the output unit 140.
[0079] According to this embodiment, the target sound and the image of the predetermined location are stored in association with each other, so that by capturing an image of a new location, a sample sound can be generated without having to record the sound at the new location. In other words, since there is no need to carry a dedicated sound recording microphone to the new location to record the sound, it can be used easily.
[0080] [Example] Next, examples of the present invention will be described with reference to Figures 6 to 15. It should be noted that the scope of the present invention is not limited to the following examples.
[0081] The target sound processing device 110 of the target sound processing system 100 generates test sounds for the six items ((a) to (f)) shown in Table 1. These are: (a) indoor noise caused by external noise, (b) spatial shielding performance (noise propagating from adjacent rooms), (c) noise propagating from inside and outside the building to the boundary of the site, (d) indoor noise caused by spatial equipment, (e) floor impact sound insulation performance (floor impact sound), and (f) indoor reverberation time (sound reverberation).
[0082] [Table 1]
[0083] Specific examples of each of the six items include (a) traffic noise (road traffic noise, railway noise, etc.), (b) airborne noise such as conversations in a conference room and hotel television sounds, (c) outdoor equipment and machinery, indoor equipment and machinery, (d) indoor air conditioning equipment, (e) floor impact noise (heavy floor impact noise and light floor impact noise), (f) conversations in a conference room, voices in a classroom, etc. For each calculation, the technologies described in Patent Documents 1 to 3, for example, are used.
[0084] The target frequencies for each of the six items are: (a) 50Hz to 5,000Hz band, (b) 100Hz to 5,000Hz band, (c) and (d) 50Hz to 5,000Hz band, (e) heavy floor impact sound: 50Hz to 630Hz band, light floor impact sound: 50Hz to 5,000Hz band, and (f) 100Hz to 5,000Hz band.
[0085] Regarding the generation of the sample sound, in (a) to (d), a volume reduction filter is generated and the sample sound is generated by filtering the sound source data (collected target sound data).
[0086] In (e), the floor impact sound level is calculated from the impact force of a standard impact source (tire, ball, tapping) according to the JIS standard, and sound source data close to the calculated conditions is extracted from the floor impact sound from the standard impact source recorded for each condition such as slab thickness, floor finish structure, and finished ceiling, and a filter is generated to reduce the volume by the difference between that floor impact sound level and the calculated value, and a trial sound is generated. Note that the heavy floor impact sound is evaluated in the 50Hz to 630Hz band, so frequencies outside the target range are not filtered and are not played.
[0087] In (f), a test sound is generated by convolving the dry source sound, such as a reading sound recorded in an anechoic chamber, with the impulse response predicted by acoustic simulation or the actual measured value of the impulse response.
[0088] Next, the microphone (sound collection unit 130 or sound collection device) and headphones (output unit 140 or speaker) used each have their own unique acoustic characteristics. Therefore, in order to faithfully reproduce the generated sample sound, these acoustic characteristics must be corrected, and the collected sound source (target sound 131) and the generated sample sound are corrected based on a correction value calculated in advance. The correction value for the acoustic characteristics of the equipment used is calculated using the following method.
[0089] <<How to correct the acoustic characteristics of a microphone>> <Microphone frequency characteristics> A method for correcting the acoustic characteristics of a microphone (sound pickup unit 130) will be described using an example in which a tablet terminal is used as the mobile terminal 120. First, it is experimentally confirmed that the microphone can pick up the sound of the frequency to be evaluated. In an anechoic chamber, pink noise (noise) is generated from a speaker (sound source), and a precision sound level meter is placed 1 m away to pick up the sound. Figure 6 shows the 1 / 3 octave band levels of the picked-up sound data.
[0090] The microphone attached to the tablet device and the precision sound level meter generally matched in the target frequency range of 50 Hz to 5,000 Hz. On the other hand, the microphone built into the tablet device showed level differences at low frequencies below 80 Hz. When capturing problematic low-frequency sounds such as heavy floor impact noise and equipment noise, it is preferable to use an external microphone.
[0091] <Microphone pickup level and linearity> When listening to the test sound, the sound pressure level to be picked up is assumed to be 30dB to 80dB, taking into consideration the user's hearing impairment and the level of noise that requires noise control measures when using the system. In addition, the frequency characteristics of the microphone used will change linearly depending on the sound pressure level to be picked up. Therefore, the linearity of the pickup level and frequency characteristics of the microphone to be used will be experimentally confirmed.
[0092] Pink noise was recorded at a volume of 30 dB to 80 dB using the tablet's built-in microphone and a precision sound level meter. The recorded sound was analyzed in 1 / 3 octave bands to examine the correspondence between the tablet's built-in microphone and the precision sound level meter. Figure 7 shows the results for the 1,000 Hz band as a representative example.
[0093] If the value of the precision sound level meter is taken as the true value, when using the built-in microphone of a tablet device, the change is roughly linear from 30dB to 80dB. Because the change is linear within the sound pressure level range expected for system use, correction can be made using a fixed correction value for each frequency band.
[0094] After confirming that the target microphone can be used with the target sound processing system 100, the 1 / 3 octave band level difference between the precision sound level meter and the target microphone is used as a correction value (first processing correction value) to ensure accurate sound pickup.
[0095] <<How to correct the acoustic characteristics of headphones>> <Headphone frequency characteristics> Many headphones are tuned according to purpose and preference, and to faithfully reproduce the generated listening sound, the frequency characteristics of each headphone must be canceled out. For this reason, headphones were attached to a dummy head in an anechoic chamber, and pink noise (sound source) was played via a tablet device. Figure 8 shows the 1 / 3 octave band levels of the pink noise and the sound reproduced by the headphones. Compared to the pink noise, which has a flat frequency response, the sound reproduced by the headphones used varies by frequency, with sounds in the 400 Hz to 1,000 Hz band being particularly emphasized.
[0096] <Headphone playback level and linearity> Pink noise is played through headphones at different levels to examine the reproducible level and linearity of the headphones. Figure 9 shows the 1 / 3 octave band level of the 1,000 Hz band as a typical example.
[0097] The headphones were able to reproduce sounds in the range of 20dB to 80dB, with a linear change in level in 5dB increments of pink noise. Similar results were obtained in other frequency bands. Because reproduction up to 20dB is possible, when considering spatial sound insulation, for example, the spatial sound insulation performance of up to Dr-55 can be expressed as a listening test sound for a sound source generated at 75dB in the sound source room. Here, Dr is the rating of the difference in sound pressure level between rooms, and the higher the Dr value, the higher the sound insulation performance of the space and the less airborne sound is transmitted.
[0098] After confirming that the target headphones can be used with the target sound processing system 100, the 1 / 3 octave band level difference between the pink noise (sound source) and the sound reproduced by the headphones is used as a correction value (second processing correction value) to cancel the acoustic characteristics unique to the headphones so that the generated trial sound can be faithfully reproduced.
[0099] <<Verification example of system accuracy in an actual building>> <Verification Overview> Using an office room, we verified the accuracy of the predicted calculations, sample sounds, and processed sample sounds generated by the target sound processing device 110 regarding spatial sound insulation performance. An overview of the measurement room is shown in Figure 10.
[0100] The partition between the conference room and the office is a dry double wall (sound insulation performance TL D -40). The corridor doors between the conference room and the office are standard steel double doors without airtight seals. The conference room was used as the sound source room, and target sounds (pink noise or a man reading aloud) were played from the speakers. The played target sounds were picked up by the target sound processing device 110 (using the built-in microphone of a tablet device), and the picked-up target sounds were used as sound source data. To verify the sound pickup accuracy, sound was also picked up by a precision sound level meter (NA-28 manufactured by RION). A dummy head was installed in the sound receiving room to pick up the sound transmitted from the conference room. The difference in sound pressure levels between rooms was also measured in accordance with JIS A1418:2000 "Method for measuring the airborne sound insulation performance of buildings."
[0101] <Prediction calculation accuracy> To confirm the accuracy of prediction calculations by the target sound processing system 100, the predicted values of the sound pressure level difference between rooms were compared with the actual measured values. Figure 11 shows the comparison results for octave band levels. The predicted values roughly correspond to the actual measured values in the 125 Hz to 4,000 Hz band, and the sound pressure level difference between rooms was predicted with good accuracy.
[0102] Because sound collection accuracy affects the accuracy of generating the sample sound, we compared the sound collection data from the target sound processing system 100 (tablet built-in microphone) with the sound collection data from a precision sound level meter. The sound pressure waveform from the precision sound level meter and the sound pressure waveform (processed target sound) collected by the target sound processing system 100 (tablet built-in microphone) are shown in Figure 12, and the octave band levels of each sound pressure waveform are shown in Figure 13.
[0103] The shape and amplitude of the sound pressure waveforms picked up by the target sound processing system 100 (tablet built-in microphone) are the same as those picked up by a precision sound level meter for both pink noise and a male reading voice. At the octave band level, the error is a maximum of about 1 dB, and the sound is picked up with the same precision as a normal sound level meter.
[0104] <Preview sound accuracy> A correction process was performed to cancel the acoustic characteristics specific to headphones on the sound pressure waveform of the sample sound, which was obtained by taking into account the predicted calculation results for the sound pickup data of the target sound processing system 100, and a processed sample sound was generated.
[0105] The sound pressure waveform of the test sound, the sound pressure waveform of the processed test sound (sound played back through headphones), and the sound pressure waveform (actual measured values) collected using a dummy head in the sound receiving room office are shown in Figure 14, and their respective octave band levels are shown in Figure 15. Note that while the background noise in the sound receiving room is shown in Figure 15, it was excluded from the evaluation when playing back a male reading, as it was affected by background noise in the sound receiving room above the 1,000 Hz band.
[0106] The processed sample sound (sound played back through headphones) has almost the same shape and amplitude as the sample sound and the actual measured sound pressure waveform, and the octave band levels are also about the same.
[0107] The sample sound generated by the target sound processing device 110 of the system can be correctly reproduced through headphones, and the processed sample sound can reproduce the actual sound.
[0108] Although the present invention has been described above with reference to the embodiments, the present invention is not limited to the above-described embodiments and can be modified as appropriate. Various modifications that can be understood by those skilled in the art can be made to the configuration and details of the present invention within the scope of the present invention. Furthermore, outer containers and packaging boxes that combine the separate features included in each embodiment in any way are also included in the scope of the present invention.
[0109] The present invention may also be applied to a system consisting of multiple devices or to a single device. Furthermore, the present invention may also be applied when an information processing program that realizes the functions of the embodiments is supplied to a system or device and executed by a built-in processor. Therefore, the technical scope of the present invention also includes a program installed on a computer to realize the functions of the present invention, a medium storing the program, a WWW (World Wide Web) server from which the program is downloaded, and a processor that executes the program. In particular, the technical scope of the present invention also includes a non-transitory computer-readable medium storing a program that causes a computer to execute at least the processing steps included in the above-described embodiments.
Claims
1. a storage unit that stores a target sound collected by a sound collection unit at a predetermined location and a predetermined location image obtained by capturing the predetermined location in association with each other; an image receiving unit that receives a listening location image, the listening location image being an image of a location different from the predetermined location, where a user wishes to listen to a listening sound that reproduces absolute sound, which is a sound that is actually heard when the target sound is listened to under a predetermined environment at the predetermined location; a search unit that searches the storage unit for the predetermined location image similar to the received listening location image; a target sound acquisition unit that acquires a target sound associated with the searched predetermined location image; a first characteristic acquisition unit that acquires a first characteristic of the sound collection unit; a processed target sound generation unit that acquires a first processing correction value for processing the target sound based on the first characteristic, and processes the target sound using the acquired first processing correction value to generate a processed target sound; a sample sound generation unit that processes the generated processed target sound to generate the sample sound that reproduces an absolute sound that is a sound that is actually heard when the target sound is listened to in a predetermined environment at a location different from the predetermined location; a second characteristic acquisition unit that acquires a second characteristic of an output unit for outputting the generated sample sound; a processed sample sound generating unit that obtains a second processing correction value for processing the sample sound based on the second characteristic, and processes the sample sound using the obtained second processing correction value to generate a processed sample sound; an output control unit that controls the generated processed sample sound and outputs it from the output unit; A target sound processing device comprising:
2. A model generation unit that generates a learned predetermined location image model by training an artificial intelligence to learn the predetermined location image, The target sound processing device according to claim 1 , wherein the search unit searches the storage unit for the predetermined location image similar to the listening location image by using the learned predetermined location image model.
3. 3. The target sound processing device according to claim 1, wherein the search unit searches the storage unit for the predetermined location image similar to the listening location image based on a degree of coincidence between the listening location image and the predetermined location image.
4. The target sound processing device according to any one of claims 1 to 3, wherein the processed sample sound includes at least one of indoor noise caused by external noise, noise propagating from an adjacent room, noise propagating from inside or outside the building to the site boundary, indoor noise caused by air conditioning equipment, floor impact sound insulation performance, and indoor reverberation time.
5. The target sound processing device according to claim 2 , wherein the model generation unit generates the trained predetermined location image model using left-right inversion as padded data.
6. The target sound processing device according to claim 2 , wherein the model generation unit generates the trained predetermined location image model using transfer learning.
7. 7. The target sound processing device according to claim 1, wherein the listening location image is an image captured according to a predetermined imaging method.
8. a storing step of storing a target sound collected by a sound collecting unit at a predetermined location and a predetermined location image obtained by capturing the predetermined location in association with each other; an image receiving step of receiving a listening location image of a location different from the predetermined location where a user wishes to listen to a listening sound that reproduces absolute sound, which is a sound that is actually heard when the user listens to the target sound under a predetermined environment at the predetermined location; a searching step of searching a storage unit for an image of the predetermined location similar to the received image of the listening location; a target sound acquisition step of acquiring a target sound associated with the searched predetermined location image; a first characteristic acquisition step of acquiring a first characteristic of the sound collection unit; a processed target sound generation step of acquiring a first processing correction value for processing the target sound based on the first characteristic, and processing the target sound using the acquired first processing correction value to generate a processed target sound; a sample sound generating step of processing the generated processed target sound to generate the sample sound that reproduces an absolute sound that is a sound that is actually heard when the target sound is listened to in a predetermined environment at a location different from the predetermined location; a second characteristic acquisition step of acquiring a second characteristic of an output unit for outputting the generated sample sound; a processed sample sound generating step of obtaining a second processing correction value for processing the sample sound based on the second characteristic, and processing the sample sound using the obtained second processing correction value to generate a processed sample sound; an output control step of controlling the generated processed sample sound to output it from the output unit; A target sound processing method including:
9. a storing step of storing a target sound collected by a sound collecting unit at a predetermined location and a predetermined location image obtained by capturing the predetermined location in association with each other; an image receiving step of receiving a listening location image of a location different from the predetermined location where a user wishes to listen to a listening sound that reproduces absolute sound, which is a sound that is actually heard when the user listens to the target sound under a predetermined environment at the predetermined location; a searching step of searching a storage unit for an image of the predetermined location similar to the received image of the listening location; a target sound acquisition step of acquiring a target sound associated with the searched predetermined location image; a first characteristic acquisition step of acquiring a first characteristic of the sound collection unit; a processed target sound generation step of acquiring a first processing correction value for processing the target sound based on the first characteristic, and processing the target sound using the acquired first processing correction value to generate a processed target sound; a sample sound generating step of processing the generated processed target sound to generate the sample sound that reproduces an absolute sound that is a sound that is actually heard when the target sound is listened to in a predetermined environment at a location different from the predetermined location; a second characteristic acquisition step of acquiring a second characteristic of an output unit for outputting the generated sample sound; a processed sample sound generating step of obtaining a second processing correction value for processing the sample sound based on the second characteristic, and processing the sample sound using the obtained second processing correction value to generate a processed sample sound; an output control step of controlling the generated processed sample sound to output it from the output unit; A target sound processing program that causes a computer to execute the above.
Citation Information
Patent Citations
Integral controller for video image and audio signal
JP1995131770A
Sensing method for noise environment, trial listening apparatus and information storage medium
JP2003156388A
Audio signal supplying apparatus, parameter providing system, television set, AV system, speaker device and audio signal supplying method
JP2009130643A
Sound signal control system and sound signal control method
JP2013197764A
Noise source search system
JP2014044083A