Target sound processing device, target sound processing method, and target sound processing program

The target sound processing system addresses the issue of inaccurate sound preview by using location-specific processing correction values to generate and output accurate sample sounds, enhancing flexibility and accuracy in sound reproduction.

JP7733605B2Active Publication Date: 2025-09-03OKUMURA CORP
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2022056008
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-03-30
Publication Date
2025-09-03
Estimated Expiration
2042-03-30

AI Technical Summary

Technical Problem

Existing sound processing technologies fail to generate accurate preview sounds based on location information due to the lack of data about the sound collection location, leading to inaccuracies in reproducing the target sound.

Method used

A target sound processing system that includes a sound pickup unit, processing correction values based on sound collection and output unit characteristics, and a storage unit associating sounds with location images, enabling the generation and output of processed sample sounds based on user-defined listening locations.

Benefits of technology

Enables the reproduction of accurate sample sounds in various environments by correcting for device characteristics and associating sounds with location information, allowing quick and flexible preview of sound environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007733605000002
    Figure 0007733605000002
  • Figure 0007733605000003
    Figure 0007733605000003
  • Figure 0007733605000004
    Figure 0007733605000004
Patent Text Reader

Abstract

To listen to a trial sound on the basis of information about a place.SOLUTION: An object sound processing apparatus comprises: a storage unit which stores a processed trial sound obtained by processing a trial sound and a prescribed place image obtained by imaging a prescribed place in association with each other by using a trial sound that reproduces an absolute sound being the sound that is actually heard and a second correction value for processing for processing the trial sound acquired on the basis of the second characteristic of an output unit for outputting the trial sound in a case where the object sound is heard under a prescribed environment; an image reception unit which receives a trial listening place image obtained by imaging a place where a user wants to trial-listen the processed trial sound; a search unit which searches for the prescribed place image similar to the received trial listening place image from the storage unit; an acquisition unit which acquires the processed trial sound associated with the searched prescribed place image from the storage unit; and an output control unit which controls the acquired processed trial sound to output it from the output unit.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a target sound processing device, a target sound processing method, and a target sound processing program. [Background technology]

[0002] The reverberation and sound insulation level are usually expressed numerically, making it difficult for the average person, who is unfamiliar with such numerical values, to visualize how a sound will sound. Therefore, for example, methods are used to calculate the reverberation and sound insulation performance based on the design specifications of a building, and generate a sample sound by incorporating the predicted calculation results into a target sound, such as noise. For example, Patent Document 1 discloses a method for predicting the attenuation of environmental noise (target sound) for each propagation path into a sound receiving room, such as a direct transmission path through a partition wall, a roundabout propagation path from an opening, and a solid propagation path through a side wall, and then convolving the impulse response waveform obtained from the predicted attenuation with the source waveform of the environmental noise to generate an evaluation sound (sample sound) (see Claim 1, etc.). [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2003-156388 [Patent Document 2] Patent No. 4307622 [Patent Document 3] Patent No. 4234257 Summary of the Invention [Problem to be solved by the invention]

[0004] However, in the technology described in Patent Document 1, the target sound to be processed is collected at the location where the user wants to hear the preview sound, and then processed to generate the preview sound. However, since the technology does not have information about the location where the sound was collected, it is not possible to generate the preview sound based on the location information. [Means for solving the problem]

[0005] In order to solve the above problems, the target sound processing device according to the present invention comprises: Target sound picked up by a sound pickup unit at a predetermined location, The target sound collection conditions, a processed target sound obtained by processing the target sound using a first processing correction value for processing the target sound acquired based on the first characteristic of the sound collection unit; A sample sound that reproduces an absolute pitch that is a sound that is actually heard when the target sound is listened to in a predetermined environment at the predetermined location; a processed sample sound obtained by processing the sample sound using a second processing correction value for processing the sample sound, the second processing correction value being acquired based on a second characteristic of an output unit for outputting the sample sound; and a storage unit that stores an image of the predetermined location in association with the predetermined location; an image receiving unit that receives a listening location image obtained by capturing an image of a location where the user wishes to listen to the processed listening location sound; a search unit that searches the storage unit for the predetermined location image similar to the received listening location image; an acquisition unit that acquires, from the storage unit, a processed sample sound associated with the searched predetermined location image; an output control unit that controls the acquired processed sample sound and outputs it from the output unit; Equipped with.

[0006] In order to solve the above problems, the target sound processing method according to the present invention includes: Target sound picked up by a sound pickup unit at a predetermined location, The target sound collection conditions, a processed target sound obtained by processing the target sound using a first processing correction value for processing the target sound acquired based on the first characteristic of the sound collection unit; A sample sound that reproduces an absolute pitch that is a sound that is actually heard when the target sound is listened to in a predetermined environment at the predetermined location; a processed sample sound obtained by processing the sample sound using a second processing correction value for processing the sample sound, the second processing correction value being acquired based on a second characteristic of an output unit for outputting the sample sound; and a storage step of storing an image of the predetermined location in association with the image; an image receiving step of receiving a listening location image of a location where the user wishes to listen to the processed listening location sound; a searching step of searching a storage unit for an image of the predetermined location similar to the received image of the listening location; an acquisition step of acquiring, from the storage unit, a processed sample sound associated with the searched predetermined location image; an output control unit that controls the acquired processed sample sound and outputs it from the output unit; Includes:

[0007] Furthermore, in order to solve the above problem, the target sound processing program according to the present invention comprises: Target sound picked up by a sound pickup unit at a predetermined location, The target sound collection conditions, a processed target sound obtained by processing the target sound using a first processing correction value for processing the target sound acquired based on the first characteristic of the sound collection unit; A sample sound that reproduces an absolute pitch that is a sound that is actually heard when the target sound is listened to in a predetermined environment at the predetermined location; a processed sample sound obtained by processing the sample sound using a second processing correction value for processing the sample sound, the second processing correction value being acquired based on a second characteristic of an output unit for outputting the sample sound; and a storage step of storing an image of the predetermined location in association with the image; an image receiving step of receiving a listening location image of a location where the user wishes to listen to the processed listening location sound; a searching step of searching a storage unit for an image of the predetermined location similar to the received image of the listening location; an acquisition step of acquiring, from the storage unit, a processed sample sound associated with the searched predetermined location image; an output control unit that controls the acquired processed sample sound and outputs it from the output unit; to be executed by the computer. [Effects of the Invention]

[0008] According to the present invention, a sample sound generated by processing a target sound is stored in association with information about the location where the target sound was collected, so that the sample sound can be listened to based on the location information. [Brief explanation of the drawings]

[0009] [Figure 1A] 1 is a diagram for explaining an overview of a target sound processing system according to a first embodiment of the present invention; [Figure 1B] 1 is a sequence diagram for explaining an outline of the operation of a target sound processing system according to a preferred embodiment of the present invention; [Figure 2] 1 is a block diagram illustrating a configuration of a target sound processing system according to a first embodiment of the present invention. [Figure 3A] 3 is a diagram showing an example of a microphone correction value table included in the target sound processing device of the target sound processing system according to the first embodiment of the present invention. FIG. [Figure 3B] 3 is a diagram showing an example of a speaker correction value table included in the target sound processing device of the target sound processing system according to the first embodiment of the present invention. FIG. [Figure 3C] 1 is a diagram showing an example of a target sound image table included in a target sound processing device of a target sound processing system according to a first embodiment of the present invention. FIG. [Figure 3D] 2 is a diagram showing an example of a stored data table included in a target sound processing device of the target sound processing system according to the first embodiment of the present invention. FIG. [Figure 4] 1 is a diagram illustrating a hardware configuration of a target sound processing device of a target sound processing system according to a first embodiment of the present invention. [Figure 5A] 3 is a flowchart illustrating a processing procedure of a target sound processing device of the target sound processing system according to the first embodiment of the present invention. [Figure 5B] 5 is a flowchart for explaining a processing procedure of a mobile terminal of the target sound processing system according to the first embodiment of the present invention. [Figure 6]FIG. 10 is a block diagram illustrating the configuration of a target sound processing device of a target sound processing system according to a second embodiment of the present invention. [Figure 7] FIG. 10 is a diagram for explaining an example of a correction amount table included in the target sound processing device of the target sound processing system according to the second embodiment of the present invention. [Figure 8] FIG. 10 is a diagram illustrating a hardware configuration of a target sound processing device of a target sound processing system according to a second embodiment of the present invention. [Figure 9] 10 is a flowchart illustrating a processing procedure of a target sound processing device of a target sound processing system according to a second embodiment of the present invention. [Figure 10] 10 is a graph showing frequency characteristics of collected sound data. [Figure 11] 10 is a graph showing the microphone's pickup level and linearity. [Figure 12] 10 is a graph showing frequency characteristics of sound reproduced through headphones. [Figure 13] 1 is a graph showing the reproducible level and linearity of headphones. [Figure 14] FIG. 1 is a diagram showing an outline of experimental conditions. [Figure 15] 10 is a graph showing the difference in sound pressure level between rooms. [Figure 16] 10 is a graph showing a sound pressure waveform on the conference room (sound source room) side. [Figure 17] 10 is a graph showing octave band levels on the conference room (sound source room) side. [Figure 18] 10 is a graph showing a sound pressure waveform on the office (sound receiving room) side. [Figure 19] This is a graph showing the octave band levels on the office (receiving room) side. DETAILED DESCRIPTION OF THE INVENTION

[0010] Hereinafter, embodiments of the present invention will be described in detail by way of example with reference to the drawings. However, the configurations, numerical values, processing flows, functional elements, etc. described in the following embodiments are merely examples, and are open to modification and alteration, and are not intended to limit the technical scope of the present invention to the following description.

[0011] [First embodiment] A target sound processing system 100 according to a first embodiment of the present invention will be described with reference to Figures 1A to 5B. The target sound processing system 100 is used, for example, to evaluate how a target sound collected at a predetermined location sounds after undergoing performance determined by architectural specifications, etc. Figure 1A is a diagram for explaining an overview of the target sound processing system 100 according to this embodiment.

[0012] The target sound processing system 100 includes a target sound processing device 110 and a mobile terminal 120. A sound collection unit 130 (microphone) and an output unit 140 (speaker) are wired to the mobile terminal 120 using a cable from the outside of the mobile terminal 120. The sound collection unit 130 and the output unit 140 may be wirelessly connected to the mobile terminal 120, or the sound collection unit 130 and the output unit 140 may be built into the mobile terminal 120.

[0013] For example, a worker or user who owns the mobile terminal 120 collects a target sound at a predetermined location using the sound collection unit 130 (microphone) connected to the mobile terminal 120. Here, the target sound includes, for example, indoor sounds, outdoor sounds, etc., but is not limited to these. The predetermined location is, for example, a planned construction site for a detached house, an apartment building, etc., or an existing building, and is a location where stakeholders of the predetermined location or potential buyers want to know the sound environment of the predetermined location.

[0014] First, a sound collection worker at a predetermined location collects the target sound 131 at the predetermined location using the sound collection unit 130, and saves the collected target sound 131 in the mobile terminal 120. Then, the sound collection worker (or the owner of the mobile terminal 120, etc.) transmits the target sound data saved in the mobile terminal 120 to the target sound processing device 110. Note that the mobile terminal 120 and the target sound processing device 110 are connected by wireless connection. Also, the target sound processing device 110 may be a cloud server installed on the cloud, etc.

[0015] Then, the target sound processing device 110 processes the target sound 131 collected at a predetermined location based on the acquired target sound data of the target sound 131 to generate a processed target sound, generates a preview sound from the generated processed target sound, and processes the generated preview sound to generate a processed preview sound 141. Then, the target sound processing device 110 generates the processed target sound, the preview sound, and the processed preview sound 141, for example, according to the following procedure.

[0016] That is, the target sound processing device 110 processes the target sound 131 collected at a predetermined location based on the acquired target sound data of the target sound 131, to generate processed sample sound 141. When transmitting the target sound data to the target sound processing device 110, the mobile terminal 120 preferably also transmits data relating to the characteristics of the sound collection unit 130 (microphone) and output unit 140 (speaker) connected to the mobile terminal 120, but may transmit such data in response to a request from the target sound processing device 110.

[0017] This is because the sound collection characteristics of the sound collection unit 130 cut off some frequency components of the collected target sound 131, and the output characteristics of the output unit 140 output a sample sound having different frequency components from the sound that is actually heard. For this reason, the target sound processing system 100 processes the target sound 131 while taking into account the characteristics of the microphone and speaker to generate a processed target sound, a sample sound, and a processed sample sound.

[0018] For example, built-in microphones and speakers built into mobile terminal 120 are smaller than external microphones and speakers, and may have limited performance or reduced functionality, in order to keep the price of mobile terminal 120 down or to conserve space within the housing of mobile terminal 120. Furthermore, even with external microphones and speakers, the available functions may be limited to certain functions, certain performance may be restricted, or conversely, certain functions may be enhanced, depending on the intended use and price. For this reason, there are many microphones and speakers with a variety of functions and performance capabilities.

[0019] In this way, if a microphone dedicated to collecting target sound 131 or a speaker dedicated to outputting a sample sound is used, there is no need to adjust for variations between microphones or speakers. However, if a dedicated microphone or speaker must be used, the dedicated microphone must be carried to a predetermined location each time target sound 131 is collected, and when the sample sound is to be listened to, the dedicated microphone must be transported to the location where the dedicated speaker is installed, making it difficult to respond flexibly and quickly.

[0020] Therefore, in the target sound processing system 100, in order to eliminate errors that depend on the characteristics of each device, at each stage of processing the picked-up target sound 131, absolute sounds (processed target sound, trial sound, processed trial sound 141) are generated that eliminate errors caused by the devices, and by processing these absolute sounds, sounds that are no different from real sounds can be reproduced so that the listener (user) can experience them.

[0021] Therefore, in the target sound processing system 100, the characteristics of the sound collection unit 130 (microphone) and the characteristics of the output unit 140 (speaker) are transmitted to the target sound processing device 110 along with the target sound data of the collected target sound 131 and the predetermined place image 160. Alternatively, the characteristics of the sound collection unit 130 and the output unit 140 may be transmitted to the target sound processing device 110 separately from the target sound 131 and the predetermined place image 160. For example, they may be transmitted before or after the transmission of the target sound 131 and the predetermined place image 160. The target sound processing device 110 processes the acquired target sound data using a correction value according to the characteristics of the sound collection unit 130 that collected the sound, thereby generating a processed target sound. Next, the target sound processing device 110 generates a trial sound from the generated processed target sound, reproducing the sound that would actually be heard when the target sound 131 is listened to at a predetermined place under a predetermined environment.

[0022] The target sound processing device 110, for example, recreates a predetermined environment (such as a building) in a virtual space, and generates a sample sound by predicting the acoustic effects in the predetermined environment using data on the picked-up target sound 131. The target sound processing device 110 processes the generated sample sound using a correction value according to the characteristics of the output unit 140, thereby generating a processed sample sound 141. By processing the sample sound using a correction value based on the characteristics of the output unit 140 in this way, it is possible to reliably reproduce the sound that is actually heard in the output unit 140 without using a dedicated speaker. Note that the sample sound may be regenerated based on the results of listening to the processed sample sound 141. By regenerating the sample sound in this way, sample sounds in various environments can be reproduced. Therefore, it is possible to simulate sample sounds in various environments by, for example, changing the window area or the grade of sound insulation performance.

[0023] Here, the predetermined environment includes, for example, buildings such as apartment buildings or detached houses that are scheduled to be constructed at the predetermined location where the target sound 131 is collected, as well as the interiors of these buildings and their outdoor areas such as verandas. Furthermore, the predetermined environment may include various factors that realize the living environment of the building, such as the position and area of ​​the building's walls, sound insulation performance (sound transmission loss), the position, area, and sound insulation performance (sound transmission loss) of windows attached to the building, the area and number of air intakes, sound insulation performance (normalized sound transmission loss), indoor surface area, and sound absorption performance (sound absorption power).

[0024] Then, the target sound processing device 110 associates and stores the generated processed sample sound 141, the target sound 131, the sound collection conditions, the processed target sound, the sample sound, and the predetermined location image 160. For example, when a user transmits a sample location image 151 capturing an image of a location where the user wants to sample the processed sample sound 141 to the target sound processing device 110, the target sound processing device 110 searches for a predetermined location image 160 similar to the sample location image 151.

[0025] When the predetermined location image 160 similar to the preview location image 151 is found, the target sound processing device 110 acquires the processed preview sound 141 associated with the found predetermined location image 160 and outputs it from the output unit 140. This allows the user to preview the processed preview sound 141 at the location where the user wishes to preview the processed preview sound 141. Furthermore, since the target sound processing device 110 pre-stores the generated processed preview sound 141, it does not require time to generate the processed preview sound 141 from the target sound 131, and the processed preview sound 141 can be quickly provided to the user.

[0026] 1B, an overview of the operation of the target sound processing system 100 will be described. In step S101, the sound collection unit 130 collects the target sound 131. In step S103, the collected target sound 131 is transmitted from the sound collection unit 130 to the mobile terminal 120. The target sound 131 collected by the sound collection unit 130 is, for example, analog data.

[0027] Then, in step S105, a camera or the like attached to the mobile terminal 120 is used to capture an image of the predetermined location where the target sound 131 was picked up, thereby acquiring a predetermined location image. In step S107, the mobile terminal 120 transmits target sound data obtained by digitally converting the acquired target sound 131 and the predetermined location image data to the target sound processing device 110. Note that the predetermined location image 160 is previously converted into digital data by being captured by a digital camera or the like.

[0028] Then, in step S109, the target sound processing device 110 generates a processed target sound by processing the target sound 131 using the first processing correction value. The target sound processing device 110 also processes the processed target sound to generate a trial sound that reproduces an absolute pitch, which is a sound that is actually heard when the target sound 131 is listened to at a predetermined location under a predetermined environment. Furthermore, the target sound processing device 110 generates a processed trial sound by processing the generated trial sound using the second processing correction value. The target sound processing device 110 then associates and stores the sound collection conditions of the target sound 131, the processed target sound, the trial sound, the processed trial sound, and the predetermined location image.

[0029] In step S111, for example, the mobile terminal 120 transmits a listening location image 151 obtained by capturing an image of a listening location, which is a location different from the predetermined location and where the user wants to listen to a preview sound, with the camera 150 to the target sound processing device 110. In step S113, the target sound processing device 110 searches for a predetermined location image 160 similar to the received listening location image 151, and acquires the processed preview sound associated with the predetermined location image 160.

[0030] In step S115, the acquired processed sample sound is transmitted to mobile terminal 120 (output unit 140). In step S117, output unit 140 outputs the received processed sample sound. Here, sound collection unit 130 and output unit 140 may be built into mobile terminal 120. Furthermore, digital conversion of target sound 131 may be performed in target sound processing device 110. Note that predetermined location image 160 and listening location image 151 may be images adjusted to have the same scale. By unifying the image scale in this way, the time required to search for predetermined location image 160 similar to listening location image 151 can be shortened and the accuracy of the search can be improved. Furthermore, the imaging conditions for listening location image 151 and predetermined location image 160 may be unified. Furthermore, listening location image 151 and predetermined location image 160 may be consecutive images captured at a predetermined time interval. Furthermore, the predetermined location image 160 may be captured while the target sound 131 is being collected.

[0031] <Configuration of target sound processing device 110> Next, the configuration of the target sound processing device 110 will be described with reference to Fig. 2. The target sound processing device 110 has a storage unit 211, an image receiving unit 212, a search unit 213, an acquisition unit 214, and an output control unit 215.

[0032] Storage unit 211 associates and stores target sound, sound collection conditions, processed target sound, sample sound, processed sample sound, and a predetermined location image obtained by capturing an image of the predetermined location. Here, the target sound is, for example, sound collected at a predetermined location by a worker or user (sound collection worker) who owns mobile terminal 120 using sound collection unit 130 (microphone) connected to mobile terminal 120. Target sound includes, for example, indoor sound, outdoor sound, etc., but is not limited to these. In addition, the predetermined location is, for example, a planned construction site for a detached house or apartment building, a planned construction site for a commercial facility, office building, road, railway facility, airport, power plant, factory, etc., an existing building, etc., and is a location where stakeholders of the predetermined location or potential buyers would like to know the sound environment of the predetermined location.

[0033] First, the target sound processing device 110 acquires the first characteristic of the sound collection unit 130. The first characteristic of the sound collection unit 130 is various conditions when the target sound 131 is collected, and includes, for example, mechanical characteristics such as the frequency characteristics of the sound collection unit 130 (microphone), characteristics of software incorporated in the sound collection unit 130, and environmental characteristics such as temperature, humidity, wind direction, and wind volume at a predetermined location, but is not limited to these.

[0034] The target sound processing device 110 acquires a first processing correction value for processing the target sound 131 based on the acquired first characteristic, and processes the target sound 131 using the acquired first processing correction value to generate a processed target sound. The target sound processing device 110 acquires the first processing correction value stored in, for example, internal storage or external storage. The first processing correction value is a correction value for processing the target sound 131 in accordance with the acquired first characteristic.

[0035] Then, the target sound processing device 110 processes the target sound 131 (target sound data) using the acquired first processing correction value to generate a processed target sound. The target sound processing device 110 processes the target sound data by, for example, canceling a specific frequency component, to obtain a processed target sound.

[0036] Next, target sound processing device 110 processes the generated processed target sound to generate a trial sound that reproduces the absolute sound that will actually be heard when target sound 131 is listened to under a predetermined environment, for example, at a predetermined location or a location different from the predetermined location. Here, the predetermined environment is, for example, a building such as an apartment building or detached house that is scheduled to be constructed at the predetermined location where target sound 131 is collected, and includes the interior of these buildings and outdoor areas such as balconies. Furthermore, the predetermined environment may include various factors that realize the living environment of the building, such as the position and area of ​​the building's walls, sound insulation performance (sound transmission loss), the position, area, and sound insulation performance (sound transmission loss) of windows attached to the building, the area and number of air intakes, sound insulation performance (normalized sound transmission loss), the surface area of ​​the room, sound absorption performance (sound absorption capacity), and sound reflection from surrounding buildings.

[0037] The generated test sounds include, for example, (1) indoor noise caused by external noise, (2) spatial sound insulation performance (noise propagating from adjacent rooms), (3) noise propagating from inside and outside the building to the site boundary, (4) indoor quieting performance caused by air conditioning equipment, (5) floor impact sound insulation performance (floor impact sound), and (6) indoor reverberation time (sound reverberation).

[0038] Specifically, (1) refers to how noise from an external noise source, such as a train, sounds indoors. (2) refers to how a room with a TV or conference sound sounds in the next room. (3) refers to how noise sources such as equipment and machinery on the building premises (inside or outside the building) and operational noise sound to the boundary or neighbors. (4) refers to how the quietness of the room changes or how the sound of the air conditioning equipment sounds indoors when the noise source is an air conditioner or total heat exchanger. (5) refers to how the sound of jumping or running from the room above sounds in the room below. (6) refers to how the sound of talking or audio equipment resonates in the room during a meeting or lecture.

[0039] Then, the target sound processing device 110 acquires second characteristics of the output unit 140 for outputting the generated preview sound. The second characteristics are output conditions of the output unit 140 (speaker) for outputting the generated preview sound, and include, for example, mechanical characteristics of the output unit 140, features of software incorporated in the output unit 140, and environmental characteristics of the output location, but are not limited to these. Then, the target sound processing device 110 acquires second characteristics for all of the output units 140 that may output the sound.

[0040] The target sound processing device 110 acquires second processing correction values ​​for processing the sample sound based on the acquired second characteristics, and processes the sample sound using the acquired second processing correction values ​​to generate a processed sample sound 141. If there are multiple acquired second characteristics, the target sound processing device 110 acquires multiple second processing correction values ​​accordingly. That is, the target sound processing device 110 acquires second processing correction values ​​for all of the output sections 140 from which output may be possible. Then, if there are multiple acquired second processing correction values, the target sound processing device 110 generates multiple processed sample sounds 141.

[0041] Then, the target sound processing device 110 associates the generated processed preview sound 141 and the like with a predetermined location image 160 obtained by capturing an image of the predetermined location, and stores them. Here, the predetermined location image 160 is captured by a camera or the like. The camera may be either one built into the mobile terminal 120 such as a smartphone or tablet terminal, or a standalone camera. The captured predetermined location image 160 is saved as digital data in the camera's internal storage or external storage.

[0042] The processed sample sound 141, the predetermined location image 160, and the like stored in the storage unit 211 in the manner described above are used in the target sound processing device 110 as follows.

[0043] The image receiving unit 212 receives a listening location image 151 captured at a location different from the predetermined location where a user wishes to listen to a sample sound that reproduces absolute pitch, which is a sound that is actually heard when listening to the target sound 131 in a predetermined environment at the predetermined location. The listening location image 151 is captured using a camera 150, for example.

[0044] Note that predetermined location image 160 and listening location image 151 are images that have been adjusted to have the same scale. The scale may be adjusted, for example, by matching the imaging conditions (angle of view, etc.) when capturing predetermined location image 160 and listening location image 151, or by adjusting the scale using rendering software or the like after capturing predetermined location image 160 and listening location image 151. In this way, by previously adjusting predetermined location image 160 and listening location image 151 to have the same scale, the search speed and search accuracy by search unit 213 can be improved.

[0045] The search unit 213 searches the storage unit 211 for a predetermined location image 160 similar to the received listening location image 151. The search unit 213 searches the storage unit 211 for a predetermined location image 160 similar to the listening location image 151, for example, based on the degree of match between the listening location image 151 and the predetermined location image 160. The search unit 213 may, for example, extract feature points of both images and search for a predetermined location image 160 similar to the listening location image 151 based on the number of matching feature points among the extracted feature points.

[0046] Furthermore, the search unit 213 may input the predetermined place image 160, which is an image of the predetermined place, into artificial intelligence (AI) to perform machine learning. When the machine learning by the AI ​​is completed, the search unit 213 generates a trained predetermined place image model. Note that the search unit 213 may store the generated trained predetermined place image model in a predetermined storage or the like. In this case, the stored trained predetermined place image model may be updated each time a new learning image is acquired, machine learning is performed, and a trained initial location image model is generated.

[0047] Machine learning using artificial intelligence is performed using a known algorithm. In machine learning, a loss function specifies weights and uses the inverse of the number of events. Furthermore, the search unit 213 inflates the number of images of a predetermined location that the artificial intelligence learns in order to improve the accuracy of the machine learning using the artificial intelligence and generate a more accurate model for similar image search. The search unit 213 obtains the inflated data by, for example, flipping the images horizontally. Furthermore, the search unit 213 may use transfer learning to improve the accuracy of the machine learning using the artificial intelligence. Here, transfer learning is a technique that aims to improve the performance of a model by using a different dataset to repurpose a trained model for a different problem and performing partial learning. This technique is particularly promising for improving inference performance and reducing learning time when there is insufficient training data.

[0048] The acquisition unit 214 acquires the processed preview sound 141 associated with the searched predetermined location image 160. The storage unit 211 stores the processed preview sound 141 that has been generated in advance based on the acquired target sound 131, in association with the predetermined location image 160. Furthermore, the predetermined location image 160 further stores the sound collection conditions, the processed target sound, the preview sound, and the like in association with each other.

[0049] The output control unit 215 controls the generated processed preview sound 141 to output it from the output unit 140. For example, the output control unit 215 transmits the generated processed preview sound 141 to the mobile terminal 120, and outputs the processed preview sound 141 from the output unit 140. Alternatively, the output control unit 215 may transmit the generated processed preview sound 141 directly to the output unit 14, thereby outputting the processed preview sound 141 from the output unit 140.

[0050] <Configuration of mobile terminal 120> Next, the configuration of the mobile terminal 120 will be described with reference to Fig. 2. The mobile terminal 120 has an acquisition unit 221, a transmission unit 222, a reception unit 223, and an output control unit 224. Note that the sound collection unit 130 and the output unit 140 may be built into the mobile terminal 120.

[0051] The acquisition unit 221 acquires the target sound 131 collected by the sound collection unit 130 (microphone) at a predetermined location. Here, the target sound 131 includes at least one of indoor sound and outdoor sound. The data of the acquired target sound 131 is analog data, but if the sound collection unit 130 has an AD converter, for example, the analog data of the collected target sound 131 may be converted into digital data in the sound collection unit 130. If the sound collection unit 130 does not have an AD converter, the analog data may be converted into digital data using an AD converter included in the mobile terminal 120. If neither the sound collection unit 130 nor the mobile terminal 120 has an AD converter, the analog data of the collected target sound 131 may be converted into digital data in an AD converter included in the target sound processing device 110.

[0052] The transmitting unit 222 transmits the acquired data of the target sound 131 to the target sound processing device 110. The data of the target sound 131 transmitted to the target sound processing device 110 may be analog data or digital data.

[0053] The receiving unit 223 receives the processed sample sound 141 transmitted from the target sound processing device 110. The data of the received processed sample sound 141 is digital data.

[0054] The output control unit 224 controls the received processed sample sound 141 to output it from the output unit 140. The output control unit 224 converts the received digital data of the processed sample sound 141 into analog data, sends it to the output unit 140, and controls the output unit 140 to play back the analog data of the processed sample sound 141. As described above, the processed target sound may be sent directly from the output control unit 215 of the target sound processing device 110 to the output unit 140 for output.

[0055] Next, an example of a microphone correction value table 301 possessed by the target sound processing device 110 will be described with reference to Fig. 3A. The microphone correction value table 301 stores characteristics 312 and correction values ​​313 in association with microphone IDs (Identifiers) 311. The microphone IDs 311 are identifiers for identifying microphones capable of collecting the target sound 131. The characteristics 312 indicate the characteristics of the microphone, and include frequency characteristics and a sound collection level. The correction values ​​313 include a frequency correction value and a level correction value. The target sound processing device 110 then refers to the microphone correction value table 301, extracts a correction value according to the characteristics of the microphone (sound collection unit 130), and processes the received target sound 131.

[0056] 3B, an example of the speaker correction value table 302 possessed by the target sound processing device 110 will be described. The speaker correction value table 302 stores characteristics 322 and correction values ​​323 in association with speaker IDs 321. The speaker IDs 321 are identifiers for identifying speakers capable of outputting the processed sample sound 141 transmitted from the target sound processing device 110. The characteristics 322 indicate speaker characteristics, including frequency characteristics and available output levels. The correction values ​​323 include frequency correction values ​​and level correction values. The target sound processing device 1110 then refers to the speaker correction value table 302, extracts correction values ​​according to the characteristics of the speaker (output unit 140), and processes the generated sample sound.

[0057] 3C , an example of a target sound image table 303 held by the target sound processing device 110 will be described. The target sound image table 303 stores an image 332, a location 333, and sound collection / imaging conditions 334 in association with a target sound ID 331. The target sound ID 331 is an identifier for identifying each target sound 131 collected by the sound collection unit 130. The image 332 is an image stored in association with the collected target sound 131, and is an image of the location where the target sound 131 was collected. The location 333 is coordinate data indicating the location where the target sound 131 was collected, and multiple target sounds 131 and images may correspond to one location. The sound collection / imaging conditions 334 are the sound collection conditions when the target sound 131 was collected and the imaging conditions when an image of the sound collection location was captured, and include the distance from the subject, exposure time, weather, temperature, etc.

[0058] 3D , an example of the stored data table 304 possessed by the target sound processing device 110 will be described. The stored data table 304 stores a processed target sound 341, a preview sound 342, and a processed preview sound 343 in association with a target sound ID 331. The processed target sound 341 is data of a sound obtained by processing the target sound 131 using a first processing correction value. The preview sound 342 is data of a sound generated by processing the processed target sound. The processed preview sound 343 is data of a sound generated by processing the generated preview sound using a second processing correction value.

[0059] Then, the target sound processing device 110 refers to these tables 301 , 302 , 303 , and 304 to search for the processed sample sound 141 associated with a predetermined location image similar to the sample location image 151 .

[0060] The hardware configuration of the target sound processing device 110 will be described with reference to FIG. 4. The CPU (Central Processing Unit) 410 is a processor for arithmetic and control, and executes programs to realize the various functional components of the target sound processing device 110 shown in FIG. 2. The CPU 410 may have multiple processors and execute different programs, modules, tasks, threads, etc. in parallel. The ROM (Read Only Memory) 420 stores fixed data such as initial data and programs, as well as other programs. The network interface 430 communicates with other devices via a network. The CPU 410 is not limited to a single CPU, and may include multiple CPUs or a GPU (Graphics Processing Unit) for image processing. The network interface 430 preferably has a CPU independent of the VPU 410 and writes and reads transmitted and received data to and from an area in the RAM (Random Access Memory) 440. It is also preferable to provide a DMAC (Direct Memory Access Controller) (not shown) for transferring data between the RAM 440 and the storage 450. The CPU 410 recognizes that data has been received or transferred to the RAM 440 and processes the data accordingly. The CPU 410 also prepares the processing results in the RAM 440, and leaves the subsequent transmission or transfer to the network interface 430 or DMAC.

[0061] The RAM 440 is a random access memory used by the CPU 410 as a work area for temporary storage. A storage area for storing data necessary for implementing this embodiment is secured in the RAM 440. The preview location image data 441 is image data of a location where the user wishes to preview the processed preview sound 141. The predetermined location image data 442 is image data of a location where the target sound 131 was picked up. The processed preview sound data 443 is preview sound data obtained by processing the picked up target sound 131, and is preview sound data output from the output unit 140.

[0062] The transmission / reception data 444 is data transmitted and received via the network interface 430. The RAM 440 also has an application execution area 445 for executing various applications.

[0063] The storage 450 stores a database, various parameters, or the following data or programs required to implement this embodiment. The storage 450 stores a microphone correction value table 301, a speaker correction value table 302, a target sound image table 303, and a stored data table 304. The microphone correction value table 301 is a table that manages the relationship between the microphone ID 311 and the correction value 313, etc., shown in FIG. 3A. The speaker correction value table 302 is a table that manages the relationship between the speaker ID 321 and the correction value 323, etc., shown in FIG. 3B. The target sound image table 303 is a table that manages the relationship between the target sound ID 331 and the image 332, etc., shown in FIG. 3C. The stored data table 304 is a table that manages the relationship between the target sound ID 331 and the processed sample sound 343, etc., shown in FIG. 3D.

[0064] The storage 450 further stores an image receiving module 451, a search module 452, an acquisition module 453, and an output control module 454. The image receiving module 451 is a module that receives a listening location image 151 captured of a location where the user wishes to listen to the processed preview sound 141. The search module 452 is a module that searches the storage unit 211 for a predetermined location image 160 that is similar to the received listening location image 151. The acquisition module 453 is a module that acquires the processed preview sound 141 associated with the searched predetermined location image 160. The output control module 454 is a module that controls the acquired processed preview sound 141 and outputs it from the output unit 140. These modules 451 to 454 are read into the application execution area 445 of the RAM 440 by the CPU 410 and executed. The control program 455 is a program for controlling the entire target sound processing device 110.

[0065] The input / output interface 460 interfaces input / output data with input / output devices. A display unit 461 and an operation unit 462 are connected to the input / output interface 460. A storage medium 464 may also be connected to the input / output interface 460. A speaker 463 serving as an audio output unit, a microphone (not shown) serving as an audio input unit, or a GPS position determination unit may also be connected. Note that the RAM 440 and storage 450 shown in FIG. 4 do not include programs or data related to the general-purpose functions of the target sound processing device 110 or other feasible functions.

[0066] Next, a processing procedure of the target sound processing device 110 will be described with reference to the flowchart shown in Fig. 5A. This flowchart is executed by the CPU 410 in Fig. 4 using the RAM 440, and realizes each functional configuration of the target sound processing device 110 in Fig. 2.

[0067] In step S501, the storage unit 211 associates and stores the collected target sound 131, sound collection conditions, processed target sound, preview sound, processed preview sound 141, and predetermined location image 160. In step S503, the image receiving unit 212 receives a preview location image 151 that is an image of the location where the processed preview sound 141 was previewed. In step S505, the search unit 213 searches the storage unit 211 for a predetermined location image 160 that is similar to the received preview location image 151. The search is performed, for example, based on the degree of match between the predetermined location image 160 and the preview location image 151.

[0068] In step S507, the search unit 213 determines whether or not a predetermined location image 160 similar to the listening location image 151 has been found. If it is determined that the predetermined location image 160 has not been found (NO in step S507), the target sound processing device 110 ends the process. If it is determined that the predetermined location image 160 has been found (YES in step S507), the target sound processing device 110 proceeds to the next step.

[0069] In step S509, the acquisition unit 214 acquires the processed preview sound 141 associated with the searched predetermined location image 160. In step S511, the output control unit 215 controls the acquired processed preview sound 141 to output it from the output unit 140.

[0070] Next, a processing procedure of mobile terminal 120 will be described with reference to the flowchart shown in Fig. 5B. This flowchart is executed by a CPU (not shown) using a RAM (not shown), and realizes each functional configuration of mobile terminal 120 shown in Fig. 2.

[0071] In step S531, the mobile terminal 120 acquires the target sound 131 collected by the sound collection unit 130. Here, the mobile terminal 120 may digitally convert the acquired target sound 131. In step S533, the mobile terminal 120 transmits data of the acquired target sound 131 to the target sound processing device 110. In step S535, the mobile terminal 120 receives the processed sample sound 141 from the target sound processing device 110. The received processed sample sound 141 is output from the output unit 140.

[0072] According to this embodiment, the processed sample sound generated by processing the target sound is stored in association with information about the location where the target sound was picked up, so that the sample sound can be listened to based on the location information. Also, because the processed sample sound is generated and stored in advance, there is no need to generate the processed sample sound each time, and it becomes possible for the user to listen to the processed sample sound in a shorter time. Furthermore, because the processed sample sound is generated and stored in advance, it is possible to avoid the trouble of generating and re-processing the processed sample sound each time, and it becomes possible for the user to listen to the processed sample sound in a shorter time.

[0073] [Second embodiment] Next, a target sound processing system 600 according to a second embodiment of the present invention will be described with reference to Figs. 6 to 9. Fig. 6 is a block diagram for explaining the configuration of the target sound processing system 600 according to this embodiment. In the target sound processing system 600 according to this embodiment, the target sound processing device 610 differs from the target sound processing device 110 of the target sound processing system 100 according to the first embodiment in that it includes a distance estimation unit, an assumed distance estimation unit, and a correction unit. Since the other configurations and operations are the same as those in the first embodiment, the same configurations and operations are denoted by the same reference numerals and detailed description thereof will be omitted.

[0074] In the first embodiment, the target sound processing device 110 generates processed sample sound 141 in advance for the collected target sound 131, and stores the processed sample sound 141 in association with a predetermined location image 160 that is an image of the location where the target sound 131 is collected. Then, the target sound processing device 110 searches for a predetermined location image 160 that is similar to the sample location image 151, and causes the output unit 140 to output the processed sample sound 141 associated with the searched predetermined location image 160.

[0075] Here, when comparing the listening location image 151 captured at the location where the user wants to listen to the processed listening sound 141 with the predetermined location image 160, the distance from the sound source does not necessarily match. In other words, when the distance from the sound source is different, the way the sound sounds at the listening location will differ depending on the distance from the sound source. For example, the farther away you are from the sound source, the more the sound that reaches the listening location from the sound source attenuates, and so on, and the way it sounds will change.

[0076] Therefore, in this embodiment, even if it is not possible to capture the listening location image 151 at the same distance, the pre-generated processed listening sound 141 is corrected by a correction amount according to the distance from the sound source position. As a result, in this embodiment, it is possible to adjust the way the processed listening sound 141 sounds even if the distance from the sound source varies.

[0077] Target sound processing device 610 further includes distance estimation unit 611, assumed distance estimation unit 612, and correction unit 613. Distance estimation unit 611 estimates the distance from the sound source position to sound collection unit 130 using scale-adjusted predetermined location image 160. Since scale-adjusted predetermined location image 160 has been performed, distance estimation unit 611 estimates the distance from the sound source position to sound collection unit 130 based on, for example, the size of an object such as a sound source in the image. Alternatively, an object that serves as a landmark for distance estimation may be captured in predetermined location image 160, and the distance between the sound source and sound collection unit 130 may be estimated using the size relationship between this object and an object such as a sound source.

[0078] The expected distance estimation unit 612 estimates the expected distance from the expected sound source position to the expected listening position using the scale-adjusted listening location image 151. Since the listening location image 151 has been scale-adjusted, the expected distance estimation unit 612 estimates the expected distance from the sound source position to the expected listening position based on, for example, the size of a subject such as a sound source in the image, in the same way as the distance estimation unit 611. Note that an object that serves as a landmark when estimating the expected distance may be placed at the expected listening position.

[0079] The correction unit 613 corrects the processed sample sound 141 associated with the predetermined location image 160 searched for by the search unit 213, based on the estimated distance and the estimated assumed distance. The correction unit 613 corrects, for example, the volume, pitch, etc. of the processed sample sound 141, according to the distance estimated by the distance estimation unit 611. The parameters corrected by the correction unit 613 are not limited to the volume, pitch, etc.

[0080] Next, an example of the correction amount table 701 possessed by the target sound processing device 610 will be described with reference to Fig. 7. The correction amount table 701 stores an estimated distance 712 and a correction amount 713 in association with an estimated distance 711. The estimated distance 711 is a distance estimated by the distance estimation unit 611. The estimated distance 712 is a distance estimated by the estimated distance estimation unit 612. The correction amount 713 is a correction value of a parameter for correcting the processed sample sound 141 in accordance with the relationship between the estimated distance 711 and the estimated distance 712. The correction amount 713 includes a frequency correction value, a level correction value, etc. Then, the correction unit 613 corrects the processed sample sound by referring to the correction amount table 701.

[0081] The hardware configuration of the target sound processing device 610 will be described with reference to Fig. 7. The RAM 840 is a random access memory used by the CPU 410 as a work area for temporary storage. A storage area for storing data necessary for implementing this embodiment is secured in the RAM 840. The estimated distance data 841 is data on the estimated distance from the sound source position to the sound collection unit 130 in the predetermined location image 160. The assumed distance 842 is data on the estimated distance from the assumed sound source position to the assumed listening position in the listening location image 151. The correction amount data 843 is data on the amount by which the processed listening sound 141 should be corrected according to the estimated distance and the assumed distance.

[0082] The storage 850 stores a database, various parameters, or the following data or programs required to implement this embodiment. The storage 850 stores a correction amount table 701. The correction amount table 701 is a table that manages the relationship between the estimated distance 711 and the correction amount 713 shown in FIG. 7 .

[0083] The storage 850 further stores a distance estimation module 851, an assumed distance estimation module 852, and a correction module 853. The distance estimation module 851 is a module that estimates the distance from the sound source position to the sound collection unit 130 using the predetermined location image 160. The assumed distance estimation module 852 is a module that estimates the assumed distance from the assumed sound source position to the assumed listening position using the listening location image 151. The correction module 853 is a module that corrects the processed listening sound 141 associated with the predetermined location image 160 searched for by the search unit 213, based on the estimated distance. These modules 851 to 853 are read into the application execution area 445 of the RAM 840 by the CPU 410 and executed.

[0084] Next, a processing procedure of the target sound processing device 610 will be described with reference to the flowchart shown in Fig. 9. This flowchart is executed by the CPU 410 in Fig. 8 using the RAM 840, and realizes each functional configuration of the target sound processing device 610 in Fig. 6.

[0085] In step S901, the distance estimation unit 611 acquires the predetermined location image 160 and estimates the distance from the sound source position to the sound collection unit 130. In step S903, the assumed distance estimation unit 612 estimates the assumed distance from the assumed sound source position to the assumed listening position using the listening location image 151. In step S905, the correction unit 613 corrects the processed listening sound associated with the predetermined location image 160 searched for by the search unit 213, based on the estimated distance and the estimated assumed distance.

[0086] According to this embodiment, the processed preview sound is corrected based on the assumed distance and the estimated distance, so that the sound that is actually heard can be reproduced even if the specified location image and the preview location image are not captured under the same imaging conditions.

[0087] [Example] Next, examples of the present invention will be described with reference to Figures 10 to 19. It should be noted that the scope of the present invention is not limited to the following examples.

[0088] The target sound processing devices 110, 610 of the target sound processing systems 100, 600 generate test sounds for the six items ((a) to (f)) shown in Table 1. These are: (a) indoor noise caused by external noise, (b) room-to-room sound insulation performance (noise propagating from adjacent rooms), (c) noise propagating from inside and outside the building to the boundary of the site, (d) indoor noise caused by air conditioning equipment, (e) floor impact sound insulation performance (floor impact sound), and (f) indoor reverberation time (sound reverberation).

[0089] [Table 1]

[0090] Specific examples of each of the six items include (a) traffic noise (road traffic noise, railway noise, etc.), (b) airborne noise such as conversations in a conference room and hotel television sounds, (c) outdoor equipment and machinery, indoor equipment and machinery, (d) indoor air conditioning equipment, (e) floor impact noise (heavy floor impact noise and light floor impact noise), (f) conversations in a conference room, voices in a classroom, etc. For each calculation, the technologies described in Patent Documents 1 to 3, for example, are used.

[0091] The target frequencies for each of the six items are: (a) 50Hz to 5,000Hz band, (b) 100Hz to 5,000Hz band, (c) and (d) 50Hz to 5,000Hz band, (e) heavy floor impact sound: 50Hz to 630Hz band, light floor impact sound: 50Hz to 5,000Hz band, and (f) 100Hz to 5,000Hz band.

[0092] Regarding the generation of the sample sound, in (a) to (d), a volume reduction filter is generated and the sample sound is generated by filtering the sound source data (collected target sound data).

[0093] In (e), the floor impact sound level is calculated from the impact force of a standard impact source (tire, ball, tapping) according to the JIS standard, and sound source data close to the calculated conditions is extracted from the floor impact sound from the standard impact source recorded for each condition such as slab thickness, floor finish structure, and finished ceiling, and a filter is generated to reduce the level difference between that floor impact sound level and the calculated value, and a trial sound is generated. Note that the heavy floor impact sound is evaluated in the 50Hz to 630Hz band, so frequencies outside the target range are not filtered and are not played.

[0094] In (f), a test sound is generated by convolving the dry source sound, such as a reading sound recorded in an anechoic chamber, with the impulse response predicted by acoustic simulation or the actual measured value of the impulse response.

[0095] Next, the microphone (sound collection unit 130 or sound collection device) and headphones (output unit 140 or speaker) used each have their own unique acoustic characteristics. Therefore, in order to faithfully reproduce the generated sample sound, these acoustic characteristics must be corrected, and the collected sound source (target sound 131) and the generated sample sound are corrected based on a correction value calculated in advance. The correction value for the acoustic characteristics of the equipment used is calculated using the following method.

[0096] <<How to correct the acoustic characteristics of a microphone>> <Microphone frequency characteristics> A method for correcting the acoustic characteristics of a microphone (sound pickup unit 130) will be described using an example in which a tablet terminal is used as the mobile terminal 120. First, it is experimentally confirmed that the microphone can pick up the sound of the frequency to be evaluated. In an anechoic chamber, pink noise (static noise) is generated from a speaker (sound source), and the sound is picked up using a precision sound level meter placed 1 m away. Figure 10 shows the 1 / 3 octave band levels of the picked-up sound data.

[0097] The microphone attached to the tablet device and the precision sound level meter generally matched in the target frequency range of 50 Hz to 5,000 Hz. On the other hand, the microphone built into the tablet device showed level differences at low frequencies below 80 Hz. When capturing problematic low-frequency sounds such as heavy floor impact noise and equipment noise, it is preferable to use an external microphone.

[0098] <Microphone pickup level and linearity> When listening to the test sound, the sound pressure level to be picked up is assumed to be 30dB to 80dB, taking into consideration the user's hearing impairment and the level of noise that requires noise control measures when using the system. In addition, the frequency characteristics of the microphone used will change linearly depending on the sound pressure level to be picked up. Therefore, the linearity of the pickup level and frequency characteristics of the microphone to be used will be experimentally confirmed.

[0099] Pink noise was recorded at a volume of 30 dB to 80 dB using the tablet's built-in microphone and a precision sound level meter. The recorded sound was analyzed in 1 / 3 octave bands to examine the correspondence between the tablet's built-in microphone and the precision sound level meter. Figure 11 shows the results for the 1,000 Hz band as a representative example.

[0100] If the value of the precision sound level meter is taken as the true value, when using the built-in microphone of a tablet device, the change is roughly linear from 30dB to 80dB. Because the change is linear within the sound pressure level range expected for system use, correction can be made using a fixed correction value for each frequency band.

[0101] After confirming that the target microphone can be used with the target sound processing system 100,600, the 1 / 3 octave band level difference between the precision sound level meter and the target microphone is used as a correction value (first processing correction value) to ensure accurate sound pickup.

[0102] <<How to correct the acoustic characteristics of headphones>> <Headphone frequency characteristics> Many headphones are tuned according to purpose and preference, and to faithfully reproduce the generated listening sound, the frequency characteristics of each headphone must be canceled out. For this reason, headphones were attached to a dummy head in an anechoic chamber, and pink noise (sound source) was played via a tablet device. Figure 12 shows the 1 / 3 octave band levels of the pink noise and the sound reproduced by the headphones. Compared to the pink noise, which has a flat frequency response, the sound reproduced by the headphones used varies by frequency, with sounds in the 400 Hz to 1,000 Hz band being particularly emphasized.

[0103] <Headphone playback level and linearity> Pink noise is played through headphones at different levels to examine the reproducible level and linearity of the headphones. Figure 13 shows the 1 / 3 octave band level of the 1,000 Hz band as a typical example.

[0104] The headphones were able to reproduce sounds in the range of 20dB to 80dB, with a linear change in level in response to a 50dB change in pink noise level. Similar results were obtained in other frequency bands. Because reproduction down to 20dB is possible, for example, when considering the sound insulation of a space, the spatial sound insulation performance of up to Dr-55 can be expressed as a listening test sound for a sound source generated at 75dB in the sound source room. Here, Dr is the rating of the difference in sound pressure level between rooms, and the higher the Dr value, the higher the sound insulation performance of the space and the less airborne sound is transmitted.

[0105] After confirming that the target headphones can be used with the target sound processing system 100,600, the 1 / 3 octave band level difference between the pink noise (sound source) and the sound reproduced by the headphones is used as a correction value (second processing correction value) to cancel the acoustic characteristics unique to the headphones so that the generated trial sound can be faithfully reproduced.

[0106] <<Verification example of system accuracy in an actual building>> <Verification Overview> Using an office room, we verified the accuracy of predicted calculations, sample sounds, and processed sample sounds generated by the target sound processing devices 110 and 610 regarding spatial sound insulation performance. An overview of the measurement room is shown in Figure 14.

[0107] The partition wall between the conference room and the office is a dry double wall (sound insulation performance TL D -40). The corridor doors between the conference room and the office are standard steel double doors without airtight seals. The conference room was used as the sound source room, and target sounds (pink noise or a man reading aloud) were played from the speakers. The played target sounds were picked up by target sound processing devices 110 and 610 (using the built-in microphones of tablet devices), and the picked-up target sounds were used as sound source data. To verify the sound pickup accuracy, sound was also picked up by a precision sound level meter (NA-28 manufactured by RION). A dummy head was placed in the sound receiving room to pick up the sound transmitted from the conference room. The difference in sound pressure levels between rooms was also measured in accordance with JIS A1418:2000 "Method for measuring the airborne sound insulation performance of buildings."

[0108] <Prediction calculation accuracy> To confirm the accuracy of the prediction calculations made by the target sound processing system 100,600, the predicted values ​​of the sound pressure level difference between rooms were compared with the actual measured values. Figure 15 shows the comparison results for octave band levels. The predicted values ​​roughly correspond to the actual measured values ​​in the 125 Hz to 4,000 Hz bands, and the sound pressure level difference between rooms was predicted with good accuracy.

[0109] Because sound pickup accuracy affects the accuracy of generating the sample sound, we compared the sound pickup data from the target sound processing system 100,600 (microphone built into the tablet) with the sound pickup data from a precision sound level meter. The sound pressure waveform from the precision sound level meter and the sound pressure waveform (processed target sound) picked up by the target sound processing system 100,600 (microphone built into the tablet) are shown in Figure 16, and the octave band levels of each sound pressure waveform are shown in Figure 17.

[0110] The shape and amplitude of the sound pressure waveforms picked up by the target sound processing system 100,600 (tablet built-in microphone) are the same as those picked up by a precision sound level meter for both pink noise and a male reading voice. At the octave band level, the error is a maximum of about 1 dB, and the sound is picked up with the same precision as a normal sound level meter.

[0111] <Preview sound accuracy> A correction process was performed to cancel the acoustic characteristics specific to headphones on the sound pressure waveform of the test sound, which was generated by taking into account the predicted calculation results for the sound pickup data of the target sound processing system 100,600, and a processed test sound was generated.

[0112] The sound pressure waveform of the test sound, the sound pressure waveform of the processed test sound (sound played back through headphones), and the sound pressure waveform (actual measured values) collected using a dummy head in the sound receiving room office are shown in Figure 18, and their respective octave band levels are shown in Figure 19. Note that while the background noise in the sound receiving room is shown in Figure 19, it was excluded from the evaluation when playing back a male reading, as it is affected by background noise in the sound receiving room above the 1,000 Hz band.

[0113] The processed sample sound (sound played back through headphones) has almost the same shape and amplitude as the sample sound and the actual measured sound pressure waveform, and the octave band levels are also about the same.

[0114] The sample sounds generated by the target sound processing devices 110 and 610 of the system can be correctly played back through headphones, and the processed sample sounds can reproduce the actual sounds.

[0115] Although the present invention has been described above with reference to the embodiments, the present invention is not limited to the above-described embodiments and can be modified as appropriate. Various modifications that can be understood by those skilled in the art can be made to the configuration and details of the present invention within the scope of the present invention. Furthermore, outer containers and packaging boxes that combine the separate features included in each embodiment in any way are also included in the scope of the present invention.

[0116] The present invention may also be applied to a system consisting of multiple devices or to a single device. Furthermore, the present invention may also be applied when an information processing program that realizes the functions of the embodiments is supplied to a system or device and executed by a built-in processor. Therefore, the technical scope of the present invention also includes a program installed on a computer to realize the functions of the present invention, a medium storing the program, a WWW (World Wide Web) server from which the program is downloaded, and a processor that executes the program. In particular, the technical scope of the present invention also includes a non-transitory computer-readable medium storing a program that causes a computer to execute at least the processing steps included in the above-described embodiments.

Claims

1. Target sound picked up by a sound pickup unit at a predetermined location, The target sound collection conditions, a processed target sound obtained by processing the target sound using a first processing correction value for processing the target sound acquired based on the first characteristic of the sound collection unit; A sample sound that reproduces an absolute pitch that is a sound that is actually heard when the target sound is listened to in a predetermined environment at the predetermined location; a processed sample sound obtained by processing the sample sound using a second processing correction value for processing the sample sound, the second processing correction value being acquired based on a second characteristic of an output unit for outputting the sample sound; and a storage unit that stores an image of the predetermined location in association with the image; an image receiving unit that receives a listening location image obtained by capturing an image of a location where the user wishes to listen to the processed listening location sound; a search unit that searches the storage unit for the predetermined location image similar to the received listening location image; an acquisition unit that acquires, from the storage unit, a processed sample sound associated with the searched predetermined location image; an output control unit that controls the acquired processed sample sound and outputs it from the output unit; A target sound processing device comprising:

2. The target sound processing device according to claim 1 , wherein the predetermined location image and the listening location image are images adjusted to have the same scale.

3. a distance estimation unit that estimates a distance from a sound source position to the sound collection unit using the scale-adjusted predetermined location image; an estimated distance estimation unit that estimates an estimated distance from an estimated sound source position to an estimated listening position using the scale-adjusted listening location image; a correction unit that corrects the processed sample sound associated with the predetermined location image searched for by the search unit based on the estimated distance and the estimated assumed distance; The target sound processing device according to claim 2, further comprising:

4. 4. The target sound processing device according to claim 1, wherein the predetermined location image and the listening location image are consecutive images captured at predetermined time intervals.

5. 5. The target sound processing device according to claim 1, wherein the predetermined location image is captured while the target sound is being collected.

6. The target sound processing device according to any one of claims 1 to 5, wherein the processed sample sound includes at least one of indoor noise caused by external noise, noise propagating from an adjacent room, noise propagating from inside or outside the building to the site boundary, indoor noise caused by air conditioning equipment, floor impact sound insulation performance, and indoor reverberation time.

7. A model generation unit that generates a learned predetermined location image model by training an artificial intelligence to learn the predetermined location image, The target sound processing device according to any one of claims 1 to 5, wherein the search unit uses the learned predetermined location image model to search the storage unit for the predetermined location image similar to the listening location image.

8. The target sound processing device according to claim 7 , wherein the model generation unit generates the trained predetermined location image model using left-right inversion as padded data.

9. The target sound processing device according to claim 7 or 8, wherein the model generation unit generates the trained predetermined location image model using transfer learning.

10. Target sound picked up by a sound pickup unit at a predetermined location, The target sound collection conditions, a processed target sound obtained by processing the target sound using a first processing correction value for processing the target sound acquired based on the first characteristic of the sound collection unit; A sample sound that reproduces an absolute pitch that is a sound that is actually heard when the target sound is listened to in a predetermined environment at the predetermined location; a processed sample sound obtained by processing the sample sound using a second processing correction value for processing the sample sound, the second processing correction value being acquired based on a second characteristic of an output unit for outputting the sample sound; and a storage step of storing an image of the predetermined location in association with the image; an image receiving step of receiving a listening location image of a location where the user wishes to listen to the processed listening location sound; a searching step of searching a storage unit for an image of the predetermined location similar to the received image of the listening location; an acquisition step of acquiring, from the storage unit, a processed sample sound associated with the searched predetermined location image; an output control unit that controls the acquired processed sample sound and outputs it from the output unit; A target sound processing method including:

11. Target sound picked up by a sound pickup unit at a predetermined location, The target sound collection conditions, a processed target sound obtained by processing the target sound using a first processing correction value for processing the target sound acquired based on the first characteristic of the sound collection unit; A sample sound that reproduces an absolute pitch that is a sound that is actually heard when the target sound is listened to in a predetermined environment at the predetermined location; a processed sample sound obtained by processing the sample sound using a second processing correction value for processing the sample sound, the second processing correction value being acquired based on a second characteristic of an output unit for outputting the sample sound; and a storage step of storing an image of the predetermined location in association with the image; an image receiving step of receiving a listening location image of a location where the user wishes to listen to the processed listening location sound; a searching step of searching a storage unit for an image of the predetermined location similar to the received image of the listening location; an acquisition step of acquiring, from the storage unit, a processed sample sound associated with the searched predetermined location image; an output control unit that controls the acquired processed sample sound and outputs it from the output unit; A target sound processing program that causes a computer to execute the above.

Citation Information

Patent Citations

  • Sound insulating simulator

    JP1997258750A

  • Method for evaluating sound insulation of building by computer

    JP2001060211A

  • Acoustic environment bodily sensing equipment

    JP2001134272A

  • Sensing method for noise environment, trial listening apparatus and information storage medium

    JP2003156388A

  • Sound insulation structure design device

    JP4234257B2