Sound reproduction method, sound reproduction device, and sound reproduction program
A sound reproduction method combining multiple natural sounds with high complexity and low volume effectively masks noise, ensuring comfort and harmony in a space.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- PANASONIC INTELLECTUAL PROPERTY MANAGEMENT CO LTD
- Filing Date
- 2025-11-04
- Publication Date
- 2026-05-15
AI Technical Summary
Conventional noise masking technologies impair comfort in a space due to the use of disruptive masking sounds, which can bother listeners.
A sound reproduction method that combines multiple types of natural sounds, ensuring high sonic complexity and outputting at a volume equal to or lower than ambient noise, effectively masking noise without discomfort.
The method achieves effective noise masking in a space without compromising comfort, using a combination of natural sounds that maintain a harmonious environment.
Smart Images

Figure JP2025038554_15052026_PF_FP_ABST
Abstract
Description
Sound reproduction method, sound reproduction device, and sound reproduction program
[0001] The present disclosure relates to a technique for masking noise.
[0002] For example, Patent Document 1 discloses a mask sound generation device including mask sound generation means for generating a mask sound composed of a disturbing sound that disturbs speech, a background sound that continuously occurs, and an effect sound that intermittently occurs.
[0003] However, in the above conventional technology, there is a risk that the comfort in the space may be impaired, and further improvement has been required.
[0004] Japanese Patent No. 5747490
[0005] The present disclosure has been made to solve the above problems, and an object thereof is to provide a technique capable of masking noise in a space without impairing the comfort in the space.
[0006] A sound reproduction method according to an aspect of the present disclosure is a sound reproduction method executed by a computer, including acquiring sound data in which a plurality of types of natural sounds are combined, reproducing the acquired sound data, and outputting the reproduced sound data to a speaker.
[0007] According to the present disclosure, it is possible to mask noise in a space without impairing the comfort in the space.
[0008] This figure shows the configuration of the sound reproduction system according to this embodiment. This figure shows the aggregated results of a questionnaire regarding the degree of masking, the degree of unnoticeability, and the degree of likability for each of the 12 types of masking sounds. This figure illustrates a method for generating UpperRP from sound data that combines multiple types of natural sounds. This figure shows the complexity index value (UD) for each of the 12 types of masking sounds. This figure shows the relationship between the type of natural sound, its frequency characteristics, and information indicating whether the natural sound is a stationary or transient sound. This figure shows the relationship between the frequency of bird calls and the relative sound pressure level. This figure shows the relationship between the frequency of insect sounds and the relative sound pressure level. This figure shows the relationship between the frequency of flowing water sounds and the relative sound pressure level. This figure shows the relationship between the frequency of spring water sounds and the relative sound pressure level. This figure shows the relationship between the frequency of pink noise sounds and the relative sound pressure level. This is a flowchart for explaining the sound reproduction process by the sound reproduction device in the embodiment of this disclosure. This figure shows the mean and standard deviation of the scores for each of the multiple evaluation items in each of the multiple evaluation periods.
[0009] (Knowledge forming the basis of this disclosure) The above-mentioned prior art discloses masking human voices with a masking sound consisting of a disruptive sound that interferes with speech, a continuously occurring background sound, and an intermittently occurring sound for effect. Conventional masking masks noise by outputting a masking sound at a volume higher than the volume of ambient noise. In this case, the noise becomes less audible due to the masking sound, but the listener may be bothered by the masking sound, potentially impairing their comfort in the space.
[0010] To address the above challenges, the following technologies are disclosed.
[0011] (1) A sound reproduction method according to one aspect of the present disclosure is a sound reproduction method performed by a computer, which includes acquiring sound data in which a plurality of types of natural sounds are combined, reproducing the acquired sound data, and outputting the reproduced sound data to a speaker.
[0012] With this configuration, sounds combining multiple types of natural sounds are output from the speaker, making it possible to mask noise in the space without compromising comfort.
[0013] (2) In the sound reproduction method described in (1) above, the sound data may satisfy the conditions relating to the complexity of the sound.
[0014] With this configuration, multiple types of natural sounds are combined, and sounds that satisfy the conditions regarding sound complexity are output from the speaker, thus more effectively masking noise in the space.
[0015] (3) In the sound reproduction method described in (2) above, the condition may be that the index value relating to the complexity of the sound indicates that the complexity of the sound is high.
[0016] With this configuration, multiple types of natural sounds are combined, and sounds indicating high sonic complexity are output from the speaker, thus more effectively masking noise in the space.
[0017] (4) In the sound playback method described in any one of (1) to (3) above, the playback of the sound data may include playing the sound data at a volume equal to or lower than a predetermined volume.
[0018] With this configuration, sound at a volume equal to or lower than a predetermined volume is output from the speaker, thus masking noise in the space without compromising comfort within the space.
[0019] (5) In the sound reproduction method described in (4) above, the predetermined volume may be the volume of ambient noise in the space in which the speaker is installed.
[0020] With this configuration, sounds at a volume lower than the ambient noise level are output from the speakers, thus masking the ambient noise in the space without compromising comfort within the space.
[0021] (6) In the sound reproduction method described in any one of (1) to (5) above, the sound data may include sound data that combines the plurality of types of natural sounds and noise sounds.
[0022] With this configuration, a combination of multiple types of natural sounds and noise is output from the speaker, which allows for more effective masking of noise in the space.
[0023] (7) In the sound reproduction method described in (6) above, the noise sound may be pink noise.
[0024] With this configuration, a sound combining multiple types of natural sounds and pink noise is output from the speaker, which can more effectively mask noise in the space.
[0025] (8) In the sound reproduction method described in any one of (1) to (7) above, the plurality of natural sounds may include one or more stationary sounds and one or more non-stationary sounds.
[0026] This configuration combines multiple types of natural sounds, including one or more stationary sounds and one or more non-stationary sounds, increasing the complexity of the sound and allowing for more effective masking of noise in the space.
[0027] (9) In the sound reproduction method described in (8) above, the one or more stationary sounds may include at least one of the sounds of flowing water and the sounds of spring water, and the one or more non-stationary sounds may include at least one of the sounds of birdsong and the sounds of insects.
[0028] With this configuration, the complexity of the sound is increased by combining at least one of the sounds of flowing water and the sounds of spring water with at least one of the sounds of birdsong and insects, thereby more effectively masking noise in the space.
[0029] (10) In the sound reproduction method described in any one of (1) to (9) above, the plurality of types of natural sounds may include a plurality of natural sounds that are related to each other.
[0030] With this configuration, sounds combining multiple related natural sounds are output from the speaker, allowing for the masking of ambient noise in the space with sounds that are less jarring to the user.
[0031] Furthermore, this disclosure can be implemented not only as a sound reproduction method that performs the characteristic processing described above, but also as a sound reproduction device having a characteristic configuration corresponding to the characteristic processing performed by the sound reproduction method. It can also be implemented as a computer program that causes a computer to execute the characteristic processing included in such a sound reproduction method. Therefore, the same effects as the sound reproduction method described above can be achieved in the following other embodiments.
[0032] (11) A sound reproduction device according to another aspect of the present disclosure is a sound reproduction device comprising a processor, the processor acquiring sound data which is a combination of a plurality of types of natural sounds, playing back the acquired sound data, and outputting the played back sound data to a speaker.
[0033] (12) A sound playback program according to another aspect of the present disclosure acquires sound data which is a combination of several types of natural sounds, plays the acquired sound data, and causes a computer to function to output the played sound data to a speaker.
[0034] A non-temporary computer-readable recording medium relating to another aspect of this disclosure records the sound playback program described in (12) above.
[0035] Embodiments of this disclosure will be described below with reference to the attached drawings. Note that the embodiments described below are all specific examples of this disclosure. The numerical values, shapes, components, steps, and order of steps shown in the following embodiments are examples only and are not intended to limit this disclosure. Furthermore, components in the following embodiments that are not described in the independent claim representing the highest-level concept will be described as optional components. Also, in all embodiments, the contents of each can be combined.
[0036] (Embodiment) Figure 1 is a diagram showing the configuration of the sound reproduction system 100 according to this embodiment.
[0037] The sound reproduction system 100 shown in Figure 1 comprises a sound reproduction device 1 and a speaker 2.
[0038] Speaker 2 is located outside the main body of the sound playback device 1. Speaker 2 may also be connected to the sound playback device 1 by wire or wireless connection, or it may be integrated with the sound playback device 1.
[0039] Speaker 2 converts the sound data reproduced by sound playback device 1 into sound and outputs the converted sound to the outside. Speaker 2 is installed in a predetermined space. The predetermined space is a place where multiple people can gather, such as an office workspace, a building entrance, a building break room, a cafe, a restaurant, a hospital, and a nursing home.
[0040] The sound reproduction system 100 may also include an amplifier that amplifies the sound data input from the sound reproduction device 1. In this case, the speaker 2 may convert the sound data amplified by the amplifier into sound and output the converted sound to the outside.
[0041] The sound playback device 1 includes, for example, a computer system comprising a control program, a processing circuit such as a processor or logic circuit that executes the control program, and a recording device such as an internal memory or an accessible external memory that stores the control program. The sound playback device 1 may be realized, for example, by hardware implementation using a processing circuit, by execution of a software program held in memory by a processing circuit or distributed from an external server, or by a combination of these hardware and software implementations. The sound playback device 1 may also be a server, a terminal, or a system comprising a server and a terminal. The terminal may be a smartphone or a smart speaker.
[0042] Furthermore, if the sound playback device 1 is a server, the sound playback device 1 may be connected to the speaker 2 via a network so as to be able to communicate with it. The network is, for example, the internet.
[0043] The sound reproduction device 1 includes a sound data storage unit 11, an acquisition unit 12, a reproduction unit 13, and an output unit 14.
[0044] The sound data storage unit 11 stores in advance sound data in which a plurality of types of natural sounds are combined.
[0045] The acquisition unit 12 acquires sound data in which a plurality of types of natural sounds are combined. The acquisition unit 12 reads out the sound data from the sound data storage unit 11.
[0046] The reproduction unit 13 reproduces the sound data acquired by the acquisition unit 12. The reproduction unit 13 reproduces the sound data at a volume equal to or lower than a predetermined volume. The predetermined volume is the volume of ambient noise in a predetermined space. The predetermined space is, for example, an office work area, and the volume of ambient noise is, for example, 45 to 50 dB. For example, it is preferable that the reproduction unit 13 reproduces the sound data at a volume of 45 dB or less. Also, for example, it is preferable that the reproduction unit 13 reproduces the sound data at a volume of 35 dB or more and 45 dB or less.
[0047] Incidentally, the predetermined volume may be, for example, the volume of human conversation sound in a predetermined space. Also, the reproduction unit 13 may reproduce the sound data at the same volume as the volume of ambient noise, or may reproduce the sound data at a volume slightly higher than the volume of ambient noise (for example, ambient noise volume + 5 dB). The reproduction unit 13 may reproduce the sound data at a volume of 50 dB or less.
[0048] The output unit 14 outputs the sound data reproduced by the reproduction unit 13 to the speaker 2. The speaker 2 converts the sound data output by the output unit 14 into sound and outputs the converted sound into a predetermined space.
[0049] Subsequently, the sound data in which a plurality of types of natural sounds are combined will be described in more detail.
[0050] First, I will describe an experiment to verify the optimal combination of several types of natural sounds. In the experiment, a non-stationary sound noise source was placed 4 meters away from the subjects in an anechoic chamber, and a masking sound source was placed 0.5 meters away from the subjects. The non-stationary sound noise and masking sound were emitted from behind the subjects. The subjects were five men aged 20 to 50. The non-stationary sound noise was a combination of human voices and the sound of air conditioning equipment. The subjects were exposed to 12 different masking sounds.
[0051] The 12 types of masking sounds are: artificial sounds for comparison, insect sounds, bird sounds, spring sounds, flowing water sounds, pink noise, a combination of flowing water sounds and spring sounds, a combination of flowing water sounds and bird sounds, a combination of flowing water sounds and insect sounds, a combination of flowing water sounds, spring sounds and insect sounds, a combination of flowing water sounds, spring sounds and bird sounds, and a combination of flowing water sounds, spring sounds, bird sounds and pink noise. The artificial sound is a human voice played in reverse. The spring sound represents the sound of water gushing out. The flowing water sound represents the sound of water flowing, such as the sound of a babbling brook.
[0052] First, the subjects performed a 100-square calculation while listening to non-stationary noise and masking sounds for 30 seconds. The 100-square calculation is a calculation method in which a grid of 10 squares in each row and column is used, with 10 numbers from 0 to 9 written once on the left and top sides of each grid. The respondent then writes the results of addition, subtraction, multiplication, or division in the intersecting squares.
[0053] Next, the subjects answered a questionnaire for 30 seconds. In the questionnaire, the subjects evaluated the degree of masking, the degree of inattention, and the degree of likability.
[0054] The degree of masking is evaluated on a four-point scale. Participants give a score of 1 if they can almost hear the human voice in the non-stationary noise, a score of 2 if they can hear the human voice in the non-stationary noise but can almost not understand the content unless they concentrate on the voice, a score of 3 if they can hear the human voice in the non-stationary noise but can almost not understand the content even when they concentrate on the voice, and a score of 4 if the non-stationary noise is almost inaudible.
[0055] The degree to which the masking sound is not noticeable is evaluated on a five-point scale. Participants give a score of 1 if the masking sound is very noticeable, 2 if it is quite noticeable, 3 if it is not so noticeable, 4 if it is not noticeable, and 5 if it is not noticeable at all.
[0056] The degree of likability is rated on a five-point scale. Participants give a score of 1 if the masking sound is very undesirable, 2 if it is quite undesirable, 3 if it is somewhat desirable, 4 if it is desirable, and 5 if it is very desirable.
[0057] After completing a 30-second questionnaire, a non-stationary noise and another type of masking sound will be output 30 seconds later. Subsequently, the process of listening to the non-stationary noise and masking sounds, and answering the questionnaire, will be repeated until all 12 types of masking sounds have been output.
[0058] Figure 2 shows the results of a questionnaire survey regarding the degree of masking, the degree of unnoticeability, and the degree of likability for each of the 12 types of masking sounds. Note that the score for the degree of masking has been multiplied by 1.25. The scores shown are the average scores of 5 people.
[0059] As shown in Figure 2, artificial sounds received the lowest scores. Furthermore, the scores for single sounds such as insect chirps, bird chirps, spring water sounds, and pink noise were all below the threshold of 9 points. Conversely, the scores for combinations of flowing water sounds and spring water sounds, flowing water sounds and bird chirps, flowing water sounds and insect chirps, flowing water sounds, spring water sounds and insect chirps, flowing water sounds, spring water sounds and bird chirps, and flowing water sounds, spring water sounds, bird chirps, and pink noise were all above the threshold of 9 points. In particular, the score for the combination of flowing water sounds, spring water sounds, bird chirps, and pink noise was the highest.
[0060] From the above experiments, it can be seen that sounds combining multiple types of natural sounds are suitable for masking non-stationary noise. Furthermore, it is preferable that the sound data includes sound data that combines multiple types of natural sounds and noise sounds. The noise sound may be pink noise, white noise, or brown noise. Pink noise is more preferable.
[0061] Furthermore, sound data that combines multiple types of natural sounds satisfies the conditions related to sound complexity. The condition is that the index value related to sound complexity indicates high sound complexity. More specifically, the condition is that the index value, which indicates that sound complexity increases as the value decreases, falls below a predetermined threshold.
[0062] Here, we will explain the index values related to sound complexity. These index values are calculated based on the recurrence plot information obtained from the recurrence plot.
[0063] A recurrence plot is a method of nonlinear time series analysis, and the recurrence plot information obtained is represented by a planar diagram. Recurrence plot information can be described as two-dimensional array information.
[0064] In a recurrence plot, the same time series data is associated with both the vertical and horizontal axes. Points are plotted where two time series data are close together (i.e., corresponding to a digital value of 1), and points are not plotted where two time series data are far apart (i.e., corresponding to a digital value of 0), thereby generating recurrence plot information. Here, distance can be defined using Euclidean distance, etc., if the time series data is represented as a vector (or scalar).
[0065] The recurrence plot information indicates that the time series data has periodicity when lines parallel to the central line (Line of Identity) are aligned. The distance to the central line indicates the period.
[0066] In this embodiment, an index relating to the complexity of sound is calculated using UpperRP (Recurrence Plot), which is recurrence plot information obtained by a hierarchical recurrence plot.
[0067] Furthermore, recurrence plots and hierarchical recurrence plots are disclosed in non-patent literature (Miwa Fukino et al., "Coarse-Graining Time Series Data: Recurrence Plot of Recurrence Plots and Its Application for Music", Chaos: An Interdisciplinary Journal of Nonline Science, vol. 2, no. 26, pp. 0-12, doi: 10.1063 / 1.4941371).
[0068] UpperRP is generated based on sound data. The following describes how to generate UpperRP from sound data that combines multiple types of natural sounds.
[0069] Figure 3 illustrates a method for generating UpperRP from sound data that combines multiple types of natural sounds.
[0070] The sound signal 201 shown in Figure 3 represents the time waveform of a sound combining the sounds of flowing water, spring water, birdsong, and pink noise.
[0071] First, the sound signal 201 is divided into n processing units defined by a window width T1 and a shift width T2. The window width T1 is, for example, 2.0 sec, the shift width T2 is, for example, 0.5 sec, and the number of processing units n is, for example, several tens to several hundred. The specific numerical values of the window width T1, the shift width T2, and the number of processing units n are not particularly limited.
[0072] Next, LowerRP202 is generated from each of the n processing units. For example, if the time-series data of the sound signal corresponding to one processing unit is mapped to the vertical and horizontal axes, and the i-th state on the vertical axis (specifically, the amplitude of the sound signal 201) is represented as s(i), and the j-th state on the horizontal axis is represented as s(j), and LowerRP is represented as LRP, then LRP(i, j) = d(s(i), s(j)). Note that 1 ≤ i, j ≤ m (m is a natural number greater than or equal to 2). d is a function that indicates distance, for example, a function that calculates the absolute value of the difference between two states. Thus, LowerRP202 is, for example, matrix data composed of m × m elements. In Figure 3, LowerRP202 is schematically illustrated in grayscale.
[0073] Next, the n LowerRP202 generated from each of the n processing units are mapped to the vertical and horizontal axes to generate UpperRP203. UpperRP203 is, for example, an n x n matrix data. In Figure 3, UpperRP203 is schematically illustrated in grayscale.
[0074] If we denote UpperRP203 as URP, then URP(p, q) = D(LRP(p), LRP(q)). Here, 1 ≤ p and q ≤ n (where n is a natural number greater than or equal to 2). D is a function that represents distance, for example, a function that calculates the Euclidean distance between two LowerRPs (i.e., between matrices).
[0075] Furthermore, thresholding is performed on UpperRP203 to generate UpperRP204 after thresholding. If each of the n x n elements of the original data in UpperRP203 is below the threshold, its position is plotted; if it is above the threshold, its position is not plotted. This generates UpperRP204 after thresholding.
[0076] From the UpperRP204 after thresholding, the index value UpperDET (hereinafter also referred to as UD) is calculated using DET, one of the Recurrence Quantification Analysis (RQA) methods. DET is a measure that quantitatively represents determinism and is calculated using a histogram of the lengths of points drawn continuously in the upper right diagonal direction. The range of UD is from 0 to 1, and the larger the UD value, the higher the regularity and the lower the complexity. In other words, the smaller the UD value, the higher the complexity. In this embodiment, UD is used as an index value related to the complexity of sound.
[0077] Figure 4 shows the complexity index (UD) for each of the 12 types of masking sounds.
[0078] The lower the UD (Uniform Scale) value, the higher the complexity. As shown in Figure 4, the UD value for artificial sounds was higher than the threshold of 0.5. Also, the UD values for insect sounds, bird sounds, spring sounds, and pink noise were all higher than the threshold of 0.5. Furthermore, the scores for combinations of flowing water sounds and spring sounds, flowing water sounds and bird sounds, flowing water sounds and insect sounds, flowing water sounds, spring sounds and insect sounds, flowing water sounds, spring sounds and bird sounds, and flowing water sounds, spring sounds, bird sounds and pink noise were all lower than the threshold of 0.5.
[0079] Based on these results, sounds composed of multiple types of natural sounds satisfy the condition that the index value related to sound complexity is higher than a predetermined threshold.
[0080] Furthermore, it is preferable that the multiple types of natural sounds include one or more stationary sounds and one or more non-stationary sounds.
[0081] Figure 5 shows the relationship between the type of natural sound, its frequency characteristics, and information indicating whether the natural sound is a stationary or transient sound. Figure 6 shows the relationship between the frequency of bird calls and relative sound pressure level, Figure 7 shows the relationship between the frequency of insect sounds and relative sound pressure level, Figure 8 shows the relationship between the frequency of flowing water sounds and relative sound pressure level, Figure 9 shows the relationship between the frequency of spring water sounds and relative sound pressure level, and Figure 10 shows the relationship between the frequency of pink noise sounds and relative sound pressure level.
[0082] Birdsong and insect sounds have a frequency characteristic with a peak above 1 kHz and are non-stationary sounds. The sounds of flowing water and springs have a frequency characteristic as broadband sound sources and are stationary sounds. Pink noise is a broadband sound source and has a frequency characteristic as 1 / f noise and is a stationary sound.
[0083] In this embodiment, one or more stationary sounds include at least one of the sounds of flowing water and the sounds of spring water. One or more stationary sounds may be other stationary sounds other than the sounds of flowing water and spring water. For example, the stationary sounds may be the sound of wind, the sound of leaves, plants, or trees rubbing together continuously, the sound of rain, or the sound of a waterfall. In this embodiment, one or more non-stationary sounds include at least one of the sounds of birdsong and the sounds of insects. One or more non-stationary sounds may be other non-stationary sounds other than the sounds of birdsong and insects. For example, the non-stationary sounds may be the sound of leaves, plants, or trees rubbing together suddenly, animal sounds, the sound of splashing water, the sound of water droplets falling on the water surface, or the sound of waves.
[0084] Furthermore, it is preferable that the multiple types of natural sounds include multiple natural sounds that are related to each other. For example, the multiple types of natural sounds may include natural sounds related to water, such as the sound of flowing water and the sound of spring water. Alternatively, for example, the multiple types of natural sounds may include natural sounds that exist in the same natural environment, such as a forest, such as the sound of flowing water, the sound of spring water, and the calls of birds.
[0085] Figure 11 is a flowchart illustrating the sound reproduction process by the sound reproduction device 1 in the embodiment of this disclosure.
[0086] First, in step S1, the acquisition unit 12 acquires sound data, which is a combination of multiple types of natural sounds, from the sound data storage unit 11.
[0087] In this embodiment, sound data is stored in the sound data storage unit 11, but the disclosure is not limited thereto, and the communication unit of the sound playback device 1 may receive sound data from the server. The acquisition unit 12 may acquire sound data from the communication unit.
[0088] Next, in step S2, the playback unit 13 plays back the sound data acquired by the acquisition unit 12.
[0089] Next, in step S3, the output unit 14 outputs the sound data reproduced by the playback unit 13 to the speaker 2. The speaker 2 converts the sound data output by the output unit 14 into sound and outputs the converted sound to a predetermined space.
[0090] In this way, since the sound output from speaker 2 is a combination of multiple types of natural sounds, it is possible to mask noise in the space without compromising comfort within the space.
[0091] Next, I will explain the impression evaluation experiment conducted by introducing the sound playback system 100 in the office work area.
[0092] In this experiment, sound data combining multiple types of natural sounds was used, specifically a combination of flowing water sounds, spring sounds, bird songs, and pink noise. The volume of the output sound data was lower than the ambient noise level in the experimental area, and the signal-to-noise ratio was -5 dB. There were approximately 20 participants. The evaluation response periods were one week before implementation, one week immediately after implementation, one week three months after implementation, and one week five months after implementation.
[0093] Participants answered a questionnaire regarding the following evaluation items: relaxation, refreshment, quietness, sense of distraction, ease of conversation, communication, fatigue resistance, and motivation. Each item was rated on a 7-point scale.
[0094] In the relaxation assessment, participants were asked to rate their level on a scale of 1 to 2, 2 to 3, 4 to 4, 5 to 5, 6 to 6, and 7 to 7.
[0095] Furthermore, in the refreshment evaluation item, subjects gave a score of 1 if they were very unrefreshed, 2 if they were unrefreshed, 3 if they were somewhat unrefreshed, 4 if neither, 5 if they were somewhat refreshed, 6 if they were refreshed, and 7 if they were very refreshed.
[0096] In addition, for the quietness evaluation item, subjects gave a score of 1 if it was very noisy, 2 if it was noisy, 3 if it was somewhat noisy, 4 if it was neither, 5 if it was somewhat quiet, 6 if it was quiet, and 7 if it was very quiet.
[0097] Furthermore, in the evaluation of the sense of disturbance, participants gave a score of 1 if the surrounding voices were very disruptive, 2 if they were disruptive, 3 if they were somewhat disruptive, 4 if neither, 5 if they were not somewhat disruptive, 6 if they were not disruptive, and 7 if they were not very disruptive.
[0098] Furthermore, in the evaluation of ease of conversation, participants gave a score of 1 if it was very difficult to talk, 2 if it was difficult, 3 if it was somewhat difficult, 4 if it was neither, 5 if it was somewhat easy, 6 if it was easy, and 7 if it was very easy.
[0099] Furthermore, in the communication evaluation items, participants gave a score of 1 if communication was very inactive, 2 if communication was inactive, 3 if communication was somewhat inactive, 4 if neither, 5 if communication was somewhat active, 6 if communication was active, and 7 if communication was very active.
[0100] Furthermore, in the evaluation item for resistance to fatigue, subjects gave a score of 1 if they were likely to get very tired, 2 if they were likely to get tired, 3 if they were slightly likely to get tired, 4 if they were neither, 5 if they were not likely to get tired, 6 if they were not likely to get tired, and 7 if they were likely to get very little tired.
[0101] Furthermore, in the motivation evaluation items, participants were asked to rate their motivation as follows: 1 point if they were very unlikely to be motivated, 2 points if they were unlikely to be motivated, 3 points if they were somewhat unlikely to be motivated, 4 points if they were neither, 5 points if they were somewhat likely to be motivated, 6 points if they were likely to be motivated, and 7 points if they were very likely to be motivated.
[0102] Then, for each of the multiple evaluation periods, the mean and standard deviation of the scores for each of the multiple evaluation items were calculated.
[0103] Figure 12 shows the mean and standard deviation of the scores for each of the multiple evaluation items during each of the multiple evaluation periods.
[0104] In Figure 12, the square marks represent the average score before implementation, the diamond marks represent the average score immediately after implementation, the triangle marks represent the average score three months after implementation, and the circle marks represent the average score five months after implementation. The error bars represent the standard deviation of each score.
[0105] As shown in Figure 12, the average score after implementation is higher than the average score before implementation for all evaluation items. Furthermore, the average score tends to increase over time for all evaluation items.
[0106] Furthermore, even when the volume of sound data combining multiple types of natural sounds is lower than the volume of ambient noise, a sufficient masking effect is obtained, and the impression of the space (quietness and sense of disturbance), ease of resting (relaxation and refreshment), ease of conversation (ease of talking and communication), and impression of work (reduced fatigue and motivation) are improved compared to before implementation. With conventional masking, if the volume of the output masking sound is lower than the volume of ambient noise, it is not possible to suitably mask the noise. In contrast, in this embodiment, by combining multiple types of natural sounds, it was possible to sufficiently mask the noise even when the volume of the output sound is lower than the volume of ambient noise (in this experiment, the ambient noise in the experimental area).
[0107] The acquisition unit 12 may acquire the volume of ambient noise in the space where the speaker 2 is installed. The volume of ambient noise may be stored in memory beforehand. Alternatively, the acquisition unit 12 may acquire the volume of ambient noise measured by a measuring instrument installed in the space where the speaker 2 is installed. Alternatively, the acquisition unit 12 may acquire a predetermined volume input by the user. The user may input a predetermined volume measured by a measuring instrument. The playback unit 13 may play back the sound data at a volume below the predetermined volume acquired by the acquisition unit 12. The predetermined volume may be, for example, the volume of ambient noise in a predetermined space, or the volume of human conversation in a predetermined space.
[0108] Furthermore, the acquisition of ambient noise levels may be performed periodically, for example, every 5 minutes or every 30 minutes. The playback unit 13 may update the playback volume of the sound data each time the ambient noise level is acquired. In addition, when acquiring the ambient noise level in the space where the speaker 2 is installed, the average value of the ambient noise level over a predetermined period (for example, 5 minutes) may be acquired, or the instantaneous volume level of the measured ambient noise may be acquired.
[0109] Furthermore, the green view ratio of the space where speaker 2 is installed is preferably 10-15%. The green view ratio (%) indicates the proportion of green in a person's field of vision and is calculated using the formula (area of green) / (area of field of vision). By setting the green view ratio of the space to 10-15%, it is possible to create a space where a sense of harmony is always felt, and the effects obtained by outputting sounds that combine natural sounds into the space may continue for a long period of time.
[0110] Furthermore, the memory provided by the sound playback device 1 may store multiple sound data sets, each containing a combination of multiple types of natural sounds. The sound playback device 1 may further include an index value calculation unit. The index value calculation unit may calculate an index value relating to the complexity of the multiple sound data sets. The acquisition unit 12 may select the sound data set with the highest complexity index value from among the multiple sound data sets.
[0111] Furthermore, in this embodiment, the sound data storage unit 11 stores a single sound data that combines multiple types of natural sounds, but this disclosure is not particularly limited thereto. The sound data storage unit 11 may store sound data for each of the multiple types of natural sounds. The acquisition unit 12 may acquire two or more sound data from the sound data storage unit 11, and the playback unit 13 may synthesize the acquired two or more sound data and play back the synthesized sound data. The sound playback device 1 may also further include a user input receiving unit. The user input receiving unit may accept the user's selection of two or more sound data from the multiple sound data. This makes it possible to generate sound data that combines multiple types of natural sounds to suit the user's preferences.
[0112] Furthermore, some or all of the functions of the apparatus according to the embodiments of this disclosure may be realized by a processor such as a CPU executing a program.
[0113] Furthermore, all figures used above are illustrative examples provided to illustrate this disclosure, and this disclosure is not limited to these illustrative figures.
[0114] Furthermore, the order in which the steps shown in the flowchart above are performed is illustrative for the purpose of specifically illustrating this disclosure, and other orders are acceptable as long as similar effects are achieved. Also, some of the steps above may be performed simultaneously (in parallel) with other steps.
[0115] The technology disclosed herein is useful as a noise masking technology because it can mask noise in a space without impairing comfort within that space.
Claims
1. A sound playback method performed by a computer, comprising: acquiring sound data which is a combination of multiple types of natural sounds; playing back the acquired sound data; and outputting the played back sound data to a speaker.
2. The sound reproduction method according to claim 1, wherein the sound data satisfies the conditions relating to the complexity of the sound.
3. The sound reproduction method according to claim 2, wherein the condition is that the index value relating to the complexity of the sound indicates that the complexity of the sound is high.
4. The sound playback method according to any one of claims 1 to 3, wherein the playback of the sound data includes playing the sound data at a volume equal to or lower than a predetermined volume.
5. The sound reproduction method according to claim 4, wherein the predetermined volume is the volume of ambient noise in the space in which the speaker is installed.
6. The sound reproduction method according to any one of claims 1 to 3, wherein the sound data includes sound data in which the plurality of types of natural sounds and noise sounds are combined.
7. The sound reproduction method according to claim 6, wherein the noise sound is pink noise.
8. The sound reproduction method according to any one of claims 1 to 3, wherein the plurality of types of natural sounds include one or more stationary sounds and one or more non-stationary sounds.
9. The sound reproduction method according to claim 8, wherein the one or more stationary sounds include at least one of a flowing water sound indicating the sound of water flowing and a spring sound indicating the sound of water gushing out, and the one or more non-stationary sounds include at least one of a bird's song and an insect's song.
10. The sound reproduction method according to any one of claims 1 to 3, wherein the plurality of types of natural sounds include a plurality of natural sounds that are related to each other.
11. A sound reproduction device comprising a processor, wherein the processor acquires sound data comprising a combination of multiple types of natural sounds, reproduces the acquired sound data, and outputs the reproduced sound data to a speaker.
12. A sound playback program that acquires sound data composed of multiple types of natural sounds, plays back the acquired sound data, and causes a computer to function to output the played sound data to a speaker.