Sound processing method and sound processing device
The sound processing method addresses discomfort by analyzing and selectively outputting sound data to align virtual and real-world auditory inputs, enhancing immersion.
Patent Information
- Application Number
- JP2025128092
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-07-31
- Publication Date
- 2025-10-03
AI Technical Summary
Existing sound processing methods can create a sense of discomfort by mismatching virtual and real-world auditory inputs, reducing immersion.
A sound processing method that analyzes environmental sounds, compares them with pre-created sound data, and selectively outputs sound data to avoid mismatches, ensuring harmonious auditory experiences.
Enables immersive sound experiences without discomfort by adjusting virtual sounds based on real-world auditory inputs, maintaining a consistent and harmonious auditory environment.
Smart Images

Figure 2025146994000001_ABST
Abstract
Description
[Technical Field]
[0001] An embodiment of the present invention relates to a sound processing method and a sound processing device. [Background technology]
[0002] Patent Document 1 discloses an audio playback device equipped with a means for detecting the direction in which a user is facing relative to a specific area and a means for detecting the positional relationship between the specific area and the user. The audio playback device of Patent Document 1 changes the sound to be played based on the direction in which the user is facing and the positional relationship between the specific area and the user. This makes it easier for the user to imagine the impression of the specific area. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] International Publication No. 2018-088450 Summary of the Invention [Problem to be solved by the invention]
[0004] There is a need for a method that allows for more immersive listening without creating a sense of discomfort.
[0005] Therefore, an object of one embodiment of the present invention is to provide a sound processing method that allows sounds to be heard with a more immersive feeling without any sense of discomfort. [Means for solving the problem]
[0006] A sound processing method according to an embodiment of the present invention includes: Get the first sound, obtaining a second sound composed of pre-created sound data; The first sound is analyzed, performing a first comparison of comparing the analysis result of the first sound with the sound data; Based on the result of the first comparison, sound data relating to a second sound of a type that does not match the first sound is reproduced; The sound corresponding to the reproduced sound data is output.
[0007] A sound processing method according to an embodiment of the present invention includes: Get the first sound, Acquire a second sound composed of sound data created in advance by a sound content creator and a playback condition for the sound data related to the second sound; A predetermined analysis is performed on the first sound, A second comparison is made to compare the results of the first sound analysis with the playback conditions. Based on the result of the second comparison, sound data relating to a second sound of a type that satisfies the playback condition is played back; The sound corresponding to the reproduced sound data is output. [Effects of the Invention]
[0008] According to one embodiment of the present invention, sounds can be heard with a more immersive feeling without any sense of discomfort. [Brief explanation of the drawings]
[0009] [Figure 1A] FIG. 1A is a block diagram showing an example of the main configuration of a sound processing device 1 according to the first embodiment. [Figure 1B] FIG. 1B is a block diagram showing an example of the main configuration of the sound processing device 1 according to the first embodiment, and is a diagram showing an example different from FIG. 1A. [Figure 2] FIG. 2 is a flowchart showing an example of the operation of the sound processing device 1 according to the first embodiment. [Figure 3] FIG. 3 is a diagram showing an example of movement of sound data in the sound processing device 1 according to the first embodiment. [Figure 4] FIG. 4 is a diagram showing an example of analysis in the analysis unit 200a. [Figure 5] FIG. 5 is a flowchart showing an example of the operation of the sound processing device 1a according to the second embodiment. [Figure 6]FIG. 6 is a diagram showing an example of movement of sound data in the sound processing device 1a according to the second embodiment. [Figure 7] FIG. 7 is a block diagram showing an example of the main configuration of a sound processing device 1b according to the third embodiment. [Figure 8] FIG. 8 is a conceptual diagram illustrating the operation of the sound processing device 1b according to the third embodiment. [Figure 9] FIG. 9 is a diagram showing an example of movement of sound data in the sound processing device 1b according to the third embodiment. [Figure 10] FIG. 10 is a flowchart showing an example of the operation of the sound processing device 1c according to the fourth embodiment. [Figure 11] FIG. 11 is a diagram showing an example of movement of sound data in the sound processing device 1c according to the fourth embodiment. [Figure 12A] FIG. 12A is a diagram showing an example of output of a first sound and reproduction of sound data. [Figure 12B] FIG. 12B is a diagram showing an example of a case where the sound processing devices 1, 1a, 1b, and 1c do not play back sound data. [Figure 12C] FIG. 12C is a diagram showing an example of eliminating the sound of a sound source of a type that does not match the second sound, among the first sounds. DETAILED DESCRIPTION OF THE INVENTION
[0010] (First embodiment) The configuration of the sound processing device 1 according to the first embodiment will be described below with reference to the drawings. FIG. 1A is a block diagram showing an example of the main configuration of the sound processing device 1 according to the first embodiment. FIG. 1B is a block diagram showing an example of the main configuration of the sound processing device 1 according to the first embodiment, and is a diagram showing an example different from FIG. 1A. FIG. 2 is a flowchart showing an example of the operation of the sound processing device 1 according to the first embodiment. FIG. 3 is a diagram showing an example of sound data movement in the sound processing device 1 according to the first embodiment. FIG. 4 is a diagram showing an example of analysis in the analysis unit 200a.
[0011] As shown in FIG. 1A, the sound processing device 1 includes a terminal 20 and headphones 30. The terminal 20 includes a microphone 10, a CPU 200, a ROM 201, a RAM 202, and an output I / F 203. The terminal 20 and the headphones 30 are connected to each other via a wired or wireless connection. Note that, as shown in FIG. 1B, the headphones 30 may include the microphone 10. In other words, the headphones 30 may be headphones with a microphone.
[0012] The microphone 10 acquires environmental sounds around the location where the microphone 10 is installed (in other words, environmental sounds around the user). The microphone 10 converts the acquired environmental sounds into sound signals. The microphone 10 outputs the sound signals obtained by the conversion to the CPU 200 of the terminal 20. Examples of environmental sounds include the sound of a car engine and the sound of thunder. The environmental sounds around the location where the microphone 10 is installed correspond to the first sound. Furthermore, the microphone 10 corresponds to the first sound acquisition unit in the present invention.
[0013] The terminal 20 stores sound data created in advance on another PC or the like by a content creator (hereinafter referred to as a creator). The terminal 20 is, for example, a mobile device such as a smartphone. In this case, the microphone 10 is a built-in microphone provided in the smartphone or the like. The terminal 20 corresponds to a second sound acquisition unit in the present invention.
[0014] Sound data is data that records a specific sound. Examples of specific sounds include the sound of waves and the sound of cicadas. That is, the sound data includes data to which sound source information (second sound source information in this embodiment) indicating the type of sound source (second sound source in this embodiment) is added in advance. The terminal 20 stores the sound data as multi-track content data having tracks. For example, the terminal 20 stores multi-track content data having two tracks: sound data of the sound of waves and sound data of cicadas. Sound data (content data) created by a creator corresponds to the second sound in this embodiment. That is, the second sound is composed of sound data. Hereinafter, sound data set via the terminal 20 will be referred to as sound data related to the second sound.
[0015] Creators create content to give users a particular impression. For example, if a creator wants to create content that gives users the impression of summer, the creator creates sound data of summer-related sounds such as the sound of waves and the chirping of cicadas.
[0016] The ROM 201 stores various data, such as a program for operating the terminal 20, environmental sound data input from the microphone 10, and content data received from other PCs.
[0017] The RAM 202 temporarily stores the predetermined data stored in the ROM 201 .
[0018] The CPU 200 controls the operation of the terminal 20. The CPU 200 performs various operations by reading out predetermined programs stored in the ROM 201 into the RAM 202. The CPU 200 includes an analysis unit 200a, a comparison unit 200b, and a playback unit 200c. The CPU 200 performs various processes on the input first sound (environmental sound). The various processes include analysis processing by the analysis unit 200a, comparison processing by the comparison unit 200b, and playback processing by the playback unit 200c in the CPU 200. In other words, the CPU 200 executes programs including an analysis processing program by the analysis unit 200a, a comparison processing program by the comparison unit 200b, and a playback processing program by the playback unit 200c.
[0019] The analysis unit 200a performs a predetermined analysis process on data related to the environmental sound. In other words, the analysis unit 200a analyzes the first sound. The predetermined analysis process in the analysis unit 200a is, for example, sound source recognition processing using artificial intelligence such as a neural network. In this case, the analysis unit 200a calculates feature quantities of the environmental sound based on the input data related to the environmental sound. The feature quantities are parameters that indicate the characteristics of the sound source. For example, the feature quantities include at least power or cepstrum coefficients. Power is the power of the sound signal. The cepstrum coefficient is the logarithm of the amplitude of the discrete cosine transform of the sound signal on the frequency axis. Note that the feature quantities of the sound are not limited to only power and cepstrum coefficients.
[0020] The analysis unit 200a recognizes the sound source (estimates the type of sound source) based on the feature amount of the sound source. For example, if the environmental sound contains the feature amount of the chirping of cicadas, the analysis unit 200a recognizes the type of sound source as the chirping of cicadas. As shown in FIG. 3, the analysis unit 200a outputs the analysis result D (sound source recognition result) to the comparison unit 200b. For example, if the analysis unit 200a recognizes that "the sound source is the chirping of cicadas," the analysis unit 200a outputs the analysis result D that "the sound source is the chirping of cicadas" to the comparison unit 200b. In other words, the analysis result D includes sound source information (first sound source information in this embodiment) that indicates the type of sound source (first sound source in this embodiment) included in the environmental sound, which is the first sound.
[0021] Here, a detailed description will be given of a case where the analysis unit 200a performs sound source recognition processing using a neural network. The following description will be given taking as an example a case where the analysis unit 200a uses a neural network NN1 as shown in FIG.
[0022] The terminal 20 has a trained neural network NN1 that outputs the type of sound source when a feature of the sound source is input. As shown in FIG. 4, the neural network NN1 recognizes a sound source based on the feature of a plurality of sounds. As shown in FIG. 4, the sound feature used by the neural network NN1 for the sound source recognition process is, for example, power P1 and cepstrum coefficient P2. The neural network NN1 outputs the degree of match between the feature of each sound source and various feature of the environmental sound. Then, the neural network NN1 outputs the type of sound source with the highest degree of match as the analysis result D.
[0023] More specifically, first, a trained model of the neural network NN1 is prepared, which has been trained (for example, parameter tuning such as weighting in artificial intelligence has been completed) using information indicating the type of sound source (hereinafter referred to as third sound source information) and a dataset indicating the relationship between the feature quantities of the third sound source as training data. Then, the neural network NN1 inputs the feature quantities contained in the first sound calculated in the analysis of the first sound into the trained model. After inputting the feature quantities into the trained model, the neural network NN1 outputs information on the type of sound source corresponding to the input feature quantities as an analysis result. For example, the neural network NN1 outputs the degree of match between the environmental sound and the trained sound source based on the input feature quantities of the environmental sound. Then, the neural network NN1 outputs information on the type of sound source with the highest degree of match (for example, label information of cicada) among the trained sound sources to the comparison unit 200b.
[0024] For example, as shown in FIG. 4, the neural network NN1 calculates the degree of match between environmental sounds and cicadas and the degree of match between environmental sounds and car engine sounds. In the example shown in FIG. 4, the neural network NN1 calculates that the environmental sounds and cicadas match with a 60% probability, and that the environmental sounds and car engine sounds match with a 30% probability. At this time, the neural network NN1 also calculates the probability that the environmental sounds do not match any of the recognition data. In the example shown in FIG. 4, the neural network NN1 calculates that the probability that the environmental sounds do not match either cicadas or car engine sounds is 10%. In the above calculation results, the sound source type with the highest degree of match is cicadas. Therefore, the neural network NN1 outputs an analysis result D that states, "The sound source is cicadas." In this way, the neural network NN1 can estimate the type of sound source that matches the environmental sounds (recognize the sound source) when the features of the sound source are input. The type of sound source to be output is specified in advance by the creator, etc.
[0025] The method for recognizing the type of sound source is not limited to a method using a neural network. For example, the analysis unit 200a may perform matching by comparing the waveforms of sound signals. In this case, waveform data (template data) for each type of sound source is pre-recorded in the terminal 20 as recognition data. The analysis unit 200a then determines whether the waveform of the environmental sound matches the template data. If the analysis unit 200a determines that the waveform of the template data matches the waveform of the environmental sound, it recognizes that the environmental sound is a sound source of the type of template data. For example, if the waveform of the chirping of cicadas is recorded in the terminal 20 as template data, the analysis unit 200a determines whether the waveform of the environmental sound matches the waveform of the chirping of cicadas. If it determines that they match, the analysis unit 200a outputs an analysis result D indicating that the environmental sound is the chirping of cicadas. The matching in the analysis unit 200a is not limited to a method of comparing waveform data. For example, the analysis unit 200a may perform matching by comparing feature quantities of sound sources. In this case, the features of the sound source (power, cepstrum coefficients, etc.) are pre-recorded as recognition data in the terminal 20. Then, the analysis unit 200a determines whether the features of the environmental sound match the features of the sound source.
[0026] As shown in FIG. 3, the comparison unit 200b performs a comparison process between the analysis result D and sound data relating to the second sound. Information indicating the type of sound source (for example, information that the sound data is the chirping of cicadas) is added to the sound data relating to the second sound. As shown in FIG. 3, the comparison unit 200b compares the analysis result D of the analysis unit 200a with the information indicating the type of sound source added to each piece of sound data. Then, if the analysis result D matches the sound data relating to the second sound (specifically, if the environmental sound and the sound data relating to the second sound are the same type of sound source), the comparison unit 200b excludes the sound data relating to the second sound that matches the environmental sound from the sound to be reproduced. For example, if the analysis result D indicates that "the environmental sound is the chirping of cicadas" and the sound data is the chirping of cicadas, the comparison unit 200b excludes the sound data of the chirping of cicadas from the sound to be reproduced. The comparison unit 200b then outputs the sound data other than the excluded sound data to the playback unit 200c. In other words, sound data related to the second sound that matches the analysis result D is not output to the playback unit 200c. Note that exclusion in the comparison unit 200b means distinguishing between sound data to be output to the playback unit 200c and sound data that is not to be output. Therefore, exclusion does not mean deleting sound data from the terminal 20. Note that, in the present invention, sound data related to a second sound of a type that matches the first sound refers to second sound data that is excluded from playback targets as a result of the comparison process. Note that sound data related to a second sound of a type that does not match the first sound refers to second sound data that is not excluded from playback targets as a result of the comparison process. In other words, sound data related to the second sound is divided into data related to a second sound of a type that matches the first sound and sound data related to a second sound of a type that does not match the first sound.
[0027] In this embodiment, "a match between the sound data relating to the first sound and the second sound" means that the information on the type of the environmental sound output by the analysis unit 200a matches the information on the type of the sound data. For example, if the analysis unit 200a recognizes that the type information of the environmental sound is "the chirping of cicadas" using a neural network or the like, the type information of the first sound will be "the chirping of cicadas." In this case, if the type information of the sound data relating to the second sound is recorded as "the chirping of cicadas," it is estimated that the sound data relating to the first sound and the second sound match. However, the comparison unit 200b may determine that the sound data relating to the first sound and the second sound match if at least some of the feature amounts match.
[0028] The reproduction unit 200c receives the sound data relating to the second sound from the comparison unit 200b. The reproduction unit 200c decodes the sound data relating to the second sound. The reproduction unit 200c then outputs the decoded sound data relating to the second sound as a sound signal to the output I / F 203. The output I / F 203 is, for example, an audio terminal, a USB terminal, a communication I / F, etc. The output I / F 203 corresponds to the output unit of the present invention. The output I / F 203 receives the sound signal and outputs the sound signal to the headphones 30.
[0029] The headphones 30 output the sound signal input from the output I / F 203 as sound. The headphones 30 are, for example, headphones owned by the user. The user hears a sound based on a sound signal relating to the second sound through the headphones 30. Note that the headphones 30 in this embodiment are devices that emit sound from a sound generator (such as a speaker) located close to a person's ear. Therefore, in this embodiment, the headphones 30 include devices such as bone-powered earphones and shoulder speakers.
[0030] A series of processes of the sound processing device 1 will be described below with reference to Figs. 2 and 3. Fig. 2 is a flowchart showing an example of the operation of the sound processing device 1. Fig. 3 is a diagram showing the movement of sound data in the sound processing device 1. In Fig. 3, excluded sound data is indicated by a dotted square. In Fig. 3, the comparison between the environmental sound data and the sound data relating to the second sound in the comparison unit 200b is indicated by a double-headed arrow. In Fig. 3, the RAM 202 and the output I / F 203 are not shown.
[0031] First, the microphone 10 acquires an environmental sound (first sound) around the user (FIG. 2: S10). The microphone 10 converts the acquired environmental sound into a sound signal. The microphone 10 outputs the sound signal obtained by the conversion to the analysis unit 200a of the terminal 20.
[0032] Next, the terminal 20 acquires content data (second sound) composed of sound data (FIG. 2: S11). The acquired content data is stored in the ROM 201. In the example shown in FIG. 3, sound data A, sound data B, and sound data C are stored in the ROM 201 as content data.
[0033] Next, the analysis unit 200a performs a predetermined analysis on the acquired environmental sound (FIG. 2: S12). The analysis result of the environmental sound is output to the comparison unit 200b. In the example shown in FIG. 3, the analysis unit 200a outputs data called analysis result D as the analysis result of the environmental sound to the comparison unit 200b.
[0034] Next, the comparison unit 200b reads the content data from the ROM 201. In the example shown in FIG.
[0035] Next, the comparison unit 200b compares the analysis result of the environmental sound related to the first sound with the content data (sound data related to the second sound) (FIG. 2: S13). In the example shown in FIG. 3, each of sound data A, sound data B, and sound data C is compared with analysis result D.
[0036] Next, if the analysis result of the environmental sound matches the sound data related to the second sound (FIG. 2: S13 Yes), the comparison unit 200b excludes the sound data related to the second sound that matches the analysis result of the environmental sound from the content data (FIG. 2: S14). In other words, the comparison unit 200b selects sound data related to the second sound other than the sound data related to the second sound that matches the analysis result of the environmental sound. In the example shown in FIG. 3, the analysis result D matches the sound data B, so the comparison unit 200b excludes the sound data B from the content data.
[0037] If the comparison result shows that there is no sound data related to the second sound that matches the analysis result of the environmental sound (FIG. 2: S13 No), the second sound is not excluded from the content data. The comparison by the comparing unit 200b between the data of the environmental sound (the analysis result of the first sound) and the sound data related to the second sound corresponds to the first comparison in the present invention.
[0038] Next, the comparison unit 200b selects sound data relating to the second sound other than the excluded sound data (FIG. 2: S15). If there is no sound data to exclude (FIG. 2: S13 No), all sound data relating to the second sound is selected (FIG. 2: S16). In the example shown in FIG. 3, sound data A and sound data C are selected by the comparison unit 200b.
[0039] Next, the comparison unit 200b outputs the sound data relating to the selected second sound to the reproduction unit 200c (FIG. 2: S17). In the example shown in FIG. 3, the comparison unit 200b outputs sound data A and sound data C (content data excluding sound data B) to the reproduction unit 200c.
[0040] Next, the playback unit 200c decodes (plays back) the content data input from the comparison unit 200b and outputs it to the output I / F 203 as an audio signal. The output I / F 203, which has received the audio signal, outputs the audio signal to the headphones 30. The headphones 30, which has received the audio signal, output the input audio signal as sound. In the example shown in FIG. 3, the playback unit 200c decodes the audio data A and audio data C into audio signals A2 and C2, respectively. The playback unit 200c then outputs the decoded audio signals A2 and C2 to the headphones 30 via the output I / F 203. In other words, the playback unit 200c plays back audio data (sound data related to the second sound) of a type that does not match the first sound (ambient sound) based on the result of the first comparison.
[0041] Finally, the headphones 30, to which the sound signals A2 and C2 have been input, output a sound A3 based on the sound signal A2 and a sound C3 based on the sound signal C2. In other words, the headphones 30 output the reproduced sound data (sound data related to the second sound).
[0042] The sound processing device 1 repeats the operations from S10 to S17. Therefore, when an environmental sound that matches the second sound is being generated, the headphones 30 do not output the second sound that matches the environmental sound. When an environmental sound that matches the second sound is not being generated, the headphones 30 output the second sound. In this way, the sound processing device 1 can switch whether or not to output the second sound in accordance with changes in the environmental sound.
[0043] With the above configuration, the sound processing device 1 is capable of processing sound so that the sound can be heard without any sense of incongruity and with a greater sense of immersion. Below, the sound processing device 1 will be described in comparison with a sound processing device that does not perform sound processing according to this embodiment (hereinafter referred to as Comparative Example 1). In comparing the sound processing device 1 with Comparative Example 1, an example will be described in which there is one river around the user. In other words, this is a case in which the environmental sounds around the user include one river sound, and the river sound intrudes from outside the headphones 30 (headphones in the case of Comparative Example 1). That is, the user hears the sound output from the headphones 30 (headphones in the case of Comparative Example 2) and the sound of the river, which is the surrounding environmental sound.
[0044] In Comparative Example 1, whether or not to output sound from the headphones is not switched in response to changes in the surrounding environmental sounds. Therefore, when the sound of a river is output from the headphones, the user hears both the sound of the river output from the headphones (the sound of the river in the virtual space) and the sound of the river entering from outside the headphones (the sound of the river). In other words, in Comparative Example 1, the user hears a combination of sounds from the virtual space and sounds from the real space. Meanwhile, the user only sees one river. That is, in the user's perception, a mismatch occurs between visual information (the user sees one river) and auditory information (the user hears the sounds of two rivers). Therefore, Comparative Example 1 may cause the user to feel uncomfortable. As a result, the user's sense of immersion may be reduced.
[0045] On the other hand, the sound processing device 1 in this embodiment switches whether or not to output sound from the headphones 30 depending on changes in the surrounding environmental sounds. Therefore, if the surrounding environmental sounds include the sound of a river, the sound processing device 1 does not play back sound data including the sound of the river. Therefore, the user can hear the sound of the river entering from outside the headphones 30, but cannot hear the sound of the river output from the headphones 30. In other words, there is no mismatch in the user's perception between visual information (the user sees one river) and auditory information (the user hears the sounds of two rivers). Therefore, the sound processing device 1 is unlikely to cause the user discomfort. As a result, it is possible to prevent a decrease in the user's sense of immersion.
[0046] (Second embodiment) The configuration of the sound processing device 1a according to the second embodiment will be described below with reference to the drawings. Fig. 5 is a flowchart showing the operation of the sound processing device 1a according to the second embodiment. Fig. 6 is a diagram showing an example of sound data movement in the sound processing device 1a according to the second embodiment.
[0047] As shown in FIG. 6, the sound processing device 1a differs from the sound processing device 1 in that it determines whether or not to reproduce the second sound based on reproduction conditions created by a creator.
[0048] The playback conditions are data that record playback conditions for the second sound. Specifically, the playback conditions set whether or not overlapping playback of the environmental sound and the second sound is permitted. For example, if the playback conditions set "Playback condition: overlapping playback permitted" for the sound data related to the second sound, the sound processing device 1a outputs the sound data related to the second sound regardless of the comparison result in the comparison unit 200b. On the other hand, if the playback conditions set "Playback condition: overlapping playback not permitted" for the sound data related to the second sound, the sound processing device 1a does not output sound data of the same type as the environmental sound. The playback conditions are acquired via the terminal 20 in the same way as acquiring the second sound. After the playback conditions are acquired, the playback conditions are stored in the ROM 201.
[0049] A series of operations of the sound processing device 1a will be described below. In the example shown in Fig. 6, sound data B is the same type of sound as sound source d (one of the sound sources included in the environmental sound). In the example shown in Fig. 6, sound data C is a different type of sound from each of sound source d and sound source e. Note that the processes of S11, S12, S15, S16, and S17 are the same as those of the sound processing device 1, and therefore descriptions thereof will be omitted.
[0050] After the analysis unit 200a performs a predetermined analysis of the first sound (after S12 in FIG. 5), the comparison unit 200b acquires the playback conditions created in advance by the creator (S20 in FIG. 5). In the example shown in FIG. 6, the comparison unit 200b acquires the playback conditions (sound data A: overlapping playback permitted, sound data B: overlapping playback not permitted, sound data C: overlapping playback permitted) from the ROM 201.
[0051] Next, the comparison unit 200b compares whether the analysis result D matches the playback conditions (FIG. 5: S21). Specifically, as shown in FIG. 6, the comparison unit 200b compares whether the sound sources d and e included in the analysis result D match the sound data A, B, and C included in the playback conditions. For example, if the comparison unit 200b receives the analysis result D from the analysis unit 200a, which states that "the environmental sound is the sound of ocean waves," and if the data included in the playback conditions is set to "sound of ocean waves," the comparison unit 200b determines that the analysis result D of the environmental sound matches the playback conditions. The comparison of the first sound analysis result by the comparison unit 200b with the playback conditions corresponds to the second comparison in the present invention. Note that a match between the analysis result D and the playback conditions means, for example, that the information on the type of environmental sound output by the analysis unit 200a matches the information on the type of sound data.
[0052] If the environmental sound analysis result D matches the playback conditions (FIG. 5: S21 Yes), the comparison unit 200b determines whether overlapping playback of the environmental sound analysis result data and sound data related to the second sound is permitted (FIG. 5: S22). For example, if the environmental sound analysis result data matches the playback conditions for the sound of waves, the comparison unit 200b determines whether overlapping playback of the sound of waves is permitted based on the playback conditions. If the environmental sound analysis result data does not match the playback conditions (FIG. 5: S21 No), the comparison unit 200b selects all sound data related to the second sound (FIG. 5: S16).
[0053] If the comparison unit 200b determines based on the playback conditions that overlapping playback of the sound data is not permitted (FIG. 5: S22 Yes), the comparison unit 200b excludes the sound data determined to be overlapping playback not permitted (sound data of a type that does not satisfy the playback conditions) from the content data (FIG. 5: S23). For example, if the sound of waves matches the environmental sound data and the playback conditions, the comparison unit 200b excludes the sound data of the sound of waves from the content data. In the example shown in FIG. 6, sound data B matches sound source d. Therefore, the comparison unit 200b excludes sound data B from the content data. Next, the comparison unit 200b selects sound data other than the excluded sound data (sound data of a type that satisfies the playback conditions) (FIG. 5: S15). If the comparison unit 200b determines based on the playback conditions that overlapping playback of the sound is permitted (FIG. 5: S22 No), the comparison unit 200b selects sound data related to the second sound (FIG. 5: S16).
[0054] Finally, the comparison unit 200b outputs the selected sound data to the reproduction unit 200c (FIG. 5: S17). Note that the processing after the comparison unit 200b outputs the selected sound data to the reproduction unit 200c is the same as that of the sound processing device 1, and therefore the description thereof will be omitted.
[0055] As a result, the sound processing device 1a determines whether or not to play back sound data relating to the second sound based on the playback conditions. In the example shown in Fig. 6, sound data B, for which duplicate playback is not permitted, is not played back. Therefore, as shown in Fig. 6, if a creator does not want a specific sound to be played back repeatedly, the creator can create a playback condition to prevent the sound processing device 1a from playing back sound data containing the specific sound repeatedly.
[0056] With the above configuration, the sound processing device 1a enables sound processing that allows sounds to be heard with a more immersive feeling without creating a sense of incongruity. Specifically, creators can reproduce sounds that would be out of place if they overlap without overlapping. Below, we will explain an example in which a creator creates sound data of the sound of waves and sound data of the chirping of cicadas as content data, and the sound of waves and the chirping of cicadas are included as environmental sounds.
[0057] In this case, by setting the playback conditions, the creator can set the playback not to play sounds that are considered problematic (unnatural) when played overlapping. Furthermore, the creator can set the playback to play sounds that are considered unproblematic (unnatural) when played overlapping. In other words, the creator can choose whether to use sounds from the real space or the virtual space. For example, if the creator determines that the sound of waves will cause discomfort to the user when heard overlapping, the creator sets the sound data of the sound of waves as a playback condition: overlap not permitted. Furthermore, if the creator determines that the sound of cicadas will not cause discomfort to the user when heard overlapping, the creator sets the sound data of cicadas as a playback condition: overlap permitted. In this case, the user can hear the sound of waves without overlapping in the real space, and can also hear multiple cicadas (cicadas in the real space and cicadas in the virtual space). In other words, the sound processing device 1a can use sounds present at the sound playback location and can supplement sounds that are considered to be lacking with sounds from the virtual space. This allows the sound processing device 1a to provide users with content as intended by creators. Therefore, the sound processing device 1a is unlikely to cause users to feel uncomfortable. As a result, it is possible to prevent a decrease in the user's sense of immersion.
[0058] (Third embodiment) The configuration of a sound processing device 1b according to the third embodiment will be described below with reference to the drawings. Fig. 7 is a block diagram showing the configuration of the sound processing device 1b according to the third embodiment. Fig. 8 is a flowchart showing the operation of the sound processing device 1b according to the third embodiment. Fig. 9 is a diagram showing an example of sound data movement in the sound processing device 1b according to the third embodiment.
[0059] 7 and 9, the CPU 200 of the sound processing device 1b differs from the CPU 200 of the sound processing device 1 in that it includes an external environment data acquisition unit 200d. Also, as shown in Fig. 8, the sound processing device 1b differs from the sound processing device 1 in that it compares the acquired external environment data with sound data related to a second sound, and selects a sound related to the second sound data in accordance with the external environment data.
[0060] The external environment data acquisition unit 200d acquires data (hereinafter referred to as external environment data) on information about the environment around the terminal 20 (the environment around the user). As shown in FIG. 7, the external environment data is acquired by the sensor 40a. The external environment data acquisition unit 200d acquires the external environment data from the sensor 40a. The sensor 40a is, for example, a thermometer (temperature data), an illuminance meter (illuminance data), a hygrometer (humidity data), or a GPS (latitude and longitude data). In other words, the external environment data includes information other than sound. The external environment data acquisition unit 200d corresponds to the environment data acquisition unit in the present invention.
[0061] 7, the external environment data acquiring unit 200d may acquire external environment data via a server 40b connected to a network. In this case, the external environment data acquiring unit 200d acquires, for example, weather information (temperature data, humidity data, etc.) or map information (latitude and longitude data) from the server 40b. The network may specifically be a LAN (Local Area Network), WAN (Wide Area Network), etc.
[0062] Note that the source of external environment data obtained via a network is not limited to the server 40b. Specifically, the external environment data obtaining unit 200d may obtain external environment data from a sensor connected via a network. For example, the source of external environment data is a terminal 20 installed indoors and a thermometer (an example of a sensor) installed outdoors. In this case, the thermometer transmits the obtained data to the terminal 20 via a wireless LAN.
[0063] The comparison unit 200b of the sound processing device 1b compares the acquired external environment data with sound data related to the second sound. Specifically, the sound processing device 1b pre-stores output conditions (hereinafter referred to as conditions between the external environment and sound data) that change the second sound to be output in accordance with the external environment. Then, when the external environment data satisfies the conditions between the external environment and sound data, the sound processing device 1b outputs the sound data. For example, if an air temperature of 25 degrees or higher is set as a condition between the external environment and sound data for the sound of cicadas, the sound processing device 1b outputs the sound of cicadas when an air temperature of 25 degrees or higher is acquired from the external environment data acquisition unit 200d (thermometer).
[0064] A series of operations of the sound processing device 1b will be described below. Note that the processes from S11 to S17 are the same as those of the sound processing device 1, and therefore descriptions thereof will be omitted.
[0065] After selecting sound data related to the second sound (after S15 or S16 in FIG. 8), the external environment data acquisition unit 200d acquires external environment data (S30 in FIG. 8). In the example shown in FIG. 9, the external environment data acquisition unit 200d acquires external environment data from the sensor 40a and the server 40b. The external environment data acquisition unit 200d outputs the acquired external environment data to the comparison unit 200b. In the example shown in FIG. 9, the external environment data acquisition unit 200d outputs external environment data X and external environment data Y to the comparison unit 200b.
[0066] Next, the comparison unit 200b compares the external environment data with the condition between the external environment and the sound data (FIG. 8: S31). For example, in the example shown in FIG. 9, if the sound data A is sound data of cicadas, the creator sets the season as summer in the condition between the external environment sound data. Then, the sound processing device 1b determines whether the season is summer based on information acquired from the external environment data acquisition unit 200d (specifically, calendar information of the server, etc.).
[0067] If the external environment data matches the condition between the external environment and the sound data (FIG. 8: S31 Yes), the comparison unit 200b selects the sound data corresponding to the external environment data (FIG. 8: S32). For example, in FIG. 9, if data indicating that the season is summer is acquired as the external environment data X and the condition between the external environment sound data of the sound data A is set to season: summer, the comparison unit 200b selects the sound data A.
[0068] On the other hand, if the external environment data and the condition between the external environment and the sound data do not match (FIG. 8: S31 No), the comparison unit 200b does not select the sound data corresponding to the external environment data (FIG. 8: S33). For example, in FIG. 9, if data indicating a temperature of 25 degrees is acquired as the external environment data Y and the condition between the external environment sound data of the sound data C is set to a temperature of 15 degrees or less, the comparison unit 200b does not select the sound data C.
[0069] Next, the comparison unit 200b outputs the selected sound data to the reproduction unit 200c (FIG. 8: S17). In the example shown in FIG. 9, the comparison unit 200b outputs sound data A to the reproduction unit 200c. Note that the processing after the comparison unit 200b outputs the selected sound data to the reproduction unit 200c is the same as that of the sound processing device 1, and therefore a description thereof will be omitted.
[0070] With the above configuration, the sound processing device 1b can process sounds that can be heard with a more immersive feeling without creating a sense of incongruity. Specifically, the sound processing device 1b can switch whether or not to output sound data in response to changes in the external environment. This reduces the possibility of outputting sound data that does not harmonize with the external environment. An example will be described below in which the sound data includes river sound data. In this case, the comparison unit 200b acquires map information about the area around the terminal 20 from the external environment data acquisition unit 200d to determine whether or not there is a river around the terminal 20 (whether or not there is a river in the acquired map). If there is a river in the map, the sound processing device 1b determines that there is a river near the user. Then, the sound processing device 1b does not output sound data of the river sound to avoid overlapping the river sound. Furthermore, if the acquired map information changes from a state in which there is a river to a state in which there is no river due to the user's movement, the sound processing device 1b determines that there is no river near the user. Then, the sound processing device 1b outputs sound data of the river sound to prevent a lack of river sounds. Therefore, the sound processing device 1b makes it possible for the user to hear the necessary sounds from the virtual space and the real space without excess or deficiency. Therefore, the sound processing device 1b is even less likely to make the user feel uncomfortable. As a result, it is possible to further prevent the user's sense of immersion from decreasing.
[0071] (Fourth embodiment) The configuration of the sound processing device 1c according to the fourth embodiment will be described below with reference to the drawings. Fig. 10 is a flowchart showing an example of the operation of the sound processing device 1c according to the fourth embodiment, and Fig. 11 is a diagram showing an example of the movement of sound data in the sound processing device 1c according to the fourth embodiment.
[0072] As shown in Fig. 11, the CPU 200 of the sound processing device 1c differs from the CPU 200 of the sound processing device 1 in that it includes a specific sound elimination unit 200e. Also, as shown in Fig. 10, the sound processing device 1c differs from the sound processing device 1 in that it acquires elimination conditions for environmental sound data. Also, the sound processing device 1c differs from the sound processing device 1 in that it compares whether there is any environmental sound that matches the elimination conditions. Note that in Fig. 11, sound sources included in environmental sounds that match the elimination conditions are circled.
[0073] When a specific sound is included in the environmental sound, the specific sound elimination unit 200e eliminates the specific sound. For example, the specific sound is the sound of a car engine. That is, when a specific sound (for example, the sound of a car engine) is included in the sound entering from outside the headphones 30, the sound processing device 1c eliminates the specific sound that has entered from outside. For example, when the sound of a car engine is set as the specific sound to be eliminated, the sound processing device 1c eliminates the sound of the car engine. The specific sound is eliminated, for example, by outputting a sound having an opposite phase to the specific sound from the headphones 30.
[0074] The ROM 201 of the sound processing device 1c stores elimination conditions that set conditions for elimination of specific sounds. For example, if the elimination condition is set to be a car engine sound, the sound processing device 1c performs an operation to eliminate the car engine sound as the specific sound to be eliminated. The elimination conditions are stored in the terminal 20 in advance.
[0075] A series of operations of the sound processing device 1c will be described below. Note that the processes from S11 to S16 are the same as those of the sound processing device 1, and therefore descriptions thereof will be omitted.
[0076] After selecting the sound data relating to the second sound (after S15 or S16 in FIG. 10), the specific sound elimination unit 200e acquires the elimination condition (S40 in FIG. 10). In the example shown in FIG. 9, the specific sound elimination unit 200e acquires the elimination condition from the ROM 201.
[0077] Next, the specific sound elimination unit 200e compares whether there are any environmental sounds that match the elimination conditions (whether there are any overlapping sounds) (FIG. 10: S41). In the example shown in FIG. 11, the specific sound elimination unit 200e compares each of the sound sources d and e included in the analysis result D with the elimination conditions.
[0078] If there is a sound source that meets the elimination condition (FIG. 10: S41 Yes), the specific sound elimination unit 200e creates cancellation data for erasing the sound source that meets the elimination condition (S42). In the example shown in FIG. 11, the specific sound elimination unit 200e creates cancellation data CD based on the sound source d that meets the elimination condition. If there is no environmental sound data that meets the elimination condition (FIG. 10: S41 No), the specific sound elimination unit 200e does not create cancellation data.
[0079] Next, the specific sound elimination unit 200e outputs the cancellation data to the playback unit 200c. In the example shown in Fig. 11, the specific sound elimination unit 200e outputs the cancellation data CD to the playback unit 200c.
[0080] Next, the playback unit 200c outputs the sound data relating to the second sound input from the comparison unit 200b and the cancellation data CD input from the specific sound elimination unit 200e as sound signals to the headphones 30 (FIG. 10: S43). In the example shown in FIG. 11, the playback unit 200c outputs the sound data A and sound data C (input from the comparison unit 200b) as sound signals A2 and C2, respectively, and outputs the cancellation data CD (input from the specific sound elimination unit 200e) as cancellation signal CD2 to the headphones 30.
[0081] Finally, the headphones 30 output a sound A3 based on the sound signal A2, a sound C3 based on the sound signal C2, and a cancellation sound CD3 based on the cancellation signal CD2.
[0082] With the above configuration, the sound processing device 1c enables sound processing that allows the user to hear sounds with a more immersive feeling without any sense of incongruity. Specifically, when external environmental sounds include noise, the sound processing device 1c can eliminate the noise. For example, a creator sets a car engine sound (one example of a noise sound) in the sound processing device 1c as a specific sound to be eliminated. In this case, when the sound processing device 1c determines that the car engine sound is included in the external environmental sounds, it eliminates the car engine sound. Therefore, the user can experience the content without the noisy car engine sound. In this way, the sound processing device 1c prevents the user from having their immersion impaired by noise. Therefore, the sound processing device 1c is even less likely to cause the user to feel uncomfortable. As a result, it is possible to further prevent the user's immersion from being reduced.
[0083] Alternatively, the elimination conditions may be created in advance by the content creator. In this case, the elimination conditions created by the creator are stored in the ROM 201. The specific sound elimination unit 200e then eliminates specific sounds from the environmental sounds based on the elimination conditions created by the creator. In this case, the user cannot hear environmental sounds that the creator did not intend. Therefore, the user can listen to the sounds with a more immersive feeling without feeling uncomfortable.
[0084] (Variation) Modifications will be described below. By using the sound processing devices 1, 1a, 1b, and 1c according to the modifications, it is possible to record the sound of a sound source at a travel destination (hereinafter referred to as the local area) and bring back content based on the sound of the recorded sound source. For example, if a user travels to a specific location (e.g., Waikiki Beach in Hawaii) while listening to specific content (e.g., tropical sound content), the sound processing devices 1, 1a, 1b, and 1c record the sound of waves at Waikiki Beach. The next time the sound processing device 1 plays back the same tropical sound content, it plays back the recorded sound data of waves at Waikiki Beach instead of the pre-recorded wave sound data. In this way, the sound processing devices 1, 1a, 1b, and 1c can switch the sound to be played. This allows the sound processing devices 1, 1a, 1b, and 1c to motivate the user to go to a specific location.
[0085] (Other variations) The terminal 20 (second sound acquisition unit) may further acquire localization processing data used in processing the sound image localization related to the second sound. The localization processing data is, for example, information on the positional relationship between the sound source and the user in a virtual space (three-dimensional space). This makes it possible to perform sound image localization processing that localizes the sound at a predetermined position intended by the creator. For example, if a creator wants to localize the sound of a river to the right of the user's position, the creator sets the position information of the sound data of the river sound to the right of the user. In this case, the user can hear the sound of the river as if the river were located to the user's right. This allows the user to naturally recognize the direction of surrounding objects, etc. Therefore, the user can hear the sound with a greater sense of immersion without feeling uncomfortable.
[0086] If the second sound is multi-track, the terminal 20 may acquire a track (sound data) switching condition. The switching condition is set in advance by a creator via the terminal 20. In this case, the sound processing devices 1, 1a, 1b, and 1c play back the sound data of the track specified by the switching condition. Switching based on the switching condition means, for example, when a specific sound is included in the environmental sound, switching of sound data is triggered by the specific sound. Below, an example will be described in which the sound processing devices 1, 1a, 1b, and 1c have conditions (1) and (2).
[0087] (1) When the sound processing devices 1, 1a, 1b and 1c record sound data of the sound of waves and sound data of the sound of a ship's whistle.
[0088] (2) When the sound processing devices 1, 1a, 1b, and 1c acquire the sound of waves in the real space, they have a switching condition of switching from sound data of the sound of waves to sound data of the sound of a ship's whistle.
[0089] Under conditions (1) and (2), if there is no sound of waves in the real space, the sound processing devices 1, 1a, 1b, and 1c reproduce sound data of the sound of waves. That is, the user hears the sound of waves in the virtual space. However, if the sound processing devices 1, 1a, 1b, and 1c acquire the sound of waves in the real space, the switching condition is met, and the sound data is switched to the sound of a ship's whistle. As a result, the user hears the sound of waves in the real space and the sound of a ship's whistle in the virtual space. That is, the sound processing devices 1, 1a, 1b, and 1c utilize sounds in the virtual space while utilizing sounds in the real space as much as possible, thereby enhancing the user's sense of immersion. This allows the sound processing devices 1, 1a, 1b, and 1c to perform effects to enhance the sense of immersion without the user being aware of it. In this way, by performing an effect of switching between multiple second sounds, the sound processing devices 1, 1a, 1b, and 1c can output sounds appropriate for the scene being played. Therefore, the user can listen to the sound with a more immersive feeling without feeling any discomfort.
[0090] The microphone 10 may be connected to the terminal 20 via a wire. In this case, even if the terminal 20 and the headphones 30 do not have the microphone 10, the terminal 20 can acquire environmental sounds using the microphone 10 connected via a wire.
[0091] The terminal 20 may be provided with an application program that can edit sound data. In this case, for example, the user can edit the sound data in real time by operating the terminal 20.
[0092] Note that, when the type of the first sound matches the type of the second sound in the first comparison, the sound processing devices 1, 1a, 1b, and 1c may output the acquired first sound to the headphones 30 without reproducing sound data related to the second sound of the type that matches the first sound. In this case, the headphones 30 have a hear-through mode that outputs sound acquired by their own microphone. In the hear-through mode, the sound acquired by the microphone of the headphones 30 is output from the speaker of the headphones 30. That is, in this case, the headphones 30 output the environmental sound acquired by their own microphone and the second sound of a type that does not match the environmental sound. A detailed description will be given below with reference to FIG. 12A. FIG. 12A is a diagram showing an example of outputting the first sound and reproducing sound data. For example, as shown in FIG. 12A, if the microphone 10 acquires the sounds of cicadas and a ship's whistle, the headphones 30 output the sounds of cicadas and a ship's whistle acquired by their own microphone. In this case, if the sound data relating to the second sound includes the chirping of cicadas, the sound data of the chirping of cicadas is not played back, as shown in Fig. 12A. However, as shown in Fig. 12A, the sound data of the river and the car engine, which do not match the first sound, are played back. This allows the user to hear the sound intended by the creator.
[0093] Note that, if the headphones 30 are equipped with a hear-through mode, the sound processing devices 1, 1a, 1b, and 1c do not necessarily have to play back sound data (sounds based on the sound data may not be heard by the user). Hereinafter, a detailed description will be given with reference to FIG. 12B. FIG. 12B is a diagram showing an example of a case in which the sound processing devices 1, 1a, 1b, and 1c do not play back sound data. As shown in FIG. 12B, if only the sounds of cicadas are set as sound data and if the sounds of cicadas are acquired as the first sound, the sound processing devices 1, 1a, 1b, and 1c do not play back the sounds of cicadas in the sound data. In this case, the sound processing devices 1, 1a, 1b, and 1c output only the sounds of cicadas in the real space. Therefore, for example, if the sounds of cicadas in the real space are continuously acquired for 30 seconds, the sound processing devices 1, 1a, 1b, and 1c do not play back sound data for 30 seconds. Then, when the sounds of cicadas in the real space are no longer acquired, the sound processing devices 1, 1a, 1b, and 1c reproduce the sound data.
[0094] Furthermore, the sound processing devices 1, 1a, 1b, and 1c may cause the headphones 30 to erase a sound of a sound source type that does not match the second sound from among the first sounds acquired by the microphone 10. This will be described in detail below with reference to FIG. 12C. FIG. 12C is a diagram showing an example of erasing a sound of a sound source type that does not match the second sound from among the first sounds. For example, if the microphone 10 acquires the sounds of cicadas, a ship's whistle, and an airplane's engine, and if the sound data related to the second sound includes cicadas, the sound processing devices 1, 1a, 1b, and 1c may cause the headphones 30 to perform processing to erase the ship's whistle and the airplane's engine (sounds that do not match the second sound). Alternatively, the sound processing devices 1, 1a, 1b, and 1c may transmit a sound signal obtained by erasing the ship's whistle and the airplane's engine from the sound acquired by the microphone 10 to the headphones 30, and output the sound signal from the headphones 30. As a result, the only environmental sound output by the headphones 30 is the chirping of cicadas. Therefore, the sound processing devices 1, 1a, 1b, and 1c can output only the environmental sounds intended by the creator while playing back the sound data intended by the creator. As a result, the user can hear even more of the sounds intended by the creator. Note that if the sound processing devices 1, 1a, 1b, and 1c are equipped with a specific sound elimination unit 200e, the specific sound elimination unit 200e may also eliminate the sounds of ship whistles and airplane engines.
[0095] The description of the present embodiment should be considered to be illustrative in all respects and not restrictive. The scope of the present invention is defined not by the above-described embodiments but by the claims. Furthermore, the scope of the present invention is intended to include all modifications within the meaning and scope of the claims. [Explanation of symbols]
[0096] 1, 1a, 1b, 1c...sound processing device 10. Mike 20...Terminal 200...CPU 200a…Analysis department 200b…Comparison section 200c…Reproduction part 200d…External environment data acquisition section 200e…Specific sound elimination section 201...ROM 202...RAM 203...Output I / F 30...Headphones 40a...sensor 40b...server
Claims
1. Obtaining a first sound; recording the acquired first sound as first sound data; acquiring a second sound composed of second sound data created in advance; performing an analysis of the first sound; performing a first comparison of comparing the analysis result of the first sound with the second sound data; Based on the result of the first comparison, the second sound data relating to the second sound of a type that does not match the first sound is reproduced, and in a case where the first sound data is recorded and the first sound and the second sound match, the first sound data is reproduced instead of the second sound data; a sound processing method for outputting sound related to the reproduced second sound data;
2. When the second sound data is reproduced, if there is a sound source of the same type as the second sound that is pre-recorded at a specific location, The second sound is converted into the sound source, and the second sound data is reproduced. The sound processing method according to claim 1 .
3. acquiring the first sound and analyzing the first sound while the second sound data is being played back; 3. The sound processing method according to claim 1 or 2.
4. the first sound includes a plurality of analysis targets; 4. The sound processing method according to claim 1.
5. the analysis result includes first sound source information indicating a type of a first sound source included in the acquired first sound, the second sound data includes data to which second sound source information indicating a type of the second sound source is added in advance; 5. A sound processing method according to claim 1.
6. preparing a trained model that has been trained using a dataset that indicates a relationship between third sound source information that indicates a type of a third sound source and a feature amount of the third sound source as training data; In the analysis of the first sound, Calculating the feature amount contained in the first sound; After calculating the feature amount, the feature amount is input to the trained model, and the third sound source information corresponding to the feature amount is output as an analysis result of the first sound.
6. A sound processing method according to claim 1.
7. Acquires surrounding environmental data, performing a process of reproducing the second sound data related to the second sound based on the acquired environmental data; 7. A sound processing method according to claim 1.
8. Connect to the network, acquiring the environmental data via the network; The sound processing method according to claim 7.
9. acquiring localization processing data to be used in processing of sound image localization related to the second sound; A sound processing method according to any one of claims 1 to 8.
10. the second sound is multi-track; obtaining a switching condition for the multi-track; reproducing the second sound data related to the second sound of the type that satisfies the switching condition; 10. A sound processing method according to claim 1.
11. a first sound acquisition unit that acquires a first sound; a recording unit that records the acquired first sound as first sound data; a second sound acquisition unit configured with second sound data created in advance; an analysis unit that analyzes the first sound; A first comparison is performed to compare the analysis result of the first sound with the second sound data. a comparison section; a playback unit that plays back the second sound data; an output unit that outputs the sound related to the second sound data reproduced by the reproduction unit; Equipped with the reproduction unit reproduces the second sound data relating to the second sound of a type that does not match the first sound based on a result of the first comparison, and when the first sound data is recorded and the first sound and the second sound match, reproduces the first sound data instead of the second sound data; outputting the sound related to the reproduced second sound data; Sound processing device.
12. The playback unit When the second sound data is reproduced, if there is a sound source of the same type as the second sound that is pre-recorded at a specific location, The second sound is converted into the sound source, and the second sound data is reproduced. The sound processing device according to claim 11 .
13. the acquisition of the first sound performed by the first sound acquisition unit and the analysis of the first sound performed by the analysis unit are performed during playback of the second sound data; The sound processing device according to claim 11 or 12.
14. the first sound includes a plurality of analysis targets; 14. The sound processing device according to claim 11.
15. the analysis result includes first sound source information indicating a type of a first sound source included in the acquired first sound, the second sound data includes data to which second sound source information indicating a type of the second sound source is added in advance; 15. The sound processing device according to claim 11.
16. the sound processing device further includes a trained model trained using a data set indicating a relationship between third sound source information indicating a type of a third sound source and a feature amount of the third sound source as training data; the analysis unit calculates a feature amount included in the first sound in the analysis of the first sound, After calculating the feature, the trained model inputs the feature into the trained model, thereby outputting the third sound source information corresponding to the feature as an analysis result of the first sound.
16. A sound processing device according to any one of claims 11 to 15.
17. Further, an external environment data acquisition unit is provided to acquire surrounding environment data, the reproduction unit performs processing to reproduce the second sound data related to the second sound based on the environmental data acquired from the external environment data acquisition unit.
17. A sound processing device according to any one of claims 11 to 16.
Citation Information
Patent Citations
Speech providing device, speech reproducing device, speech providing method, and speech reproducing method
WO2018088450A1