Masking device
The masking device addresses the limitation of pre-stored data by using real-time analysis of speech features to generate masking sounds, effectively concealing human speech in real-time scenarios.
Patent Information
- Application Number
- JP2021126014
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-07-30
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2041-07-30
AI Technical Summary
Existing masking devices do not generate masking sounds in real time corresponding to human speech, relying instead on pre-stored sample data for voice concealment.
A masking device that includes a detection unit for identifying speech in audio signals, an analysis unit for generating feature data on the speech, and a generation unit for creating masking data based on this feature data to produce music that effectively masks human speech.
Enables real-time generation of masking sounds that effectively conceal human speech, improving the ability to maintain privacy in environments where conversation needs to be obscured.
Smart Images

Figure 0007687120000001 
Figure 0007687120000002 
Figure 0007687120000003
Abstract
Description
Technical Field
[0001] The present invention relates to a masking device.
Background Art
[0002] Conventionally, in order to prevent third parties from understanding the content of conversation voices between people in a car, at a counter in a store or a hospital, etc., a technique of outputting a masking sound that erases the conversation voice has been used.
[0003] For example, Patent Document 1 discloses a concealment device for concealing conversation voices. The concealment device includes a storage device in which voice data indicating general conversation voices and music data indicating music are stored in advance. The concealment device includes a concealment data generation device that generates concealment data in which the voice data and music data read from the storage device are synthesized. Further, the concealment device includes a music playback device that plays back the concealment data. By playing back this concealment data, for example, at a bank window, the conversation between a bank clerk and a user can be concealed so that it cannot be heard by a third party.
Prior Art Documents
Patent Documents
[0004]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0005] However, the concealment device according to Patent Document 1 generated concealment data as a masking sound by synthesizing voice data and music data as pre-stored sample data. That is, the technique according to Patent Document 1 did not generate a masking sound in real time corresponding to human speech.
[0006] In view of the above circumstances, an aspect of the present disclosure aims to provide a masking device that generates masking data for responding to human speech in real time and plays music that masks human speech based on the generated masking data.
Means for Solving the Problems
[0007] In order to solve the above problems, a masking device according to an aspect of the present disclosure includes a detection unit that detects an audio signal indicating speech from an output signal output from a microphone, an analysis unit that generates feature data indicating the features of the speech by analyzing the audio signal, and a generation unit that generates masking data indicating music that masks the speech based on the feature data.
Brief Description of the Drawings
[0008]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Mode for Carrying Out the Invention
[0009] 〔1. First Embodiment〕 〔1-1. Configuration of the First Embodiment〕 FIG. 1 is a block diagram illustrating the configuration of a masking device 1 according to a first embodiment of the present disclosure. The masking device 1 is a device that generates masking data Dm indicating music for masking the voice based on the characteristics of the collected human voice, and reproduces the music for masking the voice based on the generated masking data Dm. Specifically, the masking device 1 includes a control device 11, a storage device 12, an operation device 13, a sound collection device 14, and a reproduction device 15.
[0010] The control device 11 in FIG. 1 is, for example, one or more processors that control each element of the masking device 1. For example, the control device 11 is composed of one or more types of processors such as a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a DSP (Digital Signal Processor), an FPGA (Field Programmable Gate Array), or an ASIC (Application Specific Integrated Circuit).
[0011] The storage device 12 is, for example, one or more memories composed of a known recording medium such as a magnetic recording medium or a semiconductor recording medium. The storage device 12 stores a control program PR1 executed by the control device 11 and various data used by the control device 11, particularly music data Dx. Note that the storage device 12 may be composed of a combination of a plurality of types of recording media. Further, the storage device 12 may be a portable recording medium detachable from the masking device 1, or an external recording medium (for example, online storage) that the masking device 1 can communicate with via a communication network.
[0012] The operating device 13 is an input device that receives instructions from the user. The operating device 13 is, for example, a plurality of operators that can be operated by the user, or a touch panel that detects contact from the user. In particular, the operating device 13 has a function as a switch for instructing the start and end of the operation of the masking device 1. Further, the operating device 13 is used when storing the music data Dx supplied from the outside in the storage device 12.
[0013] The sound collection device 14 includes a sound collection unit that collects ambient sound, and is a microphone that converts the collected sound into an electrical signal. The sound collection unit may be any configuration as long as it can collect sound. For example, a windproof structure is applicable. Further, the ambient sound may include human voices. The sound collection device 14 of the present embodiment generates an analog sound signal based on the collected sound. Further, the sound collection device 14 includes an AD converter that converts the sound signal into sound data Ds. The sound data Ds is output from the sound collection device 14.
[0014] The playback device 15 plays music based on the masking data Dm generated by the control device 11 under the control of the control device 11. The masking data Dm indicates music. The playback device 15 includes a DA converter, an amplifier, and a speaker. The digital signal masking data Dm is input to the DA converter. The input masking data Dm is converted into a masking signal that is an analog signal. The masking signal is amplified in the amplifier so as to have an amplitude suitable for sound playback by the subsequent speaker. The music indicated by the masking signal with the amplified amplitude is played from the speaker as a sound playback device. The masking device 1 according to the present embodiment is preferably used, for example, in the vehicle C shown in FIG. 2. In this case, the speaker mounted on the vehicle C is used as an element provided in the playback device 15.
[0015] FIG. 2 is an example of a plan view of the vehicle C equipped with the masking device 1 according to the present embodiment, and FIG. 3 is an example of a side view of the vehicle C.
[0016] In the examples shown in FIGS. 2 and 3, in the passenger compartment R of the vehicle C, in addition to the masking device 1, there are four seats 51 to 54 arranged in a rectangle, a ceiling 6, a front right door 71, a front left door 72, a rear right door 73, and a rear left door 74. The seat 51 is the driver's seat, the seat 52 is the front passenger seat, the seat 53 is the rear right seat, and further, the seat 54 is the rear left seat. Each of the seats 51 to 54 is made of a material such as cloth or leather and has sound absorption properties. The seats 51 to 54 are facing in a common direction. Each of the seats 51 to 54 has headrests 51-1 to 54-1.
[0017] The masking device 1 includes a microphone as the above-described sound collection device 14, and first speakers 15-1, second speakers 15-2, third speakers 15-3, and fourth speakers 15-4 which are elements of the playback device 15. The sound collection device 14 is arranged on the ceiling 6 of the passenger compartment R. It is preferable that the sound collection device 14 includes a first sound collection device 14-1 and a second sound collection device 14-2. In this case, the first sound collection device 14-1 is installed near the seats 51 and 52 which are the front seats on the ceiling 6 of the passenger compartment R. Further, it is preferable that the first sound collection device 14-1 has directivity so as to easily collect the voices of the persons sitting on the seats 51 and 52. Similarly, the second sound collection device 14-2 is installed near the seats 53 and 54 which are the rear seats on the ceiling 6 of the passenger compartment R. Further, it is preferable that the second sound collection device 14-2 has directivity so as to easily collect the voices of the persons sitting on the seats 53 and 54. However, the configuration of the sound collection device 14 is not limited to this. It is preferable that the sound collection device 14 can separately collect the voices of the persons sitting on the seats 51 and 52 which are the front seats and the voices of the persons sitting on the seats 53 and 54 which are the rear seats, but the configuration is not limited.
[0018] The first speaker 15-1 is installed on the headrest 51-1. The second speaker 15-2 is installed on the headrest 52-1. The third speaker 15-3 is installed on the headrest 53-1. The fourth speaker 15-4 is installed on the headrest 54-1. Note that these installation locations are just examples and are not limited to these. For example, each of the first speaker 15-1 to the fourth speaker 15-4 may be installed at the lower part of the front right door 71, the lower part of the front left door 72, the lower part of the rear right door 73, and the lower part of the rear left door 74.
[0019] When the first sound collection device 14-1 collects the voices of the persons sitting on the front seats 51 and 52, music indicated by a masking signal is emitted from the third speaker 15-3 installed on the headrest 53-1 of the rear seat 53 and the fourth speaker 15-4 installed on the headrest 54-1 of the seat 54. This is because if music indicated by a masking signal is emitted from the first speaker 15-1 and the second speaker 15-2 which are the front seat speakers, it may interfere with the conversation between the front seats.
[0020] Thereby, for example, it becomes possible to prevent the voice of the driver who is talking from being heard by the person sitting on the rear seat. Subsequently, the driver can obtain a sense of security that his own conversation is not being heard by the rear seat and can concentrate on driving.
[0021] On the other hand, when the second sound collection device 14-2 collects the voices of the persons sitting on the rear seats 53 and 54, music indicated by a masking signal is emitted from the first speaker 15-1 installed on the headrest 51-1 of the front seat 51 and the second speaker 15-2 installed on the headrest 52-1 of the seat 52. This is because if music indicated by a masking signal is emitted from the third speaker 15-3 and the fourth speaker 15-4 which are the rear seat speakers, it may interfere with the conversation between the rear seats.
[0022] This makes it possible to prevent the driver from hearing the conversation of the person sitting in the rear seat. Furthermore, by using music as the masking sound, the driver can concentrate on driving.
[0023] In addition, when both the first sound collection device 14-1 and the second sound collection device 14-2 collect the voices of the people sitting in seats 51 to 54, the music indicated by the masking signal is not played from any of the first speakers 15-1 to the fourth speakers 15-4. This is to prevent interfering with the conversation between the front seat and the rear seat.
[0024] FIG. 4 is a block diagram illustrating a functional configuration of the control device 11. The control device 11 functions as a detection unit 111, an analysis unit 112, an acquisition unit 113, a generation unit 114, and a selection unit 115 by reading out a control program PR1 and executing the read control program PR1.
[0025] The detection unit 111 detects voice data Dv indicating human voice from the sound data Ds output from the sound collection device 14. The voice data Dv has a silent section without voice and a voice section with voice. The detection unit 111 is constituted by, for example, a band-pass filter having a voice band as a pass band. The sound indicated by the sound data Ds may include running sound, music sound, etc. in addition to the voice. The voice data Dv is extracted from the sound data Ds by the detection unit 111.
[0026] In addition, the detection unit 111 outputs a control signal S to the selection unit 115. When the detection unit 111 detects a voice section from the voice data Dv, the control signal S has a value indicating “ON”. On the other hand, when the detection unit 111 detects a silent section, the control signal S has a value indicating “OFF”.
[0027] The analysis unit 112 generates feature data Df indicating the features of the voice by analyzing the voice data Dv detected by the detection unit 111. More specifically, the analysis unit 112 generates feature data Df indicating the features of the voice by analyzing the voice data Dv in the voice section. Here, the "features of the voice" include at least one of the pitch of the voice, the level of the voice, and the formants of the voice. The "pitch of the voice" refers to the fundamental frequency of the voice. The "level of the voice" refers to the volume of the voice. The "formants of the voice" refer to the frequency bands in the frequency spectrum of the voice where the intensity is greater than that of the surroundings. These frequency bands are called the "first formant", "second formant", "third formant", etc. in order from the lowest. The quality of the voice is determined by the height of the frequency of each of the multiple formants.
[0028] In particular, when the feature data Df generated by the analysis unit 112 includes the pitch of the voice or the formants of the voice, the analysis unit 112 can determine whether the person who uttered the voice is male or female by analyzing the pitch or formants of the voice. Specifically, when the pitch of the voice is equal to or higher than a predetermined value, the analysis unit 112 determines that the main speaker of the voice is female. On the other hand, when the pitch of the voice is less than the predetermined value, the analysis unit 112 determines that the main speaker of the voice is male. Also, when the first formant and the second formant of the vowel included in the voice are equal to or higher than a predetermined value, the analysis unit 112 determines that the main speaker of the voice is female. On the other hand, when the first formant and the second formant of the vowel included in the voice are less than the predetermined value, the analysis unit 112 determines that the main speaker of the voice is male.
[0029] The acquisition unit 113 acquires music data Dx from the storage device 12. As will be described later, the music indicated by the masking data Dm generated by the masking device 1 includes a plurality of parts that correspond one-to-one with a plurality of timbres. The music data Dx includes a plurality of part data Dp1, Dp2,... Dpn that correspond one-to-one with these plurality of parts. n is an integer of 2 or more. When there is no need to distinguish each part, it is simply referred to as part data Dp.
[0030] FIG. 5 is a diagram showing an example of the frequency band of human voice and the frequency bands of a plurality of timbres corresponding to a plurality of parts included in the music indicated by the masking data Dm. In FIG. 5, the top row indicates frequency. The second row indicates the code. In the example shown in FIG. 5, it is the same C code, and an example is shown where the frequency increases by one octave from C0 to C8. The third row to the ninth row indicate the frequency band of human voice. The tenth row to the fourteenth row indicate the frequency band of the performance sound of musical instruments.
[0031] As shown in FIG. 5, human voice has a frequency band from approximately 73 Hz to approximately 1047 Hz.
[0032] In particular, the bass, which is a male voice, has a vocal range from approximately D2 to F4, that is, a frequency band from approximately 73 Hz to approximately 350 Hz. The baritone, which is a male voice, has a vocal range from approximately G2 to G4, that is, a frequency band from approximately 98 Hz to approximately 392 Hz. The tenor, which is a male voice, has a vocal range from approximately C3 to C5, that is, a frequency band from approximately 131 Hz to approximately 523 Hz. Generally, male voice has a frequency band from approximately 73 Hz to approximately 523 Hz.
[0033] Alto, which is a female voice, has a vocal range of approximately F3 to E5, that is, a frequency band of approximately 175 Hz to approximately 659 Hz. Mezzo-soprano, which is a female voice, has a vocal range of approximately A3 to A5, that is, a frequency band of approximately 220 Hz to approximately 880 Hz. Soprano, which is a female voice, has a vocal range of approximately C4 to C6, that is, a frequency band of approximately 262 Hz to approximately 1047 Hz. Generally speaking, female voices have a frequency band of approximately 175 Hz to approximately 1047 Hz.
[0034] On the other hand, as shown in FIG. 5, the performance sound of musical instruments has a frequency band of approximately 25 Hz to approximately 4400 Hz. For example, the double bass corresponding to the part data Dp1 has a pitch range of approximately E1 to G3, that is, a frequency band of approximately 41 Hz to approximately 196 Hz. The cello corresponding to the part data Dp2 has a pitch range of approximately C2 to C5, that is, a frequency band of approximately 65 Hz to approximately 523 Hz. The viola corresponding to the part data Dp3 has a pitch range of approximately C3 to C6, that is, a frequency band of approximately 131 Hz to approximately 1047 Hz. The violin corresponding to the part data Dp4 has a pitch range of approximately G3 to E7, that is, a frequency band of approximately 196 Hz to approximately 2637 Hz.
[0035] Comparing the frequency band of human voices with the frequency band of the performance sound of musical instruments, it can be said that the frequency band of male voices is generally included in the frequency band of the performance sound of the cello. On the other hand, it can be said that the frequency band of female voices is generally included in the frequency band of the performance sound of the viola.
[0036] Each of the plurality of parts included in the music indicated by the masking data Dm is associated with the pitch or formant of human voice. As an example, the cello part may be associated with the pitch indicating a male voice among the pitches of human voice. Alternatively, the cello part may be associated with the formant indicating a male voice among the formants of human voice. Similarly, the viola part may be associated with the pitch indicating a female voice among the pitches of human voice. Alternatively, the viola part may be associated with the formant indicating a female voice among the formants of human voice.
[0037] The music data Dx may be MIDI (Musical Instrument Digital Interface) data. When the music data Dx is MIDI data, the music data Dx of a predetermined piece of music includes a plurality of part data Dp each corresponding to each timbre. Here, the timbre corresponding to each part data Dp includes not only the timbre of musical instruments but also the timbre of voices other than musical instruments such as human voices and synthesized sounds. Alternatively, the music data Dx may be PCM data obtained by sampling a music signal. Also, when the music data Dx is PCM data, the music data Dx may be composed of a plurality of PCM data corresponding one-to-one to a plurality of timbres. When the music data Dx is PCM data in which a plurality of timbres are mixed, the music data Dx may be decomposed into PCM data of a plurality of timbres by a well-known sound source separation technique, and a predetermined timbre (cello, viola, etc.) may be selected from them and used for masking. The plurality of PCM data corresponds to the part data Dp1 to Dpn.
[0038] Returning to FIG. 4, the generation unit 114 generates masking data Dm indicating music that masks the voice based on the feature data Df generated by the analysis unit 112. In particular, in the present embodiment, the generation unit 114 selects one part data Dp from among a plurality of part data Dp1 to Dpn included in the music data Dx acquired by the acquisition unit 113 based on the feature data Df. Next, the generation unit 114 generates the masking data Dm by correcting at least one of the pitch of the sound indicated by the selected part data Dp and the level of the sound. When the selected part data Dp is Dps, the masking data Dm includes one part data Dps' in which the part data Dps is corrected and the part data Dp excluding the one part data Dps that is the target of the correction among the plurality of part data Dp1 to Dpn described above.
[0039] More specifically, when the formant is included in the features of the voice, the generation unit 114 selects the part data Dps in the sound range that overlaps the formant of the voice. Alternatively, when the pitch is included in the features of the voice, the generation unit 114 selects the part data Dps in the sound range that includes the same frequency as the pitch of the voice. As an example, when the formant or pitch of the voice corresponds to a male voice, the generation unit 114 selects the cello part. On the other hand, when the formant or pitch of the voice corresponds to a female voice, the generation unit 114 selects the viola part.
[0040] If there is no appropriate part, the generation unit 114 selects the part data Dps in the sound range that is closest to the formant of the voice among the features of the voice from among the existing part data Dp1 to Dpn. Alternatively, the generation unit 114 selects the part data Dps in the sound range that has the frequency closest to the pitch of the voice among the features of the voice from among the existing part data Dp1 to Dpn.
[0041] On top of that, the generation unit 114 corrects the selected part data Dps so as to change the level of the sound indicated by the selected part data Dps according to the level of the voice, and generates part data Dps'. More specifically, when the music based on the part data Dps' is played from the speaker, the generation unit 114 corrects the part data Dps so that the voice indicated by the voice data Dv can be masked by the played music. Further, the generation unit 114 generates masking data Dm from the corrected part data Dps' and the part data Dp among the plurality of part data Dp excluding the part data Dps that was the target of the correction. In particular, when the level of the voice detected by the detection unit 111 is high, the generation unit 114 corrects the part data Dp so as to increase the level of the sound indicated by the selected part data Dp according to the loudness of the voice.
[0042] Also, in the present embodiment, when the voice section of the voice data Dv is detected by the detection unit 111 during the period in which the acquisition unit 113 reads the music data Dx from the storage device 12, the generation unit 114 executes the above correction to generate the masking data Dm. Further, the generation unit 114 outputs the generated masking data Dm to the selection unit 115.
[0043] Also, the generation unit 114 outputs the music data Dx acquired from the acquisition unit 113 to the selection unit 115 in parallel with the output of the masking data Dm.
[0044] The selection unit 115 selects one of the masking data Dm and the music data Dx based on the control signal S input from the detection unit 111, and outputs it to the playback device 15. More specifically, when the control signal S is a value indicating "ON", the selection unit 115 selects the masking data Dm and outputs the selected masking data Dm to the playback device 15. On the other hand, when the control signal S is a value indicating "OFF", the selection unit 115 selects the music data Dx and outputs the selected music data Dx to the playback device 15.
[0045] The playback device 15 has a function of converting the format of MIDI data or PCM data into the format of music data. As a result, the playback device 15 always plays the music indicated by the music data Dx, and during that process, it switches its operation to play the music indicated by the masking data Dm. At this time, the generation unit 114 corrects the part data Dps indicating a part of the music that was originally being played. Therefore, there is no sense of discomfort for a person who was listening to the music played by the playback device 15.
[0046] FIG. 6 is a diagram showing the levels of each part data Dp included in the music data Dx and the masking data Dm output by the generation unit 114. Note that the example shown in FIG. 6 shows the case where the human voice indicated by the voice data Dv is a male voice. At time t1, the generation unit 114 outputs, in advance, as music data Dx, the cello part data Dp2 and the other part data Dp1, Dp3, and Dp4 in parallel to the selection unit 115. During this period, the selection unit 115 outputs the music data Dx to the playback device 15. When the detection unit 111 detects a human voice at time t2, the analysis unit 112 generates feature data Df indicating the level of the voice and the features of the voice including at least one of the pitch of the voice and the formant. The generation unit 114 selects the cello part data Dp2 as the part data Dp whose sound range includes the same frequency as the pitch of the voice, or the part data Dp whose sound range overlaps with the formant of the voice. Further, the generation unit 114 corrects the part data Dp2 so as to increase the level of the sound indicated by the cello part data Dp2 according to the level of the voice, and generates part data Dp2'. The generation unit 114 outputs the masking data Dm including the cello part data Dp2 with the increased sound level to the playback device 15. Regarding the other part data Dp1, Dp3, and Dp4 included in the masking data Dm, the sound level is not continuously changed. The selection unit 115 selects the masking data Dm from the music data Dx and the masking data Dm based on the control signal S, and outputs the selected masking data Dm to the playback device 15. When the detection unit 111 stops detecting a human voice at time t3, the generation unit 114 returns the level of the cello part data Dp2 to its original level. Then, the generation unit 114 continues to output the music data Dx including the cello part data Dp2 and the other part data Dp1, Dp3, and Dp4 to the playback device 15.
[0047] Instead of, or in addition to, correcting the level of the sound, the generation unit 114 corrects the selected part data Dp so that the pitch of the sound indicated by the selected part data Dp approaches the pitch of the sound indicated by the audio data Dv detected by the detection unit 111, and may generate part data Dp'. The difference between the pitch of the sound and the pitch of the corrected sound is smaller than the difference between the pitch of the sound and the pitch of the sound before correction. Therefore, the pitch of the sound and the pitch of the corrected sound may not match.
[0048] More specifically, the generation unit 114 may correct the selected part data Dp so as to raise or lower the key of the sound indicated by the selected part data Dp in octave units according to the pitch of the human voice, and generate part data Dp'. Thereby, the generation unit 114 can correct only the selected part data Dp in a state where the music is established as a piece of music without changing the melody of the music indicated by the music data Dx.
[0049] Alternatively, the generation unit 114 may correct the selected part data Dp so as to raise or lower the chord of the sound indicated by the selected part data Dp in semitone units according to the pitch of the human voice, and generate part data Dp'. Thereby, although the melody of the music indicated by the music data Dx changes, the generation unit 114 can finely adjust the pitch of the sound indicated by the selected part data Dp. By correcting the pitch of the sound in this way, the pitch of the sound approaches the pitch of the voice, so the masking effect is improved.
[0050] [1-2. Operation of the First Embodiment] FIG. 7 is a flowchart showing the operation of the masking device 1 according to the first embodiment. Hereinafter, the operation of the masking device 1 according to the first embodiment will be described with reference to FIG. 7.
[0051] In step S1, the acquisition unit 113 acquires music data Dx from the storage device 12.
[0052] In step S2, the generation unit 114 outputs the music data Dx acquired from the acquisition unit 113 to the selection unit 115. The selection unit 115 outputs the music data Dx to the playback device 15.
[0053] In step S3, when human speech is detected by the detection unit 111 (S3: YES), the masking device 1 executes the process of step S4. When human speech is not detected by the detection unit 111 (S3: NO), the masking device 1 executes the process of step S2.
[0054] In step S4, the analysis unit 112 generates feature data Df indicating the features of the speech by analyzing the speech signal detected by the detection unit 111.
[0055] In step S5, the generation unit 114 generates masking data Dm indicating the music for masking the speech based on the feature data Df generated by the analysis unit 112. More specifically, in step S5, the generation unit 114 selects one part data Dps from among the plurality of part data Dp included in the music data Dx acquired by the acquisition unit 113 based on the feature data Df. Next, the generation unit 114 corrects the selected part data Dps so as to change the sound level indicated by the selected part data Dps according to the feature data Df, and generates part data Dps'. Further, the generation unit 114 generates masking data Dm from the part data Dps' and the part data Dp among the plurality of part data Dp excluding the part data Dps that was the target of the correction. Note that the generation unit 114 may change the pitch of the sound according to the feature data Df instead of or in addition to the sound level indicated by the selected part data Dps. In particular, when a speech section is detected by the detection unit 111 during the period in which the acquisition unit 113 reads the music data Dx from the storage device 12, the generation unit 114 executes the above-described correction.
[0056] In step S6, the generation unit 114 outputs the generated masking data Dm to the selection unit 115. The selection unit 115 outputs the masking data Dm to the playback device 15.
[0057] 〔2. Second Embodiment〕 Hereinafter, the masking device 1 according to the second embodiment of the present disclosure will be described. Among the components provided in the masking device 1 according to the second embodiment, the same components as those provided in the masking device 1 according to the first embodiment are denoted by the same reference numerals, and the description of their functions is omitted.
[0058] 〔2-1. Configuration of the Second Embodiment〕 FIG. 8 is a block diagram illustrating a functional configuration of the control device 11 included in the masking device 1 according to the second embodiment. The masking device 1 according to the second embodiment includes a generation unit 114A instead of the generation unit 114 included in the masking device 1 according to the first embodiment.
[0059] The generation unit 114A outputs a predetermined part data Dp among a plurality of part data Dp as music data Dx to the selection unit 115. On the other hand, the generation unit 114A executes the same correction as the generation unit 114. Then, the generation unit 114A outputs masking data Dm including the above-mentioned predetermined part data Dp and one corrected part data Dps’ to the selection unit 115.
[0060] FIG. 9 is a diagram showing the levels of the respective part data Dp included in the masking data Dm generated by the generation unit 114A. Note that the example shown in FIG. 9 shows the case where the human voice is a male voice. At time t1, the generation unit 114 outputs other part data Dp other than the cello as music data Dx to the selection unit 115 in advance. The "other part data" is, for example, the part data Dp4 of the violin. During this period, the selection unit 115 outputs the music data Dx to the playback device 15. When the detection unit 111 detects a human voice at time t2, the analysis unit 112 generates voice feature data Df including at least one of the pitch, level, and formant of the voice. The generation unit 114 selects the cello part data Dp2 as part data Dp whose sound range includes the same frequency as the pitch of the voice, or part data Dp whose sound range overlaps with the formant of the voice. Further, the generation unit 114 corrects the cello part data Dp2 so as to change the level of the sound indicated by the cello part data Dp2 according to the level of the voice, and generates part data Dp2'. The generation unit 114 outputs masking data Dm including the corrected cello part data Dp2' and the uncorrected violin part data Dp4 to the selection unit 115. The selection unit 115 selects the masking data Dm from the music data Dx and the masking data Dm based on the control signal S, and outputs the selected masking data Dm to the playback device 15. When the detection unit 111 stops detecting a human voice at time t3, the generation unit 114 stops outputting the corrected part data Dp2' of the cello. Then, the generation unit 114 continues to output the violin part data Dp4, which is other part data Dp, to the selection unit 115 as music data Dx. The selection unit 115 outputs the music data Dx to the playback device 15.
[0061] [2-2. Operations of the Second Embodiment] The operation of the masking device 1 according to the second embodiment is basically the same as the operation of the masking device 1 according to the first embodiment, and thus the illustration thereof is omitted.
[0062] In step S2, the generation unit 114A outputs a predetermined part data Dp among a plurality of part data Dp included in the music data Dx to the selection unit 115. The selection unit 115 outputs the predetermined part data Dp to the playback device 15 as the music data Dx.
[0063] In step S5, the generation unit 114A executes the same correction as the generation unit 114, and generates masking data Dm including the predetermined part data Dp in step S2 and one corrected part data Dps'.
[0064] 〔3. Modification Example〕 The above embodiments can be variously modified. Specific modification modes are exemplified below. Two or more modes arbitrarily selected from the following examples can be appropriately combined as long as they do not conflict.
[0065] 〔3-1. Modification Example 1〕 In the above first and second embodiments, the generation units 114 and 114A generated the masking data Dm by correcting the music data Dx acquired from the storage device 12 by the acquisition unit 113. However, the method for generating the masking data Dm in the embodiments of the present invention is not limited to this. For example, the generation units 114 and 114A may generate a new song and generate the masking data Dm corresponding to the generated song. For example, the generation units 114 and 114A may generate a new song by applying a conventional technique of automatically composing or accompanying music based on the specified key and chord. In this case, the generation units 114 and 114A may determine the key based on the pitch of the human voice detected by the detection unit 111 and automatically generate a new song based on a preselected chord.
[0066] 〔3-2. Modification Example 2〕 In the above-described first and second embodiments, the playback device 15 played music as masking sound based on the masking data Dm output from the generation unit 114. In this modification, the playback device 15 may further play music as masking sound specifically when human speech is detected by the detection unit 111.
[0067] [4. Supplementary Note] From the above-described embodiments and the like, for example, the following aspects can be grasped.
[0068] The masking device 1 according to an aspect (first aspect) of the present disclosure includes a detection unit 111 that detects a voice signal indicating voice from an output signal output from the sound collection device 14. Further, the masking device 1 includes an analysis unit 112 that generates feature data Df indicating the features of the voice by analyzing the voice signal. Furthermore, the masking device 1 includes a generation unit 114 that generates masking data Dm indicating music that masks the voice based on the feature data Df.
[0069] By having this configuration, it becomes possible to detect human voice in real time by the detection unit 111, extract the features of the voice by the analysis unit 112, and generate music data Dx corresponding to the features of the voice by the generation unit 114. For this reason, the masking device 1 can generate masking data Dm that responds to human speech in real time, and based on the generated masking data Dm, play music that masks human voice. Also, since the sound used for masking is music, there is an advantage that it does not cause fatigue even when listened to for a long time.
[0070] Also, in an example of the first aspect (second aspect), the features of the voice include at least one of the pitch of the voice, the level of the voice, and the formant of the voice.
[0071] By having this configuration, as a specific feature, it becomes possible to generate masking data Dm indicating music as a masking sound according to at least one of the pitch, level, and formant of human speech. For example, according to the pitch or formant of human speech, it is possible to determine whether the person who uttered the speech is male or female, and according to the determination result, it is possible to generate a masking sound.
[0072] Also, the example of the first aspect (the third aspect) further includes an acquisition unit 113 that acquires music data Dx indicating music. The generation unit 114 generates masking data Dm by correcting the music data Dx based on the feature data Df.
[0073] By having this configuration, it becomes possible to easily generate a masking sound by correcting the pre-stored music data Dx to generate masking data Dm indicating the masking sound.
[0074] Also, in the example of the first aspect (the fourth aspect), the above-mentioned music includes a plurality of timbres and a plurality of parts that correspond one-to-one. Also, the above-mentioned music data Dx includes a plurality of part data Dp that correspond one-to-one to the plurality of parts. Also, the generation unit 114 selects one part data Dps from among the plurality of part data Dp based on the feature data Df. Further, the generation unit 114 generates masking data Dm by correcting at least one of the pitch of the sound indicated by the one part data Dps and the level of the sound indicated by the one part data Dps based on the feature data Df.
[0075] By having this configuration, it becomes possible to generate masking data Dm indicating a masking sound by correcting at least one of the pitch and level of the sound emitted within the music indicated by the music data Dx according to the characteristics of human speech.
[0076] Also, in the example of the first aspect (the fifth aspect), the features of the voice include the formant of the voice and the level of the voice. The generation unit 114 selects one part data Dps among the plurality of part data Dp whose pitch range overlaps with the formant of the voice indicated by the feature data Df, and changes the pitch of the sound indicated by the selected one part data Dps according to the pitch of the voice indicated by the feature data Df, so as to correct the selected one part data Dps.
[0077] By having this configuration, for example, it is possible to select the part data Dps according to whether the human voice is a male voice or a female voice, and adjust the level of the selected part data Dps to match the level of the human voice.
[0078] Also, in the example of the first aspect (the sixth aspect), the features of the voice include the pitch of the voice and the level of the voice. The generation unit 114 selects one part data Dps among the plurality of part data Dp whose pitch range includes the same frequency as the pitch of the voice indicated by the feature data Df, and changes the level of the sound indicated by the selected one part data Dps according to the level of the voice indicated by the feature data Df, so as to correct the selected one part data Dps.
[0079] By having this configuration, for example, it is possible to select the part data Dps according to whether the human voice is a male voice or a female voice, and adjust the level of the selected part data Dps to match the level of the human voice.
[0080] Also, in the example of the first aspect (the seventh aspect), the features of the voice include the formant of the voice and the level of the voice. The generation unit 114 selects one part data Dps among the plurality of part data Dp whose pitch range overlaps with the formant of the voice indicated by the feature data Df, and changes the pitch of the selected one part data Dps according to the pitch of the voice indicated by the feature data Df, so as to correct the selected one part data Dps.
[0081] By having this configuration, for example, according to whether a human voice is a male voice or a female voice, part data Dp can be selected, and the pitch of the selected part data Dp can be adjusted to match the pitch of the human voice.
[0082] Also, in an example of the first aspect (eighth aspect), the feature of the voice includes the pitch of the voice. The generation unit 114 selects one part data Dps among the plurality of part data Dp that includes the same frequency as the pitch of the voice indicated by the feature data Df, and corrects the selected one part data Dps so as to change the pitch of the selected one part data Dps according to the pitch of the voice indicated by the feature data Df.
[0083] By having this configuration, for example, according to whether a human voice is a male voice or a female voice, a part can be selected, and the pitch of the selected part can be adjusted to match the pitch of the human voice.
[0084] Also, in an example of the first aspect (ninth aspect), the generation unit 114 raises or lowers the key of the selected one part data Dp in octave units.
[0085] By having this configuration, the generation unit 114 can correct only the selected part data Dps in a state where the music is established as a piece of music without changing the melody of the music indicated by the music data Dx.
[0086] Also, in an example of the first aspect (tenth aspect), the generation unit 114 raises or lowers the chord of the selected one part data Dps in semitone units.
[0087] By having this configuration, the generation unit 114 can finely adjust the pitch of the selected part data Dps.
[0088] Also, in an example of the first aspect (eleventh aspect), the masking data Dm includes the corrected one part data Dps' and the part data Dp among the plurality of part data Dp described above excluding the one part data Dps.
[0089] By having this configuration, it becomes possible to correct part data Dps indicating the performance sound of one musical instrument, and generate masking data Dm from the corrected part data Dps' and part data Dp indicating the performance sound of a musical instrument different from the musical instrument for which the part data Dp was corrected.
[0090] Also, the example of the first aspect (the 12th aspect) further includes a storage device 12 that stores music data Dx. The acquisition unit 113 reads the music data Dx from the storage device 12. When a voice signal is detected by the detection unit 111 during the period in which the acquisition unit 113 is reading the music data Dx, the generation unit 114 executes the above correction.
[0091] By having this configuration, the masking device 1 can play a piece of music that includes the performance sounds of a plurality of musical instruments in advance, and only when it senses a human voice, it can, for example, increase the performance sound of some musical instruments according to the characteristics of the voice. As a result, it becomes possible to suppress the sense of discomfort felt by the speaking human when a masking sound is suddenly output at the same time as the human speaks.
[0092] Also, the example of the first aspect (the 13th aspect) further includes a storage device 12 that stores music data Dx. The acquisition unit 113 reads the music data Dx from the storage device 12. When the detection unit 111 does not detect a voice signal, the generation unit 114A outputs a predetermined part of the plurality of part data Dp as the music data Dx. Also, when the detection unit 111 detects a voice signal, the generation unit 114A executes the above correction and outputs masking data Dm including the predetermined part data Dp and the corrected one part data Dps'.
[0093] By having this configuration, the masking device 1 can play the music indicated by certain part data Dp in advance and can insert the music indicated by other part data Dps according to the characteristics of the human voice only when the human voice is sensed. As a result, when a masking sound is suddenly output at the same time as a human speaks, it is possible to suppress the sense of discomfort felt by the speaking human.
[0094] Also, in the example of the first aspect (the 14th aspect), the music data Dx may be MIDI data.
[0095] By having this configuration, it is possible to generate masking data Dm indicating a masking sound by correcting the MIDI data as the music data Dx.
[0096] Alternatively, in the example of the first aspect (the 15th aspect), the music data Dx may be an audio signal.
[0097] By having this configuration, it is possible to generate masking data Dm indicating a masking sound by correcting the audio signal as the music data Dx.
[0098] Also, in the example of the first aspect (the 16th aspect), the generation unit 114 generates a new piece of music and generates masking data Dm corresponding to the generated piece of music.
[0099] By having this configuration, it is possible to automatically generate the melody of the masking sound.
[0100] Also, the example of the first aspect (the 17th aspect) further includes a playback device 15 that plays music based on the masking data Dm.
[0101] By having this configuration, it is possible to play the music as a masking sound.
[0102] Also, in the example of the first aspect (the 18th aspect), when the audio is detected by the detection unit 111, the playback device 15 plays music.
[0103] By having this configuration, it becomes possible to play music as masking sound in accordance with the timing of human speech.
Explanation of Signs
[0104] 11… Control device, 12… Storage device, 13… Operating device, 14… Sound collection device, 14-1… First sound collection device, 14-2… Second sound collection device, 15… Playback device, 15-1… First speaker, 15-2… Second speaker, 15-3… Third speaker, 15-4… Fourth speaker, 51~54… Seats, 71… Front right door, 72… Front left door, 73… Rear right door, 74… Rear left door, 111… Detection unit, 112… Analysis unit, 113… Acquisition unit, 114, 114A… Generation unit
Claims
1. A detection unit that detects an audio signal indicating audio from an output signal output from a microphone; An analysis unit that generates feature data indicating the characteristics of the audio by analyzing the audio signal; A generation unit that generates masking data indicating music that masks the audio based on the feature data; An acquisition unit that acquires music data indicating music; Comprising: The music includes a plurality of timbres and a plurality of parts that correspond one-to-one; The music data includes a plurality of part data that correspond one-to-one to the plurality of parts; The generation unit: Selects one piece of part data from the plurality of pieces of part data based on the feature data; Generates the masking data by correcting at least one of the pitch of the sound indicated by the one piece of part data and the level of the sound indicated by the one piece of part data based on the feature data. A masking device.
2. The masking device according to claim 1, wherein the characteristics of the audio include at least one of the pitch of the audio, the level of the audio, and the formant of the audio.
3. The characteristics of the audio include the formant of the audio and the level of the audio, The generation unit: Selects one piece of part data from the plurality of pieces of part data, the pitch range of which overlaps the formant of the audio indicated by the feature data; The masking device according to claim 1, wherein the selected one piece of part data is corrected so as to change the level of the sound indicated by the selected one piece of part data according to the level of the audio indicated by the feature data.
4. The characteristics of the audio include the pitch of the audio and the level of the audio, The generation unit: Selects one piece of part data from the plurality of pieces of part data, the pitch range of which includes the same frequency as the pitch of the audio indicated by the feature data; The masking device according to claim 1, wherein the selected one piece of part data is corrected so as to change the level of the sound indicated by the selected one piece of part data according to the level of the audio indicated by the feature data.
5. The characteristics of the audio include the formant of the audio and the pitch of the audio, The generation unit: Selects one piece of part data from the plurality of pieces of part data, the pitch range of which overlaps the formant of the audio indicated by the feature data; The masking device according to any one of claims 1 to 4, wherein the selected one part data is corrected so as to change the pitch of the sound indicated by the selected one part data according to the pitch of the voice indicated by the feature data.
6. The features of the voice include the pitch of the voice, The generation unit, selects one part data among the plurality of part data that includes a frequency whose pitch range is the same as the pitch of the voice indicated by the feature data, The masking device according to any one of claims 1 to 4, wherein the selected one part data is corrected so as to change the pitch of the sound indicated by the selected one part data according to the pitch of the voice indicated by the feature data.
7. The masking device according to claim 5 or claim 6, wherein the generation unit raises or lowers the key of the selected one part data in octave units.
8. The masking device according to any one of claims 5 to 7, wherein the generation unit raises or lowers the code of the selected one part data in semitone units.
9. The masking device according to any one of claims 1 to 8, wherein the masking data includes the corrected one part data and part data other than the one part data among the plurality of part data.
10. further comprising a storage unit for storing the music data, The acquisition unit reads the music data from the storage unit, The masking device according to any one of claims 1 to 9, wherein the generation unit executes the correction when the voice signal is detected by the detection unit during the period when the acquisition unit reads the music data.
11. further comprising a storage unit for storing the music data, The acquisition unit reads the music data from the storage unit, The generation unit, when the voice signal is not detected by the detection unit, outputs a predetermined part data among the plurality of part data as the music data, The masking device according to any one of claims 1 to 9, wherein when the voice signal is detected by the detection unit, the correction is executed and the masking data including the predetermined part data and the corrected one part data is output.
12. The masking device according to any one of claims 1 to 11, wherein the music data is MIDI data.
13. The masking device according to any one of claims 1 to 11, wherein the music data is an audio signal.
14. The masking device according to claim 1 or claim 2, wherein the generation unit generates a new piece of music and generates the masking data corresponding to the generated piece of music.
15. The masking device according to any one of claims 1 to 14, further comprising a playback unit that plays the music based on the masking data.
16. The masking device according to claim 15, wherein the playback unit plays the music when the voice is detected by the detection unit.
Citation Information
Patent Citations
Vehicle noise reduction device and method
CN108944749A
Sound output system
JP2007256606A
Encrypted data generation device, encrypted data generation method, encryption device, encryption method, and program
JP2012141524A
Signal processing device, signal processing method, and storage medium
JP2014174255A
Generation of masking signals on electronic devices
JP2014520284A