Information processing method, information processing device, and information processing program

The method generates visualized sound images using power spectra and learning models to identify specific sounds, addressing privacy concerns and reducing processing needs, enabling effective masking sound management.

WO2025150292A1PCT designated stage expired Publication Date: 2025-07-17PANASONIC INTELLECTUAL PROPERTY MANAGEMENT CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2024/042474
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-09
Filing Date
2024-12-02
Publication Date
2025-07-17

AI Technical Summary

Technical Problem

Existing technologies struggle to accurately determine whether an acquired sound includes a specific sound while ensuring privacy protection, often leading to unnecessary masking sounds being output and misjudgments.

Method used

An information processing method that generates a visualized first image of the sound based on its power spectrum, uses a learning model to compare with a second image, and determines the presence of a specific sound, reducing data volume and processing requirements to protect privacy.

Benefits of technology

Accurately identifies specific sounds while minimizing data processing and power consumption, allowing for comfortable and privacy-protected masking sound adjustments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024042474_17072025_PF_FP_ABST
    Figure JP2024042474_17072025_PF_FP_ABST
Patent Text Reader

Abstract

This information processing device: acquires a first sound; generates a first image obtained by visualizing the first sound, on the basis of a power spectrum obtained through frequency characteristic analysis of the acquired first sound; determines, by using a trained model trained by using a second image obtained by visualizing a prescribed second sound, whether the first image includes the second image; and determines whether the first sound includes the second sound, on the basis of the determination result.
Need to check novelty before this filing date? Find Prior Art

Description

Information processing method, information processing device, and information processing program

[0001] The present disclosure relates to a technique for determining whether an acquired sound contains a specific sound.

[0002] For example, Patent Document 1 discloses an audio output system in which a control device analyzes the usage status of a network and controls the output of a masking sound from an audio output device connected to the control device via the network.

[0003] Furthermore, for example, Patent Document 2 discloses a belt conveyor abnormality detection device that includes a microphone that collects the operating sounds of the rollers of the belt conveyor, and a personal computer that converts the operating sounds collected by the microphone into a visualized image, performs machine learning using AI from the visualized image, and, based on the results of the machine learning, performs an abnormality determination using AI to convert the operating sounds currently collected by the microphone into a visualized image.

[0004] However, with the above-mentioned conventional technology, it is difficult to accurately determine whether or not the acquired sound contains a specific sound while protecting privacy, and further improvement is needed.

[0005] Japanese Patent No. 5565280 Japanese Patent Application Laid-Open No. 2023-83737

[0006] The present disclosure has been made to solve the above problems, and aims to provide a technology that can accurately determine whether or not an acquired sound contains a specific sound while protecting privacy.

[0007] An information processing method according to the present disclosure is an information processing method executed by a computer, and includes acquiring a first sound, generating a first image that visualizes the first sound based on a power spectrum obtained by frequency characteristic analysis of the acquired first sound, determining whether the first image includes a second image that visualizes a predetermined second sound using a learning model trained with the second image, and determining whether the first sound includes the second sound based on the determination result.

[0008] According to the present disclosure, it is possible to accurately determine whether or not an acquired sound contains a specific sound while protecting privacy.

[0009] Fig. 1 is a diagram illustrating a configuration of a sound reproduction device according to the present embodiment. Fig. 2 is a diagram illustrating an example of a first image that visualizes a first sound including a human voice in the present embodiment. Fig. 3 is a flowchart for explaining masking processing by an information processing device according to an embodiment of the present disclosure. Fig. 4 is a diagram illustrating examples of a first sound, a first masking sound, and a second masking sound in the present embodiment.

[0010] (Findings that form the basis of the present disclosure) Conventionally, a person working alone in a space with multiple people, such as an office, may feel uncomfortable when hearing other people's voices, air conditioning noise, etc. Therefore, a continuous noise has been output into the space as a masking sound (masker sound) to mask other people's voices, air conditioning noise, etc.

[0011] In the above-mentioned Patent Document 1, for example, if a predetermined number or more of DHCP addresses have been assigned, it is determined that there are a predetermined number or more of users, and control is performed to output a masking sound from an audio output device.

[0012] However, in Patent Document 1, even if a predetermined number of users are present, that number does not necessarily mean that the predetermined number of users are conversing. Therefore, there is a risk that unnecessary masking sounds may be output even when the predetermined number of users are not conversing. Furthermore, Patent Document 1 does not determine whether the acquired sound includes a specific sound such as a user's voice.

[0013] Furthermore, since Patent Document 2 is an anomaly detection device for detecting anomalies in a belt conveyor, it does not take into consideration the protection of privacy, making it difficult for Patent Document 2 to accurately determine whether or not the acquired sound contains a specific sound while protecting privacy.

[0014] In order to solve the above problems, the following techniques are disclosed.

[0015] (1) An information processing method according to one aspect of the present disclosure is an information processing method executed by a computer, comprising: acquiring a first sound; generating a first image that visualizes the first sound based on a power spectrum obtained by frequency characteristic analysis of the acquired first sound; determining whether the first image includes a second image that visualizes a predetermined second sound using a learning model trained using a second image that visualizes the second sound; and determining whether the first sound includes the second sound based on the determination result.

[0016] According to this configuration, a first image visualizing the first sound is generated based on a power spectrum obtained by frequency characteristic analysis of the acquired first sound, and a learning model is used to determine from the first image whether the first sound includes a second sound. Here, since the first image is an image visualized based on the power spectrum obtained by frequency characteristic analysis of the first sound, which is personal information, it is difficult to identify the original first sound from the first image. Therefore, it is possible to accurately determine whether the acquired sound includes a specific sound while protecting privacy.

[0017] Furthermore, images have a smaller amount of data than sounds. With this configuration, a first image that visualizes the first sound is generated, and a determination process is performed on the first image, so the amount of calculation required can be reduced compared to when determination process is performed on the first sound, and it is possible to achieve a smaller device and lower power consumption.

[0018] (2) In the information processing method described in (1) above, generating the first image may include generating the first image based on the continuous power spectrum contained in the first sound for a predetermined period, and the first image may represent the time changes of the frequency components and sound pressure intensity of the first sound in the power spectrum using a plurality of colors.

[0019] According to this configuration, the first image, which represents the frequency components of the first sound in the power spectrum and the time change in the intensity of the sound pressure using multiple colors, is used in the determination process, making it possible to accurately determine whether the acquired sound contains a specific sound.

[0020] (3) In the information processing method described in (1) or (2) above, determining whether the first image includes the second image may include inputting the first image into the learning model and determining whether the similarity between the first image and the second image output from the learning model is greater than or equal to a threshold value.

[0021] According to this configuration, by determining whether the similarity between the first image and the second image is equal to or greater than a threshold value, it is possible to easily determine whether the first sound contains the second sound.

[0022] (4) In the information processing method described in any one of (1) to (3) above, acquiring the first sound may include acquiring the first sound at predetermined intervals, and determining whether the first sound includes the second sound may include determining that the first sound includes the second sound when it is determined that the first image includes the second image a predetermined number of times or more in succession.

[0023] There is a risk of erroneous determination if the first image is determined to contain the second image only once. Therefore, by determining that the first sound contains the second sound when the first image is determined to contain the second image a predetermined number of times or more in succession, erroneous determination can be prevented.

[0024] (5) In the information processing method described in any one of (1) to (4) above, acquiring the first sound may include acquiring the first sound at a predetermined cycle, and determining whether the first sound includes the second sound may include determining that the first sound does not include the second sound if it is determined that the first image does not include the second image a predetermined number of times in succession.

[0025] For example, when people are talking, the conversation may be interrupted, and if the first masking sound and the second masking sound are switched each time the conversation is interrupted, the switching process becomes frequent, which may cause the user to feel uncomfortable. Therefore, if it is determined that the first image does not include the second image a predetermined number of times in succession, it may be determined that the first sound does not include the second sound, and the first masking sound may be switched to the second masking sound. This reduces the number of switching processes, allowing the user to hear the first masking sound and the second masking sound without feeling uncomfortable.

[0026] (6) The information processing method according to any one of (1) to (5) above may further include performing a noise reduction process on the acquired first sound to reduce noise other than the second sound, and generating the first image may include generating the first image based on the power spectrum obtained by analyzing the frequency characteristics of the first sound after the noise reduction process has been performed.

[0027] According to this configuration, noise other than the second sound is reduced, and therefore it is possible to improve the accuracy of determining whether the first sound includes the second sound.

[0028] (7) In the information processing method described in (6) above, the noise reduction process may include reducing the noise by subtracting a pre-stored power spectrum of the noise from a power spectrum of the first sound, and further including, if it is determined that the first sound does not include the second sound, updating the pre-stored noise to the first sound that does not include the second sound.

[0029] According to this configuration, the noise contained in the first sound can be reduced by the spectral subtraction method. Furthermore, since the pre-stored noise is updated to the first sound that does not contain the second sound, the noise contained in the first sound can be reduced using noise obtained in real time.

[0030] (8) In the information processing method described in (6) or (7) above, the method may further include passing a sound of a frequency band corresponding to the second sound from the first sound after the noise reduction processing using a band-pass filter, and generating the first image may include generating the first image based on the power spectrum obtained by frequency characteristic analysis of the sound passed through the band-pass filter.

[0031] According to this configuration, after the noise reduction process is performed, only the sound in the frequency band corresponding to the second sound is extracted from the first sound, thereby further improving the accuracy of determining whether the first sound includes the second sound.

[0032] (9) In the information processing method described in any one of (1) to (8) above, the learning model may be created by machine learning using the second image that visualizes the second sound and a third image that visualizes noise other than the second sound as training data.

[0033] According to this configuration, by inputting a first image that visualizes the acquired first sound into a learning model, it is possible to determine whether the first sound contains the second sound or whether the first sound contains noise other than the second sound.

[0034] (10) In the information processing method described in any one of (1) to (8) above, after it is determined whether the first sound includes the second sound, it may further include: accepting a user input as to whether the first sound includes the second sound; and, if it is determined that the first sound includes the second sound and the user inputs that the first sound does not include the second sound, machine learning the learning model using the first image labeled with a label indicating that the second sound is not included as training data; and, if it is determined that the first sound does not include the second sound and the user inputs that the first sound includes the second sound, machine learning the learning model using the first image labeled with a label indicating that the second sound is included as training data.

[0035] If a computer determines that a first sound contains a second sound, but a user determines that the first sound does not contain the second sound, the determination result of the learning model may be incorrect. Therefore, in such a case, the learning model is trained by machine learning using a first image labeled as not containing the second sound as training data, thereby improving the determination accuracy of the learning model.

[0036] Furthermore, if the computer determines that the first sound does not contain the second sound, but the user determines that the first sound does contain the second sound, the learning model's determination result may be incorrect. Therefore, in such cases, the learning model is trained by machine learning using, as training data, first images labeled to indicate that the second sound is included, thereby improving the determination accuracy of the learning model.

[0037] (11) In the information processing method described in any one of (1) to (10) above, the second sound may be a human voice.

[0038] With this configuration, it is possible to determine whether the first sound includes a human voice.

[0039] (12) In the information processing method described in any one of (1) to (10) above, the second sound may be a non-stationary sound.

[0040] With this configuration, it is possible to determine whether the first sound includes a non-steady sound.

[0041] (13) The information processing method described in any one of (1) to (12) above may further include, when it is determined that the first sound includes the second sound, reproducing a first masking sound corresponding to the frequency or sound pressure of the second sound.

[0042] According to this configuration, the second sound is masked by the first masking sound, so that the discomfort that the second sound gives to people can be reduced.

[0043] (14) The information processing method described in (13) above may further include, when it is determined that the first sound does not include the second sound, reproducing a second masking sound corresponding to the frequency or sound pressure of noise other than the second sound.

[0044] According to this configuration, noises other than the second sound are masked by the second masking sound, so that the discomfort that noises other than the second sound give to people can be reduced.

[0045] (15) In the information processing method described in (14) above, the first masking sound and the second masking sound may be natural sounds of different types or natural sounds of different volumes.

[0046] According to this configuration, the first masking sound and the second masking sound are natural sounds of different types or different volumes, so that a comfortable space can be provided to the user.

[0047] (16) In the information processing method described in any one of (1) to (15) above, the frequency range of the power spectrum may be 1 Hz to 8 kHz.

[0048] According to this configuration, the first image can be generated based on a power spectrum in the frequency range of 1 Hz to 8 kHz, and the detection frequency is narrowed down, thereby improving the judgment accuracy of the learning model.

[0049] Furthermore, the present disclosure can be realized not only as an information processing method that executes the characteristic processes described above, but also as an information processing device having a characteristic configuration corresponding to the characteristic processes executed by the information processing method. Furthermore, the present disclosure can also be realized as a computer program that causes a computer to execute the characteristic processes included in such an information processing method. Therefore, the same effects as those of the above information processing method can also be achieved in the following other aspects.

[0050] (17) An information processing device according to another aspect of the present disclosure includes an acquisition unit that acquires a first sound; a generation unit that generates a first image that visualizes the first sound based on a power spectrum obtained by frequency characteristic analysis of the acquired first sound; and a determination unit that determines whether the first image includes a second image using a learning model trained using a second image that visualizes a predetermined second sound, and determines whether the first sound includes the second sound based on the determination result.

[0051] (18) An information processing program according to another aspect of the present disclosure causes a computer to acquire a first sound, generate a first image that visualizes the first sound based on a power spectrum obtained by frequency characteristic analysis of the acquired first sound, determine whether the first image includes the second image using a learning model trained using a second image that visualizes a predetermined second sound, and determine whether the first sound includes the second sound based on the determination result.

[0052] (19) A non-transitory computer-readable recording medium according to another aspect of the present disclosure records an information processing program, and the information processing program causes a computer to acquire a first sound, generate a first image that visualizes the first sound based on a power spectrum obtained by frequency characteristic analysis of the acquired first sound, determine whether the first image includes a second image that visualizes a predetermined second sound using a learning model trained using a second image that visualizes the second sound, and determine whether the first sound includes the second sound based on the determination result.

[0053] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. Note that each of the embodiments described below represents a specific example of the present disclosure. The numerical values, shapes, components, steps, and order of steps shown in the following embodiments are merely examples and are not intended to limit the present disclosure. Furthermore, among the components in the following embodiments, components that are not described in the independent claims that represent the highest concept are described as optional components. Furthermore, in all embodiments, the respective contents can be combined.

[0054] (Embodiment) FIG. 1 is a diagram showing the configuration of a sound reproduction device 10 according to this embodiment.

[0055] The sound reproducing device 10 shown in FIG. 1 includes a microphone 1, an information processing device 2, an amplifier 3, and a speaker 4.

[0056] The microphone 1 and the speaker 4 are provided outside the main body of the sound reproduction device 10, and the information processing device 2 and the amplifier 3 are provided inside the main body of the sound reproduction device 10. The amplifier 3 and the speaker 4 may also be provided outside the main body of the sound reproduction device 10, or the amplifier 3 and the speaker 4 may be integrated. The information processing device 2 may be a server. In this case, the information processing device 2 is connected to a terminal device equipped with the microphone 1, the amplifier 3, and the speaker 4 via a network so as to be able to communicate with each other.

[0057] The microphone 1 collects a first sound within a predetermined space. The sound reproduction device 10 is placed in a position where it can collect the sound within the predetermined space. The predetermined space is a place where multiple people can gather, such as an office, a building entrance, a building break room, a cafe, a restaurant, a hospital, or a nursing home. The first sound is environmental noise around the microphone 1. The environmental noise includes background noise, which is a steady sound such as the sound of an air conditioner, and a second sound, which is a non-steady sound such as a human voice.

[0058] The microphone 1 is connected to the information processing device 2 by wire or wirelessly so that the collected first sound can be input to the information processing device 2. The microphone 1 may be communicably connected to the information processing device 2 via a network. The microphone 1 outputs the collected first sound to the information processing device 2.

[0059] The information processing device 2 includes at least a computer system including, for example, a control program, a processing circuit such as a processor or logic circuit that executes the control program, and a recording device such as an internal memory or an accessible external memory that stores the control program. Note that the information processing device 2 may be realized, for example, by hardware implementation using the processing circuit, or by execution of a software program stored in the memory by the processing circuit or distributed from an external server, or by a combination of these hardware and software implementations.

[0060] The information processing device 2 includes a sound acquisition unit 21 , a background noise storage unit 22 , a noise reduction unit 23 , an image generation unit 24 , a sound determination unit 26 , a sound source storage unit 27 , a sound selection unit 28 , a sound reproduction unit 29 , and a background noise update unit 30 .

[0061] The sound acquisition unit 21 acquires the first sound collected by the microphone 1. The sound acquisition unit 21 acquires the first sound for a predetermined period at predetermined intervals. For example, the sound acquisition unit 21 acquires the first sound for one second every second.

[0062] The background noise storage unit 22 stores in advance background noise to be reduced from the first sound by noise reduction processing.

[0063] The noise reduction unit 23 performs noise reduction processing on the first sound acquired by the sound acquisition unit 21 to reduce background noise other than the second sound. The noise reduction processing uses a spectral subtraction method. That is, the noise reduction unit 23 reduces the background noise by subtracting the power spectrum of the background noise stored in advance in the background noise storage unit 22 from the power spectrum of the first sound.

[0064] The image generation unit 24 generates a first image that visualizes the first sound, based on the power spectrum obtained by the frequency characteristic analysis of the first sound acquired by the sound acquisition unit 21. Note that the image generation unit 24 generates the first image based on the power spectrum obtained by the frequency characteristic analysis of the first sound after the noise reduction process has been performed by the noise reduction unit 23.

[0065] The image generating unit 24 includes a power spectrum transforming unit 241 and an image transforming unit 242 .

[0066] The power spectrum transform unit 241 transforms the first sound in the time domain into a power spectrum in the frequency domain. The power spectrum transform unit 241 performs a Fourier transform on the first sound to transform the audio waveform of the first sound into a power spectrum that is obtained by decomposing the audio waveform into frequency components.

[0067] The image conversion unit 242 converts the power spectrum converted by the power spectrum conversion unit 241 into a first image.

[0068] FIG. 2 is a diagram showing an example of a first image that visualizes a first sound including a human voice in this embodiment.

[0069] As shown in Figure 2, in the first image, the horizontal axis represents time, the vertical axis represents frequency, and color represents the intensity of sound pressure. Note that color is expressed in either the RGB (red, green, and blue) or CMYK (cyan, magenta, yellow, and black) color space. The frequency range of the power spectrum is from 1 Hz to 8 kHz.

[0070] The image conversion unit 242 generates a first image based on the continuous power spectrum contained in the first sound for a predetermined period. The image conversion unit 242 generates the first image by chronologically arranging the power spectra converted by the power spectrum conversion unit 241. The first image represents the time changes in the frequency components and sound pressure intensity of the first sound in the power spectrum using multiple colors.

[0071] The sound determination unit 26 determines whether the first image includes a second image using a learning model trained using a second image that visualizes a predetermined second sound, and determines whether the first sound includes the second sound based on the determination result. The second sound is a human voice. Note that the second sound may be a non-stationary sound other than a human voice. If the sound determination unit 26 determines that the first image includes the second image, it determines that the first sound includes the second sound. Furthermore, if the sound determination unit 26 determines that the first image does not include the second image, it determines that the first sound does not include the second sound.

[0072] Whether or not the first image includes the second image can be determined based on the similarity between the first image and the second image. Here, the sound determination unit 26 inputs the first image into a learning model and determines whether or not the similarity between the first image and the second image output from the learning model is equal to or greater than a threshold. The similarity is expressed as a numerical value between 0 and 1.00. The threshold is, for example, 0.98. If the similarity output from the learning model is equal to or greater than the threshold, the sound determination unit 26 determines that the first sound includes the second sound. Furthermore, if the similarity output from the learning model is smaller than the threshold, the sound determination unit 26 determines that the first sound does not include the second sound.

[0073] The similarity may be the rate of match between the first image and the second image when comparing the first image and the second image. Alternatively, the similarity may be a prediction score that indicates the likelihood that the first image includes the second image. For example, depending on the types of the first sound and the second sound or the sound collection conditions, the rate of match between the first image and the second image may be low even if the first sound includes the second sound. Even in such a case, if an image element specific to the second sound included in the second image is also included in the first image, the prediction score may be high, making it possible to determine without omission whether the first sound includes the second sound.

[0074] Furthermore, if the first sound in the specified period includes the second sound for the entire period, the sound determination unit 26 determines that the first sound includes the second sound, and if the first sound in the specified period includes the second sound only for a portion of the period, the sound determination unit 26 determines that the first sound does not include the second sound.

[0075] The learning model is created by machine learning using, as training data, a second image that visualizes the second sound based on a power spectrum obtained by frequency characteristic analysis of the second sound, and a third image that visualizes background noise based on a power spectrum obtained by frequency characteristic analysis of background noise other than the second sound. The second sound is a human voice. A label indicating that it is a human voice is assigned to the second image, and a label indicating that it is not a human voice is assigned to the third image. Machine learning is performed on the learning model so that when the second image is input, it outputs a value "1" indicating that it is a human voice, and when the third image is input, it outputs a value "0" indicating that it is not a human voice. A plurality of second images and a plurality of third images are used to train the learning model.

[0076] The sound source storage unit 27 pre-stores a first masking sound corresponding to the frequency or sound pressure of the second sound, and a second masking sound corresponding to the frequency or sound pressure of background noise other than the second sound. The first masking sound includes a first natural sound, a second natural sound, and a third natural sound that are different from each other. The first masking sound is a sound obtained by combining the first natural sound, the second natural sound, and the third natural sound. The second masking sound includes the second natural sound and the third natural sound that are different from each other. The second masking sound is a sound obtained by combining the second natural sound and the third natural sound.

[0077] The frequency of the first natural sound is higher than the frequency of the second natural sound, which is higher than the frequency of the third natural sound. The frequency band of the first natural sound is, for example, 2 kHz to 15 kHz, the frequency band of the second natural sound is, for example, 500 Hz to 6 kHz, and the frequency band of the third natural sound is, for example, 200 Hz to 3 kHz. For example, the first natural sound is the chirping of birds, the second natural sound is the sound of spring water indicating the sound of gushing water, and the third natural sound is the sound of running water indicating the sound of flowing water. The first natural sound may also be the chirping of insects.

[0078] Because the frequency of the first natural sound is higher than the frequencies of the second and third natural sounds, the first natural sound can mask human voices (unsteady sounds).Furthermore, the second and third natural sounds can mask background noises (steady sounds) other than human voices.

[0079] The first masking sound may include the first natural sound and at least one of the second natural sound and the third natural sound. The second masking sound may include at least one of the second natural sound and the third natural sound.

[0080] The sound selection unit 28 selects either the first masking sound or the second masking sound stored in the sound source storage unit 27 based on the determination result made by the sound determination unit 26. The sound selection unit 28 selects the first masking sound when the sound determination unit 26 determines that the first sound includes the second sound. Furthermore, the sound selection unit 28 selects the second masking sound when the sound determination unit 26 determines that the first sound does not include the second sound. The sound selection unit 28 reads out either the selected first masking sound or the second masking sound from the sound source storage unit 27 and outputs it to the sound playback unit 29.

[0081] The sound reproducing unit 29 reproduces either the first masking sound or the second masking sound selected by the sound selecting unit 28. When the sound determining unit 26 determines that the first sound includes the second sound, the sound reproducing unit 29 reproduces the first masking sound corresponding to the frequency or sound pressure of the second sound. Furthermore, when the sound determining unit 26 determines that the first sound does not include the second sound, the sound reproducing unit 29 reproduces the second masking sound corresponding to the frequency or sound pressure of background noise other than the second sound. The sound reproducing unit 29 outputs either the reproduced first masking sound or the second masking sound to the speaker 4.

[0082] The sound reproducing unit 29 may set the sound pressure of the first masking sound to be greater than the sound pressure of the second masking sound. The first masking sound and the second masking sound are natural sounds with different volumes. The sound pressures of the second and third natural sounds included in the first masking sound are greater than the sound pressures of the second and third natural sounds included in the second masking sound.

[0083] The first masking sound and the second masking sound are different types of natural sounds, i.e., the first masking sound includes a first natural sound for masking the second sound and second and third natural sounds for masking background noise other than the second sound, whereas the second masking sound does not include the first natural sound but includes only the second and third natural sounds.

[0084] Furthermore, when the sound determination unit 26 determines that the first sound includes the second sound, the sound reproduction unit 29 performs temporal masking by making the sound pressure of the first masking sound greater than the sound pressure of the first sound that includes the second sound. When the sound determination unit 26 determines that the first sound includes the second sound, the sound reproduction unit 29 reproduces the first masking sound at a sound pressure greater than the sound pressure of the first sound that includes the second sound. In this way, the second sound may be masked by temporal masking.

[0085] Furthermore, when the sound determination unit 26 determines that the first sound does not include the second sound, the sound reproduction unit 29 performs temporal masking by making the sound pressure of the second masking sound greater than the sound pressure of the first sound, which includes background noise other than the second sound. When the sound determination unit 26 determines that the first sound does not include the second sound, the sound reproduction unit 29 reproduces the second masking sound at a sound pressure greater than the sound pressure of the first sound, which includes background noise other than the second sound. In this way, background noise other than the second sound may be masked by temporal masking.

[0086] The sound reproduction unit 29 may use frequency masking to mask background noise other than the second sound. That is, the sound reproduction unit 29 reproduces a second masking sound (the second natural sound and the third natural sound) having a frequency band of a critical bandwidth centered on the frequency of the background noise to be masked. The critical bandwidth CB is calculated using the frequency fc of the masking target according to the following equation (1):

[0087] CB=25+75{1+1.4(fc / 1000) 2} 0.69 ...(1)

[0088] By reproducing the second masking sounds (the second natural sound and the third natural sound) having a frequency band of a critical bandwidth centered on the frequency of the background noise other than the second sound, the background noise can be masked.

[0089] When the sound determination section 26 determines that the first sound does not include the second sound, the background noise update section 30 updates the background noise stored in advance in the background noise storage section 22 to the first sound that does not include the second sound.

[0090] The amplifier 3 amplifies the first masking sound or the second masking sound input from the sound reproducing unit 29 of the information processing device 2 .

[0091] The speaker 4 outputs the first masking sound or the second masking sound amplified by the amplifier 3 .

[0092] Next, the masking process performed by the information processing device 2 according to the embodiment of the present disclosure will be described.

[0093] 3 is a flowchart illustrating the masking process performed by the information processing device 2 according to the embodiment of the present disclosure. In the masking process shown in FIG. 3, the second sound is a human voice.

[0094] First, in step S1, the sound acquisition unit 21 acquires a first sound collected by the microphone 1. The masking process shown in Fig. 3 is performed at each predetermined period for acquiring the first sound. If the predetermined period is, for example, one second, the masking process shown in Fig. 3 is performed every second.

[0095] Next, in step S2, the noise reduction unit 23 performs noise reduction processing on the first sound acquired by the sound acquisition unit 21 to reduce background noise other than human voices.

[0096] Next, in step S3, the power spectrum conversion unit 241 converts the first sound into a power spectrum.

[0097] Next, in step S4, the image conversion unit 242 converts the power spectrum converted by the power spectrum conversion unit 241 into a first image.

[0098] Next, in step S5, the sound determination unit 26 inputs the first image into the trained learning model, determines whether the similarity between the first image and the second image output from the learning model is a threshold value, and determines whether the first sound includes a human voice based on the determination result.

[0099] If it is determined that the first sound includes a human voice (YES in step S5), then in step S6, the sound selector 28 selects a first masking sound for masking the human voice.

[0100] On the other hand, if it is determined that the first sound does not include a human voice (NO in step S5), in step S7, the sound selection unit 28 selects a second masking sound for masking background noise other than the second sound.

[0101] Next, in step S8, the background noise update unit 30 updates the background noise stored in advance in the background noise storage unit 22 to the acquired first sound that does not include the second sound.

[0102] Next, in step S9, the sound reproducing unit 29 reproduces the first masking sound or the second masking sound selected by the sound selecting unit 28. The sound reproducing unit 29 outputs the reproduced first masking sound or the second masking sound to the speaker 4 via the amplifier 3. The amplifier 3 amplifies the first masking sound or the second masking sound reproduced by the sound reproducing unit 29. The speaker 4 outputs the first masking sound or the second masking sound.

[0103] According to this embodiment, a first image visualizing the first sound is generated based on a power spectrum obtained by frequency characteristic analysis of the acquired first sound, and a learning model is used to determine from the first image whether the first sound includes a second sound. Here, since the first image is an image visualized based on the power spectrum obtained by frequency characteristic analysis of the first sound, which is personal information, it is difficult to identify the original first sound from the first image. Therefore, it is possible to accurately determine whether the acquired sound includes a specific sound while protecting privacy.

[0104] Furthermore, images have a smaller amount of data than sounds. According to this embodiment, a first image that visualizes the first sound is generated, and a determination process is performed on the first image, so the amount of calculation required can be reduced compared to when determination process is performed on the first sound, and it is possible to achieve a smaller device and lower power consumption.

[0105] Next, the relationship between the first sound, the first masking sound, and the second masking sound in this embodiment will be described.

[0106] 4 is a diagram showing an example of the first sound, the first masking sound, and the second masking sound in the present embodiment, in which the horizontal axis represents time (t) and the vertical axis represents sound pressure (dB).

[0107] As shown in Figure 4, the first sound, which is environmental noise, is divided into periods 102 and 104 that include human voices and periods 101, 103, and 105 that do not include human voices. The first sound in periods 102 and 104 includes human voices and background noise other than human voices. The first sound in periods 101, 103, and 105 includes only background noise other than human voices. Note that in Figure 4, the waveform representing human voices is represented by a white line.

[0108] During periods 101, 103, and 105 that do not include human voices, the speaker 4 outputs a second masking sound that is a combination of a second natural sound, which is the sound of spring water, and a third natural sound, which is the sound of running water. The second masking sound masks background noise other than human voices.

[0109] On the other hand, during periods 102 and 104 that include human voices, the speaker 4 outputs a first masking sound that is a combination of a first natural sound, which is a bird song, a second natural sound, which is the sound of spring water, and a third natural sound, which is the sound of running water. The first masking sound masks the human voice and background noise other than the human voice. In particular, the first natural sound included in the first masking sound changes the human voice from a meaningful sound to a meaningless sound, thereby protecting the privacy of people in the vicinity.

[0110] Furthermore, the sound pressures of the first natural sound, the second natural sound, and the third natural sound in periods 102 and 104 are greater than the sound pressures of the second natural sound and the third natural sound in periods 101, 103, and 105.

[0111] When switching from the second masking sound to the first masking sound, the sound replay unit 29 fades out the second masking sound and fades in the first masking sound. When switching from the first masking sound to the second masking sound, the sound replay unit 29 fades out the first masking sound and fades in the second masking sound. This makes it possible to switch between the first masking sound and the second masking sound without giving the user a sense of discomfort.

[0112] In the present embodiment, the information processing device 2 does not necessarily have to include the background noise storage unit 22 , the noise reduction unit 23 , and the background noise update unit 30 .

[0113] The information processing device 2 may further include a band-pass filter. The band-pass filter passes sounds in a frequency band corresponding to the first sound to the second sound after noise reduction processing by the noise reduction unit 23. For example, the band-pass filter passes sounds in a frequency band ranging from 100 Hz to 4 kHz or 200 Hz to 4 kHz corresponding to human voices. In this case, the image generation unit 24 may generate the first image based on a power spectrum obtained by analyzing the frequency characteristics of the sound that has passed through the band-pass filter.

[0114] Furthermore, in the present embodiment, the sound determination unit 26 determines that the first sound includes the second sound when it determines that the first image includes the second image, but the present disclosure is not particularly limited to this. The sound acquisition unit 21 may acquire the first sound at predetermined intervals. Then, the sound determination unit 26 may determine that the first sound includes the second sound when it determines that the first image includes the second image a predetermined number of times or more consecutively. The predetermined number of times is, for example, three times. This makes it possible to prevent the first masking sound from being reproduced when a human voice is momentarily uttered or when it is erroneously determined that the first image includes the second image.

[0115] Furthermore, the sound determination unit 26 may determine that the first sound does not include the second sound when it has determined that the first image does not include the second image a predetermined number of times or more consecutively. The predetermined number of times is, for example, five times. For example, when people are talking, the conversation may be interrupted. If the first masking sound and the second masking sound are switched each time the conversation is interrupted, the switching process may become frequent, which may cause the user to feel uncomfortable. Therefore, when it has been determined that the first image does not include the second image a predetermined number of times or more consecutively, it may be determined that the first sound does not include the second sound, and the first masking sound may be switched to the second masking sound. This reduces the number of switching processes, allowing the user to hear the first masking sound and the second masking sound without feeling uncomfortable.

[0116] In the present embodiment, the types of the first masking sound and the second masking sound are predetermined, but the present disclosure is not limited to this. The user may determine the types of the first masking sound and the second masking sound. In this case, the sound reproduction device 10 may further include an input unit. The information processing device 2 may further include a sound source type selection unit.

[0117] The input unit is, for example, a keyboard, a mouse, a touch panel, or a microphone, and accepts information input by a user. The input unit may accept a user's selection of natural sounds to be used as the first masking sound and the second masking sound from among a plurality of different types of natural sounds. The sound source type selection unit may store the first masking sound and the second masking sound, including the natural sound selected by the user, in the sound source storage unit 27.

[0118] In addition, in the present embodiment, a trained learning model is used to determine the first image, but the present disclosure is not particularly limited to this, and the information processing device 2 may train the learning model while playing the first masking sound or the second masking sound. In this case, the sound playback device 10 may further include an input unit. Furthermore, the information processing device 2 may further include a learning unit.

[0119] The input unit is, for example, a keyboard, a mouse, a touch panel, or a microphone, and accepts information input by a user. After it is determined whether the first sound includes the second sound, the input unit may accept an input by the user regarding whether the first sound includes the second sound. That is, after it is determined whether the first sound includes the second sound and the first masking sound or the second masking sound is played, the input unit may accept an input by the user regarding whether the first sound includes the second sound. The user listens to the first masking sound or the second masking sound output from the speaker 4 and inputs whether the first sound includes the second sound.

[0120] If the first sound is determined to include the second sound and then the user inputs that the first sound does not include the second sound, the learning unit may machine-train the learning model using, as training data, a first image labeled with a label indicating that the second sound is not included. That is, if the first sound is determined to include the second sound and the first masking sound is output, but the user determines that the first sound does not include the second sound, the learning model may make an erroneous judgment. Therefore, in such a case, the learning model is machine-trained using, as training data, a first image labeled with a label indicating that the second sound is not included, thereby improving the judgment accuracy of the learning model.

[0121] Furthermore, if the first sound is determined not to include the second sound and then the user inputs that the first sound includes the second sound, the learning unit may machine-train the learning model using, as training data, a first image labeled with a label indicating that the second sound is included. That is, if the first sound is determined not to include the second sound and the second masking sound is output, but the user determines that the first sound includes the second sound, the learning model may make an erroneous judgment. Therefore, in such a case, the learning model is machine-trained using, as training data, a first image labeled with a label indicating that the second sound is included, thereby improving the judgment accuracy of the learning model.

[0122] Furthermore, the sound reproduction device 10 according to the present embodiment may be provided inside an autonomous vehicle. The space inside the autonomous vehicle includes a first space in which passengers other than the driver are present, and a second space in which the driver is present. The sound reproduction device 10 is provided in each of the first space and the second space.

[0123] For example, in the first space, the second sound includes at least one of the following: the sound of passing through road joints, the sound of a car horn, the siren of an emergency vehicle such as an ambulance, the exhaust sound of another vehicle such as a truck or a motorcycle, the sound of an air conditioner inside the vehicle, the sound of a car navigation system, the sound of the driver talking, and the sound of windshield wipers. The sound determination unit 26 inputs a first image that visualizes the first sound into a learning model and determines whether the first sound includes at least one of the above multiple sounds based on the output result from the learning model. In this case, the sound determination unit 26 may use one learning model to determine whether the first sound includes each of the above multiple sounds, or may use multiple learning models corresponding to each of the above multiple sounds to determine whether the first sound includes each of the above multiple sounds. When it is determined that the first sound includes at least one of the above multiple sounds, the sound playback unit 29 plays a first masking sound corresponding to the frequency or sound pressure of at least one of the above multiple sounds. On the other hand, if it is determined that the first sound does not include the above-mentioned multiple sounds, the sound reproducing unit 29 may reproduce a second masking sound corresponding to the frequency or sound pressure of noise other than the multiple sounds.

[0124] Furthermore, for example, in the second space, the second sound includes at least one of the exhaust sound of another vehicle, such as a truck or motorcycle, the sound of an air conditioner inside the vehicle, and the sound of windshield wipers. The sound determination unit 26 inputs a first image that visualizes the first sound into a learning model and determines whether the first sound includes at least one of the above-mentioned multiple sounds based on the output result from the learning model. In this case, the sound determination unit 26 may determine whether the first sound includes each of the above-mentioned multiple sounds using one learning model, or may determine whether the first sound includes each of the above-mentioned multiple sounds using multiple learning models corresponding to each of the above-mentioned multiple sounds. If it is determined that the first sound includes at least one of the above-mentioned multiple sounds, the sound reproduction unit 29 reproduces a first masking sound corresponding to the frequency or sound pressure of at least one of the above-mentioned multiple sounds. On the other hand, if it is determined that the first sound does not include the above-mentioned multiple sounds, the sound reproduction unit 29 may reproduce a second masking sound corresponding to the frequency or sound pressure of noise other than the above-mentioned multiple sounds.

[0125] In each of the above embodiments, each component may be configured with dedicated hardware, or may be realized by executing a software program suitable for that component. Each component may be realized by a program execution unit such as a CPU or processor reading and executing a software program recorded on a recording medium such as a hard disk or semiconductor memory. Furthermore, the program may be executed by another independent computer system by recording the program on a recording medium and transferring it, or by transferring the program via a network.

[0126] Some or all of the functions of the device according to the embodiment of the present disclosure are typically realized as an LSI (Large Scale Integration), which is an integrated circuit. These may be individually integrated into a single chip, or some or all of them may be integrated into a single chip. Furthermore, the integrated circuit is not limited to an LSI, and may be realized using a dedicated circuit or a general-purpose processor. It is also possible to use an FPGA (Field Programmable Gate Array), which can be programmed after LSI manufacturing, or a reconfigurable processor, which can reconfigure the connections and settings of circuit cells within an LSI.

[0127] Furthermore, some or all of the functions of the device according to the embodiment of the present disclosure may be realized by a processor such as a CPU executing a program.

[0128] Furthermore, all the numbers used above are merely examples to specifically explain the present disclosure, and the present disclosure is not limited to the numbers used as examples.

[0129] The order in which the steps are executed in the above flowchart is merely an example for specifically explaining the present disclosure, and other orders may be used as long as similar effects are obtained. Also, some of the steps may be executed simultaneously (in parallel) with other steps.

[0130] The technology disclosed herein is useful as a technology for determining whether an acquired sound contains a specific sound, since it can accurately determine whether an acquired sound contains a specific sound while protecting privacy.

Claims

1. An information processing method executed by a computer, comprising: obtaining a first sound; generating a first image that visualizes the first sound based on a power spectrum obtained by frequency characteristic analysis of the obtained first sound; determining whether the first image includes the second image using a learning model learned by a second image that visualizes a predetermined second sound, and determining whether the first sound includes the second sound based on the determination result.

2. The generation of the first image includes generating the first image based on continuous power spectra included in the first sound for a predetermined period, and the first image represents temporal changes in frequency components and sound pressure intensities of the first sound in the power spectrum using a plurality of colors. The information processing method according to claim 1.

3. The determination of whether the first image includes the second image includes inputting the first image into the learning model and determining whether a similarity between the first image and the second image output from the learning model is equal to or greater than a threshold value. The information processing method according to claim 1 or 2.

4. The obtaining of the first sound includes obtaining the first sound at predetermined intervals, and the determination of whether the first sound includes the second sound includes determining that the first sound includes the second sound when it is determined that the first image includes the second image continuously for a predetermined number of times or more. The information processing method according to claim 1 or 2.

5. The obtaining of the first sound includes obtaining the first sound at predetermined intervals, and the determination of whether the first sound includes the second sound includes determining that the first sound does not include the second sound when it is determined that the first image does not include the second image continuously for a predetermined number of times or more. The information processing method according to claim 1 or 2.

6. Further, it includes performing noise reduction processing for reducing noise other than the second sound on the obtained first sound, and the generation of the first image includes generating the first image based on the power spectrum obtained by frequency characteristic analysis of the first sound after the noise reduction processing is performed. The information processing method according to claim 1 or 2.

7. The noise reduction process includes reducing the noise by subtracting the power spectrum of the noise, which is stored in advance, from the power spectrum of the first sound. Further, when it is determined that the first sound does not include the second sound, the noise stored in advance is updated to the first sound that does not include the second sound. The information processing method according to claim 6.

8. Further, it includes passing, through a band-pass filter, the sound in the frequency band corresponding to the second sound from the first sound after the noise reduction process is performed. The generation of the first image includes generating the first image based on the power spectrum obtained by analyzing the frequency characteristics of the sound that has passed through the band-pass filter. The information processing method according to claim 6.

9. The learning model is created by machine learning using, as teacher data, the second image visualizing the second sound and the third image visualizing the noise other than the second sound. The information processing method according to claim 1 or 2.

10. Further, after determining whether the first sound includes the second sound, it includes receiving an input from the user as to whether the first sound includes the second sound. Further, after it is determined that the first sound includes the second sound, when the user inputs that the first sound does not include the second sound, the learning model is machine-learned using, as teacher data, the first image labeled to indicate that the second sound is not included. Further, after it is determined that the first sound does not include the second sound, when the user inputs that the first sound includes the second sound, the learning model is machine-learned using, as teacher data, the first image labeled to indicate that the second sound is included. The information processing method according to claim 1 or 2 including the above.

11. The second sound is a human voice. The information processing method according to claim 1 or 2.

12. The second sound is a non-stationary sound. The information processing method according to claim 1 or 2.

13. Further, when it is determined that the first sound includes the second sound, it includes reproducing a first masking sound according to the frequency or sound pressure of the second sound. The information processing method according to claim 1 or 2.

14. The information processing method according to claim 13, further comprising: when it is determined that the first sound does not include the second sound, playing a second masking sound according to the frequency or sound pressure of noise other than the second sound.

15. The information processing method according to claim 14, wherein the first masking sound and the second masking sound are natural sounds of different types from each other or natural sounds of different volumes from each other.

16. The information processing method according to claim 1 or 2, wherein the frequency range of the power spectrum is in the range of 1 Hz to 8 kHz.

17. An information processing apparatus comprising: an acquisition unit that acquires a first sound; a generation unit that generates a first image obtained by visualizing the first sound based on a power spectrum obtained by analyzing the frequency characteristics of the acquired first sound; and a determination unit that determines whether the first image includes the second image by using a learning model learned by a second image obtained by visualizing a predetermined second sound, and determines whether the first sound includes the second sound based on the determination result.

18. An information processing program that causes a computer to function so as to acquire a first sound, generate a first image obtained by visualizing the first sound based on a power spectrum obtained by analyzing the frequency characteristics of the acquired first sound, determine whether the first image includes the second image by using a learning model learned by a second image obtained by visualizing a predetermined second sound, and determine whether the first sound includes the second sound based on the determination result.

Citation Information

Patent Citations

  • Apparatus and method for detecting voice activity period

    JP2007094388A

  • Identification device and utterance detector

    JP2011170266A

  • Work environment improvement system

    JP2017146517A

  • Ear-mounted device and reproduction method

    WO2022259589A1