Control method, control device, and program
A system using multiple microphones to detect user movement and adjust sound output ensures effective feedback by dynamically switching to the nearest speaker, addressing the issue of users moving out of audio range.
Patent Information
- Application Number
- JP2021052696
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-03-26
- Publication Date
- 2025-12-23
- Estimated Expiration
- 2041-03-26
AI Technical Summary
Existing systems fail to provide effective feedback to users who have moved away from the location of a home console, as they rely on audio output that the user cannot hear.
A system that utilizes multiple microphones to detect user movement and position, selecting the appropriate speaker to output sound based on the user's location and adjusting output as the user moves.
Ensures that users receive audio feedback even when they have moved away from the initial location by dynamically switching the sound output to the nearest speaker.
Smart Images

Figure 0007790028000001 
Figure 0007790028000002 
Figure 0007790028000003
Abstract
Description
[Technical Field]
[0001] The present invention relates to a control method, a control device, and a program. [Background technology]
[0002] A technique has been disclosed for an information processing device that can respond in accordance with a user's intention (for example, Patent Document 1). [Prior art documents] [Patent documents]
[0003] [Patent Document 1] International Publication No. 2019 / 082630 Summary of the Invention [Problem to be solved by the invention]
[0004] In the technology described in Patent Document 1, a user interacts with a home console installed in a predetermined location, and feedback from the home console is output as audio. For example, if the user says, "Please heat up the bath," and the bath is not going to be finished heating by the time the user gets home, feedback to that effect is given to the user. However, it is considered that the user often moves away from the location after giving instructions to the home console. In such cases, there is a problem in that the user is no longer near the home console and is unable to hear feedback from the home console.
[0005] The present invention has been made in consideration of such circumstances, and its purpose is to provide a control method, a control device, and a program that can select a speaker to output sound to so that the user can hear the sound that is notified to the user even when the user moves. [Means for solving the problem]
[0006] In order to solve the above-mentioned problem, one aspect of the present invention is a method for detecting a sound by collecting a detected sound by a plurality of microphones provided in the plurality of spaces. repetition a control method for detecting a user's movement from the detected sound, extracting a characteristic sound from the detected sound that is attributable to the user's position or movement, identifying a first space in which the user is present based on the characteristic sound, selecting a speaker from the plurality of speakers that is installed near the first space in which the identified user is present, outputting sound from the selected speaker, determining whether the user has moved from the first space to a second space different from the first space based on a time-series change in the volume of the characteristic sound, and if the user has moved to the second space, outputting sound from a speaker installed near the second space to which the user has moved, and stopping the output of sound that was being output from a speaker installed near the first space from which the user moved.
[0007] Another aspect of the present invention is a control device that controls sounds to be output from a plurality of speakers provided in a plurality of spaces, and detects sounds collected by a plurality of microphones provided in the plurality of spaces. repetition an extraction unit that extracts a characteristic sound caused by the position or movement of the user from the detected sound acquired by the acquisition unit; an identification unit that identifies a first space in which the user is present based on the characteristic sound; a selection unit that selects from the plurality of speakers a speaker installed near the first space in which the user identified by the identification unit is present; and an output unit that outputs sound from the speaker selected by the selection unit, wherein the identification unit determines whether the user has moved from the first space to a second space different from the first space based on a time-series change in the volume of the characteristic sound; and when the user has moved to the second space, the output unit outputs sound from a speaker installed near the second space where the user has moved, and stops the output of sound that was being output from a speaker installed near the first space from which the user moved.
[0008] In one aspect of the present invention, a computer is provided as a control device for controlling sounds to be output from a plurality of speakers provided in each of a plurality of spaces, and the computer is configured to receive detected sounds collected by a plurality of microphones provided in the plurality of spaces. repetition a selection means for selecting from the plurality of speakers a speaker installed near the first space where the user identified by the identification means is present; and an output means for outputting sound from the speaker selected by the selection means, wherein the identification means determines whether the user has moved from the first space to a second space different from the first space based on a time-series change in the volume of the characteristic sound; and wherein, when the user has moved to the second space, the output means outputs sound from a speaker installed near the second space where the user has moved, and stops the output of sound that was being output from a speaker installed near the first space where the user originated. [Effects of the Invention]
[0009] According to the present invention, even when a user moves, it is possible to select a speaker to which a sound is to be output so that the user can hear the sound that is notified to the user. [Brief explanation of the drawings]
[0010] [Figure 1] 1 is a diagram showing an overview of a speaker system 1 according to an embodiment. [Figure 2] 1 is a block diagram showing an example of the configuration of a speaker system 1 according to an embodiment. [Figure 3] 2 is a block diagram showing an example of the configuration of a control device 20 in the embodiment. FIG. [Figure 4] FIG. 2 is a diagram showing an example of continuous sound information 220 in the embodiment. [Figure 5]FIG. 2 is a diagram showing an example of informative sound information 221 in the embodiment. [Figure 6] FIG. 2 is a diagram showing an example of installation information 222 in the embodiment. [Figure 7] FIG. 2 is a diagram showing an example of installation information 222 in the embodiment. [Figure 8] 10 is a diagram illustrating a process performed by a specification unit 231 in the embodiment. FIG. [Figure 9] 10 is a diagram illustrating a process performed by a specification unit 231 in the embodiment. FIG. [Figure 10] 10A and 10B are diagrams illustrating stationary sounds according to an embodiment. [Figure 11] 10A and 10B are diagrams illustrating characteristic sounds according to an embodiment. [Figure 12] 3 is a diagram illustrating a process performed by a control device 20 in the embodiment. FIG. [Figure 13] 3 is a diagram illustrating a process performed by a control device 20 in the embodiment. FIG. [Figure 14] 10 is a flowchart illustrating a flow of processing performed by a control device 20 of the embodiment. [Figure 15] FIG. 10 is a diagram illustrating a process for generating a learning model according to an embodiment. [Figure 16] FIG. 10 is a diagram illustrating additional learning according to a first modified example of the embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0011] Hereinafter, an embodiment of the present invention will be described with reference to the drawings.
[0012] FIG. 1 is a diagram illustrating an overview of a speaker system 1 according to an embodiment. As illustrated in the example of FIG. 1, the speaker system 1 is applied to a space where a user lives, such as a house H. When the speaker system 1 is applied to the house H, the house H is provided with a plurality of sensors 10 (sensors 10-1 to 10-3), a control device 20, and a plurality of speakers 30 (speakers 30-1 to 30-3). Each of the plurality of sensors 10 is provided in a different space in the house H. Specifically, the sensors 10 are provided in the entrance, kitchen, living room, bedroom, etc. of the house H. Each of the plurality of speakers 30 is provided in a different space in the house H. Both the sensor 10 and the speaker 30 may be provided in the same space, or only one of the sensor 10 and the speaker 30 may be provided in a space such as the entrance, kitchen, living room, or bedroom of the house H. The number of sensors 10 connected to the control device 20 and the number of speakers 30 may be determined arbitrarily.
[0013] Fig. 2 is a block diagram showing an example of the configuration of a speaker system 1 according to an embodiment. As shown in Fig. 2, the speaker system 1 includes, for example, a plurality of sensors 10, a control device 20, and a plurality of speakers 30. In the speaker system 1, the plurality of sensors 10, the control device 20, and the plurality of speakers 30 are communicatively connected to each other via a communication network NW.
[0014] The communication network NW may be a wide area network, i.e., a WAN (Wide Area Network), the Internet, or a combination of these, or may be a communication network that is connected to enable communication via a wired connection such as a cable, or a wireless connection such as a wireless LAN.
[0015] The sensor 10 includes, for example, a sensor unit and a communication unit. The sensor unit acquires information that can detect the position and movement of the user. The communication unit transmits the information acquired by the sensor unit to the control device 20.
[0016] In this embodiment, the sensor unit is a microphone. When the sensor unit is a microphone, the sensor unit collects sound propagating in the space in which the sensor unit is provided. This makes it possible to detect the position and movement of a user by analyzing the sound collected by the microphone. A method for detecting the position and movement of a user based on the sound collected by the microphone will be described in detail later.
[0017] The sensor unit is not limited to a microphone. The sensor unit may be any sensor capable of acquiring information capable of detecting the user's position and movement. For example, the sensor unit may be an image sensor, an infrared sensor, a temperature sensor, an optical sensor, or the like. For example, if the sensor unit is an image sensor, the sensor unit captures an image of the space in which the sensor 10 is installed. This makes it possible to detect the user's position and movement by analyzing the image. Furthermore, if the sensor unit is an infrared sensor, a temperature sensor, an optical sensor, or the like, the sensor unit detects the temperature detected by infrared rays or a thermometer, or the phase difference between irradiated light and the reflected light thereof. This makes it possible to detect the user's position and movement based on changes in temperature, phase difference, or the like.
[0018] The speaker 30 is connected to the control device 20 and outputs sound based on the control of the control device 20. The speaker 30 may be any speaker device as long as it can output sound at least based on the control of the control device 20.
[0019] The control device 20 is a computer device such as a PC (Personal Computer) or a server device. The control device 20 receives information acquired by each of the sensors 10. The control device 20 identifies the location of the user based on the acquired information. The control device 20 outputs sound from a speaker 30 provided near the identified location where the user is located.
[0020] 3 is a block diagram showing an example of the configuration of the control device 20 in the embodiment. The control device 20 includes, for example, a communication unit 21, a storage unit 22, and a control unit 23. The communication unit 21 is realized by, for example, a general-purpose communication IC (Integrated Circuit). The communication unit 21 communicates with the sensor 10 and the speaker 30 via a communication network NW.
[0021] The storage unit 22 is realized by, for example, a storage device (a storage device having a non-transitory storage medium) such as an HDD (Hard Disk Drive) or a flash memory, or a combination of these. The storage unit 22 stores programs for implementing each component of the control device 20, variables used when executing the programs, and various types of information. The storage unit 22 stores, for example, steady sound information 220, alarm sound information 221, installation information 222, and trained model information 223. The details of the information stored in the storage unit 22 will be described later.
[0022] The control unit 23 is realized by causing a CPU provided as hardware in the control device 20 to execute a program. The control unit 23 includes, for example, an acquisition unit 230, an identification unit 231, a selection unit 232, an output unit 233, a device control unit 234, and a learning unit 235.
[0023] The acquisition unit 230 acquires information acquired by the sensor 10 via the communication unit 21 and outputs the acquired information to the identification unit 231.
[0024] The identification unit 231 identifies the position of the user based on the information acquired from the acquisition unit 230. A method by which the identification unit 231 identifies the position of the user will be described in detail later. The identification unit 231 outputs information indicating the identified position of the user to the selection unit 232.
[0025] Furthermore, the identification unit 231 identifies an alert sound based on the information acquired from the acquisition unit 230. An alert sound is a sound that should be notified to the user, such as a ringtone from an intercom. The identification unit 231 identifies, as an alert sound, a sound that has frequency characteristics similar to those of a sound that has been stored in advance in the storage unit 22 as alert sound information 221. The identification unit 231 is an example of a "determination unit."
[0026] The selection unit 232 selects a speaker 30 that is close to the user's location based on the information indicating the user's location acquired from the identification unit 231. The selection unit 232, for example, refers to the installation information 222 based on the information indicating the user's location. The installation information 222 is information indicating the locations where the speakers 30 are installed. The selection unit 232 compares the user's location with each of the locations where the speakers 30 are installed, and selects a speaker 30 that is close to the user's location. The selection unit 232 outputs information indicating the selected speaker 30 to the output unit 233.
[0027] The output unit 233 causes the speaker 30 indicated in the information to output a sound based on the information indicating the speaker 30 acquired from the selection unit 232. Here, the sound that the output unit 233 causes the speaker 30 to output may be any sound that should notify the user. For example, the output unit 233 causes the speaker 30 to output an alarm sound such as an intercom, music that the user is listening to, or the like.
[0028] Furthermore, the output unit 233 may set any of the multiple speakers 30 that is not outputting sound to a standby mode.
[0029] The device control unit 234 performs overall control of the control device 20. For example, the device control unit 234 outputs information received by the communication unit 21, i.e., information acquired by the sensor 10, to the acquisition unit 230. The device control unit 234 also outputs information output by the output unit 233, i.e., information indicating the sound to be output from the speaker 30, to the communication unit 21. As a result, sound is output from the speaker 30.
[0030] The learning unit 235 generates a trained model by performing machine learning on a machine learning model. The learning unit 235 generates a model that estimates the user's position (hereinafter referred to as a position estimation model). The position estimation model is a model that estimates the user's position from a sound collected by a microphone serving as the sensor 10. A method by which the learning unit 235 creates a position estimation model will be described in detail later.
[0031] Furthermore, the learning unit 235 generates a model for estimating a characteristic sound (hereinafter referred to as a characteristic sound estimation model). The characteristic sound estimation model is a model for estimating whether or not a characteristic sound is included in a sound. The characteristic sound here is a characteristic sound resulting from the position or movement of a user. The characteristic sound may be a characteristic sound of an unspecified user, or may be a characteristic sound of a specific user.
[0032] The characteristic sounds of unspecified users are sounds that are characteristic of the positions and movements of users but cannot identify individual users when multiple users reside in the house H. For example, the characteristic sounds of unspecified users include the sound of a living room door being opened and closed, and the startup and operating sounds that are generated when a television or other device is operated.
[0033] The characteristic sound of a particular user is a sound that is characteristic of the position or movement of each user when multiple users reside in the house H. For example, the characteristic sound of a particular user is the sound of the particular user speaking, the footsteps of the particular user, etc.
[0034] The characteristic sound estimation model outputs an estimation result of whether or not a sound input to the model is a characteristic sound that indicates the position or movement of a specific user. The method by which the learning unit 235 creates the characteristic sound estimation model will be described in detail later.
[0035] Furthermore, the learning unit 235 generates a model for estimating an informative sound (hereinafter referred to as an informative sound estimation model). The informative sound estimation model is a model for estimating a sound to be notified to the user from a sound collected by a microphone serving as the sensor 10. A method by which the learning unit 235 creates the informative sound estimation model will be described in detail later.
[0036] FIG. 4 is a diagram showing an example of steady sound information 220 in an embodiment. The steady sound information 220 is information related to steady sounds. Steady sounds are sounds that occur steadily in a space, regardless of whether a user is present in the space. For example, if the space is a kitchen, the rotating sound of a kitchen ventilation fan that is constantly running is a steady sound. The steady sound information 220 includes items such as a steady sound number and frequency characteristics. The steady sound number is identification information such as a number that uniquely identifies a steady sound. The frequency characteristics are information indicating the frequency characteristics of the steady sound identified by the steady sound number. The frequency characteristics are information indicating, for example, the magnitude (gain) of each frequency band of the sound components contained in the steady sound.
[0037] FIG. 5 is a diagram showing an example of the alarm sound information 221 in the embodiment. The alarm sound information 221 is information related to an alarm sound. An alarm sound is a sound that should be notified to the user. For example, a ring tone that sounds when an intercom is operated is an alarm sound. The alarm sound information 221 includes items such as an alarm sound number and frequency characteristics. The alarm sound number is identification information such as a number that uniquely identifies the alarm sound. The frequency characteristics is information indicating the frequency characteristics of the alarm sound identified by the alarm sound number.
[0038] The notification sound information 221 may also be generated for each user. For example, if a parent and child live in the house H, the child's voice will be the notification sound for the parent.
[0039] 6 and 7 are diagrams showing examples of installation information 222 in the embodiment. The installation information 222 is information indicating the installation positions of the sensor 10 and the speaker 30 installed in the house H. Fig. 6 shows an example of installation information 222A indicating the installation position of the sensor 10. Fig. 7 shows an example of installation information 222B indicating the installation position of the speaker 30.
[0040] 6 includes items such as a sensor number and an installation location. The sensor number is identification information such as a number that uniquely identifies the sensor 10. The installation location is information that indicates the installation location of the sensor 10 identified by the sensor number.
[0041] 7 includes items such as a speaker number and an installation position. The speaker number is identification information such as a number that uniquely identifies the speaker 30. The installation position is information that indicates the installation position of the speaker 30 identified by the speaker number.
[0042] (How to determine the user's location) Here, a method for the identification unit 231 to identify the user's position will be described with reference to Figures 8 and 9. Figures 8 and 9 are diagrams for explaining the processing performed by the identification unit 231 in the embodiment.
[0043] As shown in FIG. 8, the identification unit 231 includes, for example, a plurality of steady sound reduction units 2310 (steady sound reduction units 2310, 2311, 2312), a characteristic sound extraction unit 2313, and a position identification unit 2314.
[0044] Each of the plurality of steady sound reduction units 2310 (steady sound reduction units 2310, 2311, 2312) receives input of information detected by each of the plurality of sensors 10. In the following description, when there is no need to distinguish between the plurality of steady sound reduction units 2310 (steady sound reduction units 2310, 2311, 2312), they will be referred to as steady sound reduction unit 231N.
[0045] The steady sound reduction unit 231N reduces steady sounds from sounds detected by the sensor 10. For example, the steady sound reduction unit 231N periodically acquires the sound detected by the sensor 10 and converts the acquired signal indicating changes in the time series of the sound into frequency characteristics by Fourier transforming it to generate sensor sound frequency characteristics. The steady sound reduction unit 231N refers to the steady sound information 220 to acquire the frequency characteristics of the steady sound and subtracts the acquired frequency characteristics of the steady sound from the sensor sound frequency characteristics.
[0046] Alternatively, the steady sound reduction unit 231N may reduce steady sounds by applying a filter to the sound detected by the sensor 10. The filter here has characteristics that reduce a frequency band corresponding to steady sounds. In this case, for example, the steady sound information 220 includes items such as filter characteristics. The filter characteristics indicate the characteristics of a filter that reduces frequency components corresponding to steady sounds. For example, if the filter is a digital filter, the filter characteristics are information indicating the filter configuration and coefficients. The steady sound reduction unit 231N refers to the steady sound information 220 to obtain the filter characteristics that reduce steady sounds, and generates a filter that reduces steady sounds based on the obtained filter characteristics. The steady sound reduction unit 231N applies the generated filter to the sound detected by the sensor 10, thereby reducing steady sounds from the sound detected by the sensor 10.
[0047] The steady sound reduction unit 231N outputs the sound detected by the sensor 10, from which the steady sound has been reduced, to the characteristic sound extraction unit 2313 and the position identification unit 2314.
[0048] The characteristic sound extraction unit 2313 acquires the sound from which the steady sound has been reduced from the steady sound reduction unit 231N. The characteristic sound extraction unit 2313 determines whether or not the sound acquired from the steady sound reduction unit 231N contains a characteristic sound, and outputs the determination result. The characteristic sound extraction unit 2313 uses, for example, a characteristic sound estimation model to determine whether or not the sound acquired from the steady sound reduction unit 231N contains a characteristic sound.
[0049] The characteristic sound extraction unit 2313 inputs the sound acquired from the steady sound reduction unit 231N into a characteristic sound estimation model, and thereby acquires a result obtained from the characteristic sound estimation model. The result obtained from the characteristic sound estimation model is an estimation result of whether or not a characteristic sound is included in the sound input to the model. The characteristic sound extraction unit 2313 outputs the result obtained from the characteristic sound estimation model to the position identification unit 2314.
[0050] The position identification unit 2314 acquires the sound from which the steady sound has been reduced from the steady sound reduction unit 231N. The position identification unit 2314 also acquires information indicating whether the sound contains a characteristic sound from the characteristic sound extraction unit 2313. If the sound contains a characteristic sound, the position identification unit 2314 estimates the position and movement of the user based on the sounds acquired from each of the steady sound reduction units 231N.
[0051] The process of estimating the position and movement of the user by the position identification unit 2314 will be described with reference to Fig. 9. Fig. 9 schematically shows four microphones as sensors 10-1 to 10-4 provided at the four corners of a space in a house H. Fig. 9 also schematically shows a state in which user U# moves from the position of user U to the position of user U in the center of the space.
[0052] In the example shown in this figure, when user U# moves from his / her position to user U's position, the characteristic sounds (user movement sounds) picked up by the microphones of sensors 10-1 and 10-2 gradually become quieter. Meanwhile, the characteristic sounds (user movement sounds) picked up by the microphones of sensors 10-3 and 10-4 gradually become louder. Position identification unit 2314 detects the user's position and movement based on such changes in the volume of the characteristic sounds.
[0053] Here, steady sounds will be explained using FIG. 10. Steady sounds are, for example, sounds (such as noise) that are constantly occurring at a certain location. Steady sounds are detected by averaging the frequency characteristics of sounds collected for a certain period of time (for example, one hour) by a microphone installed at that location. Steady sounds detected in this way include, for example, sounds with the same frequency that are continuously output for a certain period of time, such as the hum of a ventilation fan or natural sounds such as running water. In the example shown in this figure, steady sound is output from sound source SA1, and the frequency characteristic of this steady sound is shown to be characteristic N1.
[0054] For example, each of the microphones provided as sensors 10 in the house H may be configured to detect in advance stationary sounds in the locations where the microphones are provided. Information that associates the stationary sounds detected by each of the microphones with the locations where the sensors 10 are provided may be stored in the storage unit 22 as stationary sound information 220. In this case, the stationary sound reduction unit 231N may be configured to reduce the stationary sounds that correspond to the locations where the sensors 10 are provided from the sounds detected by the sensors 10.
[0055] Here, characteristic sounds will be explained using FIG. 11. Characteristic sounds are sounds that have characteristics different from steady sounds. Specifically, while steady sounds are sounds that have similar frequency components that are continuously output, characteristic sounds are sounds that are caused by the presence of a user or the like and are output suddenly over a short period of time. For example, a characteristic sound is a sound that corresponds to a human voice extracted based on the formant structure (the frequency characteristics of the voice of a person speaking). In the example of this figure, a characteristic sound such as a voice is output from user SA2, and the frequency characteristics of this characteristic sound are shown to be characteristic N2. Also, a characteristic sound such as a cry is output from animal SA3, and the frequency characteristics of this characteristic sound are shown to be characteristic N3.
[0056] Furthermore, the characteristic sound may include a sine wave that does not exist in nature, such as the sound of operating a touch panel. For example, if the sound of operating a touch panel is detected, and this indicates that a user is operating the touch panel, then the sound of operating the touch panel is the characteristic sound.
[0057] Information related to the characteristic sound may be stored in the storage unit 22. In this case, it may be possible to identify an individual corresponding to the characteristic sound. Specifically, the harmonic structure (frequency characteristics) of each of the voices of multiple users is detected in advance. The multiple users may include family members, housemates, and pet animals such as a pet cat. The detected harmonic structure of each of the voices is stored in the storage unit 22 as the characteristic sound of each user. Furthermore, if a touch panel or the like is operated only by a specific individual user, the operation sound of the touch panel is stored in the storage unit 22 as the characteristic sound of the specific individual user. In this case, the characteristic sound extraction unit 2313 determines whether the sound acquired from the steady sound reduction unit 231N includes a characteristic sound, based on the frequency characteristics of the characteristic sounds stored in the storage unit 22.
[0058] Here, the operation of this embodiment will be described with reference to Fig. 12 and Fig. 13. Fig. 12 and Fig. 13 are diagrams for explaining the processing performed by the control device 20 of this embodiment. As shown in Fig. 12, it is assumed that microphones (sensors 10-1 to 10-4) serving as sensors 10 are installed at four positions in a house H. It is also assumed that speakers 30 (speakers 30-1 to 30-4) are installed near each microphone. In this space, a steady sound is output from a sound source SA1. It is also assumed that a user SA2 is present near the sensor 10-1. It is also assumed that an animal SA3 is present near the sensor 10-3. It is also assumed that a user SA4 is present near the sensor 10-4.
[0059] Here, as shown by the arrow J in Fig. 13, consider a case where the cry of animal SA3 is detected as an alarm sound and a notification is sent to user SA2. The microphone of sensor 10-1 detects the steady sound and the voice of user SA2. The microphone of sensor 10-2 detects only the steady sound. The microphone of sensor 10-3 detects the steady sound and the cry of animal SA3. The microphone of sensor 10-4 detects the steady sound and the voice of user SA4.
[0060] Specifically, as shown in Fig. 13, the microphone of sensor 10-1 detects a sound exhibiting frequency characteristics T1, which is a combination of a stationary sound and the voice of user SA2. The microphone of sensor 10-2 detects a sound exhibiting frequency characteristics T2, which is a combination of only the stationary sound. The microphone of sensor 10-3 detects a sound exhibiting frequency characteristics T3, which is a combination of the stationary sound and the cry of animal SA3. The microphone of sensor 10-4 detects a sound exhibiting frequency characteristics T4, which is a combination of the stationary sound and the voice of user SA4.
[0061] The steady sound reduction unit 231N reduces steady sounds from the sounds detected by each of the sensors 10-1 to 10-4. As a result, only the voice of the user SA2 is extracted from the microphone of the sensor 10-1. Nothing is extracted from the microphone of the sensor 10-2. Only the cry of the animal SA3 is extracted from the microphone of the sensor 10-3. Only the voice of the user SA4 is extracted from the microphone of the sensor 10-4.
[0062] The characteristic sound extraction unit 2313 extracts a characteristic sound from the sounds extracted from each of the sensors 10-1 to 10-4. Specifically, the voice of the user SA2 is extracted as the characteristic sound from the microphone of the sensor 10-1. No characteristic sound is extracted from the microphone of the sensor 10-2. The cry of the animal SA3 is extracted as the characteristic sound from the microphone of the sensor 10-3. The voice of the user SA4 is extracted as the characteristic sound from the microphone of the sensor 10-4.
[0063] The position specifying unit 2314 specifies that a user SA2 is present at the position of the sensor 10-1, an animal SA3 is present at the position of the sensor 10-3, and a user SA4 is present at the position of the sensor 10-4.
[0064] The selection unit 232 selects the speaker 30-1 that is installed near the location of the notification destination user, i.e., user SA2. The output unit 233 causes the speaker 30-1 selected by the selection unit 232 to output the cry of the animal SA3 extracted from the microphone 10-3.
[0065] 13, when the level of a steady sound is high, if the steady sound is reduced by the steady sound reduction unit 231N, the level of the sound after the steady sound reduction will be low, and it may become difficult to extract a characteristic sound from the sound after the steady sound reduction. As a countermeasure to this, if the level of the steady sound is higher than a predetermined threshold, the steady sound reduction unit 231N may output the sound detected by the microphone as is to the characteristic sound extraction unit 2313 without reducing the steady sound.
[0066] When extracting a characteristic sound from a sound detected by a microphone using a characteristic sound estimation model, the characteristic sound extraction unit 2313 uses different characteristic sound estimation models depending on whether or not steady sounds have been reduced. For example, when steady sounds have been reduced, the characteristic sound extraction unit 2313 uses a model that estimates a characteristic sound from sound in which steady sounds have been reduced. On the other hand, when steady sounds have not been reduced, the characteristic sound extraction unit 2313 uses a model that estimates a characteristic sound from sound in which steady sounds have not been reduced.
[0067] 14 is a flowchart illustrating the flow of processing performed by the control device 20 of the embodiment. The control device 20 acquires information (sensor information) acquired by the sensor 10 (step S1). When the sensor information includes sound, the control device 20 determines whether or not the sound is an alarm sound (step S2). When the sensor 10 is a microphone, or when a microphone is included as part of the multiple sensors 10, the control device 20 determines whether or not the sound is an alarm sound by, for example, comparing the frequency characteristics of the sound collected by the microphone with the sound stored in the storage unit 22 as alarm sound information 221.
[0068] If the control device 20 determines that the sound is an alarm sound, it extracts all sensor information used for location identification to identify the user's location (step S3). For example, when the control device 20 identifies the user's location using an image, it extracts image information acquired by an image sensor. Alternatively, when the control device 20 identifies the user's location using temperature, it extracts information indicating the temperature acquired by an infrared sensor and a temperature sensor. When the control device 20 identifies the user's location using sound, it extracts information indicating the sound collected by a microphone. The following flow will explain the case where the control device 20 identifies the user's location using sound.
[0069] The control device 20 reduces steady sounds from the sound collected by the microphone (step S4). The control device 20 extracts characteristic sounds from the sound from which steady sounds have been reduced (step S5). The control device 20 identifies the user's position from the extracted characteristic sounds (step S6). The control device 20 selects a speaker 30 that will output an alarm sound based on the identified user's position (step S7). The control device 20 outputs the alarm sound from the selected speaker 30 (step S8).
[0070] On the other hand, if it is determined in step S2 that the sound is not an informative sound, the control device 20 ends the process.
[0071] Here, a method for the learning unit 235 to estimate the learned models (the position estimation model, the characteristic sound estimation model, and the informative sound estimation model) will be described with reference to Fig. 15. Fig. 15 is a flowchart illustrating the flow of a process for generating a machine learning model according to the embodiment.
[0072] (Location estimation model) First, a position estimation model will be described. The position estimation model is a model for estimating the position of a user. Here, a case will be described in which the position estimation model estimates whether a sound collected by a microphone is a sound originating from the user.
[0073] The learning unit 235 acquires sensor information for machine learning, in this case, information on sounds collected by a microphone (step S11). The learning unit 235 generates a learning dataset (step S12). The learning dataset here is information in which sounds collected by a microphone are labeled to indicate whether the sounds are caused by the user or not. The learning unit 235 determines whether all of the sensor information for machine learning has been acquired (step S13). If all of the sensor information for machine learning has not been acquired, the learning unit 235 returns to step S1.
[0074] When all the sensor information for machine learning has been acquired, the learning unit 235 generates a position estimation model (trained model) (step S14). The learning unit 235 generates the position estimation model (trained model) by having a machine learning model such as a CNN (Convolutional Neural Network) learn the training data set. When the learning unit 235 inputs a sound collected by a microphone from the training data set into the machine learning model such as a CNN, the learning unit 235 repeatedly causes the machine learning model to learn the training data set while adjusting the parameters of the machine learning model so that the value output from the machine learning model approaches the label assigned to the sound (whether the sound is caused by the user or not). This makes it possible to generate a model that can accurately estimate whether the sound collected by the microphone is caused by the user or not. The learning unit 235 stores the generated position estimation model (trained model) in the storage unit 22 as trained model information 223 (step S15).
[0075] (Feature sound estimation model) Next, a characteristic sound estimation model will be described. The characteristic sound estimation model is a model that estimates characteristic sounds. Here, a case will be described in which the characteristic sound estimation model estimates whether a sound collected by a microphone is a characteristic sound.
[0076] The learning unit 235 acquires sensor information for machine learning, in this case, information on sounds collected by a microphone (step S11). The learning unit 235 generates a learning dataset (step S12). The learning dataset here is information in which sounds collected by a microphone are labeled to indicate whether or not the sounds are characteristic sounds. The learning unit 235 determines whether or not all of the sensor information for machine learning has been acquired (step S13). If all of the sensor information for machine learning has not been acquired, the learning unit 235 returns to step S1.
[0077] When all the sensor information for machine learning has been acquired, the learning unit 235 generates a characteristic sound estimation model (trained model) (step S14). The learning unit 235 generates the characteristic sound estimation model (trained model) by having a machine learning model such as a CNN learn the training data set. When a sound collected by a microphone from the training data set is input to the machine learning model such as a CNN, the learning unit 235 repeatedly causes the machine learning model to learn the training data set while adjusting the parameters of the machine learning model so that the value output from the machine learning model approaches the label assigned to the sound (whether it is a characteristic sound or not). This makes it possible to generate a model that can accurately estimate whether the sound collected by the microphone is a characteristic sound or not. The learning unit 235 stores the generated characteristic sound estimation model (trained model) in the storage unit 22 as trained model information 223 (step S15).
[0078] (Auditory sound estimation model) Next, the informative sound estimation model will be described. The informative sound estimation model is a model that estimates informative sounds. Here, a case will be described in which the informative sound estimation model estimates whether a sound collected by a microphone is an informative sound.
[0079] The learning unit 235 acquires sensor information for machine learning, in this case, information on sounds collected by a microphone (step S11). The learning unit 235 generates a learning dataset (step S12). The learning dataset here is information in which sounds collected by a microphone are labeled to indicate whether the sounds are alarm sounds or not. The learning unit 235 determines whether all of the sensor information for machine learning has been acquired (step S13). If all of the sensor information for machine learning has not been acquired, the learning unit 235 returns to step S1.
[0080] When all the sensor information for machine learning has been acquired, the learning unit 235 generates an informative sound estimation model (trained model) (step S14). The learning unit 235 generates the informative sound estimation model (trained model) by having a machine learning model such as a CNN learn the training data set. When a sound collected by a microphone from the training data set is input to the machine learning model such as a CNN, the learning unit 235 repeatedly trains the machine learning model on the training data set while adjusting the parameters of the machine learning model so that the value output from the machine learning model approaches the label assigned to the sound (whether it is an informative sound or not). This makes it possible to generate a model that can accurately estimate whether the sound collected by the microphone is an informative sound or not. The learning unit 235 stores the generated informative sound estimation model (trained model) in the storage unit 22 as trained model information 223 (step S15).
[0081] As described above, the control device 20 according to the embodiment controls the sound to be output from each of the multiple speakers 30 provided in a space. The control device 20 includes an acquisition unit 230, an identification unit 231, a selection unit 232, and an output unit 233. The acquisition unit 230 acquires sensor information (detection results) detected by the multiple sensors 10 provided in the space. The identification unit 231 identifies the position of the user present in the space based on the detection results acquired by the acquisition unit 230. The selection unit 232 selects, from the multiple speakers 30, a speaker 30 installed near the user's position identified by the identification unit 231. The output unit 233 outputs sound from the speaker selected by the selection unit 232.
[0082] In the control device 20 according to the embodiment, the identification unit 231 identifies the user position based on the output obtained by inputting the detection result into the position estimation model. The position estimation model is a trained model that has learned the correspondence between the detection result and whether or not the detection result was detected due to the presence of a user, using a training dataset in which the detection result is labeled with a label indicating whether or not the detection result was detected due to the presence of a user.
[0083] In the control device 20 according to the embodiment, the sensor 10 may be a microphone. The identification unit 231 identifies the user position based on an output obtained by inputting the detection result into a position estimation model. The position estimation model is a model generated by performing learning using a training dataset in which sounds for machine learning are labeled with labels indicating whether the sounds are caused by the presence of a user, and is a model that estimates whether a sound detected by the microphone is caused by the presence of a user.
[0084] In the control device 20 according to the embodiment, the sensor 10 may be a microphone. When the microphone collects a sound having characteristics different from the steady sound (predetermined steady sound) stored in the storage unit 22 as the steady sound information 220, the identification unit 231 determines that a user is present near the microphone and identifies the installation position of the microphone as the position of the user.
[0085] In the control device 20 according to the embodiment, the sensor 10 may include a microphone. The identification unit 231 determines whether or not a sound collected by the microphone is an alarm sound to notify the user, based on the characteristics of the sound. The output unit 233 outputs the sound determined to be an alarm sound by the identification unit 231 (an example of a determination unit) from the speaker 30 selected by the selection unit 232.
[0086] In the control device 20 according to the embodiment, the sensor 10 may include a microphone. The identification unit 231 determines whether or not the sound collected by the microphone is an informative sound based on an output obtained by inputting the sound collected by the microphone into the informative sound estimation model. The informative sound estimation model is a model generated by performing learning using a training dataset in which sounds for machine learning are labeled with labels indicating whether or not the sounds are informative sounds that notify the user, and is a model that estimates whether or not a sound detected by the microphone is an informative sound.
[0087] (Modification 1 of the embodiment) Here, a first modification of the embodiment will be described. This modification differs from the above-described embodiment in that the informative sound estimation model undergoes additional learning. Fig. 16 is a flowchart illustrating the flow of processing for performing additional learning in the first modification of the embodiment.
[0088] The control device 20 acquires sensor information for estimation, in this case, information on sound collected by a microphone (step S21). The control device 20 estimates whether the sound collected by the microphone is an informative sound using an informative sound estimation model (trained model) (step S22). The control device 20 determines whether the estimation by the informative sound estimation model is correct (step S23). The control device 20 determines whether the estimation by the informative sound estimation model is correct, for example, based on information input by the user by operating a keyboard or the like.
[0089] If the estimation by the informative sound estimation model is incorrect, the control device 20 generates a training data set for additional learning (step S24). The training data set for additional learning is information in which a correct label indicating whether the sound is an informative sound is attached to the sound that has been erroneously estimated by the informative sound estimation model.
[0090] The control device 20 determines whether or not to perform additional learning (step S25). For example, the control device 20 determines to perform additional learning when the number of learning data sets for additional learning generated in step S24 reaches a predetermined number. Alternatively, the control device 20 may determine to perform additional learning when the probability that the informative sound estimation model makes an erroneous estimation is equal to or greater than a predetermined value.
[0091] When performing additional learning, the control device 20 performs additional learning by machine learning a learning data set for additional learning, and updates the informative sound estimation model (trained model) (step S26). The control device 20 stores the updated informative sound estimation model (trained model) (step S27).
[0092] As described above, in the control device 20 according to the first modification of the embodiment, the acquisition unit 230 acquires the determination result of whether or not the sound output by the output unit 233 is an informative sound, which is determined by the user. The learning unit 235 performs additional learning on the informative sound estimation model based on the determination result acquired by the acquisition unit 230. As a result, the control device 20 according to the first modification of the embodiment can perform additional learning on the informative sound estimation model, and can correct any estimation errors. Alternatively, when a new informative sound is added, the added informative sound can be machine-learned into the informative sound estimation model.
[0093] Although the above description has been given by way of example of a case where additional learning is performed on the informative sound estimation model, the present invention is not limited to this. The control device 20 can also perform additional learning on the position estimation model and the characteristic sound estimation model using a similar method.
[0094] (Modification 2 of the embodiment) Here, a second modification of the embodiment will be described. This modification differs from the above-described embodiment in that the user carries a transmitting terminal. The transmitting terminal is a terminal device that transmits a signal indicating the presence of the user, and is, for example, a smartphone or a beacon terminal. An acquisition unit 230 acquires a signal (location signal) transmitted from the transmitting terminal. An identification unit 231 identifies the location of the user based on the signal (location signal) acquired by the acquisition unit 230.
[0095] As described above, in the control device 20 according to the second modification of the embodiment, a plurality of users may exist in a space. At least one of the plurality of users possesses a transmitting terminal that transmits a location signal indicating the user's location. The acquisition unit 230 acquires the signal (location signal) transmitted from the transmitting terminal. The identification unit 231 identifies the user's location based on the signal (location signal) acquired by the acquisition unit 230. As a result, in the control device 20 according to the second modification of the embodiment, for a user who possesses a transmitting terminal, the user's location can be accurately identified based on the signal transmitted from the transmitting terminal.
[0096] (Modification 3 of the embodiment) Here, a third modification of the embodiment will be described. This modification assumes that the user listens to a sound that continues for a predetermined time, such as music. This modification differs from the above-described embodiment in that when the user moves within a space, sound is output in accordance with the user's movement.
[0097] The identification unit 231 (an example of a determination unit) determines, based on the characteristics of the sound collected by the microphone, whether or not the sound is a sound to be output in accordance with the movement of the user. When the sound collected by the microphone is a sound to be output in accordance with the movement of the user, the identification unit 231 identifies a user who is near the microphone that collected the sound. For example, the identification unit 231 identifies a user who is near the microphone based on a characteristic sound collected together with the sound to be output in accordance with the movement of the user. The identification unit 231 repeatedly identifies the position of the user while the sound to be output in accordance with the movement of the user is being acquired by the microphone. If the previously identified user position and the currently identified user position are different, the selection unit 232 selects the speaker 30 installed near the currently identified position. The output unit 233 stops the sound that was being output from the speaker 30 previously selected by the selection unit 232, and outputs the sound that is to be output in accordance with the movement of the user from the speaker 30 that is currently selected.
[0098] As described above, in the control device 20 according to the third modification of the embodiment, the sensor 10 may include a microphone. The identification unit 231 determines whether a sound collected by the microphone is a sound to be output in accordance with the user's movement. The identification unit 231 identifies a user who is near the microphone that collected the sound determined to be a sound to be output in accordance with the user's movement. The identification unit 231 associates the identified user's position with the sound to be output in accordance with the user's movement, and repeatedly identifies the user's position. If the previously identified user's position differs from the currently identified user's position, the selection unit 232 selects a speaker 30 installed near the currently identified user's position. The output unit 233 stops the sound that was previously output from the speaker 30 selected by the selection unit 232 and outputs the sound to be output in accordance with the user's movement from the currently selected speaker 30. As a result, in the control device 20 according to the third modification of the embodiment, when a user is listening to music or the like and moves away from the installation position of the speaker 30 that is outputting the sound, the control device 20 can output the sound from the speaker 30 at the new location, and can change the speaker 30 that is outputting the sound in accordance with the user's movement. This allows the user to continue listening to sound while moving around in space.
[0099] The speaker system 1 and the control device 20 in the above-described embodiment may be implemented in whole or in part by a computer. In this case, a program for implementing the functions may be recorded on a computer-readable recording medium, and the program may be loaded into a computer system and executed. The term "computer system" as used herein includes hardware such as an OS and peripheral devices. The term "computer-readable recording medium" refers to portable media such as flexible disks, optical magnetic disks, ROMs, and CD-ROMs, as well as storage devices such as hard disks built into a computer system. The term "computer-readable recording medium" may also include media that dynamically store programs for a short period of time, such as communication lines used when transmitting programs over a network such as the Internet or a telephone line, or media that store programs for a fixed period of time, such as volatile memory within a computer system serving as a server or client. The program may be a program for implementing some of the functions described above, or may be a program that can be implemented in combination with a program already stored in the computer system, or may be implemented using a programmable logic device such as an FPGA.
[0100] Although several embodiments of the present invention have been described, these embodiments are presented as examples and are not intended to limit the scope of the invention. These embodiments can be implemented in various other forms, and various omissions, substitutions, and modifications can be made without departing from the spirit of the invention. These embodiments and their modifications are included within the scope and spirit of the invention, as well as within the scope of the invention described in the claims and their equivalents. [Explanation of symbols]
[0101] 1...speaker system, 10...sensor, 20...control device, 230...acquisition unit, 231...identification unit (determination unit), 232...selection unit, 233...output unit, 234...device control unit, 235...learning unit, 30...speaker
Claims
1. A control method performed by a computer for controlling sounds to be output from a plurality of speakers provided in a plurality of spaces, the method comprising: repeatedly acquiring detected sounds collected by a plurality of microphones installed in the plurality of spaces; extracting a characteristic sound caused by a user's position or movement from the detected sound; Identifying a first space where the user is present based on the characteristic sound; selecting a speaker installed near the first space in which the identified user is present from the plurality of speakers; outputting sound from the selected speaker; determining whether the user has moved from the first space to a second space different from the first space based on a time-series change in the volume of the characteristic sound; When the user moves to the second space, sound is output from a speaker installed near the second space as a destination, and sound output from a speaker installed near the first space as a source of the user's movement is stopped. Control method.
2. the plurality of spaces are a plurality of rooms corresponding to spaces in which the user lives, Identifying the room where the user is present based on the detected sound; selecting a speaker installed in a room where the identified user is present from the plurality of speakers; The control method according to claim 1 .
3. A steady sound that is constantly occurring regardless of whether the user is present in the space is removed from the detected sound; identifying the room where the user is present based on the detected sound from which the steady sound has been reduced; The control method according to claim 2 .
4. If the level of the steady sound is greater than a threshold, the room in which the user is present is identified based on the detected sound in which the steady sound has not been reduced, and if the level of the steady sound is less than a threshold, the room in which the user is present is identified based on the detected sound in which the steady sound has been reduced. The control method according to claim 3 .
5. Identifying the user's location based on an output obtained by inputting the detected sound into a location estimation model; the location estimation model is a model generated by performing learning using a learning dataset in which the detected sound for machine learning is labeled with a label indicating whether the detected sound was detected due to the presence of the user, and is a model that estimates whether the detected sound is due to the presence of the user. The control method according to any one of claims 1 to 4.
6. There are a plurality of users in the space, At least one of the users possesses a transmitting terminal that transmits a location signal indicating the location of the user; Acquire the location signal transmitted from the transmitting terminal; determining a location of the user based on the location signal; The control method according to any one of claims 1 to 5.
7. Identifying the user's location based on an output obtained by inputting the sound detected by the microphone into a location estimation model; The location estimation model is a model generated by performing learning using a learning dataset in which sounds for machine learning are labeled with labels indicating whether the sounds are caused by the presence of the user, and is a model that estimates whether the sounds detected by the microphone are caused by the presence of the user. The control method according to any one of claims 1 to 6.
8. When a sound having characteristics different from a predetermined steady sound is picked up by the microphone, it is determined that the user is present in the vicinity of the microphone, and the installation position of the microphone is identified as the position of the user. The control method according to any one of claims 1 to 7.
9. determining whether the sound collected by the microphone is an alarm sound to be notified to the user based on characteristics of the sound collected by the microphone as the detected sound; outputting the sound determined to be the notification sound from the selected speaker; The control method according to any one of claims 1 to 8.
10. determining whether the sound collected by the microphone is an informative sound to be notified to the user based on an output obtained by inputting the sound collected by the microphone into an informative sound estimation model; The informative sound estimation model is a model generated by performing learning using a learning dataset in which sounds for machine learning are labeled with labels indicating whether the sounds are informative sounds to be notified to the user, and is a model that estimates whether a sound detected by the microphone is the informative sound. The control method according to any one of claims 1 to 9.
11. acquiring a determination result by the user as to whether the output sound is an alarm sound; performing additional learning on the informative sound estimation model based on the determination result; The control method according to claim 10.
12. setting a speaker that is not outputting sound among the plurality of speakers to a standby mode; A control method according to any one of claims 1 to 11.
13. determining whether the sound collected by the microphone is a sound to be output in accordance with the movement of the user based on characteristics of the sound collected by the microphone; Identifying the user who is in the vicinity of a microphone that collected the sound determined to be the sound to be output in accordance with the movement of the user, associating the identified position of the user with the sound to be output in accordance with the movement of the user, and repeatedly identifying the position of the user; If the previously identified position of the user is different from the currently identified position of the user, a speaker installed near the currently identified position of the user is selected from the plurality of speakers; outputting a sound from the currently selected speaker in accordance with the movement of the user; A control method according to any one of claims 1 to 12.
14. Stopping the sound that is output from the previously selected speaker in accordance with the user's movement; The control method according to claim 13.
15. A control device that controls sounds to be output from a plurality of speakers provided in a plurality of spaces, an acquisition unit that repeatedly acquires detected sounds collected by a plurality of microphones installed in the plurality of spaces; an extraction unit that extracts a characteristic sound caused by a position or movement of a user from the detected sound acquired by the acquisition unit; an identification unit that identifies a first space in which the user is present based on the characteristic sound; a selection unit that selects, from the plurality of speakers, a speaker that is installed near the first space where the user identified by the identification unit is present; an output unit that outputs sound from the speaker selected by the selection unit; Equipped with the identification unit determines whether the user has moved from the first space to a second space different from the first space, based on a time-series change in the volume of the characteristic sound; When the user moves to the second space, the output unit outputs sound from a speaker installed near the second space as a destination of the user, and stops output of sound that has been output from a speaker installed near the first space as a source of the user. Control device.
16. A computer is a control device that controls sounds to be output from a plurality of speakers provided in each of a plurality of spaces, an acquisition means for repeatedly acquiring the detected sounds collected by a plurality of microphones provided in the plurality of spaces; an extraction means for extracting a characteristic sound caused by a position or movement of the user from the detected sound acquired by the acquisition means; an identification means for identifying a first space where the user is present based on the characteristic sound; a selection means for selecting, from the plurality of speakers, a speaker installed in the vicinity of the first space in which the user identified by the identification means is present; an output means for outputting sound from the speaker selected by the selection means; A program for executing the identifying means determines whether the user has moved from the first space to a second space different from the first space based on a time-series change in the volume of the characteristic sound; When the user moves to the second space, the output means outputs sound from a speaker installed near the second space as a destination, and stops output of sound that has been output from a speaker installed near the first space as a source of the user. program.
Citation Information
Patent Citations
JP1974000406A
Magnetic recording and reproducing device
JP1985025037A
Engine with pressure wave supercharger
JP1987048930A
Voice guide device and voice guide system provided with the same
JP2005318149A
Information presenting system
JP2011204061A