Sound reproduction method, sound reproduction device, and sound reproduction program
By acquiring information about multiple people in the space, segmenting and synchronously reproducing sound signals, and controlling volume and location, the problem of gathering and active dialogue in public places is solved, creating a pleasant dialogue environment for multiple participants.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- PANASONIC INTELLECTUAL PROPERTY MANAGEMENT CO LTD
- Filing Date
- 2024-07-22
- Publication Date
- 2026-04-24
AI Technical Summary
Existing technologies have failed to effectively bring together multiple people in a given space and activate their conversations, especially in public places, and have been unable to effectively attract non-working groups or individuals to enter and participate in the conversation.
By acquiring information about multiple people in a given space, the sound signals are segmented and reproduced synchronously. The segmented sound signals are output through a speaker, and the volume and position are controlled according to personal information to ensure that the sound superimposes to form a pleasant dialogue environment.
It effectively brings together multiple individuals within a given space, fostering lively dialogue, enhancing the pleasant atmosphere of the space, and adapting to the participation status and needs of different individuals.
Smart Images

Figure CN121925863A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to techniques for reproducing sound. Background Technology
[0002] For example, Patent Document 1 discloses a system for preventing eavesdropping on private conversations in public spaces. This system detects working or leisure groups and individuals. Furthermore, in the area where working or leisure groups are located, the volume of music is reduced, and the low-frequency characteristics of the music are also reduced. This means that people in the group can clearly hear each other's voices without having to speak loudly above the music, and the low-frequency sounds do not mask their conversations. In addition, in the area where individuals are located, the volume of music is increased, and a low-frequency sound (the sound of flowing water) is added to prevent individuals from clearly hearing conversations between working or leisure groups.
[0003] However, the aforementioned prior technologies did not take into account the situation where multiple people gather in a given space, and further improvements are needed.
[0004] Prior art literature
[0005] Patent documents
[0006] Patent Document 1: Japanese Patent Publication No. 2011-528445 Summary of the Invention
[0007] This disclosure was made to solve the above-mentioned problems, and its purpose is to provide a technique that can bring together multiple characters in a given space and enable the dialogue of multiple characters in a given space to be active.
[0008] The sound reproduction method disclosed herein is a computer-based sound reproduction method, comprising: acquiring information relating to multiple persons located in a given space; synchronously reproducing multiple segmented sound signals obtained by segmenting a given sound; and outputting the multiple segmented sound signals to a speaker based on the information relating to the multiple persons, wherein the speaker emits multiple segmented sounds obtained by transforming the multiple segmented sound signals into the given space.
[0009] According to this disclosure, it is possible to gather multiple characters in a given space and to activate the dialogue between multiple characters located in the given space. Attached Figure Description
[0010] Figure 1 This is a diagram showing the overall structure of the sound reproduction system in Embodiment 1.
[0011] Figure 2 This is a block diagram showing the structure of the sound reproduction device in Embodiment 1.
[0012] Figure 3 This is a diagram illustrating an example of the information stored in the personal information database in Embodiment 1.
[0013] Figure 4 This is a first flowchart for describing the sound reproduction process performed by the sound reproduction apparatus in Embodiment 1 of this disclosure.
[0014] Figure 5 This is a second flowchart for describing the sound reproduction process performed by the sound reproduction apparatus in Embodiment 1 of this disclosure.
[0015] Figure 6 This is a diagram showing the overall structure of the sound reproduction system in Embodiment 2.
[0016] Figure 7 This is a block diagram showing the structure of the sound reproduction device in Embodiment 2.
[0017] Figure 8 This is a first flowchart for describing the sound reproduction process performed by the sound reproduction apparatus in Embodiment 2 of this disclosure.
[0018] Figure 9 This is a second flowchart for describing the sound reproduction process performed by the sound reproduction apparatus in Embodiment 2 of this disclosure.
[0019] Figure 10 This is a diagram showing the overall structure of the sound reproduction system in Embodiment 3.
[0020] Figure 11 This is a block diagram showing the structure of the sound reproduction device in Embodiment 3.
[0021] Figure 12 This is a first flowchart for describing the sound reproduction process performed by the sound reproduction apparatus in Embodiment 3 of this disclosure.
[0022] Figure 13 This is a second flowchart for describing the sound reproduction process performed by the sound reproduction apparatus in Embodiment 3 of this disclosure.
[0023] Figure 14 This is a diagram showing the overall structure of the sound reproduction system in Embodiment 4.
[0024] Figure 15 This is a block diagram showing the structure of the sound reproduction device in Embodiment 4.
[0025] Figure 16 This is a first flowchart for describing the sound reproduction process performed by the sound reproduction apparatus in Embodiment 4 of this disclosure.
[0026] Figure 17 This is a second flowchart for describing the sound reproduction process performed by the sound reproduction apparatus in Embodiment 4 of this disclosure. Detailed Implementation
[0027] (The knowledge that forms the basis of this disclosure)
[0028] In the past, the challenge was how to gather more people into a meeting, and how to make the dialogue among the participants more active.
[0029] In Patent Document 1 mentioned above, when there are group customers with working relationships or leisure time in the first area and individual customers in the second area adjacent to the first area, the volume of music in the second area is increased, and the sound of running water is added to the second area. As a result, conversations between people in the first area are masked in the second area, preventing eavesdropping on private conversations in a public place.
[0030] However, in Patent Document 1, people located outside the area may not necessarily want to go into the area, and the case of multiple people gathering in a given space is not taken into consideration.
[0031] To address the above issues, the following technology is disclosed.
[0032] (1) One aspect of the present disclosure relates to a sound reproduction method in a computer, comprising: acquiring information relating to multiple persons located in a given space; synchronously reproducing multiple segmented sound signals obtained by segmenting a given sound; and outputting the multiple segmented sound signals to a speaker based on the information relating to the multiple persons, wherein the speaker emits multiple segmented sounds obtained by transforming the multiple segmented sound signals into the given space.
[0033] According to this configuration, multiple segmented sound signals that are synchronously reproduced are output to a loudspeaker based on information related to multiple characters, and multiple segmented sounds are emitted from the loudspeaker within a given space.
[0034] Therefore, due to the gathering of multiple characters, multiple segmented voices superimpose on each other within a given space, resulting in the audible harmonization of a single voice. Thus, it is possible to gather multiple characters within a given space and to activate the dialogue between multiple characters located within that space.
[0035] (2) In the sound reproduction method described in (1) above, the information related to the multiple people may include the identification result information of each of the multiple people being identified, and the sound reproduction method may also include: controlling the volume of each of the multiple segmented sound signals according to whether the multiple people are identified respectively.
[0036] According to this configuration, the volume of each of the multiple segmented sound signals is controlled based on whether each of the multiple characters is located within a given space. Therefore, when multiple characters are located within a given space, a given sound obtained by superimposing the multiple segmented sounds can be played within that given space.
[0037] (3) In the sound reproduction method described in (2) above, the volume control may include: when a person is identified, the volume of the segmented sound signal corresponding to the person is determined to be a given volume, and when the person is not identified, the volume of the segmented sound signal corresponding to the person is determined to be zero.
[0038] Based on this configuration, the volume of the segmented sound signals is controlled so that when a character is not located in the given space, the segmented sound corresponding to that character is not played; and when a character is located in the given space, the segmented sound corresponding to that character is played. Therefore, as the number of characters in the given space increases, the given space becomes a space with a pleasant atmosphere where more segmented sounds overlap, thus enabling multiple characters to gather within the given space.
[0039] (4) In the sound reproduction method described in (1) above, the information related to the plurality of people may include audio information representing the average volume of the speech of each of the plurality of people during a given period, and the sound reproduction method may further include: controlling the volume of each of the plurality of segmented sound signals according to the average volume of the speech of each of the plurality of people during the given period.
[0040] According to this configuration, the volume of each of the multiple segmented sound signals is controlled based on the average volume of the speech of each of the multiple characters within a given period. Therefore, it is possible to change the volume of the segmented sounds based on whether a character is speaking, and to prompt a character who has not spoken to speak based on the presence or absence of segmented sounds in a given space.
[0041] (5) In the sound reproduction method described in (4) above, the volume control may include: when the average volume of the voice of the person is less than a threshold, the volume of the segmented sound signal corresponding to the person is determined as a given volume; when the average volume of the voice of the person is above the threshold, the volume of the segmented sound signal corresponding to the person is determined as zero.
[0042] According to this configuration, when a character speaks little, the corresponding segmented sound is played; when a character speaks a lot, the corresponding segmented sound is not played. Therefore, it can encourage characters who speak little to speak.
[0043] (6) In the sound reproduction method described in (1) above, the information related to the plurality of people may include posture information related to the posture of each of the plurality of people, and the sound reproduction method may further include: controlling the volume of each of the plurality of segmented sound signals according to the posture of each of the plurality of people.
[0044] Based on this configuration, the volume of each of the multiple segmented sound signals is controlled according to the individual postures of the multiple characters. Therefore, it is possible to determine whether to emit each of the multiple segmented sounds from a speaker based on the individual postures of the multiple characters.
[0045] (7) In the sound reproduction method described in (6) above, the posture information may indicate whether each of the plurality of characters is standing or not, and the volume control may include: when the character is standing, determining the volume of the segmented sound signal to a given volume, and when the character is not standing, determining the volume of the segmented sound signal to zero.
[0046] Based on this configuration, if multiple characters are standing, multiple split sounds are played, creating an environment where multiple characters can easily gather in a given space. Furthermore, if multiple characters are sitting, multiple split sounds are not played, thus creating an environment where multiple characters can easily converse in a given space.
[0047] (8) In the sound reproduction method described in (7) above, each of the multiple segmented sound signals may correspond to a multiple location of the multiple characters in the given space.
[0048] Based on this configuration, it is possible to establish a segmentation of sound corresponding to the location of a person in a given space, which is then emitted from the speaker.
[0049] (9) In the sound reproduction method described in (6) above, the posture information may indicate whether each of the multiple characters is sitting in multiple positions within the given space, and whether at least one of the multiple characters is in a drowsy posture. The volume control includes: when the character is sitting, determining the volume of the segmented sound signal corresponding to the position where the character is sitting as a given volume; when the character is not sitting, determining the volume of the segmented sound signal corresponding to the position where the character is not sitting as zero; when at least one of the multiple characters is in a drowsy posture, determining the volume of another segmented sound signal, different from the multiple segmented sound signals corresponding to each of the multiple positions, used to suppress drowsiness as a given volume; and when none of the multiple characters are in a drowsy posture, determining the volume of the other segmented sound signal as zero.
[0050] According to this configuration, when there are drowsy characters in a given space, additional segmented sounds are played to suppress drowsiness, thus stimulating the drowsy characters and making them focus on the conversation.
[0051] (10) In any of the sound reproduction methods described in (1) to (9) above, the given sound may be divided according to the high and low frequencies.
[0052] Based on this configuration, multiple segmented sound signals from each frequency band in multiple frequency bands can be emitted in a given space, and by superimposing multiple segmented sounds, a single sound that has achieved harmony can be emitted.
[0053] (11) In any of the sound reproduction methods described in (1) to (9) above, the given sound may be music and segmented according to the multiple performance parts constituting the music.
[0054] According to this configuration, multiple segmented sounds of each of the multiple performance parts constituting music can be emitted in a given space, and by superimposing multiple segmented sounds, a single piece of music that has achieved harmony can be emitted.
[0055] Furthermore, this disclosure not only enables a sound reproduction method that performs the aforementioned characteristic processing, but also enables a sound reproduction apparatus or the like that possessing a characteristic structure corresponding to the characteristic processing performed by the sound reproduction method. Additionally, it enables a computer program that causes a computer to execute the characteristic processing included in such a sound reproduction method. Therefore, in the following other embodiments, the same effects as the sound reproduction method described above can also be achieved.
[0056] (12) Another aspect of the sound reproduction apparatus disclosed herein includes: an acquisition unit for acquiring information relating to a plurality of persons located in a given space; a reproduction unit for synchronously reproducing a plurality of segmented sound signals obtained by segmenting a given sound; and an output unit for outputting the plurality of segmented sound signals to a loudspeaker based on the information relating to the plurality of persons, wherein the loudspeaker emits a plurality of segmented sounds obtained by transforming the plurality of segmented sound signals into the given space.
[0057] (13) Another aspect of the sound reproduction program disclosed herein enables a computer to: acquire information relating to multiple persons located in a given space; reproduce multiple segmented sound signals obtained by segmenting a given sound synchronously; and output the multiple segmented sound signals to a speaker based on the information relating to the multiple persons, wherein the speaker emits multiple segmented sounds obtained by transforming the multiple segmented sound signals into the given space.
[0058] (14) Another aspect of this disclosure relates to a non-transient computer-readable recording medium recording a sound reproduction program, the sound reproduction program enabling a computer to: acquire information relating to multiple persons located in a given space; reproduce synchronously multiple segmented sound signals obtained by segmenting a given sound; and output the multiple segmented sound signals to a speaker based on the information relating to the multiple persons, the speaker emitting multiple segmented sounds obtained by transforming the multiple segmented sound signals into the given space.
[0059] The embodiments of this disclosure will now be described with reference to the accompanying drawings. Furthermore, each of the embodiments described below illustrates a specific example of this disclosure. The numerical values, shapes, constituent elements, steps, and order of steps shown in the following embodiments are examples and are not intended to limit this disclosure. Additionally, any constituent elements in the following embodiments that are not described in the independent claims representing the highest-level concept are described as arbitrary constituent elements. Moreover, the contents of all embodiments can be combined.
[0060] (Implementation Method 1)
[0061] Figure 1 This is a diagram showing the overall structure of the sound reproduction system in Embodiment 1. Figure 2 This is a block diagram showing the structure of the sound reproduction device 2 in Embodiment 1.
[0062] The sound reproduction system includes a camera 1, a sound reproduction device 2, an amplifier 3, and a speaker 4.
[0063] Camera 1 films a given space. Camera 1 is a fixed camera such as a surveillance camera, positioned to film the entire given space. The given space is, for example, a room 100 such as a conference room where multiple people (person 101, person 102, person 103, and person 104) can gather. In this embodiment 1, camera 1 films the room 100. Furthermore, the sound reproduction system may have one camera 1 or multiple cameras 1.
[0064] For example, in room 100, a meeting is held by Person 101, Person 202, Person 303, and Person 404. The meeting is one in which each person presents their ideas. Music is played during the meeting, thus transforming room 100 into a space with a pleasant atmosphere, and enlivening the conversation among the participants.
[0065] Camera 1 connects to the sound reproduction device 2 itself, or to a hub (not shown) such as a communication device or server, via wired or wireless means, enabling it to input captured images to the sound reproduction device 2. Camera 1 can also connect to the sound reproduction device 2 communicatively via a network. Camera 1 outputs captured images to the sound reproduction device 2.
[0066] In addition, the images captured by camera 1 can be output in real time, or the images can be first recorded in an external storage device such as a memory or cloud server, and then output from these external storage devices.
[0067] The sound reproduction device 2 acquires information related to multiple people located in a given space, and synchronously reproduces multiple segmented sound signals obtained by segmenting a given sound. Based on the information related to multiple people, it outputs multiple segmented sound signals to the speaker 4, which emits multiple segmented sounds obtained by transforming the multiple segmented sound signals into the given space.
[0068] The sound reproduction device 2 detects multiple individuals (person 101, person 102, person 103, and person 104) located in a given space (room 100) based on images captured by the camera 1. The sound reproduction device 2 identifies each individual (person 101, person 102, person 103, and person 104) and controls the volume of each of the multiple segmented sound signals based on whether each individual (person 101, person 102, person 103, and person 104) is identified. That is, the sound reproduction device 2 adjusts the volume of each of the first to fourth segmented sound signals based on whether each individual (person 101, person 102, person 103, and person 104) is located in the given space (room 100).
[0069] The sound reproduction device 2 includes at least a computer system, which includes, for example, a control program, a processor or logic circuit for executing the control program, and a recording device such as internal memory or accessible external memory for storing the control program. Alternatively, the sound reproduction device 2 may be implemented, for example, through hardware installation based on the processing circuit, through the execution of a software program stored in memory or distributed from an external server by the processing circuit, or through a combination of these hardware and software installations.
[0070] The sound reproduction device 2 includes a person detection unit 201, a feature extraction unit 202, a personal information database 203, a personal identification unit 204, a memory 205, a sound source reproduction unit 206, a first volume control unit 207, a second volume control unit 208, a third volume control unit 209, a fourth volume control unit 210, and a mixing unit 211.
[0071] The person detection unit 201 detects people in room 100 based on images captured by camera 1.
[0072] The feature extraction unit 202 extracts the facial features of the person detected by the person detection unit 201.
[0073] The personal information database 203 stores the facial features of multiple individuals in a corresponding manner with identification information (user ID) used to identify these individuals. Here, the method for registering facial features of individuals into the personal information database 203 will be explained. First, a person is photographed using a camera. Next, the person is detected based on the photographed image, and the facial features of the detected person are extracted. Next, the person's identification information is input via an input device such as a touch panel. Finally, the facial features of the person and the person's identification information are stored in the personal information database 203 in a corresponding manner.
[0074] In addition, the personal information database 203 will establish corresponding storage of identification information for multiple individuals and musical instruments preferred by multiple individuals.
[0075] Figure 3 This is a diagram illustrating an example of the information stored in the personal information database 203 in this embodiment 1.
[0076] like Figure 3 As shown, each person's preferred musical instrument is associated with their user ID (identification information). For example, the instrument "drum" is associated with the user ID "A". Furthermore, the sound reproduction device 2 can also store the received user ID and preferred instrument in the personal information database 203 based on the person's input of their user ID and preferred instrument. Additionally, in the personal information database 203, at least the person's facial features and identification information need to be stored in a corresponding manner; preferred instruments may not be registered.
[0077] The personal recognition unit 204 outputs a reproduction start indication to the sound source reproduction unit 206 to begin reproducing a given sound. The personal recognition unit 204 outputs a reproduction start indication to the sound source reproduction unit 206 when a person is detected by the person detection unit 201. Furthermore, the personal recognition unit 204 can also output a reproduction end indication to the sound source reproduction unit 206 to end the reproduction of a given sound when no person is detected by the person detection unit 201. Additionally, the personal recognition unit 204 can also output a reproduction start indication to the sound source reproduction unit 206 when only one person out of multiple persons is identified.
[0078] Furthermore, the personal identification unit 204 can select a sound corresponding to the current time period from multiple sounds of different types, and output a reproduction start instruction to the sound source reproduction unit 206 to begin reproducing the selected sound. For example, the personal identification unit 204 can select a sound corresponding to the current time period from among the following: a first sound corresponding to the morning period (8:00 AM to 12:00 PM), a second sound corresponding to the afternoon period (12:00 AM to 6:00 PM), and a third sound corresponding to the night period (6:00 PM to 10:00 PM). However, the time period is not limited to the above.
[0079] The personal identification unit 204 acquires information related to multiple individuals located in a given space. The personal identification unit 204 identifies each of the multiple individuals located in the given space individually. The personal identification unit 204 controls the volume of each of the multiple segmented sound signals based on whether multiple individuals are identified individually.
[0080] The personal identification unit 204 compares the facial features extracted by the feature extraction unit 202 with the facial features of multiple individuals stored in the personal information database 203. If there is a feature among the facial features of multiple individuals stored in the personal information database 203 that matches the facial features extracted by the feature extraction unit 202, the personal identification unit 204 reads the identification information corresponding to that feature from the personal information database 203 and determines the individual based on the read identification information.
[0081] Multiple individuals are pre-assigned to multiple segmented sound signals. The personal identification unit 204 refers to the personal information database 203, extracts the preferred musical instrument corresponding to the user ID of the identified individual, and uses the segmented sound signal corresponding to the extracted preferred musical instrument from among the multiple segmented sound signals as the individual's segmented sound signal and establishes a correspondence.
[0082] In addition, if the personal information database 203 does not register the preferred musical instrument corresponding to the user ID of the identified person, the personal identification unit 204 may also use any one of the multiple segmented sound signals, except for the segmented sound signal that has already been established as the segmented sound signal of any person, as the segmented sound signal of that person and establish a corresponding one.
[0083] The personal identification unit 204 determines the volume of the first segmented sound signal corresponding to the first person 101 based on whether the first person 101 is in room 100, and instructs the determined volume to the first volume control unit 207. The personal identification unit 204 determines the volume of the second segmented sound signal corresponding to the second person 102 based on whether the second person 102 is in room 100, and instructs the determined volume to the second volume control unit 208. The personal identification unit 204 determines the volume of the third segmented sound signal corresponding to the third person 103 based on whether the third person 103 is in room 100, and instructs the determined volume to the third volume control unit 209. The personal identification unit 204 determines the volume of the fourth segmented sound signal corresponding to the fourth person 104 based on whether the fourth person 104 is in room 100, and instructs the determined volume to the fourth volume control unit 210.
[0084] When the personal recognition unit 204 recognizes a person, it sets the volume of the segmented sound signal corresponding to that person as a given volume. The given volume is, for example, the maximum volume. Furthermore, when the personal recognition unit 204 does not recognize a person, it sets the volume of the segmented sound signal corresponding to that person to zero. Additionally, the given volume is not limited to the maximum volume; it can be 50% or 80% of the maximum volume, as long as it is greater than zero.
[0085] If the personal identification unit 204 identifies the first person 101 in room 100, it sets the volume of the first segmented sound signal corresponding to the first person 101 to a given volume. Conversely, if the personal identification unit 204 does not identify the first person 101 in room 100, it sets the volume of the first segmented sound signal corresponding to the first person 101 to zero.
[0086] If the personal identification unit 204 identifies a second person 102 within room 100, it sets the volume of the second segmented sound signal corresponding to the second person 102 to a given volume. Conversely, if the personal identification unit 204 does not identify a second person 102 within room 100, it sets the volume of the second segmented sound signal corresponding to the second person 102 to zero.
[0087] If the personal identification unit 204 identifies a third person 103 within room 100, it sets the volume of the third segmented sound signal corresponding to the third person 103 to a given volume. Conversely, if the personal identification unit 204 does not identify the third person 103 within room 100, it sets the volume of the third segmented sound signal corresponding to the third person 103 to zero.
[0088] If the personal identification unit 204 identifies the fourth person 104 within room 100, it sets the volume of the fourth segmented sound signal corresponding to the fourth person 104 to a given volume. Conversely, if the personal identification unit 204 does not identify the fourth person 104 within room 100, it sets the volume of the fourth segmented sound signal corresponding to the fourth person 104 to zero.
[0089] In addition, in this embodiment 1, the multiple characters are 4 people, but this disclosure is not particularly limited to this, and the multiple characters can be 2 or more people.
[0090] The memory 205 pre-stores multiple segmented sound signals obtained by segmenting a given sound. The memory 205 pre-stores a first segmented sound signal, a second segmented sound signal, a third segmented sound signal, and a fourth segmented sound signal. Alternatively, the memory 205 may pre-store multiple sounds. The number of segmented sound signals may be the same as the number of characters, or it may be different from the number of characters.
[0091] The given sound is, for example, a series of musical instruments played together. Furthermore, the given sound is not limited to music, as long as it is a series of audio content; for example, it could be a series of natural sounds such as wind, flowing water, insect chirping, or birdsong.
[0092] Furthermore, a given sound can also be segmented according to the performance portion of the reproduced music. For example, a given sound can also be segmented into a first segmented sound signal representing the melody portion playing the main melody, a second segmented sound signal representing the harmonic portion playing chords or auxiliary melodies relative to the melody portion, a third segmented sound signal representing the rhythm portion responsible for the rhythm of the music, and a fourth segmented sound signal representing the bass portion responsible for the low notes.
[0093] A given sound can also be segmented according to the type of instrument played. For example, a given sound can be segmented into a first segment representing the sound of a drum, a second segment representing the sound of a bass, a third segment representing the sound of a keyboard, and a fourth segment representing the sound of a guitar.
[0094] Furthermore, a given sound can also be segmented according to its frequency. For example, a given sound can be segmented into a first segment with a frequency below 50Hz, a second segment with a frequency between 50Hz and 500Hz, a third segment with a frequency between 500Hz and 1kHz, and a fourth segment with a frequency above 1kHz.
[0095] Furthermore, a given sound can also be segmented based on the multiple elements that constitute the given sound. For example, if the given sound is a series of natural sounds, it can also be segmented into a first segment representing the sound of wind, a second segment representing the sound of flowing water, a third segment representing the sound of insects chirping, and a fourth segment representing the sound of birds calling.
[0096] The sound source reproduction unit 206 synchronously reproduces multiple segmented sound signals obtained by segmenting a given sound. The sound source reproduction unit 206 is an example of a reproduction unit. The sound source reproduction unit 206 reads multiple segmented sound signals from the memory 205. If a reproduction start instruction is input from the personal identification unit 204, the sound source reproduction unit 206 begins reproducing the multiple segmented sound signals obtained by segmenting the given sound. Furthermore, if a reproduction end instruction is input from the personal identification unit 204, the sound source reproduction unit 206 ends the reproduction of the multiple segmented sound signals obtained by segmenting the given sound.
[0097] Regardless of whether any of the characters 101 to 104 are present in room 100, the sound source reproduction unit 206 synchronously reproduces all of the first to fourth segmented sound signals. That is, as long as at least one of the characters 101 to 104 is present in room 100, the sound source reproduction unit 206 synchronously reproduces all of the first to fourth segmented sound signals.
[0098] The sound source reproduction unit 206 outputs the first split sound signal to the first volume control unit 207, the second split sound signal to the second volume control unit 208, the third split sound signal to the third volume control unit 209, and the fourth split sound signal to the fourth volume control unit 210.
[0099] Volume control units 207 to 210 output multiple segmented sound signals reproduced by sound source reproduction unit 206 to speaker 4. Speaker 4 emits multiple segmented sounds derived from the multiple segmented sound signals into a given space (room 100). Volume control units 207 to 210 are examples of output units. Furthermore, volume control units 207 to 210 control the volume of each of the multiple segmented sound signals according to whether each of the multiple characters is located in the given space.
[0100] The first volume control unit 207 controls the volume of the first segmented sound signal based on the recognition result of the first person 101. That is, the first volume control unit 207 controls the volume of the first segmented sound signal based on whether the first person 101 is located in the room 100. The first volume control unit 207 sets the volume of the first segmented sound signal input from the sound source reproduction unit 206 to the volume indicated by the personal recognition unit 204.
[0101] The second volume control unit 208 controls the volume of the second segmented sound signal based on the recognition result of the second person 102. That is, the second volume control unit 208 controls the volume of the second segmented sound signal based on whether the second person 102 is located in the room 100. The second volume control unit 208 sets the volume of the second segmented sound signal input from the sound source reproduction unit 206 to the volume indicated by the personal recognition unit 204.
[0102] The third volume control unit 209 controls the volume of the third segmented sound signal based on the recognition result of the third person 103. That is, the third volume control unit 209 controls the volume of the third segmented sound signal based on whether the third person 103 is located in the room 100. The third volume control unit 209 sets the volume of the third segmented sound signal input from the sound source reproduction unit 206 to the volume indicated by the personal recognition unit 204.
[0103] The fourth volume control unit 210 controls the volume of the fourth segmented sound signal based on the recognition result of the fourth person 104. That is, the fourth volume control unit 210 controls the volume of the fourth segmented sound signal based on whether the fourth person 104 is located in the room 100. The fourth volume control unit 210 sets the volume of the fourth segmented sound signal input from the sound source reproduction unit 206 to the volume indicated by the personal recognition unit 204.
[0104] The mixing unit 211 combines the first segmented sound signal input from the first volume control unit 207, the second segmented sound signal input from the second volume control unit 208, the third segmented sound signal input from the third volume control unit 209, and the fourth segmented sound signal input from the fourth volume control unit 210. The mixing unit 211 outputs the combined sound signal obtained by combining the first to fourth segmented sound signals to the amplifier 3. The mixing unit 211 combines multiple segmented sound signals according to the volume set for each segmented sound signal.
[0105] The sound reproduction device 2 is connected via wired or wireless means to multiple amplifiers, communication devices, or servers (not shown) to input a synthesized sound signal obtained by combining multiple segmented sound signals into amplifier 3. The sound reproduction device 2 can also be communicatively connected to amplifier 3 via a network. The sound reproduction device 2 outputs the synthesized sound signal obtained by combining multiple segmented sound signals to amplifier 3.
[0106] Amplifier 3 amplifies the synthesized sound signal input from the mixing section 211 of the sound reproduction device 2.
[0107] In addition, amplifier 3 can be installed inside room 100 or outside room 100.
[0108] Speaker 4 is positioned in room 100 and converts the synthesized sound signal amplified by amplifier 3 into synthesized sound, which is then emitted into room 100. When the volume of the first segmented sound signal is set to a given volume by the first volume control unit 207, speaker 4 emits the first segmented sound at the given volume. When the volume of the first segmented sound signal is set to zero by the first volume control unit 207, speaker 4 does not emit the first segmented sound. Furthermore, when the volume of the second segmented sound signal is set to a given volume by the second volume control unit 208, speaker 4 emits the second segmented sound at the given volume. When the volume of the second segmented sound signal is set to zero by the second volume control unit 208, speaker 4 does not emit the second segmented sound. Furthermore, when the volume of the third segmented sound signal is set to a given volume by the third volume control unit 209, speaker 4 emits the third segmented sound at the given volume. When the volume of the third segmented sound signal is set to zero by the third volume control unit 209, speaker 4 does not emit the third segmented sound. Furthermore, when the volume of the fourth segment sound signal is set to a given volume by the fourth volume control unit 210, the speaker 4 emits the fourth segment sound at the given volume. When the volume of the fourth segment sound signal is set to zero by the fourth volume control unit 210, the speaker 4 does not emit the fourth segment sound.
[0109] Next, the sound reproduction process performed by the sound reproduction apparatus 2 in Embodiment 1 of this disclosure will be described.
[0110] Figure 4 This is a first flowchart for describing the sound reproduction process performed by the sound reproduction apparatus 2 in Embodiment 1 of this disclosure. Figure 5 This is a second flowchart for describing the sound reproduction process performed by the sound reproduction apparatus 2 in Embodiment 1 of this disclosure.
[0111] First, in step S1, the personal identification unit 204 determines whether a person has been detected by the person detection unit 201. Here, if it is determined that no person has been detected (no in step S1), the determination process in step S1 is repeated.
[0112] On the other hand, if it is determined that a person has been detected (yes in step S1), in step S2, the personal identification unit 204 determines whether the first to fourth segmented sound signals obtained by segmenting the given sound are being reproduced.
[0113] Here, if it is determined that the first to fourth segmented sound signals are being reproduced (yes in step S2), the process is transferred to step S4.
[0114] On the other hand, if it is determined that the first to fourth segmented sound signals are not being reproduced (not in step S2), in step S3, the sound source reproduction unit 206 synchronously reproduces the first to fourth segmented sound signals obtained by segmenting the given sound. At this time, the personal identification unit 204 outputs a reproduction start instruction to the sound source reproduction unit 206 to start the reproduction of the given sound. If a reproduction start instruction is input from the personal identification unit 204, the sound source reproduction unit 206 reads the first to fourth segmented sound signals from the memory 205 and synchronously reproduces the read first to fourth segmented sound signals. The sound source reproduction unit 206 outputs the reproduced first to fourth segmented sound signals to the first volume control unit 207 to the fourth volume control unit 210, respectively.
[0115] Furthermore, in the case of a meeting being held in room 100, multiple individuals are pre-selected as participants. Therefore, if an individual is detected, the sound source reproduction unit 206 can synchronously reproduce multiple segmented sound signals. For example, the sound reproduction device 2 can also accept input for identifying multiple individuals using room 100 when accepting a reservation for its use.
[0116] Next, in step S4, the personal identification unit 204 identifies multiple individuals located in room 100 based on images captured by camera 1. The person detection unit 201 detects individuals based on images captured by camera 1. The feature extraction unit 202 extracts facial features of the individuals detected by the person detection unit 201. The personal identification unit 204 compares the facial features extracted by the feature extraction unit 202 with the facial features of multiple individuals stored in the personal information database 203. If a feature in the facial features of multiple individuals stored in the personal information database 203 matches the feature extracted by the feature extraction unit 202, the personal identification unit 204 reads the corresponding identification information from the personal information database 203 and determines the individual based on the read identification information.
[0117] Next, in step S5, the personal identification unit 204 determines whether the first person 101 is identified in room 100. That is, the personal identification unit 204 determines whether the first person 101 is in room 100.
[0118] Here, if it is determined that the first person 101 has been identified (yes in step S5), in step S6, the personal identification unit 204 determines the volume of the first segmented audio signal corresponding to the first person 101 as a given volume. The given volume is, for example, the maximum volume. The personal identification unit 204 instructs the first volume control unit 207 to set the volume of the first segmented audio signal to the given volume. Based on the instruction from the personal identification unit 204, the first volume control unit 207 sets the volume of the first segmented audio signal to the given volume. The first volume control unit 207 outputs the first segmented audio signal with the set volume to the mixing unit 211.
[0119] On the other hand, if it is determined that the first person 101 has not been identified (no in step S5), in step S7, the personal identification unit 204 sets the volume of the first segmented audio signal corresponding to the first person 101 to zero. The personal identification unit 204 instructs the first volume control unit 207 to set the volume of the first segmented audio signal to zero. Based on the instruction from the personal identification unit 204, the first volume control unit 207 sets the volume of the first segmented audio signal to zero. The first volume control unit 207 outputs the first segmented audio signal with the set volume to the mixing unit 211.
[0120] Next, in step S8, the personal identification unit 204 determines whether the second person 102 is identified in room 100. That is, the personal identification unit 204 determines whether the second person 102 is in room 100.
[0121] Here, if it is determined that the second person 102 has been identified (yes in step S8), in step S9, the personal identification unit 204 determines the volume of the second segmented audio signal corresponding to the second person 102 as a given volume. The given volume is, for example, the maximum volume. The personal identification unit 204 instructs the second volume control unit 208 to set the volume of the second segmented audio signal to the given volume. Based on the instruction from the personal identification unit 204, the second volume control unit 208 sets the volume of the second segmented audio signal to the given volume. The second volume control unit 208 outputs the second segmented audio signal with the set volume to the mixing unit 211.
[0122] On the other hand, if it is determined that the second person 102 has not been identified (no in step S8), in step S10, the personal identification unit 204 sets the volume of the second segmented audio signal corresponding to the second person 102 to zero. The personal identification unit 204 instructs the second volume control unit 208 to set the volume of the second segmented audio signal to zero. Based on the instruction from the personal identification unit 204, the second volume control unit 208 sets the volume of the second segmented audio signal to zero. The second volume control unit 208 outputs the second segmented audio signal with the set volume to the mixing unit 211.
[0123] Next, in step S11, the personal identification unit 204 determines whether a third person 103 is identified in room 100. That is, the personal identification unit 204 determines whether a third person 103 is in room 100.
[0124] Here, if it is determined that a third person 103 has been identified (yes in step S11), in step S12, the personal identification unit 204 determines the volume of the third segmented audio signal corresponding to the third person 103 as a given volume. The given volume is, for example, the maximum volume. The personal identification unit 204 instructs the third volume control unit 209 to set the volume of the third segmented audio signal to the given volume. Based on the instruction from the personal identification unit 204, the third volume control unit 209 sets the volume of the third segmented audio signal to the given volume. The third volume control unit 209 outputs the third segmented audio signal with the set volume to the mixing unit 211.
[0125] On the other hand, if it is determined that the third person 103 has not been identified (no in step S11), in step S13, the personal identification unit 204 sets the volume of the third segmented audio signal corresponding to the third person 103 to zero. The personal identification unit 204 instructs the third volume control unit 209 to set the volume of the third segmented audio signal to zero. Based on the instruction from the personal identification unit 204, the third volume control unit 209 sets the volume of the third segmented audio signal to zero. The third volume control unit 209 outputs the third segmented audio signal with the set volume to the mixing unit 211.
[0126] Next, in step S14, the personal identification unit 204 determines whether a fourth person 104 is identified in room 100. That is, the personal identification unit 204 determines whether a fourth person 104 is in room 100.
[0127] Here, if it is determined that the fourth person 104 has been identified (yes in step S14), in step S15, the personal identification unit 204 determines the volume of the fourth segmented audio signal corresponding to the fourth person 104 as a given volume. The given volume is, for example, the maximum volume. The personal identification unit 204 instructs the fourth volume control unit 210 to set the volume of the fourth segmented audio signal to the given volume. Based on the instruction from the personal identification unit 204, the fourth volume control unit 210 sets the volume of the fourth segmented audio signal to the given volume. The fourth volume control unit 210 outputs the fourth segmented audio signal with the set volume to the mixing unit 211.
[0128] On the other hand, if it is determined that the fourth person 104 has not been identified (no in step S14), in step S16, the personal identification unit 204 sets the volume of the fourth segmented audio signal corresponding to the fourth person 104 to zero. The personal identification unit 204 instructs the fourth volume control unit 210 to set the volume of the fourth segmented audio signal to zero. Based on the instruction from the personal identification unit 204, the fourth volume control unit 210 sets the volume of the fourth segmented audio signal to zero. The fourth volume control unit 210 outputs the fourth segmented audio signal with the set volume to the mixing unit 211.
[0129] Next, in step S17, the mixing unit 211 outputs a synthesized sound signal obtained by combining the first to fourth segmented sound signals, whose volumes have been set, to the amplifier 3. The amplifier 3 amplifies the synthesized sound signal and outputs the amplified synthesized sound signal to the speaker 4. The speaker 4 converts the synthesized sound signal into synthesized sound and emits the synthesized sound obtained by combining the first to fourth segmented sound signals, whose respective volumes have been set by the first volume control unit 207 to the fourth volume control unit 210, into the room 100.
[0130] Therefore, if person 101 is present in room 100, a first segment of sound at a given volume is played from speaker 4; if person 101 is not present in room 100, the first segment of sound is not played from speaker 4. Similarly, if person 102 is present in room 100, a second segment of sound at a given volume is played from speaker 4; if person 102 is not present in room 100, the second segment of sound is not played from speaker 4. Furthermore, if person 103 is present in room 100, a third segment of sound at a given volume is played from speaker 4; if person 103 is not present in room 100, the third segment of sound is not played from speaker 4. Finally, if person 104 is present in room 100, a fourth segment of sound at a given volume is played from speaker 4; if person 104 is not present in room 100, the fourth segment of sound is not played from speaker 4.
[0131] In this way, because multiple people are gathered in room 100, the multiple segmented sounds are superimposed on each other, resulting in a single, harmonious sound. Therefore, it is possible to gather multiple people in room 100 and to activate the dialogue between them. For example, in the case of a meeting being held in room 100, it is possible to gather a large number of people who will be attending the meeting.
[0132] Furthermore, in this embodiment 1, a person located in a given space is identified based on facial features extracted from an image captured by camera 1. However, this disclosure is not particularly limited to this, and other identification methods can also be used. For example, a person can be identified by fingerprints, iris scans, voice features (voiceprints), hand movements or gestures, body shape or outline, beacons using near-field wireless communication, or ID cards.
[0133] Furthermore, in this embodiment 1, the personal identification unit 204 identifies persons located in a given space at given intervals (e.g., every minute), but is not limited to this. For example, when an ID card is used for person identification, a person can also be identified by reading the information on the ID card using a card reader or the like at the start of a meeting or upon entering the given space. Moreover, it can also be configured such that the personal identification unit 204 identifies the person and outputs a segmented audio signal corresponding to that person during the period from when the person is identified until the end of the meeting or when the person leaves the given space, when the information on the ID card is read again using a card reader or the like.
[0134] Furthermore, in a variation of Embodiment 1, the personal identification unit 204 may select one sound from multiple sounds of different kinds (types) and instruct the sound source reproduction unit 206 to reproduce multiple segmented sound signals obtained by segmenting the selected sound.
[0135] For example, the types of sounds include jazz, pop, orchestral, and classical music. The sound reproduction device 2 can also be configured to accept a single sound from multiple sounds of different types. Then, the personal identification unit 204 can output a reproduction start instruction to the sound source reproduction unit 206 to begin reproducing the sound selected by the user. Furthermore, the personal identification unit 204 can also select a sound corresponding to the current time period from multiple sounds of different types and output a reproduction start instruction to the sound source reproduction unit 206 to begin reproducing the selected sound.
[0136] Furthermore, each of the multiple sounds can be segmented according to the performance part of the reproduced music. For example, each of the multiple sounds can be segmented into a first segmented sound signal representing the melody part of the main melody, a second segmented sound signal representing the harmonic part of the melody part of the chord or auxiliary melody, a third segmented sound signal representing the rhythm part of the music, and a fourth segmented sound signal representing the bass part of the music.
[0137] Furthermore, each sound can be segmented according to the type of instrument being played. For example, if the "jazz" sound is selected, the first segment representing the drum sound, the second segment representing the bass sound, the third segment representing the keyboard sound, and the fourth segment representing the guitar sound can be emitted to room 100.
[0138] Furthermore, a single segmented audio signal can also include the sounds of multiple instruments. For example, the first segmented audio signal can also include the sounds of a bass and a guitar.
[0139] Furthermore, in this embodiment 1, the sound source reproduction unit 206 can also always reproduce a background sound signal that is different from the multiple segmented sound signals. The sound source reproduction unit 206 can also output the background sound signal to a fifth volume control unit that is different from the first volume control unit 207 to the fourth volume control unit 210. The fifth volume control unit can also set the volume of the background sound signal to a given volume (e.g., the maximum volume). The mixing unit 211 can also output a synthesized sound signal obtained by synthesizing the first segmented sound signal, the second segmented sound signal, the third segmented sound signal, the fourth segmented sound signal, and the background sound signal to the amplifier 3. The speaker 4 can also always emit background sound and emit the first segmented sound to the fourth segmented sound in superimposed with the background sound.
[0140] In addition, background sounds can be natural sounds such as wind, flowing water, insect chirping, or birdsong. Natural sounds have a calming effect, and by consistently reproducing natural sounds within a given space, people can be drawn to that space.
[0141] Furthermore, the background sound can be, for example, the sound of an instrument that covers the frequency band of human speech. The frequency band of the electronic piano sound is superimposed on the frequency band of human speech, and the sound source reproduction unit 206 can also independently reproduce the electronic piano sound signal as a background sound signal, along with multiple segmented sound signals.
[0142] Furthermore, in this embodiment 1, when a meeting is held in room 100, it is not necessary to pre-determine multiple individuals as participants. For example, if a participant in the given space is a person registered in the personal information database 203, then the segmented audio signal corresponding to that person is reproduced. On the other hand, if a participant is not registered in the personal information database 203, then any one of the multiple segmented audio signals, except for the segmented audio signal already corresponding to any person in the given space, is reproduced as that person's segmented audio signal. Furthermore, the person's identification information and the segmented audio signal can be registered in the personal information database 203 in a corresponding manner. In this case, the registration information of persons who were not initially registered in the personal information database 203 can be deleted from the personal information database 203 after a given period. Additionally, the given period can be, for example, the duration of the meeting, or a period of one day or several days.
[0143] (Implementation Method 2)
[0144] In the sound reproduction device 2 shown in Embodiment 1, multiple people are identified separately, and the volume of each of the multiple segmented sound signals is controlled according to whether multiple people are identified separately. However, in the sound reproduction device shown in Embodiment 2, audio information representing the average volume of the speech of each of the multiple people within a given period is acquired, and the volume of each of the multiple segmented sound signals is controlled according to the audio information of each of the multiple people.
[0145] Figure 6 This is a diagram showing the overall structure of the sound reproduction system in Embodiment 2. Figure 7 This is a block diagram showing the structure of the sound reproduction device 2A in Embodiment 2.
[0146] The sound reproduction system in this embodiment 2 includes a microphone 5, a sound reproduction device 2A, an amplifier 3, and a speaker 4. Furthermore, in this embodiment 2, structures identical to those in embodiment 1 are labeled with the same reference numerals, and descriptions are omitted.
[0147] Microphone 5 collects audio within a given space. Microphone 5 is positioned to collect audio within the given space. The given space is, for example, a room 100 such as a conference room where multiple people (person 101, person 102, person 103, and person 104) can gather. In this embodiment 2, microphone 5 collects the voices of the multiple people (person 101, person 102, person 103, and person 104) within room 100.
[0148] Microphone 5 is connected via wired or wireless means to the sound reproduction device 2A itself, or to a hub (not shown) such as a communication device or server, enabling it to input collected audio to the sound reproduction device 2A. Microphone 5 can also be connected to the sound reproduction device 2A communicatively via a network. Microphone 5 outputs the collected audio to the sound reproduction device 2A.
[0149] In addition, the audio collected by microphone 5 can be output in real time, or the audio can be first recorded in an external storage device such as a memory or cloud server, and then the image can be output from these external storage devices.
[0150] The sound reproduction device 2A acquires information related to multiple individuals located in a given space, and synchronously reproduces multiple segmented sound signals obtained by segmenting a given sound. Based on the information related to the multiple individuals, it outputs the multiple segmented sound signals to a speaker, which emits the multiple segmented sounds obtained from the multiple segmented sound signals into the given space. The sound reproduction device 2A detects the speech of multiple individuals (person 101, person 102, person 103, and person 104) located in the given space (room 100) based on the audio collected by the microphone 5. The sound reproduction device 2A controls the volume of each of the multiple segmented sound signals according to the average volume of the speech of each individual (person 101, person 102, person 103, and person 104) within a given period. The sound reproduction device 2A adjusts the volume of the first to fourth segmented sound signals based on the average volume of the speech of each of the multiple characters (character 101, character 102, character 103, and character 104) within a given period.
[0151] The sound reproduction device 2A includes at least a computer system, which includes, for example, a control program, a processor or logic circuit for executing the control program, and a recording device such as internal memory or accessible external memory for storing the control program. Alternatively, the sound reproduction device 2A may be implemented, for example, through hardware installation based on the processing circuit, through the execution of a software program stored in memory or distributed from an external server by the processing circuit, or through a combination of these hardware and software installations.
[0152] The sound reproduction device 2A includes an audio detection unit 212, a feature extraction unit 202A, a speaker information database 213, a speaker estimation unit 214, a memory 205, a sound source reproduction unit 206, a first volume control unit 207, a second volume control unit 208, a third volume control unit 209, a fourth volume control unit 210, and a mixing unit 211.
[0153] The audio detection unit 212 detects the voice of a person located in room 100 based on the audio collected by the microphone 5.
[0154] The feature extraction unit 202A extracts the features of the speech of the person detected by the audio detection unit 212.
[0155] Speaker information database 213 stores the characteristic values of the speech of multiple individuals and the identification information (user ID) used to identify these individuals in a corresponding manner. Here, the method for registering the characteristic values of a person's speech in speaker information database 213 will be described. First, audio including the person's speech is collected by a microphone. Next, the person's speech is detected based on the collected audio, and the characteristic values of the detected person's speech are extracted. Next, the person's identification information is input via an input device such as a touch panel. Finally, the characteristic values of the person's speech and the person's identification information are stored in speaker information database 213 in a corresponding manner.
[0156] In addition, the speaker information database 213 will establish corresponding storage of identification information for multiple people and the musical instruments preferred by multiple people.
[0157] The speaker estimation unit 214 outputs a reproduction start instruction to the sound source reproduction unit 206 to begin reproducing a given sound. The speaker estimation unit 214 outputs a reproduction start instruction to the sound source reproduction unit 206 when the audio detection unit 212 detects a person's voice. Furthermore, the speaker estimation unit 214 can also output a reproduction end instruction to the sound source reproduction unit 206 when the audio detection unit 212 does not detect a person's voice. Additionally, the speaker estimation unit 214 can also output a reproduction start instruction to the sound source reproduction unit 206 when it recognizes the voice of one of multiple persons.
[0158] Furthermore, the speaker estimation unit 214 can select a sound corresponding to the current time period from multiple sounds of different types, and output a reproduction start instruction to the sound source reproduction unit 206 to begin reproducing the selected sound. For example, the speaker estimation unit 214 can select a sound corresponding to the current time period from among the following: a first sound corresponding to the morning period (8:00 AM to 12:00 PM), a second sound corresponding to the afternoon period (12:00 AM to 6:00 PM), and a third sound corresponding to the night period (6:00 PM to 10:00 PM). However, the time period is not limited to the above.
[0159] Speaker estimation unit 214 acquires information related to multiple individuals located in a given space. Speaker estimation unit 214 acquires audio information representing the average volume of the speech of each of the multiple individuals within a given period. Speaker estimation unit 214 controls the volume of each of the multiple segmented sound signals based on the audio information of each of the multiple individuals.
[0160] The speaker estimation unit 214 estimates the speaker detected by the audio detection unit 212 and calculates the average volume of the estimated speaker's speech over a given period. The speaker estimation unit 214 compares the speech features extracted by the feature extraction unit 202A with the speech features of multiple speakers stored in the speaker information database 213. If a feature value in the speech features of multiple speakers stored in the speaker information database 213 matches the feature value extracted by the feature extraction unit 202A, the speaker estimation unit 214 reads the corresponding recognition information from the speaker information database 213 and estimates the speaker based on the read recognition information.
[0161] Additionally, the speaker estimation unit 214 can also store audio data of the speech of each of the multiple individuals during a given past period. Furthermore, the speaker estimation unit 214 can also store volume data of the speech of each of the multiple individuals during a given past period.
[0162] Multiple individuals are pre-assigned to multiple segmented sound signals. The speaker estimation unit 214 refers to the speaker information database 213, extracts the preferred musical instrument corresponding to the user ID of the estimated speaker, and uses the segmented sound signal corresponding to the extracted preferred musical instrument from among the multiple segmented sound signals as the speaker's segmented sound signal and establishes a correspondence.
[0163] The speaker estimation unit 214 determines the volume of the first segmented sound signal corresponding to the first speaker 101 based on the average volume of the first speaker 101's speech within a given period, and instructs the determined volume to the first volume control unit 207. The speaker estimation unit 214 determines the volume of the second segmented sound signal corresponding to the second speaker 102 based on the average volume of the second speaker 102's speech within a given period, and instructs the determined volume to the second volume control unit 208. The speaker estimation unit 214 determines the volume of the third segmented sound signal corresponding to the third speaker 103 based on the average volume of the third speaker 103's speech within a given period, and instructs the determined volume to the third volume control unit 209. The speaker estimation unit 214 determines the volume of the fourth segmented sound signal corresponding to the fourth speaker 104 based on the average volume of the fourth speaker 104's speech within a given period, and instructs the determined volume to the fourth volume control unit 210.
[0164] The speaker estimation unit 214 determines the volume of the segmented sound signal corresponding to the speaker as a given volume when the average volume of the speaker's speech during a given period is less than a threshold. The given volume is, for example, the maximum volume. Furthermore, if the average volume of the speaker's speech during a given period is greater than or equal to the threshold, the speaker estimation unit 214 determines the volume of the segmented sound signal corresponding to the speaker as zero. Additionally, the given volume is not limited to the maximum volume; it can be 50% or 80% of the maximum volume, as long as it is greater than 0.
[0165] If the average volume of the speech of the first person 101 during a given period is less than a threshold, the speaker estimation unit 214 determines the volume of the first segmented sound signal corresponding to the first person 101 as a given volume. Furthermore, if the average volume of the speech of the first person 101 during a given period is greater than or equal to the threshold, the speaker estimation unit 214 determines the volume of the first segmented sound signal corresponding to the first person 101 as zero.
[0166] If the average volume of the speech of the second person 102 during a given period is less than a threshold, the speaker estimation unit 214 determines the volume of the second segmented sound signal corresponding to the second person 102 as a given volume. Furthermore, if the average volume of the speech of the second person 102 during a given period is greater than or equal to the threshold, the speaker estimation unit 214 determines the volume of the second segmented sound signal corresponding to the second person 102 as zero.
[0167] If the average volume of the speech of the third person 103 during a given period is less than a threshold, the speaker estimation unit 214 determines the volume of the third segment sound signal corresponding to the third person 103 as a given volume. Furthermore, if the average volume of the speech of the third person 103 during a given period is greater than or equal to the threshold, the speaker estimation unit 214 determines the volume of the third segment sound signal corresponding to the third person 103 as zero.
[0168] If the average volume of the speech of the fourth person 104 during a given period is less than a threshold, the speaker estimation unit 214 determines the volume of the fourth segmented sound signal corresponding to the fourth person 104 as a given volume. Furthermore, if the average volume of the speech of the fourth person 104 during a given period is greater than or equal to the threshold, the speaker estimation unit 214 determines the volume of the fourth segmented sound signal corresponding to the fourth person 104 as zero.
[0169] In addition, in this embodiment 2, the multiple characters are 4 people, but this disclosure is not particularly limited to this, and the multiple characters can be 2 or more people.
[0170] If a playback start instruction is input to the speaker estimation unit 214, the sound source reproduction unit 206 begins to reproduce multiple segmented sound signals obtained by segmenting the given sound. Furthermore, if a playback end instruction is input to the speaker estimation unit 214, the sound source reproduction unit 206 ends the reproduction of the multiple segmented sound signals obtained by segmenting the given sound.
[0171] Regardless of whether any of the characters 101 to 104 are present in room 100, the sound source reproduction unit 206 synchronously reproduces all of the first to fourth segmented sound signals. That is, as long as the voice of at least one of the characters 101 to 104 is detected, the sound source reproduction unit 206 synchronously reproduces all of the first to fourth segmented sound signals.
[0172] Volume control units 1 to 4 control units 210 control the volume of each of the multiple segmented sound signals based on the average volume of each person's voice within a given period.
[0173] The first volume control unit 207 controls the volume of the first segmented sound signal based on the average volume of the first person 101's speech over a given period. The first volume control unit 207 sets the volume of the first segmented sound signal input from the sound source reproduction unit 206 to the volume indicated by the speaker estimation unit 214.
[0174] The second volume control unit 208 controls the volume of the second segmented sound signal based on the average volume of the second person 102's speech over a given period. The second volume control unit 208 sets the volume of the second segmented sound signal input from the sound source reproduction unit 206 to the volume indicated by the speaker estimation unit 214.
[0175] The third volume control unit 209 controls the volume of the third segmented sound signal based on the average volume of the speech of the third person 103 within a given period. The third volume control unit 209 sets the volume of the third segmented sound signal input from the sound source reproduction unit 206 to the volume indicated by the speaker estimation unit 214.
[0176] The fourth volume control unit 210 controls the volume of the fourth segmented sound signal based on the average volume of the speech of the fourth person 104 within a given period. The fourth volume control unit 210 sets the volume of the fourth segmented sound signal input from the sound source reproduction unit 206 to the volume indicated by the speaker estimation unit 214.
[0177] Next, the sound reproduction process performed by the sound reproduction apparatus 2A in Embodiment 2 of this disclosure will be described.
[0178] Figure 8 This is a first flowchart for describing the sound reproduction process performed by the sound reproduction apparatus 2A in Embodiment 2 of this disclosure. Figure 9 This is a second flowchart for describing the sound reproduction process performed by the sound reproduction apparatus 2A in Embodiment 2 of this disclosure.
[0179] First, in step S31, the speaker estimation unit 214 determines whether the audio detection unit 212 has detected a person's voice. Here, if it is determined that no person's voice has been detected (no in step S31), the determination process in step S31 is repeated.
[0180] On the other hand, if it is determined that a person's voice has been detected (yes in step S31), in step S32, the speaker estimation unit 214 determines whether the first to fourth segmented sound signals obtained by segmenting the given sound are being reproduced.
[0181] Here, if it is determined that the first to fourth segmented sound signals are being reproduced (yes in step S32), the process proceeds to step S34.
[0182] On the other hand, if it is determined that the first to fourth segmented sound signals are not being reproduced (no in step S32), in step S33, the sound source reproduction unit 206 synchronously reproduces the first to fourth segmented sound signals obtained by segmenting the given sound. At this time, the speaker estimation unit 214 outputs a reproduction start instruction to the sound source reproduction unit 206 to start the reproduction of the given sound. If the reproduction start instruction is input from the speaker estimation unit 214, the sound source reproduction unit 206 reads the first to fourth segmented sound signals from the memory 205 and synchronously reproduces the read first to fourth segmented sound signals. The sound source reproduction unit 206 outputs the reproduced first to fourth segmented sound signals to the first volume control unit 207 to the fourth volume control unit 210, respectively.
[0183] Next, in step S34, the speaker estimation unit 214 estimates which person the speaker detected by the audio detection unit 212 is. The speaker estimation unit 214 compares the speech feature data extracted by the feature extraction unit 202A with the speech feature data of multiple people stored in the speaker information database 213. If there is a feature data in the speech feature data of multiple people stored in the speaker information database 213 that is consistent with the speech feature data extracted by the feature extraction unit 202A, then the speaker estimation unit 214 reads the recognition information corresponding to that feature data from the speaker information database 213, and estimates which person the speaker is based on the read recognition information.
[0184] Next, in step S35, the speaker estimation unit 214 calculates the average volume of the speech of each of the multiple speakers during a given period. For example, the speaker estimation unit 214 calculates the average volume of the speech during the period from the current moment to the moment one minute ago.
[0185] Next, in step S36, the speaker estimation unit 214 determines whether the average volume of the first person 101's speech during a given period is less than a threshold.
[0186] Here, if it is determined that the average volume of the speech of the first person 101 during a given period is less than a threshold (Yes in step S36), in step S37, the speaker estimation unit 214 determines the volume of the first segmented sound signal corresponding to the first person 101 as a given volume. The given volume is, for example, the maximum volume. The speaker estimation unit 214 instructs the first volume control unit 207 to set the volume of the first segmented sound signal determined to be the given volume. Based on the instruction from the speaker estimation unit 214, the first volume control unit 207 sets the volume of the first segmented sound signal to the given volume. The first volume control unit 207 outputs the first segmented sound signal with the set volume to the mixing unit 211.
[0187] On the other hand, if it is determined that the average volume of the speech of the first person 101 within a given period is above a threshold (not in step S36), in step S38, the speaker estimation unit 214 determines the volume of the first segmented sound signal corresponding to the first person 101 to be zero. The speaker estimation unit 214 instructs the first volume control unit 207 to set the volume of the first segmented sound signal determined to be zero. Based on the instruction from the speaker estimation unit 214, the first volume control unit 207 sets the volume of the first segmented sound signal to zero. The first volume control unit 207 outputs the first segmented sound signal with the set volume to the mixing unit 211.
[0188] Next, in step S39, the speaker estimation unit 214 determines whether the average volume of the second person 102's speech during a given period is less than a threshold.
[0189] Here, if it is determined that the average volume of the speech of the second person 102 during a given period is less than a threshold (Yes in step S39), in step S40, the speaker estimation unit 214 determines the volume of the second segmented sound signal corresponding to the second person 102 as a given volume. The given volume is, for example, the maximum volume. The speaker estimation unit 214 instructs the second volume control unit 208 to set the volume of the second segmented sound signal determined to be the given volume. Based on the instruction from the speaker estimation unit 214, the second volume control unit 208 sets the volume of the second segmented sound signal to the given volume. The second volume control unit 208 outputs the second segmented sound signal with the set volume to the mixing unit 211.
[0190] On the other hand, if it is determined that the average volume of the speech of the second person 102 within a given period is above a threshold (not in step S39), in step S41, the speaker estimation unit 214 determines the volume of the second segmented sound signal corresponding to the second person 102 to be zero. The speaker estimation unit 214 instructs the second volume control unit 208 to set the volume of the second segmented sound signal determined to be zero. Based on the instruction from the speaker estimation unit 214, the second volume control unit 208 sets the volume of the second segmented sound signal to zero. The second volume control unit 208 outputs the second segmented sound signal with the set volume to the mixing unit 211.
[0191] Next, in step S42, the speaker estimation unit 214 determines whether the average volume of the speech of the third person 103 during a given period is less than a threshold.
[0192] Here, if it is determined that the average volume of the speech of the third person 103 during a given period is less than a threshold (Yes in step S42), in step S43, the speaker estimation unit 214 determines the volume of the third segment sound signal corresponding to the third person 103 as a given volume. The given volume is, for example, the maximum volume. The speaker estimation unit 214 instructs the third volume control unit 209 to set the volume of the third segment sound signal determined to be the given volume. Based on the instruction from the speaker estimation unit 214, the third volume control unit 209 sets the volume of the third segment sound signal to the given volume. The third volume control unit 209 outputs the third segment sound signal with the set volume to the mixing unit 211.
[0193] On the other hand, if it is determined that the average volume of the speech of the third person 103 within a given period is above a threshold (not in step S42), in step S44, the speaker estimation unit 214 sets the volume of the third segmented sound signal corresponding to the third person 103 to zero. The speaker estimation unit 214 instructs the third volume control unit 209 to set the volume of the third segmented sound signal to zero. Based on the instruction from the speaker estimation unit 214, the third volume control unit 209 sets the volume of the third segmented sound signal to zero. The third volume control unit 209 outputs the third segmented sound signal with the set volume to the mixing unit 211.
[0194] Next, in step S45, the speaker estimation unit 214 determines whether the average volume of the fourth person 104's speech during a given period is less than a threshold.
[0195] Here, if it is determined that the average volume of the speech of the fourth person 104 within a given period is less than a threshold (Yes in step S45), in step S46, the speaker estimation unit 214 determines the volume of the fourth segment sound signal corresponding to the fourth person 104 as a given volume. The given volume is, for example, the maximum volume. The speaker estimation unit 214 instructs the fourth volume control unit 210 to set the volume of the fourth segment sound signal determined to be the given volume. Based on the instruction from the speaker estimation unit 214, the fourth volume control unit 210 sets the volume of the fourth segment sound signal to the given volume. The fourth volume control unit 210 outputs the fourth segment sound signal with the set volume to the mixing unit 211.
[0196] On the other hand, if it is determined that the average volume of the speech of the fourth person 104 within a given period is above a threshold (not in step S45), in step S47, the speaker estimation unit 214 sets the volume of the fourth segmented sound signal corresponding to the fourth person 104 to zero. The speaker estimation unit 214 instructs the fourth volume control unit 210 to set the volume of the fourth segmented sound signal to zero. Based on the instruction from the speaker estimation unit 214, the fourth volume control unit 210 sets the volume of the fourth segmented sound signal to zero. The fourth volume control unit 210 outputs the fourth segmented sound signal with the set volume to the mixing unit 211.
[0197] Next, in step S48, the mixing unit 211 outputs a synthesized sound signal obtained by combining the first to fourth segmented sound signals, whose volumes have been set, to the amplifier 3. The amplifier 3 amplifies the synthesized sound signal and outputs the amplified synthesized sound signal to the speaker 4. The speaker 4 converts the synthesized sound signal into synthesized sound and emits the synthesized sound obtained by combining the first to fourth segmented sound signals, whose respective volumes have been set by the first volume control unit 207 to the fourth volume control unit 210, into the room 100.
[0198] Therefore, if the average volume of the speech of the first person 101 during a given period is less than a threshold, a first segment of sound at a given volume is played from speaker 4; if the average volume of the speech of the first person 101 during a given period is greater than or equal to the threshold, the first segment of sound is not played from speaker 4. Similarly, if the average volume of the speech of the second person 102 during a given period is less than the threshold, a second segment of sound at a given volume is played from speaker 4; if the average volume of the speech of the second person 102 during a given period is greater than or equal to the threshold, the second segment of sound is not played from speaker 4. Furthermore, if the average volume of the speech of the third person 103 during a given period is less than the threshold, a third segment of sound at a given volume is played from speaker 4; if the average volume of the speech of the third person 103 during a given period is greater than or equal to the threshold, the third segment of sound is not played from speaker 4. Finally, if the average volume of the speech of the fourth person 104 during a given period is less than the threshold, a fourth segment of sound at a given volume is played from speaker 4; if the average volume of the speech of the fourth person 104 during a given period is greater than or equal to the threshold, the fourth segment of sound is not played from speaker 4.
[0199] The average volume of speech over a given period can also be considered as the amount of speech a character gives. In this way, when a character gives little speech, the segmented sound corresponding to the character is played, and when a character gives much speech, the segmented sound corresponding to the character is not played. Therefore, characters who give little speech can be made to notice that they give little speech, thus prompting them to speak.
[0200] Furthermore, the speaker estimation unit 214 can also select a given sound based on the average volume of the speech of all multiple characters within a given period. Alternatively, if the average volume of the speech of all multiple characters within a given period is less than a threshold, the speaker estimation unit 214 can select music with a fast tempo (BPM) above the threshold. This provides a sound environment conducive to dialogue between multiple characters. Conversely, if the average volume of the speech of all multiple characters within a given period is above the threshold, the speaker estimation unit 214 can select music with a slow tempo (BPM) below the threshold. This provides a sound environment conducive to calm dialogue between multiple characters.
[0201] Furthermore, the speaker estimation unit 214 can also determine the volume of the segmented sound signal corresponding to the speaker as a given volume if the average volume of the speaker's speech during a given period is above a threshold. The given volume is, for example, the maximum volume. Alternatively, the speaker estimation unit 214 can also determine the volume of the segmented sound signal corresponding to the speaker as zero if the average volume of the speaker's speech during a given period is below a threshold. The given volume is not limited to the maximum volume; it can be 50% or 80% of the maximum volume, as long as it is greater than 0. Thus, for example, if the volume of the segmented sound signal of a person who has been silent for a while becomes zero, it can encourage the person who has been silent for a while to speak.
[0202] (Implementation Method 3)
[0203] In the sound reproduction device shown in Embodiment 3, posture information related to the postures of multiple characters is acquired, and the volume of multiple segmented sound signals is controlled according to the postures of the multiple characters.
[0204] Figure 10 This is a diagram showing the overall structure of the sound reproduction system in Embodiment 3. Figure 11 This is a block diagram showing the structure of the sound reproduction device 2B in Embodiment 3.
[0205] The sound reproduction system in this embodiment 3 includes a camera 1, a sound reproduction device 2B, an amplifier 3, and a speaker 4. Furthermore, in this embodiment 3, the same reference numerals are used for structures identical to those in embodiment 1, and descriptions are omitted.
[0206] In this embodiment 3, multiple segmented audio signals are each associated with multiple locations of multiple people within a given space. The given space is, for example, a room 100 such as a conference room where multiple people can gather. The given space includes multiple locations. For example, each of the multiple locations is a seat for a person. Room 100 includes a first location 111, a second location 112, a third location 113, and a fourth location 114. The seat locations for the multiple people are not determined. At each location, there is one person.
[0207] Camera 1 captures images of a given space comprising multiple locations. Camera 1 is a fixed camera such as a surveillance camera, positioned to capture images of all multiple locations as a whole. In this embodiment 3, the given space comprises four locations, but this disclosure is not particularly limited to this, and may also include two or more locations. In this embodiment 3, camera 1 captures images of location 111, location 112, location 113, and location 114.
[0208] The sound reproduction device 2B acquires information related to multiple people located in a given space, and synchronously reproduces multiple segmented sound signals obtained by segmenting a given sound. Based on the information related to the multiple people, it outputs the multiple segmented sound signals to a speaker, which emits the multiple segmented sounds obtained from the multiple segmented sound signals into the given space. The sound reproduction device 2B detects people located at multiple positions (position 111, position 112, position 113, and position 114) in the given space (room 100) based on images captured by the camera 1. The sound reproduction device 2B controls the volume of each of the multiple segmented sound signals according to the posture of each person. The sound reproduction device 2B adjusts the volume of each of the first to fourth segmented sound signals according to the posture of the people located at positions 111, 212, 313, and 414. For example, the sound reproduction device 2B estimates the standing and sitting postures of the people located at the first position 111, the second position 112, the third position 113, and the fourth position 114, respectively.
[0209] The sound reproduction device 2B includes at least a computer system, which includes, for example, a control program, a processor or logic circuit for executing the control program, and a recording device such as internal memory or accessible external memory for storing the control program. Alternatively, the sound reproduction device 2B may be implemented, for example, through hardware installation based on the processing circuit, through the execution of a software program stored in memory or distributed from an external server by the processing circuit, or through a combination of these hardware and software installations.
[0210] The sound reproduction device 2B includes a person detection unit 201B, a feature extraction unit 202B, a posture information database 215, a posture estimation unit 216, a memory 205, a sound source reproduction unit 206, a first volume control unit 207, a second volume control unit 208, a third volume control unit 209, a fourth volume control unit 210, and a mixing unit 211.
[0211] The person detection unit 201B detects people located at multiple positions in room 100 based on images captured by camera 1. These multiple positions are predetermined. Therefore, the person detection unit 201B detects people present in the regions corresponding to each of the multiple positions within the image.
[0212] The feature extraction unit 202B extracts the body features of the person detected by the person detection unit 201.
[0213] The posture information database 215 stores multiple body feature quantities and posture information in a corresponding manner. Here, the method for registering a person's body feature quantities into the posture information database 215 will be described. First, a person is photographed using a camera. Next, the person is detected based on the photographed image, and the detected person's body feature quantities are extracted. Next, the person's posture information is input using an input device such as a touch panel. Finally, the person's body feature quantities and posture information are stored in the posture information database 215 in a corresponding manner.
[0214] For example, when estimating the standing and sitting postures of a person, the feature quantities of the standing person's body and the feature quantities of the sitting person's body are extracted. The feature quantities of the standing person's body are mapped to the posture information indicating that the person is standing, and the feature quantities of the sitting person's body are mapped to the posture information indicating that the person is sitting, and these are stored in the posture information database 215.
[0215] The posture estimation unit 216 outputs a reproduction start indication to the sound source reproduction unit 206 to begin reproducing a given sound. The posture estimation unit 216 outputs a reproduction start indication to the sound source reproduction unit 206 when a person is detected by the person detection unit 201B at any of the multiple positions. Furthermore, the posture estimation unit 216 can also output a reproduction end indication to the sound source reproduction unit 206 to end the reproduction of a given sound when no person is detected by the person detection unit 201B at any of the multiple positions.
[0216] Furthermore, the posture estimation unit 216 can select a sound corresponding to the current time period from multiple sounds of different types, and output a reproduction start instruction to the sound source reproduction unit 206 to begin reproducing the selected sound. For example, the posture estimation unit 216 can select a sound corresponding to the current time period from among the following: a first sound corresponding to the morning period (8:00 AM to 12:00 PM), a second sound corresponding to the afternoon period (12:00 AM to 6:00 PM), and a third sound corresponding to the night period (6:00 PM to 10:00 PM). However, the time period is not limited to the above.
[0217] The posture estimation unit 216 acquires information related to multiple characters located in a given space. The information related to the multiple characters includes posture information related to the individual postures of each character. The posture estimation unit 216 estimates the individual postures of each character. Based on the individual posture information of each character, the posture estimation unit 216 controls the volume of each of the multiple segmented sound signals.
[0218] The posture estimation unit 216 compares the body features extracted by the feature extraction unit 202B with multiple body features stored in the posture information database 215. If there is a feature among the multiple body features stored in the posture information database 215 that matches the body features extracted by the feature extraction unit 202B, the posture estimation unit 216 reads the posture information corresponding to that feature from the posture information database 215 and estimates the person's posture based on the read posture information.
[0219] Additionally, the pose estimation unit 216 can also input the body features of the person extracted by the feature extraction unit 202B into the pose estimation model, and obtain the person's pose from the pose estimation model. The pose estimation model is created through machine learning using the body features and the person's pose as training data. If the body features of the person extracted by the feature extraction unit 202B are input into the pose estimation model, the pose estimation model outputs the person's pose.
[0220] In this embodiment 3, the posture information indicates whether each of the multiple characters is standing. The posture estimation unit 216 estimates whether each character in a different position is standing or sitting.
[0221] Multiple segmented sound signals are each associated with multiple locations of multiple characters within a given space. The first segmented sound signal is associated with the first location 111, the second segmented sound signal is associated with the second location 112, the third segmented sound signal is associated with the third location 113, and the fourth segmented sound signal is associated with the fourth location 114.
[0222] The posture estimation unit 216 determines the volume of the first segmented sound signal corresponding to the first position 111 within the room 100 based on the posture of the person located at the first position 111, and instructs the determined volume to the first volume control unit 207. The posture estimation unit 216 determines the volume of the second segmented sound signal corresponding to the second position 112 within the room 100 based on the posture of the person located at the second position 112, and instructs the determined volume to the second volume control unit 208. The posture estimation unit 216 determines the volume of the third segmented sound signal corresponding to the third position 113 within the room 100 based on the posture of the person located at the third position 113, and instructs the determined volume to the third volume control unit 209. The posture estimation unit 216 determines the volume of the fourth segmented sound signal corresponding to the fourth position 114 within the room 100 based on the posture of the person located at the fourth position 114, and instructs the determined volume to the fourth volume control unit 210.
[0223] When the person is standing, the posture estimation unit 216 determines the volume of the segmented sound signal as a given volume. The given volume is, for example, the maximum volume. Furthermore, when the person is not standing, the posture estimation unit 216 determines the volume of the segmented sound signal to be zero. Additionally, the given volume is not limited to the maximum volume; it can be 50% or 80% of the maximum volume, as long as it is greater than zero.
[0224] When the person at position 111 in room 100 is standing, the posture estimation unit 216 determines the volume of the first segmented sound signal corresponding to position 111 to a given volume. Furthermore, when the person at position 111 in room 100 is not standing, i.e., when the person at position 111 in room 100 is sitting, the posture estimation unit 216 determines the volume of the first segmented sound signal corresponding to position 111 to zero.
[0225] When the person in the second position 112 within the room 100 is standing, the posture estimation unit 216 determines the volume of the second segmented sound signal corresponding to the second position 112 to a given volume. Furthermore, when the person in the second position 112 within the room 100 is not standing, i.e., when the person in the second position 112 within the room 100 is sitting, the posture estimation unit 216 determines the volume of the second segmented sound signal corresponding to the second position 112 to zero.
[0226] When the person in the third position 113 within the room 100 is standing, the posture estimation unit 216 determines the volume of the third segment sound signal corresponding to the third position 113 to a given volume. Furthermore, when the person in the third position 113 within the room 100 is not standing, i.e., when the person in the third position 113 within the room 100 is sitting, the posture estimation unit 216 determines the volume of the third segment sound signal corresponding to the third position 113 to zero.
[0227] When the person in the fourth position 114 within the room 100 is standing, the posture estimation unit 216 determines the volume of the fourth segment sound signal corresponding to the fourth position 114 to a given volume. Furthermore, when the person in the fourth position 114 within the room 100 is not standing, i.e., when the person in the fourth position 114 within the room 100 is sitting, the posture estimation unit 216 determines the volume of the fourth segment sound signal corresponding to the fourth position 114 to zero.
[0228] In addition, when there are no people in any of the positions within the room 100, the posture estimation unit 216 sets the volume of the segmented sound signal corresponding to each position to zero.
[0229] Furthermore, in this embodiment 3, there are four locations, but this disclosure is not particularly limited to this; there may be two or more locations. Additionally, the number of multiple segmented audio signals may be the same as or different from the number of locations.
[0230] Furthermore, in this embodiment 3, the posture estimation unit 216 estimates the standing posture and the sitting posture of the person, but this disclosure is not particularly limited to this, and it may only estimate the standing posture of the person. That is, the posture estimation unit 216 may determine the volume of the segmented sound signal to a given volume when the person's posture is standing, and determine the volume of the segmented sound signal to zero when the person's posture is not standing.
[0231] If a reproduction start instruction is input to the posture estimation unit 216, the sound source reproduction unit 206 begins reproducing multiple segmented sound signals obtained by segmenting a given sound. Furthermore, if a reproduction end instruction is input to the posture estimation unit 216, the sound source reproduction unit 206 ends the reproduction of the multiple segmented sound signals obtained by segmenting a given sound.
[0232] Regardless of whether there are people at positions 111 to 414 within room 100, the sound source reproduction unit 206 synchronously reproduces all of the first to fourth segmented sound signals. That is, as long as a person is detected at at least one of positions 111 to 414, the sound source reproduction unit 206 synchronously reproduces all of the first to fourth segmented sound signals.
[0233] Volume control units 1 to 4 control the volume of each of the multiple segmented sound signals according to the posture of each of the multiple characters.
[0234] The first volume control unit 207 controls the volume of the first segmented sound signal based on the posture of the person located at the first position 111. That is, the first volume control unit 207 controls the volume of the first segmented sound signal based on whether the person located at the first position 111 is standing. The first volume control unit 207 sets the volume of the first segmented sound signal input from the sound source reproduction unit 206 to the volume indicated by the posture estimation unit 216.
[0235] The second volume control unit 208 controls the volume of the second segmented sound signal based on the posture of the person located at the second position 112. That is, the second volume control unit 208 controls the volume of the second segmented sound signal based on whether the person located at the second position 112 is standing. The second volume control unit 208 sets the volume of the second segmented sound signal input from the sound source reproduction unit 206 to the volume indicated by the posture estimation unit 216.
[0236] The third volume control unit 209 controls the volume of the third segment sound signal based on the posture of the person located at the third position 113. That is, the third volume control unit 209 controls the volume of the third segment sound signal based on whether the person located at the third position 113 is standing. The third volume control unit 209 sets the volume of the third segment sound signal input from the sound source reproduction unit 206 to the volume indicated by the posture estimation unit 216.
[0237] The fourth volume control unit 210 controls the volume of the fourth segmented sound signal based on the posture of the person located at the fourth position 114. That is, the fourth volume control unit 210 controls the volume of the fourth segmented sound signal based on whether the person located at the fourth position 114 is standing. The fourth volume control unit 210 sets the volume of the fourth segmented sound signal input from the sound source reproduction unit 206 to the volume indicated by the posture estimation unit 216.
[0238] Next, the sound reproduction process performed by the sound reproduction device 2B in Embodiment 3 of this disclosure will be described.
[0239] Figure 12This is a first flowchart for describing the sound reproduction process performed by the sound reproduction apparatus 2B in Embodiment 3 of this disclosure. Figure 13 This is a second flowchart for describing the sound reproduction process performed by the sound reproduction apparatus 2B in Embodiment 3 of this disclosure.
[0240] First, in step S61, the posture estimation unit 216 determines whether the person detection unit 201B has detected a person at any of the multiple positions. Here, if it is determined that no person has been detected at any of the multiple positions (no in step S61), the determination process of step S61 is repeated.
[0241] On the other hand, if it is determined that a person is detected at any of the multiple locations (yes in step S61), in step S62, the posture estimation unit 216 determines whether the first to fourth segmented sound signals obtained by segmenting the given sound are being reproduced.
[0242] Here, if it is determined that the first to fourth segmented sound signals are being reproduced (yes in step S62), the process proceeds to step S64.
[0243] On the other hand, if it is determined that the first to fourth segmented sound signals are not being reproduced (no in step S62), in step S63, the sound source reproduction unit 206 synchronously reproduces the first to fourth segmented sound signals obtained by segmenting the given sound. At this time, the posture estimation unit 216 outputs a reproduction start instruction to the sound source reproduction unit 206 to start the reproduction of the given sound. If a reproduction start instruction is input from the posture estimation unit 216, the sound source reproduction unit 206 reads the first to fourth segmented sound signals from the memory 205 and synchronously reproduces the read first to fourth segmented sound signals. The sound source reproduction unit 206 outputs the reproduced first to fourth segmented sound signals to the first volume control unit 207 to the fourth volume control unit 210, respectively.
[0244] Next, in step S64, the posture estimation unit 216 estimates the posture of the person in positions 111 to 414 in room 100 based on the image captured by camera 1. More specifically, the posture estimation unit 216 estimates whether the person in positions 111 to 414 in room 100 is standing.
[0245] The person detection unit 201B detects people located at positions 111 to 414 based on images captured by the camera 1. The feature extraction unit 202B extracts the body features of the people detected by the person detection unit 201B. The pose estimation unit 216 compares the body features extracted by the feature extraction unit 202B with multiple body features stored in the pose information database 215. If there is a feature among the multiple body features stored in the pose information database 215 that matches the body feature extracted by the feature extraction unit 202B, the pose estimation unit 216 reads the pose information corresponding to that feature from the pose information database 215 and estimates the pose of the person based on the read pose information.
[0246] Next, in step S65, the posture estimation unit 216 determines whether there is a person at the first position 111 in the room 100.
[0247] Here, if it is determined that there is no person at position 111 (no in step S65), the process proceeds to step S68.
[0248] On the other hand, if it is determined that there is a person at position 111 (yes in step S65), in step S66, the posture estimation unit 216 determines whether the person at position 111 in room 100 is standing.
[0249] Here, if it is determined that the person at position 111 is standing (yes in step S66), in step S67, the posture estimation unit 216 determines the volume of the first segmented sound signal corresponding to position 111 as a given volume. The given volume is, for example, the maximum volume. The posture estimation unit 216 instructs the first volume control unit 207 to set the volume of the first segmented sound signal to the given volume. Based on the instruction from the posture estimation unit 216, the first volume control unit 207 sets the volume of the first segmented sound signal to the given volume. The first volume control unit 207 outputs the first segmented sound signal with the set volume to the mixing unit 211.
[0250] On the other hand, if it is determined that the person at position 111 is not standing, that is, if it is determined that the person at position 111 is sitting (not in step S66), in step S68, the posture estimation unit 216 sets the volume of the first segmented sound signal corresponding to position 111 to zero. The posture estimation unit 216 instructs the first volume control unit 207 to set the volume of the first segmented sound signal to zero. Based on the instruction from the posture estimation unit 216, the first volume control unit 207 sets the volume of the first segmented sound signal to zero. The first volume control unit 207 outputs the first segmented sound signal with the set volume to the mixing unit 211.
[0251] Next, in step S69, the posture estimation unit 216 determines whether there is a person at the second position 112 in the room 100.
[0252] Here, if it is determined that there is no person at position 2 112 (no in step S69), the process moves to step S72.
[0253] On the other hand, if it is determined that there is a person in the second position 112 (yes in step S69), in step S70, the posture estimation unit 216 determines whether the person in the second position 112 in the room 100 is standing.
[0254] Here, if it is determined that the person at position 112 is standing (yes in step S70), in step S71, the posture estimation unit 216 determines the volume of the second segment sound signal corresponding to the second position 112 as a given volume. The given volume is, for example, the maximum volume. The posture estimation unit 216 instructs the second volume control unit 208 to set the volume of the second segment sound signal to the given volume. Based on the instruction from the posture estimation unit 216, the second volume control unit 208 sets the volume of the second segment sound signal to the given volume. The second volume control unit 208 outputs the second segment sound signal with the set volume to the mixing unit 211.
[0255] On the other hand, if it is determined that the person in the second position 112 is not standing, that is, if it is determined that the person in the second position 112 is sitting (no in step S70), in step S72, the posture estimation unit 216 sets the volume of the second segmented sound signal corresponding to the second position 112 to zero. The posture estimation unit 216 instructs the second volume control unit 208 to set the volume of the second segmented sound signal to zero. Based on the instruction from the posture estimation unit 216, the second volume control unit 208 sets the volume of the second segmented sound signal to zero. The second volume control unit 208 outputs the second segmented sound signal with the set volume to the mixing unit 211.
[0256] Next, in step S73, the posture estimation unit 216 determines whether there is a person at the third position 113 in the room 100.
[0257] Here, if it is determined that there is no person at position 3 113 (no in step S73), the process moves to step S76.
[0258] On the other hand, if it is determined that there is a person in the third position 113 (yes in step S73), in step S74, the posture estimation unit 216 determines whether the person in the third position 113 in the room 100 is standing.
[0259] Here, if it is determined that the person at position 3 113 is standing (yes in step S74), in step S75, the posture estimation unit 216 determines the volume of the third segment sound signal corresponding to the third position 113 as a given volume. The given volume is, for example, the maximum volume. The posture estimation unit 216 instructs the third volume control unit 209 to set the volume of the third segment sound signal to the given volume. Based on the instruction from the posture estimation unit 216, the third volume control unit 209 sets the volume of the third segment sound signal to the given volume. The third volume control unit 209 outputs the third segment sound signal with the set volume to the mixing unit 211.
[0260] On the other hand, if it is determined that the person in position 3 113 is not standing, that is, if it is determined that the person in position 3 113 is sitting (not in step S74), in step S76, the posture estimation unit 216 sets the volume of the third segment sound signal corresponding to position 3 113 to zero. The posture estimation unit 216 instructs the third volume control unit 209 to set the volume of the third segment sound signal to zero. Based on the instruction from the posture estimation unit 216, the third volume control unit 209 sets the volume of the third segment sound signal to zero. The third volume control unit 209 outputs the third segment sound signal with the set volume to the mixing unit 211.
[0261] Next, in step S77, the posture estimation unit 216 determines whether there is a person at the fourth position 114 in the room 100.
[0262] Here, if it is determined that there is no person at position 4 114 (no in step S77), the process proceeds to step S80.
[0263] On the other hand, if it is determined that there is a person in the 4th position 114 (yes in step S77), in step S78, the posture estimation unit 216 determines whether the person in the 4th position 114 in the room 100 is standing.
[0264] Here, if it is determined that the person at position 4 114 is standing (yes in step S78), in step S79, the posture estimation unit 216 determines the volume of the fourth segment sound signal corresponding to position 4 114 as a given volume. The given volume is, for example, the maximum volume. The posture estimation unit 216 instructs the fourth volume control unit 210 to set the volume of the fourth segment sound signal to the given volume. Based on the instruction from the posture estimation unit 216, the fourth volume control unit 210 sets the volume of the fourth segment sound signal to the given volume. The fourth volume control unit 210 outputs the fourth segment sound signal with the set volume to the mixing unit 211.
[0265] On the other hand, if it is determined that the person in position 4 114 is not standing, that is, if it is determined that the person in position 4 114 is sitting (no in step S78), in step S80, the posture estimation unit 216 sets the volume of the fourth segment sound signal corresponding to position 4 114 to zero. The posture estimation unit 216 instructs the fourth volume control unit 210 to set the volume of the fourth segment sound signal to zero. Based on the instruction from the posture estimation unit 216, the fourth volume control unit 210 sets the volume of the fourth segment sound signal to zero. The fourth volume control unit 210 outputs the fourth segment sound signal with the set volume to the mixing unit 211.
[0266] Next, in step S81, the mixing unit 211 outputs a synthesized sound signal obtained by combining the first to fourth segmented sound signals, whose volumes have been set, to the amplifier 3. The amplifier 3 amplifies the synthesized sound signal and outputs the amplified synthesized sound signal to the speaker 4. The speaker 4 converts the synthesized sound signal into synthesized sound and emits the synthesized sound obtained by combining the first to fourth segmented sound signals, whose respective volumes have been set by the first volume control unit 207 to the fourth volume control unit 210, into the room 100.
[0267] Therefore, if the person in position 111 is standing, the first segment of sound at a given volume is played from speaker 4; if the person in position 111 is sitting, the first segment of sound is not played from speaker 4. Similarly, if the person in position 212 is standing, the second segment of sound at a given volume is played from speaker 4; if the person in position 212 is sitting, the second segment of sound is not played from speaker 4. Likewise, if the person in position 313 is standing, the third segment of sound at a given volume is played from speaker 4; if the person in position 313 is sitting, the third segment of sound is not played from speaker 4. Finally, if the person in position 414 is standing, the fourth segment of sound at a given volume is played from speaker 4; if the person in position 414 is sitting, the fourth segment of sound is not played from speaker 4.
[0268] In this way, if the characters in each position are standing, the sound is split, creating an environment where multiple characters can easily gather in a given space; if the characters in each position are sitting, the sound is split, creating an environment where it is easy to have a conversation in a given space.
[0269] Furthermore, in this embodiment 3, the posture of a person located in a given space is estimated based on the body feature quantities of the person extracted from the image captured by camera 1. However, this disclosure is not particularly limited to this, and other estimation methods may also be used. For example, the posture of a person may be estimated using the distance measured by an ultrasonic sensor, the posture of a person may be estimated using the distance or shape measured by a laser sensor such as LiDAR (Light Detection and Ranging), the posture of a person may be estimated using the acceleration measured by an accelerometer, the posture of a person may be estimated using the pressure value measured by a pressure sensor, or the posture of a person may be estimated using the distance measured by an infrared sensor.
[0270] (Implementation Method 4)
[0271] In the sound reproduction device shown in Embodiment 4, posture information related to the postures of multiple characters is acquired, and the volume of each of the multiple segmented sound signals is controlled based on the posture information of the multiple characters. The posture information in Embodiment 4 indicates whether each of the multiple characters is sitting in multiple positions within a given space, and indicates whether at least one of the multiple characters is in a drowsy posture.
[0272] Figure 14 This is a diagram showing the overall structure of the sound reproduction system in Embodiment 4. Figure 15 This is a block diagram showing the structure of the sound reproduction device 2C in Embodiment 4.
[0273] The sound reproduction system in this embodiment 4 includes a camera 1, a sound reproduction device 2C, an amplifier 3, and a speaker 4. Furthermore, in this embodiment 4, structures identical to those in embodiments 1 and 3 are labeled with the same reference numerals, and descriptions are omitted.
[0274] In this embodiment 4, the given space is, for example, a room 100 such as a conference room that can accommodate multiple people. The given space includes multiple locations. For example, each of the multiple locations is a seat for a person. Room 100 includes a first location 111, a second location 112, and a third location 113. The seating locations for multiple people are not determined. At each location, there is one person.
[0275] In this embodiment 4, the given space includes three positions, but this disclosure is not particularly limited to this and may include two or more positions. In this embodiment 4, camera 1 captures images of position 111, position 112, and position 113.
[0276] The sound reproduction device 2C acquires information related to multiple people located in a given space, and synchronously reproduces multiple segmented sound signals obtained by segmenting a given sound. Based on the information related to the multiple people, it outputs the multiple segmented sound signals to a speaker 4, which emits the multiple segmented sounds obtained from the multiple segmented sound signals into the given space. The sound reproduction device 2C detects people located at multiple positions (position 111, position 112, and position 113) in the given space (room 100) based on images captured by the camera 1. The sound reproduction device 2C controls the volume of each of the multiple segmented sound signals according to the posture of each person. The sound reproduction device 2C changes the volume of each of the first to fourth segmented sound signals according to the posture of the people located at positions 111, 212, and 313. For example, the sound reproduction device 2C estimates the standing posture, sitting posture, and drowsy posture of the people located at positions 111, 212, and 313.
[0277] The sound reproduction device 2C includes at least a computer system, which includes, for example, a control program, a processor or logic circuitry for executing the control program, and a recording device such as internal memory or accessible external memory for storing the control program. Alternatively, the sound reproduction device 2C may be implemented, for example, through hardware installation based on the processing circuitry, through the execution of software programs stored in memory or distributed from an external server by the processing circuitry, or through a combination of these hardware and software installations.
[0278] The sound reproduction device 2C includes a person detection unit 201B, a feature extraction unit 202B, a posture information database 215C, a posture estimation unit 216C, a memory 205, a sound source reproduction unit 206, a first volume control unit 207, a second volume control unit 208, a third volume control unit 209, a fourth volume control unit 210, and a mixing unit 211.
[0279] The posture information database 215C stores multiple body feature quantities and posture information in a corresponding manner. Here, the method for registering a person's body feature quantities into the posture information database 215C is explained. First, a person is photographed using a camera. Next, the person is detected based on the captured image, and the detected person's body feature quantities are extracted. Next, the person's posture information is input using an input device such as a touch panel. Finally, the person's body feature quantities and posture information are stored in the posture information database 215C in a corresponding manner.
[0280] For example, when estimating a person's standing posture, sitting posture, and drowsy posture, feature quantities of the body of the standing person, the sitting person, and the drowsy person are extracted. The feature quantities of the standing person's body are mapped to posture information indicating that the person is standing, and the feature quantities of the sitting person's body are mapped to posture information indicating that the person is sitting, and these are stored in posture information database 215C. Furthermore, the feature quantities of the drowsy person's body are mapped to posture information indicating that the person is drowsy, and these are also stored in posture information database 215C.
[0281] Postures that indicate drowsiness include a seated person hunching over, head bobbing back and forth, frequent yawning, frequent eye rubbing, or frequent blinking.
[0282] The posture estimation unit 216C outputs a reproduction start indication to the sound source reproduction unit 206 to begin reproducing a given sound. The posture estimation unit 216C outputs a reproduction start indication to the sound source reproduction unit 206 when a person is detected by the person detection unit 201B at any of the multiple positions. Furthermore, the posture estimation unit 216C can also output a reproduction end indication to the sound source reproduction unit 206 when a person is not detected by the person detection unit 201B at any of the multiple positions.
[0283] Furthermore, the posture estimation unit 216C can select a sound corresponding to the current time period from multiple sounds of different types, and output a reproduction start instruction to the sound source reproduction unit 206 to begin reproducing the selected sound. For example, the posture estimation unit 216C can select a sound corresponding to the current time period from among the following: a first sound corresponding to the morning period (8:00 AM to 12:00 PM), a second sound corresponding to the afternoon period (12:00 AM to 6:00 PM), and a third sound corresponding to the night period (6:00 PM to 10:00 PM). However, the time period is not limited to the above.
[0284] The posture estimation unit 216C acquires information related to multiple characters located in a given space. This information includes posture information related to the individual postures of each character. The posture estimation unit 216C estimates the individual postures of each character. Based on the individual posture information of each character, the posture estimation unit 216C controls the volume of each of the multiple segmented sound signals.
[0285] The posture estimation unit 216C compares the body features extracted by the feature extraction unit 202B with multiple body features stored in the posture information database 215C. If there is a feature among the multiple body features stored in the posture information database 215C that matches the body features extracted by the feature extraction unit 202B, the posture estimation unit 216C reads the posture information corresponding to that feature from the posture information database 215C and estimates the person's posture based on the read posture information.
[0286] Additionally, the pose estimation unit 216C can also input the body features of the person extracted by the feature extraction unit 202B into the pose estimation model to obtain the person's pose. The pose estimation model is created through machine learning using the body features and the person's pose as training data. If the body features of the person extracted by the feature extraction unit 202B are input into the pose estimation model, the pose estimation model outputs the person's pose.
[0287] In this embodiment 4, the posture information indicates whether each of the multiple individuals is sitting in multiple positions within a given space, and whether at least one of the multiple individuals is in a drowsy posture. The posture estimation unit 216C estimates whether the individuals in the multiple positions are in a standing or sitting posture. Furthermore, the posture estimation unit 216C estimates whether at least one of the multiple individuals is in a drowsy posture.
[0288] Multiple segmented sound signals are each associated with multiple locations of multiple characters within a given space. The first segmented sound signal is associated with the first location 111, the second segmented sound signal is associated with the second location 112, and the third segmented sound signal is associated with the third location 113.
[0289] The posture estimation unit 216C determines the volume of the first segmented sound signal corresponding to the first position 111 within the room 100 based on the posture of the person located at the first position 111, and instructs the determined volume to the first volume control unit 207. The posture estimation unit 216C determines the volume of the second segmented sound signal corresponding to the second position 112 within the room 100 based on the posture of the person located at the second position 112, and instructs the determined volume to the second volume control unit 208. The posture estimation unit 216C determines the volume of the third segmented sound signal corresponding to the third position 113 within the room 100 based on the posture of the person located at the third position 113, and instructs the determined volume to the third volume control unit 209. The posture estimation unit 216C determines the volume of the fourth segmented sound signal based on the postures of the persons located at the first to fourth positions 114 within the room 100, and instructs the determined volume to the fourth volume control unit 210.
[0290] When a person is sitting, the posture estimation unit 216C determines the volume of the segmented sound signal corresponding to the person's sitting position as a given volume. The given volume is, for example, the maximum volume. Furthermore, when a person is not sitting, the posture estimation unit 216C determines the volume of the segmented sound signal corresponding to the position the person is not sitting at as zero. Furthermore, when at least one of the multiple people is in a drowsy posture, the posture estimation unit 216C determines the volume of a separate segmented sound signal, different from the multiple segmented sound signals corresponding to each of the multiple positions, used to suppress drowsiness, as a given volume. Furthermore, when none of the multiple people are in a drowsy posture, the posture estimation unit 216C determines the volume of another segmented sound signal as zero. Additionally, the given volume is not limited to the maximum volume; it can be 50% or 80% of the maximum volume, as long as it is greater than zero.
[0291] When the person in the first position 111 within the room 100 is sitting, the posture estimation unit 216C determines the volume of the first segmented sound signal corresponding to the first position 111 to a given volume. Furthermore, when the person in the first position 111 within the room 100 is not sitting, i.e., when the person in the first position 111 within the room 100 is standing, the posture estimation unit 216C determines the volume of the first segmented sound signal corresponding to the first position 111 to zero.
[0292] When the person in the second position 112 within the room 100 is sitting, the posture estimation unit 216C determines the volume of the second segmented sound signal corresponding to the second position 112 to a given volume. Furthermore, when the person in the second position 112 within the room 100 is not sitting, i.e., when the person in the second position 112 within the room 100 is standing, the posture estimation unit 216C determines the volume of the second segmented sound signal corresponding to the second position 112 to zero.
[0293] When the person in the third position 113 within the room 100 is sitting, the posture estimation unit 216C determines the volume of the third segment sound signal corresponding to the third position 113 to a given volume. Furthermore, when the person in the third position 113 within the room 100 is not sitting, i.e., when the person in the third position 113 within the room 100 is standing, the posture estimation unit 216C determines the volume of the third segment sound signal corresponding to the third position 113 to zero.
[0294] When at least one person in positions 111 to 414 within room 100 feels drowsy, the posture estimation unit 216C determines the volume of the fourth segment sound signal used to suppress drowsiness to a given volume. Furthermore, when none of the persons in positions 111 to 414 within room 100 feel drowsy, the posture estimation unit 216C determines the volume of the fourth segment sound signal to zero.
[0295] The fourth segment sound signal used to suppress drowsiness includes a brighter sound than the first to third segment sound signals, the sound of a percussion instrument with an appropriate rhythm, or an effect sound to attract the attention of the person. In addition, the posture estimation unit 216C can also make the given volume of the fourth segment sound signal higher than the given volume of the first to third segment sound signals.
[0296] In addition, when there are no people in any of the positions within the room 100, the posture estimation unit 216C sets the volume of the segmented sound signal corresponding to each position to zero.
[0297] Furthermore, in this embodiment 4, there are three locations, but this disclosure is not particularly limited to this, and there may be two or more locations.
[0298] Furthermore, in this embodiment 4, the posture estimation unit 216C estimates the posture of a person sitting, standing, and feeling drowsy. However, this disclosure is not particularly limited to this, and it is also possible to estimate the posture of a person sitting and feeling drowsy. That is, the posture estimation unit 216C may determine the volume of the segmented sound signal to a given volume when the person's posture is sitting, and determine the volume of the segmented sound signal to zero when the person's posture is not sitting.
[0299] Volume control units 1 to 4 control the volume of each of the multiple segmented sound signals according to the posture of each of the multiple characters.
[0300] The first volume control unit 207 controls the volume of the first segmented sound signal based on the posture of the person located at the first position 111. That is, the first volume control unit 207 controls the volume of the first segmented sound signal based on whether the person located at the first position 111 is sitting. The first volume control unit 207 sets the volume of the first segmented sound signal input from the sound source reproduction unit 206 to the volume indicated by the posture estimation unit 216C.
[0301] The second volume control unit 208 controls the volume of the second segmented sound signal based on the posture of the person located at the second position 112. That is, the second volume control unit 208 controls the volume of the second segmented sound signal based on whether the person located at the second position 112 is sitting. The second volume control unit 208 sets the volume of the second segmented sound signal input from the sound source reproduction unit 206 to the volume indicated by the posture estimation unit 216C.
[0302] The third volume control unit 209 controls the volume of the third segment sound signal based on the posture of the person located at the third position 113. That is, the third volume control unit 209 controls the volume of the third segment sound signal based on whether the person located at the third position 113 is sitting. The third volume control unit 209 sets the volume of the third segment sound signal input from the sound source reproduction unit 206 to the volume indicated by the posture estimation unit 216C.
[0303] The fourth volume control unit 210 controls the volume of the fourth segmented sound signal based on the posture of at least one person located at positions 111 to 414. That is, the fourth volume control unit 210 controls the volume of the fourth segmented sound signal based on whether the at least one person located at positions 111 to 414 is drowsy. The fourth volume control unit 210 sets the volume of the fourth segmented sound signal input from the sound source reproduction unit 206 to the volume indicated by the posture estimation unit 216C.
[0304] Next, the sound reproduction process performed by the sound reproduction apparatus 2C in Embodiment 4 of this disclosure will be described.
[0305] Figure 16 This is a first flowchart for describing the sound reproduction process performed by the sound reproduction apparatus 2C in Embodiment 4 of this disclosure. Figure 17 This is a second flowchart for describing the sound reproduction process performed by the sound reproduction apparatus 2C in Embodiment 4 of this disclosure.
[0306] Processing in steps S91 to S93 Figure 12 The processes in steps S61 to S63 are the same, so the explanation is omitted.
[0307] Next, in step S94, the posture estimation unit 216C estimates the posture of the person in positions 111 to 313 within the room 100 based on the image captured by the camera 1. More specifically, the posture estimation unit 216C estimates whether the person in positions 111 to 313 within the room 100 is sitting, and whether the person in positions 111 to 313 within the room 100 is feeling drowsy.
[0308] Next, in step S95, the posture estimation unit 216C determines whether there is a person at the first position 111 in the room 100.
[0309] Here, if it is determined that there is no person at position 111 (no in step S95), the process proceeds to step S98.
[0310] On the other hand, if it is determined that there is a person in the first position 111 (yes in step S95), in step S96, the posture estimation unit 216C determines whether the person in the first position 111 in the room 100 is sitting.
[0311] Here, if it is determined that the person at position 111 is sitting (yes in step S96), in step S97, the posture estimation unit 216C determines the volume of the first segmented sound signal corresponding to position 111 as a given volume. The given volume is, for example, the maximum volume. The posture estimation unit 216C instructs the first volume control unit 207 to set the volume of the first segmented sound signal to the given volume. Based on the instruction from the posture estimation unit 216C, the first volume control unit 207 sets the volume of the first segmented sound signal to the given volume. The first volume control unit 207 outputs the first segmented sound signal with the set volume to the mixing unit 211.
[0312] On the other hand, if it is determined that the person in position 111 is not sitting, that is, if it is determined that the person in position 111 is standing (no in step S96), in step S98, the posture estimation unit 216C sets the volume of the first segmented sound signal corresponding to position 111 to zero. The posture estimation unit 216C instructs the first volume control unit 207 to set the volume of the first segmented sound signal to zero. Based on the instruction from the posture estimation unit 216C, the first volume control unit 207 sets the volume of the first segmented sound signal to zero. The first volume control unit 207 outputs the first segmented sound signal with the set volume to the mixing unit 211.
[0313] Next, in step S99, the posture estimation unit 216C determines whether there is a person at the second position 112 in the room 100.
[0314] Here, if it is determined that there is no person at position 2 112 (no in step S99), the process moves to step S102.
[0315] On the other hand, if it is determined that there is a person in the second position 112 (yes in step S99), in step S100, the posture estimation unit 216C determines whether the person in the second position 112 in the room 100 is sitting.
[0316] Here, if it is determined that the person at position 112 is sitting (yes in step S100), in step S101, the posture estimation unit 216C determines the volume of the second segmented sound signal corresponding to the second position 112 as a given volume. The given volume is, for example, the maximum volume. The posture estimation unit 216C instructs the second volume control unit 208 to set the volume of the second segmented sound signal to the given volume. Based on the instruction from the posture estimation unit 216C, the second volume control unit 208 sets the volume of the second segmented sound signal to the given volume. The second volume control unit 208 outputs the second segmented sound signal with the set volume to the mixing unit 211.
[0317] On the other hand, if it is determined that the person in position 112 is not sitting, that is, if it is determined that the person in position 112 is standing (no in step S100), in step S102, the posture estimation unit 216C sets the volume of the second segmented sound signal corresponding to the second position 112 to zero. The posture estimation unit 216C instructs the second volume control unit 208 to set the volume of the second segmented sound signal to zero. Based on the instruction from the posture estimation unit 216C, the second volume control unit 208 sets the volume of the second segmented sound signal to zero. The second volume control unit 208 outputs the second segmented sound signal with the set volume to the mixing unit 211.
[0318] Next, in step S103, the posture estimation unit 216C determines whether there is a person at the third position 113 in the room 100.
[0319] Here, if it is determined that there is no person at position 3 113 (no in step S103), the process moves to step S106.
[0320] On the other hand, if it is determined that there is a person in the third position 113 (yes in step S103), in step S104, the posture estimation unit 216C determines whether the person in the third position 113 in the room 100 is sitting.
[0321] Here, if it is determined that the person in position 3 113 is sitting (yes in step S104), in step S105, the posture estimation unit 216C determines the volume of the third segment sound signal corresponding to position 3 113 as a given volume. The given volume is, for example, the maximum volume. The posture estimation unit 216C instructs the third volume control unit 209 to set the volume of the third segment sound signal to the given volume. Based on the instruction from the posture estimation unit 216C, the third volume control unit 209 sets the volume of the third segment sound signal to the given volume. The third volume control unit 209 outputs the third segment sound signal with the set volume to the mixing unit 211.
[0322] On the other hand, if it is determined that the person in position 3 113 is not sitting, that is, if it is determined that the person in position 3 113 is standing (no in step S104), in step S106, the posture estimation unit 216C sets the volume of the third segment sound signal corresponding to position 3 113 to zero. The posture estimation unit 216C instructs the third volume control unit 209 to set the volume of the third segment sound signal to zero. Based on the instruction from the posture estimation unit 216C, the third volume control unit 209 sets the volume of the third segment sound signal to zero. The third volume control unit 209 outputs the third segment sound signal with the set volume to the mixing unit 211.
[0323] Next, in step S107, the posture estimation unit 216C determines whether there are any people feeling sleepy in the first position 111 to the fourth position 114 in the room 100.
[0324] Here, if it is determined that a person is feeling drowsy at positions 111 to 414 (yes in step S107), in step S108, the posture estimation unit 216C determines the volume of the fourth segment sound signal used to suppress drowsiness as a given volume. The given volume is, for example, the maximum volume. The posture estimation unit 216C instructs the fourth volume control unit 210 to set the volume of the fourth segment sound signal to the given volume. Based on the instruction from the posture estimation unit 216C, the fourth volume control unit 210 sets the volume of the fourth segment sound signal to the given volume. The fourth volume control unit 210 outputs the fourth segment sound signal with the set volume to the mixing unit 211.
[0325] On the other hand, if it is determined that no person in positions 111 to 414 is feeling drowsy (not in step S107), in step S109, the posture estimation unit 216C sets the volume of the fourth segment sound signal used to suppress drowsiness to zero. The posture estimation unit 216C instructs the fourth volume control unit 210 to set the volume of the fourth segment sound signal to zero. Based on the instruction from the posture estimation unit 216C, the fourth volume control unit 210 sets the volume of the fourth segment sound signal to zero. The fourth volume control unit 210 outputs the fourth segment sound signal with the set volume to the mixing unit 211.
[0326] Next, in step S110, the mixing unit 211 outputs a synthesized sound signal obtained by combining the first to fourth segmented sound signals, whose volumes have been set, to the amplifier 3. The amplifier 3 amplifies the synthesized sound signal and outputs the amplified synthesized sound signal to the speaker 4. The speaker 4 converts the synthesized sound signal into synthesized sound and emits the synthesized sound obtained by combining the first to fourth segmented sound signals, whose respective volumes have been set by the first volume control unit 207 to the fourth volume control unit 210, into the room 100.
[0327] Therefore, if the person in position 111 is sitting, the first segment of sound at a given volume is played from speaker 4; if the person in position 111 is standing, the first segment of sound is not played from speaker 4. Similarly, if the person in position 212 is sitting, the second segment of sound at a given volume is played from speaker 4; if the person in position 212 is standing, the second segment of sound is not played from speaker 4. Furthermore, if the person in position 313 is sitting, the third segment of sound at a given volume is played from speaker 4; if the person in position 313 is standing, the third segment of sound is not played from speaker 4. Additionally, if any person in positions 111 to 414 is drowsy, the fourth segment of sound at a given volume is played from speaker 4; if no person in positions 111 to 414 is drowsy, the fourth segment of sound is not played from speaker 4.
[0328] For example, when only one person is seated in position 111, the first segmented sound representing the sound of a harp is emitted from speaker 4. Furthermore, when both positions 111 and 212 are seated, the first segmented sound representing the sound of a harp and the second segmented sound representing the sound of a violin are emitted from speaker 4. Furthermore, when positions 111 through 313 are seated, the first segmented sound representing the sound of a harp, the second segmented sound representing the sound of a violin, and the third segmented sound representing the sound of a double bass are emitted from speaker 4. Further, when positions 111 through 313 are seated and the person in position 212 is drowsy, the first segmented sound representing the sound of a harp, the second segmented sound representing the sound of a violin, the third segmented sound representing the sound of a double bass, and the fourth segmented sound representing the sound of a marimba and a drum are emitted from speaker 4.
[0329] In this way, if a character is drowsy at any of the multiple locations, a segmented sound is emitted to suppress drowsiness, thus stimulating the drowsy character and making him focus on the conversation.
[0330] Furthermore, no split sound is emitted if the person is standing, but a split sound is emitted if the person is seated, thus prompting the participants in the dialogue (meeting) to take their seats.
[0331] Furthermore, the posture estimation unit 216C can also change the given sound to fast-paced music with a BPM above a threshold if it determines that a character is feeling drowsy in positions 111 to 414. This can further suppress drowsiness in the character.
[0332] Furthermore, in the above embodiments, each component may be constructed using dedicated hardware, or implemented by executing software programs suitable for each component. Each component may also be implemented by a program execution unit such as a CPU or processor reading and executing software programs recorded on a recording medium such as a hard disk or semiconductor memory. Alternatively, the program may be recorded on a recording medium and transferred, or transferred via a network, thereby allowing the program to be implemented by a separate computer system.
[0333] The devices described in this disclosure typically implement part or all of their functionality as LSIs (Large Scale Integration) as integrated circuits. They can be individually monolithized or monolithized to include part or all of them. Furthermore, the integration is not limited to LSIs; it can also be implemented by dedicated circuits or general-purpose processors. FPGAs (Field Programmable Gate Arrays) that are programmable after LSI fabrication, or reconfigurable processors capable of reconfiguring the connections and settings of circuit cells within the LSI, can also be utilized.
[0334] In addition, some or all of the functions of the apparatus involved in the embodiments of this disclosure can be implemented by executing programs through a processor such as a CPU.
[0335] Furthermore, all the figures used above are illustrative for the purpose of specifically illustrating this disclosure, and this disclosure is not limited to the illustrative figures.
[0336] Furthermore, the order in which the steps shown in the flowchart above are executed is illustrative for the purpose of specifically illustrating this disclosure, and other orders may also be used to achieve the same effect. Additionally, some of the above steps may be executed simultaneously (in parallel) with other steps.
[0337] Industrial availability
[0338] The technology disclosed herein is useful as a technology for reproducing sound because it can bring together multiple characters in a given space and can activate the dialogue of multiple characters in a given space.
Claims
1. A sound reproduction method, which is a sound reproduction method in a computer, comprising: Obtain information about multiple people located in a given space; To enable the synchronous reproduction of multiple segmented sound signals obtained by segmenting a given sound; and Based on information related to the multiple individuals, the multiple segmented sound signals are output to a speaker, and the speaker emits multiple segmented sounds obtained from the multiple segmented sound signals into the given space.
2. The sound reproduction method according to claim 1, wherein, The information related to the multiple individuals includes the identification results obtained by identifying each of the multiple individuals. The sound reproduction method further includes controlling the volume of each of the multiple segmented sound signals based on whether the multiple characters are identified.
3. The sound reproduction method according to claim 2, wherein, The volume control includes: If a person is identified, the volume of the segmented sound signal corresponding to that person is set to a given volume; if the person is not identified, the volume of the segmented sound signal corresponding to that person is set to zero.
4. The sound reproduction method according to claim 1, wherein, Information relating to the plurality of persons includes audio information representing the average volume of the speech of each of the plurality of persons during a given period. The sound reproduction method further includes controlling the volume of each of the plurality of segmented sound signals based on the average volume of the speech of each of the plurality of persons during the given period.
5. The sound reproduction method according to claim 4, wherein, The volume control includes: If the average volume of a person's voice is less than a threshold, the volume of the segmented sound signal corresponding to the person is determined as a given volume; if the average volume of a person's voice is greater than or equal to the threshold, the volume of the segmented sound signal corresponding to the person is determined as zero.
6. The sound reproduction method according to claim 1, wherein, Information relating to the multiple characters includes posture information related to the individual postures of each character. The sound reproduction method further includes controlling the volume of each of the multiple segmented sound signals according to the respective postures of the multiple characters.
7. The sound reproduction method according to claim 6, wherein, The posture information indicates whether each of the multiple figures is standing. The volume control includes: setting the volume of the segmented sound signal to a given volume when the person is standing, and setting the volume of the segmented sound signal to zero when the person is not standing.
8. The sound reproduction method according to claim 7, wherein, Each of the multiple segmented sound signals is respectively associated with a multiple location of the multiple characters in the given space.
9. The sound reproduction method according to claim 6, wherein, The posture information indicates whether each of the multiple individuals is sitting in multiple positions within the given space, and whether at least one of the multiple individuals is in a drowsy posture. The volume control includes: When a person is sitting, the volume of the segmented sound signal corresponding to the position where the person is sitting is determined to be a given volume. When the person is not sitting, the volume of the segmented sound signal corresponding to the position where the person is not sitting is determined to be zero. When at least one of the multiple people is in a drowsy posture, the volume of another segmented sound signal, which is different from the multiple segmented sound signals corresponding to each of the multiple positions and is used to suppress drowsiness, is determined to be a given volume. When none of the multiple people are in a drowsy posture, the volume of the other segmented sound signal is determined to be zero.
10. The sound reproduction method according to any one of claims 1 to 9, wherein, The given sound is segmented according to its frequency.
11. The sound reproduction method according to any one of claims 1 to 9, wherein, The given sound is music, and is segmented according to the multiple performance parts that constitute the music.
12. A sound reproduction device, comprising: The Acquisition Department is responsible for acquiring information related to multiple individuals located in a given space. The reproduction unit synchronously reproduces multiple segmented sound signals obtained by segmenting a given sound; and The output unit outputs the multiple segmented sound signals to a speaker based on information related to the multiple characters, and the speaker emits the multiple segmented sounds obtained from the multiple segmented sound signals into the given space.
13. A sound reproduction program that enables a computer to perform functions such as: Obtain information about multiple people located in a given space; To enable the synchronous reproduction of multiple segmented sound signals obtained by segmenting a given sound; Based on information related to the multiple individuals, the multiple segmented sound signals are output to a speaker, and the speaker emits multiple segmented sounds obtained from the multiple segmented sound signals into the given space.
Citation Information
Patent Citations
Methods and systems to prevent eavesdropping on private conversations in public places
JP2011528445A