Information processing device and information processing program

The information processing device addresses communication challenges in quiet spaces by analyzing speech emotions, adding reverberation to positive utterances, and controlling background sounds to enhance emotional responses and communication ease.

JP7870718B2Active Publication Date: 2026-06-05TAKENAKA CORP

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
TAKENAKA CORP
Filing Date
2022-12-01
Publication Date
2026-06-05

AI Technical Summary

Technical Problem

Existing technologies fail to improve ease of communication in open-plan offices and similar spaces where sound environments make it difficult to converse due to excessive sound absorption leading to quietness and reduced acoustic transmission.

Method used

An information processing device that performs emotion analysis on speech segments, adds a reverberation effect to positive utterances to create background sound information, and controls its playback to enhance positive emotional responses and improve communication.

Benefits of technology

Enhances communication by creating unrecognizable background sounds with positive emotional effects, broadening attention and receptivity, thus improving ease of conversation in target spaces.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007870718000003
    Figure 0007870718000003
  • Figure 0007870718000004
    Figure 0007870718000004
  • Figure 0007870718000005
    Figure 0007870718000005
Patent Text Reader

Abstract

To obtain information processing devices and information processing programs that can improve ease of communication in a target space.SOLUTION: An information processing device 10 includes: a creation unit 11A that performs emotion analysis for each segment of speech information indicating speech uttered by a person, and creates background sound information indicating background sound whose content cannot be recognized and which positively affects the emotions of at least some of the plurality of persons present in the target space, by adding a reverberation effect to the speech that gives a positive impression; and a control unit 11C that controls playback of the background sound indicated by the background sound information in the target space.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an information processing apparatus and an information processing program.

Background Art

[0002] Conventionally, the following techniques have been available as techniques that can be applied to improve the acoustic environment.

[0003] Patent Document 1 discloses a work environment improvement system aimed at generating more accurate masking sounds according to noise and more accurately reducing the influence of noise on workers.

[0004] This work environment improvement system is a work environment improvement system for making the work environment regarding the noise of a worker working in a predetermined space better by outputting masking sounds from masking sound output means. It has a worker position recognition means for recognizing the location of the worker in the space, and a noise information acquisition means for acquiring noise information such as the type and location of the noise. Further, this work environment improvement system has a masking sound determination means for determining the type of masking sound for weakening the sensitivity of the worker to the noise based on the noise information acquired by the noise information acquisition means, and good sound image localization information for weakening the sensitivity of the worker to the noise set in advance. And it has a masking sound control unit for performing sound image localization of the masking sound by the masking sound output means based on the sound image localization information. And this work environment improvement system is characterized in that the masking sound control unit performs sound image localization of the masking sound based on the location information of the worker from the worker position recognition means, the type information of the masking sound from the masking sound determination means, and the good sound image localization information.

[0005] Furthermore, Patent Document 2 discloses an intellectual productivity improvement support system aimed at improving the work environment and thereby stimulating the intellectual activities of individual workers such as office workers, engineers, and working women, and improving the intellectual productivity of each worker.

[0006] This intellectual productivity improvement support system is formed from a small space of a predetermined volume and an additional sound generation means that emits a predetermined additional sound in the small space at a sound pressure level that allows a worker present in the small space to hear the ambient noise generated in the surroundings, and which allows the worker to unconsciously refresh themselves and activate their intellectual activities. The sound pressure level of the additional sound is in the range of -6dB to +8dB relative to the sound pressure level of the ambient noise, and the additional sound emitted by the additional sound generation means, in addition to the ambient noise, relaxes the feelings of the worker present in the small space or activates the intellectual activities of the worker.

[0007] Furthermore, Patent Document 3 discloses a work environment adjustment system aimed at preventing a decrease in the concentration of workers when work involving conversation and work without conversation are mixed together.

[0008] This work environment adjustment system is a work environment adjustment system for adjusting the work environment of an office having multiple work areas, and comprises: an ambient lighting unit that illuminates the entire office; multiple task lighting units that illuminate each of the multiple work areas with a higher illuminance than the ambient lighting; and multiple sound masking units that emit sound towards each of the multiple work areas. Furthermore, this work environment adjustment system comprises: a selection unit that selects the current work area where work is currently being performed from among the multiple work areas; and a control unit that turns on the task lighting unit for the current work area and sounds the sound masking unit for the current work area or an adjacent area adjacent to the current work area. [Prior art documents] [Patent Documents]

[0009] [Patent Document 1] Japanese Patent Publication No. 2017-146517 [Patent Document 2] Japanese Patent Publication No. 2020-181539 [Patent Document 3] Japanese Patent Publication No. 2014-154483 [Overview of the Initiative] [Problems that the invention aims to solve]

[0010] Incidentally, in recent years, in open-plan offices and similar spaces where various people work individually, the sound environment can sometimes be bothersome due to the presence of other people's voices, making it difficult to have a conversation.

[0011] In particular, in open-plan offices that have adopted Activity-Based Working (ABW), if the reverberation time standards set by the Architectural Institute of Japan and ISO (International Organization for Standardization) are met, the sound-absorbing interior will create a quiet environment, but this can result in excessive sound transmission to the surroundings, making it difficult to speak.

[0012] In other words, while such a sound environment allows one to enjoy a quiet environment, it also presents the problem of making it difficult to communicate with others.

[0013] In response to this problem, the technologies disclosed in Patent Documents 1 to 3 do not take into consideration the ease of communication and therefore cannot necessarily improve the ease of communication in the target space.

[0014] This disclosure is made in view of the above circumstances and aims to provide an information processing device and an information processing program that can improve the ease of communication in the target space. [Means for solving the problem]

[0015] The information processing apparatus according to claim 1 comprises: a creation unit that performs emotion analysis on each segment of speech information representing speech uttered by a person, and adds a reverberation effect to utterances that give a positive impression, thereby creating background sound information that makes the content of the speech unrecognizable and has a positive effect on the emotions of at least some of the multiple people present in the target space; and a control unit that performs control to play the background sound indicated by the background sound information in the target space.

[0016] According to the information processing device of the present invention as described in claim 1, emotional analysis is performed on each segment of speech information representing speech uttered by a person, and a reverberation effect is added to utterances that give a positive impression, thereby creating background sound information that represents background sounds in which the content of the speech cannot be recognized and which have a positive effect on the emotions of at least some of the multiple people present in the target space. By controlling the playback of the background sounds indicated by the background sound information in the target space, the enhancement of positive emotional extension (the range of attention, cognition, and behavior broadened by positive emotions, improved receptivity: Estrada, Isen, &Young, 1997, broadening the scope of attention: Isen, 2002, etc. have been demonstrated) is brought about in multiple people who hear the background sounds, and the ease of communication in the target space can be improved.

[0017] The information processing apparatus according to claim 2 is the information processing apparatus according to claim 1, wherein the creation unit creates the background sound information by adding a reverberation effect to a random arrangement of utterances that give a positive impression.

[0018] According to the information processing device of the present invention as described in claim 2, background sound information is created by adding a reverberation effect to randomly placed utterances that give a positive impression, thereby making it impossible to recognize the content of the speech, and thus more effectively improving the ease of communication in the target space.

[0019] The information processing apparatus according to the present invention described in claim 3 is the information processing apparatus described in claim 1 or claim 2, further comprising an acquisition unit that acquires target space voice information indicating voice in the target space, and the control unit adjusts the playback volume of the background sound according to the volume of the voice indicated by the target space voice information.

[0020] According to the information processing apparatus according to the present invention described in claim 3, by acquiring target space voice information indicating voice in the target space and adjusting the playback volume of the background sound according to the volume of the voice indicated by the target space voice information, it is possible to more effectively improve the ease of communication in the target space.

[0021] The information processing apparatus according to the present invention described in claim 4 is the information processing apparatus described in claim 1, wherein the utterance that gives a positive impression is an utterance that gives at least one kind of impression of joy, gratitude, peace, interest, happiness, and hope.

[0022] According to the information processing apparatus according to the present invention described in claim 4, by making the utterance that gives a positive impression an utterance that gives at least one kind of impression of joy, gratitude, peace, interest, happiness, and hope, it is possible to give a positive impression corresponding to the applied type to the people present in the target space.

[0023] The information processing program according to the present invention described in claim 5 causes a computer to execute a process of performing emotional analysis for each segment of voice information indicating voice uttered by a person, adding a reverberation effect to an utterance that gives a positive impression, creating background sound information indicating background sound that cannot recognize the content of the voice and has a positive influence on the emotions of at least some of a plurality of people present in the target space, and controlling the playback of the background sound indicated by the background sound information in the target space.

[0024] According to the information processing program of the present invention described in claim 5, by performing sentiment analysis for each segment of voice information indicating the voice uttered by a person and adding a reverberation effect to the utterance that gives a positive impression, the content of the voice cannot be recognized, and background sound information indicating background sound that has a positive impact on the emotions of at least some of the plurality of people existing in the target space is created. By controlling the playback of the background sound indicated by the background sound information in the target space, an expansion function of positive emotions can be provided to the plurality of people who listened to the background sound, and the ease of communication in the target space can be improved.

Effect of the Invention

[0025] As described above, according to the present invention, the ease of communication in the target space can be improved.

Brief Description of the Drawings

[0026] [Figure 1] It is a block diagram showing an example of the hardware configuration of the information processing apparatus according to the embodiment. [Figure 2] It is a block diagram showing an example of the functional configuration of the information processing apparatus according to the embodiment. [Figure 3] It is a schematic diagram showing an example of the configuration of the voice information database according to the embodiment. [Figure 4] It is a flowchart showing an example of the first information processing according to the embodiment. [Figure 5] It is a schematic diagram for explaining the creation of background sound in the first information processing according to the embodiment. [Figure 6] It is a schematic diagram for explaining the creation of background sound in the first information processing according to the embodiment. [Figure 7] It is a schematic diagram for explaining the creation of background sound in the first information processing according to the embodiment. [Figure 8] It is a flowchart showing an example of the second information processing according to the embodiment. [Figure 9]These graphs illustrate the effect of background sound according to the embodiment. They represent time-series data of speech frequencies measured when the same person (actor) pronounces the same sentence in a way that conveys different emotions. The left graph shows the application of the emotion "normal," the center graph shows the application of the emotion "joy," and the right graph shows the application of the emotion "sadness." [Figure 10] This graph illustrates the effect of background sound according to the embodiment, and shows the results of a questionnaire survey investigating listeners' emotions when they heard the sound. [Modes for carrying out the invention]

[0027] Hereinafter, examples of embodiments for carrying out the present invention will be described in detail with reference to the drawings. In this embodiment, the present invention will be described in the case where it is applied to an open-plan office that has introduced Activity-Based Working (ABW). However, the application of the present invention is not limited to such an office, and it can be applied to any space that is basically quiet but where conversation is not restricted, such as a fixed-plan office, an indoor space other than an office, or an outdoor space.

[0028] First, the configuration of the information processing device 10 according to this embodiment will be described with reference to Figure 1. Figure 1 is a block diagram showing an example of the hardware configuration of the information processing device 10 according to this embodiment. Examples of information processing devices 10 include personal computers and server computers.

[0029] As shown in Figure 1, the information processing device 10 according to this embodiment includes a CPU (Central Processing Unit) 11, a memory 12 as a temporary storage area, a non-volatile storage unit 13, an input unit 14 such as a keyboard and mouse, a display unit 15 such as a liquid crystal display, a media read / write (R / W) device 16, and a communication interface (I / F) unit 18. The CPU 11, memory 12, storage unit 13, input unit 14, display unit 15, media read / write device 16, and communication I / F unit 18 are connected to each other via bus B. The media read / write device 16 reads information written on the recording medium 17 and writes information to the recording medium 17.

[0030] The storage unit 13 in this embodiment is implemented by an HDD (Hard Disk Drive), SSD (Solid State Drive), flash memory, etc. The storage unit 13, as a storage medium, stores a first information processing program 13A and a second information processing program 13B. The first information processing program 13A and the second information processing program 13B are stored in the storage unit 13 when the recording medium 17 on which the programs are written is set in the media read / write device 16, and the media read / write device 16 reads the programs from the recording medium 17. The CPU 11 reads the first information processing program 13A and the second information processing program 13B from the storage unit 13 as appropriate, loads them into the memory 12, and sequentially executes the processes of each program.

[0031] Furthermore, the memory unit 13 stores the audio information database 13C and the background sound information 13D. The audio information database 13C and the background sound information 13D will be described in detail later.

[0032] Meanwhile, the communication interface unit 18 is connected to a microphone (hereinafter also referred to as "microphone") 30 and a speaker 40, which are installed in the space targeted by the information processing device 10 (hereinafter referred to as "target space").

[0033] In this embodiment, the microphone 30 is positioned in a location in the target space where a relatively large number of people gather, and the speaker 40 is positioned in a location where the sound emitted can be clearly heard in that location. However, the embodiment is not limited to this configuration. For example, the microphone 30 may be positioned to collect sound from as wide an area as possible within the target space, and the speaker 40 may also be positioned in a location where the sound emitted can be clearly heard in that wide area.

[0034] Furthermore, although an omnidirectional microphone is used as microphone 30 in this embodiment, it is not limited to this, and other directional microphones such as unidirectional microphones or bidirectional microphones may be used as microphone 30. Also, although a dynamic speaker is used as speaker 40 in this embodiment, it is not limited to this, and other types of speakers such as electrostatic or magnetic speakers may be used as speaker 40.

[0035] Next, the functional configuration of the information processing device 10 according to this embodiment will be described with reference to Figure 2. Figure 2 is a block diagram showing an example of the functional configuration of the information processing device 10 according to this embodiment.

[0036] As shown in Figure 2, the information processing device 10 according to this embodiment includes a creation unit 11A, an acquisition unit 11B, and a control unit 11C. The CPU 11 of the information processing device 10 executes the first information processing program 13A and the second information processing program 13B, respectively, thereby enabling the creation unit 11A, the acquisition unit 11B, and the control unit 11C to function.

[0037] In this embodiment, the creation unit 11A performs emotion analysis on each segment of the speech information representing the voice spoken by a person, and adds a reverberation effect to utterances that give a positive impression. As a result, the creation unit 11A creates background sound information that represents background sounds in which the content of the voice cannot be recognized and which have a positive effect on the emotions of at least some of the multiple people present in the target space.

[0038] In particular, the creation unit 11A according to this embodiment creates background sound information by randomly arranging the above-mentioned utterances that give a positive impression and adding a reverberation effect. However, it is not limited to this form, and a form in which background sound information is created by not performing the random arrangement, that is, by simply adding a reverberation effect. In short, any processing can be applied to the creation of background sound information as long as it can create background sound information that is distorted to the extent that the content of the reproduced sound is unrecognizable, but can have a positive effect on people's emotions.

[0039] Furthermore, the control unit 11C according to this embodiment controls the playback of background sounds indicated by the background sound information created by the creation unit 11A in the target space.

[0040] In this embodiment, the acquisition unit 11B acquires target space audio information indicating the sound in the target space, and the control unit 11C adjusts the volume of background sound playback according to the volume of the sound indicated by the target space audio information acquired by the acquisition unit 11B. In this embodiment, as an adjustment for the volume of background sound playback, an adjustment is applied in which the volume of background sound playback increases linearly as the volume of the sound indicated by the target space audio information increases, but this is not the only option. For example, an adjustment may be applied in which the volume of background sound playback increases non-linearly as the volume of the sound indicated by the target space audio information increases.

[0041] Furthermore, in this embodiment, we have applied utterances that convey one of six types of positive impressions: joy, gratitude, peace, interest, happiness, and hope, but we are not limited to these. For example, we may apply utterances that convey only one of these impressions, or a combination of two to five impressions, as utterances that convey a positive impression. Also, the types of utterances that convey a positive impression are not limited to the above six types; pride, amusement, awe, love, etc., may also be included as types of utterances that convey a positive impression.

[0042] Here, with reference to Figures 9 and 10, the effects of background sound indicated by the background sound information created by the creation unit 11A according to this embodiment will be explained. Figure 9 is time-series data of speech frequencies measured when the same person (actor) spoke the same sentence in a way that sounded like different emotions. In Figure 9, the left graph applies the emotion of "normal," the center graph applies the emotion of "joy," and the right graph applies the emotion of "sadness." Figure 10 is a graph showing the results of a questionnaire survey investigating the emotions of listeners when they heard the sound. The horizontal axis represents the proportion of the presented emotion to the total presentation time (labeled "emotional characteristics of the sound" in Figure 10), and the vertical axis represents the proportion of the total number of listeners who showed the corresponding emotional response (labeled "emotional response of listeners" in Figure 10).

[0043] In the case of the emotion of "joy," shown in the center of Figure 9, the pitch (frequency) is higher and the high-frequency sound pressure is higher compared to the emotion of "normal," shown in the left of Figure 9. Conversely, in the case of the emotion of "sadness," shown in the right of Figure 9, the pitch is lower and the high-frequency sound pressure is lower compared to the emotion of "normal." From this, it can be seen that the sound characteristics change even for the same word when expressing different emotions.

[0044] On the other hand, Figure 10 is a graph showing the correlation between emotions (the upper graph shows listener "happy" - voice "happy", and the lower graph shows listener "happy" - voice "angry"), which were found to significantly influence the listener's emotional response through regression analysis (significance level 5%).

[0045] Figure 10 shows that the longer the audio of "happiness" is presented, the more people feel "happy." Although the correlation is not very high, listeners who hear the audio of "happiness" tend to feel "happy." From this, it can be inferred that by presenting background sounds with audio estimated to represent positive emotions such as "happiness," positive emotions can be induced in workers through the mirror effect, and a certain number of workers can perform their work under the influence of the enhanced functioning provided by positive emotions.

[0046] Next, the audio information database 13C according to this embodiment will be described with reference to Figure 3. Figure 3 is a schematic diagram showing an example of the configuration of the audio information database 13C according to this embodiment. The audio information database 13C is a database in which audio information is stored, which is used by the information processing device 10 according to this embodiment when creating background sound information.

[0047] As shown in Figure 3, the voice information database 13C according to this embodiment stores multiple voice information items S1, S2, ...

[0048] In this embodiment, the audio information used is recorded in a space different from the target space and obtained from multiple people regardless of gender or age group, but it is not limited to this. For example, the audio information could be recorded in the target space or restricted by gender or age group.

[0049] Furthermore, a microphone is used to record the above audio information, and after it is recorded on a storage medium such as a data recorder, it is registered in the audio information database 13C. The recording can be done in stereo or monaural, but considering that emotion analysis will be performed in detail later, the sampling interval should be set so that frequencies that affect the determination of emotion (up to about 10 kHz) can be analyzed.

[0050] Next, the operation of the information processing device 10 according to this embodiment will be explained with reference to Figures 4 to 8. First, the operation of the information processing device 10 when executing the first information processing will be explained with reference to Figures 4 to 7. When a user gives an instruction input via the input unit 14 to start the execution of the first information processing program 13A, the CPU 11 of the information processing device 10 executes the program 13A, thereby executing the first information processing shown in Figure 4. Here, in order to avoid confusion, we will explain the case where a sufficient amount of audio information for creating background sound information has already been registered in the audio information database 13C.

[0051] In step 100 of Figure 4, the CPU 11 reads all audio information from the audio information database 13C, and in step 102, the CPU 11 performs sentiment analysis on the read audio information as shown below.

[0052] In other words, in this embodiment, the impression that the acquired audio information gives to the listener is estimated (emotion analysis) using a machine learning model. For this emotion analysis, various existing machine learning models such as AI Suite from NTT Resonant Corporation and Emo Value Generator from Empath Co., Ltd. can be applied.

[0053] In the emotion analysis according to this embodiment, a score (hereinafter also referred to as the "emotion score") is output numerically for each emotion, such as anger, disgust, joy, surprise, happiness, calmness, and sadness, based on the audio indicated by the input audio information. For each emotion analysis, the analysis result is output for each segment (one word or one utterance (which varies depending on the machine learning model)).

[0054] In step 104, CPU 11 extracts positive speech using the analysis results output from the machine learning model, as shown below.

[0055] In other words, in this embodiment, the segment with the highest score for emotions classified as positive among the emotion scores output from the machine learning model is extracted as "positive voice".

[0056] Generally, positive emotions have received less attention until now, and therefore, there are fewer emotional items to evaluate compared to negative emotions. For this reason, comparing the total score of positive emotions with the total score of negative emotions makes it difficult to determine whether a segment gives a positive impression to the listener. Furthermore, the strength of positive emotions is unlikely to be directly proportional to that of negative emotions; a high score for positive emotions suggests that a positive impression is dominant. Therefore, when comparing the scores of each emotion, if the score of a particular emotion classified as positive is the highest, it can be concluded that the impression of the audio in that segment is positive.

[0057] For example, according to one emotion analysis method, the analysis outputs emotion scores from 1 to 10 for each of the four emotions: calmness, anger, joy, and sadness. In the example shown in Figure 5, the score for "joy," the only positive emotion, is 7.1, and since this is the highest score among the other emotion scores, this segment is judged to be positive audio.

[0058] Although the concept of positive emotions has attracted attention in recent years, opinions are divided within the Japanese Psychological Association and among other academic societies dealing with psychology, and there is no unified and clear definition. For this reason, here we define positive emotions as "joy, gratitude, peace, interest, hope, pride, amusement, inspiration, awe, and love," as proposed by Fredrickson, B. L. (University of North Carolina, Chapel Hill, North Carolina, USA), a leading figure in positive psychology, and emotions expressed using equivalent words (Reference: Barbara L. Fredrickson, The broaden-and-build theory of positive emotions, The Royal Society, https: / / www.ncbi.nlm.nih.gov / pmc / articles / PMC1693418 / pdf / 15347528.pdf).

[0059] In step 106, the CPU 11 randomly arranges the extracted positive voices to form a single audio data. More specifically, the extracted positive voices are numbered, and the positive voices are connected in the order of random numbers generated by various random number generators to create the aforementioned audio data.

[0060] In step 108, CPU 11 adds a reverberation effect to the created audio data and then superimposes it to prevent recognition of the audio text.

[0061] Specifically, first, the audio data created is convolved and integrated with an impulse response that includes reverberation components, based on equation (1) below. The impulse response used is either one measured in an actual space or one that is artificially created.

[0062]

number

[0063] In equation (1), y represents the output signal, h represents the impulse response, and x represents the input signal (the generated audio data).

[0064] The impulse response used is adjusted for the space where the sound is ultimately presented (in this embodiment, the target space), taking into account the room's volume and intended use, by adjusting the reverberation components and reflected sound structure. Specifically, as shown in the following table, the space in which the impulse response is measured is a large space such as a multipurpose hall or auditorium, which is 1.3 to 1.5 times larger than the space where the background sound is presented (the target space), and is intended for speeches. Furthermore, regarding the reverberation components, the frequency range of speech is generally said to be 80 to 10,000 Hz, and since low-frequency components are unnecessary in this background sound using speech data, components below approximately 80 Hz are cut off in the impulse response using a high-pass filter (HPF).

[0065] [Table 1]

[0066] In this embodiment, background sound information is created by superimposing the output signal after adding a reverberation effect. This superposition is performed by linearly adding the signals, for example, by using the method shown in Figure 6 (sliding and adding the signals) or the method shown in Figure 7 (dividing the signals, rearranging their order, and then adding them) to create the background sound information. In addition to these methods, other methods such as inverting the time axis of the signal or randomizing the division length may also be applied in combination.

[0067] In step 110, the CPU 11 stores (registers) the background sound information 13D obtained through the above processing in the storage unit 13, and then terminates this first information processing.

[0068] Next, with reference to Figure 8, the operation of the information processing device 10 when executing the second information processing will be explained. When a user gives an instruction input via the input unit 14 to start the execution of the second information processing program 13B, the CPU 11 of the information processing device 10 executes the program 13B, thereby executing the second information processing shown in Figure 8. Here, in order to avoid confusion, we will explain the case where background sound information 13D is registered in the storage unit 13.

[0069] In step 200 of Figure 8, the CPU 11 reads background sound information 13D from the storage unit 13, and in step 202, the CPU 11 acquires audio information (corresponding to the "target space audio information" described above) from the microphone 30.

[0070] In step 204, the CPU 11 derives a playback level for the background sound indicated by the acquired background sound information 13D, in accordance with the volume of the sound indicated by the acquired audio information, as described above, that is, increasing linearly as the volume of the sound indicated by the audio information increases. In step 206, the CPU 11 sets the playback level for the background sound indicated by the background sound information 13D to the derived playback level.

[0071] In step 208, the CPU 11 starts playing the background sound through the speaker 40 at the set playback level, and in step 210, the CPU 11 waits until a predetermined time (5 minutes in this embodiment) has elapsed.

[0072] In step 212, the CPU 11 determines whether a predetermined termination timing has arrived. If the determination is negative, it returns to step 202; if the determination is positive, it proceeds to step 214. In this embodiment, the termination timing is set to the timing when the user gives an instruction to stop the execution of the second information processing program 13B via the input unit 14, but it goes without saying that this is not the only termination timing.

[0073] In step 214, the CPU 11 stops playing the background sound through the speaker 40, and then terminates this second information processing.

[0074] Furthermore, in this second information processing, a predetermined upper limit level for the playback level may be set, and a process may be included to adjust the playback level so that the playback level of the background sound does not exceed that upper limit level.

[0075] As described above, according to this embodiment, emotion analysis is performed on each segment of speech information representing the voice spoken by a person, and a reverberation effect is added to utterances that give a positive impression. This creates background sound information that indicates background sounds that are unrecognizable in terms of the content of the voice and have a positive effect on the emotions of at least some of the multiple people present in the target space. The system then controls the playback of the background sounds indicated by the background sound information in the target space. Consequently, it is possible to enhance the ability to communicate in the target space by providing an extended positive emotional response to multiple people who hear the background sounds.

[0076] Furthermore, according to this embodiment, background sound information is created by randomly arranging utterances that give a positive impression and adding a reverberation effect. As a result, the content of the speech becomes more reliably unrecognizable, which can more effectively improve the ease of communication in the target space.

[0077] Furthermore, according to this embodiment, target space audio information indicating the sound in the target space is acquired, and the volume of background sound playback is adjusted according to the volume of the sound indicated by the target space audio information. Therefore, it is possible to more effectively improve the ease of communication in the target space.

[0078] Furthermore, according to this embodiment, utterances that give a positive impression are defined as utterances that give at least one of the following impressions: joy, gratitude, peace, interest, happiness, and hope. Therefore, it is possible to give a person present in the target space a positive impression according to the type to which it is applied.

[0079] In the above embodiment, we described a case where background sound is reproduced by applying one microphone and speaker to a single target space, but the invention is not limited to this. For example, background sound may be reproduced by providing multiple microphones and speakers, at least one of each, to a single target space.

[0080] Furthermore, the various numerical values ​​and calculation formulas applied in the above embodiments are merely examples, and it goes without saying that they can be modified without departing from the spirit of the present invention.

[0081] Furthermore, in the above embodiment, for example, the hardware structure of the processing unit that executes the creation unit 11A, the acquisition unit 11B, and the control unit 11C can be any of the following types of processors. As mentioned above, these types of processors include a CPU, which is a general-purpose processor that executes software (programs) and functions as a processing unit, as well as programmable logic devices (PLDs), such as FPGAs (Field-Programmable Gate Arrays), which are processors whose circuit configuration can be changed after manufacturing, and dedicated electrical circuits, such as ASICs (Application Specific Integrated Circuits), which are processors with circuit configurations specifically designed to execute specific processes.

[0082] The processing unit may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the processing unit may consist of a single processor.

[0083] Examples of configuring a processing unit with a single processor include, firstly, a configuration where one or more CPUs and software combine to form a single processor, as is common in client and server computers, and this processor functions as the processing unit. Secondly, a configuration using a processor that realizes the functions of the entire system, including the processing unit, on a single IC (Integrated Circuit) chip, as is common in System-on-a-Chip (SoC) systems. Thus, the processing unit is configured, in terms of hardware structure, using one or more of the above-mentioned types of processors.

[0084] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits, which are combinations of circuit elements such as semiconductor devices. [Explanation of Symbols]

[0085] 10 Information Processing Devices 11 CPU 11A Creation Section 11B Acquisition Department 11C Control Unit 12 memory 13 Storage section 13A First Information Processing Program 13B Second Information Processing Program 13C Voice Information Database 13D Background Sound Information 14 Input section 15 Display 16. Media reading / writing device 17 Recording media 18 Communication I / F Section 30 microphones 40 speakers

Claims

1. A creation unit performs emotion analysis on each segment of audio information representing human speech, and adds reverberation effects to utterances that give a positive impression, thereby creating background sound information that represents background sounds in which the content of the speech cannot be recognized, and which have a positive effect on the emotions of at least some of the multiple people present in the target space. A control unit that performs control to reproduce the background sound indicated by the background sound information in the target space, Equipped with an information processing device.

2. The creation unit creates the background sound information by randomly arranging the utterances that give a positive impression and adding a reverberation effect. The information processing apparatus according to claim 1.

3. The system further includes an acquisition unit that acquires target space audio information indicating the sound in the aforementioned target space. The control unit adjusts the volume of the background sound playback according to the volume of the sound indicated by the target spatial sound information. The information processing apparatus according to claim 1 or claim 2.

4. The aforementioned utterances that give a positive impression are those that give an impression of at least one of the following: joy, gratitude, comfort, interest, happiness, and hope. The information processing apparatus according to claim 1.

5. By performing emotion analysis on each segment of audio information representing human speech, and adding reverberation effects to utterances that give a positive impression, background sound information is created that represents background sounds that are unrecognizable in terms of their content, yet have a positive effect on the emotions of at least some of the multiple people present in the target space. Control is performed to play the background sound indicated by the background sound information in the target space. An information processing program that causes a computer to perform a task.