Information processing device, information processing system, and program
The information processing apparatus enhances user interactions in meetings by generating ambient sounds tailored to user activities and environmental conditions, addressing the lack of suitable sound output in existing technologies.
Patent Information
- Application Number
- JP2022007840
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-01-21
- Publication Date
- 2026-08-25
- Estimated Expiration
- 2042-01-21
AI Technical Summary
Existing technologies lack a method to output environmental sounds suitable for interactions between users in meetings or online conferences, particularly in response to user activities and environmental conditions.
An information processing apparatus that acquires user activity information, specifies speaking users, and generates ambient sounds based on user interactions and environmental data to output sounds suitable for the context.
Enables the output of ambient sounds that enhance user interactions by adapting to user activities and environmental conditions, improving communication and ambiance in meeting settings.
Smart Images

Figure 0007910306000001 
Figure 0007910306000002 
Figure 0007910306000003
Abstract
Description
Technical Field
[0004] , If the identified user continues to speak for a predetermined period of time or longer, that user is determined to be a user who performs a refrain. ,
[0006] , , , judgement , , and before , , , , , ,
[0005] , Used in the aforementioned refrain performance
[0001] The present invention relates to an information processing apparatus, an information processing system, and a program.
Background Art
[0002] [[ID=`12]]In a meeting where multiple users gather in a room or conduct conversations via a communication network, a technique of utilizing environmental sounds such as BGM (Background Music) to smoothly progress the meeting has been conventionally known.
[0003] <00000`12>Also, a technique for modifying environmental background noise based on the mood and / or behavior information of a user is known (see, for example, Patent Document 1).
Summary of the Invention
Problems to be Solved by the Invention
[0004] However, for example, in a case where an interaction occurs between users in a meeting, a technique for outputting environmental sounds suitable for the interaction between users is not known. Note that Patent Document 1 does not describe a technique for outputting environmental sounds suitable for the interaction between users.
[0005] An embodiment of the present invention aims to provide an information processing apparatus that outputs environmental sounds suitable for the interaction between users.
Means for Solving the Problems
[0006] In order to solve the above-described problems, claim 1 of the present application includes an acquisition unit that acquires activity information of a plurality of users who are having a conversation, a means for assigning melody information corresponding to each user among the plurality of users, a specifying unit that specifies a user who is speaking among the plurality of users, If the identified user continues to speak for a predetermined period of time or longer, that user is determined to be a user who performs a refrain. the activity information of the specified user and before rec judgement orded to the user Used in the aforementioned refrain performance melody i and The information processing device is characterized by having a generation means for generating sound data based on and a sound output control means for causing an output device to output ambient sound corresponding to the sound data. [Effects of the Invention]
[0007] According to an embodiment of the present invention, it is possible to output ambient sounds suitable for communication between users. [Brief explanation of the drawing]
[0008] [Figure 1] This is a diagram illustrating an example of an information processing system according to this embodiment. [Figure 2] This is a diagram illustrating an example of a conference room according to this embodiment. [Figure 3] This is a hardware configuration diagram of an example of a computer according to this embodiment. [Figure 4] This is a hardware configuration diagram of an example of a smartphone according to this embodiment. [Figure 5] This is a functional configuration diagram of an example of an information processing system according to this embodiment. [Figure 6] This is a diagram illustrating an example of reservation information. [Figure 7] This is a diagram illustrating an example of sound source information. [Figure 8] This is a diagram illustrating an example of sound number information. [Figure 9] This is a diagram illustrating an example of beat count information. [Figure 10] This is a diagram illustrating an example of timbre information. [Figure 11] This is a diagram illustrating an example of melody information. [Figure 12] This is an example flowchart illustrating the processing procedure of the information processing system according to this embodiment. [Figure 13] This is a flowchart illustrating an example of a process for generating sound data. [Figure 14] This is a diagram illustrating an example of sound number information. [Figure 15]This is a configuration diagram of an example of beat information. [Figure 16] This is a configuration diagram of an example of tone color information. [Figure 17] This is a flowchart of an example of a process for generating sound data. [Figure 18] This is a configuration diagram of an example of an information processing system according to this embodiment. [Figure 19] This is a functional configuration diagram of an example of an information processing system according to this embodiment. [Figure 20] This is a configuration diagram of an example of tone color information. [Figure 21] This is a flowchart showing an example of a processing procedure of the information processing system according to this embodiment. [Figure 22] This is a flowchart of an example of a process for generating sound data.
Mode for Carrying Out the Invention
[0009] Hereinafter, embodiments of the present invention will be described with reference to the drawings. In this embodiment, as an example where interactions occur between users, examples of multiple users in a conference room having a conversation and multiple users in an online conference having a conversation via a communication network will be described, but the examples are not limited to conferences. This embodiment can be applied to various scenes where interactions occur between users, such as seminars, meetings, discussions, conversations, presentations, or brainstorming.
[0010] [First Embodiment] [System Configuration] FIG. 1 is a configuration diagram of an example of an information processing system according to this embodiment. FIG. 2 is a diagram for explaining an example of a conference room according to this embodiment. The information processing system 1 in FIG. 1 has an information processing device 10, a video display device 12, a sensor device 14, a speaker 16, a camera 18, a microphone 20, and an information processing terminal 22 connected by wire or wirelessly so as to be communicable via a network N such as the Internet or a LAN.
[0011] The conference room is equipped with a video display device 12, a sensor device 14, a speaker 16, a camera 18, a microphone 20, and an information processing terminal 22. The conference room may also be equipped with temperature sensors, humidity sensors, illuminance sensors, etc., that acquire at least some of the environment-dependent information and notify the information processing device 10. Furthermore, although Figure 1 shows an example where the information processing device 10 is located outside the conference room, it may also be located inside the conference room.
[0012] For example, a user entering a conference room carries a tag that emits radio waves, such as a beacon. The conference room sensor device 14 receives the radio waves emitted from the user's tag as a signal to detect the user's location and notifies the information processing device 10. The sensor device 14 can be any positioning system sensor capable of outputting a signal to detect the user's location. The tag on the measurement target side can be a dedicated tag, a smartphone, or various BLE (Bluetooth Low Energy) sensors. The information processing device 10 detects the location information of each user in the conference room based on the signals for detecting the user's location notified from one or more sensor devices 14. Note that the tag described above is just one example of a transmitting device; any device that emits a signal to detect the user's location does not have to be in the form of a tag.
[0013] The information processing terminal 22 is a device operated by a user in the conference room. For example, the information processing terminal 22 may be a notebook PC (Personal Computer), mobile phone, smartphone, tablet device, game console, PDA (Personal Digital Assistant), digital camera, wearable PC, desktop PC, or a device specifically designed for the conference room. The information processing terminal 22 may be brought into the conference room by the user or may be provided in the conference room.
[0014] Furthermore, the information processing terminal 22 may also be the target of measurement by the positioning system. For example, the sensor device 14 in the conference room may receive radio waves emitted from the tag of the information processing terminal 22 and notify the information processing device 10. The sensor device 14 can notify the information processing device 10 of a signal to detect the location information of the user operating the information processing terminal 22 within the conference room, as shown in Figure 2, for example. The tag may be built into the information processing terminal 22 or in another form. In addition, the information processing terminal 22 may be equipped with a sensor to measure the user's heart rate and may notify the information processing device 10 of the measured user's heart rate.
[0015] The conference room camera 18 captures images of the conference room and transmits the captured video data as an output signal to the information processing device 10. Camera 18 can be, for example, a Kinect® video camera. A Kinect® video camera is an example of a video camera having a distance image sensor, an infrared sensor, and an array microphone. When using a video camera with a distance image sensor, an infrared sensor, and an array microphone, the user's movements and posture can be recognized.
[0016] The microphone 20 in the conference room converts the user's voice into an electrical signal. The microphone 20 transmits the electrical signal converted from the user's voice to the information processing device 10 as an output signal. Alternatively, the microphone on the information processing terminal 22 may be used instead of, or in conjunction with, the microphone 20 in the conference room.
[0017] The conference room speaker 16 converts electrical signals into physical signals and outputs sounds such as ambient noise. The speaker 16 outputs sounds such as ambient noise under the control of the information processing device 10. Alternatively, the speaker of the information processing terminal 22 may be used instead of, or in conjunction with, the conference room speaker 16. The conference room microphone 20 and the microphone of the information processing terminal 22 are examples of input devices. The conference room speaker 16 and the speaker of the information processing terminal 22 are examples of output devices.
[0018] An example of multiple video display devices 12 in a conference room is a projector, which can display images on a partitioning surface of the conference room, as shown in Figure 2, under the control of the information processing device 10. The partitioning surfaces of the conference room include, for example, the front wall, back wall, right wall, left wall, floor, and ceiling. Note that the video display device 12 is just one example of a display device that displays images; any display device that has at least the function of displaying images is applicable.
[0019] Note that the shape of the conference room in Figure 2 is just one example, and other shapes are also possible. Furthermore, as mentioned above, a conference room does not necessarily need to be partitioned on all surfaces such as walls, floors, and ceilings; it may be an open conference room with some surfaces not partitioned. Also, a conference room is just one example of a single space in which multiple users are present, and includes various spaces such as rooms for seminars or lectures, meeting spaces, and event spaces. Thus, the space described in this embodiment is a concept that includes places or rooms where multiple users are present.
[0020] The information processing device 10 outputs ambient sounds suitable for interactions between users in the conference room (conversations, meetings, etc.) based on the user's location information detected by the signal notified from the sensor device 14, the output signal from the camera 18, and the output signal from the microphone 20, as described below.
[0021] Note that the configuration of the information processing system 1 shown in Figure 1 is just one example. The information processing device 10 may be implemented using a single computer or multiple computers, or it may be implemented using a cloud service. Furthermore, the information processing device 10 may be, for example, an output device such as a projector, a display device with electronic whiteboard functionality, or a digital signage display, a HUD (Head Up Display) device, industrial machinery, an imaging device, a sound collection device, medical equipment, networked home appliances, an automobile (Connected Car), a notebook PC, a mobile phone, a smartphone, a tablet device, a game console, a PDA, a digital camera, a wearable PC, or a desktop PC.
[0022] <Hardware Configuration> "computer" The information processing device 10 is implemented by a computer 500 with the hardware configuration shown in Figure 3, for example. Similarly, if the information processing terminal 22 is a PC, it is implemented by a computer 500 with the hardware configuration shown in Figure 3, for example.
[0023] Figure 3 is a hardware configuration diagram of an example of a computer according to this embodiment. As shown in Figure 3, the computer 500 includes a CPU (Central Processing Unit) 501, ROM (Read Only Memory) 502, RAM (Random Access Memory) 503, HD 504, HDD (Hard Disk Drive) controller 505, display 506, external device connection I / F (Interface) 508, network I / F 509, data bus 510, keyboard 511, pointing device 512, DVD-RW (Digital Versatile Disk Rewritable) drive 514, and media I / F 516.
[0024] Of these components, the CPU 501 controls the overall operation of the computer 500. The ROM 502 stores programs used to drive the CPU 501, such as the IPL. The RAM 503 is used as the work area for the CPU 501. The HD 504 stores various data, such as programs. The HDD controller 505 controls the reading or writing of various data to the HD 504 according to the control of the CPU 501.
[0025] The display 506 displays various information such as cursors, menus, windows, characters, or images. The external device connection interface 508 is an interface for connecting various external devices. In this case, external devices include, for example, USB (Universal Serial Bus) memory and printers. The network interface 509 is an interface for data communication using network N. The data bus 510 is an address bus and data bus for electrically connecting various components such as the CPU 501.
[0026] The keyboard 511 is a type of input means equipped with multiple keys for inputting characters, numbers, and various instructions. The pointing device 512 is a type of input means for selecting and executing various instructions, selecting processing targets, and moving the cursor. The DVD-RW drive 514 controls the reading or writing of various data to the DVD-RW 513, which is an example of a removable recording medium. Note that it is not limited to DVD-RW, but may also be DVD-R, etc. The media interface 516 controls the reading or writing (storage) of data to the recording medium 515, such as flash memory.
[0027] Smartphone The information processing terminal 22 may be implemented, for example, by a smartphone 600 with the hardware configuration shown in Figure 4. Even if the information processing terminal 22 is a notebook PC, mobile phone, smartphone, tablet device, game console, PDA, digital camera, wearable PC, desktop PC, or dedicated conference room device, it may be implemented with a hardware configuration similar to that shown in Figure 4. Furthermore, while some of the hardware configurations shown in Figure 4 may be omitted, some of the configurations may be added to the hardware configuration shown in Figure 4.
[0028] Figure 4 is a hardware configuration diagram of an example of a smartphone according to this embodiment. As shown in Figure 4, the smartphone 600 includes a CPU 601, ROM 602, RAM 603, EEPROM 604, CMOS sensor 605, image sensor I / F 606, acceleration / direction sensor 607, media I / F 609, and GPS receiver 611.
[0029] Of these components, the CPU 601 controls the overall operation of the smartphone 600. The ROM 602 stores programs used to drive the CPU 601, such as the CPU 601 and the IPL. The RAM 603 is used as the work area for the CPU 601. The EEPROM 604 reads or writes various data, such as smartphone programs, according to the control of the CPU 601.
[0030] The CMOS (Complementary Metal Oxide Semiconductor) sensor 605 is a type of built-in imaging means that captures an image of a subject (mainly a self-portrait) and obtains image data according to the control of the CPU 601. Note that an imaging means such as a CCD (Charge Coupled Device) sensor may be used instead of the CMOS sensor 605. The image sensor I / F 606 is a circuit that controls the driving of the CMOS sensor 605. The acceleration / direction sensor 607 is a type of sensor that detects the Earth's magnetic field, such as an electronic magnetic compass, gyrocompass, or acceleration sensor.
[0031] The media interface 609 controls the reading or writing (storage) of data to or from the recording medium 608, such as flash memory. The GPS receiver 611 receives GPS signals from GPS satellites.
[0032] The smartphone 600 also includes a long-range communication circuit 612, a CMOS sensor 613, an image sensor interface 614, a microphone 615, a speaker 616, an audio input / output interface 617, a display 618, an external device connection interface 619, a short-range communication circuit 620, an antenna 620a for the short-range communication circuit 620, and a touch panel 621.
[0033] Of these, the long-distance communication circuit 612 is a circuit that communicates with other devices via the network N. The CMOS sensor 613 is a type of built-in imaging means that captures an image of a subject and obtains image data according to the control of the CPU 601. The image sensor interface 614 is a circuit that controls the driving of the CMOS sensor 613. The microphone 615 is a built-in circuit that converts voice and sound into electrical signals. The speaker 616 is a built-in circuit that converts electrical signals into physical vibrations to produce sounds such as ambient sounds, music, or speech.
[0034] The audio input / output interface 617 is a circuit that processes the input and output of audio signals between the microphone 615 and the speaker 616 according to the control of the CPU 601. The display 618 is a type of display means, such as a liquid crystal or organic EL (electroluminescence), that displays images of the subject or various icons.
[0035] The external device connection interface 619 is an interface for connecting various external devices. The near-field communication circuit 620 is a communication circuit such as NFC (Near Field Communication) or Bluetooth (registered trademark). The touch panel 621 is a type of input means that allows the user to operate the smartphone 600 by pressing the display 618.
[0036] Furthermore, the smartphone 600 is equipped with a bus line 610. The bus line 610 is an address bus, data bus, etc., for electrically connecting the various components, such as the CPU 601 shown in Figure 4.
[0037] <Functional Configuration> The information processing system 1 according to this embodiment is implemented by a functional configuration such as that shown in Figure 5. Figure 5 is a functional configuration diagram of an example of the information processing system according to this embodiment. In the functional configuration of Figure 5, components that are not necessary for the explanation of this embodiment have been appropriately omitted.
[0038] The information processing device 10 in Figure 5 has a configuration comprising a video display control unit 30, an acquisition unit 32, a generation unit 34, a sound output control unit 36, an authentication processing unit 38, a user detection unit 40, a communication unit 42, and a storage unit 50. The storage unit 50 stores reservation information 52, sound source information 54, number of sounds information 56, beat count information 58, timbre information 60, and melody information 62, which will be described later.
[0039] The sensor device 14 has an output signal transmission unit 70. The speaker 16 has an output unit 110. The camera 18 has an output signal transmission unit 80. The microphone 20 has an output signal transmission unit 90. The information processing terminal 22 has an output signal transmission unit 100 and an output unit 102.
[0040] The output signal transmission unit 70 of the sensor device 14 transmits a signal to the information processing device 10 as an output signal for detecting multiple users in the conference room. The output signal transmission unit 80 of the camera 18 transmits the image captured inside the conference room as an output signal to the information processing device 10. The output signal transmission unit 90 of the microphone 20 transmits an electrical signal converted from the voices of multiple users in the conference room as an output signal to the information processing device 10.
[0041] Furthermore, the output signal transmission unit 100 of the information processing terminal 22 transmits an electrical signal converted by the microphone 615 from the voice of the user operating the information processing terminal 22 as an output signal to the information processing device 10. The output unit 102 of the information processing terminal 22 outputs sounds such as ambient sounds according to the sound data received from the information processing device 10. The output unit 110 of the speaker 16 outputs sounds such as ambient sounds according to the sound data received from the information processing device 10.
[0042] Note that the output signal transmission units 70, 80, 90, and 100 shown in Figure 5 are examples of input devices. Output units 102 and 110 are examples of output devices.
[0043] The communication unit 42 of the information processing device 10 receives a signal from the output signal transmission unit 70 of the sensor device 14 to detect the user's location information. The communication unit 42 receives the image captured inside the conference room from the output signal transmission unit 80 of the camera 18 as an output signal. The communication unit 42 receives an electrical signal converted from the voices of multiple users in the conference room from the output signal transmission unit 90 of the microphone 20 as an output signal. The communication unit 42 receives an electrical signal converted by the microphone 615 from the voice of the user operating the information processing terminal 22 as an output signal from the output signal transmission unit 100 of the information processing terminal 22. The communication unit 42 also receives operation signals received by the information processing terminal 22 from the user.
[0044] The user detection unit 40 detects users inside the conference room from signals for detecting the user's location information received from the sensor device 14. The user detection unit 40 also detects the location information of users inside the conference room. The authentication processing unit 38 performs authentication processing for users inside the conference room. The video display control unit 30 controls the video displayed by the video display device 12.
[0045] The acquisition unit 32 acquires activity information of users in the conference room. Examples of user activity information acquired by the acquisition unit 32 include the amount of speech spoken by multiple users in the conference room. Another example of user activity information acquired by the acquisition unit 32 is the frequency of speaker changes among multiple users in the conference room. Yet another example of user activity information acquired by the acquisition unit 32 is information on users who have been speaking continuously for a predetermined period of time or longer in the conference room. Speech volume, speaker change frequency, and information on users who have been speaking continuously can be measured from the output signals of microphone 20 or microphone 615.
[0046] Furthermore, the acquisition unit 32 acquires environment-dependent information inside or outside the conference room. Examples of environment-dependent information acquired by the acquisition unit 32 include weather, temperature, humidity, illuminance, equipment operating noise, noise level, or time of day. For example, the acquisition unit 32 may acquire environment-dependent information such as weather and temperature, which are publicly available on the internet, by sending a request to an external server that provides environment-dependent information. If an API (Application Programming Interface) is provided when acquiring information from an external server, the acquisition unit 32 may use the API to acquire the information. The acquisition unit 32 may also acquire information such as the heart rate of a user inside the conference room as an example of user activity information.
[0047] The generation unit 34 generates sound data as described later, based on activity information of multiple users in the conference room and environment-dependent information of the inside or outside of the conference room. The generation unit 34 may also generate sound data as described later, based on activity information of multiple users in the conference room without using environment-dependent information. The sound output control unit 36 controls the output of ambient sound corresponding to the generated sound data to the output unit 102 of the information processing terminal 22 or the output unit 110 of the speaker 16.
[0048] The memory unit 50 stores reservation information 52, sound source information 54, number of notes information 56, beat information 58, timbre information 60, and melody information 62 in a table format, for example, as shown in Figures 6 to 11.
[0049] Note that the reservation information 52, sound source information 54, note count information 56, beat count information 58, timbre information 60, and melody information 62 do not necessarily have to be in the table format shown in Figures 6 to 11; it is sufficient if similar information can be stored and managed.
[0050] Figure 6 is a diagram illustrating an example of reservation information. The reservation information in Figure 6 includes the following items: Reservation ID, Room ID, Reservation Time, and Participating Users. The Reservation ID is an example of identification information that identifies the reservation information. The Room ID is an example of identification information for the meeting room reserved by the reservation information. The Reservation Time is an example of date and time information for the meeting reserved by the reservation information. The Participating Users is an example of participant information for the meeting reserved by the reservation information.
[0051] For example, in the example in Figure 6, the first record contains information about a meeting scheduled to be held in room ID "room001" with a reservation time of "2022 / 01 / 12 13:00~14:00" and participating users "User 1, User 2, User 3, User 4".
[0052] Figure 7 is a diagram illustrating an example of sound source information. The sound source information in Figure 7 includes the following items: reservation ID, multiple time slots A-D, and assigned sound sources for the multiple time slots A-D. The reservation ID is an example of identification information that identifies reservation information. The multiple time slots A-D are an example of time slot information that divides the meeting reservation time into four parts. For example, in the example in Figure 7, the time slot of the first record is divided into time slot A "13:00-13:10", time slot B "13:10-13:30", time slot C "13:30-13:50", and time slot D "13:50-14:00". For example, Figure 7 is an example where the reservation time is divided into time slot A "17%", time slot B "33%", time slot C "33%", and time slot D "17%". Figure 7 is just one example and does not limit the number and percentage of multiple time slots.
[0053] Additionally, sound source sets are assigned to multiple time slots A through D. These sound source sets can be automatically assigned to time slots A through D, or they can be assigned manually by a meeting administrator or similar person.
[0054] Figure 8 is a diagram illustrating an example of sound count information. The sound count information in Figure 8 includes the following items: sound count class, utterance volume, and sound count. The sound count class is an example of identification information for classification. The utterance volume is an example of information representing the frequency of speech by users in the conference room. In Figure 8, as an example, the utterance volume is represented by the number of seconds during a predetermined period (e.g., the last 60 seconds) in which at least one user in the conference room was speaking (the state in which there is a speaker in the conference room). The sound count represents the number of sounds used in combination in the ambient sound.
[0055] According to the sound count information in Figure 8, the longer the speaker is in the conference room, the higher the sound count class becomes, resulting in a greater number of sounds being used as ambient sounds. According to the sound count information in Figure 8, the shorter the speaker is in the conference room, the lower the sound count class becomes, resulting in a smaller number of sounds being used as ambient sounds.
[0056] Figure 9 is a diagram illustrating an example of beat count information. The beat count information in Figure 9 includes the following items: beat count class, speaker change frequency, and beat count. The beat count class is an example of identification information for classification. The speaker change frequency is an example of information that represents the level of conversation activity among multiple users in a conference room, expressed by the frequency of speaker changes. In Figure 9, as an example, the speaker change frequency is represented by the number of times the speaker changed within a predetermined time period (e.g., the last 60 seconds). The beat count represents the beat used in ambient sound.
[0057] According to the beat count information in Figure 9, the higher the frequency of speaker changes, the higher the beat count class, and therefore the higher the beat count of ambient sounds. According to the beat count information in Figure 9, the lower the frequency of speaker changes, the lower the beat count class, and therefore the lower the beat count of ambient sounds.
[0058] Figure 10 is a diagram illustrating an example of timbre information. The timbre information in Figure 10 includes the following items: timbre class, weather information, and timbre. The timbre class is an example of identification information for classification. The weather information is an example of environment-dependent information outside the conference room. In the example in Figure 10, the weather outside the conference room is represented as, for example, sunny, cloudy, or rainy. The timbre represents the timbre used for ambient sound.
[0059] According to the timbre information in Figure 10, the timbre used for ambient sounds can be changed depending on the weather outside the conference room. Note that the timbre may be changed according to the timbre information in Figure 10 for all of the time periods A to D shown in Figure 7, or it may be changed according to the timbre information in Figure 10 for only a portion of the time periods A to D (for example, time periods A and D).
[0060] Figure 11 is a diagram illustrating an example of melody information. The melody information in Figure 11 includes the items "participating users" and "melody." Participating users are an example of information representing the participants of a meeting that has been reserved according to the reservation information. Melody is an example of information indicating a melody used for refrain (repeated) playback, which is an ambient sound assigned to each participating user.
[0061] According to the melody information in Figure 11, if a specific user participating in a meeting speaks continuously for a predetermined period of time or longer, the melody assigned to that user can be used as ambient sound.
[0062] <Processing> The information processing system 1 according to this embodiment outputs ambient sound to the conference room using a procedure such as that shown in Figure 12. Figure 12 is a flowchart illustrating an example of the processing procedure of the information processing system according to this embodiment.
[0063] In step S100, in the information processing system 1 according to this embodiment, a user such as the meeting organizer performs preliminary preparations. These preliminary preparations include registering the reservation information shown in Figure 6, setting the sound source information shown in Figure 7, setting the number of notes information shown in Figure 8, setting the number of beats information shown in Figure 9, setting the timbre information shown in Figure 10, and setting the melody information shown in Figure 11, which are stored in the storage unit 50 of the information processing device 10. The user accesses the information processing device 10 using the information processing terminal 22, and the communication unit 42 of the information processing device 10 receives operation information from the information processing terminal 22, allowing the user to perform any of the following operations: modify, add, or delete the various information stored in the storage unit 50. Note that the setting of the sound source information shown in Figure 7, the setting of the number of notes information shown in Figure 8, the setting of the number of beats information shown in Figure 9, the setting of the timbre information shown in Figure 10, and the setting of the melody information shown in Figure 11 may be automatically set by the information processing device 10 based on the registration of the reservation information shown in Figure 6.
[0064] In step S102, the information processing system 1 according to this embodiment determines that a meeting has started according to the reservation information in Figure 6. The determination of the start of a meeting may be made by the communication unit 42 of the information processing system 10 receiving information based on operation input entered by a user, such as the meeting organizer, into the information processing terminal 22, and making a determination based on the received information, or by detecting users in the meeting room and their movements. Alternatively, the information processing system 10 may make a determination based on the output signal of microphone 20 or microphone 615 corresponding to the voice produced by the user. In this example, the determination is made when a meeting has started, but it may also be made when user interaction such as a seminar, meeting, discussion, conversation, presentation, or brainstorming has started.
[0065] In step S104, the information processing system 1 according to this embodiment acquires activity information of multiple users in the conference room using the acquisition unit 32. The user activity information acquired by the acquisition unit 32 in step S104 includes, for example, the amount of speech of multiple users in the conference room, the frequency of speaker changes, and information on users who continue to speak in the conference room for a predetermined amount of time or longer.
[0066] In step S106, the information processing system 1 according to this embodiment acquires environment-dependent information, either inside or outside the conference room, via the acquisition unit 32. Here, it is explained that the acquisition unit 32 acquires weather information outside the conference room from an external server using an API, but weather information may be acquired by other methods.
[0067] In step S108, the information processing system 1 according to this embodiment generates sound data in a procedure such as that shown in Figure 13, based on the activity information of multiple users in the conference room acquired in step S104 and the weather information acquired in step S106.
[0068] Figure 13 is a flowchart of an example of the process for generating sound data. In step S200, the generation unit 34 determines the set of sound sources to be assigned to time slots A to D of the meeting reservation time, based on the reservation information in Figure 6 and the sound source information in Figure 7.
[0069] In step S202, the generation unit 34 determines the number of sounds to be layered in the ambient sound based on the number of sounds spoken by multiple users in the conference room, based on the sound count information in Figure 8. In step S204, the generation unit 34 determines the number of beats to be used in the ambient sound based on the frequency of speaker changes among multiple users in the conference room, based on the beat count information in Figure 9. In step S206, the generation unit 34 determines the timbre to be used in the ambient sound based on the weather information outside the conference room, based on the timbre information in Figure 10.
[0070] Furthermore, in step S208, the generation unit 34 determines that a specific user participating in the meeting is the user to be played a refrain if that user has been speaking continuously for a predetermined period of time or longer. Based on the melody information in Figure 11, the generation unit 34 determines the melody to be assigned to the user to be played a refrain.
[0071] In step S210, the generation unit 34 generates sound data based on the determined sound source set, number of notes, number of beats, timbre, and melody. The process of generating sound data may be a composition process, or it may be a process of selecting sound data corresponding to a combination of sound source set, number of notes, number of beats, timbre, and melody.
[0072] Returning to step S110 in Figure 12, the sound output control unit 36 controls the output unit 102 of the information processing terminal 22 or the output unit 110 of the speaker 16 to output ambient sound corresponding to the sound data generated in step S108. Ambient sound includes sounds such as music, voice, and white noise. If multiple speakers 16 can output individual sounds, individual ambient sounds may be output for each of the multiple users in the conference room.
[0073] As described above, the information processing system 1 according to this embodiment can output ambient sounds to a conference room containing multiple users, which change according to the situation of the users' conversations. By setting the sound source set, number of sounds, number of beats, timbre, and melody used to generate the sound data in step S108 so that ambient sounds suitable for the situation of the users in the conference room are output, the information processing system 1 according to this embodiment can output ambient sounds suitable for the interaction between users in the conference room.
[0074] For example, the information processing system 1 according to this embodiment can output ambient sounds suitable for tense users to the conference room, assuming that the greater the volume of speech and the frequency of speaker changes among the multiple users in the conference room, the higher the tension level of the conference participants. The information processing system 1 according to this embodiment can output ambient sounds suitable for relaxed users to the conference room, assuming that the less the volume of speech and the frequency of speaker changes among the multiple users in the conference room, the higher the relaxation level of the conference participants.
[0075] Steps S104 to S112 are repeated until the meeting ends. When the meeting ends, the process proceeds to step S114, and the sound output control unit 36 terminates the output of ambient sound from the output unit 102 of the information processing terminal 22 or the output unit 110 of the speaker 16.
[0076] [Second Embodiment] The first embodiment described the amount of speech spoken by multiple users in a conference room and the frequency of changes in the speaker among multiple users in a conference room as examples of user activity information. The second embodiment provides examples of the amount of change in posture among multiple users in a conference room and the frequency of changes in posture among multiple users in a conference room as examples of user activity information. User activity information may also include the amount of speech spoken by multiple users in a conference room, the frequency of changes in the speaker among multiple users in a conference room, the amount of change in posture among multiple users in a conference room, and the frequency of changes in posture among multiple users in a conference room.
[0077] The change in posture of multiple users in a conference room can be measured from the change in the volume of the user's posture bounding box, which is recognized by image processing of the video data captured by camera 18. For example, the posture bounding box can be determined from the boundary or surrounding box of the 3D point cloud obtained from the Kinect® video camera, which represents the user's position.
[0078] The frequency of posture changes among multiple users in a conference room can be measured by the number of times the volume of the user's posture bounding box, recognized through image processing of video data captured by camera 18, changes by a predetermined percentage or more.
[0079] In the information processing system 1 according to the second embodiment, the number of notes information 56, the number of beats information 58, and the timbre information 60 are configured as shown in Figures 14 to 16, for example. Figure 14 is a configuration diagram of an example of the number of notes information. Figure 15 is a configuration diagram of an example of the number of beats information. Figure 16 is a configuration diagram of an example of the timbre information.
[0080] The sound count information in Figure 14 includes the following items: sound count class, posture information, and sound count. The sound count class is an example of identification information for classification. The posture information is an example of information representing the amount of change in posture of multiple users in a conference room. In Figure 14, as an example, posture information is represented by the change in the volume of the posture bounding box of multiple users in the conference room over the past 60 seconds. The sound count represents the number of sounds used in combination in the ambient sound.
[0081] According to the sound count information in Figure 14, the greater the change in posture of multiple users in the conference room, the higher the sound count class, and therefore the more sounds that can be layered as ambient sounds. According to the sound count information in Figure 14, the smaller the change in posture of multiple users in the conference room, the lower the sound count class, and therefore the fewer sounds that can be layered as ambient sounds.
[0082] The beat count information in Figure 15 includes the following items: beat count class, posture change frequency, and beat count. The beat count class is an example of identification information for classification. The posture change frequency is an example of information representing the frequency of posture changes of multiple users in a conference room. In Figure 15, as an example, the frequency of posture changes of multiple users in a conference room is represented by the number of times the volume of the posture bounding box of multiple users in the conference room changed by more than a predetermined percentage in the most recent 60 seconds. The beat count represents the beat used for ambient sound.
[0083] According to the beat count information in Figure 15, the more frequently multiple users in the conference room change their posture, the higher the beat class, and therefore the higher the beat count of ambient noise. According to the beat count information in Figure 15, the less frequently multiple users in the conference room change their posture, the lower the beat class, and therefore the lower the beat count of ambient noise.
[0084] The timbre information in Figure 16 includes the following items: timbre class, temperature information, and timbre. The timbre class is an example of identification information for classification. The temperature information is an example of environment-dependent information, whether it is outside or inside the conference room. Figure 16 shows an example of information that represents the temperature outside or inside the conference room as low, normal, and high. The timbre represents the timbre used for ambient sound.
[0085] According to the timbre information in Figure 16, the timbre used for ambient sounds can be changed depending on the temperature outside or inside the conference room.
[0086] The information processing system 1 according to the second embodiment outputs ambient sound to the conference room in the procedure shown in Figure 12 above. In step S100, the information processing system 1 according to the second embodiment has a user, such as the conference organizer, perform preliminary preparations. These preliminary preparations include registering the reservation information shown in Figure 6, setting the sound source information shown in Figure 7, setting the number of sounds information shown in Figure 14, setting the number of beats information shown in Figure 15, setting the tone information shown in Figure 16, and setting the melody information shown in Figure 11, which are stored in the storage unit 50 of the information processing device 10. The user accesses the information processing device 10 using the information processing terminal 22, and the communication unit 42 of the information processing device 10 receives operation information from the information processing terminal 22, allowing the user to perform any of the following processes: modifying, adding, or deleting the various information stored in the storage unit 50.
[0087] Note that the settings for the sound source information in Figure 7, the number of notes information in Figure 14, the number of beats information in Figure 15, the timbre information in Figure 16, and the melody information in Figure 11 may be automatically set by the information processing device 10 based on the registration of reservation information in Figure 6.
[0088] In step S102, the information processing system 1 according to the second embodiment determines that a meeting has started according to the reservation information in Figure 6. The determination of the start of a meeting may also be made by the communication unit 42 of the information processing system 10 receiving information based on operation input entered by a user, such as the meeting organizer, into the information processing terminal 22, and making a determination based on the received information. The determination of the start of a meeting may also be made by detecting users in the meeting room and their movements. Alternatively, the information processing system 10 may make a determination based on the output signal of microphone 20 or microphone 615 corresponding to the voice produced by the user. In this example, the determination is made as to the start of a meeting, but it may also be made as to the start of user interaction, such as a seminar, meeting, discussion, conversation, presentation, or brainstorming. In step S104, the information processing system 1 according to the second embodiment has the acquisition unit 32 acquire activity information of multiple users in the meeting room. The user activity information acquired by the acquisition unit 32 in step S104 of the second embodiment includes, for example, the amount of change in posture of multiple users in the conference room, the frequency of posture changes of multiple users in the conference room, and information on users who continue to talk in the conference room for a predetermined amount of time or longer.
[0089] In step S106, the information processing system 1 according to the second embodiment acquires environment-dependent information, either inside or outside the conference room, by the acquisition unit 32. Here, it is explained that the acquisition unit 32 acquires temperature information, either outside or inside the conference room.
[0090] In step S108, the information processing system 1 according to the second embodiment generates sound data in a procedure such as that shown in Figure 17, based on the activity information of multiple users in the conference room acquired in step S104 and the temperature information acquired in step S106.
[0091] Figure 17 is a flowchart of an example of the process for generating sound data. In step S300, the generation unit 34 determines the set of sound sources to be assigned to time slots A to D of the meeting reservation time, based on the reservation information in Figure 6 and the sound source information in Figure 7.
[0092] In step S302, the generation unit 34 determines the number of sounds to be layered in the ambient sound based on the number of sounds information in Figure 14 and the posture information of multiple users in the conference room. In step S304, the generation unit 34 determines the number of beats to be used in the ambient sound based on the frequency of posture changes of multiple users in the conference room based on the beat information in Figure 15. In step S306, the generation unit 34 determines the timbre to be used in the ambient sound based on the temperature information outside or inside the conference room based on the timbre information in Figure 16.
[0093] Furthermore, in step S308, the generation unit 34 determines that a specific user participating in the meeting is the user to be played a refrain if that user has been speaking continuously for a predetermined period of time or longer. Based on the melody information in Figure 11, the generation unit 34 determines the melody to be assigned to the user to be played a refrain.
[0094] In step S310, the generation unit 34 generates sound data based on the determined sound source set, number of notes, number of beats, timbre, and melody. Returning to step S110 in Figure 12, the sound output control unit 36 controls the output of ambient sound corresponding to the sound data generated in step S108 to the output unit 102 of the information processing terminal 22 or the output unit 110 of the speaker 16.
[0095] Thus, the information processing system 1 according to the second embodiment can output ambient sounds to a conference room containing multiple users, which change according to the changes in the postures of the multiple users.
[0096] By setting the sound source set, number of sounds, number of beats, timbre, and melody used to generate the sound data in step S108 so that ambient sounds suitable for the changes in posture of multiple users in the conference room are output, the information processing system 1 according to the second embodiment can output ambient sounds suitable for the interaction between users in the conference room. For example, the information processing system 1 according to the second embodiment can assume that the greater the change in posture of multiple users in the conference room, the higher the tension level of the users participating in the meeting, and output ambient sounds suitable for multiple users with a high level of tension to the conference room.
[0097] Steps S104 to S112 are repeated until the meeting ends. When the meeting ends, the process proceeds to step S114, and the sound output control unit 36 terminates the output of ambient sound from the output unit 102 of the information processing terminal 22 or the output unit 110 of the speaker 16.
[0098] [Third Embodiment] The information processing system 1 according to the first embodiment shows an example of multiple users having a conversation in a conference room. The information processing system 2 according to the third embodiment describes an example of multiple users having a conversation during an online meeting.
[0099] <System Configuration> Figure 18 is a diagram illustrating an example of an information processing system according to this embodiment. In the information processing system 2 of Figure 18, the information processing device 10 and the information processing terminal 22 are connected via wired or wireless connections so that they can communicate via a network N such as the Internet or a LAN.
[0100] The information processing terminal 22 is a device used by multiple users to participate in an online meeting. For example, the information processing terminal 22 may be a notebook PC, mobile phone, smartphone, tablet device, game console, PDA, digital camera, wearable PC, desktop PC, or a dedicated device for the meeting room.
[0101] The microphone of the information processing terminal 22 converts the user's voice into an electrical signal. The microphone of the information processing terminal 22 transmits the electrical signal converted from the user's voice to the information processing device 10 as an output signal. The speaker of the information processing terminal 22 converts the electrical signal into a physical signal and outputs sounds such as ambient noise. The speaker of the information processing terminal 22 outputs sounds such as ambient noise under the control of the information processing device 10. The microphone of the information processing terminal 22 is an example of an input device. The speaker of the information processing terminal 22 is an example of an output device.
[0102] The information processing device 10 outputs ambient sounds suitable for user interaction (conversation, meeting, etc.) during an online meeting, based on the output signal from the microphone of the information processing terminal 22, as described below.
[0103] Note that the configuration of the information processing system 2 shown in Figure 18 is just one example. The information processing device 10 may be implemented using a single computer or multiple computers, or it may be implemented using a cloud service.
[0104] The information processing device 10 may be a projector, a display device with electronic whiteboard functionality, an output device such as a digital signage display, a HUD device, industrial machinery, an imaging device, a sound collection device, medical equipment, networked home appliances, automobiles, notebook PCs, mobile phones, smartphones, tablet devices, game consoles, PDAs, digital cameras, wearable PCs, or desktop PCs, etc.
[0105] The information processing system 2 according to this embodiment is realized by a functional configuration such as that shown in Figure 19. Figure 19 is a functional configuration diagram of an example of the information processing system according to this embodiment. In the functional configuration of Figure 19, configurations that are not necessary for the description of the third embodiment have been appropriately omitted.
[0106] The information processing device 10 shown in Figure 19 has a configuration comprising a video display control unit 30, an acquisition unit 32, a generation unit 34, a sound output control unit 36, an authentication processing unit 38, a communication unit 42, and a storage unit 50. The storage unit 50 stores reservation information 52, sound source information 54, number of notes information 56, beat count information 58, timbre information 60, and melody information.
[0107] The output signal transmission unit 100 of the information processing terminal 22 transmits an electrical signal converted by the microphone 615 from the voice of the user operating the information processing terminal 22 as an output signal to the information processing device 10. The output unit 102 of the information processing terminal 22 outputs sounds such as ambient sounds according to the sound data received from the information processing device 10. Note that the output signal transmission unit 100 shown in Figure 19 is an example of an input device. The output unit 102 is an example of an output device.
[0108] The communication unit 42 of the information processing device 10 receives an electrical signal converted by the microphone 615 from the voice of the user operating the information processing terminal 22 as an output signal from the output signal transmission unit 100 of the information processing terminal 22. The communication unit 42 also receives operation signals received by the information processing terminal 22 from the user.
[0109] The authentication processing unit 38 performs authentication processing for the user operating the information processing terminal 22. The video display control unit 30 controls the video, such as the shared screen, displayed by the information processing terminal 22 during an online meeting.
[0110] The acquisition unit 32 acquires user activity information during an online meeting. Examples of user activity information acquired by the acquisition unit 32 include the amount of speech spoken by multiple users during the online meeting. Another example of user activity information acquired by the acquisition unit 32 is the frequency of speaker changes among multiple users during the online meeting. Yet another example of user activity information acquired by the acquisition unit 32 is information on users who continue to speak for a predetermined amount of time or longer during the online meeting. Speech volume, speaker change frequency, and information on users who continue to speak can be measured from the output signal of the microphone 615.
[0111] Furthermore, the acquisition unit 32 acquires environment-dependent information such as weather, temperature, humidity, illuminance, equipment operating sounds, noise, or time of day near the information processing terminal 22. The generation unit 34 generates sound data as described later, based on the activity information of multiple users during the online meeting and the environment-dependent information near the information processing terminal 22. The generation unit 34 may also generate sound data as described later, based on the activity information of multiple users during the online meeting without using environment-dependent information. The sound output control unit 36 controls the output unit 102 of the information processing terminal 22 to output ambient sounds corresponding to the generated sound data.
[0112] The memory unit 50 stores reservation information 52, sound source information 54, number of notes information 56, beat information 58, timbre information 60, and melody information 62 in a table format, for example, as shown in Figures 6 to 9, 11, and 20.
[0113] Since the reservation information 52, sound source information 54, note count information 56, beat count information 58, and melody information 62 are identical to those in the first embodiment except for a few parts, the explanation of the identical parts will be omitted.
[0114] The Room ID in the reservation information in Figure 6 is an example of identification information for an online meeting reserved by the reservation information. The reservation time is an example of date and time information for an online meeting reserved by the reservation information. The participating users are an example of participant information for an online meeting reserved by the reservation information. The multiple time slots A to D in the sound source information in Figure 7 are an example of time slot information, where the online meeting reservation time is divided into four parts.
[0115] Figure 8 shows an example of speech volume information, which represents the frequency of user speech during an online meeting. As an example, Figure 8 shows speech volume as the number of seconds during a predetermined period (e.g., the last 60 seconds) when at least one user was speaking during the online meeting.
[0116] Figure 9 shows an example of speaker change frequency information, which represents the level of conversation activity among multiple users during an online meeting, expressed by the frequency of speaker changes. As an example, Figure 9 shows the speaker change frequency in an online meeting as the number of times the speaker changed within a predetermined time period (e.g., the last 60 seconds).
[0117] Figure 20 is a diagram illustrating an example of timbre information. The timbre information in Figure 20 includes the following items: timbre class, screen change amount, and timbre. The timbre class is an example of identification information for classification. The screen change amount is an example of information representing the frequency of screen changes on the information processing terminal 22 operated by multiple users during an online meeting. In Figure 20, as an example, the frequency of screen changes on the information processing terminal 22 operated by multiple users during an online meeting is represented by the number of times the screen of the information processing terminal 22 operated by multiple users during an online meeting has changed by a predetermined percentage or more in the last 60 seconds. The timbre represents the timbre used for ambient sound. According to the timbre information in Figure 20, the timbre used for ambient sound can be changed depending on the frequency of screen changes on the information processing terminal 22 operated by multiple users during an online meeting.
[0118] The participating users in the melody information in Figure 11 are an example of information representing participants in an online meeting that has been reserved according to the reservation information. According to the melody information in Figure 11, if a specific user speaks continuously for a predetermined period of time or longer during an online meeting, the melody assigned to the speaking user can be used as ambient sound.
[0119] The information processing system 2 according to the third embodiment outputs ambient sounds to the user's information processing terminal 22 during an online meeting, for example, in the procedure shown in Figure 21. Figure 21 is a flowchart illustrating an example of the processing procedure of the information processing system according to this embodiment.
[0120] In step S400, in the information processing system 2 according to the third embodiment, a user, such as the organizer of an online meeting, performs preliminary preparations. These preparations include registering the reservation information shown in Figure 6, setting the sound source information shown in Figure 7, setting the number of notes information shown in Figure 8, setting the beat count information shown in Figure 9, setting the timbre information shown in Figure 20, and setting the melody information shown in Figure 11, all of which are stored in the storage unit 50 of the information processing device 10. The user accesses the information processing device 10 using the information processing terminal 22, and the communication unit 42 of the information processing device 10 receives operation information from the information processing terminal 22, allowing the user to perform any of the following operations: modify, add, or delete the various information stored in the storage unit 50. Note that the setting of the sound source information shown in Figure 7, the setting of the number of notes information shown in Figure 8, the setting of the beat count information shown in Figure 9, the setting of the timbre information shown in Figure 20, and the setting of the melody information shown in Figure 11 may be automatically set by the information processing device 10 based on the registration of the reservation information shown in Figure 6.
[0121] In step S402, the information processing system 2 according to the third embodiment determines that the online meeting has started according to the reservation information in Figure 6. The determination to start the online meeting may be made by the communication unit 42 of the information processing system 10 receiving information based on operation input entered by a user, such as the organizer of the online meeting, into the information processing terminal 22, and making a determination based on the received information, or it may be started automatically according to the reservation time in the reservation information in Figure 6.
[0122] In step S404, the information processing system 2 according to the third embodiment acquires activity information of multiple users during the online meeting. The user activity information acquired by the acquisition unit 32 in step S404 includes the amount of speech spoken by multiple users during the online meeting, the frequency of speaker changes, and information on users who continue to speak for a predetermined amount of time or longer during the online meeting. In addition, the user activity information acquired by the acquisition unit 32 in step S404 is the amount of screen change of the information processing terminal 22 operated by multiple users during the online meeting.
[0123] In step S406, in the information processing system 2 according to the second embodiment, the generation unit 34 generates sound data in a procedure such as that shown in Figure 22, based on the activity information of multiple users during the online meeting acquired in step S404.
[0124] Figure 22 is a flowchart of an example of the process for generating sound data. In step S500, the generation unit 34 determines the set of sound sources to be assigned to time slots A to D of the online meeting reservation time, based on the reservation information in Figure 6 and the sound source information in Figure 7.
[0125] In step S502, the generation unit 34 determines the number of sounds to be used as ambient sounds based on the number of sounds spoken by multiple users during the online meeting, using the sound count information in Figure 8. In step S404, the generation unit 34 determines the number of beats to be used as ambient sounds based on the beat count information in Figure 9, using the frequency of speaker changes among multiple users during the online meeting.
[0126] In step S506, the generation unit 34 determines the tone to be used for ambient sound based on the tone information in Figure 20 and the amount of screen change on the user's information processing terminal 22 during the online meeting.
[0127] Furthermore, in step S508, the generation unit 34 determines that a specific user participating in the online meeting is the user to be played a refrain if that user has been speaking continuously for a predetermined period of time or longer. Based on the melody information in Figure 11, the generation unit 34 determines the melody to be assigned to the user to be played a refrain. In step S510, the generation unit 34 generates sound data based on the determined sound source set, number of notes, number of beats, timbre, and melody.
[0128] Returning to step S408 in Figure 21, the sound output control unit 36 controls the output unit 102 of the information processing terminals 22 of multiple users participating in the online meeting to output ambient sounds corresponding to the sound data generated in step S406.
[0129] Thus, the information processing system 2 according to the third embodiment can output ambient sounds that change according to the status of conversations between users in an online meeting in which multiple users are participating.
[0130] By setting the sound source set, number of sounds, number of beats, timbre, and melody used to generate the sound data in step S406 so that ambient sounds suitable for the user's situation during the online meeting are output, the information processing system 2 according to the third embodiment can output ambient sounds suitable for the interaction between users during the online meeting.
[0131] For example, the information processing system 2 according to the second embodiment can output ambient sounds suitable for multiple users with high levels of tension to the online meeting, assuming that the greater the amount of speech and the more frequently the speaker changes among multiple users during the online meeting, the higher the tension level of the users participating in the online meeting.
[0132] The process in steps S404 to S410 is repeated until the online meeting ends. When the online meeting ends, the process proceeds to step S412, and the sound output control unit 36 terminates the output of ambient sound from the output unit 110 of the output unit 102 of the information processing terminal 22.
[0133] The present invention is not limited to the embodiments specifically disclosed above, and various modifications and changes are possible without departing from the scope of the claims. It goes without saying that the information processing systems 1 and 2 described in this embodiment are merely examples, and there are various system configurations depending on the application and purpose.
[0134] Each of the functions of the embodiments described above can be realized by one or more processing circuits. Hereinafter, "processing circuit" as used herein includes processors programmed to execute each function by software, such as processors implemented by electronic circuits, as well as devices such as ASICs (Application Specific Integrated Circuits), DSPs (digital signal processors), FPGAs (field programmable gate arrays), and conventional circuit modules designed to execute each of the functions described above.
[0135] The apparatus described in the examples represents only one of several computing environments for carrying out the embodiments disclosed herein. In one embodiment, the information processing apparatus 10 includes multiple computing devices, such as a server cluster. The multiple computing devices are configured to communicate with each other via any type of communication link, including a network or shared memory, and perform the processing disclosed herein.
[0136] Furthermore, the information processing device 10 can combine the disclosed processing steps in various ways. Each element of the information processing device 10 may be combined into a single device or divided into multiple devices. Also, each processing performed by the information processing device 10 may be performed by the information processing terminal 22. In addition, user activity information may include, for example, the number of users in the meeting room, the user's heart rate, etc. [Explanation of Symbols]
[0137] 1.2 Information Processing Systems 10 Information Processing Devices 16 speakers 18 Cameras 20 microphones 22 Information Processing Terminals 32 Acquisition Department 34 Generation part 36 Sound output control unit 50 Storage section 52 Reservation Information 54 Audio Source Information 56 Number of sounds information 58 beat count information 60 Tone Information 62 Melody Information 70, 80, 90, 100 Output signal transmission section 102, 110 Output section N Network [Prior art documents] [Patent Documents]
[0138] [Patent Document 1] Special Publication No. 2018-512607
Claims
1. A means of acquiring activity information of multiple users engaging in a conversation, means for assigning melody information corresponding to each user in the aforementioned multiple users, A means for identifying the user who is speaking from among the aforementioned multiple users, The generation means, when the identified user has been speaking continuously for a predetermined period of time or longer, determines that the identified user is a user to play a refrain, and generates sound data based on the activity information of the identified user and the melody to be used for the refrain performance assigned to the determined user. Sound output control means for outputting ambient sound corresponding to the aforementioned sound data to an output device, An information processing device having
2. The acquisition means acquires at least one of the speech volume and posture information of the multiple users who are in the same space as the activity information. The information processing apparatus according to claim 1.
3. The acquisition means acquires the amount of speech uttered by the multiple users who converse via the communication network as activity information. The information processing apparatus according to claim 1.
4. The acquisition means acquires at least one of the following as activity information: the frequency of speaker changes among the multiple users, the amount of screen change of the information processing terminal operated by the multiple users, the number of the multiple users, and the heart rate of the multiple users. The information processing apparatus according to claim 2 or 3.
5. The acquisition means further acquires environment-dependent information, either inside or outside the same space. The generation means generates the sound data based on the activity information and the environment-dependent information. The information processing apparatus according to claim 2.
6. The acquisition means acquires the amount of speech uttered by the plurality of users, measured based on the output signal from the microphone, as the activity information of the plurality of users. The information processing apparatus according to claim 2 or 3.
7. The acquisition means acquires the posture information of the multiple users recognized based on the output signal from the camera as the activity information of the multiple users. The information processing apparatus according to claim 2.
8. The generation means determines the status of the multiple users based on the activity information and generates sound data to output ambient sounds to the output device according to the status. The information processing apparatus according to any one of claims 1 to 7.
9. The generation means changes the sound data generated based on the activity information as time passes, when the multiple users have a conversation at a predetermined time. The information processing apparatus according to any one of claims 1 to 8.
10. The sound output control means uses at least one of the speakers installed in the space where the multiple users are located, and the information processing terminals operated by the multiple users, as output devices to output ambient sounds corresponding to the sound data. The information processing apparatus according to any one of claims 1 to 9.
11. The sound output control means changes the ambient sound output to the space where the multiple users are located, depending on the part of the space. The information processing apparatus according to claim 2.
12. The generation means generates sound data in which at least one of the number of notes, beats, timbre, and melody is different, based on the activity information. The information processing apparatus according to any one of claims 1 to 11.
13. The system further includes a determination means for determining whether a predetermined time has elapsed since the identified user began speaking. When the determination means determines that a predetermined time has elapsed since a specific user began speaking, the sound output control means causes the output device to output ambient sound corresponding to the sound data. The information processing apparatus according to any one of claims 1 to 12.
14. A means of acquiring activity information of multiple users engaging in a conversation, means for assigning melody information corresponding to each user in the aforementioned multiple users, A means for identifying the user who is speaking from among the aforementioned multiple users, The generation means, when the identified user has been speaking continuously for a predetermined period of time or longer, determines that the identified user is a user to play a refrain, and generates sound data based on the activity information of the identified user and the melody to be used for playing the refrain assigned to the determined user. Sound output control means for outputting ambient sound corresponding to the aforementioned sound data to an output device, An information processing system having
15. An information processing system having an input device, an information processing device, and an output device, The aforementioned input device is Output signal transmission means for transmitting output signals related to the activities of multiple users engaging in a conversation to the information processing device, It has, The aforementioned information processing device is An acquisition means for acquiring activity information of the plurality of users based on the output signal, means for assigning melody information corresponding to each user in the aforementioned multiple users, A means for identifying the user who is speaking from among the aforementioned multiple users, The generation means, when the identified user has been speaking continuously for a predetermined period of time or longer, determines that the identified user is a user to play a refrain, and generates sound data based on the activity information of the identified user and the melody to be used for playing the refrain assigned to the determined user. Sound output control means for outputting ambient sound corresponding to the aforementioned sound data to an output device, It has, The output device is Output means for outputting the aforementioned ambient sound An information processing system having
16. In an information processing device, Procedure for obtaining activity information of multiple users engaging in a conversation. A procedure for assigning melody information corresponding to each user in the aforementioned multiple users, A procedure for identifying the user who is speaking from among the aforementioned multiple users, A generation procedure for determining a user to be a user to play a refrain if the user has been speaking continuously for a predetermined period of time or longer, and for generating sound data based on the activity information of the user and the melody to be used for playing the refrain assigned to the determined user. A sound output control procedure that causes an output device to output ambient sound corresponding to the aforementioned sound data. A program to execute.
Citation Information
Patent Citations
Group emotion recognition support system
JP2008272019A
Information processor, information processing method and program
JP2017026568A
Communication supporting system
JP2017201479A
Methods, Systems and Mediums for Modification of Environmental Background Noise Based on Mood and / or Behavioral Information
JP2018512607A
Electronic device control method, electronic device control system, electronic device, and program
JP2020065097A