Virtual space control system, virtual space control method, and program
The virtual space control system enhances user interaction by adjusting display and audio outputs based on call information and user status, addressing reduced communication opportunities in virtual environments.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-09-24
- Publication Date
- 2026-04-03
AI Technical Summary
Existing virtual space communication systems fail to facilitate effective interaction between users due to unclear user availability and status, leading to reduced communication opportunities.
A virtual space control system that detects user interaction requests and adjusts display and audio outputs based on acquired call information, including fatigue level estimation and relationship intimacy, to enhance user interaction.
Facilitates more realistic and efficient communication by providing real-time, personalized audio and visual cues that reflect user status and relationship, reducing psychological barriers and improving interaction quality.
Smart Images

Figure 2026058058000001_ABST
Abstract
Description
[Technical Field]
[0001] This disclosure relates to a virtual space control system, a virtual space control method, and a program. [Background technology]
[0002] Patent Document 1 discloses an information processing device that displays icons corresponding to users in order to provide a virtual office. This information processing device displays an icon for each user that is shaped like a part of a person. Furthermore, the information processing device switches the user's icon depending on whether the user is in a first state where they are not able to converse or a second state where they are able to converse.
[0003] Patent Document 2 discloses a speech synthesis device that synthesizes speech from images. In Patent Document 2, a speech synthesis model is constructed by machine learning using a set of text data, facial image data, and speech data. The speech synthesis device has an image encoder that receives image data as input and generates a feature vector corresponding to the speaker. The speech synthesis device takes the feature vector and text data as input and synthesizes and outputs speech that sounds as if it were spoken by a speaker. [Prior art documents] [Patent Documents]
[0004] [Patent Document 1] Japanese Patent Publication No. 2024-31550 [Patent Document 2] Japanese Patent Publication No. 2021-99454 [Overview of the project] [Problems that the invention aims to solve]
[0005] Patent Document 1 changes the icon display depending on whether the user is in a state where they can communicate or not. Therefore, other users cannot contact a user whose icon indicates they are not in a state where they can communicate. In addition, when one user wants to contact another user, the other user may hesitate to do so because they do not know the status of the other user. This leads to a problem where opportunities for communication are reduced.
[0006] This disclosure is made in view of the above points and aims to provide a virtual space control system, a virtual space control method, and a program that can facilitate communication between users in a virtual space. [Means for solving the problem]
[0007] The virtual space control system according to this embodiment is a virtual space control system that displays icon images corresponding to users using the virtual space on the virtual space, and comprises: a detection unit that detects an operation to request a call from a second user to a first user; a call information acquisition unit that acquires call information between the first user and the second user; and an output control unit that, when an operation to request a call is detected, performs control to change at least one of the display output of the first user's icon image and the voice output of the first user based on the call information.
[0008] The virtual space control method according to this embodiment is a virtual space control method that displays an icon image corresponding to a user using the virtual space on the virtual space, and includes the steps of: detecting an operation to request a call from a second user to a first user; acquiring call information between the first user and the second user; and, when an operation to request a call is detected, performing control to change at least one of the display output of the first user's icon image and the voice output of the first user based on the call information.
[0009] The program according to this embodiment is a program that causes a computer to execute a virtual space control method for displaying icon images corresponding to users using the virtual space in the virtual space, and causes the computer to execute the following steps: detecting an operation to request a call from a second user to a first user; acquiring call information between the first user and the second user; and, when the operation to request the call is detected, performing control to change at least one of the display output of the first user's icon image and the audio output of the second user based on the call information. [Effects of the Invention]
[0010] The purpose of this disclosure is to provide a virtual space control system, a virtual space control method, and a program that can facilitate communication between users. [Brief explanation of the drawing]
[0011] [Figure 1] This diagram schematically shows the overall configuration of the virtual space control system. [Figure 2] This figure shows an example of a window displaying a virtual office. [Figure 3] This is a block diagram showing an example configuration of a server that constitutes a virtual space control system, according to a first embodiment. [Figure 4] This block diagram shows an example configuration of a second embodiment of a server constituting a virtual space control system. [Figure 5] This is a sequence diagram showing an example of a virtual space control method according to the first embodiment. [Figure 6] This is a flowchart showing the first embodiment of the virtual space control method in the second embodiment. [Figure 7] This is a flowchart showing a second embodiment of the virtual space control method in the second embodiment. [Modes for carrying out the invention]
[0012] Hereinafter, specific embodiments to which the present invention is applied will be described in detail with reference to the drawings. However, the present disclosure is not limited to the following embodiments. Also, for clarity of explanation, the following description and drawings are simplified as appropriate.
[0013] FIG. 1 is a diagram schematically showing the overall configuration of a virtual space control system according to the present embodiment. The virtual space control system 10 is a system that provides a virtual space used by a plurality of users. The virtual space control system 10 includes a server 20 and a plurality of user terminals 30. The server 20 and the user terminals 30 are connected via a network N. Also, the server 20 can be connected to an in-house database, the Internet, etc. in order to acquire various information. The server 20 is not limited to a physically single device and may be distributed among a plurality of devices.
[0014] Here, the virtual space is, for example, a virtual office used by a user U such as an employee of a company. That is, the virtual space control system 10 is used as a virtual office providing service on the Internet. Note that the virtual space is a space constructed on a computer and is a space that can be used by a plurality of users U on the network N using the user terminals 30.
[0015] The server 20 transmits data for providing a virtual office to each user terminal 30. Specifically, the server 20 provides an image of the virtual space that becomes the virtual office. Also, the server 20 receives data for controlling the virtual office from each user terminal 30. Specifically, the server 20 accepts a request content based on the input data of the user at the user terminal 30. The virtual office is utilized as a shared space where a plurality of users U use. Note that the virtual space provided by the server 20 is not limited to a virtual office. For example, the server 20 may provide a virtual space for a school, an event, etc.
[0016] Figure 2 shows an example of a window displaying Virtual Office 2. Specifically, Figure 2 shows Window 1 displayed by user terminal 30. Window 1 may be displayed on a web browser or by a dedicated application program.
[0017] Window 1 displays Virtual Office 2. Virtual Office 2, like a real office, includes multiple seats 6. Icons 3 corresponding to user U are displayed on each seat 6. Furthermore, a pointer 4 for manipulating Icon 3 is displayed in Window 1. For example, a user can move Icon 3 within Virtual Office 2 by manipulating Pointer 4 with a mouse or other device. Virtual Office 2 may also include meeting rooms, etc.
[0018] In this way, the server 20 controls the display so that each user terminal 30 displays window 1 of the virtual office 2, which includes the seat 6 and the icon 3 of user U. Here, since two users U are using the virtual office 2, two icons 3 are displayed. Of course, the number of users U and the number of icons 3 are not particularly limited. The virtual office 2 only needs to display icons 3 corresponding to the number of users U currently using it.
[0019] The user terminal 30 is a communication terminal such as a personal computer, smartphone, or tablet PC. Alternatively, the user terminal 30 may be a VR (Virtual Reality) device such as a head-mounted display. In this case, the virtual office 2 and icons 3 may be displayed in three dimensions.
[0020] The user terminal 30 includes a processor, memory, communication functions, input / output means, etc. The user terminal 30's input / output means include display means such as a display (not shown), input means such as a mouse, keyboard, touch panel, audio output means such as a speaker, and audio input means such as a microphone.
[0021] The user terminal 30 has memory in which computer programs are stored. The processor reads and executes the computer programs stored in memory to realize various functions. The user terminal 30 is not limited to a single physical device. For example, earphones, headphones, etc., provided separately from the user terminal 30 may be used as the means of audio output.
[0022] User U on user terminal 30 is a user who utilizes the services provided by virtual office 2. Each user U has a user ID and password set for logging in. For example, user U operates user terminal 30, enters their user ID and password, and sends them to server 20. Once server 20 authenticates the user, it sends display data to user terminal 30 to display the virtual office. As a result, user terminal 30 displays virtual office 2, as shown in Figure 2.
[0023] Furthermore, when user U logs into virtual office 2, their icon 3 is displayed within virtual office 2. In other words, user U can enter virtual office 2 by logging in. Each user U has an image registered to be used as their icon 3. For example, the image of icon 3 is the face image of the corresponding user U. Alternatively, the image of icon 3 may be an avatar created from user U's face image, or an avatar created by user U. In addition, each user U has user information such as name, age, affiliation, position, and face image pre-registered. User information is associated with the user ID and stored in storage devices or internal databases.
[0024] Each user U can manipulate their own icon 3 using pointer 4, etc. User U can move their own icon 3 within virtual office 2 by manipulating pointer 4 with a mouse, etc. For example, user U selects icon 3 with pointer 4 and moves pointer 4 to seat 6. This changes the display position of icon 3 in window 1. Also, in order for user U to speak to another user U, user U manipulates pointer 4 to move their own icon 3 near the other user U's icon 3. Alternatively, user U manipulates pointer 4 to select the icon 3 of the other user U they want to speak to.
[0025] Furthermore, within Virtual Office 2, User U can make calls (converse) with other users. For example, User U inputs voice data through the microphone on User Terminal 30. User Terminal 30 sends this voice data to Server 20. Server 20 sends the voice data to the User Terminal 30 of the other User U who is the other party to the call. The other User U's User Terminal 30 then outputs the voice through its speaker or earphones. Of course, it is also possible for three or more Users U to talk simultaneously. Note that calls are not limited to voice input using a microphone; they can also be made using text input with a touch panel or keyboard.
[0026] The following describes the process of two users U making a call. Here, one user makes a call request to another user. User U who receives the call request is referred to as User 1, and User U who sends the call request is referred to as User 2. In other words, User 2 operates the user terminal 30 to contact User 1 in order to make a call with User 1. Therefore, User 1 becomes the user receiving the call, and User 2 becomes the user making the call.
[0027] Automatic text generation and speech synthesis can be used for calls between users. Server 20 automatically generates text to create a response message from the first user to the second user. Server 20 generates audio data for the response message using speech synthesis. Server 20 sends the audio data of the response message to the second user's user terminal 30. The second user's user terminal 30 plays the response message using the audio data. This allows the second user to hear the automatically generated audio of the response message.
[0028] <First Embodiment> Figure 3 is a block diagram showing an example configuration of the server 20 that constitutes the virtual space control system in a first embodiment. The following describes the process of automatic text generation and speech synthesis with the first user as the speaker. Specifically, it describes the process of automatically generating text as a response message from the first user and synthesizing it as speech when the second user performs an operation to request a call to the first user's icon 3. Note that the operation to request a call can also be rephrased as the call request operation.
[0029] Server 20 includes a text generation unit 200, a speech synthesis unit 300, and a fatigue level estimation unit 400. Furthermore, Server 20 includes a user information acquisition unit 111, a time information acquisition unit 112, a weather information acquisition unit 113, an organization information acquisition unit 114, a business information acquisition unit 115, a biometric information acquisition unit 116, and a detection unit 121. The user information acquisition unit 111, time information acquisition unit 112, weather information acquisition unit 113, organization information acquisition unit 114, business information acquisition unit 115, and biometric information acquisition unit 116 are collectively referred to as the information acquisition unit 110. The information acquisition unit 110 acquires various types of information from the Internet, storage devices, user terminals 30, or internal company databases.
[0030] Server 20, in a configuration not shown, includes an information processing unit such as a CPU (Central Processing Unit) and an MPU (Micro Processing Unit), memory such as RAM (Random Access Memory), a non-volatile storage device such as flash memory, and a storage device such as an HDD (Hard Disk Drive). The storage device stores a computer program on which the processing described herein is implemented. The information processing unit can load the computer program from the storage device into memory and execute the computer program. As a result, the information processing unit realizes the functions of a user information acquisition unit 111, a time information acquisition unit 112, a weather information acquisition unit 113, an organization information acquisition unit 114, a business information acquisition unit 115, a biometric information acquisition unit 116, a detection unit 121, a text generation unit 200, a speech synthesis unit 300, and a fatigue level estimation unit 400.
[0031] Server 20 includes a communication unit, which is not shown in the diagram. The communication unit is composed of, for example, a communication module. The communication unit mediates the communication of various types of data between Server 20 and User Terminal 30 via the network N, and appropriately mediates the communication of input and output data between the Information Acquisition Unit 110, the Detection Unit 121, and the Speech Synthesis Unit 300.
[0032] In the following explanation, it is assumed that the server 20 performs all processing, but the user terminal 30 may perform at least part of the processing. For example, part of the processing of the text generation unit 200, the speech synthesis unit 300, or the fatigue level estimation unit 400 may be performed by the user terminal 30.
[0033] The user information acquisition unit 111 acquires user information of the first user. The user information includes, for example, a facial image of user U. Furthermore, the user information may include personal information such as user U's name, age, gender, place of origin, and native language. The user information may also include the voice of user U acquired by the microphone. The user information acquisition unit 111 provides the acquired user information to the speech synthesis unit 300.
[0034] The time information acquisition unit 112 acquires time information indicating the current time. This time information includes the date, day of the week, and time of day at the user's location or office location. The time information acquisition unit 112 can acquire this time information from the web or from the operating system's clock, etc.
[0035] The weather information acquisition unit 113 acquires weather information indicating the current weather. Weather information indicates the weather at the user's location or the location of the office. The weather information acquisition unit 113 can obtain weather information from the web. For example, the weather information acquisition unit 113 acquires weather information by referring to information on a weather forecast website.
[0036] The organizational information acquisition unit 114 provides information indicating the position and relationships of each user, the first user and the second user, within the organization. For example, the organizational information may include information indicating the affiliation and position of each user, the first user and the second user. Alternatively, the organizational information may include the length of service and the length of service at their current affiliation for each user, the first user and the second user. The organizational information acquisition unit 114 can acquire the organizational information of each user, the first user and the second user, from the company's internal database or similar sources.
[0037] The text generation unit 200 generates text based on time information, weather information, and organizational information. The text is a text document that is output as voice from the user terminal 30 to the second user. The text document becomes a message output to the second user. The text generation unit 200 outputs the generated text to the speech synthesis unit 300.
[0038] The text generation unit 200 can generate text containing characteristic keywords by referring to time information and weather information. For example, if the time information indicates 9:00 AM, the text generation unit 200 can output "Good morning" as the first word in the text. If the weather information indicates clear skies, the text generation unit 200 can output words such as "The weather is nice and refreshing."
[0039] The text generation unit 200 outputs text based on the organizational information of the first user and the organizational information of the second user. For example, if it is a response from a subordinate (the first user) to a superior (the second user), the text generation unit 200 generates a text message using polite language. Alternatively, if it is a response from a senior colleague (the first user) with similar years of service to a junior colleague (the second user), or a response between colleagues with similar years of service at their current workplace, the text generation unit 200 generates a text message converted to a more informal tone. The text generation unit 200 determines the hierarchical relationship between the first user and the second user within the organization using the organizational information of the first user and the second user, and can then convert the response message to an informal or polite style.
[0040] The information input to the text generation unit 200 is not limited to the information described above. The text generation unit 200 may generate text using at least one of the following: time information, weather information, and organization information. The text generation unit 200 may generate text using two or all of the following: time information, weather information, and organization information. Alternatively, the text generation unit 200 may use information other than time information, weather information, and organization information. The text generation unit 200 may generate text without using any of the time information, weather information, or organization information. Furthermore, user U may specify the information to be used for text generation.
[0041] The text generation unit 200 may generate text using a general-purpose model such as a chatbot. Alternatively, the text generation unit 200 may use a generative AI model trained using an LLM (large language model). The text generation unit 200 can generate text using text generation AI technology such as ChatGPT (Chat Generative Pre-trained Transformer).
[0042] The speech synthesis unit 300 performs speech synthesis processing to output text as speech. In other words, the speech synthesis unit 300 generates speech data based on the text generated by the text generation unit 200. The speech synthesis unit 300 may, for example, generate speech data from text using AI speech generation technology. Alternatively, the speech synthesis unit 300 may synthesize speech that sounds like it was spoken by a speaker based on the feature quantities of the face image for icon 3 included in the user information. The speech synthesis unit 300 may also generate speech data that represents the content of the text in a tone similar to that spoken by the first user.
[0043] Furthermore, the speech synthesis unit 300 may have a speech synthesis model that takes facial images and text data as input and outputs speech data. The speech synthesis model is pre-built by machine learning using a set of text data, facial image data, and speech data, for example, as described in Patent Document 2. The speech synthesis unit 300 may also have an image encoder that extracts feature vectors from facial images, and the speech synthesis unit 300 may extract the speaker's feature vectors based on the facial images and generate speech data indicating the content of the text based on the feature vectors. In that case, the speech synthesis unit 300 performs speech synthesis processing so that the feature vectors of the synthesized speech are similar to the feature vectors of the first user.
[0044] In this way, voice data is generated that mimics the speaking style of the first user. The second user can listen to the automatically generated text message as if it were actually spoken by the first user. Of course, the user information used by the speech synthesis unit 300 is not limited to facial images. For example, the user information used by the speech synthesis unit 300 may be actual voice spoken by a real user. For example, the speech synthesis unit 300 extracts feature vectors from the first user's voice data recorded by a microphone. The speech synthesis unit 300 then performs speech synthesis processing so that the feature vectors of the synthesized voice are similar to the feature vectors of the first user.
[0045] The business information acquisition unit 115 acquires business information of the first user. Business information is information indicating the user's workload. Business information includes schedule information indicating the user's plans for the day. For example, schedule information includes information such as the time and location of meetings and business trips. Schedule information may also include information such as the type and content of meetings and business trips. Business information may also include information regarding the number of emails sent and received, the number of emails read, and the number of documents created. The business information acquisition unit 115 outputs the business information to the fatigue level estimation unit 400.
[0046] The biometric information acquisition unit 116 acquires biometric information of the first user. This biometric information includes pulse rate, heart rate, step count, and activity level, and is acquired, for example, by a wearable device such as a smartwatch. Alternatively, the biometric information may be a facial image captured by the user terminal 30 or the like. For example, the fatigue level estimation unit 400 may estimate the fatigue level by analyzing the facial image. Alternatively, the biometric information may be information obtained by spectral analysis of near-infrared light.
[0047] The fatigue level estimation unit 400 estimates the fatigue level of the first user based on the first user's work information acquired by the work information acquisition unit 115. For example, the more meetings the first user attends or the more emails they send and receive, the higher the fatigue level estimation unit 400 estimates the first user's fatigue level. Alternatively, the longer the working hours, the higher the fatigue level estimation unit 400 estimates the first user's fatigue level. An example of the fatigue level estimation process is described below.
[0048] WT [min] represents the planned number of working hours for the day, MT [min] represents the total number of scheduled meetings and discussions for the day, WR [%] represents the percentage of working hours elapsed at the current time, and M represents the number of emails received up to the current time. R , the number of emails sent is M S The fatigue coefficient for receiving emails is W R The fatigue coefficient for sending emails is W S The fatigue level estimation unit 400 calculates the fatigue level F[%] using the following formula (1).
[0049] F = WR * (MT / WT) + (MR *W R +M S *W S )···(1)
[0050] Here, the fatigue coefficients W R and W S are set to 0.01 and 0.1 respectively. The first term of Equation (1) represents the fatigue degree due to meetings and consultations, and the second term represents the fatigue degree due to email sending and receiving. The fatigue coefficient W R may be set within the range of 0.01 to 0.1 for the fatigue coefficient W S . It is estimated that email sending causes a higher degree of fatigue than email receiving. Therefore, the fatigue coefficient W R is set smaller than the fatigue coefficient W S . Of course, the calculation formulas for the fatigue coefficient and the fatigue degree are not limited to the above examples.
[0051] The fatigue degree estimation unit 400 may estimate the fatigue degree using both the biological information acquired by the biological information acquisition unit 116 and the business information acquired by the business information acquisition unit 115. Alternatively, the fatigue degree estimation unit 400 may estimate the fatigue degree using only one of the biological information and the business information. Furthermore, the fatigue degree estimation unit 400 may estimate the fatigue degree based on information other than the biological information and the business information. The fatigue degree estimation unit 400 outputs the estimated fatigue degree to the speech synthesis unit 300. Of course, the fatigue degree estimation unit 400 is not limited to the method of using biological information and business information, and may estimate the fatigue degree by other methods.
[0052] The speech synthesis unit 300 performs speech synthesis processing based on the fatigue degree of the first user estimated by the fatigue degree estimation unit 400. As a result, the speech synthesis unit 300 can generate speech data with a sense of fatigue. In order to perform speech processing using the fatigue degree, the speech synthesis unit 300 may use a machine learning model. For example, a machine learning model for speech synthesis can be constructed by performing machine learning in advance using known speech data (supervised speech data) with different fatigue degrees.
[0053] Of course, the speech synthesis unit 300 may perform speech synthesis processing without using a machine learning model. For example, the speech synthesis unit 300 may perform processing to correct the tone of voice according to the level of fatigue. The speech synthesis unit 300 may perform processing to convert the pitch according to the level of fatigue. The speech synthesis unit 300 may perform processing to change the pitch (rate of decrease of the fundamental frequency) of the output speech synthesis according to the level of fatigue. If the speech synthesis unit 300 is changing the pitch of the speech synthesis according to the level of fatigue, it will gradually change the rate of decrease of the fundamental frequency according to the level of fatigue.
[0054] Alternatively, the speech synthesis unit 300 may perform processing to emphasize the non-periodic components of the voice according to the fatigue level of the first user estimated by the fatigue level estimation unit 400. In this case, the speech synthesis unit 300 changes the amount of emphasis in stages according to the fatigue level of the first user. Specifically, the speech synthesis unit 300 performs processing to emphasize the hoarseness of the voice according to the fatigue level. The speech synthesis unit 300 may generate speech data by convolving a filter that indicates the characterization of the voice tone according to the fatigue level of the first user. The speech synthesis unit 300 may prepare multiple filters according to the fatigue level of the first user and select a filter from among the multiple filters that corresponds to the fatigue level of the first user. The speech synthesis unit 300 may generate speech data by convolving the selected filter into the speech signal.
[0055] When the second user performs an operation on the user terminal 30 to request a call from the first user, the speech synthesis unit 300 of the server 20 automatically generates text that will be a response message from the first user and performs speech synthesis processing, as described above. The speech synthesis unit 300 of the server 20 sends the audio data to the second user's user terminal 30. The second user's user terminal 30 outputs an audio message that will be a response message from the first user from the speaker or other device installed in the user terminal 30. As described above, the speech synthesis unit 300 of the server 20 performs speech synthesis processing so that the speaker or other device installed in the second user's user terminal 30 outputs an audio that mimics the voice of the first user.
[0056] In this way, a second user in the virtual office can hear automatically generated text in real time using synthesized speech that resembles the user's voice. The second user can also listen to synthesized speech that resembles the first user's voice beforehand and become familiar with their speaking style. Therefore, it becomes possible to lower the psychological barrier to initiating contact.
[0057] The fatigue level estimation unit 400 automatically estimates the fatigue level of the first user, which is private information related to the first user. The server 20 can generate voice data corresponding to the estimated fatigue level of the first user through speech synthesis. Therefore, the second user can understand the first user's physical condition by listening to the synthesized voice of the first user. This prevents actions such as forcing the conversation to continue unnecessarily. As a result, there is more leeway in work, enabling more efficient work. Even in a virtual space, interactions become more realistic, and work can proceed more smoothly.
[0058] The detection unit 121 detects an operation performed by the second user on the first user's icon. This operation by the second user on the first user's icon is an operation by the second user to request a call from the first user. When the second user's user terminal 30 detects an operation by the second user to request a call from the first user, it sends the detection result to the server 20. As a result, the detection unit 121 of the server 20 detects the operation by the second user to request a call via the communication unit.
[0059] Furthermore, when the detection unit 121 detects an operation to request a call from the second user, the user making the call is identified. The detection unit 121 identifies the second user based on the information of the user terminal 30 that performed the operation on the first user's icon. The detection unit 121 identifies the first user based on the information of the icon 3 that was operated on. Note that part of the process for detecting the second user's operation on the first user's icon may be performed on the second user's user terminal 30. In this case, the user terminal 30 only needs to send a detection signal to the server 20 indicating that it has detected a call request operation.
[0060] The following describes the operation for a second user to request a call from the first user. In window 1 shown in Figure 2, the second user places pointer 4 over the first user's icon. That is, the second user uses the mouse on user terminal 30 to move pointer 4 over icon 3. This mouseover operation constitutes the operation by which the second user requests a call from the first user. The operation to request a call is not limited to a mouseover operation. For example, the operation by the second user to search for a call recipient from the address book on user terminal 30 may also be considered the operation to request a call.
[0061] Alternatively, the operation to request a call by the second user may be to display the name of the first user listed in the address book in a predetermined position for a predetermined period of time or longer. Alternatively, the operation to request a call by the second user may be to display the cursor pointing to the name of the first user listed in the address book in the same position for a predetermined period of time or longer.
[0062] Alternatively, the second user may request a call by making a gesture (action) of staring at the first user's icon displayed in Window 1 shown in Figure 2 for a certain period of time or longer. In this case, the camera or sensors installed on the second user's user terminal 30 can detect the second user's gesture. Alternatively, the second user may request a call by calling out the first user's name using voice input via the microphone installed on the second user's user terminal 30.
[0063] <Second Embodiment> Figure 4 is a block diagram showing an example configuration of the server 20' that constitutes the virtual space control system in a second embodiment. Below, the processing of displaying the icon image of the first user based on call information between the first user and the second user, and the processing of outputting the synthesized voice of the first user will be described. Note that the details of the automatic text generation and speech synthesis processes with the first user as the speaker are the same as in Figure 3, so a detailed explanation will be omitted.
[0064] Server 20' includes a call information acquisition unit 117 and a closeness calculation unit 122 in addition to the configuration of Server 20. Furthermore, Server 20' includes an output control unit 130 that includes a speech synthesis unit 300 and an icon image generation unit 500, instead of the speech synthesis unit 300 in Figure 3. The processing units that add the call information acquisition unit 117 to the user information acquisition unit 111, time information acquisition unit 112, weather information acquisition unit 113, organization information acquisition unit 114, business information acquisition unit 115, biometric information acquisition unit 116 are collectively referred to as the information acquisition unit 110'. The user information acquisition unit 111, text generation unit 200, and fatigue level estimation unit 400 are the same as in Figure 3, so detailed explanations are omitted as appropriate. Also, the time information acquisition unit 112, weather information acquisition unit 113, organization information acquisition unit 114, business information acquisition unit 115, and biometric information acquisition unit 116 are the same as in Figure 3, so their illustrations are omitted.
[0065] Server 20' includes, in a configuration not shown, an information processing unit such as a CPU (Central Processing Unit) and an MPU (Micro Processing Unit), memory such as RAM (Random Access Memory), a non-volatile storage device such as flash memory, and a storage device such as an HDD (Hard Disk Drive). The storage device stores a computer program on which the processing described herein is implemented. The information processing unit can load the computer program from the storage device into memory and execute the computer program. As a result, the information processing unit realizes the functions of the user information acquisition unit 111, the time information acquisition unit 112, the weather information acquisition unit 113, the organization information acquisition unit 114, the business information acquisition unit 115, the biometric information acquisition unit 116, the call information acquisition unit 117, the detection unit 121, the intimacy calculation unit 122, the text generation unit 200, the fatigue level estimation unit 400, and the output control unit 130.
[0066] Server 20' includes a communication unit, which is not shown in the diagram. The communication unit is composed of, for example, a communication module. The communication unit mediates the communication of various types of data between Server 20' and the user terminal 30 via the network N, and appropriately mediates the communication of input and output data between the information acquisition unit 110', the detection unit 121, and the output control unit 130.
[0067] When server 20' detects an operation to request a call from the second user to the first user, server 20' controls the display output or audio output based on the call information. Specifically, server 20' controls the display by changing the image of the first user's icon 3 based on the call information regarding the call between the first user and the second user. Alternatively, server 20' changes the audio data output from the speaker based on the call information regarding the call between the first user and the second user. The output control by server 20' will be explained below using Figure 4. It is assumed that server 20' maintains a call history between each user in its storage device.
[0068] The call information acquisition unit 117 acquires call information relating to calls between users and acquires call information relating to past calls between the first user and the second user from the acquired call information. Thus, the call information acquisition unit 117 acquires call information relating to past calls between the first user and the second user. The call information may include information relating to at least one of the number of past calls between users, the duration of the calls, and the frequency of the calls. The call information may also be a history of past calls between the first user and the second user. The call history may include information such as the number of calls, the start date and time of the calls, the end date and time of the calls, the frequency of the calls, and the duration of the calls. Furthermore, the call information may also include information relating to the content of the calls.
[0069] The intimacy calculation unit 122 calculates the intimacy between users based on the call information acquired by the call information acquisition unit 117. For example, the more calls and the more frequent the calls, the closer the relationship between the first and second users. Alternatively, the longer the call duration, the closer the relationship between the first and second users. The intimacy calculation unit 122 may also evaluate intimacy by comparing the number of calls, call frequency, and call duration with thresholds. The intimacy calculation unit 122 may also calculate a score indicating intimacy using at least one of the pieces of information: the number of calls, call frequency, and call duration.
[0070] The output control unit 130 controls at least one of the display output of the icon image of the first user and the audio output to the second user, based on the call information acquired by the call information acquisition unit 117. The output control unit 130 controls the display output to change the icon image of the first user according to the call information acquired by the call information acquisition unit 117. For example, the output control unit 130 changes the control parameters for generating the icon image according to the call information acquired by the call information acquisition unit 117 and outputs them to the icon image generation unit 500. The icon image generation unit 500 applies predetermined processing to the icon 3 of the first user and outputs the changed icon image. Then, the output control unit 130 transmits the icon image, which is the display output output by the icon image generation unit 500, to the user terminal 30 via the communication unit. Here, the predetermined processing is, for example, the process of adding a fatigue level.
[0071] Alternatively, the output control unit 130 controls the audio output to change the audio data according to the call information acquired by the call information acquisition unit 117. For example, the output control unit 130 changes the control parameters for the speech synthesis process according to the call information acquired by the call information acquisition unit 117 and outputs them to the speech synthesis unit 300. Alternatively, it changes the control parameters for the text generation process according to the call information acquired by the call information acquisition unit 117 and outputs them to the text generation unit 200. Then, the output control unit 130 transmits the audio data output by the speech synthesis unit 300 to the user terminal 30 via the communication unit. Here, the predetermined processing is, for example, the process of adding a fatigue level.
[0072] The following describes an example of processing in the output control unit 130. Here, we describe an example in which the output control unit 130 controls the display output according to the fatigue level estimated by the fatigue level estimation unit 400. For example, the output control unit 130 generates an icon image by adding information indicating the fatigue level to the face image. The fatigue level estimation unit 400 calculates the fatigue level as a percentage [ ] as shown in equation (1). Then, the output control unit 130 transmits the icon image containing the fatigue level information to the second user's user terminal 30.
[0073] The output control unit 130 changes the display mode indicating fatigue level according to the call information. The output control unit 130 changes the display mode to show fatigue level in more detail for users whose number of calls or call frequency exceeds a predetermined value. Specifically, for users whose number of calls or call frequency exceeds a predetermined value, the output control unit 130 controls the display output to become an icon image with a numerical value [%] indicating fatigue level added to the first user's face image.
[0074] On the other hand, for users whose call count and call frequency are below a predetermined value, the output control unit 130 changes the display mode so that the fatigue level is not displayed in detail. Here, the output control unit 130 changes the display mode so that the fatigue level is shown in four stages: "energetic (not tired)", "slightly tired", "tired", and "very tired". For example, the output control unit 130 uses an icon image that is a face image of the first user with a message indicating the fatigue level added to it. The messages indicating the fatigue level are, for example, "energetic (not tired)", "slightly tired", "tired", and "very tired".
[0075] Alternatively, the output control unit 130 may change the output mode by changing the face image according to the degree of fatigue. For example, the output control unit 130 changes the face image included in the icon image to a tired expression the higher the degree of fatigue. The output control unit 130 can change the face image to a tired expression by performing image processing. A machine learning model, such as supervised learning, can be used for this image processing. Alternatively, multiple images of the user's face when they are tired may be taken, and an icon image may be created from those face images.
[0076] When users make many calls or call frequently, it can be inferred that they have a close relationship. Generally, if users have a close relationship, they may disclose private information to the person requesting the call, and if they do not have a close relationship, they should not disclose too much private information to the person requesting the call. Therefore, when users have a close relationship, the output control unit 130 controls the display output so that the second user can understand the first user's fatigue level in more detail. For example, the output control unit 130 changes the face image according to the fatigue level [%].
[0077] On the other hand, among users whose number of calls or call frequency is below a predetermined value, it can be inferred that the relationship between the users is not very close. In such cases, the output control unit 130 changes the face image in four stages in order to abstract private information. The server 20 then sends an icon image of the face image with the expression changed according to the fatigue level to the user terminal 30. The output control unit 130 changes the display mode of the icon 3 according to the call information. Furthermore, the output control unit 130 may also change the color, size, or shape of the icon image according to the fatigue level. Note that the output control unit 130 only needs to send the icon image with the fatigue level added to the user terminal 30 of the second user.
[0078] The output control unit 130 may control the voice output by controlling the voice synthesis processing in the voice synthesis unit 300. If the voice synthesis unit 300 changes the pitch of voice synthesis according to the level of fatigue, the control parameters are set so that the rate of decrease of the fundamental frequency changes according to the level of fatigue. Alternatively, if the voice synthesis unit 300 is performing processing to emphasize the non-periodic components of the voice according to the level of fatigue, the voice synthesis unit 300 is set so that the amount of emphasis changes according to the level of fatigue.
[0079] Furthermore, in the case of users with a high number of calls or high call frequency, the output control unit 130 controls the speech synthesis unit 300 so that the second user can understand the first user's fatigue level in more detail. For example, in the case of users with a number of calls or call frequency exceeding a predetermined value, the output control unit 130 changes the control parameters over a wider range, and in the case of users with a number of calls or call frequency below a predetermined value, the output control unit 130 changes the control parameters over a narrower range. In this way, the speech output pattern changes in the case of users with a number of calls or call frequency exceeding a predetermined value so that the second user can understand the fatigue level in more detail.
[0080] If the speech synthesis unit 300 uses a machine learning model, it should use a machine learning model that takes fatigue level as input. The output control unit 130 may prepare multiple machine learning models and change the machine learning model based on the call information. In other words, the output control unit 130 uses different speech synthesis models depending on whether the users are close or not.
[0081] The output control unit 130 may control the audio output by controlling the text generation process in the text generation unit 200. The text generation unit 200 generates text according to the fatigue level. Furthermore, the output control unit 130 controls the text generation unit 200 so that the text generated changes according to the fatigue level. For example, between users whose number of calls or call frequency exceeds a predetermined value, the output control unit 130 controls the text generation unit 200 so that it creates a message that includes a numerical value [%] indicating the fatigue level.
[0082] For users whose number of calls or call frequency is below a predetermined value, text indicating fatigue level in four stages is generated. Examples of fatigue level text include "Energetic (not tired)", "Slightly tired", "Tired", and "Very tired". In this way, for users whose number of calls or call frequency exceeds a predetermined value, the voice output method changes so that a second user can understand the fatigue level in more detail.
[0083] Furthermore, while the above explanation defines a close relationship between users as those whose number of calls or call frequency exceeds a predetermined value, a different method may be used to determine whether a relationship between users is close. For example, the closeness calculation unit 122 may calculate the closeness based on call information. The closeness calculation unit 122 calculates the closeness by weighting and adding the number of calls, call frequency, and call duration. Alternatively, the closeness calculation unit 122 may calculate the closeness using not only call information but also organizational information, etc. The closeness calculation unit 122 determines whether the first user and the second user are close based on whether the closeness is above a threshold. Then, the output control unit 130 performs output control to change the icon display mode or the audio output mode according to the closeness.
[0084] This allows a second user to gain a more detailed understanding of the fatigue level between close users. Because the second user can understand the first user's physical condition, they can avoid actions such as forcing conversations. This creates more flexibility in work, enabling more efficient work processes. It also brings a sense of realism to interactions even in a virtual space, allowing for smoother work progress. Note that while we have described the fatigue level as having four stages for users with call frequency below a predetermined value, the number of fatigue levels used for users with call frequency below a predetermined value is not limited to these.
[0085] Figure 5 is a sequence diagram showing an example of the virtual space control method of the first embodiment. Specifically, Figure 5 shows the process of automatically providing a voice response when a second user makes a call request to the first user. Each user U is logged into the virtual office. Therefore, the user terminal 30 is displaying window 1 of the virtual office 2 as shown in Figure 2.
[0086] First, the second user, user U, places the pointer 4 in window 1 displayed by user terminal 30 over the first user's icon 3 (S10). In other words, the second user mousees over the first user's icon to request a call to the first user. Of course, the second user may perform an action other than mouseover to request a call. As a result, user terminal 30 detects the action of the second user requesting a call to the first user.
[0087] The user terminal 30 requests the server 20 to obtain the icon image of the first user to be called (S20). The user information acquisition unit 111 of the information acquisition unit 110 of the server 20 then acquires the icon image and outputs it to the speech synthesis unit 300 of the server 20 (S21).
[0088] Furthermore, the user terminal 30 requests the acquisition of information necessary for text generation, such as time information, weather information, or organization information (S30). Then, the time information acquisition unit 112, weather information acquisition unit 113, and organization information acquisition unit 114 of the information acquisition unit 110 of the server 20 acquire the time information, weather information, or organization information and output it to the text generation unit 200 of the server 20 (S31). The text generation unit 200 of the server 20 then generates text based on the time information, weather information, or organization information and outputs it to the speech synthesis unit 300 of the server 20 (S32).
[0089] The user terminal 30 requests the acquisition of work information or biometric information necessary for estimating fatigue level (S40). The work information acquisition unit 115 and biometric information acquisition unit 116 of the server 20's information acquisition unit 110 acquire the work information or biometric information and output it to the fatigue level estimation unit 400 (S41). The fatigue level estimation unit 400 of the server 20 estimates the fatigue level based on the work information or biometric information and outputs it to the speech synthesis unit 300 of the server 20 (S42).
[0090] The speech synthesis unit 300 of the server 20 performs speech synthesis processing based on the icon image, text, and fatigue level, and transmits the speech data to the user terminal 30 (S51). The user terminal 30 plays the synthesized speech based on the speech data to the second user (S52). In other words, the user terminal 30 plays the synthesized speech based on the speech data received from the server 20 through a speaker or earphones, etc. This allows the second user to hear synthesized speech that mimics the voice of the first user.
[0091] In this way, the second user can hear a synthesized voice that mimics the first user's voice before speaking to them. Therefore, the second user can perceive the first user's demeanor through the synthesized voice and content, thus lowering the barrier to initiating conversation in the virtual office. Furthermore, because the second user can communicate more easily, they can feel a sense of realism in their interactions even in a virtual space.
[0092] Figure 6 is a flowchart showing the first embodiment of the virtual space control method in the second embodiment. First, the detection unit 121 of the server 20' determines whether or not it has detected an operation to request a call from the second user to the first user (S101). If the detection unit 121 of the server 20' has not detected an operation to request a call (NO in S101), it continues processing S101 until it detects one. If the detection unit 121 of the server 20' has detected an operation to request a call (YES in S101), the call information acquisition unit 117 of the server 20' acquires call information regarding the call between the first user and the second user (S102). Here, the call information acquisition unit 117 of the server 20' acquires only the number of calls as call information.
[0093] Next, the intimacy calculation unit 122 of server 20' determines whether the number of calls is equal to or greater than a predetermined value (S103). In other words, the intimacy calculation unit 122 of server 20' determines whether the relationship between the first user and the second user is intimate. Here, the intimacy calculation unit 122 of server 20' determines the intimacy using only the number of calls, but the intimacy may also be determined by combining the number of calls, call frequency, call duration, etc.
[0094] If the number of calls is not equal to or greater than a predetermined value (NO in S103), the output control unit 130 of the server 20' generates a simplified icon image or a response audio (S104). If the number of calls is equal to or greater than a predetermined value (YES in S103), the output control unit 130 of the server 20' generates a detailed icon image or a response audio (S105). The output control unit 130 of the server 20' transmits the icon image or response audio to the user terminal 30 (S106). The user terminal 30, having received the icon image or response audio, updates and displays the icon image in window 1. Alternatively, the user terminal 30 plays the response audio from a speaker or the like. In this way, the process ends.
[0095] Figure 7 is a flowchart showing a second embodiment of the virtual space control method in the second embodiment. First, the detection unit 121 of the server 20' determines whether or not it has detected an operation to request a call from the second user to the first user (S201). If the detection unit 121 of the server 20' has not detected an operation to request a call (NO in S201), it continues processing S201 until it detects one. If the detection unit 121 of the server 20' detects a user operation that constitutes a call request operation (YES in S201), the fatigue level estimation unit 400 of the server 20' estimates the fatigue level of the first user (S202).
[0096] Next, the call information acquisition unit 117 of server 20' acquires call information regarding past calls between the first user and the second user (S203). Here, the call information acquisition unit 117 of server 20' acquires only the number of calls as call information. The intimacy calculation unit 122 of server 20' determines whether the number of calls is greater than or equal to a predetermined value (S204). In other words, the intimacy calculation unit 122 of server 20' determines whether the relationship between the first user and the second user is intimate. Here, the intimacy calculation unit 122 of server 20' determines the intimacy using only the number of calls, but the intimacy may also be determined by combining the number of calls, call frequency, call duration, etc.
[0097] If the number of calls is not equal to or greater than a predetermined value (NO in S204), the output control unit 130 of the server 20' generates an icon image that shows the fatigue level in stages (S205). Here, as described above, the fatigue level is shown in four stages. If the number of calls is equal to or greater than a predetermined value (YES in S204), the output control unit 130 of the server 20' generates an icon image that includes the fatigue level value (S206). The output control unit 130 of the server 20' sends the icon image to the user terminal 30 (S207). Therefore, the user terminal 30 displays the icon image. In this way, the process ends.
[0098] This allows the second user to know the first user's fatigue level before speaking to them. Therefore, the second user can understand the first user's physical condition and prevent unnecessary calls. This brings a sense of realism to interactions even in a virtual space, allowing things to proceed more smoothly. Furthermore, the second user can easily request a call from the first user. This increases opportunities for users to communicate with each other, thus facilitating smoother relationships within the organization.
[0099] Some or all of the above processing on the server 20 or user terminal 30 may be executed by a computer program. The above-described program can be stored and supplied to the computer using various types of non-transitory computer-readable medium. Non-transitory computer-readable medium includes various types of tangible storage medium. Examples of non-transitory computer-readable medium include magnetic recording media (e.g., flexible disks, magnetic tapes, hard disk drives), magneto-optical recording media (e.g., magneto-optical disks), CD-ROMs (Read Only Memory), CD-Rs, CD-R / Ws, and semiconductor memory (e.g., mask ROMs, PROMs (Programmable ROMs), EPROMs (Erasable PROMs), flash ROMs, RAMs (Random Access Memory)). The program may also be supplied to the computer using various types of transient computer-readable medium. Examples of transient computer-readable medium include electrical signals, optical signals, and electromagnetic waves. Temporary computer-readable media can supply programs to a computer via wired communication channels such as electric wires and optical fibers, or via wireless communication channels.
[0100] Some of the above processes may be performed in a distributed manner across multiple devices. For example, the storage device that stores various information may be a cloud server. One user may use two or more user terminals 30. For example, one user may display a virtual office on a PC and play audio on a smartphone. Furthermore, this can also be applied to processes in which three or more users make calls simultaneously. Some of the processes in the above embodiment may be omitted.
[0101] Although the present invention has been specifically described above based on embodiments, it goes without saying that the present invention is not limited to the above embodiments and can be modified in various ways without departing from its essence. [Explanation of symbols]
[0102] 1 window 2. Virtual Office 3 icons 4 pointers 6 seats 10. Virtual Space Control System 20, 20' Server 30 User terminals 110, 110' Information acquisition section 111 User Information Acquisition Unit 112-hour information acquisition unit 113 Weather Information Acquisition Department 114 Organization Information Acquisition Department 115 Business Information Acquisition Department 116 Biological Information Acquisition Unit 117 Call information acquisition section 121 Detection unit 122 Intimacy calculation part 130 Output control unit 200 Text generation unit 300 Speech Synthesis Unit 400 Fatigue level estimation unit 500 Icon Image Generation Unit U User
Claims
1. A virtual space control system that displays icon images corresponding to users using the virtual space on the virtual space, A detection unit that detects an operation to request a call from the second user to the first user, A call information acquisition unit that acquires call information between the first user and the second user, The system includes an output control unit that, upon detecting an operation to request the aforementioned call, controls the display output of the first user's icon image and the audio output of the first user based on the call information, thereby changing at least one of these. Virtual space control system.
2. The system further includes a closeness calculation unit that calculates the closeness between the first user and the second user based on the call history between the first user and the second user, The output control unit, The virtual space control system according to claim 1, which changes at least one of the display mode of the icon image and the audio output according to the intimacy level.
3. An icon image generation unit that generates an icon image corresponding to the first user, The system further comprises a fatigue level estimation unit for estimating the fatigue level of the first user, If the aforementioned level of intimacy is less than a predetermined level, The output control unit, The virtual space control system according to claim 2, wherein the icon image generation unit is controlled to generate the icon image of the first user in an output manner in which the fatigue level is divided into a predetermined number of stages.
4. A speech synthesis unit that synthesizes speech corresponding to the first user, The system further comprises a fatigue level estimation unit for estimating the fatigue level of the first user, The output control unit, The virtual space control system according to claim 1, wherein the voice synthesis unit is controlled to synthesize the voice of the first user in an output manner corresponding to the fatigue level.
5. A virtual space control method that displays an icon image corresponding to a user using the virtual space, A step of detecting an operation to request a call from the second user to the first user, The steps include obtaining call information between the first user and the second user, A virtual space control method that includes the step of, when an operation to request the aforementioned call is detected, performing control to change at least one of the display output of the first user's icon image and the audio output of the first user based on the call information.
6. A program that causes a computer to execute a virtual space control method that displays icon images corresponding to users using the virtual space, A step of detecting an operation to request a call from the second user to the first user, The steps include obtaining call information between the first user and the second user, A program that causes a computer to perform the following steps when it detects an operation to request the aforementioned call: to control the display output of the first user's icon image and the audio output of the first user based on the call information.
Citation Information
Patent Citations
Speech synthesis device, speech synthesis program, and speech synthesis method
JP2021099454A
Information processing device and program
JP2024031550A