Program, method and information processing device
Patent Information
- Application Number
- JP2022171380
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2022-10-26
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2042-06-29
AI Technical Summary
Existing technologies for reflecting facial expressions on avatars in real-time, such as changing the shape of an avatar's mouth based on voice frequencies, often result in unnatural movements that can cause discomfort to viewers due to immediate and abrupt changes.
A system that synchronizes avatar mouth movements with user utterances by acquiring an audio spectrum, determining an avatar corresponding to the speaker, and adjusting the mouth aspect to match the speaker's utterances at a slower pace than the actual changes, using a combination of audio and facial movement tracking to ensure natural-looking avatar movements.
The system enhances the natural appearance of avatar mouth movements, reducing viewer discomfort and providing a more immersive experience by smoothing the transitions between mouth shapes in response to speech.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] The present disclosure relates to a program, a method, and an information processing device. [Background technology]
[0002] There is known a technique for reflecting a user's facial expression, etc. in an avatar in real time.
[0003] Patent document 1 describes a technology that estimates which Japanese vowel (aiueo) is being spoken from the combination of the first and second formant frequencies of a person's voice, and changes the shape of an avatar's mouth to correspond to each vowel. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] JP 2016-126500 A Summary of the Invention [Problem to be solved by the invention]
[0005] The technology disclosed in Patent Document 1 estimates which Japanese vowel is being spoken based on the voice picked up from a microphone, and determines and changes the shape and size of an avatar's mouth. However, the technology in Patent Document 1 merely changes the manner in which an avatar's mouth appears in response to voice. For example, when the mouth is changed continuously, the mouth movement is different from the way a person speaks, and instantly changes to a manner that corresponds to a vowel, which may cause discomfort to the viewer / user. Therefore, there is a need for technology that can make changes in an avatar's mouth appear more natural. [Means for solving the problem]
[0006] According to one embodiment, there is provided a program executed by a computer having a processor, the program causing the processor to execute the steps of: acquiring an audio spectrum of a performer's speech; changing the mouth appearance of an avatar corresponding to the performer in accordance with the performer's speech based on the acquired audio spectrum; presenting the avatar corresponding to the performer and the performer's voice to an audience; and accepting a setting for the degree to which the mouth appearance of the avatar is changed in accordance with the performer's speech to be lower than the change in the performer's speech, wherein in the changing step, the mouth appearance of the avatar is changed in accordance with the setting. Effect of the Invention
[0007] According to the present disclosure, it is possible to provide a technology that makes changes in the state of an avatar's mouth appear even more natural. [Brief description of the drawings]
[0008] [Figure 1] FIG. 2 is a block diagram showing the overall configuration of the system 1. [Diagram 2] FIG. 2 is a diagram showing a functional configuration of a terminal device 10. [Diagram 3] FIG. 2 is a diagram showing the functional configuration of a server 20. [Figure 4] 1 shows data structures of a user information database (DB), an avatar information DB, and a wearable device information DB stored in the storage unit of server 20. [Diagram 5] 11 is a flowchart showing a series of processes for acquiring a voice spectrum of a user's speech and changing the state of the mouth of an avatar corresponding to the user in accordance with the speech of the performer based on the acquired voice spectrum. [Figure 6] 13 is an example of a screen when a user registers the voice spectrum of his / her own vowel in the system 1. [Figure 7] 13 shows an example of a screen when a user sets the degree of change in the appearance of an avatar's mouth or facial parts. [Figure 8]An example screen is shown in which one or more candidate emotions of a user are estimated from the user's speech, and the appearance of an avatar is changed based on the estimated one or more emotions of the user. [Figure 9] 13 shows an example of a screen where a user can configure various settings based on voice spectrum, etc. for an avatar with attributes different from humans. [Figure 10] A flowchart showing a series of processes involved in sensing the movement of one or more facial parts of a user and changing the appearance of one or more facial parts of an avatar corresponding to the user based on the sensed movement of the one or more facial parts. [Figure 11] 13 shows an example screen when sensing the movement of one or more facial parts of a user, and changing the appearance of one or more facial parts of a corresponding avatar based on the sensed movement of the one or more facial parts. [Figure 12] 13 shows an example screen when one or more emotion candidates of a user are estimated, and the appearance of one or more facial parts of a corresponding avatar is changed based on the emotion selected by the user. [Figure 13] 13 shows an example of a screen for setting the degree of change in the avatar's appearance when sensing results cannot be obtained for at least one associated part out of one or more parts of the user's face. [Figure 14] 13 shows an example of a screen when correcting the degree of change in the appearance of an avatar when a user is wearing a wearable device such as glasses. [Figure 15] 13 shows an example of a screen in which the state of the avatar's mouth is changed based on the degree of change in speech when the movement of the user's mouth cannot be sensed. [Figure 16] 13 shows an example screen that is displayed when a specific notification is presented to the user when the difference in degree settings between one or more pre-associated facial features of an avatar exceeds a specific threshold. [Figure 17]13 shows an example of a screen that shows the state in which at least one or more facial parts change when the degree difference is set within a predetermined range when a predetermined notification is presented to the user. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0009] Hereinafter, an embodiment of the present invention will be described with reference to the drawings. In the following description, the same components are denoted by the same reference numerals. Their names and functions are also the same. Therefore, detailed description thereof will not be repeated.
[0010] <First embodiment> <Summary> In the following embodiment, a technique for changing the state of the mouth of an avatar based on the voice spectrum of a user who is an actor operating the avatar will be described. Here, there is no limitation on the devices etc. that are used as appropriate when realizing the technology disclosed herein, and it may be a terminal device such as a smartphone or tablet terminal owned by the user, or it may be presented from a stationary PC (Personal Computer).
[0011] There is a known technology that controls the mouth movement of an avatar based on the user's voice acquired through a sound collection device such as a microphone. However, in this system, the user's actual mouth movement and the avatar's movement are not accurately synchronized, which may cause viewers to feel uncomfortable.
[0012] Therefore, system 1 provides a technology that makes changes in the avatar's mouth shape appear more natural.
[0013] System 1 can be used, for example, in situations such as live streaming distribution using an avatar that tracks the movements of a user (performer) on a video distribution site. For example, system 1 tracks the movements of a user via a camera (imaging device) provided on a terminal device (PC, etc.) used by the user, and reflects the movements of the user in the movements of an avatar. System 1 also acquires the audio spectrum of the performer's speech via a microphone (sound collection device) also provided on the user's terminal device, and changes the mouth shape of the avatar corresponding to the performer according to the performer's speech based on the acquired audio spectrum. At this time, the system 1 presents an avatar corresponding to the performer and the performer's voice to the audience, and accepts a setting for the degree to which the avatar's mouth shape is changed in response to the performer's speech to be lower than the degree of change in the performer's speech. By executing this process, the system 1 can change the avatar's mouth shape in response to the performer's speech. This allows the changes in the avatar's mouth appearance to appear even more natural.
[0014] <1 Overall system configuration> FIG. 1 shows the overall configuration of a system 1 according to the first embodiment.
[0015] As shown in Fig. 1, the system 1 includes a plurality of terminal devices (terminal device 10A and terminal device 10B are shown in Fig. 1. Hereinafter, they may be collectively referred to as "terminal device 10". Furthermore, a plurality of terminal devices 10C and the like may be further included in the configuration) and a server 20. The terminal devices 10 and the server 20 are connected for communication via a network 80.
[0016] The terminal device 10 is a device operated by each user. The terminal device 10 is realized by a mobile terminal such as a smartphone or tablet compatible with a mobile communication system. In addition, the terminal device 10 may be, for example, a stationary PC (Personal Computer) or a laptop PC. As shown as the terminal device 10B in FIG. 1, the terminal device 10 includes a communication IF (Interface) 12, an input device 13, an output device 14, a memory 15, a storage unit 16, and a processor 19. The server 20 includes a communication IF 22, an input / output IF 23, a memory 25, a storage 26, and a processor 29.
[0017] The terminal device 10 is communicatively connected to the server 20 via a network 80. The terminal device 10 is connected to the network 80 by communicating with communication devices such as a wireless base station 81 conforming to communication standards such as 5G and LTE (Long Term Evolution) and a wireless LAN router 82 conforming to wireless LAN (Local Area Network) standards such as IEEE (Institute of Electrical and Electronics Engineers) 802.11.
[0018] The communication IF 12 is an interface for inputting and outputting signals so that the terminal device 10 can communicate with an external device. The input device 13 is an input device (e.g., a touch panel, a touch pad, a pointing device such as a mouse, a keyboard, etc.) for receiving an input operation from a user. The output device 14 is an output device (a display, a speaker, etc.) for presenting information to a user. The memory 15 is for temporarily storing programs and data processed by the programs, etc., and is a volatile memory such as a DRAM (Dynamic Random Access Memory). The storage unit 16 is a storage device for saving data, and is, for example, a flash memory or a HDD (Hard Disc Drive). The processor 19 is hardware for executing an instruction set described in a program, and is composed of an arithmetic unit, a register, a peripheral circuit, etc.
[0019] The server 20 manages information set by a user when performing live streaming using an avatar, etc. The server 20 stores, for example, information about the user, information about the avatar, and information about a wearable device worn by the user.
[0020] The communication IF 22 is an interface for inputting and outputting signals so that the server 20 can communicate with an external device. The input / output IF 23 functions as an interface with an input device for receiving input operations from a user and an output device for presenting information to a user. The memory 25 is for temporarily storing programs and data processed by the programs, etc., and is a volatile memory such as a DRAM (Dynamic Random Access Memory). The storage 26 is a storage device for saving data, such as a flash memory or a HDD (Hard Disc Drive). The processor 29 is hardware for executing an instruction set described in a program, and is composed of an arithmetic unit, a register, a peripheral circuit, etc.
[0021] In this embodiment, each device (terminal device, server, etc.) can also be regarded as an information processing device. That is, a collection of the devices can be regarded as one "information processing device", and the system 1 may be formed as a collection of multiple devices. The method of allocating multiple functions required to realize the system 1 according to this embodiment to one or multiple pieces of hardware can be appropriately determined in consideration of the processing capacity of each piece of hardware and / or the specifications required for the system 1.
[0022] <1.1 Configuration of the terminal device 10> FIG. 2 is a block diagram of the terminal device 10 constituting the system 1 of the first embodiment. As shown in FIG. 2, the terminal device 10 includes a plurality of antennas (antenna 111, antenna 112), wireless communication units (first wireless communication unit 121, second wireless communication unit 122) corresponding to the respective antennas, an operation reception unit 130 (including a touch-sensitive device 1301 and a display 1302), a voice processing unit 140, a microphone 141, a speaker 142, a position information sensor 150, a camera 160, a motion sensor 170, a storage unit 180, and a control unit 190. The terminal device 10 also has functions and configurations (for example, a battery for holding power, a power supply circuit for controlling the supply of power from the battery to each circuit, etc.) that are not particularly shown in FIG. 2. As shown in FIG. 2, each block included in the terminal device 10 is electrically connected by a bus or the like.
[0023] The antenna 111 emits a signal generated by the terminal device 10 as a radio wave. The antenna 111 also receives a radio wave from space and provides the received signal to the first wireless communication unit 121.
[0024] The antenna 112 radiates a signal generated by the terminal device 10 as a radio wave. The antenna 112 also receives the radio wave from space and provides the received signal to the second radio communication unit 122.
[0025] The first wireless communication unit 121 performs modulation and demodulation processing for transmitting and receiving signals via the antenna 111 so that the terminal device 10 can communicate with other wireless devices. The second wireless communication unit 122 performs modulation and demodulation processing for transmitting and receiving signals via the antenna 112 so that the terminal device 10 can communicate with other wireless devices. The first wireless communication unit 121 and the second wireless communication unit 122 are communication modules including a tuner, a Received Signal Strength Indicator (RSSI) calculation circuit, a Cyclic Redundancy Check (CRC) calculation circuit, a high-frequency circuit, and the like. The first wireless communication unit 121 and the second wireless communication unit 122 perform modulation and demodulation and frequency conversion of wireless signals transmitted and received by the terminal device 10, and provide the received signals to the control unit 190.
[0026] The operation reception unit 130 has a mechanism for receiving an input operation by a user. Specifically, the operation reception unit 130 is configured as a touch screen, and includes a touch-sensitive device 1301 and a display 1302. The touch-sensitive device 1301 receives an input operation by a user of the terminal device 10. The touch-sensitive device 1301 detects a touch position of the user on the touch panel, for example, by using a capacitive touch panel. The touch-sensitive device 1301 outputs a signal indicating the touch position of the user detected by the touch panel to the control unit 190 as an input operation. The terminal device 10 may also include a keyboard (not shown) that can be physically inputted, and may receive an input operation by the user via the keyboard.
[0027] The display 1302 displays data such as images, videos, and text under the control of the control unit 190. The display 1302 is realized by, for example, an LCD (Liquid Crystal Display) or an organic EL (Electro-Luminescence) display.
[0028] The audio processing unit 140 modulates and demodulates an audio signal. The audio processing unit 140 modulates a signal provided from the microphone 141 and provides the modulated signal to the control unit 190. The audio processing unit 140 also provides the audio signal to the speaker 142. The audio processing unit 140 is realized by, for example, a processor for audio processing. The microphone 141 accepts audio input and provides an audio signal corresponding to the audio input to the audio processing unit 140. The speaker 142 converts the audio signal provided from the audio processing unit 140 into audio and outputs the audio to the outside of the terminal device 10.
[0029] The position information sensor 150 is a sensor that detects the position of the terminal device 10, and is, for example, a GPS (Global Positioning System) module. The GPS module is a receiving device used in a satellite positioning system. In the satellite positioning system, signals are received from at least three or four satellites, and the current position of the terminal device 10 equipped with the GPS module is detected based on the received signals. The position information sensor 150 may be a transmitting / receiving device based on a communication standard used in a short-range communication system between information devices. Specifically, the position information sensor 150 uses the 2.4 GHz band, such as a Bluetooth (registered trademark) module, to receive a beacon signal from another information device equipped with a Bluetooth (registered trademark) module.
[0030] The camera 160 is a device that receives light with a light receiving element and outputs the received light as a captured image. The camera 160 is, for example, a depth camera that can detect the distance from the camera 160 to a subject being photographed. Furthermore, the camera 160 acquires the body movements of the user who uses the terminal device 10. Specifically, for example, the camera 160 acquires the movements of the user's mouth and each part of the face (eyes, eyebrows, etc.). The acquisition of the movements may utilize any existing technology.
[0031] The motion sensor 170 is configured with a gyro sensor, an acceleration sensor, etc., and detects the inclination of the terminal device 10.
[0032] The storage unit 180 is configured, for example, with a flash memory or the like, and stores data and programs used by the terminal device 10. In one aspect, the storage unit 180 stores user information 1801, avatar information 1802, wearable device information 1803, and the like. The information may be held in the storage unit 180 of the terminal device 10, or may be stored as a database in a storage unit 202 of a server described later and acquired via the network 80.
[0033] User information 1801 is information such as an ID for identifying a user, a user name, and information on an avatar corresponding to the user. Here, the user refers to a performer who moves an avatar based on information acquired via microphone 141 or camera 160. Details of the information included in the user information will be described later.
[0034] The avatar information 1802 is various information related to an avatar corresponding to a user. The avatar information 1802 holds information such as the corresponding user and settings that the user normally uses, and is information that is referenced by the user to smoothly operate the avatar in distribution such as live streaming. Details of the information included in the avatar information will be described later. The settings that the user normally uses are parameters and conditions that the user can adjust when broadcasting using an avatar, such as the basic settings for the degree of change in the avatar's appearance, the emotions displayed as default in normal broadcasts, the user's sensing sensitivity, etc.
[0035] The wearable device information 1803 is various information related to the wearable device that the user is wearing at the time of distribution. The various information includes, for example, the following: Types of wearable devices - Size of wearable device -Transmittance of wearable devices Possibility of obtaining information electronically The wearable device information 1803 holds various information related to various instruments and devices, such as glasses, smart glasses, and other eyewear worn by the user, and a head mounted display (HMD), etc. Details of the information included in the wearable device information 1803 will be described later.
[0036] The control unit 190 reads a program stored in the storage unit 180 and executes instructions included in the program to control the operation of the terminal device 10. The control unit 190 is, for example, an application processor. The control unit 190 operates according to the program to fulfill the functions of an input operation reception unit 1901, a transmission / reception unit 1902, a data processing unit 1903, and a notification control unit 1904.
[0037] The input operation reception unit 1901 performs processing for receiving a user's input operation on an input device such as the touch-sensitive device 131. The input operation reception unit 1901 determines the type of operation, such as whether the user's operation is a flick operation, a tap operation, or a drag (swipe) operation, based on information on the coordinates where the user touches the touch-sensitive device 1301 with a finger or the like.
[0038] The transmitting / receiving unit 1902 performs processing for the terminal device 10 to transmit and receive data to and from an external device such as the server 20 in accordance with a communication protocol.
[0039] The data processing unit 1903 performs calculations on the data that the terminal device 10 accepts as input in accordance with a program, and outputs the calculation results to a memory or the like.
[0040] The data processing unit 1903 receives the movement of the user's mouth etc. acquired by the camera 160, and controls the processing for executing various processes. For example, the data processing unit 1903 executes the processing for controlling the movement of the mouth of an avatar corresponding to the user, based on the movement of the user's mouth acquired by the camera 160.
[0041] The notification control unit 1904 performs processing for displaying a display image on the display 132, processing for outputting sound from the speaker 142, and processing for causing the camera 160 to generate vibrations.
[0042] <1.2 Functional configuration of server 20> 3 is a diagram showing a functional configuration of the server 20. As shown in FIG. 3, the server 20 fulfills the functions of a communication unit 201, a storage unit 202, and a control unit 203.
[0043] The communication unit 201 performs processing for the server 20 to communicate with external devices.
[0044] The storage unit 202 stores data and programs used by the server 20. The storage unit 202 stores a user information database 2021, an avatar information database 2022, a wearable device information database 2023, and the like.
[0045] The user information database 2021 is a database for storing various information related to the performers who operate the avatars. Details of each record stored in the database will be described later.
[0046] The avatar information database 2022 is a database for holding various information related to the avatars operated by the users, as will be described in detail later.
[0047] The wearable device information database 2023 is a database for storing various information related to the eyewear worn by the user who operates the avatar. Details will be described later.
[0048] The control unit 203 is composed of, for example, a processor 29, which performs processing in accordance with a program to provide functions such as various modules including a reception control module 2031, a transmission control module 2032, a user information acquisition module 2033, an avatar information acquisition module 2034, a voice spectrum acquisition module 2035, an avatar change module 2036, an avatar presentation module 2037, a setting reception module 2038, a wearable device information acquisition module 2039, and a change correction module 2040.
[0049] The reception control module 2031 controls the process in which the server 20 receives a signal from an external device in accordance with a communication protocol.
[0050] The transmission control module 2032 controls the process in which the server 20 transmits signals to external devices in accordance with a communication protocol.
[0051] The user information acquisition module 2033 controls the process of acquiring various information about the user who is the performer operating the avatar. The various information includes, for example, the following: User name, identification ID - Avatar information corresponding to the user Devices worn by the user (glasses, etc.) Specifically, for example, the user information acquisition module 2033 may acquire the information by referring to the user information 1801 from the storage unit 180 of the terminal device 10 used by the user. Also, the user information acquisition module 2033 may acquire the information by referring to a user information database 2021 held in the storage unit 202 of the server 20, which will be described later. In addition, the user information acquisition module 2033 may acquire the information by accepting input of various pieces of information related to the user directly from the user.
[0052] The avatar information acquisition module 2034 acquires various information about the avatar operated by the user. The various information includes, for example, the following: -ID information to identify the avatar - Information on avatar attributes (human, non-human, etc.) - Information about the mouth, face, or other body part appearances set by the corresponding user as default Dedicated settings for the mouth, face or other body parts, each individually configured for each avatar Specifically, for example, the avatar information acquisition module 2034 may acquire information on the avatar associated with each user by referring to the avatar information 1802 or the user information database 2021. In addition, in a certain aspect, server 20 may receive from a user an input of a setting for the degree of change in the appearance of the avatar's mouth, face, or other body parts, and upon receiving an operation to reset the setting as a default, perform a process of updating the avatar information held in avatar information 1802, etc. This allows the user to appropriately set a frequently used setting among the degrees of change in the appearance of the avatar as the default, making it easy to operate the avatar.
[0053] Note that the number of pieces of avatar information for each user does not need to be one. For example, a plurality of pieces of avatar information may be associated with a user in advance, or additional pieces of avatar information may be associated with the user.
[0054] In addition, in one aspect, the avatar ID may include the following information: Information about the appearance of your avatar (such as gender, eye color, hairstyle, mouth, facial features, or other body features, hair color, skin color, etc.) Information about the state of the avatar's mouth, face, or other body parts (such as the number of types of changes in state, the amount of change in state, etc.) Here, the server 20 may store the avatar information in association with the type of content. Specifically, for example, the server 20 may associate the type of content (chat, song, acting, etc.) provided by the user in live distribution or the like with settings of the degree of change in the appearance of the mouth, face, and other body parts, and when the server 20 receives an operation by the user to select which content to provide, the server 20 may reflect the setting corresponding to the content in the avatar.
[0055] This allows the user to appropriately change the appearance of the avatar in accordance with the content provided to the viewer, thereby providing the viewer with a sense of immersion.
[0056] Furthermore, after accepting the selection of an avatar to be used from the user, the server 20 may accept the selection from the user and simultaneously present a predetermined notification (dialog or the like) to the user, rather than automatically reflecting the setting of the degree of change in appearance. For example, after the user selects an avatar to be used, the server 20 may present a notification such as "Is the normal setting OK?" when the user selects a general-purpose setting that is normally used, rather than a dedicated setting that is set for each avatar.
[0057] This prevents the user from reflecting incorrect settings in live distribution, thereby preventing the viewer's sense of immersion from being diminished.
[0058] Furthermore, the server 20 may accept and reflect the setting of the degree of change in appearance during live distribution. Specifically, for example, when the server 20 accepts a change in the setting of the degree of change in appearance from a user during live distribution, the server 20 may execute a process of changing the appearance of the avatar based on the changed setting when changing the appearance of the avatar based on the voice spectrum of the user acquired after a predetermined time has elapsed after accepting the change in setting or the sensing result of the user. This allows the user to change and reflect the settings for the appearance change as appropriate during live streaming, so that even if the content provided during live streaming is switched, the user can see the change in the avatar's appearance without feeling uncomfortable.
[0059] The voice spectrum acquisition module 2035 controls a process of acquiring a voice spectrum of a user's speech. Specifically, for example, the voice spectrum acquisition module 2035 controls a process of acquiring a voice spectrum from a voice uttered by the user acquired via the microphone 141. For example, the voice spectrum acquisition module 2035 acquires the user's voice via the microphone 141 and acquires the voice spectrum contained in the voice. For example, the voice spectrum acquisition module 2035 may perform a Fourier transform on the voice acquired from the microphone 141 to acquire information on the voice spectrum contained in the voice. At this time, the calculation for acquiring the voice spectrum is not limited to a Fourier transform, and may be any existing method. In addition, in a certain aspect, the voice spectrum acquisition module 2035 may acquire information on the voice spectrum of a vowel from the voice of the user. For example, the voice spectrum acquisition module 2035 accepts a vowel setting input by the user in advance, and then accepts an utterance from the user via the microphone 141, and stores the accepted vowel setting and the acquired voice spectrum in association with each other. In addition, in one aspect, the voice spectrum acquisition module 2035 may acquire information on sounds resulting from consonants, such as "t," "c," "h," "k," "m," "r," "s," "n," and "w," and combine it with the stored vowel information to estimate the words spoken by the user. This allows the system 1 to separately characterize and store the voice spectrum relating to vowels among the user's voice spectrum, thereby making it possible to change the movement of the mouth shape of the avatar more accurately.
[0060] The avatar change module 2036 controls a process of changing the state of the mouth of an avatar corresponding to a performer according to the speech of the performer based on the acquired voice spectrum. Specifically, for example, the avatar change module 2036 estimates the words spoken by the user from the user's voice spectrum acquired by the voice spectrum acquisition module 2035, and changes the state of the mouth of the avatar according to the estimated words. For example, the avatar change module 2036 estimates information on a vowel spoken by the user from the user's voice spectrum, and changes the state of the mouth according to the vowel. For example, when the user's voice spectrum acquired by the voice spectrum acquisition module 2035 is "a", the avatar change module 2036 changes the state of the mouth of the avatar to a shape corresponding to "a".
[0061] The avatar presentation module 2037 controls a process of presenting an avatar corresponding to a performer and a user's voice to a viewer. Specifically, for example, the avatar presentation module 2037 transmits an image of an avatar corresponding to a user and the user's voice to the display 1302 and the speaker 142 of the terminal device 10 used by the viewer, and presents them to the viewer. At this time, the viewer is not limited to one person, and the avatar and voice may be presented to the terminal device 10 of multiple viewers.
[0062] The setting reception module 2038 controls a process of receiving a setting of the degree to which the avatar's mouth state is changed in response to the user's speech so as to be lower than the change in the user's speech. Specifically, for example, the setting reception module 2038 receives from the user a setting of a time interval for reflecting the user's speech in the avatar's mouth state as the setting of the degree to which the avatar's mouth state is changed. For example, the setting reception module 2038 may receive settings including the following: - Setting the time required for the avatar's mouth to change to a state that corresponds to the speech sound estimated from the user's voice spectrum · Change the avatar's behavior based on the user's speech over a certain period of time · Set the update frequency (e.g. number of updates per second) Here, the change in the user's speech will be defined. The change in the user's speech is, for example, the speed of the user's speech, and may be calculated based on the following. The time interval when the vowel spoken by the user changes (for example, the time interval when the vowel changes from "a" to "i") At this time, the server 20 may simultaneously acquire sounds derived from consonants (c, k, etc.), and even if the same vowel is acquired consecutively, estimate the speech speed as if different words were being spoken. - Number of vowels uttered in a given period At this time, the setting reception module 2038 may receive the setting to a degree lower than the degree of change of the avatar estimated from the change of the user's speech. For example, the setting reception module 2038 may receive the time required to change the state of the mouth to correspond to the speech (vowel) estimated in advance from the user's voice spectrum. The server 20 acquires information on the time when the user uttered the vowel from the user's voice spectrum based on the received required time, calculates the ratio to the previously set required time, multiplies it by the amount of change in the state, and calculates the amount of change in the state of the avatar's mouth. The server 20 changes the state of the mouth based on the acquired speech time and the amount of change. For example, it is assumed that the user inputs a setting that the avatar's mouth changes to the state of "a" in a required time of "1 second". For example, the time when the state of "a" is completely changed is set to "100", and it is set to "100" in "1 second". At this time, the server 20 may also receive and record from the user the degree of change in the behavior in one second (amount of change in the mouth, speed). (That is, when the behavior of the mouth changes in one second, the difference in the amount of change in the behavior between the first 0.5 seconds and the remaining 0.5 seconds may be set.) When the user utters the sound "a" for one second, the server 20 changes the state of the avatar's mouth to that of "a" over one second based on the above-mentioned setting of the amount of change, etc. However, when the user utters "a" for only "0.5 seconds", the server 20 may perform a process of changing the amount of change in the state of the avatar's mouth up to "50".
[0063] Furthermore, when the user speaks continuously (for example, "aiueo"), the server 20 may obtain the speaking time of each vowel and perform the above process. That is, the server 20 may calculate the amount of change in the state of the avatar's mouth corresponding to each vowel from the speaking time of each vowel, and change the state of the avatar's mouth. For example, if the time required for the state of the avatar's mouth to change to that corresponding to each vowel is "1 second," and "a" is spoken for "0.2 seconds" and "i" is spoken for "0.3 seconds," the amount of change corresponding to "a" is "20," and the amount of change corresponding to "i" is "30." Furthermore, when the user speaks for a time longer than the required time, the server 20 may maintain the state of the state of the avatar's mouth even after the required time has passed. This allows the user to change the avatar's mouth appearance gradually according to the time the user is speaking, rather than instantly changing the mouth appearance to "ah" when the user utters the sound "ah." In addition, by setting a required time and multiplying the change in the mouth appearance when the utterance does not reach that time, it is possible to prevent the avatar's mouth appearance from changing significantly even when the user speaks lightly (for example, even if the mouth is opened only about 30 degrees, the avatar's mouth appearance changes to 100). This allows the user to prevent the viewer from feeling uncomfortable due to the user's speech and the change in the avatar's mouth appearance, thereby giving the viewer a greater sense of immersion.
[0064] In a certain aspect, the server 20 may change the state of the mouth of the avatar based on information such as the size and pitch of the voice spectrum acquired from the user. Specifically, for example, the server 20 may acquire information on the frequency (Hz) and sound pressure (dB) of the user's voice spectrum, and may change the state of the mouth of the avatar when the information exceeds a predetermined threshold. For example, it is assumed that the server 20 accepts a setting to change the state of the mouth of the avatar in a required time of "1 second", and the user's speaking time is "1 second". At this time, when the server 20 detects that the user has spoken with a sound pressure exceeding the threshold at the time of "0.8 seconds", it may change the state of the mouth of the avatar to a larger state than usual (to make the mouth open widely). At this time, the server 20 may reflect the same setting not only on the mouth, but also on parts of the face and parts of the body. This allows the user to have the avatar's mouth movements reflect the user's sudden loud voice, allowing the viewer to see more natural avatar movements.
[0065] Additionally, the setting receiving module 2038 may receive a setting for the degree of change in the avatar's mouth movement so that the setting is lower than the update frequency of the avatar's movement estimated from the speech speed estimated from the user's speech. The server 20 then transmits the information set by the setting reception module 2038 to the avatar change module 2036, changes the shape of the avatar's mouth according to the settings, and then presents the avatar and the user's voice to the viewer by the avatar presentation module 2037. This allows the user to change the avatar's mouth movement more smoothly to match the user's own speech by changing the avatar's mouth movement more gradually than the change in vowels, which allows the user to prevent the avatar's mouth movement from being too sensitive and unnatural, thereby presenting more natural mouth movements to the viewer and enhancing the viewer's sense of immersion.
[0066] The wearable device information acquisition module 2039 controls the process of acquiring information on the wearable device worn by the user. Specifically, for example, when the wearable device information acquisition module 2039 acquires the user information, it refers to the wearable device information database 2023 described later and acquires the information on the wearable device worn by the user. The server 20 transmits the acquired wearable device information to the change correction module 2040.
[0067] The change correction module 2040 controls a process of correcting the setting of the degree of change in appearance to be reflected in the avatar, based on the information of the wearable device acquired by the wearable device information acquisition module 2039. Specifically, for example, the change correction module 2040 corrects the setting of the degree of change in appearance of parts of the user's face that are covered or blocked by the wearable device, from the information of the wearable device acquired by the wearable device information acquisition module 2039. The server 20 acquires the sensing result of a predetermined part of the user's face (mouth, eyes, eyebrows, nose, etc.), and changes the appearance of the avatar by reflecting the sensing result and the settings received from the user (degree to reflect the sensing result, parameter settings, etc.). At this time, for example, when the user is wearing glasses, the change correction module 2040 receives from the user the degree of change in the state of the eyes of the avatar corresponding to the user based on the information, and corrects and reflects the change based on a preset correction value. Correction refers to a process in which, for example, when the accuracy of sensing parts of the user's face decreases for each wearable device, the rate of decrease (or attenuation rate) is preset, and the degree to which the movement is reflected in the avatar during sensing and tracking is corrected based on the setting. This allows the user to present changes in the avatar's appearance to the viewer in a natural way, even if the user is wearing glasses or the like.
[0068] Additionally, the change correction module 2040 may correct the degree of change in the appearance of the avatar, depending on the attributes of the avatar acquired by the avatar information acquisition module 2034 . Specifically, for example, the change correction module 2040 may acquire information on whether the avatar operated by the user is a human or a non-human whose behavior changes differently from that of a human, and may execute a process of correcting the degree of change in the behavior of the avatar based on the information. For example, if the attribute of the avatar operated by the user is "dragon," the movement of the eyes, mouth, etc. may behave differently from that of a human. In that case, the change correction module 2040 may correct the amount of change in the corners of the mouth, the amount of change in the eyeballs, etc., to a shape that matches the avatar, based on the attribute of the "dragon." This allows users to present more natural movements to viewers based on the results of their own speech and facial sensing, even when they are operating an avatar that is not human.
[0069] In the embodiment of the present disclosure, the above configuration is not essential. That is, the terminal device 10 may play the role of the server 20 and execute the same processes as the various modules constituting the control unit 203 of the server 20. Furthermore, the terminal device 10 may execute the various functions disclosed in the present invention based on information acquired via the microphone 141, the camera 160, etc. provided in the terminal device 10, without going through the network 80.
[0070] <2 Data Structure> FIG. 4 is a diagram showing the data structures of the user information database 2021, the avatar information database 2022, and the wearable device information database 2023 stored in the server 20. As shown in FIG.
[0071] As shown in FIG. 4, the user information database 2021 includes an item "ID", an item "compatible avatar", an item "device used", an item "dedicated preset (mouth)", an item "dedicated preset (face)", an item "basic settings", an item "frequently used emotions", an item "notes", etc.
[0072] The item "ID" is information for identifying each user who is an actor operating an avatar.
[0073] The item "Corresponding Avatar" is information for identifying each avatar corresponding to each user.
[0074] The item "Device in use" is information that identifies a device worn by each user, for example, each wearable device worn by the user.
[0075] The item "dedicated preset (mouth)" is information indicating conditions previously set for each user regarding the degree to which the state of the avatar's mouth is changed when each user operates the avatar. Specifically, for example, it indicates various conditions individually set for each avatar when the avatar operated by the user is in a predetermined situation (for example, when the state of the mouth is greatly changed). Information included in the preset may include, for example, information on the height of the corners of the mouth, the shape of the lips, etc. The server 20 may reflect the setting in the avatar by accepting the selection of the preset from the user and present it to the viewer. This allows the user to instantly reflect changes in the mouth shape that are specific to the avatar corresponding to the user and present it to the viewer, allowing the viewer to see the avatar moving in a more natural way.
[0076] The item "dedicated preset (face)" is information indicating conditions previously set for each user regarding the degree of change in the state of the avatar's facial parts when each user operates the avatar. Specifically, for example, it indicates various conditions individually set for each avatar when the avatar operated by the user is in a predetermined situation (for example, when the avatar's facial expression changes significantly). Information included in the preset may include, for example, the direction of eyebrows, the shape of eyes, the size of pupils, flushing of cheeks, speech, or sensing of the user's facial expression. The server 20 may reflect the setting in the avatar by accepting the selection of the preset from the user and present it to the viewer. For example, suppose that the user uses an avatar with attributes other than human (monster, inorganic object, robot, etc.). In that case, the change in the state of the various parts of the avatar (mouth, face, body) may not completely match the user's voice spectrum and sensing results. For this reason, the server 20 may accept the setting of the dedicated preset (mouth) or dedicated preset (face) exemplified above from the user. By doing so, the user can select the preset during live distribution, etc., and the user's voice spectrum and sensing results can be reflected in the change in the state of the avatar without any sense of incongruity even when using any avatar.
[0077] The server 20 may also receive dedicated preset settings according to the type of content provided by the user. For example, when a user delivers a song, the server 20 may receive settings that change the mouth, face, and body parts of the avatar more significantly than when the user normally chats. This allows the user to instantly reflect changes in facial features specific to the avatar corresponding to the user and present it to the viewer, allowing the viewer to see the avatar moving in a more natural way.
[0078] The item "basic settings" indicates the setting of the degree of change that the user normally uses. Specifically, for example, it indicates the condition of the degree of change that the user who operates the avatar normally (generally) uses when changing the state of the mouth, face, and other body parts in normal distribution, live distribution, and live streaming. For example, the condition may include the degree of the degree of tracking the user's sensing results. The degree of tracking the sensing results includes, for example, the degree of sensitivity when the sensing results are directly reflected in the change in the state of the avatar, which is set to 100, the amount of change in the state of the avatar compared to the amount of change in the user's face, the degree of change to be reflected in the movement of the avatar relative to the amount of change in the avatar per unit time estimated from the sensing results, etc. This allows the user to easily start distribution without having to set the degree of change every time distribution is performed.
[0079] At this time, the server 20 may combine the basic settings and the dedicated presets and accept them as settings according to the content. Specifically, for example, the following are examples of avatar settings according to the content. ASMR (Autonomous Sensory Meridian Response) mode (whisper mode) Use the dedicated preset for the mouth (lower sensitivity, softer voice), and use the basic settings for facial expressions. Or use the dedicated facial expression settings in combination. Action game streaming mode A special preset is used for the mouth (which increases sensitivity and causes an over-reaction), while the sensitivity of facial expressions is also increased. Horror game streaming mode Use a dedicated preset for the mouth (lower sensitivity and lower frequency threshold for detection) and use the same sensitivity settings for facial expressions. Or use dedicated settings. Chat mode (uses basic settings)
[0080] In addition, in one aspect, server 20 may present a switching button to the user for switching the above modes, and accept the user's pressing of the button to cause the avatar to reflect the setting of the degree of change in appearance based on the mode. At this time, the server 20 may present the switching button to the user in a state that is invisible to the viewer but visible to the user. Also, the server 20 may change the arrangement of the switching button in response to an operation by the user. This allows users to use different presets depending on the content they are providing to viewers, enabling a wider range of expression.
[0081] The server 20 may also store information related to the background in the virtual space in association with the object, and may change the background in response to the switching of the mode. Additionally, the server 20 may store a specific object, as exemplified below, in association with the object, and display the object in the virtual space in response to the switching of the mode. -Microphones, instruments, and other equipment for live music streaming Game console objects during game distribution - General objects (houseplants, furniture in a room, etc.) This allows the server 20 to reduce the load of the reading process when switching modes, and prevents delays and other issues that may cause viewers to feel uncomfortable.
[0082] The above settings may be used in combination with basic settings, etc. The combination may be any setting received from the user, or may be stored in the storage unit as a dedicated combination for each user. The server 20 may also acquire information on the frequency of use of multiple presets. Based on the information on the frequency of use, the server 20 may present a notification to the user as to whether to store a frequently used preset as a "frequently used setting" or a "basic setting". When the server 20 receives an instruction from the user to set the preset as a "frequently used setting", etc., the server 20 may store the preset as a "frequently used setting" in the storage unit.
[0083] The item "frequently used emotion" indicates the setting of an emotion that the user frequently uses when operating the avatar. Specifically, for example, if the user frequently uses the emotion "happiness" during distribution, the server 20 may store in advance in the database conditions for changing the avatar's appearance based on that emotion. In this case, the conditions for changing the appearance include the degree of change in the mouth appearance, the amount of change in various parts of the face that move when expressing the emotion "happiness", the degree of tracking the sensing results, etc. When the server 20 receives a selection of the stored emotion setting from the user, the server 20 may change the appearance of the avatar based on the emotion and present it to the viewer. This allows the user to instantly set changes in the appearance of the avatar according to the emotions used in everyday broadcasts, making broadcasting easy.
[0084] The item "remarks" is information that is stored when there are special notes regarding the user information.
[0085] As shown in FIG. 4, the avatar information database 2022 includes an item "ID", an item "corresponding user", an item "attributes", an item "associated body part", an item "presence or absence of special body part", an item "motion setting of special body part", an item "standard change speed", an item "frequently used emotions", an item "notes", etc.
[0086] The item "ID" is information used in distribution to identify each avatar presented to viewers.
[0087] The item "Corresponding User" is information for identifying a user corresponding to the avatar.
[0088] The item "attribute" is information for identifying attributes set for each avatar. Specifically, the attribute indicates, for example, information for specifying whether the avatar is a human or a non-human whose appearance changes differently from that of a human. Attributes include, for example, the following information: ·human - Living things other than humans (animals, plants, etc.) Imaginary creatures (dragons, angels, demons, etc.) ·machine Formless entities (such as slimes and ghosts in fantasy) In some aspects, the record may hold information specific to the defined attribute as information of a subordinate concept. Specifically, for example, when the attribute is "inorganic matter", the record may hold a subordinate concept such as "no eyes", and when the attribute is "virtual creature", the record may hold information such as "has multiple eyes". The server 20 may hold information for correcting the degree of change in the appearance of the avatar based on the information of the attribute. This allows the user to appropriately change the state of the mouth and face even when operating a non-human avatar.
[0089] The item "associated parts" is information on associated parts among one or more facial parts of an avatar. Specifically, the associated parts may hold information on the "eyebrows" among the facial parts of an avatar that are associated with each other, for example. The server 20 may accept settings of the same degree of change for the associated parts. This eliminates the need for the user to set the degree of change in appearance for each associated body part individually, thereby reducing the time and effort required for setting the degree of change in appearance.
[0090] The item "presence or absence of special body part" is information for identifying whether or not the avatar has a special body part. Specifically, for example, when the avatar has attributes other than human and has a body part such as "horns" or "tail", the server 20 may hold that information. Here, the special body part does not have to belong to the avatar's body, and may be an object floating around the avatar. The special portion is not limited to the above. For example, an object such as a living thing different from the avatar may be arranged around the avatar.
[0091] The item "special body part movement setting" is information on the setting for moving a special body part of the avatar. Specifically, for example, when the avatar has a special body part (e.g., "horns", "tail", etc.), the server 20 may store in this record information on what condition triggers the movement of that body part. For example, in the case of an avatar with a special body part "horn", when it is set to "linked with the movement of the entire eyes", the server 20 may reflect the setting of the degree of change in the state of the eyes set by the user in the horn, and change the state. In addition, in a certain aspect, the server 20 may receive a setting of the degree of change in appearance from the user for each special body part. For example, in a case where the special body part is an object that is not connected to the body of the avatar but is floating around the avatar and changes its appearance, the server 20 may receive a setting input from the user for each of the objects. However, the server 20 may also change the appearance of the objects in accordance with the settings of the avatar's body parts (mouth, face, etc.).
[0092] Furthermore, when the special body part is an object such as a living thing different from the avatar and exists around the avatar, the server 20 may accept a setting of the degree to which the body part of the object (e.g., eyes, mouth, etc.) changes its appearance based on the user's voice spectrum or sensing results. For example, the server 20 may set the change amount of the object's eyes by multiplying the change amount of the avatar by a predetermined ratio, or may accept a setting from the user for each body part of the object. This allows the user to perform operations suited to the characteristics of a non-human avatar, even when operating the avatar.
[0093] The item "Notes" is information that is stored when there are special notes regarding the avatar information.
[0094] As shown in FIG. 4, the wearable device information database 2023 includes an "ID" item, an "Type" item, an "Detection Accuracy" item, an "Amount of Correction" item, and an "Notes" item.
[0095] The item "ID" is information that identifies each wearable device worn by the user.
[0096] The item "Type" is information that indicates the type of wearable device worn by the user. The wearable device worn by the user is not particularly limited, and may be eyewear such as glasses, or a device that covers the head such as an HMD.
[0097] The item "detection accuracy" indicates the detection accuracy of sensing the movement of the user's eyes or face when the user is wearing a wearable device. Specifically, for example, the server 20 may score the detection accuracy of sensing for each wearable device worn by the user and store the information. For example, if the user is wearing glasses with high transmittance and almost the same as the naked eye, the information may be stored as a detection accuracy of "◯". In this case, the score stored by the server 20 may not be a symbol such as "◯", but may be a numerical value such as "100" based on transmittance or the like, or may be a notation such as "A" or "good", and is not limited thereto.
[0098] The item "Correction amount" indicates the correction amount of the degree of change of the avatar set for each wearable device. Specifically, for example, the server 20 sets the correction amount of the degree of change of the avatar's appearance based on the above-mentioned detection accuracy value. For example, if the user is wearing glasses, a process may be executed to multiply the degree of change by a predetermined magnification based on the transmittance of the glasses. In one aspect, when the user is wearing a device such as an HMD and the server 20 can obtain sensing results of the user's eyes or face from the device, even if the detection accuracy is low, the server 20 may not perform any correction processing. The information on the wearable device held by the server 20 may also be information on a mask, an eye patch, etc. In that case, the server 20 may reflect the movement of a part covered by a mask, an eye patch, etc., in the avatar by reflecting the user's speech or the settings of other parts that are not covered, rather than changing the appearance based on the sensing results. This allows the user to broadcast without worrying about how they look when broadcasting.
[0099] The item "Notes" is information that is stored when there are special notes regarding the wearable device information.
[0100] <3 operations> A series of processes performed by the system 1 when acquiring the voice spectrum of the user's speech and changing the state of the mouth of an avatar corresponding to the user in accordance with the speech of the performer based on the acquired voice spectrum will be described below.
[0101] 5 is a flowchart showing a series of processes when acquiring a voice spectrum of a user's speech and changing the state of the mouth of an avatar corresponding to the user according to the speech of the performer based on the acquired voice spectrum. Note that, in this flowchart, an example is disclosed in which the control unit 190 of the terminal device 10 used by the user executes the series of processes, but this is not limited thereto. That is, the terminal device 10 may transmit a part of the information to the server 20 and execute the process in the server 20, or the server 20 may execute the entire series of processes.
[0102] In step S501, the control unit 190 of the terminal device 10 acquires a voice spectrum of the speech of the user who is the performer operating the avatar. Specifically, for example, the control unit 190 of the terminal device 10 controls a process of acquiring a voice spectrum from the voice uttered by the user acquired via the microphone 141, similar to the voice spectrum acquisition module 2035 of the server 20. For example, the control unit 190 acquires the user's voice via the microphone 141 and acquires the voice spectrum contained in the voice. For example, the control unit 190 may perform a Fourier transform on the voice acquired from the microphone 141 to acquire information on the voice spectrum contained in the voice. At this time, the calculation for acquiring the voice spectrum is not limited to a Fourier transform, and may be any existing method. In addition, in a certain aspect, the control unit 190 may acquire information on the voice spectrum of a vowel from the voice of the user. For example, the control unit 190 accepts a vowel setting input by the user in advance, and then accepts an utterance from the user via the microphone 141, thereby storing the accepted vowel setting and the acquired voice spectrum in association with each other. In addition, in one aspect, the voice spectrum acquisition module 2035 may acquire information on sounds resulting from consonants, such as "t," "c," "h," "k," "m," "r," "s," "n," and "w," and combine it with the stored vowel information to estimate the words spoken by the user. This allows the system 1 to separately characterize and store the voice spectrum relating to vowels among the user's voice spectrum, thereby making it possible to change the movement of the mouth shape of the avatar more accurately.
[0103] In step S502, the control unit 190 of the terminal device 10 changes the state of the mouth of the avatar corresponding to the user according to the user's utterance based on the acquired voice spectrum. Specifically, for example, the control unit 190 of the terminal device 10, like the avatar change module 2036 of the server 20, estimates the words uttered by the user from the acquired user's voice spectrum, and changes the state of the mouth of the avatar according to the estimated words. For example, when the acquired user's voice spectrum is "a", the control unit 190 changes the state of the mouth of the avatar to a shape corresponding to "a".
[0104] In step S503, the control unit 190 of the terminal device 10 presents an avatar corresponding to the user and the user's voice to the viewer. Specifically, for example, the control unit 190 of the terminal device 10 transmits an image of an avatar corresponding to the user and the user's voice to the display 1302 and the speaker 142 of the terminal device 10 used by the viewer, similar to the avatar presentation module 2037 of the server 20, and presents them to the viewer. At this time, the viewer is not limited to one person, and the avatars and voices may be presented to the terminal devices 10 of multiple viewers.
[0105] In step S504, the control unit 190 of the terminal device 10 accepts a setting for changing the state of the avatar's mouth in response to the speech of the performer to a degree lower than the change in the speech of the user. Specifically, for example, the control unit 190 of the terminal device 10 may accept settings including the following, similar to the setting acceptance module 2038 of the server 20. · Change the avatar's behavior based on the user's speech over a certain period of time · Set the update frequency (e.g. number of updates per second) Here, the change in the user's speech will be defined. The change in the user's speech is, for example, the speed of the user's speech, and may be calculated based on the following. The time interval when the vowel spoken by the user changes (for example, the time interval when the vowel changes from "a" to "i") At this time, the control unit 190 may simultaneously acquire sounds derived from consonants (c, k, etc.), and even if the same vowel is acquired consecutively, estimate the speech speed as if different words were being spoken. - Number of vowels uttered in a given period In addition, at this time, the control unit 190 may accept the setting to be lower than the degree of change of the avatar estimated from the change in the user's speech. For example, the control unit 190 may accept the time required to change the state of the mouth to correspond to the speech (vowel) estimated in advance from the user's voice spectrum. The control unit 190 acquires information on the time when the user uttered the vowel from the user's voice spectrum based on the accepted required time, calculates a ratio to the previously set required time, multiplies it by the amount of change in the state, and calculates the amount of change in the state of the avatar's mouth. The control unit 190 changes the state of the mouth based on the acquired speech time and the amount of change. For example, it is assumed that the user inputs a setting that the avatar's mouth changes to the state of "a" in a required time of "1 second". For example, the time when the state of "a" is completely changed is set to "100", and it is set to "100" in "1 second". At this time, the control unit 190 may also receive and record from the user the degree of change in the state in one second (amount of change in the mouth, speed). (That is, when the state of the mouth changes in one second, the difference in the amount of change in the state between the first 0.5 seconds and the remaining 0.5 seconds may be set.) When the user utters the sound "a" for one second, the control unit 190 changes the state of the avatar's mouth to that of "a" over one second based on the above-mentioned setting of the amount of change, etc. However, when the user utters "a" for only "0.5 seconds", the control unit 190 may perform processing to change the amount of change in the state of the avatar's mouth up to "50".
[0106] Furthermore, when the user speaks continuously (for example, speaking "aiueo"), the control unit 190 may obtain the speaking time of each vowel and perform the above process. That is, the control unit 190 may calculate the amount of change in the state of the avatar's mouth corresponding to each vowel from the speaking time of each vowel, and change the state of the avatar's mouth. For example, if the time required for the state of the avatar's mouth to change to that corresponding to each vowel is "1 second," and "a" is spoken for "0.2 seconds" and "i" is spoken for "0.3 seconds," the amount of change corresponding to "a" is "20," and the amount of change corresponding to "i" is "30." Furthermore, when the user speaks for a time longer than the required time, the control unit 190 may maintain the state of the state of the avatar's mouth even after the required time has passed. This allows the user to change the avatar's mouth appearance gradually according to the time the user is speaking, rather than instantly changing the mouth appearance to "ah" when the user utters the sound "ah." In addition, by setting a required time and multiplying the change in the mouth appearance when the utterance does not reach that time, it is possible to prevent the avatar's mouth appearance from changing significantly even when the user speaks lightly (for example, even if the mouth is opened only about 30 degrees, the avatar's mouth appearance changes to 100). This allows the user to prevent the viewer from feeling uncomfortable due to the user's speech and the change in the avatar's mouth appearance, thereby giving the viewer a greater sense of immersion.
[0107] In a certain aspect, the control unit 190 may change the state of the mouth of the avatar based on information such as the size and pitch of the voice spectrum acquired from the user. Specifically, for example, the control unit 190 may acquire information on the frequency (Hz) and sound pressure (dB) of the user's voice spectrum, and may change the state of the mouth of the avatar when the information exceeds a predetermined threshold. For example, it is assumed that the control unit 190 accepts a setting to change the state of the mouth of the avatar in a required time of "1 second", and the user's speaking time is "1 second". At this time, when the control unit 190 detects that the user has spoken with a sound pressure exceeding the threshold at the time of "0.8 seconds", it may change the state of the mouth of the avatar to a larger state than usual (to make the mouth open widely). At this time, the control unit 190 may reflect the same setting not only on the mouth, but also on parts of the face and parts of the body. This allows the user to have the avatar's mouth movements reflect the user's sudden loud voice, allowing the viewer to see more natural avatar movements.
[0108] Additionally, control unit 190 may accept a setting for the degree of change in the state of the avatar's mouth so that the setting is lower than the update frequency of the avatar's movement estimated from the speech speed estimated from the user's speech. For example, the control unit 190 divides the user's speech into fixed time intervals and changes the avatar to a state of the mouth corresponding to the first and last vowels of the time interval. For example, when the speech changes to "aiueo" in one second, the control unit 190 may reflect in the avatar the shape of the mouth at the timing of the first "a" of "aiueo" and the state of the mouth at the "o". In addition, when the control unit 190 temporarily stores the user's speech in a buffer in a memory, the control unit 190 may change the state of the avatar's mouth from "a" to "o" over a certain time interval (for example, one second). The control unit 190 may also accept a setting so that the state of the avatar's mouth changes slower than the time that elapses when the user's vowel changes. For example, when the user's vowel changes from "a" to "u" and it takes one second for the change, the server 20 may take 1.5 seconds for the state of the avatar's mouth to change from "a" to "u". At this time, the server 20 may also execute a process to complement the change in state. That is, the server 20 may change the state of the avatar's mouth by passing through a mouth shape that is intermediate between "a" and "u", rather than immediately changing the state of the avatar from "a" to "u". This allows the user to change the movement of the avatar's mouth in a manner that is closer to the mouth movements of a real person, rather than having the avatar's mouth change instantly for each word, thereby reducing the sense of discomfort felt by viewers when viewing the avatar. At this time, the control unit 190 may accept from the user a setting of the degree that can be set lower than the user's speech rate. Specifically, for example, the control unit 190 may calculate the user's speech rate from the voice spectrum of the speech accepted from the user. After that, the control unit 190 accepts from the user a setting of the degree to be set lower than the user's speech rate by setting an upper limit value of the amount of change per unit time of the avatar's appearance that can be accepted from the user from the calculated user's speech rate. This allows the user to set the degree of change of the avatar to be slower than the change in the user's own speech, allowing the avatar's appearance to change more smoothly.
[0109] In step S505, the control unit 190 of the terminal device 10 changes the state of the mouth of the avatar according to the setting. Specifically, for example, the control unit 190 of the terminal device 10 changes the state of the mouth of the avatar according to the setting based on the information set in step S604, and then presents the avatar and the voice of the user to the viewer. This allows the user to smoothly change the state of the avatar's mouth in accordance with the user's speech, and allows the viewer to see more natural mouth movements. In one aspect, control unit 190 of terminal device 10 may change the state of the mouth of the avatar based on at least one of the group consisting of strength and weakness and pitch of the voice spectrum. Specifically, for example, the control unit 190 determines the strength and pitch by analyzing the following parameters of the audio spectrum: Audio spectrum dynamics parameters: decibels (dB) Audio spectrum high / low parameter: Hertz (Hz) For example, when the control unit 190 acquires a sound spectrum that is louder than the reference sound spectrum in decibels, the control unit 190 may change the state of the mouth of the avatar more than the change in the state of the mouth at the time of the reference. This allows the user to change the state of the avatar based on subtle changes in the voice, reducing the sense of discomfort felt by the viewer.
[0110] In a certain aspect, the control unit 190 of the terminal device 10 may receive a setting of a frequency range for detecting the voice spectrum, and may change the state of the mouth of the avatar based on a first setting of the degree in response to detecting the voice spectrum in the set range. Specifically, for example, the control unit 190 receives a setting of upper and lower limit values as a frequency range for detecting the voice spectrum from the user in step S604. The control unit 190 analyzes the voice spectrum of the user's speech acquired via the microphone 141, and determines whether the frequency of the voice spectrum is within the range. If the frequency is within the range, the control unit 190 may change the state of the avatar in step S605 based on the first setting of the degree, i.e., the setting of the degree of change of the state of the avatar, which is set in advance by the user.
[0111] In addition, in a certain aspect, the control unit 190 of the terminal device 10 may change the state of the mouth of the avatar based on a second setting, which is a setting of a predetermined degree and different from the first setting, in response to detecting a voice spectrum outside the set range. Specifically, for example, when the control unit 190 detects a frequency outside the frequency range for detecting the voice spectrum received from the user, the control unit 190 may change the state of the avatar based on a setting (second setting) different from the normal setting (first setting). For example, when the user utters a voice (e.g., an extremely shriek, etc.) outside the range of frequencies normally used, the voice spectrum falls outside the detection range. In that case, the control unit 190 may change the state of the avatar by applying a setting (second setting) that is applied only outside the detection range, instead of the setting (first setting) for change received from the user. This allows the avatar's appearance to change accordingly even if the user moves or speaks in a way that is different from normal, providing the viewer with a greater sense of immersion.
[0112] In one aspect, in response to detecting a voice spectrum outside the set range, control unit 190 may change the appearance of face parts and body parts other than the mouth. Specifically, for example, in response to detecting a voice spectrum outside the set range, control unit 190 may cause the avatar to perform the following actions. - Changing the appearance of facial parts (eyebrows, corners of the eyes, corners of the eyes, corners of the mouth, etc.) - Changing the appearance of body parts (arms, hands, shoulders, etc.) In addition, in response to detecting an audio spectrum outside the set range, the control unit 190 may display a predetermined object on the screen viewed by the viewer. This allows the control unit 190 to better convey the user's emotions to the viewer, for example, when the user suddenly shouts or screams, by changing the appearance of parts of the face or body, displaying objects, etc.
[0113] In a certain aspect, the control unit 190 of the terminal device 10 may estimate one or more candidate emotions of the user and present the estimated one or more candidate emotions of the user to the user. After that, the control unit 190 may receive an input operation from the user to select one emotion from the one or more candidate emotions, and when receiving the selection of the emotion from the user, may execute a process of changing the state of the mouth of the avatar based on the selected emotion. Specifically, for example, the control unit 190 analyzes the voice spectrum acquired from the user and estimates candidate emotions when the user speaks. In this case, the candidate emotions include, for example, the following: Anger, rage Joy, fun Surprise, fear Sadness, grief Peace and serenity Here, a process of estimating emotion candidates from a voice spectrum will be illustrated. For example, the control unit 190 may accept information on a voice spectrum corresponding to an emotion from the user in advance, and store the information in the storage unit 180 or the like, thereby associating the user's voice spectrum with the user's emotion. After that, when the control unit 190 acquires a voice spectrum from the user, it estimates emotion candidates associated with a voice spectrum having a waveform similar to that of the acquired voice spectrum. The waveforms being similar indicate, for example, that the similarity between the waveforms of a plurality of voice spectra is determined, and the waveforms match a predetermined percentage, or the waveforms of a plurality of voice spectra deviate from each other by a predetermined percentage (for example, match within a range of ±10%). In a certain aspect, a trained model may be used as a method for estimating candidates of a user's emotion from a voice spectrum. For example, the terminal device 10 may store a trained model in the storage unit 180 that associates the voice spectra of multiple users with emotions corresponding to the users. After that, when the control unit 190 of the terminal device 10 receives an input of a voice spectrum from a user, the control unit 190 may estimate candidates of emotions corresponding to the user's voice spectrum based on the trained model and present the candidates to the user.
[0114] The control unit 190 may present candidates of the estimated emotion to the user and accept a selection from the user. The control unit 190 also accepts in advance a setting of the degree of change in the state of the mouth for each emotion, and when it accepts a selection of an emotion from the user, it changes the state of the mouth of the avatar based on the setting of the corresponding emotion. This allows the user to change the appearance of the avatar based on the emotion estimated from the utterance.
[0115] At this time, if the control unit 190 cannot estimate the user's emotion, the control unit 190 may change the state of the mouth based on conditions preset by the user. Specifically, for example, if the control unit 190 cannot estimate a candidate for the user's emotion from the voice spectrum acquired from the user, that is, if a similar voice spectrum cannot be estimated, the control unit 190 may change the state of the avatar's mouth based on conditions preset by the user. For example, when the control unit 190 cannot accurately acquire a voice spectrum from the user, when the control unit 190 cannot estimate a candidate emotion similar to the acquired voice spectrum, or when the control unit 190 has received a mouth correspondence setting for “calm” from the user, the control unit 190 changes the state of the avatar's mouth to a state based on the emotion of “calm”. This allows the user to change the avatar into a preset state even when the emotion cannot be estimated, thereby reducing the sense of discomfort felt by the viewer.
[0116] Furthermore, in a certain aspect, the control unit 190 may operate a body part other than the mouth of the avatar based on the estimated emotion. Specifically, for example, the control unit 190 may operate a body part such as a shoulder, arm, or hand as a body part other than the mouth of the avatar. In addition, the control unit 190 may operate a special body part (for example, wings, a tail, or an object floating around the avatar if the avatar is not human) as a body part other than the mouth of the avatar. For example, the control unit 190 may raise the arm of the avatar when the emotion estimated from the voice spectrum acquired from the user is "anger" or the like. In addition, at this time, the control unit 190 may present candidate emotions to the user, rather than an emotion estimated from the acquired voice spectrum, and move a part of the avatar's body other than the mouth based on the emotion selected by the user.
[0117] In addition, the control unit 190 may present to the user one or more candidates of body parts other than the mouth of the avatar whose behavior is to be changed based on the emotion estimated from the acquired voice spectrum, and change the behavior of the body part in response to receiving a selection of the body part whose behavior is to be changed from the user. This allows the user to move parts of the avatar other than the mouth based on emotions estimated from the user's voice spectrum, providing a greater sense of immersion to the viewer.
[0118] In a certain aspect, when the user's speech rate deviates by a predetermined rate from the speech rate estimated from the degree of change in the mouth state set by the user, the control unit 190 of the terminal device 10 may change the mouth state based on the speech rate, not the degree set by the user. Specifically, for example, the control unit 190 calculates the speech rate from the user's speech. As a method of calculating the speech rate, for example, the control unit 190 may define the speech rate value by calculating the number of words per unit time by the user from the voice spectrum acquired from the user. In addition, the control unit 190 calculates the number of utterances per unit time from the degree of change in the mouth state set by the user, and calculates the user's speech rate estimated from the degree of change in the mouth state of the avatar. Thereafter, if there is a predetermined deviation between the speech rate calculated from the user's speech and the speech rate estimated from the degree of change in the state of the avatar's mouth, the control unit 190 may change the state of the mouth based on the speech rate calculated from the user's speech rather than the degree set by the user. This allows the user to change the avatar's mouth movement based on the speech rate when the user's own speech rate deviates too much from the speech rate estimated from the degree of change in the avatar's mouth movement, thereby reducing the sense of discomfort felt by the viewer.
[0119] In a certain aspect, the control unit 190 of the terminal device 10 may receive an attribute of the avatar from the user, and correct the amount of change in the state of the mouth of the avatar based on the attribute. Specifically, for example, the control unit 190 may receive information on either a human or a non-human whose mouth state changes differently from that of a human as an attribute of the avatar, and correct the amount of change in the state of the mouth based on the attribute. For example, the control unit 190 may obtain information on whether the avatar operated by the user is a human or a non-human whose state changes differently from that of a human, similar to the change correction module 2040 of the server 20, and execute a process of correcting the degree of change in the state of the avatar based on the information. For example, if the attribute of the avatar operated by the user is a "dragon", the movement of the eyes, mouth, etc. may behave differently from that of a human. In that case, the control unit 190 may correct the amount of change in the corners of the mouth, the amount of change in the eyeballs, etc. to a shape that matches the avatar based on the attribute of the "dragon". This allows users to present more natural movements to viewers based on the results of their own speech and facial sensing, even when they are operating an avatar that is not human.
[0120] <4 Screen example> 6 to 9 are diagrams showing various examples of screens when a user, who is a performer operating an avatar, operates an avatar using the system 1 disclosed in the present invention.
[0121] FIG. 6 is an example of a screen when a user registers the speech spectrum of his / her own vowel in the system 1.
[0122] In FIG. 6, the control unit 190 of the terminal device 10 displays a setting screen 601, an avatar 602, and the like on the display 1302. The setting screen 601 is a setting screen displayed to the user when acquiring and associating information on the voice spectrum corresponding to each vowel from the user. For example, the control unit 190 of the terminal device 10 displays a setting screen for six characters, "A", "I", "U", "E", "O", and "N", as vowels to be associated with the user's voice spectrum on the screen. At this time, the control unit 190 may display information on the vowels currently associated with the user's voice spectrum at the top of the screen. Furthermore, the control unit 190 may display information about the microphone 141 used by the user at the bottom of the setting screen 601. When frequency characteristics differ depending on the type of microphone 141 used by the user, the control unit 190 may associate the user's voice spectrum with vowel information for each microphone 141 used.
[0123] The avatar 602 is an avatar whose mouth state is changed in response to the user's speech. The control unit 190 changes the state of the mouth of the avatar 602 in response to the voice spectrum acquired from the user. For example, when the user utters the vowel "a (A)", the control unit 190 determines whether the utterance matches the voice spectrum stored as the vowel "a (A)". Thereafter, when the user utters "a (A)", the control unit 190 changes the state of the mouth of the avatar 602 to the shape of "a (A)". This allows the user to accurately change the state of the avatar's mouth for each vowel.
[0124] FIG. 7 shows an example of a screen when the user sets the degree of change in the appearance of the avatar's mouth or other facial features.
[0125] In FIG. 7, the control unit 190 of the terminal device 10 displays, on the display 1302, an information display screen 701, a user video 702, a setting screen 703, an avatar 704, and the like.
[0126] The information display screen 701 is a screen that displays the frequency of the voice spectrum acquired from the user, the range of the detectable voice spectrum, the setting of the mode when it is outside the detection range, etc. In addition, the terminal device 10 may display the user's speech speed calculated from the user's speech, the degree of change in the mode that the user can set, the sensing result of the user's face, etc. on the screen, and visually display various conditions that the user can set.
[0127] The user video 702 is a screen that displays an image of the user himself / herself captured via the camera 160 provided in the terminal device 10. When the user makes some kind of speech in front of the terminal device 10, the control unit 190 of the terminal device 10 displays an image of the user himself / herself and information such as the voice spectrum of the user's speech on the user video 702 and the information display screen 701 by the camera 160 and the microphone 141 provided in the terminal device 10.
[0128] The setting screen 703 is a screen for the user to set the degree of change in the appearance of the avatar. The control unit 190 of the terminal device 10 presents the following settings to the user, for example, and accepts input. Mouth switching speed The mouth switching speed is information regarding the time required for the voice spectrum acquired from the user to reach the maximum volume (100). Eye movement: Maximum upward movement Eye movement: Maximum downward movement Eye movement: Maximum horizontal movement Eye movement: Sensitivity (device sensing sensitivity) Sensitivity refers to the sensitivity reflected in the avatar when sensing the user's eyes, etc. Specifically, for example, the sensitivity is a parameter that sets the extent to which the avatar's eyes, etc. are reflected in the actual movement amount of the eyes, etc., when the user moves their eyes, etc., in the left-right direction when the coordinates of the position of the eyes, etc., when facing straight ahead are set to "0". In this case, the sensitivity may be a proportional function at 100, and a function that is convex downward as it approaches 0. In other words, when the sensitivity is 100, the movement of the user's eyes and the movement of the avatar's eyes are completely synchronized, and when the sensitivity is 50, etc., if the user's eyes, etc., do not move much from the center, the movement of the avatar's eyes, etc. is reflected less than the movement distance of the user's eyes, and if the eyes move to the corners of the eyes, the movement of the avatar's eyes, etc. is reflected more than the movement distance of the user's eyes. This makes it possible to prevent the avatar's eyes from moving "wandering" when the user does not move their eyes much. The sensitivity setting is not limited to the eyes, and similar settings may be accepted for other parts of the face and body other than the eyes. At this time, the control unit 190 of the terminal device 10 may accept a setting of the degree of change that can be accepted from the user that is lower than the change in the user's speech. For example, the control unit 190 may accept the setting from the user so that the setting is lower than the degree of change of the avatar (amount of change in the object, speed of change in the object) estimated from the user's speech. At this time, if the user attempts to set a value that is not within the settable range, the control unit 190 may display a predetermined alert, or if the setting screen is a slider type, may lock the value in advance so that it is not set. This allows the user to move the avatar more slowly than the user's own speech, thereby smoothing the degree of change in the avatar felt by the viewer, thereby giving the viewer a greater sense of immersion.
[0129] The avatar 704 is an avatar that changes its appearance based on settings received from the user. When the control unit 190 of the terminal device 10 receives settings on the setting screen 703 from the user, the control unit 190 may synchronize the user video 702 and the avatar 704 and display them to the user. This allows the user to check in advance for any discomfort or other issues that may arise when changing the appearance of the avatar based on the user's own settings.
[0130] FIG. 8 shows an example of a screen in which one or more candidate emotions of a user are estimated from the user's speech, and the appearance of an avatar is changed based on the estimated one or more emotions of the user.
[0131] In FIG. 8, the control unit 190 of the terminal device 10 displays an information display screen 801, a user video image 802, an avatar 803, and the like on the display 1302.
[0132] Information display screen 801, like information display screen 701 in FIG. 7, is a screen that displays the frequency of the voice spectrum obtained from the user, and in FIG. 8, it may present one or more candidate emotions estimated from the voice spectrum, as well as candidate emotion settings that the user will reflect in the appearance of the avatar. Control unit 190 may change the appearance of the avatar, for example, the appearance of the avatar's mouth or the appearance of parts of the face other than the mouth, by accepting a selection from the presented setting candidates from the user.
[0133] The user video 802 is a screen that displays an image of the user himself / herself captured via the camera 160 provided in the terminal device 10, similar to the user video 702 in FIG.
[0134] The avatar 803 is an avatar that changes its appearance based on the emotion setting received from the user, similar to the avatar 704 in FIG. 7. The control unit 190 of the terminal device 10 changes the appearance of the avatar (for example, the mouth) based on the emotion setting received from the user and displays it to the user. At this time, the control unit 190 may change or move the appearance of other parts of the avatar, not limited to the appearance of the avatar's mouth. For example, when the emotion selected and received from the user is "anger", the control unit 190 may change the appearance of the avatar's mouth based on the emotion of "anger", and may also change the appearance of other parts of the avatar, such as the eyebrows and corners of the eyes. In addition, the control unit 190 may move parts of the avatar's body (for example, swinging up the arms, etc.) based on the emotion. In addition, the control unit 190 may display a predetermined object corresponding to the emotion on the screen displaying the avatar based on the emotion. This allows the user to cause the avatar to undergo various changes and actions based on emotions estimated from the user's speech, thereby providing a greater sense of immersion to the viewer.
[0135] In addition, in a certain aspect, control unit 190 may refer to user information 1801 or user information database 2021 to obtain information on emotions frequently used by the user, and present the information as candidates for emotions to be reflected in the avatar. This allows the user to easily change the state of the avatar even when the user is trying to change the state of the avatar for dramatic effect or the like, regardless of speech.
[0136] FIG. 9 shows an example of a screen on which a user can make various settings based on the voice spectrum, etc., for an avatar with attributes different from those of a human.
[0137] In FIG. 9, the control unit 190 of the terminal device 10 displays, on the display 1302, an information display screen 901, a user video 902, a setting screen 903, an avatar 904, and the like.
[0138] 7 and 8, the information display screen 901 is a screen that displays the frequency of the voice spectrum acquired from the user, the range of the detectable voice spectrum, the setting of the mode when it is outside the detection range, etc. At this time, the control unit 190 may display information on the attributes of an avatar corresponding to the user on the information display screen 901. For example, the control unit 190 may refer to the user information 1801 or the user information database 2021 and acquire information on the avatar corresponding to the user, thereby displaying information on the attributes of the avatar on the screen.
[0139] The user video 902 is a screen that displays an image of the user himself / herself captured via the camera 160 provided in the terminal device 10, similar to the user videos 702 and 802 in FIG.
[0140] The setting screen 903 is a screen for the user to set the degree of change in the avatar's appearance, similar to the setting screen 703 in Fig. 7. In Fig. 9, the control unit 190 may display, in addition to the screen presented to the user in the setting screen 703, suggestions of settings recommended based on the attributes of the avatar. Specifically, for example, the control unit 190 may refer to the avatar information 1802 or the avatar information database 2022, etc., to acquire information on the correction amount of the degree of change in appearance caused by the avatar, and present to the user a setting obtained by multiplying the correction result by a basic setting for changing the appearance of a normal human avatar. This allows the user to set changes in appearance that do not feel unnatural even if the user's avatar has attributes different from those of a human. In addition, in a certain aspect, when the avatar has a special body part, the control unit 190 may display a screen for the user to set the degree of change in the appearance of the special body part. For example, when synchronizing with the setting of another body part, the control unit 190 may reflect the setting of the change in the body part of the other avatar, or may present the user with a screen for separately setting the degree of change in appearance in detail. This allows the user to freely set the degree of change in appearance even if the user's avatar has a special body part, providing a greater sense of immersion to the viewer.
[0141] Avatar 904 is an avatar that changes its appearance based on an emotion setting received from a user, similar to avatars 704 and 803 in Fig. 7 and Fig. 8. In Fig. 9, control unit 190 may simultaneously display a special part of the avatar on avatar 904. This allows the user to broadcast to viewers while checking changes in the avatar's appearance, even if the avatar has a special body part.
[0142] <Second embodiment> So far, a series of processes for changing the state of the avatar's mouth based on the voice spectrum of the user's speech has been described. In the invention according to the second embodiment, in addition to the voice spectrum of the user's speech, the avatar's appearance, for example, the appearance of one or more facial parts, can be changed based on the user's sensing results. The series of processes will be described below. Note that the description of parts having the same configuration as the first embodiment (for example, the terminal device 10, the server 20, etc.) will be omitted, and only the configuration and processing unique to the second embodiment will be described.
[0143] <5. Operation in the Second Embodiment> Below, we will explain a series of processes that system 1 performs when it senses the movement of one or more facial parts of a user and changes the appearance of one or more facial parts of an avatar corresponding to the user based on the sensed movement of the one or more facial parts.
[0144] 10 is a flowchart showing a series of processes when sensing the movement of one or more facial parts of a user's face and changing the state of one or more facial parts of an avatar corresponding to the user based on the sensed movement of the one or more facial parts. Note that this flowchart also discloses an example in which the control unit 190 of the terminal device 10 used by the user executes the series of processes, but is not limited to this. That is, the terminal device 10 may transmit part of the information to the server 20 and the server 20 may execute the process, or the server 20 may execute the entire series of processes.
[0145] In step S1001, the control unit 190 of the terminal device 10 senses the movement of one or more facial parts of the user. Specifically, for example, the control unit 190 of the terminal device 10 senses one or more facial parts of the user when the user moves his / her face in front of the camera 160 provided in the terminal device 10. At this time, the sensing method performed by the control unit 190 may be any existing technology. For example, the control unit 190 may sense the user's facial parts by providing the camera 160 with a sensing function, or may sense the user's facial parts by the motion sensor 170. At this time, the control unit 190 of the terminal device 10 senses at least one of the group consisting of the user's eyebrows, eyelids, inner corners of the eyes, outer corners of the eyes, eyeballs, pupils, and mouth as one or more facial parts of the user. However, the part is not limited to this and may be another facial part (cheeks, forehead, etc.).
[0146] In step S1002, the control unit 190 of the terminal device 10 changes the state of one or more facial parts of the avatar corresponding to the user based on the movement of the one or more facial parts sensed. Specifically, for example, the control unit 190 associates one or more facial parts of the user with one or more facial parts of the avatar in advance. Thereafter, the control unit 190 changes the state of the facial parts of the avatar corresponding to one or more facial parts of the user acquired by sensing based on the sensing result. For example, when the control unit 190 has associated the user's eyes with the avatar's eyes, it changes the state of the avatar's eyes based on the sensing result of the user's eyes.
[0147] In step S1003, the control unit 190 of the terminal device 10 accepts a setting of the degree to which the appearance of one or more facial parts of the avatar is made to follow the sensing result, and changes the appearance of one or more facial parts of the avatar according to the setting of the degree. Specifically, for example, the control unit 190 accepts a setting of conditions including the following as the degree to which the appearance is made to follow the user's sensing result. -Changes in the appearance of the avatar (for example, changes in the opening and closing of the eyes, etc.) This allows the user to finely adjust the degree to which the changes in the avatar's appearance follow the user's own sensing results, preventing the viewer from feeling uncomfortable with the movements.
[0148] In the second embodiment, the control unit 190 may accept a setting from the user for the degree of change in the state of the avatar's face parts and body parts other than the face, similar to the setting of the degree of change in the state of the avatar's mouth in the first embodiment. That is, the control unit 190 may acquire from the user in advance by sensing the state of the mouth, various faces, and body parts corresponding to various vowels. The control unit 190 may specify a difference from the change in the user's mouth, face parts, and body parts acquired in advance from the user's sensing result, calculate a ratio to the previously acquired sensing result, and multiply the amount of change in the state by the ratio to calculate the amount of change in the state of the avatar's mouth, face parts, and body parts. The control unit 190 may change the state of the avatar's mouth, face parts, and body parts based on the calculated amount of change. For example, if the user only moves a portion of their mouth or eyebrows (100 positions are set in advance, and sensing results show that the user only moves their mouth, eyebrows, etc. up to position 50), processing may be performed such that the avatar's mouth, eyebrows, etc. also only move up to position 50. This allows the user to gradually change the state of the avatar according to the result of his / her own sensing, and allows the viewer to see natural movements. This allows the user to prevent the viewer from feeling uncomfortable due to the user's movements and the change in the avatar's state, and thus gives the viewer a greater sense of immersion.
[0149] In one aspect, control unit 190 may accept the same settings for predetermined associated parts among one or more facial parts of an avatar. Specifically, control unit 190 may accept settings from a user that associate, for example, the following parts among one or more facial parts of an avatar, and may accept the same settings regarding the degree settings for the parts. Paired parts of the face such as eyebrows and eyes Parts that work together like eyebrows and eyes ·Facial areas and other body areas (shoulders, arms, legs, neck, etc.) Additionally, control unit 190 may accept a setting that associates a facial part with a special non-facial part according to an attribute of an avatar, which will be described later. This allows the user to easily change and distribute the appearance of the avatar without having to set individual degrees for paired parts of the face or parts that move in conjunction with each other.
[0150] In a certain aspect, the control unit 190 of the terminal device 10 may present one or more candidates for the setting of the degree of following the sensed result to the user, and may accept a selection of one or more candidate settings of the degree from the user. After that, the control unit 190 may change the state of one or more facial parts of the avatar based on the setting of the degree selected and accepted. Specifically, for example, when acquiring the sensing result of the user, the control unit 190 may present the degree of following as one or more candidates (presets) instead of accepting a setting of the degree of following from the user. In this case, as a method of presenting the candidates, the control unit 190 may accept information of one or more candidates of the degree of following to be used from the user in advance, and present the candidates based on the information. This allows the user to change the appearance of parts of the avatar's face based on the sensing results without having to set the degree of tracking each time, making distribution easier.
[0151] In addition, in a certain aspect, the control unit 190 of the terminal device 10 may receive an attribute of the avatar from the user and correct the degree based on the attribute. Here, the control unit 190 may receive information on either a human or a non-human whose change in the appearance of one or more facial parts is different from that of a human as an attribute, and correct the degree based on the attribute. For example, the control unit 190 may obtain information on whether the avatar operated by the user is a human or a non-human whose change in appearance is different from that of a human, similar to the change correction module 2040 of the server 20, and execute a process of correcting the degree of change in the appearance of the avatar based on the information. For example, if the attribute of the avatar operated by the user is a "dragon", the movement of the eyes, mouth, etc. may behave differently from that of a human. In that case, the control unit 190 may correct the amount of change in the corners of the mouth, the amount of change in the eyeballs, etc., to a shape that matches the avatar based on the attribute of the "dragon". This allows users to present more natural movements to viewers based on the results of their own speech and facial sensing, even when they are operating an avatar that is not human.
[0152] In another aspect, the control unit 190 of the terminal device 10 may acquire a voice spectrum of the user, and acquire information on the degree of change in the user's speech from the acquired voice spectrum. Thereafter, the control unit 190 may accept a degree setting that is settable within a range associated with the degree of change in the user's speech, and change the state of one or more facial parts of the avatar according to the degree setting. Specifically, for example, the control unit 190 may acquire a voice spectrum from the user's speech via the microphone 141 or the like, and acquire the following information as the degree of change in the user's speech: - The amount of words spoken by the user per unit time (speech rate) - Change in volume of user's voice - Changes in the pitch of the user's voice For example, the control unit 190 executes the following process to change the state of the avatar to a degree less than the degree of change of the avatar estimated from the change in the user's speech. · Reflects mouth movements on the avatar at regular time intervals, regardless of vowel changes in the voice spectrum obtained from the user. The control unit 190 specifies a settable range of the degree of tracking the sensing result based on the acquired information on the degree of change in speech. For example, the control unit 190 may accept a setting of the degree from the user so that the amount of change or the like described above does not exceed the degree of change in speech based on the acquired degree of change in speech. This allows the user to change the facial expression of the avatar based on not only the sensing results but also voice spectrum information, allowing the avatar to appear to viewers with more natural movements.
[0153] At this time, the control unit 190 may receive a setting of a frequency range for detecting the voice spectrum, and in response to detecting the voice spectrum in the set range, change the appearance of one or more facial parts of the avatar based on a first setting of the degree. Specifically, for example, when acquiring a voice spectrum from the user's speech, the control unit 190 may receive a setting of a detectable range from the user. When the voice spectrum acquired from the user is within the frequency range, the control unit 190 may change the appearance of the avatar's face based on the setting of the degree received from the user. Furthermore, in response to detecting a voice spectrum outside the set range, the control unit 190 may change the appearance of one or more facial parts of the avatar based on a second degree setting that is a predetermined degree setting and different from the first degree setting. In this case, the second degree setting may be, for example, when the user utters a voice with an extremely high frequency (such as a shriek), not the setting of the degree (first degree) received from the user, but a preset degree (second degree) corresponding to the frequency may be reflected to change the appearance of the avatar's face. This allows the user to change the facial expression of the avatar even when the user speaks at a frequency that is not normally spoken, thereby providing a greater sense of immersion to the viewer.
[0154] In addition, in a certain aspect, the control unit 190 of the terminal device 10 may change the state of the mouth of the avatar based on the degree of change in the user's speech when the movement of the user's mouth cannot be sensed. Specifically, for example, the control unit 190 may change the state of the mouth of the avatar based on the voice spectrum of the user's speech, rather than the result of sensing the user, as described above, in the following cases. - When the user is wearing a mask or other object over their mouth and mouth movements cannot be sensed When the mouth movement cannot be sensed due to an error in the sensing function of the terminal device 10 When mouth movements cannot be sensed due to external conditions This allows a user to change the state of the avatar's mouth to match the user's speech, for example, even when the user has to broadcast while wearing a mask.
[0155] In a certain aspect, the control unit 190 of the terminal device 10 may estimate one or more candidates of the user's emotions and present the estimated one or more candidates of the user's emotions to the user. Thereafter, the control unit 190 may accept an input operation from the user to select one of the one or more candidates of emotions, and change the state of one or more facial parts of the avatar corresponding to the user based on the selected emotion. Specifically, for example, the control unit 190 may acquire and associate in advance from the user a sensing result of a facial part corresponding to the user's emotion. Thereafter, the control unit 190 senses the user's face via the camera 160 or the like, and determine whether all or a part of the sensing result of the face included in the associated emotion matches. Thereafter, the control unit 190 may present candidates of the user's emotions based on the determination result, accept a selection from the user, and change the state of the avatar's face based on the selected emotion. At this time, if the user's emotion cannot be estimated, control unit 190 may change the appearance of one or more facial parts based on a setting preset by the user. For example, when the control unit 190 is unable to accurately sense parts of the user's face, or when it is unable to estimate a candidate emotion similar to the sensing result, if the control unit 190 has received a mouth correspondence setting for “calm” from the user, the control unit 190 changes the state of the avatar's mouth to a state based on the emotion of “calm”. This allows the user to select emotion candidates so that the change in the avatar's appearance reflects the user's emotion, even if accurate sensing cannot be performed.
[0156] In addition, in a certain aspect, when the control unit 190 of the terminal device 10 is unable to obtain sensing results for at least one associated part of one or more parts of the user's face, the control unit 190 may apply the degree of the part for which sensing results have been obtained to the associated part. Specifically, for example, when the user is wearing an eye patch or the like and sensing of one eye is difficult or impossible, the control unit 190 may reflect the degree of change in the other eye for which sensing results have been obtained. In this way, even if the user is wearing an eye patch or the like, the avatar corresponding to the user can change its appearance without being affected by the eye patch.
[0157] Furthermore, in a certain aspect, the control unit 190 of the terminal apparatus 10 may acquire information on a wearable device worn by the user, and correct the degree setting based on the acquired information on the wearable device. Furthermore, when correcting the degree setting, the control unit 190 may accept an input operation for adjusting the degree of correction from the user. Specifically, for example, the control unit 190 refers to the wearable device information 1803 or the wearable device information database 2023, and acquires information on the wearable device worn by the user. Thereafter, the control unit 190 may execute a process similar to that of the change correction module 2040 in the server 20 described above, and correct the degree setting.
[0158] In one aspect, the control unit 190 of the terminal device 10 may present a predetermined notification to the user when the difference in the degree setting between pre-associated parts of one or more facial parts of the avatar exceeds a predetermined threshold. Specifically, the control unit 190 associates, for example, paired parts such as eyebrows among one or more facial parts of the avatar, and sets the degree numerical value to be acceptable so that the degree of change between the parts does not exceed a predetermined difference. Thereafter, when accepting an input of the degree of change of the part from the user, the control unit 190 may present a notification such as an alert to the user if the input of a numerical value exceeding the threshold is accepted. This allows the user to prevent changing the state of the associated body part with its state changed in a state where there is an extreme difference in the degree of change. Furthermore, the control unit 190 may set the above settings not only for paired parts, but also for parts that change in conjunction with each other, such as the cheeks and eyebrows (which may include special parts, etc.).
[0159] At this time, when presenting a predetermined notification to the user, the control unit 190 may present the part whose difference in degree exceeds a predetermined threshold in a different manner together with the numerical value to the user. Specifically, for example, when the control unit 190 receives an input of the degree of change in the state of the eyes from the user, if the degree of change in both eyes is too large (for example, the amount of change in one eye is too large, etc.), the control unit 190 may present the eyes in a different manner (for example, in a different color manner) together with the notification to the user. At this time, the different manner presented by the control unit 190 may be, but is not limited to, a color, a pop-up notification, or changing the shape of the corresponding part. Furthermore, when presenting a predetermined notification to the user, the control unit 190 may present to the user how at least one or more facial parts change when the difference in degree is set within a predetermined range. For example, the control unit 190 may display, on a screen different from the screen displaying the above-mentioned notification, how the appearance of the avatar changes when the difference in degree is within an appropriate range (a range that does not cause discomfort to the viewer). This allows the user to check, when the degree of change in behavior that he or she has set exceeds a specified threshold, how the behavior would change if the user set it to an appropriate value.
[0160] In addition, in a certain aspect, the control unit 190 of the terminal device 10 may set the degree of a part related to one or more facial parts for which a degree setting has been accepted to a predetermined value. The control unit 190 may also accept a degree setting within a predetermined range for each of one or more parts of the avatar. Specifically, for example, the control unit 190 may associate the following parts as parts related to one or more facial parts of the avatar, and accept a degree setting from the user: -Horns, tails, wings, and other special body parts for non-human avatars -Any body part different from the avatar's face (arms, shoulders, legs, etc.) This allows the user to change the appearance of the avatar in accordance with the results of his or her own sensing, even if the avatar is non-human or an inorganic object.
[0161] <6 Screen Examples in the Second Embodiment> 11 to 17 are diagrams showing various examples of screens when the state of an avatar is changed based on the result of sensing of a user, as disclosed in the second embodiment.
[0162] FIG. 11 shows an example screen when sensing the movement of one or more facial parts of a user and changing the appearance of one or more facial parts of a corresponding avatar based on the sensed movement of the one or more facial parts.
[0163] In FIG. 11, the control unit 190 of the terminal device 10 displays, on the display 1302, an information display screen 1101, a user video 1102, a setting screen 1103, an avatar 1104, and the like.
[0164] The information display screen 1101 is a screen that displays the sensing results of the user's face parts, associated parts of the face parts, presets for the degree of change in appearance, etc. At this time, the control unit 190 of the terminal device 10 may accept the following selections from the user. -Selection of the parts of the user's face on which sensing will be performed -Selection of the parts to be associated from among the sensed parts -Select candidate degree of change This allows the user to reduce the number of sensing locations in some cases, thereby reducing the load during distribution.
[0165] The user video 1102 is a screen that displays an image of the user himself / herself captured via the camera 160 provided in the terminal device 10. The control unit 190 of the terminal device 10 displays an image of the user himself / herself in the user video 1102 by the camera 160 provided in the terminal device 10.
[0166] The setting screen 1103 is a screen for the user to set the degree of change in the appearance of the avatar. The control unit 190 of the terminal device 10 presents the following settings to the user, for example, and accepts input. Mouth switching speed Eye movement: Maximum upward movement Eye movement: Maximum downward movement Eye movement: Maximum horizontal movement Eye Movement: Sensitivity At this time, the control unit 190 of the terminal device 10 may accept a setting of the degree to which the state of one or more facial parts of the avatar is made to follow the sensed results, within a range associated with the degree of change in the performer's speech. For example, the control unit 190 may accept the setting from the user so that the setting is lower than the degree of change of the avatar (amount of change in the object, speed of change in the object) estimated from the user's speech. At this time, if the user attempts to set a value or the like that is not within the settable range, the control unit 190 may display a predetermined alert, or if the setting screen is a slider type or the like, may lock the value in advance so that it does not become that value. This allows the user to move the avatar more slowly than the user's own speech, thereby smoothing the degree of change in the avatar felt by the viewer, thereby giving the viewer a greater sense of immersion.
[0167] The avatar 1104 is an avatar that changes its appearance based on settings received from a user. When the control unit 190 of the terminal device 10 receives settings on the setting screen 1103 from the user, the control unit 190 may synchronize the user video 1102 and the avatar 1104 and display them to the user. This allows the user to check in advance for any discomfort or other issues that may arise when changing the appearance of the avatar based on the user's own settings.
[0168] FIG. 12 shows an example of a screen when one or more emotion candidates of a user are estimated, and the appearance of one or more facial parts of a corresponding avatar is changed based on the emotion selected by the user.
[0169] In FIG. 12, the control unit 190 of the terminal device 10 displays, on a display 1302, an information display screen 1201, a user video 1202, a setting screen 1203, an avatar 1204, and the like.
[0170] 11, the information display screen 1201 is a screen that displays the sensing results of the parts of the user's face, the parts of the face that are associated with each other, presets for the degree of change in appearance, etc. In addition, the control unit 190 may display information on one or more candidate emotions of the user identified from the sensing results on the screen. When control unit 190 receives a selection of emotion candidates from the user, it reflects the degree of change in the appearance of the avatar that corresponds to the emotion. For example, the control unit 190 may acquire and associate in advance from the user the sensing results of the facial parts corresponding to the user's emotion. Thereafter, the control unit 190 senses the user's face via the camera 160 or the like, and determines whether all or a part of the sensing results of the face included in the associated emotion match. Thereafter, the control unit 190 may present candidates for the user's emotion based on the determination result, accept a selection from the user, and change the facial state of the avatar based on the selected emotion. At this time, if the user's emotion cannot be estimated, control unit 190 may change the appearance of one or more facial parts based on a setting preset by the user. This allows the user to select emotion candidates so that the change in the avatar's appearance reflects the user's emotion, even if accurate sensing cannot be performed.
[0171] The user video 1202 is a screen that displays an image of the user himself / herself captured via the camera 160 provided in the terminal device 10, similar to the user video 1102 in FIG.
[0172] The setting screen 1203, like the setting screen 1103 in FIG. 11, is a screen for the user to set the degree of change in the appearance of the avatar.
[0173] Avatar 1204, like avatar 1104 in FIG. 11, is an avatar that changes appearance based on settings received from the user.
[0174] FIG. 13 shows an example of a screen for setting the degree of change in the avatar's appearance when sensing results cannot be obtained for at least one associated part out of one or more parts of the user's face.
[0175] In FIG. 13, the control unit 190 of the terminal device 10 displays, on the display 1302, an information display screen 1351, a user video 1352, a setting screen 1353, an avatar 1354, and the like.
[0176] 12, the information display screen 1351 is a screen that displays the sensing results of the parts of the user's face, the parts of the face that are associated with each other, presets for the degree of change in appearance, etc. In addition, the control unit 190 may display information on accessories, attachments, etc. that are worn by the user and that cover a part of the user's face on the screen. For example, when a user is wearing an eye patch or the like and sensing of one eye is difficult or impossible, the control unit 190 may reflect the degree of change in the other eye from which the sensing result was obtained. This allows the avatar corresponding to the user to change its appearance without being affected by an eye patch or the like even if the user is wearing an eye patch or the like.
[0177] The user video 1352 is a screen that displays an image of the user himself / herself captured via the camera 160 provided in the terminal device 10, similar to the user video 1202 in FIG.
[0178] The setting screen 1353, like the setting screen 1203 in FIG. 12, is a screen for the user to set the degree of change in the appearance of the avatar.
[0179] Avatar 1354, like avatar 1204 in FIG. 12, is an avatar that changes appearance based on settings received from the user.
[0180] FIG. 14 shows an example of a screen when correcting the degree of change in the appearance of an avatar when a user is wearing a wearable device such as glasses.
[0181] In FIG. 14, the control unit 190 of the terminal device 10 displays, on the display 1302, an information display screen 1401, a user video image 1402, a setting screen 1403, an avatar 1404, and the like.
[0182] 13, the information display screen 1401 is a screen that displays the sensing results of the parts of the user's face, the parts of the face that are associated with each other, presets for the degree of change in appearance, etc. In addition, the control unit 190 may display information on the wearable device worn by the user, information on the correction amount for the degree of change for each wearable device, etc. on the screen. For example, control unit 190 acquires information about the wearable device worn by the user by referring to wearable device information 1803 or wearable device information database 2023. Then, control unit 190 may execute a process similar to that of change correction module 2040 in server 20 described above to correct the degree setting.
[0183] The user video 1402 is a screen that displays an image of the user himself / herself captured via the camera 160 provided in the terminal device 10, similar to the user video 1352 in FIG.
[0184] The setting screen 1403, like the setting screen 1353 in FIG. 13, is a screen for the user to set the degree of change in the appearance of the avatar.
[0185] Avatar 1404, like avatar 1354 in FIG. 13, is an avatar that changes appearance based on settings received from the user.
[0186] FIG. 15 shows an example of a screen in which the state of the avatar's mouth is changed based on the degree of change in speech when the movement of the user's mouth cannot be sensed.
[0187] In FIG. 15, the control unit 190 of the terminal device 10 displays, on the display 1302, an information display screen 1501, a user video 1502, a setting screen 1503, an avatar 1504, and the like.
[0188] 14, the information display screen 1501 is a screen that displays the sensing results of the parts of the user's face, the parts of the face that are associated with each other, presets for the degree of change in appearance, etc. In addition, the control unit 190 may display information on a mask worn by the user, information on a voice spectrum acquired from the user's speech, etc. on the screen. For example, when the user is wearing a mask or the like over the mouth and the control unit 190 cannot sense the movement of the mouth, the control unit 190 may change the state of the avatar's mouth based on the voice spectrum of the user's speech rather than the results of sensing the user, as described above. This allows a user to change the state of the avatar's mouth to match the user's speech, for example, even when the user has to broadcast while wearing a mask.
[0189] The user video 1502 is a screen that displays an image of the user himself / herself captured via the camera 160 provided in the terminal device 10, similar to the user video 1402 in FIG.
[0190] The setting screen 1503, like the setting screen 1403 in FIG. 14, is a screen for the user to set the degree of change in the appearance of the avatar.
[0191] Avatar 1504, like avatar 1404 in FIG. 14, is an avatar that changes its appearance based on settings received from the user.
[0192] FIG. 16 shows an example screen that is displayed when a specific notification is presented to the user when the difference in degree settings between one or more facial parts of an avatar that are previously associated with each other exceeds a specific threshold.
[0193] In FIG. 16, the control unit 190 of the terminal device 10 displays, on the display 1302, an information display screen 1601, a user image 1602, a setting screen 1603, an avatar 1604, and the like.
[0194] Information display screen 1601, like information display screen 1501 in FIG. 15, is a screen that displays the results of sensing of parts of the user's face, associated parts of the face, and preset candidates (presets) for the degree of change in appearance that has been previously set.
[0195] The user video 1602 is a screen that displays an image of the user himself / herself captured via the camera 160 provided in the terminal device 10, similar to the user video 1502 in FIG.
[0196] The setting screen 1603 is a screen for the user to set the degree of change in the appearance of the avatar, similar to the setting screen 1503 in Fig. 15. At this time, when the control unit 190 receives an input of the degree of change in the appearance of a part of the face from the user on this screen, if the degree of change in a paired or related part (both eyes, etc.) is too large (for example, the amount of change in one eye is too large), the control unit 190 may display that the setting value for that part is abnormal and may display recommended settings.
[0197] Avatar 1604 is an avatar that changes its appearance based on settings received from the user, similar to avatar 1504 in Fig. 15. When control unit 190 receives an input of the degree of change in the appearance of the face from the user on setting screen 1503 described above, if the degree of change in a paired or related part is too large, control unit 190 may present the part in a different appearance (for example, in a different color) together with a notification to the user. In this case, the different appearance presented by control unit 190 may be, for example, a change in color, a pop-up notification, or the shape of the relevant part, but is not limited thereto.
[0198] This allows the user to visually determine if an abnormal value has been input when making settings to change the appearance of parts of the avatar's face, thereby preventing the viewer from feeling uncomfortable.
[0199] FIG. 17 shows an example of a screen displayed when presenting a predetermined notification to a user, showing how at least one or a plurality of facial parts change when the difference in degree is set within a predetermined range.
[0200] 17, the control unit 190 of the terminal device 10 displays, on the display 1302, a setting screen 1701, an avatar 1702, a setting preview screen 1703, an avatar preview screen 1704, and the like.
[0201] The setting screen 1603 is a screen for the user to set the degree of change in the appearance of the avatar, similar to the setting screen 1503 in Fig. 15. At this time, when the control unit 190 receives an input of the degree of change in the appearance of a part of the face from the user on this screen, if the degree of change in a paired or related part (both eyes, etc.) is too large (for example, the amount of change in one eye is too large), the control unit 190 may display that the setting value for that part is abnormal and may display recommended settings.
[0202] The setting preview screen 1703 is a screen that displays recommended settings when there is an abnormal value in the degree of change between paired, related parts of the avatar's face, etc., on the setting screen 1701. The control unit 190 of the terminal device 10 displays, on the setting preview screen 1703, the degree of change of a setting mode different from the setting input on the setting screen 1701. At this time, the control unit 190 may display numerical values, objects, etc., in a mode different from that displayed on the setting screen 1701 (for example, a different color, size, shape, etc.).
[0203] The avatar preview screen 1704 is a screen that displays an avatar that reflects the settings recommended on the setting preview screen 1703. For example, the control unit 190 of the terminal device 10 may display, on a screen different from the screen that displays the above-mentioned notification, how the appearance of the avatar changes when the difference in degree is within an appropriate range (a range that does not cause discomfort to the viewer). This allows the user to check, when the degree of change in behavior that he or she has set exceeds a specified threshold, how the behavior would change if the user set it to an appropriate value.
[0204] <7 Variation> A modified example of this embodiment will be described below. That is, the following aspects may be adopted. (1) An information processing device in which the program may be pre-installed or may be installed later, such a program may be stored on an external non-transitory storage medium, or may be operated by cloud computing. (2) A method in which a computer functions as an information processing device, and the program may be pre-installed on the information processing device or may be installed later, such a program may be stored on an external non-transitory storage medium, or may be run by cloud computing.
[0205] <6 Notes> The matters described in the above embodiments will be supplemented below.
[0206] (Appendix 1) A program executed by a computer 20 having a processor 29, the program causing the processor 29 to execute the steps of: acquiring an audio spectrum of a performer's speech (S501); changing the mouth appearance of an avatar corresponding to the performer in accordance with the performer's speech based on the acquired audio spectrum (S502); presenting the avatar corresponding to the performer and the performer's voice to an audience (S503); and accepting a setting for the degree to which the mouth appearance of the avatar is changed in accordance with the performer's speech to be lower than the change in the performer's speech (S504), wherein in the changing step (S502), the mouth appearance of the avatar is changed in accordance with the setting.
[0207] (Appendix 2) 2. The program according to claim 1, wherein in the receiving step (S504), a setting of the degree is received from the performer, allowing the setting to be lower than the performer's speaking rate.
[0208] (Appendix 3) The program according to claim 1 or 2, wherein in the receiving step (S504), a setting of a frequency range for detecting the voice spectrum is received, and in the changing step, in response to detecting the voice spectrum in the set range, the state of the avatar's mouth is changed based on a first setting of the degree.
[0209] (Appendix 4) The program described in Appendix 3, wherein in the changing step (S502), in response to detecting a voice spectrum outside a set range, the state of the avatar's mouth is changed based on a second setting that is a predetermined degree setting and different from the first setting.
[0210] (Appendix 5) 5. The program according to any one of appendixes 1 to 4, wherein in the changing step (S502), the state of the mouth is changed based on at least one of the group consisting of loudness and low / highness of the voice spectrum.
[0211] (Appendix 6) The program further causes processor 29 to execute the steps of estimating one or more candidate emotions of the performer, presenting the one or more estimated candidate emotions of the performer to the performer, and accepting an input operation from the performer to select one of the one or more candidate emotions of the performer, and in the changing step (S502), if a selection of an emotion is accepted from the performer, the state of the avatar's mouth is changed based on the selected emotion. The program described in any of Appendices 1 to 6.
[0212] (Appendix 7) The program according to claim 6, wherein, if the emotion cannot be estimated in the estimating step, the mouth state is changed based on a condition preset by the performer in a changing step (S502).
[0213] (Appendix 8) The program of claim 6, further causing the processor 29 to execute a step of activating a body part other than the avatar's mouth based on the estimated emotion.
[0214] (Appendix 9) The program of claim 6, further comprising causing the processor 29 to execute a step of activating a part of the avatar's body other than the mouth based on an emotion selected by the performer.
[0215] (Appendix 10) The program further comprises causing the processor to change the mouth posture based on the speaking rate rather than the degree of change set by the performer when the speaking rate of the performer deviates by a predetermined rate from the speaking rate estimated from the degree of change in the mouth posture set by the performer.
[0216] (Appendix 11) 11. The program according to any one of appendices 1 to 10, wherein in the receiving step (S504), avatar attributes are received from the performer, and an amount of change in the mouth shape is corrected based on the attributes.
[0217] (Appendix 12) The program according to claim 11, wherein in the receiving step (S504), information on either a human or a non-human object whose mouth shape changes differently from a human's is received as an attribute, and the amount of change in the mouth shape is corrected based on the attribute.
[0218] (Appendix 13) The program described in any of Appendices 1 to 12, further causing the processor 29 to perform the steps of presenting to the performer one or more candidate body parts other than the avatar's mouth for which the behavior is to be changed based on the emotion estimated from the acquired voice spectrum, and changing the behavior of the part in response to accepting a selection from the performer of the part for which the behavior is to be changed.
[0219] (Appendix 14) A method executed by a computer 20 having a processor 29, the method comprising the steps of: a step (S501) of acquiring an audio spectrum of a performer's speech; a step (S502) of changing the mouth appearance of an avatar corresponding to the performer in accordance with the performer's speech based on the acquired audio spectrum; a step (S503) of presenting the avatar corresponding to the performer and the performer's voice to an audience; and a step (S504) of accepting a setting for the degree to which the mouth appearance of the avatar is changed in accordance with the performer's speech, the degree being lower than the change in the performer's speech; and in the changing step (S502), the mouth appearance of the avatar is changed in accordance with the setting.
[0220] (Appendix 15) An information processing device 20 equipped with a control unit 203, wherein the control unit 203 executes the steps of: acquiring an audio spectrum of a performer's speech (S501); changing the mouth appearance of an avatar corresponding to the performer in accordance with the performer's speech based on the acquired audio spectrum (S502); presenting the avatar corresponding to the performer and the performer's voice to an audience (S503); and accepting a setting for the degree to which the mouth appearance of the avatar is changed in accordance with the performer's speech to be lower than the change in the performer's speech (S504); and in the changing step (S502), changing the mouth appearance of the avatar in accordance with the setting. [Explanation of symbols]
[0221] 10 terminal device, 12 communication interface, 13 input device, 14 output device, 15 memory, 16 storage unit, 19 processor, 20 server, 22 communication interface, 23 input / output interface, 25 memory, 26 storage, 29 processor, 80 network, 1801 user information, 1802 avatar information, 1803 wearable device information, 1901 input operation reception unit, 1902 transmission / reception unit, 1903 data processing unit, 1904 notification control unit, 1302 display, 140 voice processing unit, 141 microphone, 142 speaker, 150 position information sensor, 160 camera, 170 motion sensor, 2021 user information database, 2022 avatar information database, 2023 wearable device information database, 2031 reception control module, 2032 transmission control module, 2033 user information acquisition module, 2034 an avatar information acquisition module, 2035 a voice spectrum acquisition module, 2036 an avatar change module, 2037 an avatar presentation module, 2038 a setting reception module, 2039 a wearable device information acquisition module, and 2040 a change correction module.
Claims
1. A program executed by a computer including a processor, the program causing the processor to acquire the voice of the speaker's speech; change the mode of the avatar corresponding to the speaker according to the speech of the speaker based on the acquired voice; present the avatar corresponding to the speaker and the voice of the speaker to the viewer; receive a setting of the degree of changing the mode of the avatar according to the speech of the speaker; and in the changing step, change the mode of the avatar according to the setting, the program.
2. The program according to claim 1, wherein in the receiving step, it is possible to receive from the speaker a setting that the degree of the setting is lower than the change in the speech of the speaker.
3. The program according to claim 1, wherein in the receiving step, it is possible to receive from the speaker a setting that the degree of the setting is lower than the speech rate of the speaker.
4. The program according to claim 1, wherein in the changing step, at least one of the mouth of the avatar corresponding to the speaker, a facial part other than the mouth, and a body part is changed according to the speech of the speaker.
5. The program according to claim 1, wherein in the changing step, the mouth of the avatar corresponding to the speaker and a part other than the mouth are changed according to the speech of the speaker.
6. in the receiving step, receiving a setting of a frequency range for detecting a voice spectrum included in the voice, in the changing step, in response to detecting the voice spectrum in the set range, changing the mode of the avatar based on the first setting of the degree, the program according to claim 1.
7. The program according to claim 6, wherein in the changing step, in response to detecting the voice spectrum outside the set range, changing the mode of the avatar based on a second setting that is a predetermined setting of the degree and different from the first setting.
8. The program according to claim 1, wherein in the changing step, the mode of the avatar is changed based on at least one of the strength or pitch of the voice spectrum included in the voice.
9. The program further causes the processor to perform steps of estimating one or more emotion candidates of the performer, presenting the estimated one or more emotion candidates of the performer to the performer, and receiving an input operation from the performer for selecting one emotion from the one or more emotion candidates of the performer. The program according to claim 1, wherein, in the step of changing, when the selection of the emotion is received from the performer, the aspect of the avatar is changed based on the selected emotion.
10. The program according to claim 9, wherein, in the step of estimating, when an emotion cannot be estimated, in the step of changing, the aspect of the avatar is changed based on conditions preset by the performer.
11. The program according to claim 9, further causing the processor to perform a step of operating at least one of the mouth of the avatar, a facial part other than the mouth, and a body part based on the estimated emotion.
12. The program according to claim 9, further causing the processor to perform a step of operating at least one of the mouth of the avatar, a facial part other than the mouth, and a body part based on the emotion selected by the performer.
13. The program according to claim 1, wherein, when the speech rate of the performer deviates from a predetermined rate from the speech rate estimated from the degree of change in the aspect of the avatar set by the performer, the aspect of the avatar is changed based on the speech rate instead of the degree of change set by the performer.
14. In the step of receiving, The program according to claim 1, wherein an attribute of the avatar is received from the performer, and the amount of change in the aspect of the avatar is corrected based on the attribute.
15. The program according to claim 14, wherein, as the attribute, information of either a human or a non-human whose mode of change in aspect is different from that of a human is received, and the amount of change in the aspect of the avatar is corrected based on the attribute.
16. The program further causes the processor to present to the performer at least one candidate among the mouth of the avatar that changes its aspect, facial parts other than the mouth, and body parts, based on the emotion estimated from the acquired voice. The program according to claim 1, further causing the processor to execute a step of changing the aspect of the selected part in response to receiving a selection of the part whose aspect is to be changed from the performer.
17. A method executed by a computer including a processor, the method comprising the steps of: the processor acquiring voice of the performer's speech; changing, according to the acquired voice, the aspect of the avatar corresponding to the performer in response to the performer's speech; presenting the avatar corresponding to the performer and the voice of the performer to the viewer; receiving a setting of the degree to which the aspect of the mouth of the avatar is changed in response to the performer's speech; and in the changing step, changing the aspect of the avatar according to the setting.
18. An information processing apparatus including a control unit, the control unit acquiring voice of the performer's speech; changing, according to the acquired voice, the aspect of the mouth of the avatar corresponding to the performer in response to the performer's speech; presenting the avatar corresponding to the performer and the voice of the performer to the viewer; receiving a setting of the degree to which the aspect of the mouth of the avatar is changed in response to the performer's speech; and in the changing step, changing the aspect of the avatar according to the setting.