Program, method, and information processing apparatus

By synchronizing avatar mouth movements with user speech through facial tracking and customizable adjustments, the system addresses unnatural mouth movements, enhancing viewer immersion.

JP2026020248APending Publication Date: 2026-02-06COVER CORP
View PDF 11 Cites 0 Cited by

Patent Information

Application Number
JP2025197114
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-11-18
Publication Date
2026-02-06

AI Technical Summary

Technical Problem

Existing technologies that change an avatar's mouth shape based on voice input often result in unnatural movements, causing discomfort to viewers due to discrepancies between the user's mouth movements and the avatar's responses.

Method used

A system that synchronizes avatar mouth movements with user speech by tracking facial movements and adjusting the mouth shape based on an acquired audio spectrum, allowing for customizable and gradual changes that match the user's speech patterns.

Benefits of technology

The system enhances the natural appearance of avatar mouth movements, providing a more immersive experience for viewers by aligning avatar responses with user speech dynamics.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026020248000001_ABST
    Figure 2026020248000001_ABST
Patent Text Reader

Abstract

To provide a technique for making a change in a mode of a mouth of an avatar look more natural.SOLUTION: A program executed by a computer including a processor, the program causing the processor to execute a step of acquiring a sound spectrum of a speech of a performer, a step of changing a mode of a mouth of an avatar corresponding to the performer in accordance with the speech of the performer on the basis of the acquired sound spectrum, a step of presenting the avatar corresponding to the performer and a sound of the performer to a viewer / listener, and a step of accepting a setting of a degree of changing the mode of the mouth of the avatar in accordance with the speech of the performer such that the degree of changing the mode of the mouth of the avatar is lower than a degree of change in the speech of the performer, changing a form of a mouth of the avatar according to a setting.SELECTED DRAWING: Figure 7
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to a program, a method, and an information processing device. [Background technology]

[0002] There is known a technique for reflecting a user's facial expression or the like on an avatar in real time.

[0003] Patent document 1 describes a technology that estimates which Japanese vowel (aiueo) is being spoken from the combination of the first and second formant frequencies of a person's voice, and changes the shape of an avatar's mouth to correspond to each vowel. [Prior art documents] [Patent documents]

[0004] [Patent Document 1] JP 2016-126500 A Summary of the Invention [Problem to be solved by the invention]

[0005] The technology disclosed in Patent Document 1 estimates which Japanese vowel is being spoken based on the voice picked up from a microphone, and determines and changes the shape and size of the avatar's mouth. However, the technology in Patent Document 1 simply changes the shape of the avatar's mouth in response to sound. For example, when the mouth is changed continuously, the mouth movement is different from the mouth movement when a person speaks, and it instantly changes to a shape corresponding to a vowel, which may cause discomfort to the viewer / user. Therefore, there is a need for technology that can make avatar mouth changes appear more natural. [Means for solving the problem]

[0006] According to one embodiment, there is provided a program executed by a computer having a processor, the program causing the processor to execute the steps of: acquiring an audio spectrum of a performer's speech; changing the mouth shape of an avatar corresponding to the performer in accordance with the performer's speech based on the acquired audio spectrum; presenting the avatar corresponding to the performer and the performer's voice to an audience; and accepting a setting for the degree to which the avatar's mouth shape is changed in accordance with the performer's speech, which setting can be lower than the degree to which the performer's speech changes; and in the changing step, the program changes the avatar's mouth shape in accordance with the setting. [Effects of the Invention]

[0007] According to the present disclosure, it is possible to provide a technology that makes changes in the state of an avatar's mouth appear more natural. [Brief explanation of the drawings]

[0008] [Figure 1] FIG. 1 is a block diagram showing the overall configuration of a system 1. [Figure 2] FIG. 2 is a diagram illustrating a functional configuration of a terminal device 10. [Figure 3] FIG. 2 is a diagram showing the functional configuration of a server 20. [Figure 4] 1 shows the data structures of a user information database (DB), an avatar information DB, and a wearable device information DB stored in the storage unit of server 20. [Figure 5] 10 is a flowchart showing a series of processes for acquiring a voice spectrum of a user's speech and, based on the acquired voice spectrum, changing the mouth shape of an avatar corresponding to the user in accordance with the speech of the performer. [Figure 6] 10 is an example of a screen when a user registers the audio spectrum of his or her own vowels in the system 1. [Figure 7] 10 shows an example of a screen when a user sets the degree of change in the appearance of an avatar's mouth or facial features. [Figure 8]10 shows an example screen in which one or more candidate emotions of the user are estimated from the user's speech, and the appearance of the avatar is changed based on the estimated one or more emotions of the user. [Figure 9] 10 shows an example of a screen on which a user can make various settings based on the voice spectrum, etc. for an avatar with attributes different from those of a human. [Figure 10] A flowchart showing a series of processes for sensing the movement of one or more facial parts of a user's face and changing the appearance of one or more facial parts of an avatar corresponding to the user based on the sensed movement of the one or more facial parts. [Figure 11] An example screen is shown in which the movement of one or more parts of a user's face is sensed, and the appearance of one or more parts of the face of the corresponding avatar is changed based on the sensed movement of the one or more parts of the face. [Figure 12] 10 shows an example of a screen when one or more candidate emotions of a user are estimated, and the appearance of one or more facial parts of a corresponding avatar is changed based on the emotion selected by the user. [Figure 13] 10 shows an example of a screen for setting the degree of change in the avatar's appearance when sensing results cannot be obtained for at least one associated part of one or more parts of the user's face. [Figure 14] 10 shows an example of a screen when correcting the degree of change in the appearance of an avatar when a user is wearing a wearable device such as glasses. [Figure 15] 10 shows an example of a screen in which the state of the avatar's mouth is changed based on the degree of change in speech when the user's mouth movement cannot be sensed. [Figure 16] This shows an example screen that displays a specified notification to the user when the difference in degree settings between one or more facial parts of an avatar that are pre-associated exceeds a specified threshold. [Figure 17]10 shows an example of a screen that shows the user how at least one or more facial parts change when the difference in degree is set within a predetermined range when a predetermined notification is presented to the user. DETAILED DESCRIPTION OF THE INVENTION

[0009] Hereinafter, an embodiment of the present invention will be described with reference to the drawings. In the following description, the same components are denoted by the same reference numerals. The names and functions of the components are also the same. Therefore, detailed description thereof will not be repeated.

[0010] First Embodiment <Summary> In the following embodiment, a technique for changing the state of the mouth of an avatar based on the voice spectrum of a user who is an actor operating the avatar will be described. Here, there are no limitations on the devices etc. that may be used as appropriate when realizing the technology disclosed herein, and the device may be a terminal device such as a smartphone or tablet device owned by the user, or it may be presented from a stationary PC (Personal Computer).

[0011] There is a known technology that controls the mouth movements of an avatar based on the user's voice captured through a sound collection device such as a microphone. However, in this system, the user's actual mouth movements and the avatar's movements do not accurately synchronize, which can cause viewers to feel uncomfortable.

[0012] Therefore, System 1 provides a technology that makes changes in the avatar's mouth shape appear more natural.

[0013] System 1 can be used, for example, in situations such as live streaming distribution on video distribution sites, where an avatar that tracks the movements of a user (performer) is used. For example, system 1 tracks the user's movements via a camera (image capture device) provided on a terminal device (such as a PC) used by the user and reflects the user's movements in the avatar's movements. System 1 also acquires the audio spectrum of the performer's speech via a microphone (sound collection device) also provided on the user's terminal device, and changes the mouth shape of the avatar corresponding to the performer based on the acquired audio spectrum in accordance with the performer's speech. In this case, the system 1 presents an avatar corresponding to the performer and the performer's voice to the audience, and accepts a setting for the degree to which the avatar's mouth movements change in response to the performer's speech, which can be set to a degree lower than the degree to which the performer's speech changes. By executing this process, the system 1 may change the avatar's mouth movements in response to the performer's speech. This allows the changes in the avatar's mouth shape to appear more natural.

[0014] <1 Overall system configuration> FIG. 1 shows the overall configuration of a system 1 according to the first embodiment.

[0015] As shown in Fig. 1, the system 1 includes a plurality of terminal devices (Fig. 1 shows terminal device 10A and terminal device 10B. Hereinafter, they may be collectively referred to as "terminal device 10." Furthermore, a plurality of terminal devices 10C and the like may also be included in the configuration.) and a server 20. The terminal devices 10 and the server 20 are connected for communication via a network 80.

[0016] The terminal device 10 is a device operated by each user. The terminal device 10 is realized by a mobile terminal such as a smartphone or tablet compatible with a mobile communication system. Alternatively, the terminal device 10 may be, for example, a stationary personal computer (PC) or a laptop PC. As shown as terminal device 10B in FIG. 1, the terminal device 10 includes a communication IF (Interface) 12, an input device 13, an output device 14, a memory 15, a storage unit 16, and a processor 19. The server 20 includes a communication IF 22, an input / output IF 23, a memory 25, a storage 26, and a processor 29.

[0017] The terminal device 10 is communicatively connected to the server 20 via a network 80. The terminal device 10 is connected to the network 80 by communicating with communication devices such as a wireless base station 81 conforming to communication standards such as 5G and LTE (Long Term Evolution), and a wireless LAN router 82 conforming to a wireless LAN (Local Area Network) standard such as IEEE (Institute of Electrical and Electronics Engineers) 802.11.

[0018] The communication IF 12 is an interface for inputting and outputting signals so that the terminal device 10 can communicate with external devices. The input device 13 is an input device (e.g., a touch panel, a touch pad, a pointing device such as a mouse, a keyboard, etc.) for receiving input operations from a user. The output device 14 is an output device (e.g., a display, a speaker, etc.) for presenting information to a user. The memory 15 is for temporarily storing programs and data processed by the programs, etc., and is a volatile memory such as a DRAM (Dynamic Random Access Memory). The storage unit 16 is a storage device for saving data, such as a flash memory or an HDD (Hard Disc Drive). The processor 19 is hardware for executing an instruction set written in a program, and is composed of an arithmetic unit, registers, peripheral circuits, etc.

[0019] The server 20 manages information set by users when they perform live streaming using avatars, etc. The server 20 stores, for example, information about users, information about avatars, and information about wearable devices worn by users.

[0020] The communication IF 22 is an interface for inputting and outputting signals so that the server 20 can communicate with external devices. The input / output IF 23 functions as an interface with an input device for receiving input operations from a user and an output device for presenting information to the user. The memory 25 is for temporarily storing programs and data processed by the programs, etc., and is a volatile memory such as a DRAM (Dynamic Random Access Memory). The storage 26 is a storage device for saving data, such as a flash memory or an HDD (Hard Disc Drive). The processor 29 is hardware for executing an instruction set written in a program, and is composed of an arithmetic unit, registers, peripheral circuits, etc.

[0021] In this embodiment, each device (terminal device, server, etc.) can also be considered as an information processing device. That is, a collection of devices can be considered as one "information processing device," and system 1 can be formed as a collection of multiple devices. The way in which multiple functions required to realize system 1 according to this embodiment are allocated to one or multiple pieces of hardware can be determined appropriately in consideration of the processing capacity of each piece of hardware and / or the specifications required for system 1.

[0022] <1.1 Configuration of the terminal device 10> FIG. 2 is a block diagram of a terminal device 10 constituting the system 1 of the first embodiment. As shown in FIG. 2, the terminal device 10 includes a plurality of antennas (antenna 111, antenna 112), wireless communication units (first wireless communication unit 121, second wireless communication unit 122) corresponding to the respective antennas, an operation reception unit 130 (including a touch-sensitive device 1301 and a display 1302), an audio processing unit 140, a microphone 141, a speaker 142, a position information sensor 150, a camera 160, a motion sensor 170, a storage unit 180, and a control unit 190. The terminal device 10 also has functions and configurations (e.g., a battery for storing power, a power supply circuit for controlling the supply of power from the battery to each circuit, etc.) that are not specifically shown in FIG. 2. As shown in FIG. 2, the blocks included in the terminal device 10 are electrically connected by a bus or the like.

[0023] The antenna 111 emits a signal emitted by the terminal device 10 as a radio wave. The antenna 111 also receives a radio wave from space and provides the received signal to the first radio communication unit 121.

[0024] The antenna 112 emits a signal emitted by the terminal device 10 as a radio wave. The antenna 112 also receives a radio wave from space and provides the received signal to the second radio communication unit 122.

[0025] The first wireless communication unit 121 performs modulation / demodulation processing and the like for transmitting and receiving signals via the antenna 111 so that the terminal device 10 can communicate with other wireless devices. The second wireless communication unit 122 performs modulation / demodulation processing and the like for transmitting and receiving signals via the antenna 112 so that the terminal device 10 can communicate with other wireless devices. The first wireless communication unit 121 and the second wireless communication unit 122 are communication modules including a tuner, an RSSI (Received Signal Strength Indicator) calculation circuit, a CRC (Cyclic Redundancy Check) calculation circuit, a high-frequency circuit, etc. The first wireless communication unit 121 and the second wireless communication unit 122 perform modulation / demodulation and frequency conversion of wireless signals transmitted and received by the terminal device 10, and provide the received signals to the control unit 190.

[0026] The operation reception unit 130 has a mechanism for receiving input operations from a user. Specifically, the operation reception unit 130 is configured as a touch screen and includes a touch-sensitive device 1301 and a display 1302. The touch-sensitive device 1301 receives input operations from a user of the terminal device 10. The touch-sensitive device 1301 detects the user's touch position on the touch panel, for example, by using a capacitive touch panel. The touch-sensitive device 1301 outputs a signal indicating the user's touch position detected by the touch panel to the control unit 190 as an input operation. The terminal device 10 may also be provided with a physical keyboard (not shown) that can be used for input, and may receive the user's input operations via the keyboard.

[0027] The display 1302 displays data such as images, videos, and text under the control of the control unit 190. The display 1302 is realized by, for example, an LCD (Liquid Crystal Display) or an organic EL (Electro-Luminescence) display.

[0028] The audio processing unit 140 modulates and demodulates audio signals. The audio processing unit 140 modulates a signal provided from the microphone 141 and provides the modulated signal to the control unit 190. The audio processing unit 140 also provides the audio signal to the speaker 142. The audio processing unit 140 is realized, for example, by a processor for audio processing. The microphone 141 accepts audio input and provides an audio signal corresponding to the audio input to the audio processing unit 140. The speaker 142 converts the audio signal provided from the audio processing unit 140 into audio and outputs the audio to the outside of the terminal device 10.

[0029] The position information sensor 150 is a sensor that detects the position of the terminal device 10, and is, for example, a GPS (Global Positioning System) module. The GPS module is a receiving device used in a satellite positioning system. In the satellite positioning system, signals are received from at least three or four satellites, and the current position of the terminal device 10 equipped with the GPS module is detected based on the received signals. The position information sensor 150 may be a transmitting / receiving device based on a communication standard used in a short-range communication system between information devices. Specifically, the position information sensor 150 uses the 2.4 GHz band, such as a Bluetooth (registered trademark) module, to receive beacon signals from other information devices equipped with a Bluetooth (registered trademark) module.

[0030] The camera 160 is a device that receives light with a light receiving element and outputs the received light as a captured image. The camera 160 is, for example, a depth camera that can detect the distance from the camera 160 to a subject being photographed. Furthermore, the camera 160 acquires the body movements of the user who uses the terminal device 10. Specifically, for example, the camera 160 acquires the movements of the user's mouth and the movements of each part of the face (eyes, eyebrows, etc.). Any existing technology may be used to acquire the movements.

[0031] The motion sensor 170 is configured with a gyro sensor, an acceleration sensor, etc., and detects the tilt of the terminal device 10.

[0032] The storage unit 180 is configured with, for example, a flash memory or the like, and stores data and programs used by the terminal device 10. In one aspect, the storage unit 180 stores user information 1801, avatar information 1802, wearable device information 1803, etc. This information may be held in the storage unit 180 of the terminal device 10, or may be stored as a database in the storage unit 202 of a server (described later) and acquired via the network 80.

[0033] User information 1801 includes information such as an ID for identifying a user, a user name, and information about an avatar corresponding to the user. Here, a user refers to a performer who controls an avatar based on information acquired via microphone 141 or camera 160. Details of the information included in the user information will be described later.

[0034] Avatar information 1802 is various information related to an avatar corresponding to a user. For example, avatar information 1802 holds information such as the corresponding user and settings that the user normally uses, and is information that is referenced by the user to smoothly operate the avatar during distribution such as live streaming. Details of the information included in the avatar information will be described later. The settings that the user normally uses are parameters and conditions that the user can adjust when broadcasting using an avatar, such as the basic settings for the degree of change in the avatar's appearance, the emotions displayed as default in normal broadcasts, the user's sensing sensitivity, etc.

[0035] Wearable device information 1803 is various information related to the wearable device the user is wearing at the time of distribution. The various information includes, for example, the following: Types of wearable devices Wearable device size -Transmittance of wearable devices Whether or not information can be obtained electronically The wearable device information 1803 holds various information related to various instruments and devices worn by the user, such as eyeglasses, eyewear such as smart glasses, a head-mounted display (HMD), etc. Details of the information included in the wearable device information 1803 will be described later.

[0036] The control unit 190 reads a program stored in the storage unit 180 and executes instructions included in the program to control the operation of the terminal device 10. The control unit 190 is, for example, an application processor. By operating in accordance with the program, the control unit 190 fulfills the functions of an input operation reception unit 1901, a transmission / reception unit 1902, a data processing unit 1903, and a notification control unit 1904.

[0037] The input operation receiving unit 1901 performs processing to receive a user's input operation on an input device such as the touch-sensitive device 131. The input operation receiving unit 1901 determines the type of operation, such as whether the user's operation is a flick operation, a tap operation, or a drag (swipe) operation, based on information about the coordinates where the user has touched the touch-sensitive device 1301 with a finger or the like.

[0038] The transmitting / receiving unit 1902 performs processing for the terminal device 10 to transmit and receive data to and from an external device such as the server 20 in accordance with a communication protocol.

[0039] The data processing unit 1903 performs calculations on data that the terminal device 10 has received as input in accordance with a program, and outputs the calculation results to a memory or the like.

[0040] The data processing unit 1903 receives the movement of the user's mouth and the like acquired by the camera 160, and controls the execution of various processes. For example, the data processing unit 1903 executes a process of controlling the movement of the mouth of an avatar corresponding to the user, based on the movement of the user's mouth acquired by the camera 160.

[0041] The notification control unit 1904 performs processing for displaying a display image on the display 132, processing for outputting sound from the speaker 142, and processing for causing the camera 160 to generate vibrations.

[0042] <1.2 Functional configuration of server 20> 3 is a diagram showing the functional configuration of the server 20. As shown in FIG. 3, the server 20 functions as a communication unit 201, a storage unit 202, and a control unit 203.

[0043] The communication unit 201 performs processing for the server 20 to communicate with external devices.

[0044] The storage unit 202 stores data and programs used by the server 20. The storage unit 202 stores a user information database 2021, an avatar information database 2022, a wearable device information database 2023, and the like.

[0045] The user information database 2021 is a database for storing various information related to the performers who operate the avatars. Details of each record stored in this database will be described later.

[0046] The avatar information database 2022 is a database for storing various information related to the avatars operated by the users, as will be described in detail later.

[0047] The wearable device information database 2023 is a database for storing various information about the eyewear worn by the user operating the avatar. Details will be described later.

[0048] The control unit 203 is composed of, for example, a processor 29, which performs processing according to a program to perform functions of various modules such as a reception control module 2031, a transmission control module 2032, a user information acquisition module 2033, an avatar information acquisition module 2034, an audio spectrum acquisition module 2035, an avatar change module 2036, an avatar presentation module 2037, a setting reception module 2038, a wearable device information acquisition module 2039, and a change correction module 2040.

[0049] The reception control module 2031 controls the process by which the server 20 receives signals from external devices in accordance with a communication protocol.

[0050] The transmission control module 2032 controls the process in which the server 20 transmits signals to external devices in accordance with a communication protocol.

[0051] The user information acquisition module 2033 controls the process of acquiring various information about the user who is the performer operating the avatar. The various information includes, for example, the following: User name, identification ID Avatar information corresponding to the user Devices worn by the user (glasses, etc.) Specifically, for example, the user information acquisition module 2033 may acquire the information by referring to the user information 1801 from the storage unit 180 of the terminal device 10 used by the user. Also, the user information acquisition module 2033 may acquire the information by referring to a user information database 2021 stored in the storage unit 202 of the server 20, which will be described later. Alternatively, the user information acquisition module 2033 may acquire the information by accepting input of various pieces of information about the user directly from the user.

[0052] The avatar information acquisition module 2034 acquires various information about the avatar operated by the user. The various information includes, for example, the following: ID information to identify the avatar Avatar attributes (human, non-human, etc.) Information about the mouth, face, or other body parts that the corresponding user has set as default Dedicated settings for mouth, face or other body parts, individually configured for each avatar Specifically, for example, the avatar information acquisition module 2034 may acquire information about the avatar associated with each user by referring to the avatar information 1802 or the user information database 2021. In addition, in a certain aspect, server 20 may receive from a user an input of a setting for the degree of change in the appearance of the avatar's mouth, face, or other body parts, and upon receiving an operation to reset the setting as a default, perform a process of updating the avatar information held in avatar information 1802, etc. This allows the user to appropriately set a frequently used setting of the degree of change in the appearance of the avatar as the default, making it easier to operate the avatar.

[0053] It should be noted that the avatar information does not have to be one for each user. For example, information on a plurality of avatars may be associated with a user in advance, or additional avatars may be associated with the user.

[0054] In addition, in some aspects, the avatar ID may include the following information: Information about the appearance of your avatar (gender, eye color, hairstyle, mouth, size of facial or other body parts, hair color, skin color, etc.) Information about the state of the avatar's mouth, face, or other body parts (such as the number of types of changes in state, the amount of change in state, etc.) Here, the server 20 may store the avatar information in association with the type of content. Specifically, for example, the server 20 may associate the type of content (chat, singing, acting, etc.) that the user provides in live streaming or the like with settings of the degree of change in the appearance of the mouth, face, and other body parts, and when the server 20 receives an operation by the user to select which content to provide, the server 20 may reflect the setting corresponding to the content in the avatar.

[0055] This allows the user to appropriately change the appearance of the avatar in accordance with the content provided to the viewer, thereby giving the viewer a sense of immersion.

[0056] Furthermore, after receiving a selection of an avatar to be used from the user, the server 20 may receive the selection from the user and simultaneously present a predetermined notification (dialog or the like) to the user, rather than automatically reflecting the setting of the degree of change in appearance. For example, after the user selects an avatar to be used, the server 20 may present a notification such as "Is the normal setting OK?" when the user selects a general-purpose setting that is normally used, rather than a dedicated setting that is set for each avatar.

[0057] This prevents the user from reflecting incorrect settings in live distribution, thereby preventing viewers from losing their sense of immersion.

[0058] Furthermore, the server 20 may accept and reflect the setting of the degree of change in appearance during live streaming. Specifically, for example, when the server 20 accepts a change in the setting of the degree of change in appearance from a user during live streaming, the server 20 may execute a process of changing the appearance of the avatar based on the changed setting when changing the appearance of the avatar based on the user's voice spectrum or the user's sensing results acquired a predetermined time after accepting the change in setting. This allows the user to change and reflect the settings for the appearance change as appropriate during live streaming, so that even if the content provided during live streaming is switched, the user can see the change in the avatar's appearance without feeling uncomfortable.

[0059] The voice spectrum acquisition module 2035 controls processing for acquiring a voice spectrum of a user's speech. Specifically, for example, the voice spectrum acquisition module 2035 controls processing for acquiring a voice spectrum from a voice uttered by the user acquired via the microphone 141. For example, the voice spectrum acquisition module 2035 acquires the user's voice via the microphone 141 and acquires the voice spectrum contained in the voice. For example, the voice spectrum acquisition module 2035 may perform a Fourier transform on the voice acquired from the microphone 141 to acquire information on the voice spectrum contained in the voice. In this case, the calculation for acquiring the voice spectrum is not limited to a Fourier transform and may be any existing method. In addition, in a certain aspect, the voice spectrum acquisition module 2035 may acquire information on the voice spectrum of vowels from the user's voice. For example, the voice spectrum acquisition module 2035 accepts a vowel setting input by the user in advance, and then accepts an utterance from the user via the microphone 141, thereby storing the accepted vowel setting and the acquired voice spectrum in association with each other. In addition, in one aspect, the voice spectrum acquisition module 2035 may acquire voice information resulting from consonants, such as sounds "t," "c," "h," "k," "m," "r," "s," "n," and "w," and combine it with the stored vowel information to estimate the words spoken by the user. This allows the system 1 to separately characterize and store the voice spectrum relating to vowels among the user's voice spectrum, thereby enabling more accurate changes to the movement of the avatar's mouth.

[0060] The avatar change module 2036 controls a process of changing the mouth shape of an avatar corresponding to a performer in accordance with the performer's speech based on the acquired voice spectrum. Specifically, for example, the avatar change module 2036 estimates words spoken by the user from the user's voice spectrum acquired by the voice spectrum acquisition module 2035, and changes the mouth shape of the avatar in accordance with the estimated words. For example, the avatar change module 2036 estimates information about vowels spoken by the user from the user's voice spectrum, and changes the mouth shape in accordance with the vowels. For example, if the user's voice spectrum acquired by the voice spectrum acquisition module 2035 is "a," the avatar change module 2036 changes the mouth shape of the avatar to a shape corresponding to "a."

[0061] The avatar presentation module 2037 controls the process of presenting an avatar corresponding to a performer and the user's voice to a viewer. Specifically, for example, the avatar presentation module 2037 transmits an image of an avatar corresponding to a user and the user's voice to the display 1302 and speaker 142 of the terminal device 10 used by the viewer, and presents them to the viewer. In this case, the viewer is not limited to one person, and the avatar and voice may be presented to the terminal device 10 of multiple viewers.

[0062] The setting receiving module 2038 controls a process for receiving a setting for the degree to which the avatar's mouth behavior is changed in response to the user's utterance, so that the setting can be set to a degree that is lower than the degree of change in the user's utterance. Specifically, for example, the setting receiving module 2038 receives from the user a setting for the time interval for reflecting the user's utterance in the avatar's mouth behavior as the setting for the degree to which the avatar's mouth behavior is changed. For example, the setting receiving module 2038 may receive settings including the following: - Setting the time required for the avatar's mouth to change to a state corresponding to the speech estimated from the user's voice spectrum · Change the avatar's behavior based on the user's speech within a certain period · Set the update frequency (e.g., number of updates per second) Here, the change in the user's speech will be defined. The change in the user's speech is, for example, the speed of the user's speech, and may be calculated based on the following. The time interval between changes in the vowels spoken by the user (for example, the time interval between changes in the vowels from "a" to "i") At this time, the server 20 may simultaneously acquire sounds derived from consonants (c, k, etc.), and may estimate the speech rate by assuming that different words are being spoken even when the same vowel is acquired consecutively. - Number of vowels uttered in a given period At this time, the setting receiving module 2038 may receive the setting to a degree lower than the degree of change of the avatar estimated from the change in the user's speech. For example, the setting receiving module 2038 may receive a time required to change the mouth posture to correspond to the speech (vowel) estimated in advance from the user's voice spectrum. The server 20 acquires information on the time taken for the user to utter a vowel from the user's voice spectrum based on the received required time, calculates a ratio to a predetermined required time, multiplies the ratio by the amount of change in the posture, and calculates the amount of change in the posture of the avatar's mouth. The server 20 changes the posture of the mouth based on the acquired speech time and amount of change. For example, suppose the user inputs a setting that the avatar's mouth changes to the posture of "a" in "1 second." For example, the time when the posture of "a" is completely changed is set to "100," and the setting is made so that it takes "1 second" to reach "100." At this time, the server 20 may also receive and record from the user the degree of change in the manner of speech (amount of change in the mouth, speed) in one second (i.e., a difference may be set between the amount of change in the manner of speech in the first 0.5 seconds and the remaining 0.5 seconds of the change in the manner of speech in one second). When the user utters the sound "a" for one second, the server 20 changes the state of the avatar's mouth to the state of "a" over one second based on the setting of the amount of change, etc. However, when the user utters "a" for only 0.5 seconds, the server 20 may perform processing to change the amount of change in the state of the avatar's mouth up to "50."

[0063] Furthermore, when the user speaks continuously (for example, when the user speaks "aiueo"), the server 20 may acquire the speaking time of each vowel and perform the above processing. That is, the server 20 may calculate the amount of change in the state of the avatar's mouth corresponding to each vowel from the speaking time of each vowel, and change the state of the avatar's mouth. For example, if the time required for the avatar's mouth to change to the state corresponding to each vowel is "1 second" and "a" is spoken in "0.2 seconds" and "i" is spoken in "0.3 seconds," the amount of change corresponding to "a" is "20" and the amount of change corresponding to "i" is "30." Furthermore, when the user speaks for longer than the required time, the server 20 may maintain the state of the avatar's mouth even after the required time has passed. This allows the user to change the avatar's mouth shape gradually according to the time the user is speaking, rather than instantly changing the avatar's mouth shape to "ah" when the user utters the sound "ah." Also, by setting a required time and multiplying the change in the mouth shape by the amount of change in the mouth shape when the user speaks less than that time, it is possible to prevent the avatar's mouth shape from changing significantly even when the user speaks lightly (for example, even if the mouth opens by about 30 degrees, the avatar's mouth shape will change as if it were 100 degrees). This allows the user to prevent viewers from feeling uncomfortable due to the user's speech and the change in the avatar's mouth shape, thereby giving the viewer a greater sense of immersion.

[0064] In some aspects, the server 20 may change the mouth shape of the avatar based on information such as the size and pitch of the voice spectrum acquired from the user. Specifically, for example, the server 20 may acquire information on the frequency (Hz) and sound pressure (dB) of the user's voice spectrum and change the mouth shape of the avatar when the information exceeds a predetermined threshold. For example, assume that the server 20 receives a setting to change the mouth shape of the avatar in a required time of "1 second" and the user's speaking time is "1 second." In this case, if the server 20 detects that the user has spoken at a sound pressure exceeding the threshold at "0.8 seconds," the server 20 may change the mouth shape of the avatar to a larger degree than usual (to a mouth-wide open shape). In this case, the server 20 may reflect similar settings not only on the mouth but also on other facial and body parts. This allows the user to have the avatar's mouth movements reflect the user's sudden loud voice, allowing the viewer to see more natural avatar movements.

[0065] Additionally, the setting receiving module 2038 may receive a setting for the degree of change in the avatar's mouth movement so that the setting is lower than the update frequency of the avatar's movement estimated from the speech rate estimated from the user's speech. The server 20 then sends the information set by the setting reception module 2038 to the avatar change module 2036, changes the avatar's mouth shape according to the settings, and then presents the avatar and the user's voice to the viewer by the avatar presentation module 2037. This allows the user to change the avatar's mouth movements more gradually than vowel changes, allowing the avatar's mouth movements to change more smoothly to match the user's own speech. This allows the user to prevent the avatar's mouth movements from moving too delicately and appearing unnatural, allowing the viewer to see more natural mouth movements and increasing the viewer's sense of immersion.

[0066] The wearable device information acquisition module 2039 controls the process of acquiring information about the wearable device worn by the user. Specifically, for example, after acquiring user information, the wearable device information acquisition module 2039 refers to the wearable device information database 2023 (described later) to acquire information about the wearable device worn by the user. The server 20 transmits the acquired wearable device information to the change correction module 2040.

[0067] The change correction module 2040 controls the process of correcting the setting of the degree of change in appearance to be reflected in the avatar, based on the wearable device information acquired by the wearable device information acquisition module 2039. Specifically, for example, the change correction module 2040 corrects the setting of the degree of change in appearance of parts of the user's face that are covered or obstructed by the wearable device, based on the wearable device information acquired by the wearable device information acquisition module 2039. The server 20 acquires the sensing results of predetermined parts of the user's face (mouth, eyes, eyebrows, nose, etc.), and changes the appearance of the avatar by reflecting the sensing results and settings received from the user (degree to reflect the sensing results, parameter settings, etc.). In this case, for example, if the user is wearing glasses, when the change correction module 2040 receives from the user the degree of change in the state of the eyes of the avatar corresponding to the user based on that information, it corrects and reflects the change based on a preset correction value. Correction refers to, for example, when the accuracy of sensing parts of the user's face decreases for each wearable device, setting the rate of decrease (or attenuation rate) in advance and correcting the degree to which the movement is reflected in the avatar during sensing and tracking based on that setting. This allows the user to present changes in the avatar's appearance to the viewer in a natural way, even if the user is wearing glasses or the like.

[0068] Additionally, the change correction module 2040 may correct the degree of change in the appearance of the avatar according to the attributes of the avatar acquired by the avatar information acquisition module 2034 . Specifically, for example, the change correction module 2040 may acquire information on whether the avatar operated by the user is a human or a non-human whose behavior changes differently from humans, and may execute a process to correct the degree of change in the avatar's behavior based on the information. For example, if the attribute of the avatar operated by the user is a "dragon," the movements of the eyes, mouth, etc. may behave differently from humans. In that case, the change correction module 2040 may correct the amount of change in the corners of the mouth, the amount of change in the eyeballs, etc., to a shape that is consistent with the avatar's behavior based on the attribute of the "dragon." This allows users to present more natural movements to viewers based on the results of their own speech and facial sensing, even when operating an avatar that is not human.

[0069] Note that the above configuration is not essential to the embodiments of the present disclosure. That is, the terminal device 10 may play the role of the server 20 and execute the same processes as the various modules that configure the control unit 203 of the server 20. Furthermore, the terminal device 10 may execute the various functions disclosed in the present invention based on information acquired via the microphone 141, camera 160, etc. provided in the terminal device 10, without going through the network 80.

[0070] <2 Data Structure> FIG. 4 is a diagram showing the data structures of the user information database 2021, the avatar information database 2022, and the wearable device information database 2023 stored in the server 20. As shown in FIG.

[0071] As shown in FIG. 4, the user information database 2021 includes an item "ID," an item "compatible avatar," an item "device used," an item "dedicated preset (mouth)," an item "dedicated preset (face)," an item "basic settings," an item "frequently used emotions," an item "notes," and the like.

[0072] The item "ID" is information that identifies each user who is an actor operating an avatar.

[0073] The item "corresponding avatar" is information that identifies each avatar that corresponds to each user.

[0074] The item "Device in use" is information that identifies a device worn by each user, for example, a wearable device worn by each user.

[0075] The item "Dedicated Preset (Mouth)" is information indicating conditions preset for each user regarding the degree to which the avatar's mouth shape is changed when each user operates the avatar. Specifically, for example, it indicates various conditions individually set for each avatar when the avatar operated by the user is in a predetermined situation (e.g., when the mouth shape is significantly changed). Information included in the preset may include, for example, information such as the height of the corners of the mouth and the shape of the lips. By accepting a selection of the preset from the user, the server 20 may reflect the setting in the avatar and present it to the viewer. This allows the user to instantly reflect changes in the mouth shape that are unique to the avatar corresponding to the user and present them to the viewer, allowing the viewer to see the avatar moving more naturally.

[0076] The item "Dedicated Preset (Face)" is information indicating conditions preset for each user regarding the degree to which the appearance of the avatar's facial features is changed when each user operates the avatar. Specifically, for example, it indicates various conditions individually set for each avatar when the avatar operated by the user is in a predetermined situation (e.g., when the avatar's facial expression changes significantly). Information included in the preset may include, for example, the direction of the eyebrows, the shape of the eyes, the size of the pupils, the flushing of the cheeks, and sensing information on the user's speech or facial expression. By accepting a selection of the preset from the user, the server 20 may reflect the setting in the avatar and present it to the viewer. For example, suppose a user uses an avatar with non-human attributes (such as a monster, an inorganic object, or a robot). In this case, changes in the appearance of the avatar's various parts (mouth, face, body) may not completely match the user's voice spectrum and sensing results. For this reason, the server 20 may accept settings for the dedicated preset (mouth) or dedicated preset (face) exemplified above from the user. By selecting such a preset during live streaming or the like, the user can seamlessly reflect the user's voice spectrum and sensing results in changes in the appearance of the avatar, even when using any avatar.

[0077] The server 20 may also accept dedicated preset settings according to the type of content provided by the user. For example, when a user distributes a song, the server 20 may accept settings that change the avatar's mouth, facial features, and body features more significantly than when the user is chatting. This allows the user to instantly reflect changes in the facial features specific to the avatar corresponding to the user and present them to the viewer, allowing the viewer to see the avatar moving more naturally.

[0078] The "Basic Settings" item indicates the setting for the degree of change that the user typically uses. Specifically, for example, it indicates the conditions for the degree of change that the user operating the avatar typically (generally) uses when changing the appearance of the mouth, face, or other body parts during everyday broadcasting, live broadcasting, or live streaming. For example, the conditions may include the degree of tracking of the user's sensing results. The degree of tracking of the sensing results includes, for example, the degree of sensitivity, where 100 is the value when the sensing results are directly reflected in the change in the avatar's appearance, the amount of change in the avatar's appearance compared to the amount of change in the user's face, and the degree of change in the avatar's movement relative to the amount of change in the avatar per unit time estimated from the sensing results. This allows the user to easily start distribution without having to set the degree of change each time distribution is performed.

[0079] At this time, the server 20 may combine the basic settings and the dedicated presets and accept them as settings according to the content. Specifically, for example, the following are examples of avatar settings according to the content. ASMR (Autonomous Sensory Meridian Response) mode (whisper mode) Use the dedicated preset for the mouth (lower sensitivity, softer voice), and use the basic settings for facial expressions. Alternatively, use the dedicated facial expression settings in combination. Action game streaming mode A special preset is used for the mouth (which increases sensitivity and causes over-reactions), while also increasing sensitivity for facial expressions. Horror game streaming mode Use a dedicated preset for the mouth (lower sensitivity and lower frequency threshold for detection), and set the same sensitivity for facial expressions. Alternatively, use dedicated settings. Chat mode (uses basic settings)

[0080] In addition, in one aspect, the server 20 may present a switching button to the user to switch the above modes, and by accepting the user's pressing of the button, reflect the setting of the degree of change in appearance based on the mode in the avatar. At this time, the server 20 may present the switching button to the user in a state that is invisible to the viewer but visible to the user. Also, the server 20 may change the arrangement of the switching button in response to an operation by the user. This allows users to use different presets depending on the content they wish to provide to viewers, enabling a wider range of expression.

[0081] The server 20 may also store information relating to the background in the virtual space in association with the object, and may change the background in response to a mode change. Alternatively, the server 20 may store a predetermined object, such as the following example, in association with the object, and display the object in the virtual space in response to a mode change. - Equipment objects such as microphones and instruments for live music streaming Game console objects when streaming games General objects (houseplants, room furniture, etc.) This allows the server 20 to reduce the load of the reading process when switching modes, and prevents delays and other issues that may cause discomfort to viewers.

[0082] The above settings may be used in combination with basic settings, etc. The combination may be any setting accepted from the user, or may be stored in a storage unit as a dedicated combination for each user. The server 20 may also acquire information on the frequency of use of multiple presets. Based on the information on the frequency of use, the server 20 may notify the user whether to save a frequently used preset as a "frequently used setting" or a "basic setting." When the server 20 accepts an instruction from the user to set a "frequently used setting," etc., the server 20 may save the preset as a "frequently used setting" in a storage unit.

[0083] The "frequently used emotion" item indicates the setting of an emotion that the user frequently uses when operating the avatar. Specifically, for example, if the user frequently uses the emotion "joy" during distribution, the server 20 may store in advance in the database conditions for changing the avatar's appearance based on that emotion. In this case, the conditions for changing the appearance include the degree of change in the mouth appearance, the amount of change in various facial parts that move when expressing the emotion "joy," the degree to which the avatar follows the sensing results, etc. When the server 20 receives a selection of the stored emotion setting from the user, the server 20 may change the appearance of the avatar based on the emotion and present it to the viewer. This allows the user to instantly set changes in the avatar's appearance according to the emotions used in normal broadcasts, making broadcasting easier.

[0084] The item "Notes" is information that is stored when there are special notes or the like in the user information.

[0085] As shown in FIG. 4, the avatar information database 2022 includes an item "ID," an item "corresponding user," an item "attributes," an item "associated body part," an item "presence or absence of special body part," an item "motion setting of special body part," an item "standard change speed," an item "frequently used emotions," an item "remarks," and the like.

[0086] The item "ID" is information used for distribution and identifying each avatar presented to viewers.

[0087] The item "Corresponding User" is information that identifies the user corresponding to the avatar.

[0088] The "attribute" item is information that identifies the attribute set for each avatar. Specifically, the attribute indicates, for example, information that specifies whether the avatar is human or non-human, whose appearance changes differently from that of a human. Attributes include, for example, the following information: ·human - Living things other than humans (animals, plants, etc.) Imaginary creatures (dragons, angels, demons, etc.) ·machine Formless entities (such as slimes and ghosts in fantasy) In some cases, the record may store information specific to the defined attribute as a subordinate concept. Specifically, for example, if the attribute is "inorganic matter," the record may store a subordinate concept such as "no eyes," and if the attribute is "virtual creature," the record may store information such as "has multiple eyes." Server 20 may store information for correcting the degree of change in the avatar's appearance based on the attribute information. This allows the user to appropriately change the appearance of the mouth and face even when operating a non-human avatar.

[0089] The "associated parts" item is information about associated parts of one or more facial parts of an avatar. Specifically, the associated parts may hold information about, for example, "eyebrows" among the facial parts of an avatar that are associated with each other. The server 20 may accept settings for the same degree of change in appearance for the associated parts. This eliminates the need for the user to individually set the degree of change in appearance for each associated body part, thereby reducing the time and effort required to set the degree of change in appearance.

[0090] The "presence or absence of special body parts" item is information for identifying whether or not the avatar has a special body part. Specifically, for example, if the avatar has attributes other than human and has body parts such as "horns" or "tail," the server 20 may store this information. Here, the special body part does not have to belong to the avatar's body, but may be an object floating around the avatar. The special body part is not limited to the above. For example, an object such as a living thing different from the avatar may be placed around the avatar.

[0091] The item "Special Body Part Operation Settings" is information related to the settings for operating a special body part of the avatar. Specifically, for example, if the avatar has a special body part (e.g., "horns," "tail," etc.), the server 20 may store in this record information on what conditions trigger the operation of that body part. For example, for an avatar with a special body part "horn," if the operation is set to "linked with the movement of the entire eye," the server 20 may change the state of the horn by reflecting the setting of the degree of change in the state of the eye set by the user. In addition, in a certain aspect, the server 20 may accept a setting from the user of the degree of change in appearance for each special body part. For example, if the special body part is an object that is not connected to the body of the avatar but that floats around the avatar and changes its appearance, the server 20 may accept a setting input from the user for each of the objects. However, the server 20 may also change the appearance of the objects in accordance with the settings of the avatar's body parts (mouth, face, etc.).

[0092] Furthermore, when the special body part is an object such as a living thing different from the avatar and exists around the avatar, the server 20 may accept a setting of the degree to which the body part of the object (for example, eyes, mouth, etc.) changes its appearance based on the user's voice spectrum or sensing results. For example, the server 20 may set the amount of change in the object's eyes by multiplying the amount of change in the avatar by a predetermined percentage, or may accept a setting from the user for each body part of the object. This allows the user to perform operations that suit the characteristics of the avatar, even when operating a non-human avatar.

[0093] The item "Notes" is information that is stored when there are special notes or the like in the avatar information.

[0094] As shown in FIG. 4, the wearable device information database 2023 includes an item "ID," an item "Type," an item "Detection accuracy," an item "Correction amount," and an item "Remarks."

[0095] The item "ID" is information that identifies each wearable device worn by the user.

[0096] The "Type" item is information indicating the type of wearable device worn by the user. The wearable device worn by the user is not particularly limited, and may be eyewear such as glasses, or a head-mounted device such as an HMD.

[0097] The "detection accuracy" item indicates the detection accuracy of sensing the movement of the user's eyes or face when the user is wearing a wearable device. Specifically, for example, the server 20 may score the detection accuracy of sensing for each wearable device worn by the user and store the information. For example, if the user is wearing glasses with high transmittance that are almost the same as the naked eye, the information may be stored as a detection accuracy of "Good." In this case, the score stored by the server 20 may not be a symbol such as "Good," but may be a number such as "100" based on transmittance or the like, or may be a notation such as "A" or "Good," and is not limited thereto.

[0098] The "Correction Amount" item indicates the correction amount for the degree of change in the avatar, which is set for each wearable device. Specifically, for example, the server 20 sets the correction amount for the degree of change in the avatar's appearance based on the aforementioned detection accuracy value. For example, if the user is wearing glasses, the server 20 may execute a process of multiplying the degree of change by a predetermined magnification based on the transmittance of the glasses. In some cases, the server 20 may not perform any special correction processing if the user is wearing a device such as an HMD and the sensing results of the user's eyes or face can be obtained from the device, even if the detection accuracy is low. The information about the wearable device held by the server 20 may also be information about a mask, an eye patch, etc. In this case, the server 20 may reflect the movement of a part of the body that is covered by a mask, an eye patch, etc. in the avatar by reflecting the user's speech or the settings of other parts that are not covered, rather than by changing the appearance based on the sensing results. This allows the user to broadcast without worrying about how they look during broadcasting.

[0099] The "Notes" item is information that is stored when there are special notes or the like in the wearable device information.

[0100] <3 operations> The following describes a series of processes performed by the system 1 when acquiring the voice spectrum of a user's speech and changing the mouth shape of an avatar corresponding to the user in accordance with the speech of the performer based on the acquired voice spectrum.

[0101] 5 is a flowchart showing a series of processes for acquiring a voice spectrum of a user's speech and, based on the acquired voice spectrum, changing the mouth shape of an avatar corresponding to the user in accordance with the speech of the performer. Note that this flowchart discloses an example in which the control unit 190 of the terminal device 10 used by the user executes the series of processes, but this is not limiting. That is, the terminal device 10 may transmit some information to the server 20 and the server 20 may execute the processes, or the server 20 may execute the entire series of processes.

[0102] In step S501, the control unit 190 of the terminal device 10 acquires a voice spectrum of the speech of the user who is the performer operating the avatar. Specifically, for example, the control unit 190 of the terminal device 10 controls the process of acquiring a voice spectrum from the voice uttered by the user acquired via the microphone 141, similar to the voice spectrum acquisition module 2035 of the server 20. For example, the control unit 190 acquires the user's voice via the microphone 141 and acquires the voice spectrum included in the voice. For example, the control unit 190 may perform a Fourier transform on the voice acquired from the microphone 141 to acquire information on the voice spectrum included in the voice. In this case, the calculation for acquiring the voice spectrum is not limited to a Fourier transform and may be any existing method. In addition, in a certain aspect, control unit 190 may acquire information on the voice spectrum of vowels from the user's voice. For example, control unit 190 accepts a vowel setting input by the user in advance, and then accepts an utterance from the user via microphone 141, thereby storing the accepted vowel setting and the acquired voice spectrum in association with each other. In addition, in one aspect, the voice spectrum acquisition module 2035 may acquire voice information resulting from consonants, such as sounds "t," "c," "h," "k," "m," "r," "s," "n," and "w," and combine it with the stored vowel information to estimate the words spoken by the user. This allows the system 1 to separately characterize and store the voice spectrum relating to vowels among the user's voice spectrum, thereby enabling more accurate changes to the movement of the avatar's mouth.

[0103] In step S502, the control unit 190 of the terminal device 10 changes the mouth shape of an avatar corresponding to the user in accordance with the user's utterance based on the acquired voice spectrum. Specifically, for example, the control unit 190 of the terminal device 10, like the avatar change module 2036 of the server 20, estimates words uttered by the user from the acquired user's voice spectrum and changes the mouth shape of the avatar in accordance with the estimated words. For example, when the acquired user's voice spectrum is "a," the control unit 190 changes the mouth shape of the avatar to a shape corresponding to "a."

[0104] In step S503, the control unit 190 of the terminal device 10 presents an avatar corresponding to the user and the user's voice to the viewer. Specifically, for example, the control unit 190 of the terminal device 10, like the avatar presentation module 2037 of the server 20, transmits an image of the avatar corresponding to the user and the user's voice to the display 1302 and speaker 142 of the terminal device 10 used by the viewer, and presents them to the viewer. At this time, the viewer is not limited to one, and the avatars and voices may be presented to the terminal devices 10 of multiple viewers.

[0105] In step S504, the control unit 190 of the terminal device 10 accepts a setting for the degree to which the avatar's mouth shape is changed in response to the performer's speech, which can be set to a degree lower than the degree of change in the user's speech. Specifically, for example, the control unit 190 of the terminal device 10 may accept settings including the following, similar to the setting acceptance module 2038 of the server 20. · Change the avatar's behavior based on the user's speech within a certain period · Set the update frequency (e.g., number of updates per second) Here, the change in the user's speech will be defined. The change in the user's speech is, for example, the speed of the user's speech, and may be calculated based on the following. The time interval between changes in the vowels spoken by the user (for example, the time interval between changes in the vowels from "a" to "i") At this time, the control unit 190 may simultaneously acquire sounds derived from consonants (c, k, etc.), and may estimate the speech rate by assuming that different words are being spoken even when the same vowel is acquired consecutively. - Number of vowels uttered in a given period Furthermore, at this time, the control unit 190 may accept the setting to a degree lower than the degree of change of the avatar estimated from the change in the user's speech. For example, the control unit 190 may accept a time required for changing the mouth posture to correspond to the speech (vowel) estimated in advance from the user's voice spectrum. The control unit 190 acquires information on the time it takes the user to speak a vowel from the user's voice spectrum based on the accepted required time, calculates a ratio to a preset required time, multiplies the ratio by the amount of change in the posture, and calculates the amount of change in the posture of the avatar's mouth. The control unit 190 changes the posture of the mouth based on the acquired speech time and amount of change. For example, suppose the user inputs a setting that the avatar's mouth changes to the posture of "a" in "1 second." For example, the time when the posture of "a" is completely changed is set to "100," and the setting is made so that it takes "1 second" to reach "100." At this time, the control unit 190 may also receive and record from the user the degree of change in the manner in one second (amount of change in the mouth, speed). (That is, when the manner of the mouth changes in one second, a difference may be set in the amount of change in the manner between the first 0.5 seconds and the remaining 0.5 seconds.) When the user utters the sound "a" for one second, the control unit 190 changes the state of the avatar's mouth to the state of "a" over one second based on the setting of the amount of change, etc. However, when the user utters "a" for only "0.5 seconds," the control unit 190 may perform processing to change the amount of change in the state of the avatar's mouth up to "50."

[0106] Furthermore, when the user speaks continuously (for example, when the user speaks "aiueo"), the control unit 190 may acquire the speaking time of each vowel and perform the above processing. That is, the control unit 190 may calculate the amount of change in the state of the avatar's mouth corresponding to each vowel from the speaking time of each vowel, and change the state of the avatar's mouth. For example, if the time required for the avatar's mouth to change to the state corresponding to each vowel is "1 second" and "a" is spoken in "0.2 seconds" and "i" is spoken in "0.3 seconds," the amount of change corresponding to "a" is "20" and the amount of change corresponding to "i" is "30." Furthermore, when the user speaks for longer than the required time, the control unit 190 may maintain the state of the avatar's mouth even after the required time has passed. This allows the user to change the avatar's mouth shape gradually according to the time the user is speaking, rather than instantly changing the avatar's mouth shape to "ah" when the user utters the sound "ah." Also, by setting a required time and multiplying the change in the mouth shape by the amount of change in the mouth shape when the user speaks less than that time, it is possible to prevent the avatar's mouth shape from changing significantly even when the user speaks lightly (for example, even if the mouth opens by about 30 degrees, the avatar's mouth shape will change as if it were 100 degrees). This allows the user to prevent viewers from feeling uncomfortable due to the user's speech and the change in the avatar's mouth shape, thereby giving the viewer a greater sense of immersion.

[0107] In a certain aspect, the control unit 190 may change the state of the avatar's mouth based on information such as the size and pitch of the voice spectrum acquired from the user. Specifically, for example, the control unit 190 may acquire information on the frequency (Hz) and sound pressure (dB) of the user's voice spectrum, and change the state of the avatar's mouth when the information exceeds a predetermined threshold. For example, assume that the control unit 190 receives a setting to change the state of the avatar's mouth in a required time of "1 second," and the user's speaking time is "1 second." In this case, if the control unit 190 detects that the user has spoken at a sound pressure exceeding the threshold at "0.8 seconds," the control unit 190 may change the state of the avatar's mouth to a larger degree than usual (to a state with the mouth wide open). In this case, the control unit 190 may reflect similar settings not only on the mouth, but also on facial and body parts. This allows the user to have the avatar's mouth movements reflect the user's sudden loud voice, allowing the viewer to see more natural avatar movements.

[0108] Additionally, the control unit 190 may accept a setting for the degree of change in the state of the avatar's mouth so that the setting is lower than the update frequency of the avatar's movement estimated from the speech rate estimated from the user's speech. For example, the control unit 190 divides the user's speech into fixed time intervals and changes the avatar to have a mouth shape that corresponds to the first and last vowels in that time interval. For example, if the speech changes to "aiueo" in one second, the control unit 190 may reflect in the avatar the shape of the mouth at the timing of the first "a" of "aiueo" and the mouth shape at the timing of "o." Alternatively, when the control unit 190 temporarily stores the user's speech in a buffer memory, the control unit 190 may change the mouth posture from "a" to "o" over a certain time interval (for example, one second). The control unit 190 may also accept a setting so that the avatar's mouth posture changes slower than the time elapsed when the user's vowel changes. For example, when the user's vowel changes from "a" to "u" and the change takes one second, the server 20 may take 1.5 seconds for the avatar's mouth posture to change from "a" to "u." At this time, the server 20 may also execute a process to complement the change in posture. That is, instead of instantly changing the avatar's mouth posture from "a" to "u," the server 20 may change the mouth posture by passing through a mouth shape intermediate between "a" and "u." This allows the user to change the avatar's mouth movement in a manner that is closer to the mouth movements of a real person, rather than having the avatar's mouth change instantly for each word, thereby reducing the sense of discomfort felt by viewers when watching the avatar. At this time, the control unit 190 may accept from the user a setting of the degree that can be set lower than the user's speech rate. Specifically, for example, the control unit 190 may calculate the user's speech rate from the audio spectrum of the speech accepted from the user. Thereafter, the control unit 190 accepts from the user a setting of the degree that can be set lower than the user's speech rate by setting an upper limit value for the amount of change per unit time of the avatar's appearance that can be accepted from the user based on the calculated user's speech rate. This allows the user to set the degree of change of the avatar to be slower than the change in the user's own speech, and allows the avatar's appearance to change more smoothly.

[0109] In step S505, the control unit 190 of the terminal device 10 changes the mouth shape of the avatar in accordance with the settings. Specifically, for example, the control unit 190 of the terminal device 10 changes the mouth shape of the avatar in accordance with the settings based on the information set in step S604, and then presents the avatar and the user's voice to the viewer. This allows the user to smoothly change the state of the avatar's mouth in accordance with the user's speech, and allows the viewer to see more natural mouth movements. In one aspect, when changing the state of the mouth of the avatar, control unit 190 of terminal device 10 may change the state of the mouth based on at least one of the group consisting of strength and weakness and pitch of the voice spectrum. Specifically, for example, the control unit 190 determines the strength and pitch by analyzing the following parameters of the audio spectrum: Audio spectrum intensity parameter: decibels (dB) Audio spectrum high / low parameter: Hertz (Hz) For example, when the control unit 190 acquires an audio spectrum that is louder than the reference audio spectrum in decibels, the control unit 190 may change the state of the avatar's mouth more than the change in the state of the mouth at the time of reference. This allows the user to change the appearance of the avatar based on subtle changes in the voice, reducing the sense of discomfort felt by the viewer.

[0110] In one aspect, the control unit 190 of the terminal device 10 may accept a setting of a frequency range for detecting a voice spectrum, and in response to detecting a voice spectrum in the set range, may change the manner of the avatar's mouth based on a first setting of the degree. Specifically, for example, in step S604, the control unit 190 accepts, from the user, setting of upper and lower limit values ​​as the frequency range for detecting a voice spectrum. The control unit 190 analyzes the voice spectrum of the user's speech acquired via the microphone 141 and determines whether the frequency of the voice spectrum is within the range. If the frequency is within the range, the control unit 190 may change the manner of the avatar based on the first setting of the degree, i.e., the setting of the degree of change of the avatar's manner that is set in advance by the user, in step S605.

[0111] In addition, in a certain aspect, in response to detecting a voice spectrum outside the set range, the control unit 190 of the terminal device 10 may change the manner of the avatar's mouth based on a second setting that is a predetermined setting and different from the first setting. Specifically, for example, when the control unit 190 detects a frequency outside the frequency range for detecting the voice spectrum received from the user, the control unit 190 may change the manner of the avatar based on a setting (second setting) that is different from the normal setting (first setting). For example, if the user utters a voice outside the range of frequencies normally used (e.g., an extreme shriek), the voice spectrum falls outside the detection range. In this case, the control unit 190 may change the manner of the avatar by applying a setting (second setting) that is applied only to frequencies outside the detection range, rather than the change setting (first setting) received from the user. This allows the user to change the appearance of the avatar accordingly even when the user moves or speaks in a way that is different from normal, giving the viewer a greater sense of immersion.

[0112] In one aspect, in response to detecting an audio spectrum outside the set range, control unit 190 may change the appearance of facial parts and body parts other than the mouth. Specifically, for example, in response to detecting an audio spectrum outside the set range, control unit 190 may cause the avatar to perform the following actions. - Change the appearance of facial parts (eyebrows, corners of the eyes, inner corners of the eyes, corners of the mouth, etc.) · Change the state of body parts (arms, hands, shoulders, etc.) In addition, the control unit 190 may display a predetermined object on the screen viewed by the viewer in response to detecting an audio spectrum outside the set range. This allows the control unit 190 to better convey the user's emotions to the viewer, for example, when the user suddenly shouts or screams, by changing the appearance of facial or body parts, displaying objects, etc.

[0113] In one aspect, the control unit 190 of the terminal device 10 may estimate one or more candidate emotions of the user and present the estimated one or more candidate emotions to the user. Thereafter, the control unit 190 may receive an input operation from the user to select one of the one or more candidate emotions, and upon receiving the selection of the emotion from the user, may execute a process of changing the state of the avatar's mouth based on the selected emotion. Specifically, for example, the control unit 190 analyzes a voice spectrum acquired from the user and estimates candidate emotions when the user speaks. In this case, the candidate emotions include, for example, the following: Anger, rage Joy, fun Surprise, fear Grief, lamentation Peace and serenity Here, an example of a process for estimating emotion candidates from a voice spectrum will be described. For example, the control unit 190 may previously receive information on voice spectra corresponding to emotions from the user and store the information in the storage unit 180 or the like, thereby associating the user's voice spectrum with the user's emotion. After that, when the control unit 190 acquires a voice spectrum from the user, it estimates emotion candidates associated with voice spectra having waveforms similar to the acquired voice spectrum. "Similar waveforms" means, for example, that the similarity between the waveforms of multiple voice spectra is determined, and the waveforms match by a predetermined percentage, or the waveforms of multiple voice spectra deviate by a predetermined percentage (for example, match within a range of ±10%). In a certain aspect, a trained model may be used as a method for estimating candidate emotions of a user from a voice spectrum. For example, the terminal device 10 may store a trained model in the storage unit 180, which associates the voice spectra of multiple users with the emotions corresponding to the users. Thereafter, upon receiving input of a voice spectrum from the user, the control unit 190 of the terminal device 10 may estimate candidate emotions corresponding to the user's voice spectrum based on the trained model and present the candidate emotions to the user.

[0114] The control unit 190 may present the estimated emotion candidates to the user and accept a selection from the user. The control unit 190 also accepts in advance a setting of the degree of change in the mouth state for each emotion, and when it accepts an emotion selection from the user, it changes the state of the avatar's mouth based on the setting of the corresponding emotion. This allows the user to change the appearance of the avatar based on the emotion estimated from the utterance.

[0115] At this time, if the control unit 190 cannot estimate the user's emotion, it may change the state of the mouth based on conditions preset by the user. Specifically, for example, if the control unit 190 cannot estimate a candidate for the user's emotion from the voice spectrum acquired from the user, that is, if it cannot estimate a similar voice spectrum, it may change the state of the avatar's mouth based on conditions preset by the user. For example, when the control unit 190 cannot accurately acquire a voice spectrum from the user, when it cannot estimate a candidate emotion similar to the acquired voice spectrum, or when it has received a mouth correspondence setting for "calm" from the user, it changes the state of the avatar's mouth to a state based on the emotion of "calm." This allows the user to change the avatar into a preset state even when the emotion cannot be estimated, thereby reducing the sense of discomfort felt by the viewer.

[0116] Furthermore, in a certain aspect, the control unit 190 may operate a body part other than the mouth of the avatar based on the estimated emotion. Specifically, for example, the control unit 190 may operate a body part such as the shoulders, arms, or hands as a body part other than the mouth of the avatar. In addition, the control unit 190 may operate a special body part (for example, wings, a tail, or an object floating in the surrounding area if the avatar is not human) as a body part other than the mouth of the avatar. For example, the control unit 190 may raise the arms of the avatar when the emotion estimated from the voice spectrum acquired from the user is "anger" or the like. In addition, at this time, the control unit 190 may present candidate emotions to the user, rather than an emotion estimated from the acquired voice spectrum, and may move a body part of the avatar other than the mouth based on the emotion selected by the user.

[0117] Alternatively, the control unit 190 may present to the user one or more candidate body parts other than the mouth of the avatar whose behavior is to be changed based on the emotion estimated from the acquired voice spectrum, and change the behavior of the body part in response to receiving a selection from the user of the body part whose behavior is to be changed. This allows the user to move parts of the avatar other than the mouth based on emotions estimated from the user's own voice spectrum, giving the viewer a greater sense of immersion.

[0118] In one aspect, when the user's speech rate deviates by a predetermined rate from the speech rate estimated from the degree of change in the mouth attitude set by the user, the control unit 190 of the terminal device 10 may change the mouth attitude based on the speech rate rather than the degree set by the user. Specifically, for example, the control unit 190 calculates the speech rate from the user's speech. As a method for calculating the speech rate, for example, the control unit 190 may define the speech rate value by calculating the number of words the user speaks per unit time from a voice spectrum acquired from the user. Furthermore, the control unit 190 calculates the number of utterances per unit time from the degree of change in the mouth attitude set by the user, and calculates the user's speech rate estimated from the degree of change in the mouth attitude of the avatar. Thereafter, if there is a predetermined deviation between the speech rate calculated from the user's speech and the speech rate estimated from the degree of change in the avatar's mouth pattern, the control unit 190 may change the mouth pattern based on the speech rate calculated from the user's speech, rather than the setting by the user. This allows the user to change the avatar's mouth movements based on their own speech rate if the user's own speech rate deviates too much from the speech rate estimated from the degree of change in the avatar's mouth movements, thereby reducing the sense of discomfort felt by the viewer.

[0119] In some aspects, the control unit 190 of the terminal device 10 may receive avatar attributes from a user and correct the amount of change in the avatar's mouth shape based on the attributes. Specifically, for example, the control unit 190 may receive information on either a human or a non-human character whose mouth shape changes differently from humans as the avatar's attributes, and correct the amount of change in the avatar's mouth shape based on the attributes. For example, similar to the change correction module 2040 of the server 20, the control unit 190 may acquire information on whether the avatar operated by the user is a human or a non-human character whose mouth shape changes differently from humans, and perform processing to correct the amount of change in the avatar's shape based on the information. For example, if the attribute of the avatar operated by the user is a "dragon," the movements of the eyes, mouth, etc. may behave differently from humans. In this case, the control unit 190 may correct the amount of change in the corners of the mouth, the amount of change in the eyeballs, etc., to match the shape of the avatar based on the "dragon" attribute. This allows users to present more natural movements to viewers based on the results of their own speech and facial sensing, even when operating an avatar that is not human.

[0120] <4 Screen example> 6 to 9 are diagrams showing various examples of screens when a user, who is an actor operating an avatar, operates an avatar using the system 1 disclosed in the present invention.

[0121] FIG. 6 shows an example of a screen when a user registers the audio spectrum of his or her own vowels in the system 1.

[0122] In FIG. 6, the control unit 190 of the terminal device 10 displays a setting screen 601, an avatar 602, and the like on the display 1302. The setting screen 601 is a setting screen displayed to the user when acquiring and associating information on the voice spectrum corresponding to each vowel from the user. For example, the control unit 190 of the terminal device 10 displays a setting screen for six characters, "A," "I," "U," "E," "O," and "N," as vowels to be associated with the user's voice spectrum, on the screen. At this time, the control unit 190 may display information on the vowels currently associated with the user's voice spectrum at the top of the screen. Furthermore, the control unit 190 may display information about the microphone 141 used by the user at the bottom of the setting screen 601. When frequency characteristics differ depending on the type of microphone 141 used by the user, the control unit 190 may associate the user's voice spectrum with vowel information for each microphone 141 used.

[0123] The avatar 602 is an avatar whose mouth shape changes in response to the user's speech. The control unit 190 changes the mouth shape of the avatar 602 in response to a voice spectrum acquired from the user. For example, when the user utters the vowel "a (A)," the control unit 190 determines whether the utterance matches the voice spectrum stored for the vowel "a (A)." Thereafter, if the user utters "a (A)," the control unit 190 changes the mouth shape of the avatar 602 to the shape of "a (A)." This allows the user to accurately change the state of the avatar's mouth for each vowel.

[0124] FIG. 7 shows an example of a screen when the user sets the degree of change in the appearance of the avatar's mouth or facial features.

[0125] 7, the control unit 190 of the terminal device 10 displays an information display screen 701, a user video 702, a setting screen 703, an avatar 704, and the like on the display 1302.

[0126] The information display screen 701 is a screen that displays the frequency of the voice spectrum acquired from the user, the detectable range of the voice spectrum, the setting of the mode when the detection range is exceeded, etc. In addition, the terminal device 10 may display the user's speech rate calculated from the user's speech, the degree of mode change that the user can set, etc. The screen may also display the results of sensing the user's face, etc., to visually display various conditions that the user can set.

[0127] The user video 702 is a screen that displays an image of the user himself / herself captured by the camera 160 provided in the terminal device 10. When the user makes some kind of speech in front of the terminal device 10, the control unit 190 of the terminal device 10 displays an image of the user himself / herself and information such as the audio spectrum of the user's speech on the user video 702 and the information display screen 701 using the camera 160 and microphone 141 provided in the terminal device 10.

[0128] The setting screen 703 is a screen for the user to set the degree of change in the avatar's appearance. The control unit 190 of the terminal device 10 presents the following settings to the user, for example, and accepts input. Mouth switching speed The mouth switching speed is information about the time required for the voice spectrum acquired from the user to reach the maximum amplitude (100). Eye movement: Maximum upward movement Eye Movement: Maximum downward movement Eye movement: Maximum horizontal movement Eye Movement: Sensitivity (device sensing sensitivity) Sensitivity refers to the sensitivity with which the sensed user's eyes, etc. are reflected in the avatar. Specifically, for example, when the coordinates of the user's eyes, etc., are set to "0" when the user is facing straight ahead, sensitivity is a parameter that sets the extent to which the avatar's eyes, etc., are reflected in the actual movement of the eyes, etc., when the user moves their eyes, etc., left or right. In this case, sensitivity is a proportional function when it is 100, and may be a function that becomes more convex downward as it approaches 0. In other words, when sensitivity is 100, the user's eye movement and the avatar's eye movement are perfectly synchronized. When sensitivity is 50, for example, if the user's eyes, etc., do not move much from the center, the movement of the avatar's eyes, etc., is reflected less than the distance the user's eyes moved, and if the eyes move to the corners of the eyes, the movement of the avatar's eyes, etc., is reflected more than the distance the user's eyes moved. This prevents the avatar's eyes from moving wildly when the user does not move their eyes much. The sensitivity setting is not limited to the eyes, and similar settings may be accepted for other parts of the face and body other than the eyes. At this time, the control unit 190 of the terminal device 10 may accept a setting of the degree of change that is lower than the change in the user's utterance as the setting of the degree of change that can be accepted from the user. For example, the control unit 190 may accept the setting from the user so that the setting is lower than the degree of change of the avatar (amount of change in the object, speed of change in the object) estimated from the user's utterance. At this time, if the user attempts to set a value that is outside the settable range, the control unit 190 may display a predetermined alert, or, if the setting screen is a slider type or the like, may lock the value in advance so that it cannot be set. This allows the user to move the avatar more slowly than the user's own speech, thereby smoothing the degree of change in the avatar that is perceived by the viewer, thereby giving the viewer a greater sense of immersion.

[0129] The avatar 704 is an avatar that changes its appearance based on settings received from the user. When the control unit 190 of the terminal device 10 receives settings on the setting screen 703 from the user, the control unit 190 may synchronize the user video 702 and the avatar 704 and display them to the user. This allows the user to check in advance whether there is any discomfort when changing the appearance of the avatar according to their own settings.

[0130] FIG. 8 shows an example of a screen in which one or more candidate emotions of the user are estimated from the user's speech, and the appearance of the avatar is changed based on the estimated one or more emotions of the user.

[0131] 8, the control unit 190 of the terminal device 10 displays an information display screen 801, a user video 802, an avatar 803, and the like on the display 1302.

[0132] Similar to information display screen 701 in FIG. 7, information display screen 801 is a screen that displays the frequency of the voice spectrum acquired from the user, and in FIG. 8, it may present one or more candidate emotions estimated from the voice spectrum, as well as candidate emotion settings that the user wants to reflect in the appearance of the avatar. The control unit 190 may change the appearance of the avatar, for example, the appearance of the avatar's mouth or the appearance of parts of the face other than the mouth, by accepting a selection from the presented setting candidates from the user.

[0133] The user video 802 is a screen that displays an image of the user himself / herself captured by the camera 160 provided in the terminal device 10, similar to the user video 702 in FIG.

[0134] Avatar 803, like avatar 704 in FIG. 7, is an avatar that changes its appearance based on an emotion setting received from the user. Based on the emotion setting received from the user, the control unit 190 of the terminal device 10 changes the appearance of the avatar (e.g., the mouth) and displays it to the user. In this case, the control unit 190 may change or move the appearance of other parts of the avatar, not just the appearance of the avatar's mouth. For example, when the emotion selected and received from the user is "anger," the control unit 190 may change the appearance of the avatar's mouth based on the emotion "anger," and may also change the appearance of other parts of the avatar, such as the eyebrows and corners of the eyes. Alternatively, the control unit 190 may move a body part of the avatar (e.g., swinging its arms) based on the emotion. Based on the emotion, the control unit 190 may also display a predetermined object corresponding to the emotion on the screen displaying the avatar. This allows the user to make the avatar perform various changes and movements based on the emotions inferred from the speech, thereby providing a greater sense of immersion to the viewer.

[0135] In addition, in a certain aspect, control unit 190 may refer to user information 1801 or user information database 2021 to acquire information on emotions frequently used by the user and present it as candidates for emotions to be reflected in the avatar. This allows the user to easily change the state of the avatar even when the user is trying to change the state of the avatar for dramatic effect or the like, regardless of speech.

[0136] FIG. 9 shows an example of a screen on which a user can make various settings based on the voice spectrum, etc., for an avatar with attributes different from those of a human.

[0137] 9, the control unit 190 of the terminal device 10 displays an information display screen 901, a user video 902, a setting screen 903, an avatar 904, and the like on the display 1302.

[0138] 7 and 8, the information display screen 901 is a screen that displays the frequency of the audio spectrum acquired from the user, the range of the detectable audio spectrum, settings for when the audio spectrum is outside the detection range, etc. At this time, the control unit 190 may display information about the attributes of an avatar corresponding to the user on the information display screen 901. For example, the control unit 190 may refer to the user information 1801 or the user information database 2021 to acquire information about the avatar corresponding to the user, and thereby display information about the avatar's attributes on the screen.

[0139] The user video 902 is a screen that displays an image of the user himself / herself captured by the camera 160 provided in the terminal device 10, similar to the user videos 702 and 802 in FIGS.

[0140] Similar to the setting screen 703 in Fig. 7, the setting screen 903 is a screen for the user to set the degree of change in the avatar's appearance. In Fig. 9, the control unit 190 may display, in addition to the screen presented to the user on the setting screen 703, suggestions for settings recommended based on the attributes of the avatar. Specifically, for example, the control unit 190 may refer to the avatar information 1802 or the avatar information database 2022, etc., to acquire information on the amount of correction for the degree of change in appearance caused by the avatar, and present to the user settings obtained by multiplying the correction result by a basic setting for changing the appearance of a normal human avatar. This allows the user to set a natural change in appearance even when the user's avatar has attributes different from those of a human. In addition, in a certain aspect, when the avatar has a special body part, control unit 190 may display a screen for the user to set the degree of change in the appearance of that body part. For example, when synchronizing with the setting of another body part, control unit 190 may reflect the setting of change in the body part of the other avatar, or may present the user with a separate screen for setting the degree of change in appearance in detail. This allows the user to freely set the degree of change in appearance even when their avatar has a special body part, thereby providing a greater sense of immersion to the viewer.

[0141] Avatar 904 is an avatar that changes its appearance based on an emotion setting received from the user, similar to avatars 704 and 803 in Figures 7 and 8. In Figure 9, control unit 190 may simultaneously display special body parts of the avatar on avatar 904. This allows the user to broadcast to viewers while checking changes in the appearance of the avatar, even if the avatar has a special body part.

[0142] <Second embodiment> So far, a series of processes for changing the state of the avatar's mouth based on the voice spectrum of the user's speech has been described. In the invention according to the second embodiment, the appearance of the avatar, for example, the appearance of one or more facial features, can be changed based on the user's sensing results in addition to the voice spectrum of the user's speech. The series of processes will be described below. Note that a description of parts having the same configuration as the first embodiment (e.g., the terminal device 10, the server 20, etc.) will be omitted, and only the configuration and processing unique to the second embodiment will be described.

[0143] <5. Operation in the Second Embodiment> Below, we will explain a series of processes that system 1 performs when it senses the movement of one or more facial parts of a user's face and changes the appearance of one or more facial parts of an avatar corresponding to the user based on the sensed movement of the one or more facial parts.

[0144] 10 is a flowchart showing a series of processes for sensing the movement of one or more facial features of a user and changing the appearance of one or more facial features of an avatar corresponding to the user based on the sensed movement of the one or more facial features. This flowchart also discloses an example in which the control unit 190 of the terminal device 10 used by the user executes the series of processes, but is not limited to this. That is, the terminal device 10 may transmit some information to the server 20, and the server 20 may execute the process, or the server 20 may execute the entire series of processes.

[0145] In step S1001, the control unit 190 of the terminal device 10 senses the movement of one or more facial parts of the user. Specifically, for example, the control unit 190 of the terminal device 10 senses one or more facial parts of the user when the user moves his or her face in front of the camera 160 provided in the terminal device 10. At this time, the sensing method performed by the control unit 190 may be any existing technology. For example, the control unit 190 may sense the user's facial parts by providing the camera 160 with a sensing function, or may sense the user's facial parts using the motion sensor 170. At this time, the control unit 190 of the terminal device 10 senses at least one of the group consisting of the user's eyebrows, eyelids, inner corners of the eyes, outer corners of the eyes, eyeballs, pupils, and mouth as one or more facial parts of the user. However, the parts are not limited to these and may be other facial parts (cheeks, forehead, etc.).

[0146] In step S1002, the control unit 190 of the terminal device 10 changes the appearance of one or more facial features of the avatar corresponding to the user based on the sensed movement of one or more facial features. Specifically, for example, the control unit 190 associates one or more facial features of the user with one or more facial features of the avatar in advance. Thereafter, the control unit 190 changes the appearance of the facial features of the avatar corresponding to one or more facial features of the user acquired by sensing based on the sensing results. For example, if the control unit 190 has associated the user's eyes with the avatar's eyes, it changes the appearance of the avatar's eyes based on the sensing results of the user's eyes.

[0147] In step S1003, the control unit 190 of the terminal device 10 accepts a setting of the degree to which the appearance of one or more facial parts of the avatar is to follow the sensing result, and changes the appearance of one or more facial parts of the avatar in accordance with the setting of the degree. Specifically, for example, the control unit 190 accepts a setting of conditions including the following as the degree to which the appearance of one or more facial parts of the avatar is to follow the sensing result of the user. -Changes in the avatar's appearance (for example, changes in the opening and closing of the eyes, etc.) This allows the user to fine-tune the degree to which the changes in the avatar's appearance follow the user's own sensing results, preventing the viewer from feeling uncomfortable with the movements.

[0148] In the second embodiment, the control unit 190 may accept a user setting for the degree of change in the appearance of the avatar's facial features and non-facial body features, similar to the setting for the degree of change in the appearance of the avatar's mouth in the first embodiment. That is, the control unit 190 may acquire, in advance, from the user, the mouth appearances and various facial and body features corresponding to various vowels by sensing. The control unit 190 may identify a difference between the user's sensing results and previously acquired changes in the user's mouth, facial features, and body features, calculate a ratio to the previously acquired sensing results, and multiply the ratio by the amount of change in the appearance to calculate the amount of change in the appearance of the avatar's mouth, facial features, and body features. The control unit 190 may change the appearance of the avatar's mouth, facial features, and body features based on the calculated amount of change. For example, if the user only moves a portion of their mouth or eyebrows (100 positions are set in advance, and sensing results show that the user only moves their mouth, eyebrows, etc. up to position 50), processing can be performed such that the avatar's mouth, eyebrows, etc. also only move up to position 50. This allows the user to gradually change the avatar's appearance according to the user's own sensing results, allowing the viewer to see natural movements. This prevents the viewer from feeling uncomfortable due to the user's movements and the changes in the avatar's appearance, providing a greater sense of immersion to the viewer.

[0149] In one aspect, control unit 190 may accept the same setting for predetermined associated parts among one or more facial parts of an avatar. Specifically, control unit 190 may accept from the user a setting that associates, for example, the following parts among one or more facial parts of an avatar, and may accept the same setting regarding the degree setting for the parts: Paired parts of the face such as eyebrows and eyes Parts that move in tandem, like the eyebrows and eyes ·Facial areas and other body areas (shoulders, arms, legs, neck, etc.) Additionally, the control unit 190 may accept a setting that associates a facial part with a special non-facial part according to an attribute of the avatar, which will be described later. This allows the user to easily change and distribute the appearance of the avatar without having to set individual degrees for paired parts or parts that move in conjunction with each other among multiple facial parts.

[0150] In one aspect, the control unit 190 of the terminal device 10 may present the user with one or more candidate settings for the degree of tracking of the sensing results and may accept a selection of one or more candidate settings from the user. The control unit 190 may then change the appearance of one or more facial features of the avatar based on the selected and accepted setting of the degree. Specifically, for example, when acquiring the user's sensing results, the control unit 190 may present one or more candidate (preset) degrees of tracking rather than accepting a setting of the degree of tracking of the sensing results from the user. In this case, as a method of presenting the candidates, the control unit 190 may receive information on one or more candidate degrees of tracking to be used from the user in advance and present the candidates based on the information. This allows the user to change the appearance of the avatar's facial features based on the sensing results without having to set the degree of tracking each time, making distribution easier.

[0151] In addition, in a certain aspect, the control unit 190 of the terminal device 10 may receive avatar attributes from the user and correct the degree based on the attributes. Here, the control unit 190 may receive information on either a human or a non-human entity whose facial changes differ from those of a human in the manner in which one or more facial features change, and correct the degree based on the attributes. For example, similar to the change correction module 2040 of the server 20, the control unit 190 may acquire information on whether the avatar operated by the user is a human or a non-human entity whose facial changes differ from those of a human, and perform processing to correct the degree of change in the avatar's facial features based on the information. For example, if the attribute of the avatar operated by the user is a "dragon," the movements of the eyes, mouth, etc. may behave differently from those of a human. In this case, the control unit 190 may correct the amount of change in the corners of the mouth, the amount of change in the eyeballs, etc., to match the shape of the avatar based on the "dragon" attribute. This allows users to present more natural movements to viewers based on the results of their own speech and facial sensing, even when operating an avatar that is not human.

[0152] In another aspect, the control unit 190 of the terminal device 10 may acquire a voice spectrum of the user and acquire information on the degree of change in the user's speech from the acquired voice spectrum. Thereafter, the control unit 190 may accept a degree setting that is settable within a range associated with the degree of change in the user's speech, and change the appearance of one or more facial features of the avatar in accordance with the degree setting. Specifically, for example, the control unit 190 may acquire a voice spectrum from the user's speech via the microphone 141 or the like, and acquire the following information as the degree of change in the user's speech: - The amount of words spoken by the user per unit time (speech rate) - Changes in the volume of the user's voice - Changes in the pitch of the user's voice For example, the control unit 190 executes the following process to change the state of the avatar to a degree less than the degree of change of the avatar estimated from the change in the user's utterance. · Reflects mouth movements to the avatar at regular time intervals, regardless of vowel changes in the voice spectrum obtained from the user. The control unit 190 specifies a settable range of the degree of tracking the sensing result based on the acquired information on the degree of change in speech. For example, the control unit 190 may accept a setting of the degree from the user so that the amount of change, etc. described above does not exceed the degree of change in speech based on the acquired degree of change in speech. This allows the user to change the facial expression of the avatar based not only on the sensing results but also on the information of the audio spectrum, allowing the avatar to appear to viewers with more natural movements.

[0153] At this time, the control unit 190 may accept a setting of a frequency range in which the audio spectrum is detected, and in response to detecting the audio spectrum in the set range, change the appearance of one or more facial features of the avatar based on a first setting of the degree. Specifically, for example, when acquiring an audio spectrum from the user's speech, the control unit 190 may accept a setting of a detectable range from the user. If the audio spectrum acquired from the user is within the frequency range, the control unit 190 may change the appearance of the avatar's face based on the setting of the degree accepted from the user. Furthermore, in response to detecting an audio spectrum outside the set range, control unit 190 may change the appearance of one or more facial features of the avatar based on a second degree setting that is a predetermined degree setting and is different from the first degree setting. In this case, the second degree setting may be, for example, when the user utters a voice with an extremely high frequency (such as a shriek), to change the appearance of the avatar's face in accordance with a predetermined degree (second degree) corresponding to the frequency, rather than the degree setting (first degree) received from the user. This allows the user to change the facial expression of the avatar even when uttering a frequency that the user would not normally utter, thereby providing a greater sense of immersion to the viewer.

[0154] Furthermore, in a certain aspect, when the control unit 190 of the terminal device 10 cannot sense the movement of the user's mouth, the control unit 190 may change the state of the avatar's mouth based on the degree of change in the user's speech. Specifically, for example, in the following cases, the control unit 190 may change the state of the avatar's mouth based on the audio spectrum of the user's speech, rather than on the results of sensing the user's speech, as described above. - When the user is wearing a mask or other protective gear over their mouth and mouth movements cannot be sensed. When the mouth movement cannot be sensed due to an error in the sensing function of the terminal device 10 When mouth movements cannot be sensed due to external conditions This allows the user to change the state of the avatar's mouth to match the user's speech, even when, for example, the user has to broadcast while wearing a mask.

[0155] In one aspect, the control unit 190 of the terminal device 10 may estimate one or more candidate emotions of the user and present the estimated one or more candidate emotions to the user. The control unit 190 may then accept an input operation from the user to select one of the one or more candidate emotions, and change the appearance of one or more facial features of an avatar corresponding to the user based on the selected emotion. Specifically, for example, the control unit 190 may acquire and associate sensing results of facial features corresponding to the user's emotions in advance from the user. The control unit 190 then senses the user's face via the camera 160 or the like, and determines whether all or part of the sensing results of the face included in the associated emotion match. The control unit 190 may then present candidate emotions of the user based on the determination result, accept a selection from the user, and change the appearance of the avatar's face based on the selected emotion. Furthermore, at this time, if the user's emotion cannot be estimated, the control unit 190 may change the appearance of one or more parts of the face based on a setting preset by the user. For example, if the control unit 190 cannot accurately sense parts of the user's face, or if it cannot estimate a candidate emotion similar to the sensing result, and if it has received a mouth correspondence setting for "calm" from the user, it changes the state of the avatar's mouth to a state based on the emotion of "calm." This allows the user to select emotion candidates even when accurate sensing is not possible, allowing the user's emotions to be reflected in the changes in the avatar's appearance.

[0156] Furthermore, in a certain aspect, when the control unit 190 of the terminal device 10 is unable to obtain sensing results for at least one associated part of one or more facial parts of the user, the control unit 190 may apply the degree of change for the part for which sensing results were obtained to the associated part. Specifically, for example, when the user is wearing an eyepatch or the like and sensing one eye is difficult or impossible, the control unit 190 may reflect the degree of change in the other eye for which sensing results were obtained. This allows the avatar corresponding to the user to change its appearance without being affected by the eyepatch or the like, even when the user is wearing an eyepatch or the like.

[0157] Furthermore, in one aspect, the control unit 190 of the terminal apparatus 10 may acquire information about a wearable device worn by the user and correct the degree setting based on the acquired information about the wearable device. Furthermore, when correcting the degree setting, the control unit 190 may accept an input operation from the user to adjust the degree of correction. Specifically, for example, the control unit 190 may refer to the wearable device information 1803 or the wearable device information database 2023 to acquire information about the wearable device worn by the user. The control unit 190 may then execute processing similar to that of the change correction module 2040 in the server 20 described above to correct the degree setting.

[0158] In one aspect, the control unit 190 of the terminal device 10 may present a predetermined notification to the user when the difference in the degree setting between pre-associated parts of one or more facial parts of the avatar exceeds a predetermined threshold. Specifically, the control unit 190 associates paired parts, such as eyebrows, among one or more facial parts of the avatar, and sets the degree of change between the parts to be acceptable so that the difference does not exceed a predetermined difference. Thereafter, when accepting input from the user of the degree of change of the part, the control unit 190 may present a notification such as an alert to the user if the input exceeds the threshold. This allows the user to prevent changing the state of the associated body parts that are to be changed in a state where there is an extreme difference in the degree of change. Furthermore, the control unit 190 may also set the settings for parts that change in conjunction with each other (which may include special parts, etc.), such as cheeks and eyebrows, in addition to paired parts.

[0159] In this case, when presenting a predetermined notification to the user, the control unit 190 may present the part of the eye whose difference in degree exceeds a predetermined threshold in a different manner along with the numerical value. Specifically, for example, when the control unit 190 receives an input of the degree of change in the state of the eyes from the user, if the degree of change in both eyes is too large (for example, the amount of change in one eye is too large), the control unit 190 may present the eyes in a different manner (for example, in a different color) along with the notification to the user. In this case, the different manner presented by the control unit 190 may include, but is not limited to, changing the color, a pop-up notification, or the shape of the relevant part. Furthermore, when presenting a predetermined notification to the user, the control unit 190 may present to the user how at least one or more facial features change when the difference in degree is set within a predetermined range. For example, the control unit 190 may display, on a screen different from the screen displaying the above-mentioned notification, how the appearance of the avatar changes when the difference in degree is within an appropriate range (a range that does not cause discomfort to the viewer). This allows the user to check how the behavior will change if the degree of behavior change that the user has set exceeds a predetermined threshold value, along with how the behavior will change if the value is set to an appropriate value.

[0160] In addition, in a certain aspect, control unit 190 of terminal device 10 may set the degree of a part associated with one or more facial parts for which a degree setting has been accepted to a predetermined value. Control unit 190 may also accept a degree setting within a predetermined range for each of one or more parts of the avatar. Specifically, for example, control unit 190 may associate the following parts as parts associated with one or more facial parts of the avatar, and accept a degree setting from the user: - Special body parts for non-human avatars, such as horns, tails, and wings - Body parts other than the avatar's face (arms, shoulders, legs, etc.) This allows the user to change the appearance of the avatar in accordance with the results of his or her own sensing, even if the avatar is not human or is an inorganic object.

[0161] <6. Screen Examples in the Second Embodiment> 11 to 17 are diagrams showing various examples of screens when the state of the avatar is changed based on the sensing result of the user, as disclosed in the second embodiment.

[0162] FIG. 11 shows an example screen when sensing the movement of one or more facial parts of a user and changing the appearance of one or more facial parts of the corresponding avatar based on the sensed movement of the one or more facial parts.

[0163] In FIG. 11, the control unit 190 of the terminal device 10 displays, on the display 1302, an information display screen 1101, a user video 1102, a setting screen 1103, an avatar 1104, and the like.

[0164] The information display screen 1101 is a screen that displays the sensing results of the user's facial parts, associated parts of the facial parts, preset candidates for the degree of change in appearance, etc. At this time, the control unit 190 of the terminal device 10 may accept the following selections from the user. -Selection of the parts of the user's face to perform sensing -Selection of the parts to associate from the sensed parts -Selection of candidate degrees of change This allows the user to reduce the number of sensing locations in some cases, thereby reducing the load during distribution.

[0165] The user video 1102 is a screen that displays an image of the user himself / herself captured by the camera 160 provided in the terminal device 10. The control unit 190 of the terminal device 10 displays an image of the user himself / herself on the user video 1102 by the camera 160 provided in the terminal device 10.

[0166] The setting screen 1103 is a screen for the user to set the degree of change in the avatar's appearance. The control unit 190 of the terminal device 10 presents the following settings to the user, for example, and accepts input. Mouth switching speed Eye movement: Maximum upward movement Eye Movement: Maximum downward movement Eye movement: Maximum horizontal movement Eye Movement: Sensitivity At this time, the control unit 190 of the terminal device 10 may accept a setting for the degree to which the appearance of one or more facial parts of the avatar follows the sensed results, settable within a range associated with the degree of change in the performer's speech. For example, the control unit 190 may accept the setting from the user so that the degree of change in the avatar (amount of change in the object, speed of change in the object) is lower than the degree of change estimated from the user's speech. At this time, the control unit 190 may display a predetermined alert if the user attempts to set a value outside the settable range, or, if the setting screen is a slider type or the like, may lock the value in advance so that it cannot be set. This allows the user to move the avatar more slowly than the user's own speech, thereby smoothing the degree of change in the avatar that is perceived by the viewer, thereby giving the viewer a greater sense of immersion.

[0167] The avatar 1104 is an avatar that changes its appearance based on settings received from the user. When the control unit 190 of the terminal device 10 receives settings on the setting screen 1103 from the user, the control unit 190 may synchronize the user video 1102 and the avatar 1104 and display them to the user. This allows the user to check in advance whether there is any discomfort when changing the appearance of the avatar according to their own settings.

[0168] FIG. 12 shows an example of a screen when one or more candidate emotions of a user are estimated, and the appearance of one or more facial parts of a corresponding avatar is changed based on the emotion selected by the user.

[0169] In FIG. 12, the control unit 190 of the terminal device 10 displays an information display screen 1201, a user video 1202, a setting screen 1203, an avatar 1204, and the like on a display 1302.

[0170] 11, the information display screen 1201 is a screen that displays the sensing results of the user's facial parts, associated parts of the facial parts, preset candidates for the degree of change in appearance, etc. In addition, the control unit 190 may display information on one or more candidate emotions of the user identified from the sensing results on the screen. When the control unit 190 receives a selection of emotion candidates from the user, it reflects the degree of change in the appearance of the avatar corresponding to the emotion. For example, the control unit 190 may acquire from the user in advance sensing results of facial parts corresponding to the user's emotions and associate them with the user's emotions. Thereafter, the control unit 190 senses the user's face via the camera 160 or the like, and determines whether all or part of the sensing results of the face included in the associated emotion match. Thereafter, the control unit 190 may present candidate emotions for the user based on the determination result, accept a selection from the user, and change the facial appearance of the avatar based on the selected emotion. Furthermore, at this time, if the user's emotion cannot be estimated, the control unit 190 may change the appearance of one or more parts of the face based on a setting preset by the user. This allows the user to select emotion candidates even when accurate sensing is not possible, allowing the user's emotions to be reflected in the changes in the avatar's appearance.

[0171] The user video 1202 is a screen that displays an image of the user himself / herself captured by the camera 160 provided in the terminal device 10, similar to the user video 1102 in FIG.

[0172] Similar to the setting screen 1103 in FIG. 11, the setting screen 1203 is a screen for the user to set the degree of change in the appearance of the avatar.

[0173] Avatar 1204, like avatar 1104 in FIG. 11, is an avatar that changes its appearance based on settings received from the user.

[0174] FIG. 13 shows an example of a screen for setting the degree of change in the avatar's appearance when sensing results cannot be obtained for at least one associated part of one or more parts of the user's face.

[0175] In FIG. 13, the control unit 190 of the terminal device 10 displays, on the display 1302, an information display screen 1351, a user video 1352, a setting screen 1353, an avatar 1354, and the like.

[0176] 12, the information display screen 1351 is a screen that displays the sensing results of the user's facial parts, associated parts of the facial parts, preset candidates (presets) for the degree of change in appearance, etc. In addition, the control unit 190 may display, on the screen, information about accessories, attachments, etc. that the user is wearing and that cover part of the user's face. For example, if the user is wearing an eyepatch or the like and sensing one eye is difficult or impossible, the control unit 190 may reflect the degree of change in the other eye from which the sensing results were obtained. This allows the avatar corresponding to the user to change its appearance without being affected by the eyepatch or the like, even if the user is wearing an eyepatch or the like.

[0177] The user video 1352 is a screen that displays an image of the user himself / herself captured by the camera 160 provided in the terminal device 10, similar to the user video 1202 in FIG.

[0178] Similar to the setting screen 1203 in FIG. 12, the setting screen 1353 is a screen for the user to set the degree of change in the appearance of the avatar.

[0179] Avatar 1354 is an avatar that changes its appearance based on settings received from the user, similar to avatar 1204 in FIG.

[0180] FIG. 14 shows an example of a screen when correcting the degree of change in the appearance of an avatar when a user is wearing a wearable device such as glasses.

[0181] In FIG. 14, the control unit 190 of the terminal device 10 displays, on the display 1302, an information display screen 1401, a user video 1402, a setting screen 1403, an avatar 1404, and the like.

[0182] 13, information display screen 1401 is a screen that displays the sensing results of parts of the user's face, associated parts of the face, presets for the degree of change in appearance, etc. In addition, control unit 190 may display, on the screen, information about wearable devices worn by the user, information about the correction amount for the degree of change for each wearable device, etc. For example, the control unit 190 acquires information about the wearable device worn by the user by referring to the wearable device information 1803 or the wearable device information database 2023. Thereafter, the control unit 190 may execute the same process as the change correction module 2040 in the server 20 described above to correct the degree setting.

[0183] The user video 1402 is a screen that displays an image of the user himself / herself captured by the camera 160 provided in the terminal device 10, similar to the user video 1352 in FIG.

[0184] Similar to the setting screen 1353 in FIG. 13, the setting screen 1403 is a screen for the user to set the degree of change in the appearance of the avatar.

[0185] Avatar 1404 is an avatar that changes its appearance based on settings received from the user, similar to avatar 1354 in FIG.

[0186] FIG. 15 shows an example of a screen in which the state of the avatar's mouth is changed based on the degree of change in speech when the user's mouth movement cannot be sensed.

[0187] 15, the control unit 190 of the terminal device 10 displays, on the display 1302, an information display screen 1501, a user video 1502, a setting screen 1503, an avatar 1504, and the like.

[0188] 14, the information display screen 1501 is a screen that displays the sensing results of the user's facial parts, associated parts of the facial parts, preset candidates for the degree of change in appearance, etc. In addition, the control unit 190 may display information such as a mask worn by the user, information on the voice spectrum acquired from the user's speech, etc. For example, if the user is wearing a mask or the like over their mouth and the control unit 190 cannot sense the movement of their mouth, as described above, the control unit 190 may change the shape of the avatar's mouth based on the audio spectrum of the user's speech rather than the results of sensing the user's movements. This allows the user to change the state of the avatar's mouth to match the user's speech, even when, for example, the user has to broadcast while wearing a mask.

[0189] The user video 1502 is a screen that displays an image of the user himself / herself captured by the camera 160 provided in the terminal device 10, similar to the user video 1402 in FIG.

[0190] Similar to the setting screen 1403 in FIG. 14, the setting screen 1503 is a screen for the user to set the degree of change in the appearance of the avatar.

[0191] Avatar 1504 is an avatar that changes its appearance based on settings received from the user, similar to avatar 1404 in FIG.

[0192] FIG. 16 shows an example of a screen that displays a predetermined notification to the user when the difference in degree settings between one or more facial parts of an avatar that are pre-associated exceeds a predetermined threshold.

[0193] 16, the control unit 190 of the terminal device 10 displays, on the display 1302, an information display screen 1601, a user video 1602, a setting screen 1603, an avatar 1604, and the like.

[0194] The information display screen 1601, like the information display screen 1501 in FIG. 15, is a screen that displays the sensing results of the user's facial parts, associated facial parts, and preset candidates (presets) for the degree of change in appearance.

[0195] The user video 1602 is a screen that displays an image of the user himself / herself captured by the camera 160 provided in the terminal device 10, similar to the user video 1502 in FIG.

[0196] Settings screen 1603 is a screen for the user to set the degree of change in the appearance of the avatar, similar to settings screen 1503 in Fig. 15. At this time, when control unit 190 receives an input from the user regarding the degree of change in the appearance of a facial part, and if the degree of change in a paired or related part (both eyes, etc.) is too large (for example, the amount of change in one eye is too large), control unit 190 may display on this screen that the setting value for that part is abnormal and may display recommended settings.

[0197] Avatar 1604 is an avatar that changes its appearance based on settings received from the user, similar to avatar 1504 in Fig. 15. When control unit 190 receives an input from the user on setting screen 1503 indicating the degree of change in the facial appearance, if the degree of change in a paired or related part is too great, control unit 190 may present the part in a different appearance (for example, in a different color) along with a notification to the user. In this case, the different appearance presented by control unit 190 may be, but is not limited to, a change in color, a pop-up notification, a change in the shape of the relevant part, or the like.

[0198] This allows the user to visually determine if an abnormal value has been entered when making settings to change the appearance of parts of the avatar's face, preventing viewers from feeling uncomfortable.

[0199] FIG. 17 shows an example of a screen that shows the user how at least one or more facial parts change when the difference in degree is set within a predetermined range when a predetermined notification is presented to the user.

[0200] 17, the control unit 190 of the terminal device 10 displays, on the display 1302, a setting screen 1701, an avatar 1702, a setting preview screen 1703, an avatar preview screen 1704, and the like.

[0201] Settings screen 1603 is a screen for the user to set the degree of change in the appearance of the avatar, similar to settings screen 1503 in Fig. 15. At this time, when control unit 190 receives an input from the user regarding the degree of change in the appearance of a facial part, and if the degree of change in a paired or related part (both eyes, etc.) is too large (for example, the amount of change in one eye is too large), control unit 190 may display on this screen that the setting value for that part is abnormal and may display recommended settings.

[0202] The setting preview screen 1703 is a screen that displays recommended settings when there is an abnormal value in the degree of change between paired, related parts, such as the avatar's face, on the setting screen 1701. The control unit 190 of the terminal device 10 displays, on the setting preview screen 1703, the degree of change in the aspect of a setting that is different from the setting entered on the setting screen 1701. At this time, the control unit 190 may display numerical values, objects, etc. in a manner different from the manner displayed on the setting screen 1701 (for example, a different color, size, shape, etc.).

[0203] The avatar preview screen 1704 is a screen that displays an avatar that reflects the settings recommended on the setting preview screen 1703. For example, the control unit 190 of the terminal device 10 may display, on a screen different from the screen that displays the above-mentioned notification, how the appearance of the avatar changes when the difference in degree is within an appropriate range (a range that does not cause discomfort to the viewer). This allows the user to check how the behavior will change if the degree of behavior change that the user has set exceeds a predetermined threshold value, along with how the behavior will change if the value is set to an appropriate value.

[0204] <7 Variations> A modification of this embodiment will be described below. That is, the following aspects may be adopted. (1) An information processing device in which the program may be pre-installed or may be installed later, or such a program may be stored on an external non-transitory storage medium or may be run using cloud computing. (2) A method in which a computer functions as an information processing device, and the program may be pre-installed on the information processing device or may be installed later, or such a program may be stored on an external non-transitory storage medium, or may be operated using cloud computing.

[0205] <6 Notes> The matters described in the above embodiments will be supplemented below.

[0206] (Appendix 1) A program executed by a computer 20 having a processor 29, the program causing the processor 29 to execute the steps of: acquiring an audio spectrum of a performer's speech (S501); changing the mouth shape of an avatar corresponding to the performer in accordance with the performer's speech based on the acquired audio spectrum (S502); presenting the avatar corresponding to the performer and the performer's voice to an audience (S503); and accepting a setting for the degree to which the mouth shape of the avatar is changed in accordance with the performer's speech, to be able to set the degree to be lower than the change in the performer's speech (S504), wherein in the changing step (S502), the mouth shape of the avatar is changed in accordance with the setting.

[0207] (Appendix 2) 2. The program according to claim 1, wherein in the receiving step (S504), the degree setting is received from the performer so that it can be set lower than the performer's speaking rate.

[0208] (Appendix 3) 3. The program according to claim 1, wherein in the receiving step (S504), a setting of a frequency range for detecting the audio spectrum is received, and in the changing step, in response to detecting the audio spectrum in the set range, the state of the avatar's mouth is changed based on a first setting of the degree.

[0209] (Appendix 4) The program described in Appendix 3, wherein in the changing step (S502), in response to detecting an audio spectrum outside the set range, the program changes the shape of the avatar's mouth based on a second setting that is a predetermined setting and different from the first setting.

[0210] (Appendix 5) 5. The program according to any one of appendices 1 to 4, wherein in the changing step (S502), the mouth posture is changed based on at least one of the group consisting of strength and weakness and pitch of the voice spectrum.

[0211] (Appendix 6) The program further causes processor 29 to execute the steps of estimating one or more candidate emotions of the performer, presenting the estimated one or more candidate emotions of the performer to the performer, and accepting an input operation from the performer to select one of the one or more candidate emotions of the performer, and in the changing step (S502), if a selection of an emotion is accepted from the performer, the program changes the state of the avatar's mouth based on the selected emotion.

[0212] (Appendix 7) 7. The program according to claim 6, wherein, if the emotion cannot be estimated in the estimation step, the mouth shape is changed in a changing step (S502) based on a condition preset by the performer.

[0213] (Appendix 8) 7. The program of claim 6, further causing the processor 29 to perform a step of activating a body part of the avatar other than the mouth based on the estimated emotion.

[0214] (Appendix 9) The program of claim 6, further causing the processor 29 to perform a step of moving a body part of the avatar other than the mouth based on an emotion selected by the performer.

[0215] (Appendix 10) The program further causes the processor to change the mouth posture based on the speaking rate rather than the setting set by the performer when the speaking rate of the performer deviates by a predetermined rate from the speaking rate estimated from the degree of change in mouth posture set by the performer.

[0216] (Appendix 11) 11. The program according to any one of appendices 1 to 10, wherein in the receiving step (S504), attributes of the avatar are received from the performer, and the amount of change in the mouth shape is corrected based on the attributes.

[0217] (Appendix 12) The program described in Appendix 11, wherein in the receiving step (S504), information on either a human or a non-human entity whose mouth shape changes differently from a human's is received as an attribute, and the amount of change in the mouth shape is corrected based on the attribute.

[0218] (Appendix 13) A program described in any of Appendices 1 to 12, which program further causes processor 29 to perform the steps of presenting to the performer one or more candidate body parts other than the avatar's mouth whose appearance is to be changed based on emotions estimated from the acquired voice spectrum, and changing the appearance of the part in response to receiving from the performer a selection of the part whose appearance is to be changed.

[0219] (Appendix 14) A method executed by a computer 20 having a processor 29, the method comprising the steps of: acquiring an audio spectrum of a performer's speech (S501); changing the mouth shape of an avatar corresponding to the performer in accordance with the performer's speech based on the acquired audio spectrum (S502); presenting the avatar corresponding to the performer and the performer's voice to an audience (S503); and accepting a setting for the degree to which the avatar's mouth shape is changed in accordance with the performer's speech, to a degree that is lower than the change in the performer's speech (S504), wherein in the changing step (S502), the mouth shape of the avatar is changed in accordance with the setting.

[0220] (Appendix 15) An information processing device 20 having a control unit 203, wherein the control unit 203 executes the steps of: acquiring an audio spectrum of a performer's speech (S501); changing the mouth shape of an avatar corresponding to the performer in accordance with the performer's speech based on the acquired audio spectrum (S502); presenting the avatar corresponding to the performer and the performer's voice to an audience (S503); and accepting a setting for the degree to which the avatar's mouth shape is changed in accordance with the performer's speech, to be able to set the degree to be lower than the change in the performer's speech (S504); and in the changing step (S502), the information processing device 20 changes the mouth shape of the avatar in accordance with the setting. [Explanation of symbols]

[0221] 10 terminal device, 12 communication interface, 13 input device, 14 output device, 15 memory, 16 storage unit, 19 processor, 20 server, 22 communication interface, 23 input / output interface, 25 memory, 26 storage, 29 processor, 80 network, 1801 user information, 1802 avatar information, 1803 wearable device information, 1901 input operation acceptance unit, 1902 transmission / reception unit, 1903 data processing unit, 1904 notification control unit, 1302 display, 140 audio processing unit, 141 microphone, 142 speaker, 150 position information sensor, 160 camera, 170 motion sensor, 2021 user information database, 2022 avatar information database, 2023 wearable device information database, 2031 reception control module, 2032 transmission control module, 2033 user information acquisition module, 2034 Avatar information acquisition module, 2035 voice spectrum acquisition module, 2036 avatar change module, 2037 avatar presentation module, 2038 setting reception module, 2039 wearable device information acquisition module, 2040 change correction module.

Claims

1. A program executed by a computer having a processor, the program causing the processor to: obtaining an audio spectrum of a speaker's speech; a step of changing a mouth shape of an avatar corresponding to the performer in accordance with the speech of the performer based on the acquired voice spectrum; presenting an avatar corresponding to the performer and the performer's voice to a viewer; and accepting a setting of the degree to which the mouth shape of the avatar is changed in response to the speech of the performer so as to be lower than the degree of change in the speech of the performer; In the changing step, the state of the mouth of the avatar is changed in accordance with the setting.

2. 2. The program according to claim 1, wherein the step of accepting accepts the setting of the degree from the performer such that the setting can be set lower than the speaking rate of the performer.

3. In the receiving step, a setting of a frequency range in which the audio spectrum is detected is received, 2. The program according to claim 1, wherein in the changing step, the state of the mouth of the avatar is changed based on a first setting of the degree in response to detecting the audio spectrum in the set range.

4. The program of claim 3, wherein in the changing step, in response to detecting a voice spectrum outside the set range, the state of the avatar's mouth is changed based on a second setting that is a predetermined setting of the degree and is different from the first setting.

5. 2. The program according to claim 1, wherein the step of changing the mouth shape changes the mouth shape based on at least one of the group consisting of loudness and low pitch of the voice spectrum.

6. The program further causes the processor to execute the steps of estimating one or more candidate emotions of the performer, presenting the estimated one or more candidate emotions of the performer to the performer, and receiving an input operation from the performer to select one emotion from the one or more candidate emotions of the performer, 2. The program according to claim 1, wherein, in the changing step, when a selection of the emotion is received from the performer, the state of the mouth of the avatar is changed based on the selected emotion.

7. 7. The program according to claim 6, wherein, if the emotion cannot be estimated in the estimating step, the state of the mouth is changed in the changing step based on a condition preset by the performer.

8. The program according to claim 6 , further causing the processor to execute a step of moving a body part of the avatar other than the mouth based on the estimated emotion.

9. The program of claim 6 , further causing the processor to execute a step of moving a body part of the avatar other than the mouth based on the emotion selected from the performer.

10. 2. The program of claim 1, further comprising causing the processor to change the mouth posture based on the speech rate rather than the degree of change set by the performer when the performer's speech rate deviates by a predetermined rate from the speech rate estimated from the degree of change in the mouth posture set by the performer.

11. In the receiving step, The program according to claim 1 , further comprising: receiving attributes of the avatar from the performer; and correcting the amount of change in the mouth shape based on the attributes.

12. 12. The program according to claim 11, wherein the receiving step receives, as the attribute, information on either a human or a non-human object whose mouth shape changes differently from a human object, and corrects the amount of change in the mouth shape based on the attribute.

13. The program further instructs the processor to present to the performer one or more candidate body parts, other than the mouth, of the avatar that change appearance based on the emotion estimated from the acquired voice spectrum; 2. The program according to claim 1, further comprising: a step of changing the appearance of the part in response to receiving a selection from the performer of the part whose appearance is to be changed.

14. 1. A computer-implemented method comprising a processor, the method comprising: obtaining an audio spectrum of a speaker's speech; a step of changing a mouth shape of an avatar corresponding to the performer in accordance with the speech of the performer based on the acquired voice spectrum; presenting an avatar corresponding to the performer and the performer's voice to a viewer; and accepting a setting of a degree to which the mouth shape of the avatar is changed in response to the speech of the performer, the degree being lower than the degree of change in the speech of the performer; A method, wherein in the changing step, a mouth aspect of the avatar is changed in accordance with the setting.

15. An information processing device including a control unit, the control unit obtaining an audio spectrum of a speaker's speech; a step of changing a mouth shape of an avatar corresponding to the performer in accordance with the speech of the performer based on the acquired voice spectrum; presenting an avatar corresponding to the performer and the performer's voice to a viewer; and accepting a setting of a degree to which the mouth shape of the avatar is changed in response to the speech of the performer, the degree being lower than the degree of change in the speech of the performer; An information processing device wherein, in the changing step, a state of the mouth of the avatar is changed in accordance with the setting.

Citation Information

Patent Citations

  • Mouth shape expression animation generation method and device based on formant, and storage medium

    CN112700520A

  • Animation method and device for performing lip sinchronization

    JP2002108382A

  • Program, information storage medium, mouth shape control method, and mouth shape control device

    JP2010238133A

  • Voice synchronization processor, voice synchronization processing program, voice synchronization processing method, and voice synchronization system

    JP2015148932A

  • Cat type conversation robot

    JP2019061111A