program

The program uses acoustic features to generate avatars that reflect user individuality and tone, addressing privacy concerns and enhancing immersion by incorporating machine learning and biometric data for dynamic changes.

JP7894503B1Active Publication Date: 2026-07-23COLOPL
View PDF 7 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
COLOPL
Filing Date
2025-09-25
Publication Date
2026-07-23

Smart Images

  • Figure 0007894503000001_ABST
    Figure 0007894503000001_ABST
Patent Text Reader

Abstract

We provide a program that generates avatars that reflect the user's individuality while respecting user privacy. [Solution] The program enables the server to function as an extraction means for extracting acoustic features from the user's voice, a determination means for determining visual features corresponding to the extracted acoustic features, and a generation means for generating an avatar having the determined visual features.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a program.

Background Art

[0002] In recent years, in virtual spaces such as social networking services (SNS), online games, and the metaverse, the opportunity to use an avatar as a user's alter ego has been increasing. In such services, a function for creating an avatar that a user uses in the virtual space is often provided.

[0003] As a method for creating an avatar, a method of arbitrarily combining parts such as a face contour, eyes, nose, mouth, hairstyle, etc. prepared in advance by a user is common. Also, a method of taking a user's own face photo and generating an avatar based on the face photo is known (see Patent Document 1).

Prior Art Documents

Patent Documents

[0004]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0005] However, in the method using a user's face photo, although an avatar similar to the user can be generated relatively easily, since the appearance of the user himself / herself is directly reflected, concerns regarding privacy may arise. Also, an avatar with the appearance desired by the user is not always generated, and conversely, it may hinder the game experience. Furthermore, in this method, physical characteristics other than visual information cannot be reflected in the appearance generation of the avatar, and there is room for improvement in the user's attachment to the avatar.

[0006] The present invention aims to generate avatars that reflect the user's individuality while respecting the user's privacy. [Means for solving the problem]

[0007] To achieve the above objective, the program according to the present invention provides a computer with an extraction means for extracting acoustic features from the user's voice, a determination means for determining visual features corresponding to the extracted acoustic features, and a generation means for generating an avatar having the determined visual features. The decision means inputs the extracted acoustic features into a correlation model that has learned the relationship between acoustic and visual features, and determines the visual feature with the highest correlation. The generation means generates an avatar by adding the content of the user's voice to the determined visual feature. It is structured in this way.

[0008] Furthermore, the program according to the present invention may be configured such that the decision means determines the visual features using a machine learning model that has learned the correlation between acoustic features and visual features. [Effects of the Invention]

[0009] According to the present invention, it is possible to generate avatars that reflect the user's individuality while taking user privacy into consideration. [Brief explanation of the drawing]

[0010] [Figure 1] This figure shows the overall configuration of the information processing system according to the first embodiment of the present invention. [Figure 2] This is an external view of a user terminal according to the first embodiment of the present invention. [Figure 3] This is a block diagram showing the hardware configuration of a user terminal according to the first embodiment of the present invention. [Figure 4] A block diagram showing the hardware configuration of a server according to the first embodiment of the present invention. [Figure 5] This is a functional block diagram of a user terminal according to the first embodiment of the present invention. [Figure 6] This is a functional block diagram of a server according to the first embodiment of the present invention. [Figure 7]It is a diagram showing a conceptual example of a correlation model between acoustic features and facial features according to the first embodiment of the present invention. [Figure 8] It is a diagram showing a configuration example of various data tables stored in a server storage device according to the first embodiment of the present invention. [Figure 9] It is a diagram showing a display example of an avatar generation screen in the game according to the first embodiment of the present invention. [Figure 10] It is a diagram showing a display example of a voice input screen in the game according to the first embodiment of the present invention. [Figure 11] It is a sequence diagram showing the flow of processing in an information processing system according to the first embodiment of the present invention. [Figure 12] It is a flowchart showing the avatar generation process of a user terminal in the first embodiment of the present invention. [Figure 13] It is a flowchart showing the avatar generation control process of a server in the first embodiment of the present invention. [Figure 14] It is a functional block diagram of a user terminal according to the second embodiment of the present invention. [Figure 15] It is a functional block diagram of a server according to the second embodiment of the present invention. [Figure 16] It is a diagram showing an example of the correspondence between a mental state and avatar changes according to the second embodiment of the present invention. [Figure 17] It is a sequence diagram showing the flow of processing in an information processing system according to the second embodiment of the present invention. [Figure 18] It is a flowchart showing the avatar state change process of a user terminal in the second embodiment of the present invention. [Figure 19] It is a flowchart showing the avatar state change control process of a server in the second embodiment of the present invention. [Figure 20] It is a diagram showing an example of the state change of an avatar in the second embodiment of the present invention. [Figure 21] It is a diagram showing an example of the state change of an avatar in the second embodiment of the present invention. [Figure 22]This is a diagram showing an example of the state change of an avatar in the second embodiment of the present invention.

Mode for Carrying Out the Invention

[0011] (First Embodiment) Hereinafter, a first embodiment of the present invention will be described with reference to the drawings. Note that the following embodiments are merely one aspect for explaining the present invention, and the numerical values shown in this embodiment are also merely examples and do not limit the present invention. Also, not all of the configurations described in this embodiment are essential constituent elements of the present invention. Furthermore, in this specification and the drawings, elements having substantially the same functions and configurations are given the same reference numerals as much as possible, and redundant descriptions are omitted, and elements not directly related to the present invention are not shown.

[0012] (Configuration of Information Processing System 1) FIG. 1 is a diagram showing the overall configuration of an information processing system 1 according to the first embodiment. As shown in FIG. 1, the information processing system 1 includes a server 100 and a plurality of user terminals 200, which are communicably connected via a network 300. The network 300 is configured by, for example, the Internet, a mobile phone network, a LAN (Local Area Network), a WAN (Wide Area Network), a dedicated line, or the like, or a combination thereof.

[0013] In the first embodiment, a game is provided by the information processing system 1. Also, the information processing system 1 functions as an information processing device in which the server 100 and the user terminals 200 cooperate to provide a game. In the server 100 and the user terminals 200, the execution of various functions in the game and the control related to the progress of the game are respectively shared, and by the cooperation of the server 100 and the user terminals 200, various functions and games are executed.

[0014] Server 100 is a web server, game server, etc., that executes a program related to the game of the first embodiment (hereinafter also referred to as the game program), and performs processes such as transmitting various game-related data (player data related to the user playing the game (hereinafter also referred to as the player), avatar data related to the avatar that is the player's alter ego used in the game) in response to requests from the user terminal 200, storing and updating such various data, and managing the progress of the game.

[0015] The user terminal 200 is a terminal device operated by the player, and examples include smartphones, tablet devices, mobile phone terminals, personal computers, home game consoles, portable game consoles, and stationary game consoles installed in stores. The user terminal 200 interprets and executes game programs and various data transmitted from the server 100 to display the game screen on the display 204 and to accept input from the player. Each user terminal 200 may also be able to send and receive various data directly to other user terminals 200 without going through the network 300 or server 100, using its own wireless communication function. Furthermore, the user terminal 200 may be linked to or included in the biometric information acquisition device 230 described later.

[0016] (Hardware configuration) Next, we will describe the hardware configuration of server 100 and user terminal 200.

[0017] Figure 2 is an external view of the user terminal 200 of the first embodiment. Figure 2 illustrates the case where the user terminal 200 is a smartphone. The user terminal 200 has a housing 202 and a display 204 capable of displaying text, images, etc. The display 204 also functions as an input device that functions as a touch panel, and the player can perform various operations by touching, tapping, swiping, dragging, flicking, pinching in, pinching out, etc. on the display 204 (hereinafter also referred to as the touch panel 204).

[0018] Figure 3 is a block diagram showing the hardware configuration of the user terminal 200. As shown in Figure 3, the user terminal 200 includes a CPU (Central Processing Unit) 210, RAM (Random Access Memory) 212, storage device 214, communication interface 216, input / output interface 218, GPU (Graphics Processing Unit) 220, etc. Hereinafter, the CPU 210, RAM 212, storage device 214, communication interface 216, input / output interface 218, and GPU 220 of the user terminal 200 will also be referred to as terminal CPU 210, terminal RAM 212, terminal storage device 214, terminal communication interface 216, terminal input / output interface 218, and terminal GPU 220, respectively.

[0019] The terminal CPU 210 is a processor that comprehensively controls all parts of the user terminal 200, and performs various processes such as executing various functions in the game and progressing the game, according to the various programs and data stored in the terminal storage device 214.

[0020] Terminal RAM 212 is a volatile storage medium that functions as a work area for terminal CPU 210. It stores various programs and data used for various processes such as executing various functions in the game and progressing through the game.

[0021] The terminal storage device 214 is a non-volatile storage medium such as an HDD (Hard Disk Drive), SSD (Solid State Drive), or flash memory, and stores various programs and data used for various processes such as the OS (Operating System), execution of various functions in games, and game progression. The various programs and data stored in the terminal storage device 214 are read (loaded) into the terminal RAM 212 by the terminal CPU 210.

[0022] The terminal communication interface 216 is an interface for communicating with the server 100 and other user terminals 200 via the network 300. In the user terminal 200, various data is received from the server 100 via the terminal communication interface 216, and the received data is stored in the terminal RAM 212 and terminal storage device 214. Furthermore, information regarding the results of various processes performed by the terminal CPU 210 is transmitted to the server 100 and other user terminals 200 via the terminal communication interface 216.

[0023] The terminal input / output interface 218 is an interface for connecting input / output devices such as the display 204 (touch panel 204), a speaker 222 capable of outputting sound, a microphone 224 capable of inputting sound, and a biometric information acquisition device 230, which will be described later. In this embodiment, various processes such as storing and updating various programs and data in the terminal RAM 212 and terminal storage device 214 are executed based on operation input to the display 204. The microphone 224 is used to acquire the player's voice. Note that the input device capable of receiving operation input from the player is not limited to the touch panel 204, but may also be a keyboard, mouse, controller, physical buttons, etc. Furthermore, an acceleration sensor capable of detecting the tilt, movement, vibration, etc. of the user terminal 200 and inputting the detected information may also be installed as an input device.

[0024] Here, the biometric information acquisition device 230 is an input device for acquiring the player's biometric information, and examples include wearable devices and peripherals such as smartwatches, VR (Virtual Reality) headsets, and dedicated game controllers, which are equipped with one or more sensors capable of acquiring the player's biometric information. Examples of sensors capable of acquiring biometric information include heart rate sensors for measuring heart rate, skin electrical activity sensors for measuring skin electrical activity (sweating level), eye-tracking sensors for tracking eye movements and pupil size, electroencephalogram (EEG) sensors for measuring brain activity, and electromyogram (EMG) sensors for measuring muscle activity. The biometric information acquisition device 230 is connected to the user terminal 200 via short-range wireless communication such as Bluetooth® or wired connection such as USB (Universal Serial Bus). The user terminal 200 then acquires biometric information such as the player's heart rate, skin electrical activity (sweating level), eye movements, pupil size, EEG, and muscle activity in real time via the terminal input / output I / F 218 using these sensors. The user terminal 200 itself may also incorporate these sensors and function as a biometric information acquisition device.

[0025] The terminal GPU 220 is a processor specializing in image processing. It performs rendering processing according to instructions from the terminal CPU 210 and outputs the results to the display 204. Note that the terminal GPU 220 is not necessarily required; the terminal CPU 210 may perform the rendering processing instead.

[0026] The terminal CPU 210, terminal RAM 212, terminal storage device 214, terminal communication interface 216, terminal input / output interface 218, terminal GPU 220, etc., are connected to each other via internal buses, connecting cables, etc., which are not shown in the diagram.

[0027] Figure 4 is a block diagram showing the hardware configuration of server 100. As shown in Figure 4, server 100 includes a CPU 110, RAM 112, storage device 114, communication interface 116, input / output interface 118, etc. Hereafter, the CPU 110, RAM 112, storage device 114, communication interface 116, and input / output interface 118 of server 100 will also be referred to as server CPU 110, server RAM 112, server storage device 114, server communication interface 116, and server input / output interface 118, respectively. Since these hardware configurations are substantially the same as those of user terminal 200 in terms of configuration and function, detailed explanations will be omitted except for some parts.

[0028] The server storage device 114 stores the game program of the first embodiment, as well as data such as player data and avatar data for each player. The server 100 also performs various processes such as executing various functions in the game and progressing the game, according to the game program and data stored in the server storage device 114. The server 100 also receives various data from the user terminal 200 via the server communication interface 116, and the received data is stored in the server RAM 112 and the server storage device 114. In addition, data related to the results of various processes executed by the server CPU 110 is transmitted to the user terminal 200 via the server communication interface 116.

[0029] (Functional block) Next, the functions of the information processing system 1 of the first embodiment will be described. In the first embodiment, the system is configured to generate an avatar based on the player's voice.

[0030] Figure 5 is a functional block diagram of the user terminal 200 according to the first embodiment. Functionally, the user terminal 200 includes a data transmission / reception unit 280, a storage processing unit 282, an operation control unit 284, a game control unit 286, a screen control unit 288, and an audio acquisition unit 290. These functions are realized by the terminal CPU 210 executing various programs stored in the terminal storage device 214 and causing the various hardware components of the user terminal 200 to work together.

[0031] The data transmission / reception unit 280 transmits and receives various types of data to and from the server 100.

[0032] The memory processing unit 282 stores data received from the server 100 and data generated by the player's operations in the terminal storage device 214.

[0033] The operation control unit 284 receives and controls the player's operation input made to the display 204 (touch panel 204). The operation control unit 284 detects the type of operation input (e.g., touch, tap, swipe, drag, flick, pinch in, pinch out, etc.) based on the contact position (operation position) and the direction of movement from the contact position (operation direction) to the display 204.

[0034] The game control unit 286 controls the overall progress of the game. Based on game data transmitted from the server 100, the game control unit 286 constructs a virtual space (game space), places various objects, and advances the game by treating the operation input detected by the operation control unit 284 as instructions from the player. For example, it interprets touches and taps on various icons and objects displayed on the display 204 as "confirm" or "select" operations and performs the corresponding processing.

[0035] The screen control unit 288 controls various screens displayed on the display 204 based on various data stored in the terminal storage device 214. For example, the screen control unit 288 generates images according to the instructions of the game control unit 286 according to the progress of the game and displays them on the display 204, controls the display and operation of various objects displayed in the game space, and displays a UI (User Interface) such as icons, buttons, status, and menu screens that are used for game operation and display of game information, superimposed on the viewing area of ​​the display 204. In addition, the screen control unit 288 displays an avatar on the display 204 based on avatar data transmitted from the server 100 (described later) and stored in the terminal storage device 214 by the memory processing unit 282. It then controls the display and operation of the avatar in the game space.

[0036] The voice acquisition unit 290 receives voice input from the player via the microphone 224 mounted on the user terminal 200, acquires the player's speech, and generates voice data as digital data. For example, on a specific screen of the game, the voice acquisition unit 290 starts acquiring voice in response to a start operation from the player and stops acquiring voice in response to an end operation.

[0037] Figure 6 is a functional block diagram of the server 100 according to the first embodiment. Functionally, the server 100 includes a data transmission / reception unit 180, a memory processing unit 182, a game control unit 184, an acoustic feature extraction unit 186, a facial feature determination unit 188, and an avatar generation unit 190. These functions are realized by the server CPU 110 executing various programs stored in the server storage device 114 and causing the various hardware components of the server 100 to work together.

[0038] The data transmission / reception unit 180 transmits and receives various types of data to and from the user terminal 200. For example, when the data transmission / reception unit 180 receives login information for logging into the game from the user terminal 200, it transmits to the user terminal 200 the player data corresponding to the login information (e.g., player ID, player name, etc.), avatar data, game play data by the player (e.g., player level, experience points, possessions, event execution information, game progress information, etc.), and game data for running the game (e.g., setting information for constructing a virtual space (game space), etc.).

[0039] The memory processing unit 182 stores player data, avatar data, play data, etc., in the server storage device 114.

[0040] The game control unit 184 controls the overall progress of the game, similar to the game control unit 286 of the user terminal 200. Alternatively, the control performed by the game control unit 286 of the user terminal 200 may be performed by the game control unit 184 of the server 100.

[0041] The acoustic feature extraction unit 186 analyzes the audio data received from the user terminal 200 and extracts predetermined amounts of acoustic features (hereinafter also referred to as acoustic feature quantities). Acoustic features are physical characteristics of speech, such as the fundamental frequency which is an indicator of pitch, the formant frequency (a frequency band on the speech frequency spectrum that has a higher intensity than its surroundings) and spectral envelope (a gently fluctuating curve that represents the general shape or outline of the speech frequency spectrum) which are indicators of voice quality, and the speech rate which is an indicator of speaking speed. The acoustic feature extraction unit 186 extracts these amounts of acoustic features (acoustic feature quantities) as numerical data.

[0042] The facial feature determination unit 188 determines the facial features of the avatar based on the extracted acoustic features and a "correlation model between acoustic features and facial features" pre-built in the server memory device 114. This correlation model is based on machine learning, which has been used to determine the relationship between acoustic features and facial features. For example, this correlation model is constructed by performing machine learning (e.g., deep learning) using training data (data consisting of input data and corresponding correct answers used for training in machine learning) such as voice data from many people and survey results regarding the facial features of those people and the impressions generally associated with their voices. Facial features are parameters that define the contour of the avatar's face, the size and shape of the eyes, the shape of the nose, the shape of the mouth, the jawline, etc.

[0043] Figure 7 shows an example of the concept of a correlation model. For example, the correlation model is learned to have a high correlation between acoustic features such as "high fundamental frequency (high pitch)" and facial features such as "slender face contour" and "large eyes," and between acoustic features such as "fast speech rate" and facial features such as "almond-shaped eyes" and "sharp jawline," thus defining the relationship between such acoustic features and facial features. The facial feature determination unit 188 then determines the combination of facial features (parameters) with the highest correlation from this correlation model, according to the extracted acoustic feature quantities.

[0044] The avatar generation unit 190 generates data such as a three-dimensional model or two-dimensional image of an avatar based on the combination of facial features determined by the facial feature determination unit 188. For avatar generation, deep generative models (image generation technologies) such as generative adversarial networks (GANs) or variational autoencoders (VAEs) can be used. This dynamically generates a unique and natural avatar that evokes the owner of the input voice. The generated avatar data is transmitted to the requesting user terminal 200 via the data transmission / reception unit 180. Note that avatar generation is not limited to the above; for example, a server storage device 114 may store data of many face parts, each with different facial features, and an avatar may be generated by selecting face parts that have a high correlation with the extracted acoustic features.

[0045] Figure 8 shows an example of the configuration of various data tables stored in the server storage device 114 by the memory processing unit 182 in the first embodiment. The player data table stores data such as the player name, various player parameters (occupation, level, etc.), the avatar ID of the avatar used (an identifier that uniquely identifies the avatar), and possessions (i.e., player data, play data), associated with a player ID that uniquely identifies the player. The avatar data table stores model data that defines the appearance of the avatar (face outline, eye size, eye shape, jawline, etc.), associated with the avatar ID. The acoustic-facial correlation model table stores parameters of facial features with high correlation, associated with the type and numerical range of acoustic features. This is an example of a correlation model referenced by the facial feature determination unit 188.

[0046] (Specific example of avatar creation screen 400) Figure 9 shows an example of the avatar generation screen 400 in the game of this embodiment. The avatar generation screen 400 is displayed when creating an avatar at the start of the game or when creating a new avatar. A display area 401 is provided in the center of the screen where the generated avatar is displayed. A voice input button 404, a regenerate button 406, and a confirm button 408 are located at the bottom of the screen.

[0047] When the player taps the voice input button 404, the system transitions to the voice input screen 500, as shown in Figure 10. The voice input screen 500 displays a message 502 prompting the player to speak (for example, "Speak to generate an avatar. Please introduce yourself."). When the player speaks into the microphone 224, their voice is recorded and sent to the server 100. Once recording is complete, the system returns to the avatar generation screen 400, and the generated character 402 is displayed in the display area 401. If the player does not like the generated avatar, they can tap the regenerate button 406 to start the voice input process again. To finalize the character 402, the player taps the confirm button 408.

[0048] (Processing of the information processing system 1 in the first embodiment) Next, an example of the processing flow in the information processing system 1 of the first embodiment will be described. Here, the explanation will focus on the avatar generation process, and other processes will be omitted.

[0049] Figure 11 is a sequence diagram showing the processing flow during avatar generation in the information processing system 1 of the first embodiment. First, the user terminal 200 detects that the player has tapped the voice input button 404 on the avatar generation screen 400 (step 101). Then, the voice acquisition unit 290 acquires the player's voice via the microphone 224 and sends the voice data to the server 100 (step 102).

[0050] When the server 100 receives voice data from the user terminal 200 (step 103), the acoustic feature extraction unit 186 extracts acoustic features such as fundamental frequency and formant frequency from the received voice data (step 104). Next, the facial feature determination unit 188 refers to the acoustic-facial correlation model table and determines the facial features corresponding to the extracted acoustic features (step 105). Then, the avatar generation unit 190 generates avatar data based on the determined facial features (step 106).

[0051] The server 100 transmits the generated avatar data to the user terminal 200 (step 107). When the user terminal 200 receives the avatar data (step 108), the screen control unit 288 displays the avatar related to the received avatar data in the display area 401 of the avatar generation screen 400 (step 109). This allows the player to confirm the avatar generated based on their own voice.

[0052] Figure 12 is a flowchart showing the avatar generation process performed on the user terminal 200. This process starts, for example, when the player selects the avatar creation menu on the home screen (not specifically shown). First, the screen control unit 288 displays the avatar generation screen 400 (step 201). Then, when the player taps the voice input button 404 (YES in step 202), the screen control unit 288 displays the voice input screen 500, and the voice acquisition unit 290 accepts the voice input (step 203). The acquired voice data is transmitted to the server 100 via the data transmission / reception unit 280 (step 204). Subsequently, when avatar data is received from the server 100 (YES in step 205), the screen control unit 288 displays the avatar related to the received avatar data in the display area 401 of the avatar generation screen 400 (step 206). Then, if the player taps the confirm button 408 (YES in step 207), the generated avatar is registered as the avatar the player will use in the game (not specifically illustrated), and this process ends. If the player taps the regenerate button 406 (NO in step 207, YES in step 208), the process returns to step 202 and starts again from voice input. If neither the confirm button 408 nor the regenerate button 406 is tapped (NO in step 208), the process returns to step 207 and waits for the player's input.

[0053] Figure 13 is a flowchart showing the avatar generation control process executed on server 100. This process is triggered by the reception of audio data from user terminal 200. When the data transmission / reception unit 180 receives audio data from user terminal 200 (step 301), the acoustic feature extraction unit 186 extracts acoustic features from the received audio data (step 302). Subsequently, the facial feature determination unit 188 determines facial features corresponding to the extracted acoustic features by referring to the acoustic-facial correlation model table (step 303). Then, the avatar generation unit 190 generates avatar data using the determined facial features (step 304). The generated avatar data is transmitted to the requesting user terminal 200 via the data transmission / reception unit 180 (step 305), and this process ends.

[0054] (Examples of applying the present invention to games) The avatar generation method in the information processing system 1 of the first embodiment is applicable to various game applications. For example, it can be applied to RPGs (Roll Playing Games) in which players control their avatars and operate in a virtual space set in a fantasy world.

[0055] At the start of the game, the player creates an avatar that represents themselves. On the avatar creation screen 400 (see Figure 9), the player taps the voice input button 404, which displays a voice input screen 500 (see Figure 10), where they speak into the microphone 224, introducing themselves, describing their ideal hero, or any other words they like. For example, the player might say into the microphone 224, "I want to be a passionate wizard who mows down targets with fire magic!" in a cheerful, high-pitched voice. The user terminal 200 then sends this audio to the server 100.

[0056] In server 100, the acoustic feature extraction unit 186 extracts acoustic features such as "high fundamental frequency," "fast speech rate," and "high sound pressure level" from the received audio data. The facial feature determination unit 188 derives facial features such as "large eyes with many highlights in the pupils," "slightly upturned eyebrows," "cheerful open mouth," and "sharp, youthful contours" from these acoustic features.

[0057] Based on these facial features, and taking into account the image associated with the player's spoken content (fire, wizard), the avatar generation unit 190 generates an avatar image of a young person with, for example, reddish hair and an energetic appearance, and sends it to the user terminal 200. The avatar generation screen 400 of the requesting player displays an avatar of a passionate and cheerful wizard that reflects the impression of their own voice.

[0058] For example, if the player speaks in a low, slow voice, saying, "...an elf living with animals deep in a quiet forest...", the acoustic feature extraction unit 186 of server 100 extracts acoustic features such as "low fundamental frequency" and "slow speech rate," and derives corresponding facial features such as "calm, narrow eyes," "gentle expression," and "slender face contour," generating a quiet and mysterious elf avatar. The player then customizes the hairstyle and clothing of the generated avatar to play the RPG.

[0059] Thus, according to this embodiment, players can obtain a unique avatar that reflects their own voice tone and speaking style simply by speaking into the microphone 224, without performing any complex operations. This allows players to experience a completely new creation process, unlike conventional part-selection or photo-based avatar creation methods. Players can develop a strong sense of self-projection and attachment to avatars that reflect their inner characteristics, dramatically increasing their immersion in the game world. Furthermore, when interacting with other players via voice chat, the other person's avatar matches the impression of their voice, enabling smoother and more intuitive communication. In addition, since there is no need to register a facial photograph, players can immerse themselves in the game world with peace of mind while protecting their privacy.

[0060] (Second Embodiment) Next, we will describe the information processing system 1 of the second embodiment. In the second embodiment, the avatar's state is configured to change dynamically in conjunction with the player's biometric information. As a general rule, we will omit explanations of configurations that overlap with the first embodiment.

[0061] (Functional block) Figure 14 is a functional block diagram of the user terminal 200 according to the second embodiment. In addition to the configuration of the first embodiment (see Figure 5), the user terminal 200 in the second embodiment has a newly added biometric information acquisition unit 292.

[0062] The biometric information acquisition unit 292 acquires the player's biometric information in real time via the biometric information acquisition device 230. The acquired biometric information includes, for example, heart rate, skin electrical activity (sweating level), eye movement, pupil size (pupil diameter), electroencephalogram (EEG), muscle activity level, etc. The biometric information acquisition unit 292 transmits the acquired biometric information data to the server 100 periodically or when a specific event occurs in the game via the data transmission / reception unit 280.

[0063] Figure 15 is a functional block diagram of the server 100 according to the second embodiment. In addition to the configuration of the first embodiment (see Figure 6), the server 100 in the second embodiment newly includes a psychological state estimation unit 192 and an avatar control unit 194.

[0064] The psychological state estimation unit 192 estimates the player's psychological state based on biometric data received from the user terminal 200. This estimation is performed using a machine learning model (e.g., a biometric information-psychological state correlation model table) that has been pre-trained to recognize the correlation between biometric information patterns and psychological states (e.g., excitement, concentration, tension, fear, relaxation, fatigue, etc.). For example, the combination of biometric information "increased heart rate" and "increased sweating level" estimates the player's psychological state as "excited," and the combination of biometric information "gaze fixed on a single point" and "pupil constriction" estimates the player's psychological state as "concentrated." The psychological state estimation unit 192 outputs the estimated psychological state to the avatar control unit 194.

[0065] The avatar control unit 194 generates avatar control data to change the appearance and movements of the avatar according to the psychological state estimated by the psychological state estimation unit 192. This avatar control data includes parameters for changing the avatar's facial expression, body color, complexion, eye color, posture, hairstyle, and animation (such as breathing rate and body tremors), as well as instructions for generating effects associated with the avatar (such as an aura and sweat).

[0066] Figure 16 shows an example of the correspondence between psychological state and avatar changes. In the second embodiment, a variety of changes are defined according to psychological state. For example, the psychological state of "excitement" is associated with changes such as the avatar's hair standing on end and an aura surrounding the body, the psychological state of "concentration" is associated with changes such as the eyes shining sharply, the psychological state of "fear / tension" is associated with changes such as the face turning pale and trembling slightly, and the psychological state of "fatigue" is associated with changes such as dark circles under the eyes and a hunched posture. For example, if the psychological state is estimated to be "excitement," the avatar control unit 194 generates avatar control data that specifies changes such as the avatar's hair standing on end and an aura surrounding the body as described above. The generated avatar control data is transmitted to the corresponding user terminal 200 via the data transmission / reception unit 180. The screen control unit 288 of the user terminal 200 updates the avatar display based on the received avatar control data.

[0067] The server storage device 114 may store a biometric information-psychological state correlation model table (see Figure 16) which defines the estimated psychological state and the content of avatar changes corresponding to that psychological state, associated with combination patterns of biometric information. The psychological state estimation unit 192 may refer to this table to estimate the player's psychological state corresponding to the biometric information received from the user terminal 200, and the avatar control unit 194 may refer to this table to obtain the content of avatar changes corresponding to the estimated psychological state and generate avatar control data.

[0068] (Processing of the information processing system 1 in the second embodiment) Next, we will describe an example of the processing flow in the information processing system 1 of the second embodiment. Here, we will focus on the processing related to changes in the avatar's state during gameplay, and will omit explanations of other processes.

[0069] Figure 17 is a sequence diagram showing the processing flow when the avatar state changes in the information processing system 1 of the second embodiment. First, the biometric information acquisition unit 292 of the user terminal 200 acquires biometric information in real time from the player playing the game (step 401). The acquired biometric information is transmitted to the server 100 via the data transmission / reception unit 280 (step 402).

[0070] When the server 100 receives biometric information from the user terminal 200 (step 403), the psychological state estimation unit 192 estimates the player's psychological state based on the received biometric information (step 404). Next, the avatar control unit 194 generates avatar control data to change the appearance and movements of the avatar according to the estimated psychological state (step 405).

[0071] The server 100 transmits the generated avatar control data to the user terminal 200 (step 406). When the user terminal 200 receives the avatar control data (step 407), the screen control unit 288 changes the appearance and behavior of the avatar displayed on the display 204 in real time based on the received avatar control data (step 408). This dynamically reflects the player's psychological state in the avatar.

[0072] Figure 18 is a flowchart showing the avatar state change process executed on the user terminal 200. This process is performed continuously during gameplay. First, the biometric information acquisition unit 292 acquires the player's biometric information (step 501). The acquired biometric information is transmitted to the server 100 via the data transmission / reception unit 280 (step 502). Subsequently, when avatar control data is received from the server 100 (YES in step 503), the screen control unit 288 updates the avatar display according to the received avatar control data (step 504). After that, the process returns to step 501 and repeats the series of processes starting from the acquisition of biometric information.

[0073] Figure 19 is a flowchart showing the avatar state change control process executed on server 100. This process is triggered by the reception of biometric information from user terminal 200. When the data transmission / reception unit 180 receives biometric information from user terminal 200 (step 601), the psychological state estimation unit 192 estimates the player's psychological state based on the received biometric information (step 602). Subsequently, the avatar control unit 194 generates avatar control data corresponding to the estimated psychological state (step 603). The generated avatar control data is transmitted to the requesting user terminal 200 via the data transmission / reception unit 180 (step 604), and this process ends.

[0074] (Examples of applying the present invention to games) The method for changing the state of an avatar in the information processing system 1 of the second embodiment can be applied when playing a game using an avatar generated by the avatar generation method of the first embodiment, or an avatar created by a conventional method. Here, we assume a case where the player wears a smartwatch or a VR headset to play the aforementioned RPG.

[0075] For example, consider a scenario where the player confronts a powerful boss monster B. As the battle intensifies and the player becomes excited, the heart rate sensor built into the smartwatch detects a sudden increase in heart rate. This biometric information is transmitted to the server 100 in real time, and the psychological state estimation unit 192 estimates that the player's psychological state is "excited." As a result, based on instructions from the avatar control unit 194, the hair of avatar A on the game screen 600 stands on end, and a red aura rises from its body (see Figures 20(a)~(b)). The player feels as if their own excitement has been transferred to avatar A, further deepening their immersion in the battle.

[0076] Furthermore, when the player attempts to solve the puzzles of a complex ancient ruin, the VR headset's eye-tracking sensor detects when the player's gaze is focused on a specific pattern for an extended period. From this biometric information, the psychological state estimation unit 192 estimates that the player's psychological state is "concentrated." In response, the eyes of avatar A on the game screen 600 light up sharply, and a subtle highlight S is applied to areas that provide hints for solving the puzzle (see Figures 21(a)~(b)). The player realizes that their concentration is influencing the game's progression, allowing them to engage in puzzle-solving more actively.

[0077] Conversely, if a player's heart rate decreases and their eye movements become sluggish due to prolonged gameplay, the psychological state estimation unit 192 estimates the player's psychological state to be "fatigued" based on this biometric information. In response, an animation is played in which avatar A on the game screen 600 becomes pale, develops dark circles under its eyes, hunchs slightly, and sighs (see Figures 22(a)~(b)). This allows the player to objectively recognize their own fatigue and provides an opportunity to take a break. In addition, controls such as displaying a message encouraging a break or temporarily lowering the game difficulty may be implemented in conjunction with the playback of this animation.

[0078] Furthermore, this method of changing the state of an avatar can be applied to various genres of games other than RPGs.

[0079] For example, in an action game where players control avatars to compete against each other, if a player's heart rate suddenly increases due to taking damage from an enemy attack or inflicting significant damage on an enemy, the psychological state estimation unit 192 estimates that the player's psychological state is "excited." Then, although not specifically illustrated, an animation is played in which the avatar's hair stands on end, an aura emanates from its eyes and body, and its breathing becomes rapid. This directly reflects the player's excitement in the avatar's appearance, enhancing the sense of realism in the battle. Furthermore, by observing the changes in the opponent's avatar's state, the opposing player can infer that the opponent is excited, for example, the timing of using a special move or, conversely, being cornered, enabling a deeper level of strategy. In addition, the subtle manifestation of the opponent player's psychological state (anxiety, carelessness, etc.) in the avatar allows for new strategic possibilities.

[0080] For example, in puzzle games or shooting games that require precise operation and thinking, if the player's gaze is fixed on one spot for a long time and their pupils constrict while solving a difficult puzzle or dodging enemy bullets, the psychological state estimation unit 192 estimates that the player's psychological state is "concentration." In this case, although not specifically illustrated, an animation is played in which the avatar's eyes light up sharply and their expression changes to one of furrowed brows. Depending on the game, a slow-motion effect may be activated to make the flow of time in the game appear slower in order to assist the player's concentration. This allows the player to visually recognize their own state of concentration and encourages further improvement in their concentration.

[0081] For example, in a horror game designed to provide players with a terrifying experience, if a terrifying monster suddenly appears while the player is exploring a dark corridor, causing their heart rate to skyrocket and their skin electrical activity (sweating level) to increase, the psychological state estimation unit 192 estimates the player's psychological state to be "fear" and "anxiety." Then, although not specifically illustrated, an animation is played in which the avatar's face turns pale and its body begins to tremble slightly. This change in the avatar feeds back the player's own sense of fear into the game world, reflecting it like a mirror, creating an unprecedented sense of immersion and a terrifying experience. Furthermore, depending on the game, it is possible to add a new gameplay element of overcoming fear by directly penalizing gameplay, such as making the avatar's movements slower or narrowing its field of view, as the player's level of fear increases.

[0082] Thus, according to the second embodiment, the player's invisible inner state is reflected in the avatar's appearance and movements, causing the avatar to change dynamically. This dramatically improves the sense of unity between the player and the avatar, making it possible to achieve an unprecedented level of immersion. Furthermore, in multiplayer, simply by looking at the other player's avatar, their tension or excitement can be conveyed, realizing a new dimension of communication that does not rely on words.

[0083] (modified version) Although one embodiment of the present invention has been described above, the present invention is not limited to the above embodiment, and various modifications are possible without departing from the spirit of the invention.

[0084] In the above embodiment, the server 100 was capable of performing avatar generation and avatar state change control, but it is not limited to this. For example, the user terminal 200 may be equipped with avatar generation and avatar state change control functions, and the player may generate and change avatars using the functions provided by the user terminal 200. In other words, each process in the above embodiment only needs to be executed on at least one of the server 100 and the user terminal 200, and the timing of execution and the device on which it is executed are not particularly limited. Furthermore, the functions of the server 100 may be all provided by a single server 100, or they may be distributed among multiple servers 100. In such cases, the same effects and advantages as in the above embodiment will be achieved.

[0085] In the above embodiment, a smartphone was used as the user terminal 200, and the game could be run using a dedicated application installed on the user terminal 200. However, the type of user terminal 200 and the game execution tool are not limited to this. For example, a personal computer may be used as the user terminal 200, or the game may be run using a general web browser. In such cases, the same effects and advantages as in the above embodiment will be achieved.

[0086] In the above embodiment, the design of the avatar's face was generated according to the acoustic characteristics of the player's voice, but the design generated based on acoustic characteristics is not limited to the face. For example, the design of the entire avatar (body type, clothing, hairstyle, player's race (human, elf, dwarf, etc.)) may be generated from the acoustic characteristics of the player's voice. Furthermore, the design is not limited to human faces, and designs of animals, monsters, abstract characters, etc., may be generated from the acoustic characteristics of the voice. In the above cases, the same effects and advantages as in the above embodiment will be achieved.

[0087] In the above embodiment, the player's voice was acquired when generating the avatar, but this is not the only way to do so. For example, the avatar may be generated using voice data pre-stored in the user terminal 200. In this case, the same effects and advantages as in the above embodiment will be achieved.

[0088] Although the information processing system 1 in the above embodiment provided a game, for example, an information processing system that provides an image distribution service may also perform avatar generation as in the above embodiment. That is, the avatar displayed in the image distribution service may be generated based on the acoustic characteristics of the user's voice, as in the above embodiment. In this case, the same effects and advantages as in the above embodiment will be achieved.

[0089] Furthermore, in an information processing system that provides communication services in a virtual space such as social VR, the state change control of the avatar as in the above embodiment may be implemented. When this is done, the avatar can express subtle emotions that cannot be conveyed by words or facial expressions alone, thereby facilitating smoother communication.

[0090] The biometric information acquired in the avatar state change control of the above embodiment is not limited to heart rate, skin electrical activity, eye movement, pupil size, electroencephalogram, muscle activity, etc., but may also include any other biometric information that reflects the user's psychological state, such as respiratory rate, blood pressure, and body temperature. Furthermore, the estimated psychological state is not limited to excitement, concentration, tension, fear, relaxation, fatigue, etc., but may include a wider range of emotions and states such as joy, sadness, and surprise.

[0091] In the avatar state change control of the above embodiment, the appearance and movements of the avatar are changed according to the psychological state estimated from the acquired biometric information, but the system is not limited to this, and game effects may be executed according to the psychological state. For example, as described above, if the player's psychological state is estimated to be "fatigued," the system may display a message prompting the player to take a break or temporarily lower the difficulty of the game. Also, for example, if the player's psychological state is estimated to be "feared," the system may display effects or output background music that further enhances the feeling of fear. Doing so can improve the enjoyment of the game.

[0092] The program for executing the processing in the above embodiment may be stored in a computer-readable storage medium and provided as such. Furthermore, the above embodiment may not be a client-server type information processing system, but rather an information processing device that stores the above program and can independently execute processing similar to that of an information processing system, or it may be an information processing method that realizes each function and the steps shown in the flowchart.

[0093] The acoustic feature extraction unit 186 in the above embodiment corresponds to an example of the extraction means of the present invention. The facial features in the above embodiment correspond to an example of the visual features of the present invention. The facial feature determination unit 188 in the above embodiment corresponds to an example of the determination means of the present invention.

[0094] (Note) Some of the features of the present invention are described below.

[0095] <Challenges> For example, it can generate avatars that reflect the user's personality. <Solution> (1) A program that causes a computer to function as an extraction means for extracting acoustic features from a user's voice, a determination means for determining visual features corresponding to the extracted acoustic features, and a generation means for generating an avatar having the determined visual features. (2) The decision means is the program from (1) which determines the visual features using a machine learning model that has learned the correlation between acoustic features and visual features. Furthermore, the solutions constructed in the above program may be adapted to the fields of devices, systems, methods, and media as appropriate. <Effects> According to the programs described in (1) and (2) above, for example, it is possible to generate an avatar that reflects the user's personality. [Explanation of symbols]

[0096] 1. Information Processing System 100 servers 186 Acoustic Feature Extraction Unit 188. Parts that determine facial features 190 Avatar Generation Section 192 Psychological State Estimation Unit 194 Avatar Control Unit 200 user terminals 224 Mike 230 Biometric Information Acquisition Devices 290 Voice acquisition unit 292 Biological Information Acquisition Unit

Claims

[Claim 1] An extraction method for extracting acoustic features from the user's voice, A determination means for determining visual features corresponding to extracted acoustic features, A generation means for generating an avatar having determined visual characteristics, and a computer that functions as such. The decision-making method involves inputting the extracted acoustic features into a correlation model that has learned the relationship between acoustic and visual features, and determining the visual feature with the highest correlation. The generation method is a program that generates an avatar by adding the content of the user's voice to the determined visual features.