Avatar generation system and communication system
The centralized management of avatar generation APIs simplifies the process of creating avatars by automatically selecting the optimal service based on user input, addressing the inefficiencies of managing multiple services and reducing costs.
Patent Information
- Application Number
- PCT/JP2025/014064
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-04-08
- Filing Date
- 2025-04-08
- Publication Date
- 2025-10-16
AI Technical Summary
Users face inconvenience and inefficiency when using multiple avatar generation services due to differing specifications and capabilities, requiring trial-and-error to select the appropriate service and manage various APIs.
A system that centrally manages multiple avatar generation APIs, allowing users to input parameters and automatically select the optimal API based on user preferences, integrating facial expression, voice, and personality information to generate avatars efficiently.
Enables easy avatar generation across multiple services without understanding individual specifications, reduces costs, and enhances user convenience by automating the process and combining capabilities for high-quality results.
Smart Images

Figure JP2025014064_16102025_PF_FP_ABST
Abstract
Description
Avatar generation system and communication system
[0001] The present invention relates to a system that centrally manages multiple APIs, selects an appropriate API in response to a request, and generates an avatar, and a communication system that uses the avatar generated by the system.
[0002] In recent years, with the development of computer graphics technology, many services are being provided that generate avatars according to user preferences. These services allow users to select parameters such as hairstyle, eye color, and clothing to obtain the avatar they desire, and input them into an avatar generation API to generate the avatar.
[0003] In this technical field, a method for creating and editing an avatar and navigating an avatar selection interface (see Patent Document 1) and a faster and more efficient interface for creating and editing avatars (Patent Document 2) have been disclosed.
[0004] JP 2022-008470 A JP 2023-085356 A
[0005] However, there are many providers offering avatar generation services, and each has different avatar generation API specifications and parameter setting methods. As a result, when a user uses multiple avatar generation services, they need to understand the specifications of each service and set parameters, which can be inconvenient.
[0006] In addition, because the avatar generation capabilities of each avatar generation service differ, users must select an appropriate service to obtain the avatar they desire. However, it is not easy for users to understand the avatar generation capabilities of each service, which forces them to use the service on a trial-and-error basis.
[0007] The present invention has been made in view of the above-mentioned problems, and aims to provide a system that makes it easier to generate avatars.
[0008] According to the present invention, the following is obtained.
[0009] According to the present invention, the following effects are achieved.
[0010] By centrally managing multiple avatar generation APIs, users can easily generate avatars without having to worry about the specifications of each service.
[0011] By managing the avatar generation capabilities of each avatar generation API, it is possible to select the most suitable avatar generation API to generate an avatar according to the user's request. By utilizing avatar generation history information, it is possible to generate an avatar that suits the user's preferences.
[0012] By having a computer execute the avatar generation process, avatar generation can be automated, improving user convenience.
[0013] Centralized management allows us to understand the usage status of each avatar generation API, which can be used to improve services and provide new services.
[0014] By combining multiple avatar generation APIs, it is possible to generate avatars that cannot be achieved with a single service.
[0015] It reduces the cost of generating avatars. Compared to using multiple services individually, centralized management reduces API usage fees.
[0016] As described above, according to the present invention, a plurality of avatar generation services can be efficiently used, improving user convenience.
[0017] FIG. 1 is a diagram showing an example of the configuration of an avatar generation system (hereinafter referred to as "the system") according to an embodiment of the present invention. It is a functional block diagram of the system of FIG. 1. It is an example of a management screen provided to a user by the system avatar management server of FIG. 1. It is a sequence diagram of processing of the system of FIG. 1. It is a schematic diagram showing a fixed conversation registration function of a communication system using an avatar generated by the system of FIG. 1. It is a schematic diagram showing a fixed conversation candidate registration function of a communication system using an avatar generated by the system of FIG. 1.
[0018] The details of embodiments of the present invention will be described below. The present invention has the following configuration: [Item 1] An avatar generation system including at least a facial expression generation API provision server, a voice generation API provision server, a personality generation API provision server, and an avatar management server, wherein the avatar management server includes: a parameter receiving unit that receives input of facial expression parameters, voice parameters, and personality parameters from a user; a parameter transmitting unit that transmits the facial expression parameters, voice parameters, and personality parameters to the corresponding API provision servers; a generation information acquiring unit that receives facial expression information, voice information, and personality information generated by each of the API provision servers; and an avatar generation unit that generates an avatar by integrating the facial expression information, voice information, and personality information. [Item 2] The avatar generation system according to item 1, wherein the avatar management server further includes an avatar storage unit that stores the generated avatar. [Item 3] The avatar generation system according to item 1 or 2, wherein the avatar management server further comprises a preview unit that reads and displays a saved avatar in response to a preview request from the user. [Item 4] The avatar generation system according to any one of items 1 to 3, wherein the avatar management server further comprises an avatar editing unit that edits the generated avatar based on editing instructions from the user. [Item 5] The avatar generation system according to any one of items 1 to 4, wherein the avatar management server further comprises an avatar sharing unit that shares the generated avatar with other users based on sharing instructions from the user. [Item 6] The avatar generation system according to any one of items 1 to 5, wherein the avatar management server further comprises an API selection unit that selects an optimal API provision server from each of the plurality of facial expression generation API provision servers, the plurality of voice generation API provision servers, and the plurality of personality generation API provision servers.
[0019] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS Hereinafter, preferred embodiments of the present invention will be described with reference to the accompanying drawings.
[0020] 1, an avatar generation system (hereinafter sometimes referred to as the "system") 1 according to an embodiment of the present invention relates to a system that generates an avatar by combining multiple APIs. In the following embodiment, the system 1 is described as comprising a facial expression generation server 30, a voice generation API providing server 31, and a personality generation API providing server 33, which respectively provide a facial expression generation API, a voice generation API, and a personality generation API, and an avatar management server 20 that integrates these APIs to generate an optimal avatar.
[0021] According to the present invention, a user can easily create an avatar by simply inputting facial expression, voice, and personality parameters on the platform system provided by the avatar management server 20. Furthermore, when there are multiple API providers, the optimal API can be dynamically selected, thereby improving the performance and efficiency of the system.
[0022] <Example of hardware configuration> This system is configured to be able to communicate with each other via the Internet. Note that this system may be configured as a cloud-based / network-based system provided by a designated business operator, or as an on-premise system that is operated independently within the company that adopts the system.
[0023] Each of the above-mentioned functional blocks can be configured by, for example, hardware provided in a server device (terminal device), a DSP (Digital Signal Processor), or software. For example, when configured by software, each of the above-mentioned functional blocks is actually configured with a CPU, RAM, ROM, etc. of a computer, and is realized by the operation of a program stored in a recording medium such as RAM, ROM, a hard disk, or a semiconductor memory.
[0024] <Network Configuration> This system consists of a facial expression generation API providing server, a voice generation API providing server, a personality generation API providing server, and an avatar management server. Each server is an independent computer system, and they are connected to each other via a network so that they can communicate with each other. The avatar management server accepts requests from user terminals and sends and receives data to and from each API providing server. The avatar management server also has a database for storing the generated avatars.
[0025] <Avatar Management Server> As shown in FIG. 2, avatar management server 20 according to this embodiment has the following functions for configuring the system.
[0026] <Storage Unit> The storage unit 101 is implemented in a storage device such as a hard disk or SSD in the avatar management server. Data is managed using a relational database, NoSQL database, or the like. The storage unit 101 is configured to ensure data consistency and permanence, and to enable high-speed reading and writing.
[0027] The memory unit 101 is located within the avatar management server and stores various data necessary for the system's operation. The main items stored are as follows: ・Avatar data: Stores data such as the image, voice, personality, and movements of the generated avatar. Data received from the avatar storage unit is stored in a format that is easy to search and reuse. ・User data: Stores information about users who use the system. This includes user IDs, passwords, personal settings, and purchase history. ・API data: Stores information about servers that provide various APIs used by the system, such as facial expression generation APIs, voice generation APIs, and personality generation APIs. This includes access information, authentication information, and usage status for the API-providing servers. ・Setting data: Stores various setting information for controlling the operation of the system 1. This includes API selection criteria, data storage format, and security settings. The memory unit 101 also reads and writes data from other functional sections of the avatar management server as appropriate. For example, avatar data generated by the avatar generation unit may be stored in the memory unit 101, and API data used by the API selection unit may be read from the memory unit 101.
[0028] <Parameter Receiving Unit> The parameter receiving unit 102 receives input of facial expression parameters, voice parameters, and personality parameters required for avatar generation from a user terminal. These parameters are used by the user to customize the avatar's facial expression, voice, and personality. The parameter receiving unit provides an interface with the user terminal, analyzes and verifies the parameters entered by the user, and then makes them available within the system for subsequent processing.
[0029] <API Selection Unit> The API selection unit 103 is responsible for selecting the optimal server from among multiple API servers available for facial expression generation, voice generation, and personality generation. Selection criteria include API response speed, the quality of the generated facial expressions, voice, and personality, the cost of using the API, and the reliability of the API. The API selection unit may evaluate APIs based on these criteria and dynamically select the optimal API for each task. This allows the system to efficiently use resources and generate high-quality avatars.
[0030] <Parameter Transmission Unit> The parameter transmission unit 104 transmits corresponding parameters to the API providing server selected by the API selection unit. Facial expression parameters are transmitted to the facial expression generation API, voice parameters to the voice generation API, and personality parameters to the personality generation API. The parameter transmission unit converts the parameters according to the data format required by each API and transmits them to the API providing server via the network. The parameter transmission unit also monitors responses from the API providing server and takes appropriate action if an error occurs.
[0031] <Generation Information Acquisition Unit> The generation information acquisition unit 105 receives facial expression information, voice information, and personality information returned from the facial expression generation API, voice generation API, and personality generation API. The generation information acquisition unit manages communication with each API and verifies the consistency of the received data. The received information is passed to the avatar generation unit and integrated.
[0032] <Avatar Generation Unit> The avatar generation unit 106 generates the final avatar by integrating the facial expression information, voice information, and personality information received from the generation information acquisition unit. The avatar generation unit uses 3D modeling and rendering techniques to create a consistent avatar with facial expressions, voice, and personality. The generated avatar is displayed on the user's device and is also passed to the avatar storage unit for further processing.
[0033] <Avatar Storage Unit> The avatar storage unit 107 stores the avatars generated by the avatar generation unit in a database. Avatar data includes facial expression, voice, and personality information. The avatar storage unit stores the avatar data in a structured database in a format that is easy to search and update. The saved avatars can be used by the user to reuse or edit the avatars later.
[0034] <Avatar Reader> The avatar reader 108 reads a stored avatar from the database in response to an avatar playback request from a user terminal. When a user requests avatar playback, the avatar reader retrieves avatar data from the database and sends it to the user terminal. The avatar reader manages connections with the database to enable efficient data search and access.
[0035] <Avatar Editing Unit> The avatar editing unit 109 provides a function for editing the generated avatar based on editing instructions from the user terminal. The user transmits editing instructions to change the avatar's facial expression, voice, personality, etc. The avatar editing unit receives these instructions and modifies the avatar data. The edited avatar is saved in a database and transmitted to the user terminal.
[0036] <Avatar Sharing Unit> The avatar sharing unit 110 provides a function for sharing a generated avatar with other users based on a sharing instruction from a user terminal. A user can send an avatar to other users or share it on social media. The avatar sharing unit receives a user's sharing instruction, converts the avatar data into an appropriate format, and sends it to the specified destination. The avatar sharing unit also performs privacy settings and authority management to prevent unauthorized sharing of the avatar.
[0037] <Configuration of API Providing Server> The API providing server according to this embodiment has the following functions according to the respective characteristics.
[0038] <Facial Expression Generation Processing Unit> The facial expression generation API server generates facial expressions for the avatar based on the facial expression parameters received from the avatar management server. The facial expression parameters include the degree of eye opening, mouth shape, and eyebrow position. The facial expression generation API uses these parameters as input to generate facial expressions in the form of a 3D model or 2D image. The generated facial expression information is sent back to the avatar management server.
[0039] <Speech Generation Processing Unit> The speech generation API server generates the speech of the avatar based on the speech parameters received from the avatar management server. The speech parameters include pitch, speed, intonation, etc. The speech generation API uses these parameters to generate the speech of the avatar using text-to-speech or speech synthesis technology. The generated speech information is sent back to the avatar management server.
[0040] <Personality Generation Processing Unit> The personality generation API server generates an avatar's personality based on the personality parameters received from the avatar management server. Personality parameters include cheerfulness, positivity, friendliness, etc. The personality generation API uses these parameters as input to generate an avatar's personality using natural language processing and dialogue system technology. The generated personality information is used to control the avatar's dialogue and behavior and is sent back to the avatar management server.
[0041] In addition to the API provider server described above, the following APIs may be used in combination or independently: ・Movement Generation API: Generates the avatar's body movements and gestures. Natural movements such as walking, running, jumping, and waving can be generated. This allows for more realistic and lifelike avatars. ・Facial Expression Recognition API: Recognizes the user's facial expressions and reflects them in the avatar. When the user smiles or gets angry at the camera, the avatar reproduces the same facial expression. This allows for emotional synchronization between the user and the avatar. ・Speech Recognition API: Recognizes the user's voice and reflects it in the avatar's responses and behavior. When the user speaks, the avatar responds appropriately or takes the requested action. This allows for interactive communication with the avatar. ・Natural Language Processing API: Enables natural conversation with the avatar. Understands what the user says and generates appropriate responses based on the context. This allows for realistic dialogue with the avatar. ・Sentiment Analysis API: Analyzes emotions from text and voice and reflects them in the avatar's facial expressions and speaking style. The avatar responds empathetically by reading the user's emotional responses. This allows for more natural and empathetic interactions with the avatar. ・Rendering API: Renders avatars using high-quality 3D graphics. Avatars are smoothly animated in real time and displayed in high resolution. This enhances the visual immersion. ・Physics Simulation API: Physically simulates the movement of the avatar's hair and clothing. This allows hair to flow in the wind and clothing to sway with movement. This further enhances the realism of the avatar. ・Virtual Background API: Places avatars in various virtual environments. Generates any background, such as a room, city, or natural environment, allowing the avatar to blend seamlessly into the surroundings. This allows avatars to be used in a variety of scenes. ・Accessory Generation API: Equips avatars with hats, glasses, accessories, and more. Users can dress their avatars to their liking. This expands the variety of avatar appearances. ・Animation Control API: Controls avatar behavior. Apply animations to avatars and customize their movements through the API.This allows for fine-tuning of the avatar's movements.
[0042] By combining these APIs, it is possible to create more expressive and attractive avatars. The avatar's appearance, movement, dialogue, emotional expression, etc. can be enhanced in multiple ways, deepening interaction with the user.
[0043] <Processing Flow> The processing flow of the system 1 according to this embodiment will be described with reference to FIGS.
[0044] When a user logs in and accesses avatar management server 20, the management screen shown in Fig. 3 is displayed. As shown in the figure, the management screen displays an API selection area G and an avatar management area A. In API selection area G, buttons Btn for inputting parameters for various APIs are displayed. When each button Btn is selected, a parameter input screen for the corresponding QPI is displayed.
[0045] When the user inputs facial expression parameters, voice parameters, and personality parameters into each parameter input screen (SQ001), the parameter receiving unit of the avatar management server 20 receives the parameters, and the API selecting unit selects the most suitable API providing server (SQ002).
[0046] Next, the parameter transmission unit transmits the corresponding parameters to the selected API providing server. That is, the avatar management server 20 transmits the parameters necessary for generating facial expressions to the facial expression generation API providing server 30 (SQ003). The parameters here include not only whether or not to generate the facial expression, but also items for adjusting the generation process, such as instructions, numerical values, degree, frequency, intensity, start and end times, input by the user. The technology provided by the facial expression generation API according to this embodiment is lip synchronization technology. Specifically, the user uploads an image to which they wish to apply lip synchronization to the avatar management server. The uploaded image is sent to the selected facial expression generation API. The facial expression generation API performs the following processes on the received image: Detects a face in the image and extracts facial feature points (mouth, eyes, nose, etc.). Analyzes the specified audio data and extracts audio features (phonemes, pitch, volume, etc.). Generates lip synchronization by shifting facial feature points chronologically based on the audio features. Generates an image to which lip synchronization has been applied and returns it to the avatar management server. This technology allows users to apply lip synchronization to still images, making it possible to generate moving images that make it appear as if the images are speaking.
[0047] Similarly, the avatar management server 20 transmits the parameters necessary for generating facial expressions to the voice generation API providing server 31 (SQ005). The technology provided by the voice generation API providing server according to this embodiment allows a user to upload a voice sample of a few minutes, and the system learns the voice features of that user from the voice sample to generate all voices in the same voice as that user. Specifically, the user uploads a voice sample of the person for whom voice generation is desired via the avatar management server. The uploaded voice sample is sent to the selected voice generation API providing server. The voice generation API performs the following processes on the received voice sample: - Analyzes the voice sample and extracts the voice features (timbre, pitch, intonation, etc.) of that person. - Uses the extracted features to train a voice synthesis model to reproduce that person's voice. - When any text is input, the trained voice synthesis model generates a voice in which the text is read aloud in that person's voice. - Returns the generated voice data to the avatar management server. This technology allows a user to create a voice synthesis model that reproduces a specific person's voice from a short voice sample.
[0048] Similarly, the avatar management server 20 transmits parameters necessary for generating the avatar's personality to the personality generation API providing server 32 (SQ007). The technology provided by the personality generation API providing server according to this embodiment provides technology for generating the personality, tone of voice, etc. of the generated avatar when the generated avatar uses a generative AI to respond to chats, voice, etc. Specifically, the user specifies the personality, tone of voice, etc. to be assigned to the avatar as a prompt input via the avatar management server. The specified prompt is sent to the selected personality generation API. The personality generation API performs the following processes based on the received prompt input: ・Analyzes the prompt input and extracts characteristics such as the specified personality and tone of voice. ・Uses the extracted characteristics to generate a personality model for defining the avatar's personality and tone of voice. ・Applying the generated personality model to the generative AI enables chats and voice responses based on the personality and tone of voice. ・Returns the generated personality model to the avatar management server. This technology allows the user to freely define the avatar's personality and tone of voice with simple prompt input. Furthermore, the personality model generated by this technology can be seamlessly integrated with generative AI, enabling consistent responses based on the personality and tone of voice, providing a more immersive experience for users.
[0049] In this way, facial expression information, voice information, and personality information are generated based on the parameters received by each API providing server. Each piece of generated information is returned to the avatar management server (SQ004, SQ006, SQ008).
[0050] The generation information acquisition unit of the avatar management server receives the generation information from each API providing server, and the avatar generation unit integrates the received information to generate an avatar (SQ009).
[0051] The avatar management server stores the generated avatar in a database and provides it to the user terminal 10 (SQ010). When a user requests to play, edit, or share an avatar, the avatar reading unit, avatar editing unit, and avatar sharing unit can perform the corresponding processing.
[0052] As described above, the present invention is a system that centrally manages multiple avatar generation APIs and selects the optimal API in response to a user request to generate an avatar. Specifically, the management system of the present invention is capable of communicating with avatar generation APIs provided by multiple providers and selects the optimal API in response to a user request for avatar generation. The selected API is then assigned parameters based on the user's request and a request is made to generate an avatar. The generated avatar is acquired by the management system and provided to the user. This series of processes allows the user to obtain a desired avatar with simple operations without having to be aware of multiple avatar generation services.
[0053] According to this invention, users can easily generate avatars through a single management system, without having to understand the specifications of multiple avatar generation services individually. Furthermore, the management system understands the avatar generation capabilities of each API and selects the optimal API according to the user's requirements, allowing for efficient generation of high-quality avatars. Furthermore, when adding a new avatar generation API, it can be provided to users simply by establishing a connection with the management system, ensuring the flexibility and scalability of the system.
[0054] <Communication System> The avatars generated by the avatar generation system according to the above-described embodiment can be implemented as the following communication system. That is, the communication system of this embodiment includes a statement receiving unit that receives statements from users via chat, an answer information acquisition unit that communicates with an answer generation API providing server that generates answers in response to the statements and acquires answer information for the statements, and an avatar statement generation unit that generates statements for the generated avatar based on the answer information.
[0055] The comment receiving unit provides a chat interface between the user and the avatar, and receives text or voice comments from the user. The received comments are sent to the answer information acquisition unit.
[0056] The answer information acquisition unit communicates with the answer generation API providing server to generate an answer corresponding to the received utterance. The answer generation API providing server uses a large-scale language model to generate a natural answer to the utterance. The generated answer information is sent to the avatar utterance generation unit via the answer information acquisition unit.
[0057] The avatar utterance generator generates a utterance for the avatar generated by the avatar generation system based on the response information. The generated utterance is converted into text or voice and presented to the user via a chat interface.
[0058] The avatar utterance generator can also generate facial expressions and movements for the avatar based on the response information and display them in sync with the utterance, thereby achieving more realistic and natural communication with the avatar.
[0059] According to the communication system of this embodiment, it is possible to have natural conversations with users using avatars generated by the avatar generation system. By utilizing advanced language processing technology provided by the answer generation API server, the avatars can generate appropriate and natural answers to user comments, realizing realistic communication.
[0060] <Fixed Conversation Registration Function> As shown in Figure 5, the communication system may also include a question and answer registration unit that registers predetermined questions and answers, and a question and answer determination unit that sends the registered answer directly to the avatar statement generation unit when a user's comment matches the question. This allows a pre-prepared answer to be returned without going through the generative AI, preventing unintended answers and enabling faster response times.
[0061] The question and answer registration unit has a database for registering predetermined question contents and answer contents. For example, the following question and answer pairs are registered in this database. Question content: What is your favorite animal? Answer content: I like dogs. Question content: What is your favorite food? Answer content: Apples. The question and answer determination unit analyzes the content of the user's statement received by the statement receiving unit and compares it with the question content registered in the question and answer registration unit. If the content of the statement matches the registered question content or can be interpreted as being located, the question and answer determination unit sends the corresponding answer content registered in the question and answer registration unit directly to the avatar statement generation unit without going through the answer information acquisition unit.
[0062] For example, if a user says, "What is your favorite animal?", the question and answer determination unit detects that the content of this statement matches the question content registered in the question and answer registration unit. Then, it obtains the corresponding answer, "I like dogs," and sends it to the avatar statement generation unit.
[0063] The avatar statement generator generates a statement for the avatar using the answer received from the question and answer determination unit. In this case, the prepared answer is used as the avatar's statement without going through the answer generation API server.
[0064] By registering questions and answers in advance, the avatar can provide appropriate answers without using generative AI. This prevents unintended answers, shortens the processing time for answer generation, and enables high-speed response.
[0065] 6, the system may further include a fixed conversation registration unit that registers the user's utterance content received by the utterance receiving unit and the avatar's utterance content generated by the avatar utterance generation unit as a question and an answer in the question and answer registration unit based on an instruction from the user. This allows the user to register the question and answer at any time as a "fixed conversation" while communicating with an avatar using generative AI.
[0066] For example, if a user asks, "What are some recommended tourist spots?" and the avatar replies, "My recommendation is Kamakura. The hydrangeas at Hasedera Temple are very beautiful," and the user likes this exchange, they can register it in the fixed conversation registration section, so that the same answer will be returned for the same question in the future.
[0067] The fixed conversation registration unit has the function of registering the user's statement content received by the statement receiving unit and the avatar's statement content generated by the avatar statement generation unit as question content and answer content in the question and answer registration unit based on instructions from the user.
[0068] If a user wishes to register a question and answer during a conversation with an avatar as a fixed conversation, the user performs a predetermined operation (e.g., clicks the "Register this conversation as a fixed conversation" button). When the fixed conversation registration unit detects this operation, it obtains the user's most recent statement and the avatar's statement and registers them in the question and answer registration unit.
[0069] For example, suppose the following exchange took place: User: What's your least favorite food? Avatar: Pears. The user likes this conversation and presses the register button as a fixed conversation. The fixed conversation registration unit registers the user's statement as the question and the avatar's statement as the answer in the question and answer registration unit.
[0070] Once this registration is complete, when the user subsequently utters, "What is your least favorite food?", the question and answer determination unit will detect that the content of this utterance matches the question registered in the question and answer registration unit, and will send the registered answer to the avatar utterance generation unit. As a result, the avatar will always respond with, "Pears."
[0071] In this way, by providing a fixed conversation registration unit, users can fix their favorite exchanges from natural conversations with the generative AI, realizing reproducible communication. This makes it possible to customize communication with the avatar according to the user's preferences.
[0072] <Industrial Application Fields> The avatar generation system according to the embodiment of the present invention described above can be applied to avatar generation in various fields, such as games, entertainment, education, and communication. In particular, by combining multiple APIs, it is possible to efficiently generate high-quality avatars with different elements such as facial expressions, voices, and personalities, which can greatly contribute to the development and provision of services that utilize avatars. Furthermore, the API selection function allows the generation of an optimal avatar depending on the situation, which can also contribute to improving user satisfaction.
[0073] The above-described embodiment is merely an example for facilitating understanding of the present invention, and is not intended to limit the present invention. The present invention can be modified and improved without departing from the spirit thereof, and it goes without saying that the present invention includes equivalents thereof.
[0074] 10 Legal affairs officer terminal 20 Client terminal 100 Legal consultation support server
Claims
1. An avatar generation system including at least a facial expression generation API providing server, a voice generation API providing server, a personality generation API providing server, and an avatar management server, wherein the avatar management server comprises: a parameter receiving unit that receives input of facial expression parameters, voice parameters, and personality parameters from a user; a parameter sending unit that sends the facial expression parameters, voice parameters, and personality parameters to each of the corresponding API providing servers; a generation information acquisition unit that receives facial expression information, voice information, and personality information generated by each of the API providing servers; and an avatar generation unit that generates an avatar by integrating the facial expression information, voice information, and personality information.
2. An avatar generation system according to claim 1, wherein the avatar management server further comprises an avatar storage unit that stores the generated avatar.
3. An avatar generation system according to claim 1 or claim 2, wherein the avatar management server further comprises a preview unit that reads out and displays a saved avatar in response to a preview request from the user.
4. An avatar generation system according to any one of claims 1 to 3, wherein the avatar management server further comprises an avatar editing unit that edits the generated avatar based on editing instructions from a user.
5. An avatar generation system according to any one of claims 1 to 4, wherein the avatar management server further comprises an avatar sharing unit that shares the generated avatar with other users based on a sharing instruction from the user.
6. An avatar generation system according to any one of claims 1 to 5, wherein the avatar management server further comprises an API selection unit that selects an optimal API providing server from each of the plurality of facial expression generation API providing servers, the plurality of voice generation API providing servers, and the plurality of personality generation API providing servers.
7. A communication system for conversing with an avatar generated by the avatar generation system of any one of claims 1 to 6, comprising: a statement receiving unit that receives statements from users via chat; an answer information acquisition unit that communicates with an answer generation API providing server that generates answers in response to the statements and acquires answer information to the statements; and an avatar statement generation unit that generates statements for the generated avatar based on the answer information.
8. A communication system as claimed in claim 7, comprising: a question and answer registration unit that accepts registration of predetermined question contents and answer contents; and a question and answer determination unit that, when the content of a statement accepted by said statement acceptance unit matches the content of a question registered in said question and answer registration unit, transmits the content of the answer registered in said question and answer registration unit to said avatar statement generation unit without going through said answer information acquisition unit.
9. A communication system as described in claim 8, further comprising a fixed conversation registration unit that registers question contents and answer contents in the question and answer registration unit based on instructions from a user, and the fixed conversation registration unit registers the user's statement contents accepted by the statement acceptance unit and the avatar's statement contents generated by the avatar statement generation unit as question contents and answer contents in the question and answer registration unit in accordance with instructions from the user.
Citation Information
Patent Citations
Avatar creation user interface
JP2023085356A
Apparatus and method for emotional content services on telecommunication devices, apparatus and method for emotion recognition therefor, and apparatus and method for generating and matching the emotional content using same
WO2013027893A1