Voice metaverse chatbot system using ai-based sts
The voice metaverse chatbot system uses AI-based STS to learn and synthesize voice models, enabling dynamic voice interactions and phone calls with deceased individuals, enhancing emotional connections and memory stimulation.
Patent Information
- Application Number
- PCT/KR2024/005404
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-04-22
- Publication Date
- 2025-10-30
AI Technical Summary
Existing AI-based conversation systems are limited in their ability to provide realistic and dynamic voice interactions, particularly in scenarios where communicating with deceased or unavailable individuals, and lack the capability to make phone calls using learned voice models.
A voice metaverse chatbot system using AI-based STS that includes a user terminal with a Remember Phone application and a chatbot server, which learns a voice model from subject information, converts user voice to text, generates a response text, and synthesizes it into the subject's voice tone for phone calls.
Enables realistic voice interactions and phone calls with deceased or unavailable individuals, providing comfort and stimulating memories through personalized voice conversations.
Smart Images

Figure KR2024005404_30102025_PF_FP_ABST
Abstract
Description
A voice metaverse chatbot system using AI-based STS
[0001] The present invention relates to a voice metaverse chatbot system using an AI-based STS, which learns a voice model through a subject's voice information, and can make a phone call to a subject of the voice model by calling a specific phone number assigned to the subject based on the learned voice model.
[0002] Artificial Intelligence (AI) systems are computer systems that achieve human-level intelligence. Unlike existing rule-based smart systems, AI systems are machines that learn, make decisions, and become smarter on their own. As AI systems become more used, their recognition rates improve and their ability to more accurately understand user preferences increases. As such, existing rule-based smart systems are gradually being replaced by deep learning-based AI systems.
[0003] The existing AI-based conversation service provision technology was a task-based conversation processing technology that recognized the user's needs from the task to be serviced (e.g., phone calls, message operations, weather search, route search, schedule management, etc.) and provided the user with a preset answer corresponding to the need.
[0004] To overcome the limitations of these task-based conversation processing technologies, chatbot technology has been researched and developed. However, chatbot-based conversation processing technologies rely on rules, patterns, and example matching, and suffer from the limitation of producing identical and repetitive responses regardless of the conversation.
[0005] Meanwhile, the proliferation of mobile devices has made it possible to converse with anyone, anytime, anywhere, via phone. However, when the person you're trying to talk to is a deceased person or a celebrity, phone conversations are often impossible.
[0006] In particular, when you want to hear the voice of a deceased person, such as a relative, lover, or friend, you have no choice but to rely on pre-recorded or videotaped data.
[0007] Accordingly, as a technology that learns and provides the speech and writing style of an actual person, a conversation method and system that imitates the speech and writing style of an actual person is disclosed in Patent Publication No. 10-2441456.
[0008] The above-described conversation method of the technology provides an artificial intelligence conversation method including a step of creating a conversation model composed of the speech and writing style of a target person based on data related to the target person; and a step of providing a response corresponding to a request from a client using the conversation model.
[0009] However, the above technology has limitations in use because it is a conversation format using a chat room by connecting to the system and must be connected to the system to receive the service.
[0010] In addition, a two-way conversation system based on voice recognition and emotion recognition utilizing artificial intelligence and augmented reality technology was disclosed in Patent Publication No. 10-2021-0156145 as a two-way conversation system.
[0011] In addition, Patent Publication No. 10-2021-0117827 discloses a system and method for providing voice services using artificial intelligence.
[0012] The above technology includes a target terminal that collects target voice information including the voice of the target of service; a service providing server that receives the collected voice information of the target of service, learns a speaking model of the target of service, and provides a virtual voice of the target of service to the service user based on the speaking model; and a user terminal that allows the service user to request a voice service from the service providing server and receive the voice service provided by the service providing server.
[0013] However, the above technologies have the disadvantage that the scope of voice conversation through artificial intelligence is limited to question-and-answer format, as they simply recognize the user's (service user's) voice inquiry (or question, etc.) and provide the learned voice of the subject (deceased).
[0014] The present invention was created to solve the problems of the above-mentioned conventional technology, and the problem to be solved by the present invention is to provide a voice metaverse chatbot system using AI-based STS that allows a user to hear the voice of a deceased person by accessing a chatbot system that provides voice and making a phone call using the voice of the deceased person.
[0015] In addition, the purpose is to provide a voice metaverse chatbot system using AI-based STS that can perform voice calls from the chatbot system to the user terminal according to the event.
[0016] In order to solve the above problem, according to an embodiment of the present invention, a voice metaverse chatbot system using an AI-based STS includes: a user terminal having a phone number assigned and a Remember Phone application for providing a chatbot service installed; and a chatbot server that learns a voice model of a subject through the provided voice information of the subject, analyzes a voice transmitted from the user terminal when a call is connected to the user terminal and converts it into a query text, writes a response text for the converted query text, and converts the written response text into a voice through the learned voice model and transmits it to the user terminal, wherein the Remember Phone application stores and manages voice information of the subject, and provides the stored voice information of the subject to the chatbot server upon a user's request.
[0017] Here, the chatbot server includes an ID management unit that stores and manages an identification ID corresponding to the subject; a voice model learning unit that receives voice information about the subject and learns a tone of the subject based on the received voice information to create a voice model of the subject; an STT (speech-to-text) unit that converts a user's voice transmitted from the user terminal into a query text; a TTS (text-to-speech) unit that creates a response text based on the query text converted by the STT unit and converts the created response text into a voice; and a voice synthesis unit that applies the voice converted by the TTS unit to a voice model learned by the voice model learning unit and synthesizes it into a tone of the subject.
[0018] At this time, the Remember Phone application detects scheduler information stored in the user terminal and provides it to the chatbot server, and the chatbot server further includes an event writing unit that writes a greeting text based on the schedule information provided from the user terminal through the Remember Phone application, and the chatbot server may be configured to convert the greeting text written in the event writing unit into voice through the TTS unit, synthesize the voice converted by the TTS unit into the tone of the target through the voice synthesis unit, and transmit the voice synthesized by the voice synthesis unit when a call connection is attempted to the user terminal and a connection is made.
[0019] In addition, the STT unit can analyze emotional information based on the tremor, pitch, and speed of the voice transmitted from the user terminal, and write the query text based on the analyzed emotional information.
[0020] Additionally, the TTS unit may be configured to convert the query text created based on emotional information in the STT unit into a voice corresponding to the query text.
[0021] In addition, the voice synthesis unit is characterized by segmenting the voice converted by the TTS unit into syllables, words, particles, exclamations, and sentences and synthesizing it into the voice tone of the subject.
[0022] According to the present invention, there is an advantage in that a telephone call can be made with an actual subject using a learned voice model of the subject, thereby providing comfort to the subject in the absence of the subject.
[0023] Additionally, there is an advantage in that a phone call with the subject can stimulate the subject's memories and emotions, allowing the subject to be remembered for a long time.
[0024] Additionally, by receiving calls from the target according to the event, there is an advantage in providing a realistic experience not only through sending but also through receiving.
[0025] Figure 1 is a diagram showing the communication configuration of a voice metaverse chatbot system using an AI-based STS according to the present invention.
[0026] FIG. 2 is a drawing showing an interface according to the execution of the Remember Phone application of one embodiment applied to a voice metaverse chatbot system using AI-based STS according to the present invention.
[0027] FIG. 3 is a diagram showing a chatbot server configuration of one embodiment applied to a voice metaverse chatbot system using AI-based STS according to the present invention.
[0028] FIG. 4 is a diagram showing a management table of IDs stored and managed in the ID management section of a chatbot server applied to a voice metaverse chatbot system using an AI-based STS according to the present invention.
[0029] Figure 5 is a flowchart sequentially showing the response of a chatbot server to a voice input into a user terminal in a voice metaverse chatbot system using an AI-based STS according to the present invention.
[0030] FIG. 6 illustrates an example of a schedule stored in a user terminal in a voice metaverse chatbot system using an AI-based STS according to the present invention.
[0031] Figure 7 is a flowchart sequentially showing the process of processing a schedule transmitted to a user terminal in a voice metaverse chatbot system using an AI-based STS according to the present invention in a chatbot server.
[0032] Hereinafter, with reference to the attached drawings, embodiments of the present invention will be described in detail so that those skilled in the art can easily implement the present invention. However, this is not intended to limit the present invention to specific embodiments, and it should be understood that all modifications, equivalents, and alternatives included within the spirit and technical scope of the present invention are included.
[0033] When it is said that a component is "connected" or "connected" to another component, it should be understood that it may be directly connected or connected to that other component, but that there may be other components in between.
[0034] On the other hand, when it is said that a component is "directly connected" or "directly connected" to another component, it should be understood that there are no other components in between.
[0035] The terminology used herein is merely used to describe specific embodiments and is not intended to limit the present invention. The singular expression includes the plural expression unless the context clearly indicates otherwise. In this application, it should be understood that the terms "comprise" or "have" indicate the presence of a feature, number, process, operation, component, part, or combination thereof described in the specification, but do not preclude the presence or addition of one or more other features, numbers, processes, operations, components, parts, or combinations thereof.
[0036] Unless otherwise defined, all terms used herein, including technical or scientific terms, have the same meaning as commonly understood by those of ordinary skill in the art to which this invention pertains. Terms defined in commonly used dictionaries should be interpreted as having a meaning consistent with their meaning within the context of the relevant technology, and unless explicitly defined herein, they shall not be construed in an idealized or overly formal sense.
[0037] The term “MODULE” as used herein means a single unit that processes a specific function or operation, and may refer to hardware, software, or a combination of hardware and software.
[0038] The terms and words used in this specification and claims should not be interpreted as limited to their conventional or dictionary meanings. Based on the principle that an inventor can appropriately define the concept of a term to best explain his or her invention, the terms and words should be interpreted with meanings and concepts that conform to the technical spirit of the present invention. Furthermore, unless otherwise defined, the technical and scientific terms used have meanings commonly understood by those of ordinary skill in the art to which this invention pertains. In the following description and accompanying drawings, descriptions of well-known functions and configurations that may unnecessarily obscure the gist of the present invention are omitted. The drawings introduced below are provided as examples to ensure that the spirit of the present invention can be sufficiently conveyed to those skilled in the art. Accordingly, the present invention is not limited to the drawings presented below and may be embodied in other forms. Furthermore, like reference numerals represent like elements throughout the specification. It should be noted that like reference numerals are used throughout the drawings to represent like elements wherever possible.
[0039] The present invention relates to a voice metaverse chatbot system using an AI-based STS that learns a voice model through voice information of a subject, and can make a phone call to a subject of the voice model by calling a specific phone number assigned to the subject based on the learned voice model.
[0040] In the present invention, the subject refers to a third person who cannot make a voice call, including a deceased person, a missing person, a lover, or a fictional character, and is defined as a person who cannot be met.
[0041] Figure 1 is a diagram showing the communication configuration of a voice metaverse chatbot system using an AI-based STS according to the present invention.
[0042] Referring to the attached drawing 1, the voice metaverse chatbot system using AI-based STS according to the present invention is configured to include a user terminal (100) and a chatbot server (200) connected to the user terminal (100) through a communication network.
[0043] The user terminal (100) is a terminal for receiving services provided by the chatbot server (200) by connecting to the chatbot server (200) via a communication network. It is sufficient that the user terminal (100) be a device that connects to the chatbot server (200) via a communication network, receives each service provided by the chatbot server (200) upon request, and makes payments for receiving the services. Such a device may be a smartphone, tablet, desktop PC, laptop, etc., as the user terminal (100).
[0044] The user terminal (100) above is installed with a Remember Phone application (110) that facilitates connection to a chatbot server (200), transmits data to the chatbot server (200) while connected to a communication network, outputs data transmitted from the chatbot server (200), calculates charges for services provided by the chatbot server (200), and processes payment according to the calculated charges.
[0045] FIG. 2 is a drawing showing an interface according to the execution of the Remember Phone application of one embodiment applied to a voice metaverse chatbot system using AI-based STS according to the present invention.
[0046] Referring to the attached Figure 2, the interface according to the execution of the Remember Phone application displays various icons such as user information (111), payment information (112), upload (113), settings (114), and sending (115).
[0047] In user information (111), the user's name, user's phone number, target's name, relationship between user and target, target's identification number, etc. are entered, and the entered information is stored and managed.
[0048] The above target identification number may be composed of a specific phone number or identification ID.
[0049] Payment information (112) stores and manages information such as payment card (or bank account) information, service usage period, and service usage method according to service usage.
[0050] The above service usage method means that the service is charged according to the period (monthly, yearly, etc.) or number of uses (number of cases).
[0051] Upload (113) is for transmitting the subject's voice information and video information to the chatbot server (200), and can be configured to allow upload of any content that includes the subject's voice information.
[0052] Additionally, the upload (113) may be configured to transmit scheduler information stored by the user in the user terminal (100) to the chatbot server (200).
[0053] The above scheduler information is a schedule detected from data detected from a scheduler application stored in the user terminal (100), and can be configured to be performed with permission by the user's consent.
[0054] Settings (114) store and manage the sender's volume, receiver's volume, reception time, reception area, and transmission cycle.
[0055] The above-mentioned receivable time is configured to set the receivable time zone for phone connections sent from the chatbot server (200) in conjunction with RTS, and is configured to allow the receivable time zone to be set according to the user's operation.
[0056] The above-mentioned receivable area is configured to set a receivable area for a call connection sent from the chatbot server (200) in conjunction with GPS, and is configured to set a receivable location (city, county, district, etc.).
[0057] In the present invention, a chatbot server (200) is configured to transmit a call from the chatbot server (200) to the user terminal (100) based on the scheduler information of the user terminal (100). The user is configured to set a time zone and region in which he or she can receive a call from the chatbot server (200), so that the call can be received only when he or she is within the range of the set time zone and region.
[0058] The call cycle refers to the cycle for connecting a greeting call from the chatbot server (200) to the user terminal (100) at set intervals. For example, if a call cycle is set, the chatbot server (200) attempts to connect a call to the user terminal (100) within the range of the call cycle and attempts to call the user. Here, a call connection refers to a connection from the user terminal (100) to the chatbot server (200).
[0059] Calling (115) detects the target identification number stored in the user information and attempts to connect a call using the detected identification number. The call connection means connecting to the chatbot server (200) through a communication network from the user terminal (100).
[0060] Here, if multiple target identification numbers are registered in the user information, multiple targets are configured to be displayed on the screen according to the selection (click) of the caller (115), and the user can be configured to select a target from among the multiple targets displayed on the screen to establish a call connection.
[0061] Next, we will explain the chatbot server.
[0062] FIG. 3 is a diagram showing a chatbot server configuration of one embodiment applied to a voice metaverse chatbot system using AI-based STS according to the present invention.
[0063] Referring to the attached Figure 3, a chatbot server (200) applied to a voice metaverse chatbot system using AI-based STS is configured to include an ID management unit (210), a voice model learning unit (220), an STT (speech-to-text) unit (230), a TTS (text-to-speech) unit (240), a voice synthesis unit (250), an event creation unit (260), and a transmission unit (270).
[0064] The ID management unit (210) stores and manages the identification ID corresponding to the subject.
[0065] FIG. 4 is a diagram showing a management table of IDs stored and managed in the ID management section of a chatbot server applied to a voice metaverse chatbot system using an AI-based STS according to the present invention.
[0066] Referring to the attached Figure 4, the ID management unit (210) registers and manages multiple user phone numbers, relationships, and associated IDs for one ID.
[0067] For example, if the subject is a deceased person, the child is registered and managed as the subject's user, and if there are multiple children, multiple users are registered and managed for one ID.
[0068] Additionally, a related ID for the above subject is registered and managed.
[0069] The above-mentioned related ID refers to an ID associated with the subject, and in cases where the subject is a spouse, relative, sibling, etc., it refers to a second subject who is mutually related.
[0070] The voice model learning unit (220) receives voice information about the subject, and learns the tone of the subject based on the received voice information to create a voice model of the subject.
[0071] The chatbot server (200) can be configured to receive voice information about the subject in various ways, such as through the web, a communication network, or a network.
[0072] For example, the voice information of the subject can be uploaded to the chatbot server (200) by using a general PC, tablet, etc. connected to a communication network as well as a user terminal (100) by the user's operation, such as a voice file or video file containing the subject's voice information, and the voice model learning unit (220) of the chatbot server (200) analyzes the uploaded voice information to learn the tone of the subject and generates a voice model of the subject.
[0073] At this time, the greater the amount of voice information, the more accurate the subject's voice model can be, but the time required for learning may also increase.
[0074] The above voice model can be created by analyzing the conversation intent and objects from the input voice through natural language processing such as morphological analysis of the input voice information based on preprocessing of the uploaded voice information, morphological analysis, word dictionary creation, slang, dialect, etc.
[0075] Specifically, the above-mentioned voice model can be generated using WaveNet and a convolutional neural network (CNN). Therefore, a voice model with a natural sound can be created, and a single voice model can be used in a conversational model that produces a variety of voices. A voice model utilizing a CNN can identify the characteristics of a speaker among multiple speakers, and it outperforms other voice synthesis technologies and produces more natural-sounding voice synthesis.
[0076] The above natural language processing can utilize BERT (Bi-directional Encoder Representations for Transformers), a natural language processing language model. The BERT model is a model that pre-trains a model with large-scale data (unlabeled data), such as wiki or book data, and then performs transfer learning (transfer data) on data with specific tasks (labeled data). It is a method that designs a general-purpose solution, implements it in a scalable form, and trains it with a large number of machine resources to improve performance.
[0077] The STT (speech-to-text) unit (230) converts the user's voice transmitted from the user terminal (100) into query text.
[0078] That is, when a call connection is attempted from a user terminal (100) to a chatbot server (200) and a connection is established, or when a call connection is attempted from the chatbot server (200) to the user terminal (100) and a connection is established, and the user's voice is input to the chatbot server (200), the STT unit (230) analyzes the user's voice and converts it into text (characters) that can be recognized by the chatbot server (200).
[0079] At this time, the STT unit (230) may be configured to analyze emotional information about the voice transmitted from the user terminal based on the voice's tremor, pitch, and speed, and to write the query text based on the analyzed emotional information.
[0080] The emotional information analyzed above can be organized into emotions including joy, excitement, sadness, fear, anger, disgust, anxiety, guilt, and shame.
[0081] The TTS (text-to-speech) unit (240) creates a response text based on the query text converted by the STT unit (230) and converts the created response text into a voice corresponding to the text.
[0082] Here, the TTS unit (240) may be configured to convert the query text created based on emotional information in the STT unit (230) into a voice corresponding to the text.
[0083] For example, if the emotion of the voice analyzed in the STT unit (230) is determined to be 'joy', the TTS unit (240) creates a response text based on the determined emotion of 'joy' and converts the created response text into voice.
[0084] The voice synthesis unit (250) applies the voice converted by the TTS unit (240) to the voice model learned by the voice model learning unit (220) and synthesizes it into the voice tone of the subject.
[0085] At this time, the voice synthesis unit (250) may be configured to segment the voice converted by the TTS unit (240) into syllables, words, particles, exclamations, and sentences and synthesize them into the voice tone of the subject.
[0086] That is, when the voice analyzed by the STT unit (230) includes emotional information, the query text converted by the STT unit (230) includes particles and exclamations, etc., and the TTS unit (240) converts this into a voice including emotional information, and the voice synthesis unit (250) converts the voice converted by the TTS unit (240) into a voice including emotional information in the process of synthesizing the voice into the subject's voice.
[0087] Figure 5 is a flowchart sequentially showing the response of a chatbot server to a voice input into a user terminal in a voice metaverse chatbot system using an AI-based STS according to the present invention.
[0088] When the user terminal (100) and the chatbot server (200) are connected (connected), if the user inputs a voice into the user terminal (100), the chatbot server (200) receives the voice input into the user terminal (100) and converts the received voice into a query text in the STT unit (230). Thereafter, the TTS unit (240) writes the query text converted by the STT unit (230) as a response text and then converts it into a voice, and the voice synthesis unit (250) applies the voice converted by the STT unit (240) to the learned voice model of the subject and converts it into a voice having the tone of the subject, and the voice converted by the voice synthesis unit (250) is transmitted to the user terminal (100) so that the user can hear it.
[0089] With the above configuration, the user can make a voice call using the tone (voice) of the target person, and maintain compassion for the target person for a long time through the voice call.
[0090] The event writing unit (260) writes a greeting text based on the schedule information provided from the user terminal (100) through the Remember Phone application.
[0091] FIG. 6 illustrates an example of a schedule stored in a user terminal in a voice metaverse chatbot system using an AI-based STS according to the present invention.
[0092] At this time, the schedule information may be content described in a schedule application, a text message application, a reservation schedule application, etc., and may be detected from various text messages written by the user or received by the user terminal (100) according to payment, etc.
[0093] As shown in the attached Figure 6, the event creation unit (260) can create a greeting text such as "My son!, you're going to Busan on the 00th of the 00th month? You're going on a long business trip. Be careful on your way back." based on the schedule information entered as "Business trip to Busan on the 00th of the 00th month" in the scheduler of the user terminal (100). In addition, a greeting text such as "It's ??'s birthday on the 00th of the 00th month! Congratulations." can be created.
[0094] If the caller (270) determines that the greeting text written in the event writing unit (260) is a specific event, it attempts to connect a call to the user terminal (100).
[0095] Here, the chatbot server converts the greeting text written in the event writing unit (260) into voice through the TTS unit (240), synthesizes the voice converted in the TTS unit (240) into the voice tone of the target person through the voice synthesis unit (250), and when a call connection is attempted from the calling unit (270) to the user terminal (100) and a connection is made, the voice synthesized in the voice synthesis unit (250) is transmitted to the user terminal (100).
[0096] Here, the call connection made in the transmitter (270) is made based on the setting information (reception time, reception area, transmission cycle, etc.) stored in the Remember Phone application (110) of the user terminal (100).
[0097] Figure 7 is a flowchart sequentially showing the process of processing a schedule transmitted to a user terminal in a voice metaverse chatbot system using an AI-based STS according to the present invention in a chatbot server.
[0098] When schedule information is transmitted from the user terminal (100) to the chatbot server (200) by the Remember Phone application (110), the event creation unit (260) of the chatbot server (200) detects a schedule from the transmitted schedule information, creates a greeting text for each detected schedule, and the sending unit (270) attempts to connect a call to the user terminal (100). Thereafter, the TTS unit (240) converts the greeting text created by the event creation unit (260) into voice, and the voice synthesis unit (250) applies the voice converted by the STT unit (240) to the learned voice model of the subject and converts it into a voice having the tone of the subject, and when a call is connected to the user terminal (100) according to the call connection of the sending unit (270), the voice converted by the voice synthesis unit (250) is transmitted to the user terminal (100) and is heard by the user.
[0099] In this way, the user and the subject can communicate effectively by appropriately performing the processes of the attached Figures 5 and 7 in a manner such as crossing, alternating, and repeating.
[0100] The present invention provides the advantage of allowing users to conduct phone calls with actual individuals using a learned voice model of the individual, thereby providing comfort in the absence of the individual. Furthermore, the phone calls stimulate memories and emotions in the individual, thereby fostering long-lasting memories. Furthermore, by receiving calls from individuals based on events, the present invention provides the advantage of providing a realistic experience, both through sending and receiving calls.
[0101] Although the present invention has been described above with reference to several embodiments thereof, it will be understood by those skilled in the art that various modifications and changes may be made to the present invention without departing from the spirit and scope of the present invention as set forth in the claims below.
Claims
1. A user terminal with a phone number and the Remember Phone application installed to provide chatbot services; and A chatbot server that learns a voice model of a subject through the voice information of the subject provided, analyzes a voice transmitted from the user terminal when a call is connected to the user terminal, converts it into a query text, writes a response text for the converted query text, and converts the written response text into voice through the learned voice model and transmits it to the user terminal; Including, The above Remember Phone application is, A voice metaverse chatbot system using an AI-based STS, characterized in that it stores and manages voice information about the subject and provides the stored voice information of the subject to the chatbot server according to a user's request.
2. In claim 1, The above chatbot server, An ID management unit that stores and manages an identification ID corresponding to the above target; A voice model learning unit that receives voice information about the subject and learns the tone of the subject based on the received voice information to create a voice model of the subject; A speech-to-text (STT) unit that converts a user's voice transmitted from the user terminal into a query text; A TTS (text-to-speech) unit that creates a response text based on the query text converted by the STT unit and converts the created response text into voice; and A voice synthesis unit that applies the voice converted by the TTS unit to the voice model learned by the voice model learning unit and synthesizes it into the voice tone of the subject; A voice metaverse chatbot system using AI-based STS, characterized by including:
3. In claim 2, The above Remember Phone application is, Detects scheduler information stored in the user terminal and provides it to the chatbot server, In the above chatbot server, An event writing unit that writes a greeting text based on schedule information provided from the user terminal through the Remember Phone application; Including more, The above chatbot server, A voice metaverse chatbot system using AI-based STS, characterized in that the text of greeting written in the event writing section is converted into voice through the TTS section, the voice converted in the TTS section is synthesized into the voice tone of the subject through the voice synthesis section, and when a call connection is attempted to the user terminal and a connection is made, the voice synthesized by the voice synthesis section is transmitted.
4. In claim 2, A voice metaverse chatbot system using AI-based STS, characterized in that the STT unit analyzes emotional information based on the tremors, pitch, and speed of the voice transmitted from the user terminal, and writes the query text based on the analyzed emotional information.
5. In claim 4, A voice metaverse chatbot system using AI-based STS, characterized in that the TTS unit converts the query text into a voice corresponding to the query text based on emotional information in the STT unit.
6. In claim 1, The above voice synthesis unit, A voice metaverse chatbot system using AI-based STS, characterized in that the voice converted by the above TTS unit is segmented into syllables, words, particles, exclamations, and sentences and synthesized into the voice tone of the subject.
Citation Information
Patent Citations
System and method for virtually communicating in on-line
KR1020110033017A
System and method for providing conversation service with decedent
KR1020120002007A
Handler for testing electronic components and method of photographing electronic components therein
KR1020230030767A
Hinge device of console armrest for car
KR102419962B1
Clothing Dryer
KR102830938B1