Method, system for providing personalized speech dialogue service and trading system

TW202632643AActive Publication Date: 2026-08-01PLAYSEE INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
TW · TW
Patent Type
Applications
Current Assignee / Owner
PLAYSEE INC
Filing Date
2025-01-23
Publication Date
2026-08-01

AI Technical Summary

Technical Problem

Current natural language chatbots like ChatGPT lack the ability to provide personalized voice dialogue services that are relevant to the user's current situation, failing to account for individual needs, preferences, and real-time emotions.

Method used

A system and method that utilizes a cloud server connected to a user device to analyze real-time voice and image data, employing natural language processing and machine learning to generate personalized voice dialogues based on user emotions, environmental information, and user profiles, using generative AI to create lifelike chatbots that adapt to user needs.

Benefits of technology

Enables personalized voice dialogue services that respond to users' real-time emotions and environmental factors, providing tailored interactions through lifelike chatbots that learn and adapt over time, enhancing user engagement and relevance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure TWG2TA001069787_001
    Figure TWG2TA001069787_001
  • Figure TWG2TA001069787_002
    Figure TWG2TA001069787_002
  • Figure TWG2TA001069787_003
    Figure TWG2TA001069787_003
Patent Text Reader

Abstract

A method, a system for providing personalized speech dialogue service and a trading system are provided. A cloud server is provided for connecting with user devices, providing users a personalized speech dialogue service via a dialogue interface, performing the method by processors. The cloud server is configured to receive speech data generated by a user’s real-time voice, or also receive image data, and then extract current speech features and image features. The speech features and the image features are referred to for learning the user’s current mood. A natural language processing model is used to identify semantic meaning in the speech features. A speech dialogue based on the user’s current mood can be generated according to the semantic meaning and current environmental information. A simulated object displayed on a user device can speak the speech dialogue.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The disclosure concerns personalized intelligent voice dialogue technology, specifically a method and system that uses computer technology to provide users with personalized voice dialogue services, as well as a transaction system that provides personalized model creation. Prior Technology

[0002] Among the rapidly developing artificial intelligence (AI) technologies across various fields, one type is the natural language chatbot, which can process natural language and automatically generate content. Examples include ChatGPT (Chat Generative Pre-trained Transformer), developed by OpenAI. These chatbots utilize generative AI technology, allowing them to be trained on large amounts of data and then generate new data related to the original data. This new data is then processed through deep learning (such as generative adversarial networks, GANs) to build an intelligent model.

[0003] Take ChatGPT as an example. ChatGPT is trained by learning from a large amount of online information and can respond to users in natural language. However, the content that it usually responds to users is a standard answer derived from learning, and it cannot respond in real time to provide answers that are relevant to the user's current situation. Although it is a natural language chatbot, it lacks content that is relevant to the user and consistent with the real situation. In other words, the dialogue service provided by current chatbots is general and does not provide personalized voice dialogue services (such as personal needs, preferences, background, etc.). Summary of the Invention

[0004] In order to provide voice dialogue models that can be tailored to individual needs and learning preferences, the disclosure proposes a method and system for providing personalized voice dialogue services, as well as a transaction system for creating personalized models.

[0005] According to a system embodiment of a method for providing personalized voice dialogue services to users, the system includes a cloud server, the cloud server is connected to a user device, the user device obtains the user's voice through a human-machine interface, or adds captured images, provides personalized voice dialogue services to users through a dialogue interface, and the method for providing personalized voice dialogue services to users is executed by one or more processors.

[0006] In the method embodiment, by using a user device to receive voice data generated by a user's real-time voice and obtain the user's current voice characteristics, the user's current emotion can be learned based on the voice characteristics. Then, a natural language processing model is used to identify the semantics in the voice characteristics. Based on the semantics and the current environmental information, a voice dialogue in response to the user's real-time voice is generated. Afterwards, the voice dialogue can be issued through a virtual object displayed on the user device's screen.

[0007] By repeating the above steps, the cloud server receives the voice data generated by the user through the user's device, obtains the voice characteristics, and learns the user's current emotions, thereby continuously generating personalized voice dialogues to achieve the goal of providing personalized voice dialogue services.

[0008] Furthermore, while receiving voice data generated by the user's real-time voice, the system also receives real-time image data generated by the user's device and obtains the user's current image features, so that the system can learn the user's current emotions based on both voice and image features.

[0009] Furthermore, by querying an emotion database, the audio, speech rate, and word choice of the corresponding user's current emotion can be obtained, and then a personalized voice dialogue can be generated.

[0010] Furthermore, deep learning can be performed on the cloud server to learn the user's preferences and personality based on the voice and image data continuously received from the user's device in the past and present, in order to build a user profile.

[0011] Furthermore, the current environmental information obtained by the cloud server includes geographic information and time received from the user's device, and it can also connect to external servers to obtain weather information based on geographic information.

[0012] Furthermore, on the cloud server, after obtaining the user's current voice characteristics and semantics, and obtaining the user's feature file from the database based on the user's recognition information, a corresponding voice dialogue can be generated based on the obtained current semantics, user feature file, and / or current environmental information.

[0013] Furthermore, when running a method to provide personalized voice dialogue services to users on a cloud server, the method can obtain realistic objects generated by artificial intelligence simulation models from the database, transmit model data and image signals to the user's device, display the realistic objects on the user's device's display, and simulate voice dialogue.

[0014] The disclosure outlines a trading system that includes a computer-based intelligent model trading platform offering multiple simulated object options through an interactive interface. It also includes a model database storing simulated objects generated by AI models, and a user database storing and updating user data over time. The trading system includes a web server that allows users to customize simulated objects via a web interface.

[0015] The transaction system receives voice data generated by each user from the user's device through the computer system, obtains the user's current voice characteristics, and then performs deep learning to learn the user's preferences and personality based on the voice data continuously received in the past and present, so as to establish a user profile.

[0016] Furthermore, the transaction system can receive design parameters through a web interface, and then create the personalized virtual object based on the user profile and the design parameters.

[0017] In addition to receiving voice data generated by each user through the computer system, the transaction system can also receive image data of users and build user profiles after learning voice and image features.

[0018] To further understand the features and technical content of the present invention, please refer to the following detailed description and drawings of the present invention. However, the drawings provided are for reference and illustration only and are not intended to limit the present invention. Simple Explanation of the Diagram

[0019] Figure 1 shows a schematic diagram of a scenario that provides users with personalized voice dialogue services;

[0020] Figure 2 shows a schematic diagram of a system architecture embodiment that provides personalized voice dialogue services to users;

[0021] Figure 3 shows an embodiment of a method for providing personalized voice dialogue services to users;

[0022] Figure 4 shows one of the flowcharts of an embodiment of a method for providing personalized voice dialogue services to users;

[0023] Figure 5 shows a second embodiment flowchart of a method for providing personalized voice dialogue services to users; and

[0024] Figure 6 shows an embodiment of the trading system. Implementation

[0025] The following specific embodiments illustrate the implementation of the present invention. Those skilled in the art can understand the advantages and effects of the present invention from the content disclosed in this specification. The present invention can be implemented or applied through other different specific embodiments, and various details in this specification can also be modified and changed based on different viewpoints and applications without departing from the concept of the present invention. Furthermore, the accompanying drawings of the present invention are for simple illustrative purposes only and are not depictions of actual dimensions; this is stated in advance. The following embodiments will further describe the relevant technical content of the present invention in detail, but the disclosed content is not intended to limit the scope of protection of the present invention.

[0026] It should be understood that while terms such as "first," "second," and "third" may be used in this document to describe various components or signals, these components or signals should not be limited by these terms. These terms are primarily used to distinguish one component from another, or one signal from another. Furthermore, the term "or" as used herein should, as appropriate, include any combination of one or more of the related listed items.

[0027] In order to provide users with individual voice dialogue services and to enable natural language dialogue using lifelike chatbots, the disclosure proposes a method and system for providing personalized voice dialogue services, as well as a transaction system for creating personalized models.

[0028] The scenario for providing personalized voice dialogue services can be seen in the schematic diagram shown in Figure 1. The system providing personalized voice dialogue services is equipped with a cloud server 20, which enables users 10 to engage in natural language dialogue with a lifelike chatbot 101 generated by generative artificial intelligence (AI) through learning knowledge, human speech, behavior, and facial expressions. Notably, the intelligent model running behind the lifelike chatbot 101 learns the pitch and speed of the user 10's current conversation, allowing the chatbot 101 to respond with corresponding audio and speech speed, reflecting the user 10's current emotions.

[0029] For example, when the intelligent model analyzes the changes in tone, rhythm and volume in the voice data and determines that the user 10 is currently sad, the lifelike chatbot 101 will adjust the audio and speaking speed to reflect the user 10's current sad emotion.

[0030] The diagram illustrates user 10 operating an application on user device 100. For example, after opening a chat interface through the application and connecting to cloud server 20, user 100 can select to load a lifelike chatbot 101 from cloud server 20, including relevant model data and image signals of lifelike chatbot 101, just like two people video chatting in the real world. The diagram shows a chat interface starting on the display of user device 100, which displays two user images 103 and lifelike chatbot 101 that are having a conversation.

[0031] When providing personalized dialogue services, user device 100 connects to cloud server 20, activates the audio (microphone) and video (camera) functions on user device 100, and performs real-time audio recording and video capture. Voice and video data can be uploaded to cloud server 20 in real time. The language and image models running on cloud server 20 extract voice features, or add image features, and may also add current environmental information. Using an intelligent model, the system determines the user's current emotion and generates a voice dialogue. The criteria for determining emotion include analyzing the pitch, rhythm, and intensity changes of each segment of voice data, as well as semantic meaning, through an intelligent model.

[0032] It is worth mentioning that, when generating voice dialogue using a language model, according to one embodiment, the relevant audio data is transmitted to the user device 100 and can be displayed on the dialogue interface after being transcribed into text. Furthermore, in addition to the voice dialogue being emitted directly from the speaker of the user device 100, the voice dialogue can also be simulated by a lifelike chatbot 101 displayed on the dialogue interface. The artificial intelligence running there can generate facial expressions and emoticons to match the current emotions. For example, the mouth movements of the generated lifelike character can simulate the mouth movements of a human speaking, and the facial muscle movements can also match the emotions expressed in the current voice.

[0033] The cloud server 20 implements a system that provides personalized voice dialogue services to users. It mainly uses hardware and software collaboration to run various artificial intelligence algorithms and models, as shown in the schematic diagram of the system architecture embodiment in Figure 2.

[0034] The cloud server 20 utilizes the processing circuitry, memory, and software of a computer system to implement various functional modules, such as a natural language processing module 201 for processing the content of dialogues with the user; an instruction processing module 203 for providing corresponding services based on the user's requests generated through the dialogue interface; and a user interface module 205 for generating the interface for dialogues with the user and interfacing with the user interface generated by a specific application running on the user's device. The cloud server 20 also includes a machine learning module 207 for training a natural language model that meets the user's needs; and a database module 209 for providing content from built-in or external databases.

[0035] The cloud server 20 has a built-in or external database 22, which mainly contains user data 221 and multiple simulated objects generated by the artificial intelligence simulation model 223. It can also include an emotion database 225 derived from big data analysis. The emotion database 225 stores data such as audio, speech rate, and vocabulary required for speech under various emotions, allowing simulated objects to query and generate dialogues corresponding to specific emotions. The cloud server 20 provides personalized voice dialogue services via the network 200. Users operate the user device 100 on their terminals and use the software running on it to access the services of the cloud server 20 via the network 200.

[0036] The database module 209 in the cloud server 20 is used to access and manage the built-in or external database 22, including maintaining user data 221 generated by users using personalized voice dialogue services, and also provides simulated objects generated by artificial intelligence simulation models 223. The simulated objects can be humanoid, which may include various simulated humanoid chatbots from different fields provided by multiple users, as well as various simulated humanoid chatbots provided by the system.

[0037] According to the embodiment, the database 22 of the cloud server 20 provides realistic objects generated by the artificial intelligence simulation model 223, which can be used to generate multiple realistic humanoid chatbots (excluding realistic humans, animals, or various objects). In addition to the realistic humanoids having different appearances and accents, they can also be designed as various realistic humanoid chatbots with different professional knowledge, so as to provide users with personalized voice dialogue services according to their needs. It is worth mentioning that the method of providing users with personalized voice dialogue services proposed in the disclosure can use the realistic objects generated by the artificial intelligence simulation model 223 to perform natural language dialogue. It uses a generative artificial intelligence, which uses machine learning algorithms to learn from a large amount of human data to build a realistic human character generation model, and then applies generative artificial intelligence based on 3D graphics technology to generate images, which can generate specific realistic objects according to user needs (providing prompts).

[0038] Furthermore, to give the lifelike chatbot a professional background, this part involves using machine learning algorithms to build a domain-specific intelligent model by learning data from specific domains. For example, a Retrieval Augmented Generation (RAG) technique can be used to implement a natural language model for a specific domain based on a large language model (LLM).

[0039] The instruction processing module 203 in the cloud server 20 is used to process various requests transmitted from the user device 100, such as requests to perform a dialogue, requests to select a lifelike chatbot to perform a personalized voice dialogue service, and various requests generated during the voice dialogue.

[0040] The cloud server 20 transmits various messages, such as audio and video content, text content and image files, to the user device 100 through the user interface module 205 (including necessary encoding, decoding, compression and decompression processes). It also processes the dialogue content displayed by two or more parties through the dialog interface and uses it to generate various functions and graphics in the dialog interface.

[0041] In the cloud server 20, the machine learning module 207 runs machine learning algorithms, using neural network deep learning technology to learn the speech features of the user. It trains a large amount of speech data to create a speech model and establishes a natural language processing (NLP) model to obtain speech features from the speech data, derive semantics, and generate corresponding dialogue content. Furthermore, the machine learning module 207 can also run machine learning algorithms to learn image features. For example, it can build an image model that can learn user expressions and emotions by training on a large number of user facial image features. Thus, in addition to judging the user's emotions based on speech features, the cloud server 20 can also use the image model to process image features to assist in judging the user's emotions.

[0042] The cloud server 20 uses the natural language processing module 201 to process the voice data obtained from the user device 100. In addition to textualizing the voice content, it also runs the natural language processing model to obtain the voice features in the voice data, recognize the semantics therein, and then generate dialogue content to respond to the user in natural language.

[0043] Furthermore, in the process of obtaining semantics and judging user emotions, the speech features used include changes in pitch, rhythm, and volume. Based on these speech features (such as sound wave waveform features), the speech model can identify the user's current emotions and can also learn the individual user's personality and preferences through long-term learning, thus establishing a user profile.

[0044] By using a computer system to run machine learning algorithms to learn human language and classify it, and by deeply learning language structure and the relationships between sentences, a natural language processing model is generated. In the method disclosed, the natural language processing model can be trained using speech data generated from conversations between users to form a personalized language model. Furthermore, after training the natural language processing model to generate a personalized language model, deep learning algorithms can be used to learn features such as pitch, rhythm, and intensity changes in speech, thereby learning emotions in language.

[0045] It is worth mentioning that the natural language processing module 201 does not simply utilize existing natural language search / dialogue services such as ChatGPT, which employs a large language model (LLM). This is because existing ChatGPT can only respond to user dialogue in a one-way manner and fails to learn the user's emotions or reference environmental information. Therefore, the natural language processing module 201 further utilizes user profiles, current semantics, and real-time environmental information (such as current geographical location, weather, and time) to enable the lifelike chatbot to engage in dialogue based on the user's current emotions, generating dialogues that meet the user's immediate needs. According to an embodiment, the lifelike chatbot can query the emotion database 225 to retrieve data such as audio, speech rate, and vocabulary corresponding to various emotions, in order to issue dialogues corresponding to specific emotions.

[0046] For example, when it is noon, the system obtains the user's geographical location (such as the user's device location information obtained through the application) and weather (cold, hot, sunny or rainy) through environmental information. The lifelike chatbot can proactively suggest suitable restaurants based on the user's preferences. The user can respond whether they want to eat, what their acceptable price range is, or other needs, so that the lifelike chatbot can adjust its suggestions in real time or conduct a dialogue to confirm the needs.

[0047] According to an embodiment, in the machine learning algorithm that determines the current emotion using the image features of the user's face, a supervised machine learning method can be employed to learn the changes in image features labeled manually under different emotions, such as changes in the distance ratio between facial features or specific parts, line changes, shape changes, and color changes. Based on this, an image model for judging emotions based on image features is trained. Thus, in the cloud server 20, both the speech model and the image model can be used simultaneously to determine the user's emotion. The natural language processing module 201 then obtains the semantics and generates dialogue content, where the audio, speech rate, and word choice of the dialogue can be obtained by querying the emotion database 225.

[0048] Based on the above system architecture and implementation method for providing personalized voice dialogue services to users, Figure 3 then shows an embodiment of the method for providing personalized voice dialogue services to users.

[0049] This example demonstrates that the user's device uses the human-machine interface 303 to convert the user's voice / image 301 into text content displayed on the dialogue interface 305 or to issue a voice message, which is then received by the cloud server. The user can also use an application running on their device to request a personalized voice dialogue service from the cloud server, which then provides the personalized voice dialogue service to the user through the dialogue interface 305.

[0050] A cloud server executes a method for providing personalized voice dialogue services to users through one or more processors of its computer system, acquiring user voice / image 301 generated from the user device through a dialogue interface 305. According to an embodiment, the system utilizes various sensors in the user device to acquire the user's voice and image data. For example, after the user device initiates the dialogue interface 305, it simultaneously activates the microphone to receive the user's voice, activates the camera to capture an image of the user's face, and generates speech (or inputs text) through the human-machine interface 303, allowing interaction with realistic objects displayed in the dialogue interface 305 and dialogue in natural language.

[0051] Voice and video data are transmitted from the user's device to the cloud server. The cloud server runs a voice / video model 307, which can be divided into a voice model and a video model. For voice data, the voice model transcribes the voice data to obtain its semantics and determine the user's current emotion. For video data, the video model analyzes the video data to obtain the user's facial features. For example, it can analyze the color and lighting changes of the pixels corresponding to the user's facial organs in the image to obtain changes in the images of facial features or specific parts such as the mouth and eyes. If the pixels of each part reflect the geometric relationship between facial features or specific parts, combined with semantics and voice features, the user's current emotion can be determined.

[0052] Furthermore, the voice / image model 307 can also obtain current environmental information 313 from an external database (or server). According to the embodiment, after the cloud server connects to the user's device, it can obtain the location information in the user's device and determine the user's geographical location. Therefore, it can obtain environmental information 313 based on the user's location and time from an external database, such as local weather, surrounding store business information, geographical environment, and holidays, to facilitate providing the user with effective content. When the cloud server runs the voice / image model 307, it also queries the user database 315 to obtain the user's feature file, from which it can obtain the user's basic information, learned preferences, historical records, etc. Among them, the cloud server also stores and updates each user's data according to the time dimension by using the user database 315, recording each user's historical dialogue records.

[0053] In this way, the cloud server can effectively provide users with personalized voice dialogue services through the voice / image model 307. Furthermore, when the user conducts a personalized voice dialogue, the cloud server continuously acquires the user's voice / image features 311 and continuously learns the voice and image data generated in the user's dialogue through the feature learning algorithm 309 to update or optimize the voice / image model 307.

[0054] Based on the above system architecture and system operation embodiments, Figure 4 then shows an embodiment flowchart of a method for providing personalized voice dialogue services to users.

[0055] The method runs on a cloud server. The user operates an application running on their device to establish a connection with the cloud server and initiates a dialogue interface through the application. The user device can acquire real-time voice data generated by the user's speech through sensing elements such as a microphone (or with the addition of a camera), or add image data (step S401). Then, it learns the voice features, which may include changes in tone, rhythm, and volume of the voice transmitted by the user through the user device (step S403). Thus, the meaning can be identified based on the voice features. When receiving the real-time voice data generated by the user, according to one embodiment, real-time image data of the user is also received, and the user's current image features are obtained. Therefore, the user's current emotion can be learned simultaneously based on the voice features (including meaning) and image features (step S405).

[0056] According to the embodiment, while obtaining voice features, image features can also be obtained based on image data. Then, the formed voice and / or image features 41 are transmitted to the image generation model 43 of the cloud server to generate realistic images, such as the facial expressions, mouth and image changes of specific parts of the humanoid chatbot. They are also transmitted to the personalized natural language model 45 to generate voice dialogue with emotions derived from features 41.

[0057] In step S405, which determines the user's current emotion, the cloud server uses a natural language processing model to identify the semantics in the speech features. Based on the semantics and current environmental information, it generates a voice dialogue based on the user's current emotion, which is also a voice dialogue that responds to the user's real-time voice.

[0058] Furthermore, the cloud server uses image generation model 43 to generate images of the simulated object based on voice features or by adding image features (step S407), and displays the simulated object in the dialogue interface on the user's device display (step S409). A personalized natural language model 45 then generates dialogue content based on the user's semantics, current emotion, or environmental information to conduct a voice dialogue with the user (step S411). For example, the simulated object could be a lifelike chatbot, where the image of the simulated object primarily reflects facial expressions, mouth, and changes in specific body parts that reflect the current emotion.

[0059] Thus, by repeating the above steps, such as receiving voice data generated by the user through the user's device, or adding image data to obtain voice features, or adding image features to learn the user's current emotions, personalized voice dialogues are continuously generated.

[0060] Figure 5 then shows a flowchart of another embodiment of a method for providing personalized voice dialogue services to users.

[0061] According to one embodiment, the cloud server can generate an interactive interface through a web server program, providing users with the option to open multiple lifelike chatbots after connecting their user devices to the cloud server. Users can then select one of the lifelike chatbots (step S501) and start a voice conversation (step S503). When the user has a voice conversation with the lifelike chatbot through the conversation interface, in addition to activating the microphone to record sound, the camera can also be activated to record the user and have a video conversation with the lifelike chatbot. At the same time, the cloud server will obtain real-time voice and video data (step S505).

[0062] According to an embodiment, in a cloud server, a transformation model is used to learn the data correlations of voice data (or image data combined) and extract voice features, or add image features (step S507). The cloud server uses neural network deep learning technology, such as a transformation model employing an attention mechanism, to learn the data correlations based on the features of voice and image data. Specifically, it can learn the user's preferences and personality based on the voice and image data continuously received from the user's device in the past and present, and build a user profile describing the user's preferences and personality. Depending on the application, the neural network deep learning technology used by the cloud server can learn about the user's friends, family, and occupation through the user's online activity data, making the user profile more complete and providing voice dialogue services that are closer to the user's needs, especially emotional needs.

[0063] Thus, the software running on the cloud server can continuously extract features from voice and video data to build or update personalized profiles. Furthermore, after obtaining semantic information, it determines emotions based on semantics and user profiles (step S509). Additionally, the cloud server can obtain current environmental information, including geographic information and time received from the user's device. After connecting to an external server, it can obtain real-time weather information based on geographic information (step S511).

[0064] Having obtained the user's current voice characteristics and semantics, and based on the user's recognition information, a user feature profile is obtained. That is, based on the semantics, the user feature profile, and / or the current environmental information, a corresponding voice dialogue is generated (step S513). The process then repeats steps S503 to S513, repeatedly receiving voice data generated by the user through the user's device, obtaining voice characteristics, learning the user's current emotions, and continuously generating personalized voice dialogues.

[0065] According to another embodiment, based on the above system and method, a trading system as shown in Figure 6 is also proposed.

[0066] The trading system includes an intelligent model trading platform 600 built on a computer system. The intelligent model trading platform 600 has a model database 601, a user database 603, and a web server 605. The user database 603 is used to store user data that is stored and updated according to the time dimension. The intelligent model trading platform 600 provides users with customized personalized lifelike chatbots through a web interface initiated by the web server 605. On the other hand, it also allows users to remotely select (or obtain through trading) lifelike chatbots and import them into the user's device to execute voice dialogue applications.

[0067] The intelligent model trading platform 600 provides multiple lifelike chatbot options through an interactive interface. These lifelike chatbots, generated by artificial intelligence models, can be stored in the model database 601. Users connect to the intelligent model trading platform 600 via network 60 through various user devices 61, 62, 63, etc., and can select one of the lifelike chatbots from the options displayed on the interactive interface to perform voice conversations.

[0068] According to an embodiment, the system, implemented through a computer system, allows users to share personalized, lifelike chatbots customized by the system with other users. These customized lifelike chatbots may possess expertise in a specific field or be uniquely designed lifelike chatbots that can be listed and sold on the intelligent model trading platform 600.

[0069] In one embodiment of customizing a lifelike chatbot, the transaction system receives voice data generated by each user from the user's device via a computer system, and obtains the user's current voice characteristics. It then performs deep learning to learn the user's preferences and personality based on continuously received voice data (or including image data) from the past and present. Similarly, a user profile can be created based on voice characteristics (or image characteristics). In this way, the user can also design the appearance of a personalized lifelike chatbot. When the intelligent model transaction platform 600 receives design parameters through a web interface, it creates a personalized lifelike chatbot based on the user profile and design parameters, and stores it in the model database 601.

[0070] In this way, the user can connect to the intelligent model trading platform 600 through the user device to obtain the image signal and model data of one of the lifelike chatbots from the model database 601 in the trading system. The lifelike chatbot can be displayed on the user device's screen and a voice conversation can begin.

[0071] In summary, according to the methods, systems, and transaction systems for providing personalized voice dialogue services described in the above embodiments, end users can learn their personal data through the hardware and software collaboration functions implemented in the cloud server. Using machine learning algorithms and natural language models, an intelligent model capable of providing personalized voice dialogue services is established, and a chatbot provides natural language services. Furthermore, it can determine the user's current mood, and, in conjunction with environmental information and user preferences, provide personalized voice dialogue services based on the user's current semantics. The personalized, lifelike chatbot trained by the user on the cloud server can also be shared with others through an intelligent model trading platform implemented in the transaction system.

[0072] The content disclosed above is only a preferred and feasible embodiment of the present invention, and is not intended to limit the scope of the patent application of the present invention. Therefore, all equivalent technical changes made using the contents of the present invention specification and drawings are included in the scope of the patent application of the present invention.

[0073] 10: Users 100: User Device 101: Lifelike Chatbot 103: User Images 20: Cloud Server 201: Natural Language Processing Module 203: Instruction Processing Module 205: User Interface Module 207: Machine Learning Module 209: Database Module 200: Network 22: Database 221: User Information 223: Artificial Intelligence Simulation Model 225: Emotion Database 301: User's voice / video 303: Human-Computer Interface 305: Dialogue Interface 307: Voice / Image Model 309: Feature Learning Algorithm 311: User's voice / image characteristics 313: Environmental Information 315: User Database 41: Features 43: Image Generation Model 45: Personalized Natural Language Models 61, 62, 63: User devices 60: Internet 600: Intelligent Model Trading Platform 601: Model Database 603: User Database 605: Web server Steps S401~S411: Process for providing personalized voice dialogue services to users Steps S501~S513: Process for providing users with personalized voice dialogue services

Claims

1. A method for providing personalized voice dialogue services to users, running on a cloud server, comprising: The system receives voice data generated by a user's real-time voice through a user device and obtains the user's current voice characteristics. The system learns the user's current emotion based on the voice features; it uses a natural language processing model to identify the semantics in the voice features, and generates a voice dialogue based on the user's current emotion according to the semantics and the current environmental information. The current environmental information obtained by the cloud server includes geographical information and time received from the user's device, and weather information obtained from the geographical information by connecting to an external server; and it issues the voice dialogue through a virtual object displayed on a screen of the user's device; it continuously generates personalized voice dialogues by repeatedly receiving the voice data generated by the user through the user's device, obtaining the voice features, and learning the user's current emotion; and it performs deep learning to learn the user's preferences and personality based on the continuously received voice data and image data to establish a user profile.

2. The method for providing a personalized voice dialogue service to a user as described in claim 1, wherein the voice features include pitch, rhythm, and intensity variations in the voice transmitted by the user through the user device.

3. The method for providing personalized voice dialogue services to users as described in claim 2, wherein, When receiving the voice data generated by the user's real-time voice, the system also receives the user's real-time image data and obtains the user's current image features. At the same time, the system learns the user's current emotions based on the voice features and image features.

4. The method for providing a personalized voice dialogue service to a user as described in claim 3, wherein an emotion database is queried to obtain the audio, speech rate and word choice that should correspond to the user's current emotion, and then the personalized voice dialogue is generated.

5. The method for providing personalized voice dialogue services to users as described in claim 1, wherein, In the cloud server, the user's current voice characteristics and semantics are obtained, and the user's feature file is obtained based on the user's recognition information. That is, the voice dialogue is generated according to the semantics, the user's feature file, and / or the current environmental information.

6. A method for providing a personalized voice dialogue service to a user as described in any one of claims 1 to 5, wherein, When the cloud server runs a method to provide users with personalized voice dialogue services, it retrieves the simulated object generated by an artificial intelligence simulation model from a database, transmits image signals and model data to the user device, displays the simulated object on a display of the user device, and simulates voice dialogue.

7. The method for providing a personalized voice dialogue service to a user as described in claim 6, wherein the simulated object is a lifelike chatbot, and an interactive interface is generated by the cloud server through a web server, providing multiple options of the lifelike chatbot.

8. A system for providing personalized voice dialogue services to users, comprising: A cloud server connects to a user device, providing a personalized voice dialogue service to the user through a dialogue interface, and a method for providing a personalized voice dialogue service to the user, executed by one or more processors, includes: receiving voice data generated by a user's real-time voice through the user device and obtaining the user's current voice features; learning the user's current emotion based on the voice features; recognizing the semantics in the voice features using a natural language processing model, and generating a voice dialogue based on the user's current emotion according to the semantics and current environmental information, wherein the current environmental information obtained by the cloud server includes geographical information and time received from the user device, and weather information obtained based on the geographical information from an external server; and issuing the voice dialogue through a simulated object displayed on a screen of the user device; wherein the personalized voice dialogue is continuously generated by repeatedly receiving voice data generated by the user through the user device, obtaining the voice features, and learning the user's current emotion; and establishing a user profile by performing deep learning to learn the user's preferences and personality based on the continuously received voice data and image data.

9. A system for providing personalized voice dialogue services to users as described in claim 8, wherein, When receiving the voice data generated by the user's real-time voice, the system also receives the user's real-time image data and obtains the user's current image features. At the same time, the system learns the user's current emotions based on the voice features and image features.

10. The system for providing personalized voice dialogue services to users as described in claim 9, wherein the personalized voice dialogue is generated by querying an emotion database to obtain the audio, speech rate and word choice that should correspond to the user's current emotion.

11. A system for providing personalized voice dialogue services to users as described in claim 8, wherein, In the cloud server, the user's current voice characteristics and semantics are obtained, and the user's feature file is obtained based on the user's recognition information. That is, the voice dialogue is generated according to the semantics, the user's feature file, and / or the current environmental information.

12. A system for providing personalized voice dialogue services to users as described in any one of claims 8 to 11, wherein, When the cloud server runs a method to provide users with personalized voice dialogue services, it obtains the simulated object generated by an artificial intelligence simulation model from a database, transmits the model data and image signals to the user device, displays the simulated object on a display of the user device, and simulates voice dialogue.

13. The system for providing personalized voice dialogue services to users as described in claim 12, wherein the simulated object is a lifelike chatbot, and an interactive interface is generated by the cloud server through a web server, providing multiple options of the lifelike chatbot.

14. A trading system, comprising: An intelligent model trading platform built with a computer system provides multiple options for realistic objects through an interactive interface; A model database stores the multiple realistic objects generated by an artificial intelligence simulation model; A user database stores and updates user data based on a time dimension; and a web server provides users with customized personalized virtual objects through a web interface. The transaction system receives voice data generated by each user from a user device via a computer system, and obtains the user's current voice characteristics. Through repeated reception of the voice data generated by the user through the user device, obtaining the voice characteristics, and learning the user's current emotions, a personalized voice dialogue is continuously generated. Deep learning is then used to learn the user's preferences and personality based on the continuously received voice data from the past and present, establishing a user profile. Upon receiving design parameters through the web interface, the personalized virtual object is created based on the user profile and the design parameters. The computer system uses a natural language processing model to identify the semantics in the voice characteristics, and generates a voice dialogue based on the user's current emotions according to the semantics and current environmental information. The current environmental information includes geographical information and time received from the user device, and weather information obtained from an external server based on the geographical information.

15. The transaction system as claimed in claim 14, wherein the user device obtains the image signal and model data of the simulated object from the model database, displays the simulated object on a display of the user device, and simulates voice dialogue.