Method, system and transaction system for providing personalized voice dialogue services
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- PEIXI TECHNOLOGY CO LTD
- Filing Date
- 2025-02-06
- Publication Date
- 2026-08-07
AI Technical Summary
[0003]以ChatGPT为例,ChatGPT是通过学习大量网络信息训练而成,可以用自然语言方式与用户对答,但是常见响应使用者的内容是通过学习得出的标准答案,并无法实时应变以提供与使用者当下状态有关的答案,虽说是自然语言聊天机器人,但缺乏与使用者关连与符合实时状况的内容,也就是现行聊天机器人提供的对话服务为一般性的,而没有针对个人化(如个人需求、喜好、背景等)的语音对话服务
Smart Images

Figure CN122531368A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to personalized intelligent voice dialogue technology, and in particular to a method and system for providing personalized voice dialogue services to users using computer technology, as well as a trading system for providing personalized model creation. Background Technology
[0002] Among the rapidly developing fields of artificial intelligence (AI), one type is the natural language chatbot, which can process natural language and automatically generate content. Examples include ChatGPT (Chat Generative Pre-trained Transformer), developed by OpenAI. These chatbots utilize generative AI technology, allowing them to be trained on large amounts of data and then generate new data related to the original data. This new data is then used for deep learning (such as generative adversarial networks, GANs) to build an intelligent model.
[0003] Take ChatGPT as an example. ChatGPT is trained by learning from a large amount of network information and can respond to users in natural language. However, the content that usually responds to users is a standard answer derived from learning and cannot adapt in real time to provide answers that are relevant to the user's current state. Although it is a natural language chatbot, it lacks content that is relevant to the user and consistent with the real-time situation. In other words, the dialogue service provided by current chatbots is general and does not provide personalized voice dialogue services (such as personal needs, preferences, background, etc.). Summary of the Invention
[0004] In order to provide a voice dialogue model that can be tailored to individual needs and learning preferences, this disclosure proposes a method and system for providing personalized voice dialogue services, as well as a transaction system for creating personalized models.
[0005] According to a system embodiment of a method for providing personalized voice dialogue services to users, the system includes a cloud server, a user device connected to the cloud server, the user device acquiring the user's voice through a human-machine interface, or adding captured images, providing personalized voice dialogue services to users through a dialogue interface, and the method for providing personalized voice dialogue services to users being executed by one or more processors.
[0006] In the method embodiment, the user device receives voice data generated by a user's real-time voice and obtains the user's current voice features. The user's current emotion can be learned based on the voice features. Then, a natural language processing model is used to identify the semantics in the voice features. Based on the semantics and the current environmental information, a voice dialogue in response to the user's real-time voice is generated. Afterward, the voice dialogue can be initiated through a simulated object displayed on the user device's screen.
[0007] By repeating the above steps, the cloud server receives voice data generated by the user through the user's device, obtains voice features, and learns the user's current emotions, thereby continuously generating personalized voice dialogues to achieve the goal of providing personalized voice dialogue services.
[0008] Furthermore, while receiving voice data generated by the user's real-time voice, the system also receives real-time image data generated by the user's device and obtains the user's current image features, so that the system can learn the user's current emotions based on both voice and image features.
[0009] Furthermore, by querying an emotion database, the audio, speech rate, and word choice of the corresponding user's current emotion can be obtained, and then a personalized voice dialogue can be generated.
[0010] Furthermore, deep learning can be performed on the cloud server to learn the user's preferences and personality based on the voice and image data continuously received from the user's device in the past and present, in order to build a user profile.
[0011] Furthermore, the current environmental information obtained by the cloud server includes geographic information and time received from the user device, and it can also connect to external servers to obtain weather information based on geographic information.
[0012] Furthermore, in the cloud server, after obtaining the user's current speech characteristics and semantics, and retrieving the user feature file from the database based on the user's recognition information, a corresponding voice dialogue can be generated based on the obtained current semantics, user feature file, and / or current environmental information.
[0013] Furthermore, when running a method for providing personalized voice dialogue services to users on a cloud server, a simulated object generated by an artificial intelligence simulation model can be obtained from the database, and the model data and image signals can be transmitted to the user device to display the simulated object on the user device's display screen and simulate a voice dialogue.
[0014] This disclosure proposes a trading system, including an intelligent model trading platform built with a computer system, which provides options for multiple simulated objects through an interactive interface, and includes a model database storing multiple simulated objects generated by an artificial intelligence simulation model, and a user database storing and updating user data according to a time dimension. The trading system includes a web server that provides users with customized simulated objects through a web interface.
[0015] The transaction system receives voice data generated by each user from the user device through the computer system, obtains the user's current voice characteristics, and then performs deep learning to learn the user's preferences and personality based on the voice data continuously received in the past and present, so as to build a user feature file.
[0016] Furthermore, the transaction system can receive design parameters through a web interface, and then create the personalized simulation object based on the user feature file and the design parameters.
[0017] In addition to receiving voice data generated by each user through the computer system, the transaction system can also receive image data from users and build user feature files after learning voice and image features.
[0018] To further understand the features and technical content of the present invention, please refer to the following detailed description and accompanying drawings. However, the drawings provided are for reference and illustration only and are not intended to limit the present invention. Attached Figure Description
[0019] Figure 1 This displays a schematic diagram illustrating a scenario that provides users with personalized voice dialogue services.
[0020] Figure 2 This diagram illustrates a system architecture embodiment that provides personalized voice dialogue services to users.
[0021] Figure 3 An embodiment diagram showing a method for providing personalized voice dialogue services to users is displayed;
[0022] Figure 4 One of the flowcharts illustrating an embodiment of a method for providing personalized voice dialogue services to users is shown;
[0023] Figure 5 A second flowchart illustrating an embodiment of a method for providing personalized voice dialogue services to users; and
[0024] Figure 6 An example diagram of a transaction system is shown. Detailed Implementation
[0025] The following specific embodiments illustrate the implementation of the present invention. Those skilled in the art can understand the advantages and effects of the present invention from the content disclosed in this specification. The present invention can be implemented or applied through other different specific embodiments, and various details in this specification can also be modified and changed based on different viewpoints and applications without departing from the concept of the present invention. Furthermore, the accompanying drawings of the present invention are for simple illustrative purposes only and are not depictions of actual dimensions; this is stated beforehand. The following embodiments will further describe the relevant technical content of the present invention in detail, but the disclosed content is not intended to limit the scope of protection of the present invention.
[0026] It should be understood that while terms such as "first," "second," and "third" may be used in this document to describe various components or signals, these components or signals should not be limited by these terms. These terms are primarily used to distinguish one component from another, or one signal from another. Furthermore, the term "or" as used herein should, as appropriate, include any combination of one or more of the related listed items.
[0027] In order to provide users with individual voice dialogue services and to enable natural language dialogue using a lifelike chatbot, this disclosure proposes a method and system for providing personalized voice dialogue services, as well as a transaction system for providing personalized model creation.
[0028] For scenarios that provide personalized voice dialogue services to users, please refer to Figure 1 The diagram shows that the system providing personalized voice dialogue services includes a cloud server 20. Through the cloud server 20, user 10 can engage in natural language dialogue with a lifelike chatbot 101 generated by generative artificial intelligence (AI) through learning knowledge, human speech, behavior, and facial expressions. Notably, the intelligent model running behind the lifelike chatbot 101 learns the pitch and speed of user 10's current conversation, allowing the chatbot to adjust its audio and speech accordingly, reflecting the user 10's current emotions.
[0029] For example, when the intelligent model analyzes the changes in tone, rhythm and volume in the voice data and determines that the user 10 is currently sad, the lifelike chatbot 101 will adjust the audio and speaking speed to reflect the user 10's current sad emotion.
[0030] The diagram illustrates user 10 operating an application on user device 100. For example, by opening a chat interface through the application and connecting to cloud server 20, user 10 can select to load a lifelike chatbot 101 from cloud server 20, including relevant model data and image signals of the lifelike chatbot 101, just like two people having a video call in the real world. The diagram shows a chat interface starting on the display of user device 100, which displays two user images 103 and the lifelike chatbot 101 having a conversation.
[0031] When providing personalized dialogue services, user device 100 connects to cloud server 20, activates the audio (microphone) and video (camera) functions on user device 100, and performs real-time audio recording and image capture. Voice and image data can be uploaded to cloud server 20 in real time. The language and image models running on cloud server 20 acquire voice features, or add image features, and may also add current environmental information. After using an intelligent model to determine the user's current emotion, a voice dialogue is generated. The basis for determining emotion includes analyzing the pitch, rhythm, and intensity changes of each segment of voice data, as well as semantics, through an intelligent model.
[0032] It is worth mentioning that, in one embodiment, when generating voice dialogue using a language model, the relevant audio data is transmitted to the user device 100 and can be transcribed and displayed on the dialogue interface. Furthermore, in addition to the voice dialogue being emitted directly from the speaker of the user device 100, the voice dialogue can also be simulated by a lifelike chatbot 101 displayed on the dialogue interface. The running artificial intelligence can generate facial expressions and emoticons to match the current emotion. For example, the generated lifelike character's mouth movements can simulate human mouth movements, and the facial muscle movements can also match the emotion expressed in the voice.
[0033] The cloud server 20 implements a system that provides personalized voice dialogue services to users. This system primarily utilizes hardware and software collaboration to run various artificial intelligence algorithms and models, as can be found in [reference needed]. Figure 2 The diagram shows a system architecture embodiment.
[0034] The cloud server 20 utilizes the collaborative processing circuitry, memory, and software of a computer system to implement various functional modules, such as a natural language processing module 201 for processing content from conversations with the user; an instruction processing module 203 for providing corresponding services based on user requests generated through the dialogue interface; and a user interface module 205 for generating an interface for conversations with the user, which interfaces with the user interface generated by a specific application running on the user's device. The cloud server 20 also includes a machine learning module 207 for training a natural language model that meets user needs; and a database module 209 for providing content from built-in or external databases.
[0035] The cloud server 20 has a built-in or external database 22, which mainly contains user data 221 and multiple simulated objects generated by the artificial intelligence simulation model 223. It can also include an emotion database 225 derived from big data analysis. The emotion database 225 stores data such as audio, speech rate, and word choice for speech under various emotions, allowing the simulated objects to query and generate dialogues corresponding to specific emotions. The cloud server 20 provides personalized voice dialogue services via network 200. Users operate the user device 100 on their terminals and use the software programs running on it to access the services of the cloud server 20 via network 200.
[0036] The database module 209 in the cloud server 20 is used to access and manage the built-in or external database 22, including maintaining user data 221 generated by users using personalized voice dialogue services, and also provides simulated objects generated by the artificial intelligence simulation model 223. The simulated objects can be humanoid, which may include various simulated humanoid chatbots from different fields provided by multiple users, as well as various simulated humanoid chatbots provided by the system.
[0037] According to an embodiment, the database 22 of the cloud server 20 provides realistic objects generated by the artificial intelligence simulation model 223, which can be used to generate multiple realistic humanoid chatbots (excluding realistic humans, animals, or various objects). Besides the realistic humanoids having different appearances and accents, they can also be designed as various realistic humanoid chatbots with different professional knowledge, so as to provide personalized voice dialogue services to users as needed. It is worth mentioning that the method for providing personalized voice dialogue services proposed in this disclosure can use realistic objects generated by the artificial intelligence simulation model 223 to perform natural language dialogue. This method employs generative artificial intelligence, using machine learning algorithms to learn from a large amount of human data to build a realistic human character generation model, and then applying generative artificial intelligence based on 3D graphics technology to generate images. It can generate specific realistic objects according to user needs (providing prompts).
[0038] Furthermore, to give the lifelike chatbot a professional background, this part is achieved through machine learning algorithms. By learning from data in specific domains, a domain-specific intelligent model is built. For example, a Retrieval Augmented Generation (RAG) technique can be used to implement a natural language model for a specific domain based on a large language model (LLM).
[0039] The instruction processing module 203 in the cloud server 20 is used to process various requests transmitted from the user device 100, such as requests to execute a dialogue, requests to select a lifelike chatbot to execute a personalized voice dialogue service, and various requests generated during the voice dialogue.
[0040] The cloud server 20 transmits various information, such as audio and video content, text content and image files, to the user device 100 through the user interface module 205 (including necessary encoding, decoding, compression and decompression programs). It also processes the dialogue content displayed by two or more parties through the dialog interface and uses it to generate various functions and graphics in the dialog interface.
[0041] In cloud server 20, machine learning module 207 runs machine learning algorithms, using neural network deep learning technology to learn the speech features of the user. It trains a large amount of speech data to create a speech model and establishes a natural language processing (NLP) model to obtain speech features from the speech data, derive semantics, and generate corresponding dialogue content. Furthermore, machine learning module 207 can also run machine learning algorithms to learn image features. For example, it can build an image model that can learn user expressions and emotions by training on a large number of user facial image features. Thus, in addition to judging user emotions based on speech features, cloud server 20 can also use image models to process image features to assist in judging user emotions.
[0042] The cloud server 20 uses the natural language processing module 201 to process the voice data obtained from the user device 100. In addition to textualizing the voice content, it also runs the natural language processing model to obtain the voice features in the voice data, recognize the semantics therein, and then generate dialogue content to respond to the user in natural language.
[0043] Furthermore, in the process of obtaining semantics and judging user emotions, the speech features used include changes in pitch, rhythm, and volume. Based on these speech features (such as sound wave waveform features), the speech model can identify the user's current emotions and can also learn the individual user's personality and preferences through long-term learning, thus creating a user profile.
[0044] By using a computer system to run machine learning algorithms to learn human language and classify it, and by deep learning language structure and the relationships between sentences, a natural language processing model is generated. In the method proposed in this disclosure, the natural language processing model can be trained using speech data generated from conversations between users to form a personalized language model. Furthermore, after training the natural language processing model to generate a personalized language model, deep learning algorithms can be used to learn features such as pitch, rhythm, and volume changes in speech, thereby learning emotions in the language.
[0045] It is worth mentioning that the natural language processing module 201 does not simply utilize existing natural language search / dialogue services such as ChatGPT, which employs a large language model (LLM). This is because existing ChatGPT can only respond to user dialogue in a one-way manner and fails to learn the user's emotions or reference environmental information. Therefore, the natural language processing module 201 further utilizes user feature files, current semantics, and real-time environmental information (such as current geographical location, weather, and time) to enable the lifelike chatbot to engage in dialogue based on the user's current emotions, generating dialogues that meet the user's real-time needs. According to an embodiment, the lifelike chatbot can query the emotion database 225 to retrieve data such as audio, speech rate, and word choice corresponding to various emotions, thereby uttering dialogues corresponding to specific emotions.
[0046] For example, when it is noon, the system obtains the user's geographical location (such as the user's device location information obtained through the application) and weather (cold, hot, sunny or rainy) through environmental information. The lifelike chatbot can proactively suggest suitable restaurants based on the user's preferences. The user can respond whether they want to eat, what their acceptable price range is, or other needs, so that the lifelike chatbot can adjust its suggestions in real time or conduct dialogue to confirm the needs.
[0047] According to an embodiment, in the machine learning algorithm that determines the current emotion based on the image features of the user's face, a supervised machine learning method can be used to learn the changes in image features labeled by humans under different emotions, such as changes in the distance ratio between facial features or specific parts, line changes, shape changes, and color changes. This trains an image model for judging emotions based on image features. Thus, in the cloud server 20, both the speech model and the image model can be used simultaneously to determine the user's emotion. The natural language processing module 201 then obtains the semantics and generates dialogue content, where the audio, speech rate, and word choice of the dialogue can be obtained by querying the emotion database 225.
[0048] Based on the above system architecture and implementation method for providing personalized voice dialogue services to users, Figure 3 Next, an embodiment diagram of a method for providing personalized voice dialogue services to users is shown.
[0049] This example demonstrates that the user's terminal uses the human-machine interface 303 to convert the user's voice / image 301 into text content displayed on the dialogue interface 305 or to issue a voice message, while the cloud server receives the text content or voice data. The user also operates an application running on the user's device to request a personalized voice dialogue service from the cloud server, which then provides the user with a personalized voice dialogue service through the dialogue interface 305.
[0050] The cloud server executes a method for providing personalized voice dialogue services to users through one or more processors of its computer system, acquiring user voice / image 301 generated from the user device through the dialogue interface 305. According to an embodiment, the system utilizes various sensors in the user device to acquire the user's voice and image data. For example, after the user device initiates the dialogue interface 305, it simultaneously activates the microphone to receive the user's voice, activates the camera to capture an image of the user's face, and generates speech (or inputs text) through the human-machine interface 303, enabling interaction with realistic objects displayed in the dialogue interface 305 and dialogue in natural language.
[0051] Voice and image data are transmitted from the user device to the cloud server. The cloud server runs a voice / image model 307, which can be divided into a voice model and an image model. For voice data, the voice model transcribes the voice data to obtain its semantics and determine the user's current emotion. For image data, the image model analyzes the image data to obtain the user's facial features. For example, it can analyze the color and lighting changes of pixels corresponding to the user's facial organs in the image to obtain changes in the images of facial features or specific parts such as the mouth and eyes. If the pixels of each part reflect the geometric relationship between facial features or specific parts, combined with semantics and voice features, the user's current emotion can be determined.
[0052] Furthermore, the voice / image model 307 can also obtain current environmental information 313 from an external database (or server). According to the embodiment, after the cloud server connects to the user device, it can obtain the location information in the user device and determine the user's geographical location. Therefore, it can obtain environmental information 313 based on the user's location and time from an external database, such as local weather, surrounding store business information, geographical environment, and holidays, to facilitate providing effective content to the user. When the cloud server runs the voice / image model 307, it also queries the user database 315 to obtain user feature files, from which basic user data, learned preferences, historical records, etc., can be obtained. Among them, the cloud server also stores and updates each user's data according to the time dimension through the user database 315, recording each user's historical dialogue records.
[0053] In this way, the cloud server can effectively provide personalized voice dialogue services to users through the voice / image model 307. Furthermore, when users engage in personalized voice dialogue, the cloud server continuously acquires the user's voice / image features 311 and continuously learns the voice and image data generated in the user's dialogue through the feature learning algorithm 309 to update or optimize the voice / image model 307.
[0054] Based on the above system architecture and system operation implementation examples, Figure 4 Next, a flowchart of an embodiment of a method for providing personalized voice dialogue services to users is shown.
[0055] The method runs on a cloud server. The user operates an application running on their user device to establish a connection with the cloud server and initiate a dialogue interface through the application. The user device can acquire real-time voice data generated by the user's speech through sensing components such as a microphone (or with the addition of a camera), or add image data (step S401). Then, it learns the voice features, which may include changes in pitch, rhythm, and volume of the voice transmitted by the user through the user device (step S403). Thus, the meaning can be identified based on the voice features. When receiving the real-time voice data generated by the user, according to one embodiment, real-time image data of the user is also received, and the user's current image features are obtained. Therefore, the user's current emotion can be learned simultaneously based on the voice features (including meaning) and image features (step S405).
[0056] According to the embodiment, while obtaining voice features, image features can also be obtained based on image data. Then, the formed voice and / or image features 41 are transmitted to the image generation model 43 of the cloud server to generate realistic images, such as the facial expressions, mouth and specific parts of the image of the realistic humanoid chatbot. They are also transmitted to the personalized natural language model 45 to generate voice dialogue with emotions derived from features 41.
[0057] In step S405, which determines the user's current emotion, the cloud server uses a natural language processing model to identify the semantics in the speech features. Based on the semantics and the current environmental information, it generates a voice dialogue based on the user's current emotion, which is also a voice dialogue that responds to the user's real-time voice.
[0058] Furthermore, the cloud server uses image generation model 43 to generate an image of the simulated object based on voice features or by adding image features (step S407), and displays the simulated object in the dialogue interface on the user device's display (step S409). A personalized natural language model 45 then generates dialogue content based on the user's semantics, current emotion, or environmental information to conduct a voice dialogue with the user (step S411). For example, the simulated object could be a lifelike chatbot, and the image of the simulated object mainly reflects current emotions, such as facial expressions, mouth, and changes in specific body parts.
[0059] Thus, by repeating the above steps, such as receiving voice data generated by the user through the user device, or adding image data to obtain voice features, or adding image features to learn the user's current emotions, personalized voice dialogues are continuously generated.
[0060] Figure 5 Next, a flowchart of another embodiment of a method for providing personalized voice dialogue services to users is shown.
[0061] According to one embodiment, the cloud server can generate an interactive interface through a web server program, providing users with the option to open multiple lifelike chatbots after connecting their user devices to the cloud server. Users can then select one of the lifelike chatbots (step S501) and start a voice conversation (step S503). When the user has a voice conversation with the lifelike chatbot through the conversation interface, in addition to activating the microphone to record sound, a camera can also be activated to record the user and have a video conversation with the lifelike chatbot. At the same time, the cloud server will obtain real-time voice and image data (step S505).
[0062] According to an embodiment, in a cloud server, a transformation model is used to learn the data correlations of voice data (or image data combined) and extract voice features, or image features combined (step S507). The cloud server uses neural network deep learning technology, such as a transformation model employing an attention mechanism, to learn the data correlations based on the features of voice and image data. Specifically, it can learn the user's preferences and personality based on voice and image data continuously received from the user's device in the past and present, building a user profile to describe the user's preferences and personality. Depending on the application, the neural network deep learning technology used by the cloud server can learn about the user's friends, family, and occupation through the user's online activity data, making the user profile more complete and providing voice dialogue services that are closer to the user's needs, especially emotional needs.
[0063] Thus, the software program running on the cloud server can continuously extract features from voice and image data to build or update personalized profiles. Furthermore, after obtaining semantic information, it determines emotions based on semantics and user feature files (step S509). Further, the cloud server can also obtain current environmental information, including geographic information and time received from the user's device. After connecting to an external server, it can obtain real-time weather information based on the geographic information (step S511).
[0064] Having obtained the user's current voice characteristics and semantics, and also obtained the user feature file based on the user's recognition information, the process generates a corresponding voice dialogue based on the semantics, the user feature file, and / or the current environmental information (step S513). The process then repeats steps S503 to S513, repeatedly receiving voice data generated by the user through the user device, obtaining voice characteristics, learning the user's current emotions, and continuously generating personalized voice dialogues.
[0065] According to yet another embodiment, based on the above system and method, the following is also proposed: Figure 6 The trading system shown.
[0066] The trading system includes an intelligent model trading platform 600 built on a computer system. The intelligent model trading platform 600 has a model database 601, a user database 603, and a web server 605. The user database 603 is used to store user data stored and updated according to the time dimension. The intelligent model trading platform 600 provides users with customized personalized lifelike chatbots through a web interface initiated by the web server 605. On the other hand, it also allows users to remotely select (or obtain through trading) lifelike chatbots and import them into the user's device to execute voice dialogue applications.
[0067] The intelligent model trading platform 600 provides multiple lifelike chatbot options through an interactive interface. These lifelike chatbots, generated by artificial intelligence models, can be stored in the model database 601. Users can connect to the intelligent model trading platform 600 via network 60 through various user devices 61, 62, 63, etc., and select one of the lifelike chatbots from the options displayed on the interactive interface to perform voice conversations.
[0068] According to an embodiment, the system, implemented through a computer system, allows users to share personalized, lifelike chatbots customized by the system with other users. These user-customized lifelike chatbots may possess expertise in a specific field or be uniquely designed lifelike chatbots that can be listed and sold on the intelligent model trading platform 600.
[0069] In one embodiment of customizing a lifelike chatbot, the transaction system receives voice data generated by each user from the user's device via a computer system, and obtains the user's current voice characteristics. It then performs deep learning to learn the user's preferences and personality based on continuously received voice data (or including image data) from the past and present. Similarly, a user feature file can be created based on voice characteristics (or image characteristics). In this way, the user can also design the appearance of a personalized lifelike chatbot. When the intelligent model transaction platform 600 receives the design parameters through a web interface, it creates a personalized lifelike chatbot based on the user feature file and the design parameters, and stores it in the model database 601.
[0070] In this way, the user can operate the user device to connect to the intelligent model trading platform 600, obtain the image signal and model data of one of the lifelike chatbots from the model database 601 in the trading system, display the lifelike chatbot on the user device's display screen, and start a voice conversation.
[0071] In summary, according to the method, system, and transaction system for providing personalized voice dialogue services described in the above embodiments, end users can learn their personal data through the hardware and software collaboration functions implemented in the cloud server. Using machine learning algorithms and natural language models, an intelligent model capable of providing personalized voice dialogue services can be established, and a chatbot can provide natural language services. Furthermore, it can determine the user's current emotion, and, in conjunction with environmental information and user preferences, provide personalized voice dialogue services based on the user's current semantics. The personalized, lifelike chatbot trained by the user through the cloud server can also be shared with others through an intelligent model trading platform implemented in the transaction system.
[0072] The content disclosed above is only a preferred and feasible embodiment of the present invention, and is not intended to limit the claims of the present invention. Therefore, all equivalent technical changes made based on the content of the present invention specification and drawings are included within the scope of the claims of the present invention.
Claims
1. A method for providing personalized voice dialogue services to users, running on a cloud server, characterized in that... The method includes: The system receives real-time voice data generated by a user through a user device and obtains the user's current voice characteristics. Learn the user's current emotion based on the voice characteristics; A natural language processing model is used to identify the semantics in the speech features, and based on the semantics and current environmental information, a corresponding voice dialogue based on the user's current emotion is generated; and The voice dialogue is delivered through a realistic object displayed on a screen of the user device; The system repeatedly receives voice data generated by the user through the user's device, obtains voice features, learns the user's current emotions, and continuously generates personalized voice dialogues.
2. The method for providing personalized voice dialogue services to users as described in claim 1, characterized in that... The voice features include pitch, rhythm, and intensity variations in the voice transmitted by the user through the user device.
3. The method for providing personalized voice dialogue services to users as described in claim 2, characterized in that, While receiving the voice data generated by the user's real-time voice, the system also receives the user's real-time image data and obtains the user's current image features. At the same time, the system learns the user's current emotion based on the voice features and image features.
4. The method for providing personalized voice dialogue services to users as described in claim 3, characterized in that... By querying an emotion database, the audio, speech rate, and word choice corresponding to the user's current emotion are obtained, and then a personalized voice dialogue is generated.
5. The method for providing personalized voice dialogue services to users as described in claim 3, characterized in that, On the cloud server, deep learning is performed to learn the user's preferences and personality based on the voice and image data continuously received from the user's device in the past and present, and to build a user profile.
6. The method for providing personalized voice dialogue services to users as described in claim 5, characterized in that... The current environmental information obtained by the cloud server includes geographical information and time received from the user device, and weather information obtained based on the geographical information by connecting to an external server.
7. The method for providing personalized voice dialogue services to users as described in claim 6, characterized in that, In the cloud server, the user's current voice features and semantics are obtained, and the user's feature file is obtained based on the user's recognition information. That is, the voice dialogue is generated based on the semantics, the user feature file, and / or the current environmental information.
8. The method for providing personalized voice dialogue services to users as described in any one of claims 1 to 7, characterized in that, When the cloud server is running a method to provide personalized voice dialogue services to users, the simulated object generated by an artificial intelligence simulation model is obtained from a database, and image signals and model data are transmitted to the user device. The simulated object is displayed on a display of the user device, and a simulated voice dialogue is delivered.
9. The method for providing personalized voice dialogue services to users as described in claim 8, characterized in that... The simulated object is a lifelike humanoid chatbot, and the cloud server generates an interactive interface through a web page server program, providing multiple options for the lifelike humanoid chatbot.
10. A system for providing personalized voice dialogue services to users, characterized in that... The system includes: A cloud server, connected to a user device, providing a personalized voice dialogue service to the user through a dialog interface, and a method for providing the personalized voice dialogue service to the user, executed by one or more processors, comprising: The user device receives voice data generated by a user's real-time voice and obtains the user's current voice characteristics. Learn the user's current emotion based on the voice characteristics; A natural language processing model is used to identify the semantics in the speech features, and based on the semantics and current environmental information, a corresponding voice dialogue based on the user's current emotion is generated; and The voice dialogue is delivered through a realistic object displayed on a screen of the user device; The system repeatedly receives voice data generated by the user through the user's device, obtains voice features, learns the user's current emotions, and continuously generates personalized voice dialogues.
11. The system for providing personalized voice dialogue services as described in claim 10, characterized in that, While receiving the voice data generated by the user's real-time voice, the system also receives the user's real-time image data and obtains the user's current image features. At the same time, the system learns the user's current emotion based on the voice features and image features.
12. The system for providing personalized voice dialogue services as described in claim 11, characterized in that... By querying an emotion database, the audio, speech rate, and word choice corresponding to the user's current emotion are obtained, and then a personalized voice dialogue is generated.
13. The system for providing personalized voice dialogue services as described in claim 11, characterized in that, On the cloud server, deep learning is performed to learn the user's preferences and personality based on the voice and image data continuously received from the user's device in the past and present, and to build a user profile.
14. The system for providing personalized voice dialogue services as described in claim 13, characterized in that... The current environmental information obtained by the cloud server includes geographical information and time received from the user device, and weather information obtained based on the geographical information by connecting to an external server.
15. The system for providing personalized voice dialogue services as described in claim 14, characterized in that, In the cloud server, the user's current voice features and semantics are obtained, and the user's feature file is obtained based on the user's recognition information. That is, the voice dialogue is generated based on the semantics, the user feature file, and / or the current environmental information.
16. The system for providing personalized voice dialogue services to users as claimed in any one of claims 10 to 15, characterized in that, When the cloud server is running a method to provide personalized voice dialogue services to users, the simulated object is generated by an artificial intelligence simulation model from a database, the model data and image signals are transmitted to the user device, the simulated object is displayed on a display of the user device, and a simulated voice dialogue is delivered.
17. The system for providing personalized voice dialogue services as described in claim 16, characterized in that... The simulated object is a lifelike humanoid chatbot, and the cloud server generates an interactive interface through a web page server program, providing multiple options for the lifelike humanoid chatbot.
18. A trading system, characterized in that... The aforementioned trading system includes: A smart model trading platform built with a computer system provides options for multiple simulated objects through an interactive interface; A model database that stores the multiple realistic objects generated by an artificial intelligence simulation model; A user database that stores and updates user data based on a time dimension; and A web server that provides users with a customized, realistic virtual object through a web interface; The transaction system receives voice data generated by each user from a user device through the computer system and obtains the user's current voice characteristics. It then performs deep learning to learn the user's preferences and personality based on the voice data continuously received in the past and present, and establishes a user feature file. After receiving design parameters through the web interface, it establishes a personalized simulation object based on the user feature file and the design parameters.
19. The transaction system as described in claim 18, characterized in that... The user device obtains the image signal and model data of the simulated object from the model database, displays the simulated object on a display of the user device, and simulates voice dialogue.
20. The transaction system as described in claim 19, characterized in that, In addition to receiving voice data generated by each user through the computer system, the transaction system also receives the user's image data. After learning the voice features and image features, it creates a user feature file.