Figure duplicating interaction system and method based on artificial intelligence
By combining computer vision and natural language processing systems with generative models, realistic videos, images, voice, and text are generated, solving the problem of lack of personalization and natural interaction in existing human-computer interaction systems. This enables natural and fluent communication between intelligent characters and humans, alleviating the suffering of family members and improving mental health.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- NANTONG XIANGQINFU MEDICAL TECHNOLOGY CO LTD
- Filing Date
- 2023-12-16
- Publication Date
- 2026-04-17
AI Technical Summary
Existing human-computer interaction systems mainly rely on text or video for real-time interaction, lacking the ability to collect information and interact in real time based on images or videos of other people, and thus failing to provide a natural, smooth, and personalized interactive experience.
By employing computer vision, natural language processing, and generative modeling systems, combined with deep learning technology, realistic videos, images, audio, and text are generated, enabling natural and fluid interaction between intelligent characters and humans.
By simulating a person's appearance, voice, and language, it provides a more natural, smooth, and personalized interactive experience, helping family members alleviate pain and improve their mental health and emotional connection.
Abstract
Description
Technical Field
[0001] This invention belongs to the field of human-computer interaction technology, and more specifically, relates to an artificial intelligence-based character replication interaction system and method. Background Technology
[0002] The rapid development of Artificial Intelligence (AI) technology has brought convenience to people's daily work and life. AI is a new technical science that studies and develops theories, methods, technologies, and application systems to simulate, extend, and expand human intelligence. As a branch of computer science, AI attempts to understand the essence of intelligence and produce a new kind of intelligent machine that can react in a way similar to human intelligence. Research in this field includes robotics, speech recognition, image recognition, natural language processing, and expert systems. AI is increasingly integrated into human-computer interaction. AI-based human-computer interaction can analyze user needs and, after the user inputs interactive statements, provide feedback on the user's desired interactive information. Currently, to generate interactive information for users, the common approach is to pre-build databases for different types of interactive statements to generate different types of interactive information.
[0003] Human-computer interaction (HCI) refers to the information exchange process between humans and computers using a certain dialogue language and interactive methods to complete a specific task. With the continuous development of artificial intelligence technology, the use of intelligent devices such as intelligent robots, smart homes, and intelligent customer service is becoming increasingly widespread, and HCI is finding broad applications in multiple fields. As people's living standards gradually improve, they need to have timely access to information about various buildings, shops, restaurants, hotels, and shopping options in the cities they visit, so that they can use these signals for travel, shopping, accommodation, and payment.
[0004] Existing human-computer interaction methods only allow real-time interaction with existing users, using text or video, without collecting information from other people's images or videos for real-time human-computer interaction. Summary of the Invention
[0005] Purpose of the invention: The purpose of this invention is to address the shortcomings of existing technologies and provide an artificial intelligence-based character replication interactive system and method.
[0006] Technical solution: The present invention provides an artificial intelligence-based character replication interactive system, comprising a computer vision system, a natural language processing system, a generative model system, and a human-computer interaction system;
[0007] The computer vision system performs target detection on external images or videos, and identifies and extracts facial feature information and facial expression information from them.
[0008] The natural language processing system enables intelligent characters to communicate by recognizing and extracting speech information, constructing natural language models and machine learning algorithms.
[0009] The generative model system generates high-quality, realistic videos, images, audio, and text by learning data distribution and combining the results of computer vision and natural language processing systems. This helps intelligent characters improve the realism of their appearance, voice, and language, making them closer to real humans.
[0010] The human-computer interaction system enables real-time interaction between intelligent characters and humans, allowing intelligent characters to communicate with humans more flexibly and providing a more natural and smooth interaction effect.
[0011] In some implementations, the computer vision system includes image processing, object detection, face recognition, and facial expression recognition, thereby simulating facial expressions, body language, and visual attention to achieve a more realistic and engaging interactive experience.
[0012] In some embodiments, the computer vision system further includes a human image restoration system, which is built on a GAN generative model and trained by a game between a generator and a discriminator. The generator is trained to generate images with missing pixels, while the discriminator is trained to distinguish between generated images and real images.
[0013] In some implementations, the portrait restoration system employs semantic segmentation technology to divide the image into different regions for image segmentation and reconstruction.
[0014] In some embodiments, the natural language processing system further includes a speech restoration system, which includes a speech enhancement module, a speech segmentation module, a speech recognition module, and a speech reconstruction module.
[0015] In some implementations, the generative model system employs deep learning techniques, using a large amount of training data to train neural networks, thereby enabling the replication of a person's behavior, language, and thought patterns.
[0016] In some implementations, the human-computer interaction system includes a speech recognition module, a command classification module, and an action recognition module, enabling intelligent characters to communicate with humans more flexibly and providing a more natural and fluent interaction effect. At the same time, the human-computer interaction system can also help intelligent characters perceive and adapt to user feedback in real time, and adjust their behavior and language output accordingly to achieve a better interaction effect.
[0017] In some implementations, the human-computer interaction system is based on natural language processing and machine learning techniques. Through a self-attention mechanism, it makes different weighted contributions to each position in the input sequence, thereby capturing the dependencies between different positions in the sequence, and learns dialogue generation skills through pre-training.
[0018] On the other hand, the present invention also discloses a method for using an artificial intelligence-based character replication interactive system, comprising the following steps:
[0019] (1) Accessing the program: Users register and log in through various means to access the software;
[0020] (2) Create a person: Collect data, process data, train the model, deploy the model, and create a person in sequence;
[0021] (3) Human-computer interaction: After deployment and customer settings are completed, users can communicate and interact directly with the intelligent character through various means.
[0022] In some implementations, step (3) human-computer interaction specifically includes the following steps:
[0023] (1) Input processing: Converting the natural language input by the user into a language that the computer can understand, which is accomplished by natural language processing algorithms, including word segmentation, part-of-speech tagging, named entity recognition, and syntactic analysis;
[0024] (2) Intent recognition: Based on understanding user input, determine the operation or intent that the user wants to perform;
[0025] (3) Dialogue flow management: Based on the user's intent, the dialogue is assigned to the appropriate robot module or corresponding processing logic to generate a response containing the answer;
[0026] (4) Generate response: Automatically generate corresponding text or voice response based on user input and robot module;
[0027] (5) Training of dialogue AI networks: For long-term goals, deep learning techniques are used to train specific neural networks to better simulate human dialogue.
[0028] (6) Personalized answers: By learning from users' historical dialogue records, we can understand users' interests, habits, preferences and other information, thereby improving the personalization of our answers.
[0029] Beneficial effects: This invention utilizes advanced artificial intelligence (AI) technology to process and store information such as photos, videos, and voices of people, creating a virtual individual that replicates the person's appearance and voice. By engaging in dialogue with this intelligent person, it helps family members alleviate pain, receive psychological support, and enhance their emotional identification.
[0030] This invention employs deep learning technology, using a large amount of training data to train a neural network, thereby enabling the replication of human behavior, language, and thought patterns. Specifically, a large amount of information about human behavior, language, and thought patterns is input into the neural network for training. Through this process, the intelligent character can gradually "learn" the behavior and thought patterns of a human, thus being able to imitate the human's behavior and thought processes.
[0031] This invention aims to help family members of intelligent characters gain psychological support and a sense of belonging, alleviate their suffering, improve their mental health, thereby enhancing their self-confidence and improving their quality of life. We hope users can utilize it to obtain maximum psychological comfort. Detailed Implementation
[0032] The technical solution of the present invention will be clearly and completely described below with reference to specific embodiments. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0033] In the description of this invention, it should be noted that the terms "center", "upper", "lower", "left", "right", "inner", "outer", etc., indicate the orientation or positional relationship shown, and are only for the convenience of describing this invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of this invention.
[0034] In the description of this invention, it should be noted that, unless otherwise explicitly specified and limited, the terms "installation," "connection," and "linking" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances.
[0035] Example 1
[0036] This invention utilizes advanced artificial intelligence (AI) technology to process and store information such as photos, videos, and audio of a person, creating a virtual individual that replicates the person's appearance and voice. Through dialogue with this intelligent person, the software helps family members alleviate pain, receive psychological support, and enhance emotional connection.
[0037] This embodiment of an artificial intelligence-based character replication interactive system includes a computer vision system, a natural language processing system, a generative model system, and a human-computer interaction system.
[0038] The computer vision system performs target detection on external images or videos, and identifies and extracts facial feature information and facial expression information from them.
[0039] The natural language processing system enables intelligent characters to communicate by recognizing and extracting speech information, constructing natural language models and machine learning algorithms.
[0040] The generative model system generates high-quality, realistic videos, images, audio, and text by learning data distribution and combining the results of computer vision and natural language processing systems. This helps intelligent characters improve the realism of their appearance, voice, and language, making them closer to real humans.
[0041] The human-computer interaction system enables real-time interaction between intelligent characters and humans, allowing intelligent characters to communicate with humans more flexibly and providing a more natural and smooth interaction effect.
[0042] Example 2
[0043] This embodiment of an artificial intelligence-based character replication interactive system includes a computer vision system, a natural language processing system, a generative model system, and a human-computer interaction system.
[0044] The computer vision system detects targets in external images or videos, identifying and extracting facial features and expressions. This system includes image processing, target detection, face recognition, and facial expression recognition, providing the foundation for intelligent characters to perceive their environment and engage in subjective communication. Based on these technologies, intelligent characters can simulate facial expressions, body language, and visual attention, achieving more realistic and engaging interactive effects.
[0045] This embodiment also includes a portrait restoration system, which primarily relies on a GAN generative model. GAN is a deep learning technique used to generate adversarial examples using generative models, trained by a game between a generator and a discriminator. In the ARC portrait restoration system, the generator is trained to generate images with missing pixels, while the discriminator is trained to distinguish between the generated images and real images.
[0046] The portrait restoration system employs a self-attention mechanism to increase the focus on the area to be repaired, thereby improving the generator's restoration effect. The self-attention mechanism effectively selects pixels with more important information, preventing the network from focusing too much on surrounding important information, reducing processing complexity, and avoiding overfilling in certain areas.
[0047] Most importantly, the portrait restoration system utilizes semantic segmentation technology to improve restoration results and the semantic consistency of the image. Semantic segmentation divides an image into different regions for image segmentation and reconstruction. The portrait restoration system uses semantic segmentation to segment the image and repair missing pixels in the segmented regions, thus making the restored image more semantically consistent. The system also employs various techniques to enhance restoration results, such as an improved discriminator network structure, multi-scale processing, and hybrid training of the generator and discriminator. These techniques work together to enable the portrait restoration system to achieve high-quality image restoration while maintaining semantic consistency.
[0048] In summary, the core technologies of portrait restoration systems include GANs, self-attention mechanisms, semantic segmentation, and various improvement techniques. These technologies enable ARC portrait restoration technology to perform portrait restoration efficiently and accurately, achieving better image restoration results.
[0049] The Natural Language Processing (NLP) system enables intelligent characters to communicate verbally by recognizing and extracting speech information, constructing natural language models, and using machine learning algorithms. NLP is another core technology for intelligent characters. It involves technologies such as machine translation, dialogue generation, natural language understanding, and sentiment analysis, and achieves communication capabilities through the construction of natural language models and machine learning algorithms. These technologies can be combined with computer vision technology, enabling intelligent characters to solve problems not only through language but also through non-verbal means such as facial expressions and body language.
[0050] In this embodiment, a natural language processing system can also be used to extract key sentences and perform sentiment analysis from the character's text. This information can then be used to depict the character's emotions, personality, and thought patterns. Character replication requires a large amount of data, and further in-depth analysis of the character's own characteristics is necessary to ensure the accuracy and authenticity of the replication effect.
[0051] This embodiment also includes a voice restoration system, which includes a voice enhancement module, a voice segmentation module, a voice recognition module, and a voice reconstruction module.
[0052] The speech enhancement module is a core technology of the speech restoration system. It mainly relies on deep learning models to denoise and enhance noise signals, making the speech signals clearer and more natural. This technology primarily uses deep learning-based noise recognition and decomposition methods. By analyzing and learning from noise signals, it gains the basic ability to judge noise signals, and then uses noise reduction methods to restore the dialogue audio.
[0053] The speech segmentation module, a crucial component of the speech restoration system, segments the audio signal, removing unwanted segments such as noise and background noise, and then splices the remaining audio together. The speech restoration system primarily utilizes deep learning models, such as convolutional neural networks (CNNs) and recurrent neural networks (RNNs), to segment the audio signal and isolate or remove unwanted audio signals during the restoration process.
[0054] The speech recognition module converts speech signals into text, and then analyzes and repairs the converted text after speech reconstruction. By using deep learning models, such as recurrent neural networks (RNNs) and transformers, speech recognition technology can better perform processes such as speech segmentation, noise reduction, and reconstruction.
[0055] The speech reconstruction module primarily analyzes the frequency, phase, and intensity features of speech signals to generate complex speech signals, thereby restoring distortion and noise. The deep learning models used in speech reconstruction technology include CNN-based WaveNet and transformer-based Tacotron 2, which provide a stable and accurate foundation for speech signal generation and restoration.
[0056] These technologies utilize deep learning models, leveraging powerful self-learning capabilities and optimization strategies to improve the accuracy and speed of speech signal processing, thereby enhancing the practical effectiveness of speech restoration systems.
[0057] The generative model system described above generates high-quality, realistic videos, images, audio, and text by learning data distributions and combining the results processed by computer vision and natural language processing systems. This helps intelligent characters improve the realism of their appearance, voice, and speech, making them more like real humans. Intelligent characters utilize powerful model generation technologies, such as GANs and VAEs, which can learn data distributions to generate high-quality, realistic images, audio, and text. These technologies can help intelligent characters improve the realism of their appearance, voice, and speech, making them more like real humans.
[0058] The human-computer interaction system enables real-time interaction between intelligent characters and humans, allowing the intelligent characters to communicate more flexibly and provide a more natural and fluid interaction experience. Through human-computer interaction technologies, including speech recognition, command classification, and action recognition, the intelligent characters can communicate more flexibly and provide a more natural and fluid interaction experience. Simultaneously, human-computer interaction technologies can also help the intelligent characters perceive and adapt to user feedback in real time, adjusting their behavior and language output accordingly to achieve better interaction results.
[0059] In this embodiment, the human-computer interaction system is mainly based on natural language processing and machine learning technologies.
[0060] Human-computer interaction systems, through self-attention mechanisms, can assign different weighted contributions to each position in the input sequence, thereby capturing the dependencies between different positions in the sequence. In dialogue generation tasks, by learning the context of the dialogue, each sentence in the dialogue can be seen as a response to the previous sentence. In this way, the self-attention mechanism can be used to model the content and context of the dialogue.
[0061] Human-computer interaction systems learn dialogue generation skills through pre-training. Pre-training refers to unsupervised training on a large-scale corpus, pre-training the model into a general language model. In dialogue generation tasks, dialogue data can be constructed into language sequences, and these sequences can then be used for pre-training. In this way, the model can learn features such as syntax, semantics, context, and dialogue flow in the dialogue, providing strong support for subsequent dialogue generation tasks.
[0062] Human-computer interaction systems offer greater flexibility and naturalness. Compared to traditional rule-based or template-based dialogue systems, the YiNian-human-computer dialogue technology can better handle complex scenarios and more accurately understand user needs and intentions, thus providing a more natural and personalized dialogue experience. Furthermore, YiNian-human-computer dialogue technology can utilize multi-turn dialogue techniques, generating more coherent dialogue content by remembering the dialogue context, providing a smoother interactive experience.
[0063] In summary, the human-computer interaction system of this invention employs deep learning technology and uses a large amount of training data to train a neural network, thereby enabling it to replicate a person's behavior, language, and thought patterns. Specifically, a large amount of information about a person's behavior, language, and thought patterns is input into the neural network for training. Through this process, the intelligent character can gradually "learn" the behavior and thought patterns of a person, thus being able to imitate the behavior and thought processes of a person.
[0064] Example 3
[0065] A method for using an AI-based character replication and interaction system includes the following steps:
[0066] (1) Accessing the program: Users can access the software by registering and logging in, regardless of the system, APP, mini-program, or other carriers they use.
[0067] (2) Create a person: Collect data, process data, train the model, deploy the model in sequence, and create a person.
[0068] Data collection: First, users need to collect a large amount of data related to the person they want to replicate, including text (person's resume, diary, books, etc.), pictures (photos), videos, and audio.
[0069] Data Processing: The system cleans, labels, and classifies data to train AI models. Furthermore, data restoration techniques are used to optimize the accuracy and clarity of low-quality images, audio, and video, improving data labeling accuracy so that the model can learn human characteristics and behavioral traits.
[0070] Model training: Staff use deep learning algorithms, combined with user-uploaded data, to train a neural network model so that it can learn from the data the characteristics of a person's behavior, language, and thinking patterns, and generate relevant content based on these characteristics.
[0071] Model Deployment: Deploy the trained model to the system's client interface, where users can configure and optimize it according to their usage habits.
[0072] (3) Human-computer interaction: After deployment and customer setup, users can directly communicate with the intelligent character through short video technology and naked-eye 3D technology. The intelligent character can handle various complex dialogue scenarios and provide users with a high-quality interactive experience, such as supporting multi-turn dialogue, situational dialogue, and task-oriented dialogue.
[0073] Adaptive learning: Intelligent dialogue can use machine learning technology to gradually learn user input information, intentions and behavioral characteristics, and continuously summarize and provide feedback to improve the quality of dialogue interaction.
[0074] Intelligent dialogue: Based on the user's historical records and behavioral characteristics, the intelligent assistant will proactively initiate dialogues with the user based on trending topics such as weather, historical communication records, and social news, offering suggestions, expressions of concern, reminders, and warnings. It can also recommend information the user may like, such as products, news, and music.
[0075] Sentiment Analysis: Intelligent dialogue can analyze users' emotions and expressions, such as happiness, sadness, and anger, to more accurately understand users' intentions and emotional needs, and provide more humanized services and communication.
[0076] Multimedia Interaction: Intelligent character dialogue can support the input and output of multimedia languages such as voice, images, and video, providing a more vivid and interactive interactive experience.
[0077] Security Guarantee: YiNian - Intelligent Character Dialogue ensures user privacy and information security by employing encryption, anti-plagiarism, and other measures to protect users' personal information and data security.
[0078] Step (3) of human-computer interaction specifically includes the following steps:
[0079] (1) Input processing: Converting the natural language input by the user into a language that the computer can understand. This is usually done using natural language processing algorithms, including word segmentation, part-of-speech tagging, named entity recognition, and syntactic analysis.
[0080] (2) Intent recognition: Based on understanding user input, determine the operation or intent that the user wants to perform. This is usually accomplished using machine learning-based classification algorithms.
[0081] (3) Dialogue flow management: Based on the user's intent, the dialogue is assigned to the appropriate robot module or corresponding processing logic to generate a response containing an answer. This process typically involves the design of the dialogue flow and business logic.
[0082] (4) Response Generation: Based on user input and the robot module, the system automatically generates corresponding text or speech responses. It can generate coherent and natural answers based on the context of the dialogue. By memorizing contextual information, it selects the best response and utilizes natural language generation technology to construct natural and reasonable statements, making the responses more consistent with human communication habits and norms. This typically uses natural language generation algorithms, such as template-based text generation, rule-based text generation, neural network-based text generation, and speech synthesis. It also enables multi-turn dialogues by memorizing dialogue history. The system can understand the flow and theme of the dialogue, as well as the user's needs and intentions, and provide corresponding responses in subsequent dialogues, making the dialogue process more natural and fluent.
[0083] (5) Training of dialogue AI networks: For long-term goals, deep learning techniques are used to train specific neural networks to better simulate human dialogue.
[0084] (6) Personalized answers: By learning from users' historical dialogue records, we can understand users' interests, habits, preferences and other information, thereby improving the personalization of our answers.
[0085] In summary, the system and method of this invention, combining artificial intelligence with image replication, aims to help the families of intelligent characters gain psychological support and a sense of belonging, alleviate their suffering, improve their mental health, thereby enhancing their self-confidence and contributing to a higher quality of life. We hope users can utilize it to obtain maximum psychological comfort.
[0086] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any way. Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some modifications or alterations to the above-disclosed technical content to create equivalent embodiments without departing from the scope of the present invention. Any simple modifications, equivalent changes, and alterations made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the scope of the present invention.
Claims
1. An artificial intelligence-based character replication interactive system, characterized in that: This includes computer vision systems, natural language processing systems, generative modeling systems, and human-computer interaction systems. The computer vision system performs target detection on external images or videos, and identifies and extracts facial feature information and facial expression information from them. The natural language processing system enables intelligent characters to communicate by recognizing and extracting speech information, constructing natural language models and machine learning algorithms. The generative model system generates high-quality, realistic videos, images, audio, and text by learning data distribution and combining the results of computer vision and natural language processing systems. This helps intelligent characters improve the realism of their appearance, voice, and language, making them closer to real humans. The human-computer interaction system enables real-time interaction between intelligent characters and humans, allowing intelligent characters to communicate with humans more flexibly and providing a more natural and smooth interaction effect.
2. The character replication interactive system based on artificial intelligence according to claim 1, characterized in that: The computer vision system includes image processing, object detection, face recognition, and facial expression recognition, thereby simulating facial expressions, body language, and visual attention to achieve a more realistic and engaging interactive experience.
3. The character replication interactive system based on artificial intelligence according to claim 1, characterized in that: The computer vision system also includes a human image restoration system, which is built on a GAN generative model. The model is trained by a game between the generator and the discriminator. The generator is trained to generate images with missing pixels, while the discriminator is trained to distinguish between the generated images and real images.
4. The character replication interactive system based on artificial intelligence according to claim 3, characterized in that: The portrait restoration system uses semantic segmentation technology to divide the image into different regions for image segmentation and reconstruction.
5. The character replication interactive system based on artificial intelligence according to claim 1, characterized in that: The natural language processing system also includes a speech restoration system, which comprises a speech enhancement module, a speech segmentation module, a speech recognition module, and a speech reconstruction module.
6. The character replication interactive system based on artificial intelligence according to claim 1, characterized in that: The generative model system employs deep learning technology, using a large amount of training data to train neural networks, thereby enabling it to replicate a person's behavior, language, and thought patterns.
7. The character replication interactive system based on artificial intelligence according to claim 1, characterized in that: The human-computer interaction system includes a voice recognition module, a command classification module, and a motion recognition module, enabling intelligent characters to communicate with humans more flexibly and providing a more natural and smooth interaction effect. At the same time, the human-computer interaction system can also help intelligent characters perceive and adapt to user feedback in real time, and adjust their behavior and language output accordingly to achieve a better interaction effect.
8. The character replication interactive system based on artificial intelligence according to claim 7, characterized in that: The human-computer interaction system is based on natural language processing and machine learning technology. Through a self-attention mechanism, it makes different weighted contributions to each position in the input sequence, thereby capturing the dependencies between different positions in the sequence, and learns dialogue generation skills through pre-training.
9. A method of using an artificial intelligence-based character replication interactive system according to any one of claims 1-8, characterized in that: Includes the following steps: (1) Accessing the program: Users register and log in through various means to access the software; (2) Create a person: Collect data, process data, train the model, deploy the model, and create a person in sequence; (3) Human-computer interaction: After deployment and customer settings are completed, users can communicate and interact directly with the intelligent character through various means.
10. The method according to claim 9, characterized in that: Step (3) of human-computer interaction specifically includes the following steps: (1) Input processing: Converting the natural language input by the user into a language that the computer can understand, which is accomplished by natural language processing algorithms, including word segmentation, part-of-speech tagging, named entity recognition, and syntactic analysis; (2) Intent recognition: Based on understanding user input, determine the operation or intent that the user wants to perform; (3) Dialogue flow management: Based on the user's intent, the dialogue is assigned to the appropriate robot module or corresponding processing logic to generate a response containing the answer; (4) Generate response: Automatically generate corresponding text or voice response based on user input and robot module; (5) Training of dialogue AI networks: For long-term goals, deep learning techniques are used to train specific neural networks to better simulate human dialogue. (6) Personalized answers: By learning from users' historical dialogue records, we can understand users' interests, habits, preferences and other information, thereby improving the personalization of our answers.