Human-computer Interaction Method, System, Device and Medium Based on Asynchronous Art Creation
By adopting a human-computer interaction method based on asynchronous art creation in art therapy, using artificial intelligence to process visitors' art works and conversation information, and generating co-created art works and conversation summary, the problems of high threshold for artistic creation and difficult tracking in the existing technology are solved, and more efficient and personalized art therapy effects are achieved.
Patent Information
- Application Number
- CN202510147305.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-11
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2045-02-11
AI Technical Summary
In existing art therapy, the lack of guidance from service providers leads to a high psychological threshold for artistic creation, which increases the difficulty of performing tasks, is not effective, and it is difficult for service providers to track the creative process of visitors.
A human-computer interaction method based on asynchronous art creation is adopted to collect the visitors’ original art works and multimodal conversation information in real time, use artificial intelligence models to perform dynamic analysis and processing, generate human-computer co-creation art works and conversation summary, and feed these results to the service providers so that the service provider can generate the next interactive task with the assistance of artificial intelligence.
It reduces the difficulty of the client performing artistic creation tasks, improves the efficiency and effectiveness of the tasks, helps the client establish long-term trust and intimacy, and the server can better track and customize artistic creation interactive tasks.
Smart Images

Figure CN119626465B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of human - computer interaction technology, and particularly relates to a human - computer interaction method, system, device and medium based on asynchronous art creation. Background Art
[0002] Art therapy is a multi - modal therapy method that combines creative art activities and verbal expression, which can help people release emotions and improve mental health. In this therapy, the service provider and the client cooperate through the interaction during art therapy sessions and interactive tasks.
[0003] In the prior art, the service provider recommends interactive tasks that combine art creation and oral expression. For example, the service provider encourages the client to complete daily home creations, such as observing, painting, taking pictures and writing about the sky; combining art creation with oral practice in interactive tasks to help the client improve daily coping abilities. This method can not only enable the client to deeply explore inner thoughts and find ways to release emotions through language, but also bring a unique healing experience through the creativity and expressiveness of art creation. However, the lack of guidance from the service provider may increase the psychological threshold of the client's art creation. Coupled with individual differences in emotional understanding and expression, it becomes more challenging to express past emotions and experiences in language. At the same time, it is difficult for the service provider to track the treatment situation, and it is difficult for the service provider to understand the specific situation of the client's creative process. Therefore, ensuring the treatment effect of art creation is a problem that needs to be solved.
[0004] It can be seen that the art therapy in the prior art has problems such as high difficulty in effective implementation, poor effect and high tracking difficulty due to the high psychological threshold of art creation. Summary of the Invention
[0005] In view of the above - mentioned deficiencies of the prior art, the purpose of the present invention is to provide a human - computer interaction method, system, device and medium based on asynchronous art creation, aiming to solve the problems of high difficulty in effective implementation, poor effect and high tracking difficulty in the art therapy in the prior art due to the high psychological threshold of art creation.
[0006] To achieve the above purpose, the first aspect of the present invention provides a human - computer interaction method based on asynchronous art creation, including:
[0007] Real - time collecting the original art works and multi - modal conversation information input by the client according to the current interactive task;
[0008] Using an artificial intelligence model combination to dynamically analyze and process the original art works and multi - modal conversation information respectively, and generating a human - machine co - created art work and a conversation summary;
[0009] Feed the human-machine co-created art work and the conversation summary back to the service provider, and in response to receiving the next interaction task generated by the service provider with the assistance of artificial intelligence based on the conversation summary and the human-machine co-created art work, feed the next interaction task back to the visitor.
[0010] Optionally, the original art work and multi-modal conversation information input by the visitor according to the current interaction task are collected in real time, including:
[0011] Based on the conversation principles and example questions corresponding to the current interaction task, construct a series of interaction questions;
[0012] Feed the interaction questions to the visitor in sequence, and collect the original art work and multi-modal conversation information corresponding to each interaction question feedback by the visitor in sequence.
[0013] Optionally, the artificial intelligence model combination is used to dynamically analyze and process the original art work and multi-modal conversation information respectively to generate a human-machine co-created art work and a conversation summary, including:
[0014] Use the first module in the artificial intelligence model combination to analyze the information other than the art creation information in the multi-modal conversation information to generate several types of conversation instructions;
[0015] Feed the conversation instructions back to the visitor, and in response to receiving the feedback information of the visitor on the conversation instructions, generate several types of conversation information;
[0016] Fuse all the conversation information to generate a conversation summary.
[0017] Optionally, using the first module in the artificial intelligence model combination to analyze the information other than the art creation information in the multi-modal conversation information to generate several types of conversation instructions, including:
[0018] Perform standardization processing and / or data conversion on the information other than the art creation information in the multi-modal conversation information to obtain data in the target format;
[0019] Use the first module in the artificial intelligence model combination to analyze the data in the target format to generate several types of conversation instructions.
[0020] Optionally, the artificial intelligence model combination is used to dynamically analyze and process the original art work and multi-modal conversation information respectively to generate a human-machine co-created art work and a conversation summary, including:
[0021] Analyze the artistic creation information in the multimodal conversation information using the second module in the artificial intelligence model combination to generate an artistic creation instruction and an artistic creation summary;
[0022] Feed back the artistic creation instruction to the visitor, and use the third module in the artificial intelligence model combination to dynamically analyze and process the artistic creation information and the artistic creation summary fed back by the visitor according to the artistic creation instruction to generate a human-machine co-created artistic work.
[0023] Optionally, the use of the third module in the artificial intelligence model combination to dynamically analyze and process the artistic creation information and the artistic creation summary fed back by the visitor according to the artistic creation instruction to generate a human-machine co-created artistic work includes:
[0024] In response to receiving the artistic creation information fed back by the visitor according to the artistic creation instruction, obtain the artistic creation interface information;
[0025] Use the third module in the artificial intelligence model combination to perform color segmentation on the artistic creation interface information to obtain information on several color modules;
[0026] Generate a human-machine co-created artistic work based on the information of all the color modules and the artistic creation summary.
[0027] Optionally, the generation process of the next interaction task includes:
[0028] Collect historical interaction results within a preset time period;
[0029] Automatically generate the current interaction result based on the conversation summary and the human-machine co-created artistic work;
[0030] In response to receiving at least one piece of information among the current interaction result, the human-machine co-created artistic work, and the historical interaction results from the service provider, generate the next interaction task with the assistance of artificial intelligence.
[0031] Optionally, the collection of historical interaction results within a preset time period includes:
[0032] Collect the personal information of the visitor;
[0033] Collect the number of times the visitor has performed the interaction tasks within the preset time period;
[0034] Collect the conversation summary of each interaction task performed by the visitor within the preset time period;
[0035] Generate historical interaction results based on the personal information, the number of times, and / or the conversation summary.
[0036] Optionally, automatically generating a current interaction result based on the session summary and the human-machine co-created art work, including:
[0037] Collecting the original art work, the human-machine co-created art work, and the session summary corresponding to several target time nodes in the current interaction task;
[0038] Automatically generating a current interaction result based on the original art work, the human-machine co-created art work, and / or the session summary corresponding to all the target time nodes.
[0039] Optionally, generating a next interaction task with the assistance of artificial intelligence, including:
[0040] Responding to the session reflection content fed back by the visitor according to the current interaction result, generating optimized dialogue principles and optimized example questions;
[0041] Generating a next interaction task with the assistance of artificial intelligence based on the optimized dialogue principles and the optimized example questions.
[0042] Optionally, the multimodal session information includes:
[0043] At least one of audio data, video data, text data, art creation information, physiological information, and environmental information.
[0044] A second aspect of the present invention provides a human-machine interaction system based on asynchronous art creation, and the system includes:
[0045] An information collection module, configured to collect in real time the original art work and multimodal session information input by a visitor according to the current interaction task;
[0046] An art creation module, configured to respectively perform dynamic analysis and processing on the original art work and the multimodal session information by using an artificial intelligence model combination to generate a human-machine co-created art work and a session summary;
[0047] An interaction module, configured to feed back the human-machine co-created art work and the session summary to a service provider, respond to the next interaction task generated by the service provider with the assistance of artificial intelligence based on the session summary and the human-machine co-created art work, and feed back the next interaction task to the visitor.
[0048] A third aspect of the present invention provides a human-machine interaction device, and the human-machine interaction device includes a memory for storing executable instructions; a processor for calling and running the executable instructions in the memory to execute the steps of any one of the above-mentioned human-machine interaction methods based on asynchronous art creation.
[0049] The fourth aspect of the present invention provides a computer-readable storage medium, in which program instructions are stored. When the program instructions are run by a processor, the steps of any one of the above-mentioned human-computer interaction methods based on asynchronous art creation are implemented.
[0050] Compared with the prior art, the beneficial effects of this solution are as follows:
[0051] The present invention combines art creation and multi-modal conversation information through human-computer interaction, helping visitors externalize their feelings and experiences during the art creation process. With the assistance of artificial intelligence, the service provider provides structured guidance based on the conversation summary and the co-created artworks by humans and machines, so as to integrate the professional beliefs, personal styles, emotional support, etc. of the service provider into the interaction tasks, and can provide customized and personalized interaction needs for visitors. The present invention realizes the asynchronous collaboration between the service provider and the visitor through a human-computer interaction system, enables the service provider and the visitor to maintain interaction between two meetings, and helps the visitor establish long-term trust and intimacy; at the same time, under the guidance of the artificial intelligence model, the difficulty for the visitor to perform art creation tasks is significantly reduced, which helps to improve the confidence of the visitor and the efficiency of asynchronous interaction. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0053] Figure 1 It is a flowchart of the human-computer interaction method based on asynchronous art creation of the present invention;
[0054] Figure 2 It is a schematic diagram of the process of the human-computer interaction system based on asynchronous art creation of the present invention;
[0055] Figure 3 It is a schematic diagram of the application program interface for visitors of the present invention;
[0056] Figure 4 It is a schematic diagram of the application program interface for visitors of the present invention;
[0057] Figure 5 It is a schematic diagram of the application program interface for service providers of the present invention;
[0058] Figure 6 It is a schematic diagram of the application program interface for service providers of the present invention;
[0059] Figure 7It is the interaction flowchart of the first stage for the example of performing the artistic creation interaction task of the present invention;
[0060] Figure 8 It is the interaction flowchart of the second stage for the example of performing the artistic creation interaction task of the present invention;
[0061] Figure 9 It is the interaction flowchart of the third stage for the example of performing the artistic creation interaction task of the present invention;
[0062] Figure 10 It is the structural schematic diagram of the human-computer interaction device of the present invention. Detailed implementation manners
[0063] In the following description, for the purpose of illustration rather than limitation, specific details such as specific system structures, technologies, etc. are presented to thoroughly understand the embodiments of the present invention. However, those skilled in the art should clearly understand that the present invention can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present invention.
[0064] It should be understood that when used in this specification and the appended claims, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.
[0065] It should also be understood that the terms used in the specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in the specification of the present invention and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" are intended to include the plural forms.
[0066] It should be further understood that the term "and / or" used in the specification of the present invention and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0067] As used in this specification and the appended claims, the term "if" can be interpreted according to the context as "when", "once", "in response to a determination", or "in response to a detection". Similarly, the phrase "if a determination" or "if a [described condition or event] is detected" can be interpreted according to the context as meaning "once a determination is made", "in response to a determination", "once a [described condition or event] is detected", or "in response to a detection of a [described condition or event]".
[0068] Combined with the accompanying drawings of the embodiments of the present invention, the technical solutions in the embodiments of the present invention are clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts fall within the scope of protection of the present invention.
[0069] In the following description, many specific details are set forth in order to provide a thorough understanding of the present invention. However, the present invention may be implemented in other ways different from those described herein. Those skilled in the art can make similar generalizations without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below.
[0070] In the art therapy in the prior art, service providers often recommend interactive tasks that combine art creation and oral expression. For example, service providers encourage clients to complete home creations every day, such as observing, painting, taking pictures, and writing about the sky; combining art creation with oral exercises in the interactive tasks to help clients improve their daily coping abilities. This method can not only enable clients to deeply explore their inner thoughts and find ways to release emotions through words, but also bring unique healing experiences through the creativity and expressiveness of art creation. However, the lack of guidance from service providers may increase the psychological threshold of art creation. Coupled with individual differences in emotional understanding and expression, it becomes more challenging to express past emotions and experiences in words. For example, a client is assigned to create a painting expressing emotions as an interactive task without clear guidance. The lack of structural support may lead the client to have negative beliefs such as "not having the ability to express emotions" or "the work will not be understood", thus feeling frustrated. Such experiences will not only make the client lose motivation for creation, but may also affect his trust in the entire therapy.
[0071] Meanwhile, service providers face numerous challenges in tracking and customizing artistic creation interaction tasks. The interaction tasks support asynchronous collaboration between service providers and clients, enabling clients to remain engaged between sessions and helping to build long-term trust and intimacy. This asynchronous communication requires service providers to track the records and history of interaction tasks, and this valuable data can provide insights into the client's mindset during the treatment interval, serving as a basis for adjusting the treatment plan. However, managing and tracking these records increases the workload of service providers and is not easily achievable. In addition, customizing personalized creations while recording treatment progress is crucial for strengthening the treatment course. However, creating structured family creations that take into account the emotional needs of clients remains a major challenge in treatment. For example: ① The service provider learned that the client had difficulty expressing emotions and suggested that he record his emotions through a diary, but found that the client felt pressured by writing. So the service provider adjusted the family creation to "drawing emotion cards", which made it easier for the client to participate and express emotions intuitively. However, customizing family creations for different clients remains challenging, especially when emotional reactions are unpredictable. The service provider may not be able to accurately evoke the target emotion, affecting the treatment effect. ② The client was enthusiastic about the first graffiti creation, but due to the service provider's failure to provide sufficient encouragement between sessions, the client lost interest in the creation and eventually stopped. This example illustrates the importance of consistent encouragement in maintaining the implementation of family creations. Lack of support may reduce the client's motivation and affect the compliance of the creation. ③ The client was assigned emotion exploration family creations after each session, but due to the service provider's lack of systematic recording, it took a lot of time to search for the client's creation records each time, affecting the coherence of the session. This increases the workload of the service provider and also makes it difficult to integrate and analyze family creation data, failing to detect changes in the client's emotion management in a timely manner.
[0072] As can be seen from the above analysis, in the prior art, in the execution of art interaction tasks, the problems faced by service providers and visitors mainly include: Problem 1, facing challenges in executing art creation interaction tasks: Without the guidance of a service provider, therapeutic activities based on art creation may bring about creative obstacles; visitors may lack confidence in completing interaction tasks or having an emotional response due to the lack of guidance, resulting in reduced compliance; visitors may have difficulty interpreting works in a therapeutic manner without support, reducing the motivation for in-depth reflection; without a clear direction, it is difficult for visitors to create meaningful drawings to promote reflection. Problem 2, facing challenges in customizing therapeutic activities: Service providers usually customize interaction tasks based on practical experience and therapeutic techniques (such as cognitive behavioral therapy, mindfulness). In the current oral or written form, it is difficult to flexibly adjust structured instructions, resulting in visitors forgetting or abandoning the guidance; it is difficult to provide personalized encouragement to visitors outside of the communication time. Problem 3, facing challenges in tracking interaction history: Due to the difficulty in counting a large amount of historical interaction task data, it is difficult to track the interaction task history, which may affect the accuracy of service providers in formulating interaction tasks; using the existing face-to-face interaction method to execute interaction tasks is likely to miss valuable data on the emotional or mental state of visitors when they complete the interaction tasks.
[0073] Based on this, the present invention applies human-computer interaction technology to art interaction tasks and proposes an exemplary human-computer interaction system, which mainly consists of the front ends of the application programs for visitors and service providers, and the back end including an artificial intelligence model combination agent module, an AI art creation module, and a memory module.
[0074] For the front end: The front ends of the visitor and service provider application programs are both web-based user interfaces built using the Quasar framework, which is an open-source framework based on Vue.js. Vue.js is a progressive JavaScript framework for building user interfaces. The dialogue interaction of the application program for the end uses a text-to-speech model of an artificial intelligence company to process client messages, generate default auto-playing voice messages, and uses the existing speech dictation service to implement the speech-to-text function. The application program for the service provider interacts with the back-end server through hypertext transfer protocol requests, and the data is stored and transmitted in JSON format. The service provider can select visitors from the name list, view session logs, assign custom home creations, and modify the guidelines of the agent.
[0075] For the backend: It is implemented using Flask (Flask is a lightweight web application framework written in Python), responsible for generating artworks, handling the workflow of the AI model combination, and managing data. The AI art creation module includes a control network segmentation unit and an artwork generation AI unit. The control network segmentation unit is used to segment the colors, sounds, etc. of artworks, and the artwork generation AI unit. The system uses queue-based asynchronous processing and multi-threading to handle multiple concurrent artwork generation requests. The backend of the workflow of the AI model proxy module consists of five AI model proxies that use the same or different ones, mainly including: The first module is the session proxy, which is used during the discussion phase to interact with the visitor by combining the principles of the service provider and customized questions. The second module is the art creation input proxy, which is used during the art creation phase to convert the visitor's drawing and description into multi-modal input for the AI art creation module. The third module is the art creation generation proxy, which is used during the art creation phase to perform generative creation of AI artworks using the visitor's original artworks and the multi-modal input of the AI art creation module. The fourth module is the session record proxy, which is used to process the multi-modal data of the human-computer session of the visitor and generate summary and record information for easy retrieval and query for the service provider. The fifth module is the creation record proxy, which is used to integrate the multi-modal data in the entire art creation interaction process and provide relevant information and analysis of the visitor to the service provider. Additionally, the backend processes data storage and retrieval, manages user-generated content such as drawings, images, drawing logs, and conversation history. Images are stored in the Portable Network Graphics format, and logs and settings are saved in JSON format. All services use RESTful API (RESTful API is a web service interface based on the REST (Representational State Transfer) architectural style) for communication. Multiple APIs are set up for the backend service to ensure the smooth interaction between the front and back ends. The data storage module of the backend is used to store all data related to interaction tasks (including conversation principles, conversation content, original artworks, human-computer interaction artworks, and AI-compiled interaction result logs) in a database to facilitate tracking historical interaction records and provide references for formulating subsequent interaction tasks.
[0076] Meanwhile, the human-computer interaction system at least includes the following functional modules: an information collection module, which is used to collect in real time the original artworks and multimodal conversation information input by the visitor according to the current interaction task; an art creation module, which is used to dynamically analyze and process the original artworks and multimodal conversation information by using a number of artificial intelligence models to generate human-computer co-created artworks and conversation summaries; an interaction module, which is used to feedback the human-computer co-created artworks and the conversation summaries to the service provider, respond to the next interaction task generated by the service provider with the assistance of artificial intelligence based on the conversation summaries and the human-computer co-created artworks, and feedback the next interaction task to the visitor.
[0077] Furthermore, the human-computer interaction system of this embodiment can be installed on a variety of devices, allowing visitors to select different devices to access art interaction tasks in different natural environments, thereby reducing the difficulty of participating in the interaction.
[0078] Based on the above human-computer interaction system, the present invention proposes a human-computer interaction method based on asynchronous art creation, which is deployed on electronic devices such as computers and servers and is applied to scenarios where asynchronous interaction between a service provider and a visitor needs to be achieved through human-computer interaction, targeting situations where art creation needs to be completed in an interactive manner. The above scenarios can be various scenarios that require both parties to participate in the interaction, such as intelligent medical care, educational tutoring, elderly care, art creation teaching, etc. The present invention does not make specific restrictions on the application scenarios. Among them, the relationship between the service provider and the visitor includes objects such as medical staff and patients, art tutors and students, etc., who achieve interaction tasks through art creation and language dialogue. Specifically, as Figure 1 and Figure 2 shown, the steps of the method in this embodiment include:
[0079] Step S100: Collect in real time the original artworks and multimodal conversation information input by the visitor according to the current interaction task;
[0080] Specifically, during an interaction task, the human-computer interaction system first feeds back the current interaction task to the visitor. The visitor then successively inputs original artworks (such as two-dimensional paintings, two-dimensional animations, three-dimensional models, games, music, videos, etc.) and multimodal conversation information (such as text, voice, video calls, etc.) according to the content of the received current interaction task. The human-computer interaction system collects the original artworks and multimodal conversation information in real time. Among them, the current interaction task refers to an interaction task customized based on the result of the previous interaction. The original artwork refers to an artwork file generated by receiving and storing the artistic creation information input by the visitor through a brush, gesture, key, etc. using a system-defined or existing artistic creation software or platform. For multimodal conversation information, a speech recognition engine can be used to convert speech into text, a text analysis engine to extract key information, and a video analysis engine to capture actions and expressions in video calls. In particular, during the process of collecting the original artwork information and multimodal conversation information successively input by the visitor, the system needs to ensure the security of data by using a secure transmission protocol, and ensure the integrity, timeliness of data transmission, as well as the reliability and stability of the monitoring system. Further, during the entire collection and processing process, encryption technology is used to ensure the security and compliance of the transmitted and stored data to protect the privacy of the visitor. It should be noted that this embodiment does not specifically limit the type of interaction task, which can be presented in the form of an interaction task, or in the form of a questionnaire, etc.
[0081] Further, the multimodal conversation information includes at least one of audio data, video data, text data, artistic creation information, physiological information, and environmental information.
[0082] Specifically, since the multimodal conversation information includes one or several of audio data, video data, text data, artistic creation information, physiological information, and environmental information. Among them, the artistic creation information includes data information such as stroke pressure, thickness of painting lines, color, and shape during the artistic creation process using a brush, information such as movement amplitude and smoothness during the artistic creation process using gestures, information such as pressing force, frequency, and selection tendency during the artistic creation process using operation keys, as well as information such as head tracking, eye movement tracking, expression analysis, gesture information detection, posture detection, and relative displacement information of the visitor's body during the painting process in various painting forms; the physiological information includes information such as the visitor's heart rate, brain waves, body temperature, etc.; the environmental information includes information such as location and navigation information, light environment, sound environment, etc.
[0083] This embodiment expands new forms for the input of creative instructions, incorporating multimodal information into the system's analysis and the considerations of service providers. In addition to converting voice information into text, the emotional state of the user is analyzed through voice intonation to adjust the dialogue principles of the conversation agent. In addition to manual typing, text information also adds the direct input of handwritten text and handwritten text pictures into the conversation agent. By identifying factors such as the magnitude of the pen pressure, the approximation of the written characters to the standard font to judge the neatness, etc., it adapts to the habits and personalities of the visitors, and inputs the factors to the service provider side for the service provider to refer to and judge the personality of the visitors, and customize personalized interaction tasks or dialogue principles. The information on the magnitude of the pen pressure of the handwritten text is collected through a pressure sensor. Similarly, during the process of creating an artistic work, the touch screen operation situation can also reflect the above information; while on devices that are not operated through a touch screen, such as external input devices like a mouse and a keyboard, the usage situations, such as the key press frequency, key press duration, and time interval, can also depict the general image of the visitor during the creation process, such as relatively impatient, decisive, calm, or indecisive personality traits. Action information can provide a way for the system to judge the emotional state of the visitor other than the visitor's written self-report. Head tracking, eye movement tracking, facial expression analysis, gesture information detection, posture detection, and the relative displacement information of the visitor's body, etc., can all provide sufficient information for the system and service providers. Suppose during the text input process, the visitor inputs relatively positive content, while facial expression analysis identifies that the visitor's expression at this time is sad. The system learns about the situation of "saying one thing but meaning another", adjusts the dialogue principle of the conversation agent to be warm, prompts the service provider about the contradiction of the current multimodal information of the visitor to adjust the customized interaction task, and guides the visitor to face the true emotions and thoughts in their heart, which can provide a better artistic creation interaction effect for the visitor.
[0084] Physiological information can not only provide a similar effect to the input of action information, but also play a role in providing timely warnings during the artistic creation interaction process of specific visitors. For example, for visitors who are afraid to express themselves and have low self-confidence, when they are emotionally tense, the system identifies that the visitor's heart rate increases and body temperature rises, asks the visitor if they need to take a break, breathe deeply, or listen to soothing music, and timely inputs the visitor's emotional situation into the reminder window of the service provider. The service provider judges whether artificial intervention is required in the artistic creation interaction process. Environmental situation information can be used as an auxiliary means for the above information identification and judgment, such as location and navigation information, light environment, sound environment, etc. A certain visitor is creating an artistic work in a noisy environment. The sensor identifies that their heart rate increases and body temperature rises. If only physiological information is used for analysis, the system may issue an emotional warning prompt at this time. By combining the environmental sound situation and the previous personality trait of the visitor being easily shy, the system analyzes that the visitor is currently only in a strange and noisy environment, and will not issue a warning prompt, but records the multimodal information and recognition results for the subsequent review and reference of the service provider.
[0085] It is easily understandable that the multi-modal interaction information includes but is not limited to the above aspects. The key lies in that the fusion of various types of interaction information can provide more accurate, effective and comprehensive support information for the system to form a summary, judge the current state of the visitor, and provide intelligent prompts for the service provider to customize interaction tasks.
[0086] Step S200: Dynamically analyze and process the original art work and the multi-modal conversation information respectively by using an artificial intelligence model combination to generate a human-machine co-created art work and a conversation summary;
[0087] Specifically, according to the task requirements, select suitable artificial intelligence model methods, which have the capabilities of natural language processing and a preliminary understanding of the visitor's art creation process and phased achievements; for art work analysis, combine different modal inputs such as images and audio to extract key elements in art creation. According to the preset interaction rules or algorithms, guide each artificial intelligence model to dynamically analyze and process the collected original art work and multi-modal conversation information according to the task division of labor, including first digitizing the original art work to extract analyzable features (such as color, shape, texture, etc.), and preprocessing the multi-modal conversation information, such as speech recognition, text tokenization, noise removal, etc.; then identify the theme, style, emotion, etc. contained in the original art work according to the features of the original art work, and identify the theme, emotion, key points, behavior pattern recognition, etc. contained in the multi-modal conversation information according to the speech, text or expression obtained by preprocessing; finally, use the artificial intelligence model to combine the recognition results of the original art work and the multi-modal conversation information, comprehensively analyze the context relationship in the interaction process, understand the intentions and expectations of the visitor, generate a human-machine co-created art work that can match the content of the original art work and the intentions of the visitor, and extract the main viewpoints, emotional tendencies and key information, etc. in the multi-modal conversation information to generate a concise and clear conversation summary and record. Further, short explanations or background information can also be provided for the human-machine co-created art work and the conversation record to help the service provider better understand.
[0088] Step S300: Feed back the human-machine co-created art work and the conversation record to the service provider, respond to the next interaction task formulated by the service provider based on the conversation record and the human-machine co-created art work with the assistance of artificial intelligence, and feed back the next interaction task to the visitor.
[0089] Specifically, the human-computer interaction system sends the co-created artworks and conversation records between humans and machines to the service provider, along with a short description or request, to remind the service provider of the confidentiality of the materials and inform the basic situation and needs of the visitor. After receiving the current interaction information feedback by the visitor, based on the conversation records and co-created artworks in the current interaction information, as well as the pre-recorded personal information and communication situation of the visitor, the service provider determines the current mental, emotional, behavioral, and / or cognitive status of the visitor with the assistance of artificial intelligence, formulates personalized interaction principles based on the status of all aspects, and feeds back the interaction principles and interaction methods to the human-computer interaction system. With the assistance of artificial intelligence, the next interaction task is generated and fed back to the visitor, so that the visitor can execute the next interaction task according to the process of executing the current interaction task under the guidance of the human-computer interaction system, thereby meeting the asynchronous interaction needs between the visitor and the service provider.
[0090] In this embodiment, the asynchronous collaboration between the service provider and the visitor is realized through the human-computer interaction system, enabling the service provider and the visitor to maintain interaction between two meetings and helping the visitor establish long-term trust and intimacy. At the same time, under the guidance of the artificial intelligence model, the difficulty of the visitor executing the art creation task is significantly reduced, which helps to improve the visitor's creation and expression abilities and the effect of asynchronous interaction.
[0091] In a preferred implementation manner, the real-time collection of the original artworks and multimodal conversation information input by the visitor according to the current interaction task in step S100 includes:
[0092] Step S110: Construct a series of interaction questions based on the conversation principles and example questions corresponding to the current interaction task;
[0093] Specifically, after receiving the current interaction information feedback by the visitor, the service provider determines the current mental, emotional, behavioral and other aspects of the visitor based on the conversation summary, the human-machine co-created art work, as well as the personal information and communication situation of the visitor, and formulates personalized conversation principles and example questions by synthesizing the conditions of all aspects, and feeds back the conversation principles and example questions to the human-computer interaction system. The human-computer interaction system customizes the interaction process according to the conversation principles and example questions, and generates a series of interaction questions according to the interaction process, so as to guide the visitor to successfully complete the art creation process. Among them, the conversation principle refers to the key interaction content adopted to solve the visitor's appeal to participate in the creation activity; the example question refers to the guiding question corresponding to the interaction process set for completing the key interaction content in the conversation principle. For example, if the current interaction task is to let the visitor draw a two-dimensional painting with a paintbrush, the conversation principle can be set to inform the visitor of the content of the art creation, guide the visitor to name their art work, guide the visitor to describe their feelings about the art creation, etc. Then the example questions can be set as follows: 1) Dear user, please draw the spring in your mind (for example, the visitor draws clouds and trees on the electronic painting screen with a paintbrush); 2) What shapes and colors do you hope the clouds and trees to be?; 3) What do you hope to have on the trees?; 4) Please name this painting; 5) Please describe your feelings about creating this painting.
[0094] Step S120: Feed the interaction questions to the visitor in sequence, and collect the original art works and multimodal conversation information corresponding to each of the interaction questions feedback by the visitor in sequence.
[0095] Specifically, the setting order of the interaction questions of the human-computer interaction system gradually guides the forms of art creation, the types of art creation, and the content of art creation that the visitor can adopt, etc., and according to the original art works and multimodal conversation information corresponding to each of the interaction questions feedback by the visitor during the art creation process, gradually guides the visitor to introduce the content and mental state of their art creation, laying a foundation for the subsequent human-computer interaction system to deeply understand the mental, emotional and behavioral states and tendencies of the visitor. Among them, the forms of art creation include but are not limited to manual drawing, gesture actions, button operations, etc., and the types of art creation include but are not limited to two-dimensional painting, two-dimensional animation, three-dimensional model, interactive game, music, video, etc. The content of art creation includes but is not limited to content of types such as human figures, landscapes, animals, still lifes, scenes, etc.
[0096] In a preferred implementation manner, the dynamic analysis and processing of the original art work and the multimodal conversation information respectively by using the artificial intelligence model combination in step S200 to generate a human-machine co-created art work and a conversation summary and record includes:
[0097] Step S210: Analyze the multimodal conversation information using the first module in the artificial intelligence model combination to generate several types of conversation instructions;
[0098] Specifically, since the multimodal conversation information may include one or several of various types of data such as audio data, video data, text data, artistic creation information, physiological information, and environmental information, and the processing methods for different data are different, in this embodiment, the information other than the artistic creation information in the multimodal conversation information is subjected to standardization processing and / or data conversion to obtain data in a target format, so as to eliminate the dimensionality influence between different types of data, make different data indicators comparable, and be conducive to improving the accuracy of comprehensive evaluation. For example, for standardization processing, the audio signal is normalized by scaling the amplitude of the audio signal to a fixed range (such as [-1, 1]) to eliminate the amplitude difference under different recording conditions and improve the effectiveness of feature extraction; for the color information in video data, the color information is adjusted to a unified range to make it consistent under different lighting conditions and eliminate the color difference under different devices or shooting conditions; for text data, the norm of the text vector is normalized so that vectors of different lengths can be compared on the same scale; for the light environment information in the environmental information, features such as illuminance and color temperature are calculated and normalized to eliminate the difference under different light conditions. For data conversion, speech is converted to text or text is converted to speech to implement the function of speech and text conversion in the artistic creation dialogue agent window.
[0099] Then, the first module in the artificial intelligence model combination is used to analyze and reason the data in the target format to generate several types of conversation instructions. For example, if the human-computer interaction system extracts sad emotional information from the expression of the visitor in the video data uploaded by the visitor, then the human-computer interaction system will generate a conversation instruction for asking the visitor what thing or person makes them sad; another example is that if the human-computer interaction system extracts pleasant emotional information from the text data uploaded by the visitor, then the human-computer interaction system will generate a conversation instruction for asking the visitor what thing they most want to do or what person they most want to see.
[0100] Step S220: Feed back the conversation instructions to the visitor, and in response to receiving the feedback information of the visitor on the conversation instructions, generate several types of conversation information;
[0101] Specifically, the session instructions are fed back to the visitor, and the visitor makes verbal or non-verbal expressions under the guidance of the session instructions. In each interaction task, the AI model constructs multiple session instructions in the form of questions and answers according to the corresponding conversation principles and sample questions of the current interaction task, so as to gradually guide the visitor to actively complete the interaction task, and help the visitor regulate emotions, express themselves, and solve the visitor's confusion.
[0102] Step S230: Integrate all the session information to generate a session summary and record.
[0103] Specifically, in each interaction task, multiple Q&A contents of the communication between the AI model and the visitor are integrated in chronological order or by topic type, etc., to generate a complete session content, and key information in the complete session content is extracted to generate a session summary, so that the service provider can quickly and accurately grasp the mental, emotional, behavioral, and / or cognitive state information of the visitor in this interaction task.
[0104] In a preferred implementation manner, the dynamic analysis and processing of the original art work and the multi-modal session information respectively by using the AI model combination in step S200 to generate a human-machine co-created art work and a session summary and record includes:
[0105] Step S240: Use the second module in the AI model combination to analyze the art creation information in the multi-modal session information to generate art creation instructions and art creation records;
[0106] Specifically, use the second module in the AI model combination to analyze the art creation information in the multi-modal session information. For example, by analyzing data information such as the brush stroke pressure, line thickness, color, and shape of the brush, understand the visitor's creative content, style, and intention; by identifying information such as the amplitude and smoothness of the gesture movements, infer the visitor's creative process and the emotions that may be expressed; by analyzing information such as the pressing force, frequency, and selection tendency, understand the visitor's creative decisions and preferences; by parsing the text descriptions related to art creation, understand the theme and style of the art work created by the visitor, etc.
[0107] Then, the second module generates a series of artistic creation instructions based on the analysis results of the artistic creation information in the multimodal conversation information to guide the visitor or assist the human-computer interaction system to continue or improve the artistic creation process. For example, it is recommended that the visitor adjust the creation style to better conform to a specific theme or requirement; based on the analysis of gestures and expressions, it is recommended that the visitor better express a specific emotion or atmosphere in the work; the feedback from the visitor is converted into specific creation instructions to guide the visitor to make modifications or improvements, etc. At the same time, the second module extracts the artistic creation summary of the visitor during the entire artistic creation process based on the analysis results of the artistic creation information in the multimodal conversation information to outline the theme, style, composition, color and other characteristics of the artistic works created by the visitor, and analyze and summarize the emotions and intentions expressed by the visitor during the creation process.
[0108] Step S250: Feed back the artistic creation instructions to the visitor, and use the third module in the artificial intelligence model combination to dynamically analyze and process the artistic creation information and the artistic creation summary fed back by the visitor according to the artistic creation instructions to generate a human-machine co-created artistic work.
[0109] Specifically, the artistic creation instructions are fed back to the visitor. In response to receiving the artistic creation information fed back by the visitor according to the artistic creation instructions, the artistic creation interface information is obtained. If the artistic creation is digital, the image data can be directly extracted as the artistic creation interface information; if the artistic creation is physical, it can be converted into a digital image by taking a photo or scanning, and then the image data is extracted. Then, the third module in the artificial intelligence model combination is used to identify the color and semantic information in the artistic creation interface information, and perform color and semantic segmentation on the artistic creation interface information to obtain information such as the shape, position, and size of several color and semantic segmentation. Finally, based on all the artistic creation summaries and the description information about the theme, style, emotion, and intention of the work provided by the record, and all the color and semantic chunk information providing the detailed information about the color and semantic segmentation of the work, the third module is used for generative creation to create a human-machine co-created artistic work that matches the above input information.
[0110] In a preferred embodiment, the generation process of the next interaction task in step S300 includes:
[0111] Step S310: Collect historical interaction results within a preset time period;
[0112] Specifically, collect the personal information of the visitor, such as height, weight, name, email, contact information, etc.; collect the number of times the visitor performs interaction tasks within a preset time period, that is, by setting up a logging system or database, record the number of times the visitor performs various interaction tasks within the preset time period; collect the session summaries and records of each interaction task performed by the visitor within the preset time period, and use the fourth module to analyze the situation of the visitor performing interaction tasks, generating summary information compiled by AI, including information such as task type, completion rate, error rate, and summary of interaction content; and integrate the personal information, the number of times, original artworks, human-machine co-created artworks, and / or session summaries to generate historical interaction results, thereby forming a complete visitor interaction dataset or historical interaction result report to achieve the evaluation of indicators such as visitor activity, task completion efficiency, and AI assistance effect. Further, intuitively display the historical interaction results in the form of charts, dashboards, etc. to facilitate in-depth understanding and improve the evaluation accuracy.
[0113] Step S320: Automatically generate the current interaction result based on the session summary and the human-machine co-created artwork.
[0114] Specifically, identify several key time nodes in the current interaction task and use these nodes as reference points for collecting original artworks; collect the original artworks, human-machine co-created artworks, and session summaries corresponding to each target time node in the current interaction task; integrate the original artworks, human-machine co-created artworks, and session summaries into a unified dataset to ensure the integrity and interconnection of the data for each time node; use technologies such as image recognition and text analysis to deeply analyze the artworks and session summaries, extracting key features such as style, theme, and emotion; analyze the change trends of artworks and session content at different time nodes to identify the visitor's creative preferences, the effect of AI assistance, and potential problems during the creation process; based on the integrated data and analysis results, comprehensively evaluate the original artworks, human-machine co-created artworks, and session summaries for each time node, and generate a report or summary of the current interaction result according to the evaluation results, including the advantages and disadvantages of the artworks, the effectiveness of human-machine co-creation, and the highlights of the visitor's feedback. Further, use forms such as charts, images, and videos to intuitively present the interaction result to the visitor and the service provider to facilitate a better understanding of the entire interaction process and the final outcome. It should be noted that the current interaction result only refers to the result generated based on the original artworks, human-machine co-created artworks, session summaries, etc. of the visitor corresponding to the current interaction task, while the historical interaction result is the comprehensive evaluation result of the multiple task results completed by the visitor corresponding to the current interaction task in history.
[0115] Step S330: In response to receiving at least one piece of information among the current interaction result, the human-machine co-created art work, and the historical interaction result from the service provider, generate a next interaction task with the assistance of artificial intelligence.
[0116] Specifically, the human-machine interaction system, in response to receiving at least one piece of information among the current interaction result, the human-machine co-created art work, and the historical interaction result from the service provider (even combined with the service provider's experience), deeply understands the changes in aspects such as the emotions, behaviors, and cognitions of the visitor contained in the current interaction result, the depth of emotion expressed in the art work, the improvement of interpersonal interaction, the elements, color application, and composition layout in the human-machine co-created art work, etc., to understand the physical and mental states of the visitor such as emotions and cognitions. At the same time, by reviewing the historical interaction results, analyze the change trends of the visitor within the target time period, including progress, challenges, and recurring themes, and compare the current interaction result with the historical data to find connections and differences, so as to guide the service provider to generate the next interaction task. To achieve an effective feedback mechanism, ensure that the visitor can obtain timely guidance and support from the service provider, and enable the service provider to closely observe the performance of the visitor, record the key events and the visitor's reactions during the task execution process. It is beneficial for the service provider to flexibly adjust the task content, difficulty, or interaction method according to the task execution situation and the real-time feedback of the visitor, to ensure the effectiveness and safety of the interaction process.
[0117] In a preferred implementation manner, generating a next interaction task with the assistance of artificial intelligence in step S330 includes:
[0118] Step S331: In response to receiving the session reflection content fed back by the visitor according to the current interaction result, generate an optimized conversation principle and optimized example questions;
[0119] Step S332: Based on the optimized conversation principle and the optimized example questions, generate a next interaction task with the assistance of artificial intelligence.
[0120] Specifically, in this embodiment, through asking or the initiative feedback of the visitor on the session reflection content of the current interaction result, an artificial intelligence model is used to identify the emotions and needs expressed by the visitor in the reflection, such as seeking more understanding, hoping to obtain more specific guidance, hoping for a smoother conversation, etc., so as to generate optimized conversation principles and optimized example questions, making the optimized conversation principles clearer, more concise, easier for the visitor to understand and follow, encouraging the visitor to actively participate in the conversation to enhance the interactivity and effectiveness of the conversation; making the optimized example questions better guide the visitor to think deeply and express. Then, with the assistance of artificial intelligence, the optimized conversation principles and optimized example questions are integrated into the next interaction task to perform real-time update and flexible adjustment on the next interaction task, thereby improving the quality and effect of the conversation and better meeting the needs and expectations of the visitor.
[0121] Furthermore, the artificial intelligence model can be fine-tuned or retrained according to the historical interaction results within a period of time to improve its accuracy and efficiency. Or, more training data, especially data related to artworks and conversation information, can be introduced to enhance the generalization ability of the model, and the optimized model is applied to new artworks and conversation information, continuously iterating and optimizing the entire process to improve the quality and efficiency of human-machine co-created artworks and conversation summaries.
[0122] It should be noted that in this embodiment, whenever personal information-related data of visitors or service providers, etc. is applied to specific products or technologies, it is used under the condition of obtaining user permission or consent, and the collection, use, and processing of relevant data comply with relevant standards.
[0123] The following provides an exemplary embodiment of a human-computer interaction system and method, specifically as follows:
[0124] To support the art creation interaction task of the visitor and the service provider-visitor cooperation around it, the human-computer interaction system designed in this embodiment includes: an application program for the visitor, which combines the art creation co-created by humans and artificial intelligence with dialogue interaction to promote the interaction task in the daily environment. An application program for the service provider, which provides the historical record of the AI-compiled interaction task and the customization of the visitor interaction task agent to obtain customized guidance, see Figures 3 - 6 .
[0125] To address three key challenges existing in the prior art, the core design functions of the human-computer interaction system provided in this embodiment are specifically as follows:
[0126] First, combine human-artificial intelligence co-created art creation with dialogue interaction (application program for the visitor). To solve Problem 1, the application program for the visitor utilizes the human-AI co-creation canvas (such as Figure 3) to lower the threshold of art creation for visitors and provide structured guidance to visitors in the "Art Creation Stage" and "Discussion Stage" using a two-stage dialogue workflow. As Figure 3 shown, the application for visitors has three panels: the AI Brush Window 305, the User Creation Canvas 303, and the AI Generated Image Window 304. Among them, the AI Brush Window 305 includes various creation tools that utilize the capabilities of AI to generate objects with matching shapes by creating semantic brush elements such as color blocks and geometric bodies; the User Creation Canvas 303, a canvas for visitors' drawings using color segmentation; the AI Generated Image Window 304, used to preview human-AI co-created artworks. The Art Creation Stage dialog box includes a Painting Description Editing Area 301 and an Art Creation Dialogue Agent Window 302, which are used to extract keywords for generating paintings and prompt visitors to verbally describe the work in detail. The Discussion Stage dialog box includes a Figure 4 Discussion Stage Dialogue Agent Window 401 as shown, which is used to facilitate visitors' in-depth self-exploration and reflection according to the task instructions. The dialogue principles and example questions for the initial task guidance are from existing art interaction programs and are refined through the feedback of professional art service providers.
[0127] Second, support for customizing the client creation agent (application for service providers).
[0128] To solve Problem 2, the application for service providers has three panels for customization, as Figure 6 shown, including: the AI Dialogue Principle Customization Area 603, the Interaction Task Customization 601, and the Personal Message Customization Area 602. Among them, the AI Dialogue Principle Customization Area 603: allows service providers to customize dialogue principles, modify example questions, and adjust the dialogue process based on practical experience. The Interaction Task Customization Area 601: used to support service providers in assigning interaction tasks. The Personal Message Customization Area 602: tailors personal information for visitors and provides encouragement and emotional support in art therapy interaction tasks.
[0129] Third, enable the AI-compiled interaction task history (application for service providers).
[0130] To solve Problem 3, the application for service providers also has the following panels: the Interaction Task Overview Area, the Interaction Task Record Area for each session, and the AI Compiled Summary Area. As Figure 5 shown, the Interaction Task Overview Area includes the Personal Information Area 501, the Usage Log Area 502, and the Summary of Created Works Area 503. Among them, Figure 5The personal information area 501 in it is hidden to avoid involving user privacy. In actual applications, this part of the interface will display the personal information of the visitor. The interaction task record area for each section includes the original creative work area 504, the AI-generated creative work area 506, the creative process description area 505, and the AI conversation record area 507. AI Compilation Summary: Used to summarize the image description, feelings, and experiences of the visitor, providing insights for the service provider.
[0131] The system provided based on the exemplary embodiments below implements the following four different types of interaction tasks including but not limited to the following using the method of the present invention, specifically as follows:
[0132] Embodiment 1: An interaction task of creating a two-dimensional art work using voice and text.
[0133] The creative tools included in the AI paintbrush window 305 are: paintbrush elements with semantics: ocean, cloud, earth, lawn, mountain, tree, hill, lake, person, wall, road, etc. Drawn in the form of color blocks; paintbrush thickness adjustment bar; eraser tool; color palette; painting style selector: watercolor, oil painting, illustration, etc. Details can be adjusted during the interaction through the art creation dialogue agent window 302.
[0134] This embodiment is based on the visitor's initial two-dimensional hand-drawn work. The art creation dialogue agent window is used to input detailed descriptions, and the target two-dimensional image art work is generated through multiple iterations. The specific steps are as follows:
[0135] In the previous communication, the service provider learned that the visitor currently has certain anxiety emotions in the marital relationship and customized a task of creating a two-dimensional image art work for the visitor. The image should include two plants, one representing herself and the other representing her partner. Through the art creation dialogue agent window 302, the visitor can understand the task of this interaction task.
[0136] In the startup stage, the visitor selects the corresponding "paintbrush element with semantics". In this embodiment, the visitor selects the "tree" paintbrush element and then draws the shape of a tree on the user creation canvas 303; the color palette is used to select the paintbrush color, and the paintbrush thickness adjustment lever is used to adjust the paintbrush thickness; for the parts drawn wrongly or to be cancelled, the visitor uses the eraser tool to erase them.
[0137] After the system recognizes the drawing traces on the user's creation canvas 303, the dialogue agent prompts the visitor to describe the specific details of the elements through voice input, and this prompt is displayed in the art creation dialogue agent window 302. At this time, it enters the detail adjustment stage. The visitor describes the shape of the painting as two apple trees in full bloom through voice input. The dialogue agent generates a summary of "apple trees, blooming, two" based on the input and displays it on the painting description editing area. The visitor sequentially selects the brush elements of "mud", "cloud", and "lawn" and smears them on the user's creation canvas 303 to add more details. For the parts drawn wrongly or cancelled, the visitor uses the eraser tool to erase them. After completing the detail supplement, the visitor selects the watercolor style option on the painting style selector.
[0138] Click to generate and enter the art work preview stage. In the AI-generated image window 304, the original art work generated at this stage is displayed. The visitor uses voice input and creation tools to generate a human-machine co-created art work based on the original art work using an artificial intelligence model. After adjusting the human-machine co-created art work multiple times, until the visitor sees a work that meets the expectations in the AI-generated image window 304, click OK to export the final human-machine co-created art work and generate the interactive result compiled by AI to end this interaction task.
[0139] Example 2: An interactive task of creating an animated art work using voice and text.
[0140] The creation tools included in the AI brush window 305 are: brush elements with semantics: ocean, cloud, earth, lawn, mountain, tree, hill, lake, person, wall, road, etc. Drawn in the form of color blocks; brush thickness adjustment bar; eraser tool; color palette; painting style selector: watercolor, oil painting, illustration, etc. Detail adjustment is realized during the interaction through the art creation dialogue agent window 302.
[0141] Object selector: point selection (click on the color block), range selection, select all, invert selection, deselect, screen selection, etc., which is convenient for adding actions or special effects to a single part later; action brush: basic movements: move, rotate, scale, jump; deformation effects: stretch and compress, distort, wave effect; loop and interaction: repeat action, path movement; special effect brush: light and shadow category: glow, shadow; particle category: smoke, spark; weather and environment: raindrop, snowflake; dynamics: blur, jitter; color category: filter, flicker; space: virtual depth, starry sky; editing and playback tools: timeline editor: track manager; keyframe editor; audio editor: input dialogue and sound effects, editable. Detail adjustment is realized during the interaction through the art creation dialogue agent window 302.
[0142] This embodiment is based on the visitor's initial 2D animation work. Details are input through the art creation dialogue agent window 302, and the target animation art work is generated through multiple iterations. The specific steps are as follows:
[0143] In the previous communication, the service provider learned that the visitor is currently experiencing seasonal affective disorder and customized a task of creating an animation art work for the visitor. The animation should include four elements representing spring, summer, autumn, and winter and be able to show the environmental changes in different seasons. Through the art creation dialogue agent window, the visitor can learn about the task of this interaction.
[0144] In the initial stage, the visitor selects the corresponding "semantic brush elements". In this embodiment, the visitor selects the brush elements of "tree", "flower", "fallen leaf", and "snowflake", and then paints the shapes of trees, flowers, fallen leaves, and snowflakes on the user creation canvas; selects the brush color using the color palette and adjusts the brush thickness using the "brush thickness adjustment lever"; for the parts painted wrongly or cancelled, the visitor uses the eraser tool to erase them. Basic motion action brushes such as "move" and "rotate" are selected and applied to these objects, and deformation motion action brushes such as "stretch" and "compress" are selected and applied to the fallen leaf object; loop and interaction action brushes such as "repeat action" are selected and applied to all objects.
[0145] After the system recognizes that there are drawing traces on the user creation canvas 303, the dialogue agent prompts the visitor to describe the specific details of the object through voice input, and this prompt is displayed in the art creation dialogue agent window 302. At this time, it enters the detail adjustment stage. The visitor describes the painted shapes as a tree slowly swaying in spring, a flower gently shaking in summer, a pile of fallen leaves flying in autumn, and a snowflake floating in winter through voice input; the dialogue agent generates a summary of "trees in spring, flowers in summer, fallen leaves in autumn, snowflakes in winter" according to the input and displays it in the painting description editing area; the visitor sequentially selects the "glow" special effect brush of the light and shadow category and applies it to the tree element, the "smoke" special effect brush of the particle category and applies it to the flower element, the "raindrop" special effect brush of the weather and environment category and applies it to the fallen leaf object, and the "jitter" special effect brush of the dynamic category and applies it to the snowflake element to add special effect details; after completing the detail supplement, the visitor selects the realistic style option on the painting style selector.
[0146] Entering the animation editing stage, the visitor opens the timeline editor. Uses the "track manager" to arrange the actions and special effects of each element to appropriate time nodes. The visitor uses the "keyframe editor" to finely adjust the duration and transition effect of each action to make the process of seasonal changes smoother.
[0147] In order to make the work more expressive, visitors added background music in the "audio editor" and added natural sounds of seasonal changes, such as birds singing in spring and cicadas chirping in summer, and other "sound effects" to enhance the integration of vision and hearing with the pictures.
[0148] Click Generate to enter the artwork preview stage. Visitors use the playback controller to preview the animation effect of the original artwork. In the AI generated image window 304, the original artwork animation generated at this stage appears; based on the original artwork, visitors adjust the AI generated animation multiple times through voice input, creation tools and editing tools. After the visitor uses the playback controller in the AI generated image window 304 to see the expected work, click OK to export the final human-computer co-creation artwork, generate the AI-compiled interaction results, and end this interaction task.
[0149] Example 3: An interactive task of creating a three-dimensional model artwork using voice and text.
[0150] The creation tools included in the AI brush window 305 are: brush elements with semantics: preset models: ocean, cloud, earth, lawn, mountain, tree, hill, lake, person, wall, road, etc.; three-dimensional primitives: cube, cylinder, cone, sphere, etc.; definition brush: ocean, cloud, earth, lawn, mountain, tree, hill, lake, person, wall, road, etc., the definition brush defines the object properties for the established three-dimensional primitives and their combinations; three-dimensional model design tools include: model parameter adjustment bar: parameter display: center coordinates, three-axis dimensions; adjustment lever and value input box: three-axis scaling lever, center adjustment lever, rotation lever; model combination options: combination: alignment, group; Boolean operation: union, difference, intersection, exclusion; deformation tools: stretching, twisting, squeezing; material and texture color palette. Detail adjustment is achieved during the interaction process through the art creation dialogue agent window 302.
[0151] This embodiment is based on the initial creation of the 3D model work by the visitor, and uses the art creation dialogue agent window 302 to input detailed descriptions and iterate multiple times to generate the target 3D model art work. The specific steps are as follows:
[0152] In the previous communication, the service provider learned that the visitor was currently experiencing emotional stress, and customized a task for the visitor to create a three-dimensional model artwork. The animation must include mountains, lakes and lawns to form a natural landscape model. Through the art creation dialogue agent window 302, the visitor can understand the task of this interactive task.
[0153] In the initial stage, the visitor selected the brush element with semantics in the creation tool. In this embodiment, the visitor selected the "mountain peak" and "lawn" brush elements in the preset model, and then placed the shapes of the mountain peak and the lawn on the user creation canvas; selected the "ellipsoid" brush element in the three-dimensional basic body and named the object "lake"; used the "material and texture palette" to select the model color; and formed the initial scene.
[0154] Entering the three-dimensional model adjustment stage, the visitor used the "three-axis scaling lever" to adjust the size of the lawn; used the "center adjustment lever" to adjust the position of the lake; and used the "union" function of Boolean operation to combine multiple mountain peaks into one model.
[0155] After the system recognized that there was a model with material or texture in the user creation canvas 303, the dialogue agent prompted the visitor to describe the specific details of the object through voice input, and this prompt was displayed in the art creation dialogue agent window 302. At this time, the detail adjustment stage was entered. The visitor described the shape of the model as a very high continuous snow mountain, a clean lake, and a green lawn through voice input; the dialogue agent generated a summary of "continuous snow mountain, lake, green lawn" according to the input and displayed it in the painting description editing area; the visitor used the "stretch" and "twist" functions in the deformation tool to refine the shape of the mountain peak; and used the "extrusion tool" on the lake model to increase the water depth performance. After completing the detail supplement, the visitor selected the realistic style option on the painting style selector.
[0156] Clicking Generate, the art work preview stage is entered. The visitor uses the three-dimensional view observation bar to preview the scene effect of the original art work. In the AI-generated image window 304, the three-dimensional model generated at this stage appears. Based on the original art work, the visitor adjusts the scene of the AI-generated three-dimensional model through voice input, creation tools, and three-dimensional model design tools multiple times. After seeing the work that meets the expectations in the AI-generated image window 304 using the three-dimensional view observation bar, the visitor can click OK to export the final human-computer co-created art work and generate the interactive result compiled by AI, ending this interaction task.
[0157] Embodiment 4: The interaction task of creating a three-dimensional interactive scene art work using voice and text.
[0158] The creative tools included in the AI painting brush window 305 are as follows: Painting brush elements with semantics: preset models: ocean, cloud, earth, lawn, mountain, tree, hill, lake, person, wall, road, etc.; 3D basic solids: cube, cylinder, cone, sphere, etc.; Defined painting brushes: ocean, cloud, earth, lawn, mountain, tree, hill, lake, person, wall, road, etc., and the defined painting brushes define object attributes for the established 3D basic solids and their combinations; 3D model design tools: Model parameter adjustment bar: Parameter display: central coordinates, three-axis dimensions; Adjustment levers and numeric input boxes: three-axis scaling lever, central adjustment lever, rotation lever; Model combination options: Combination: alignment, grouping; Boolean operations: union, difference, intersection, exclusion; Deformation tools: stretching, twisting, extrusion; Material and texture palettes. Action painting brushes: Action application painting brushes: Basic movements: moving, rotating, scaling, jumping; Deformation effects: stretching and compression, twisting, wave effects; Looping and interaction: repeating actions, path movement; Skeleton binding selection bar and character action addition. Special effect painting brushes: Collision and strike effects: flash effects, vibration and shock waves, debris splashing, sparks and smoke; Physical reaction special effects: falling effects, bouncing effects, deformation animations; Visual feedback: color change, high-light contour, transparency; Particle effects: fire and explosion, smoke and dust, magic special effects; Sound and vision combination: sound effect synchronization, screen flickering; Dynamic change special effects: size change, rotation and displacement, splitting and combination; Energy and light effects: energy flow, halo and flash; Disappearance and generation special effects: dissolution special effects; Portal effects; UI prompts: floating text, arrows or markers; Trigger event special effects: beam connection, chain reaction. Audio sound effects: Background music application bar: preset music / local music / AI music creation; Sound effect painting brushes: define interactive sound effects for the established 3D basic solids and their combinations; Audio editor: input dialogue and sound effects, editable. Details can be adjusted during the interaction through the art creation dialogue agent window 302.
[0159] This embodiment is based on the three-dimensional interactive scene work initially created by the visitor. The art creation dialogue agent window 302 is used to input detailed descriptions, and the target three-dimensional interactive scene art work is generated through multiple iterations. The specific steps are as follows:
[0160] In the previous communication, the service provider learned that the visitor is currently experiencing a confused period and wants to seek inner peace and a sense of belonging. The service provider customized a task of creating a three-dimensional interactive scene art work for the visitor. The scene should include mountains, forests, and houses to form a dynamic interactive scene. Through the art creation dialogue agent window 302, the visitor can understand the task of this interaction.
[0161] In the initial stage, the visitor selected the brush elements with semantics in the creation tools. In this embodiment, the visitor selected the brush elements of "mountain", "tree", "medieval folk house" and "muddy path" in the preset model, and then placed the shapes of mountains, trees, houses and roads on the user creation canvas 303; selected the model colors using the "material and texture palette"; and formed the initial scene.
[0162] Entering the three-dimensional model adjustment stage, the visitor used the "three-axis scaling lever" to adjust the size of the mountain; used the "center adjustment lever" to adjust the positions of the tree, house and road; used the "union" function of Boolean operation to combine multiple trees into one model to form a forest; and used "rotation" to adjust the orientation of the road.
[0163] Entering the dynamic effect adding stage, the visitor used the "move" tool in basic motion to add actions to the trees in the forest to simulate the effect of the trees being blown by a gentle breeze; and used "smoke and dust" in particle effects to add cooking smoke on the user creation canvas 303.
[0164] After the system recognizes that there is a model with material or texture and dynamic effects in the user creation canvas 303, the dialogue agent prompts the visitor to describe the specific details of the object through voice input, and this prompt is displayed in the art creation dialogue agent window 302. At this time, it enters the detail adjustment stage. The visitor describes the scene as a big mountain, a lush forest, an ancient village, a muddy path between the village and the mountain, moist air, and mysterious light on the mountaintop through voice input, and automatically adds other details suitable for the scene; the dialogue agent generates a summary of "big mountain, forest, muddy path, fog, mysterious light, other details" according to the input and displays it in the painting description editing area; after completing the detail supplement, the visitor selects the option of oil painting style on the painting style selector. Click Generate, and the visitor can use the three-dimensional view observation bar to preview the dynamic scene.
[0165] After the system recognizes that the preview scene has been generated, the dialogue agent prompts the visitor to create an interaction subject and interaction effects, and at this time it enters the interaction design stage.
[0166] First, create the interaction subject, that is, the role played by the visitor. The visitor selects the "farmer" role in the preset model; adds the preset bone binding to the role in the "bone binding selection bar"; uses "role action addition", including "walking", "running", "jumping", "raising the hand", "sitting down"; and sets the role interaction action to "raising the hand".
[0167] Secondly, set the scene interaction effects. The visitor selects the "Rotation" tool in the action brush and applies it to the doors and windows of the house, connecting to interaction key one; selects "Smoke and Dust" in the particle effect and applies it to the muddy path. When the interactive subject, i.e., the player's character, passes by, a dust - rising effect can be produced on the road surface; selects "Flame and Explosion" in the particle effect and applies it to the trees and the combustible items such as firewood added by the AI during the detailed adjustment stage, connecting to interaction key two for releasing magic; connects the actions of the trees to the "running" and "jumping" actions of the character and the character model's coordinate position, and the action effect is enhanced when triggered, so that when the character runs or jumps in the forest, the shaking degree of the trees is greater.
[0168] To make the work more expressive, the visitor adds background music in the "Audio Editor"; the visitor adds sound effects to the action keys of the character, such as the sound effects of walking, running, jumping, and releasing magic; adds sound effects to the interactive actions of the scene, such as the sound effects of the trees shaking slightly and violently, the sound effect of items burning, and the sound effect of opening and closing doors and windows; coordinates with the picture to enhance the integration of vision and hearing.
[0169] Click to generate and enter the preview stage of the art work. The visitor uses the action keys in the perspective switching and scene interaction bar to preview the three - dimensional scene display effect of the original art work. In the AI - generated image window 304, the three - dimensional interactive scene generated at this stage appears; the visitor adjusts the three - dimensional interactive scene generated by the AI multiple times through voice input, creation tools, three - dimensional model design tools, interactive design tools, and audio sound effect editing tools. After the visitor sees the work that meets the expectations using the "action key" and "perspective switching option" in the scene interaction bar in the AI - generated image window 304, click OK to export the final human - machine co - created art work and generate the interactive result compiled by the AI, ending this interaction task.
[0170] Finally, for the method of the present invention, a three - party interaction process for an online art interaction task is provided as follows:
[0171] Reason description: The visitor has been trying to deal with the emotional problems with the partner, but feels misunderstood by the partner.
[0172] Execution steps:
[0173] As Figure 7 shown in Phase 1: Pre - communication (basic situation communication, personalized problem customization, support information customization).
[0174] 1. The visitor explains the personal basic situation of the problems encountered to the service provider;
[0175] 2. The service provider incorporates the dialogue principles and sample questions into the system to guide the dialogue agent to ask the correct questions;
[0176] 3. The service provider customizes questions according to the individual needs of the visitor: Customizes interaction tasks related to the marital relationship;
[0177] 4. The service provider provides emotional care for the individual: Adds supportive information, "Your sensitivity and ability to put yourself in others' shoes are truly a gift", to provide encouragement during the visitor's interaction tasks;
[0178] Such as Figure 8 shown in Stage 2: Co - creation of art (detailed situation communication and corresponding art creation tasks).
[0179] 5. After the visitor encounters a helpless emotion during another quarrel, she informs the service provider;
[0180] 6. The service provider assigns an interaction task: Draw two plants, one representing herself and the other representing her partner;
[0181] 7. The visitor opens the client application → Enters the "Art Creation Stage" on the tablet → Selects the "Tree" brush element from the toolbox → Draws a tree;
[0182] 8. The art creation agent prompts to be described through voice input;
[0183] 9. The visitor describes this tree as an apple tree in full bloom;
[0184] 10. The agent generates a summary based on the input;
[0185] 11. The visitor adds more semantic brush elements, such as "soil", "clouds", and "lawn", and completes her work;
[0186] 12. The visitor selects the watercolor painting style;
[0187] 13. The visitor clicks to generate and creates a human - machine co - created work;
[0188] Enter the discussion stage:
[0189] 14. The agent guides the visitor to reflect on the work and asks questions such as "Do you want to describe your tree?";
[0190] 15. The visitor shares her feelings and realizes independently that she and her boyfriend are like two different plants - independent but in need of mutual understanding;
[0191] Such as Figure 9 shown in Stage 3: Progress review (retrospective thinking on the created work and descriptive information).
[0192] 16. The service provider reviews the visitor's interaction task history on the service provider's application;
[0193] 17. The service provider examined the original works and the co-created artworks, as well as the conversation records, and gained a deeper understanding of the visitor's situation.
[0194] 18. The agent prompted the service provider to consider the visitor's experience of arguing with their partner.
[0195] 19. The service provider decided to address these insights in the upcoming meeting.
[0196] Through the above steps 1 - 19, it is possible to gain an in-depth understanding of the current emotional problems between the visitor and their partner, and through the execution of multiple interactive tasks, it is expected to solve the emotional problems encountered by the visitor.
[0197] Based on the above embodiments, the present invention also provides a human-computer interaction device, the principle block diagram of which can be as Figure 10 shown. This human-computer interaction device can be used to execute the human-computer interaction method based on asynchronous art creation provided in the above embodiments. For the sake of brevity, it will not be elaborated here. This human-computer interaction device includes: a processor, the processor is coupled to a memory, the memory is used to store computer programs or instructions, and the processor is used to execute the computer programs or instructions stored in the memory, so that the methods in the above method embodiments are executed.
[0198] The present invention also provides a computer-readable storage medium, on which computer instructions for implementing the methods in the above method embodiments are stored.
[0199] For example, when the computer program is executed by a computer, the computer can implement the methods in the above method embodiments.
[0200] The embodiments of the present application also provide a computer program product containing instructions, which when executed by a computer cause the computer to implement the methods in the above method embodiments.
[0201] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of the present application.
[0202] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated here.
[0203] In several embodiments provided by this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.
[0204] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0205] In addition, in each embodiment of this application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.
[0206] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art or part of this technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of this application. The foregoing storage medium includes: USB flash drives, mobile hard disks, read-only memories, random access memories, magnetic disks, or optical discs and other various media that can store program codes.
Claims
1. A human-computer interaction method based on asynchronous art creation, characterized in that: The following steps are involved: Real-time collection of original artwork and multimodal conversation information input by visitors according to the current interactive task; Using a combination of artificial intelligence models to dynamically analyze and process the original artwork and multimodal conversation information, respectively, to generate human-machine co-created artwork and conversation summaries; Feeding back the human-machine co-created artwork and the conversation summary to the service provider, and in response to receiving a next interactive task generated by the service provider based on the conversation summary and the human-machine co-created artwork with the assistance of artificial intelligence, feeding back the next interactive task to the visitor; The method of using the artificial intelligence model combination to dynamically analyze and process the original artwork and the multimodal conversation information respectively to generate a human-machine co-created artwork and a conversation summary includes: using the second module in the artificial intelligence model combination to analyze the artistic creation information in the multimodal conversation information to generate an artistic creation instruction and an artistic creation summary; feeding back the artistic creation instruction to the visitor, and using the third module in the artificial intelligence model combination to dynamically analyze and process the artistic creation information and the artistic creation summary fed back by the visitor according to the artistic creation instruction to generate a human-machine co-created artwork; The method of using the third module in the artificial intelligence model combination to dynamically analyze and process the art creation information and the art creation summary fed back by the visitor according to the art creation instruction to generate a human-machine co-created art work includes: obtaining art creation interface information in response to receiving the art creation information fed back by the visitor according to the art creation instruction; using the third module in the artificial intelligence model combination to perform color segmentation on the art creation interface information to obtain information of several color modules; and generating a human-machine co-created art work based on the information of all the color modules and the art creation summary.
2. The human-computer interaction method based on asynchronous artistic creation according to claim 1, characterized in that: The real-time collection of original artworks and multimodal conversation information input by the visitor according to the current interactive task includes: Constructing a series of interactive questions based on the dialogue principles and example questions corresponding to the current interactive task; The interactive questions are fed back to the visitor in sequence, and the original artworks and multimodal conversation information corresponding to each of the interactive questions fed back by the visitor are collected in sequence.
3. The human-computer interaction method based on asynchronous artistic creation according to claim 1 or 2, characterized in that: The method of using the artificial intelligence model combination to dynamically analyze and process the original artwork and multimodal conversation information to generate human-machine co-created artwork and conversation summary includes: Analyzing information other than artistic creation information in the multimodal conversation information using the first module in the artificial intelligence model combination to generate several types of conversation instructions; Feeding back the conversation instruction to the visitor, and generating several types of conversation information in response to receiving the visitor's feedback information on the conversation instruction; All the session information is integrated to generate a session summary.
4. The human-computer interaction method based on asynchronous artistic creation according to claim 3, characterized in that: The first module in the artificial intelligence model combination is used to analyze the information other than the artistic creation information in the multimodal conversation information to generate several types of conversation instructions, including: Performing standardization processing and / or data conversion on the information other than the artistic creation information in the multimodal conversation information to obtain data in a target format; The first module in the artificial intelligence model combination is used to analyze the data in the target format to generate several types of conversation instructions.
5. The human-computer interaction method based on asynchronous artistic creation according to claim 1, characterized in that: The generation process of the next interactive task includes: Collect historical interaction results within a preset time period; Automatically generate a current interaction result based on the conversation summary and the human-machine co-created artwork; In response to receiving at least one of the information from the server based on the current interaction result, the human-machine co-created artwork and the historical interaction result, a next interaction task is generated with the assistance of artificial intelligence.
6. The human-computer interaction method based on asynchronous artistic creation according to claim 5, characterized in that: The collecting of historical interaction results within a preset time period includes: Collecting personal information of the visitor; Collecting the number of times the visitor performs the interactive task within the preset time period; Collecting a conversation summary of each interactive task performed by the visitor within the preset time period; Based on the personal information, the times and / or the session summary, a historical interaction result is generated.
7. The human-computer interaction method based on asynchronous artistic creation according to claim 5, characterized in that: The automatically generating a current interaction result based on the conversation summary and the human-machine co-created artwork includes: Collecting original artworks, human-computer co-created artworks and conversation summaries corresponding to several target time nodes in the current interactive task; Based on the original artwork, the human-machine co-created artwork and / or the conversation summary corresponding to all the target time nodes, the current interaction result is automatically generated.
8. The human-computer interaction method based on asynchronous artistic creation according to claim 5, characterized in that: Generate the next interactive task with the help of artificial intelligence, including: In response to receiving the conversation reflection content fed back by the visitor according to the current interaction result, generating optimized conversation principles and optimized example questions; Based on the optimized dialogue principles and the optimized example questions, the next interactive task is generated with the assistance of artificial intelligence.
9. The human-computer interaction method based on asynchronous artistic creation according to claim 1, characterized in that: The multimodal session information includes: At least one of audio data, video data, text data, artistic creation information, physiological information and environmental information.
10. A human-computer interaction system based on asynchronous art creation, characterized in that: The system comprises: The information collection module is used to collect original artworks and multimodal conversation information input by visitors according to the current interactive task in real time; An art creation module, for dynamically analyzing and processing the original artwork and multimodal conversation information using a combination of artificial intelligence models, to generate human-machine co-created artworks and conversation summaries; An interaction module, configured to feed back the human-machine co-created artwork and the conversation summary to the service provider, and in response to receiving a next interaction task generated by the service provider based on the conversation summary and the human-machine co-created artwork with the assistance of artificial intelligence, feed back the next interaction task to the visitor; The art creation module is further used to analyze the art creation information in the multimodal conversation information using the second module in the artificial intelligence model combination to generate an art creation instruction and an art creation summary; feed back the art creation instruction to the visitor, and use the third module in the artificial intelligence model combination to dynamically analyze and process the art creation information and the art creation summary fed back by the visitor according to the art creation instruction to generate a human-machine co-created art work; The art creation module is also used to obtain art creation interface information in response to receiving the art creation information fed back by the visitor according to the art creation instructions; use the third module in the artificial intelligence model combination to perform color segmentation on the art creation interface information to obtain information of several color modules; and generate human-machine co-created art works based on the information of all the color modules and the art creation summary.
11. A human-computer interaction device, characterized in that: include: A memory for storing executable instructions; A processor is used to call and run the executable instructions in the memory to execute the steps of the human-computer interaction method based on asynchronous artistic creation as described in any one of claims 1-9.
12. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores program instructions, and when the program instructions are executed by the processor, the human-computer interaction method based on asynchronous artistic creation as described in any one of claims 1 to 9 is implemented.
Citation Information
Patent Citations
Human-computer interaction type abstract picture generation method
CN112529978A
Psychological counseling providing system in metaverse environment
KR1020230167972A