Conversation processing method and device, electronic equipment and storage medium
Through dynamic state analysis driven by a large language model, the problem of the single interaction mode of virtual characters in voice interaction is solved, and emotional human-computer interaction and in-depth gaming experience are achieved.
Patent Information
- Application Number
- CN202510637854.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-16
- Publication Date
- 2025-09-19
AI Technical Summary
In the existing technology, in the user interaction experience of the voice dialogue system, the interaction mode of the voice dialogue system is mainly based on menu-based mechanical responses, which lacks emotional immersion and multimodal interaction experience.
Through dynamic state analysis driven by the Large Language Model (LLM), a dynamic narrative space is generated to achieve continuous evolution and emotional interaction of virtual characters, breaking through the limitations of traditional voice interaction.
It realizes the continuous evolution of virtual characters and emotional human-computer interaction in voice interaction scenarios, breaks through the limitations of traditional menu-based interaction and mechanical responses, and provides a deep gaming experience and emotional immersion.
Smart Images

Figure CN120673753A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data processing, and in particular to a conversation processing method, device, electronic device, and computer-readable storage medium. Background Art
[0002] With advancements in artificial intelligence (AI) technology, the demand for human-computer interaction systems in various application scenarios, such as voice assistants and virtual agents, is increasing. With the rapid development of deep learning technology, data-driven conversational systems are gaining popularity. In related technologies, intelligent hardware products centered around voice conversations primarily utilize menu-based interactions within voice user interfaces (VUIs) to provide mechanical responses. Currently, with increasing demand for a superior user experience, improvements to conversation processing methods are urgently needed. Summary of the Invention
[0003] The embodiments of the present application provide a dialogue processing method, device, electronic device and computer-readable storage medium, which can create a dynamic narrative space through the generative content of LLM, and feed back the dynamic state analysis results driven by LLM to the intelligent voice device so that the intelligent voice device can present the generated chat feedback to the user.
[0004] The technical solution of the embodiment of the present application is implemented as follows:
[0005] This embodiment of the present application provides a conversation processing method, applied to a server, including:
[0006] Receiving first voice dialogue information of a user sent by an intelligent voice device, wherein the first voice dialogue information is used to instruct a virtual character to perform a target task initiated by the user;
[0007] Acquiring the initial activity state of the virtual character when the user initiates the target task;
[0008] Determining a target activity state of the virtual character according to the target task and the initial activity state;
[0009] Sending first dialogue reply information of the target task to the intelligent voice device, wherein the first dialogue reply information includes information obtained by the server analyzing at least the target activity state through a large language model.
[0010] The present application also provides a method for processing a conversation, which is applied to an intelligent voice device and includes:
[0011] Acquiring first voice dialogue information of a user, wherein the first voice dialogue information is used to instruct a virtual character to perform a target task initiated by the user;
[0012] Sending the first voice dialogue information to the server;
[0013] Receive first dialogue reply information of the target task sent by the server, wherein the first dialogue reply information includes information obtained by the server analyzing at least the target activity state of the virtual character through a large language model.
[0014] An embodiment of the present application provides a conversation processing device, which is applied to a server and includes:
[0015] A first receiving module is configured to receive first voice dialogue information of a user sent by an intelligent voice device, wherein the first voice dialogue information is used to instruct the virtual character to perform a target task initiated by the user;
[0016] A first acquisition module is used to acquire the initial activity state of the virtual character when the user initiates the target task;
[0017] A first processing module, configured to determine the target activity state according to the target task and the initial activity state;
[0018] A first sending module is used to send first dialogue reply information of the target task to the intelligent voice device, wherein the first dialogue reply information includes information obtained by the server through analyzing at least the target activity state through a large language model.
[0019] In the above scheme, the first processing module is used to control the activity state of the virtual character to be converted from the initial activity state to the activity state indicated by the state identifier when the activity state indicated by the state identifier contained in the target task is different from the initial activity state and the state migration condition is met between the initial activity state and the activity event contained in the target task, wherein the target activity state includes the activity state indicated by the state identifier.
[0020] In the above solution, the first processing module is configured to perform one of the following:
[0021] When the target activity state is an in-transit state, controlling the activity state of the virtual character to transition from the in-transit state to the at-location state after a first duration, wherein the in-transit state indicates a state in which the virtual character performs activities according to the activity event but has not yet reached the destination included in the target task, and the at-location state indicates a state in which the virtual character has reached the destination;
[0022] When the target activity state is the on-site state, controlling the activity state of the virtual character to switch from the on-site state to the on-the-go state after a second time period;
[0023] When the target activity state is the on-site state, the activity state of the virtual character is controlled to be converted from the on-site state to the at-home state after a third time period, wherein the at-home state indicates that the virtual character is at the departure place.
[0024] In the above solution, the first acquisition module is used to acquire the real-time activity status and real-time location information of the virtual character, wherein the real-time activity status includes the en route status or the on-site status;
[0025] The first processing module is configured to analyze the real-time activity state and the real-time location information through the large language model to generate the first dialogue reply information, wherein the real-time activity state includes the target activity state and / or the activity state after the target activity state is converted.
[0026] In the above solution, the first conversation reply information includes postcard text, wherein the postcard text includes one or more of the following: text corresponding to the real-time activity status and / or the real-time location information, and a postcard picture corresponding to the real-time location information.
[0027] In the above scheme, the first sending module is used to send the postcard text to the intelligent voice device in response to the activity status of the virtual character being in transit and the virtual character leaving the scenic spot along the way and heading for the destination included in the target task, or in response to the activity status of the virtual character being in transit and the virtual character leaving the scenic spot along the way and heading for the departure place.
[0028] In the above solution, the first receiving module is used to receive the second voice dialogue information of the user sent by the intelligent voice device;
[0029] The first processing module is configured to generate, by using the large language model, a second dialogue reply message based on content indexed by the key information corresponding to the location information included in the scenic spot list, when the acquired real-time location information of the virtual character is included in the location information included in the scenic spot list, and key information of the user's intention included in the second voice dialogue message is the same as key information corresponding to the location information included in the scenic spot list;
[0030] The first sending module is used to send the second dialogue reply information to the intelligent voice device.
[0031] The present application also provides a dialogue processing device, which is applied to an intelligent voice device and includes:
[0032] A second acquisition module is configured to acquire first voice dialogue information of a user, wherein the first voice dialogue information is used to instruct the virtual character to execute a target task initiated by the user;
[0033] A second sending module, configured to send the first voice dialogue information to a server;
[0034] The second receiving module is used to receive the first dialogue reply information of the target task sent by the server, wherein the first dialogue reply information includes information obtained by the server through analyzing at least the target activity state of the virtual character through a large language model.
[0035] In the above solution, the first dialogue reply information includes postcard text, wherein the postcard text includes one or more of the following: text corresponding to the real-time activity status and / or real-time location information of the virtual character, and a postcard picture corresponding to the real-time location information.
[0036] In the above solution, the second acquisition module is used to acquire the third voice dialogue information of the user, and the third voice dialogue information is used to display the postcard text;
[0037] The presentation module is used to display the postcard picture included in the postcard text on the human-computer interaction interface of the intelligent voice device, and the output device is used to broadcast the text through voice synthesis.
[0038] In the above solution, the second acquisition module is used to obtain the second voice dialogue information of the user;
[0039] A second sending module, configured to send the second voice dialogue information to the server;
[0040] A second receiving module is used to receive a second dialogue reply message sent by the server, wherein the second dialogue reply message includes information generated by the server through the large language model according to the content of the key information index corresponding to the location information contained in the scenic spot list, and the key information corresponding to the location information contained in the scenic spot list is the same as the key information of the user intention contained in the second voice dialogue message.
[0041] An embodiment of the present application provides an electronic device, including:
[0042] a memory for storing computer-executable instructions or computer programs;
[0043] The processor is used to implement a dialogue processing method provided in an embodiment of the present application when executing the computer-executable instructions or computer program stored in the memory.
[0044] An embodiment of the present application provides a computer-readable storage medium storing a computer program or computer-executable instructions, which, when executed by a processor, implements a conversation processing method provided in an embodiment of the present application.
[0045] An embodiment of the present application provides a computer program product, including a computer program or computer-executable instructions. When the computer program or computer-executable instructions are executed by a processor, a conversation processing method provided in an embodiment of the present application is implemented.
[0046] The embodiments of the present application have the following beneficial effects:
[0047] The server receives a first voice conversation message from a user sent by an intelligent voice device, wherein the first voice conversation message is used to instruct a virtual character to execute a target task initiated by the user; obtains the initial activity state of the virtual character when the user initiates the target task; determines the target activity state of the virtual character based on the target task and the initial activity state; and sends a first conversation reply message for the target task to the intelligent voice device, wherein the first conversation reply message includes information obtained by the server analyzing at least the target activity state using a large language model. In this way, in a voice interaction scenario, the server analyzes at least the target activity state of the virtual character using a large language model to obtain the first conversation reply message. The first conversation reply message is the result of a dynamic state analysis driven by a large language model (LLM), which enables the virtual character to have the characteristics of a continuously evolving "digital life" and opens up a new paradigm for emotional human-computer interaction. In this application, the intelligent voice device receives a first voice dialogue message initiated by a user and forwards it to a server. The server creates a dynamic narrative space through the generative content of LLM for the target task initiated by the user in the first voice dialogue message, and feeds back the dynamic state analysis results driven by LLM to the intelligent voice device so that the intelligent voice device can present the generated chat feedback to the user. In this way, in products with AI voice interactive intelligent hardware as the main body, it breaks through the menu-based interaction limitations of traditional VUI and the mechanical response limitations of traditional voice assistants, and realizes an emotional immersion experience that goes beyond the GUI graphic interactive products, and realizes a deep gaming experience through multimodal perception and generative narrative. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] Figure 1 1 is a schematic diagram of the architecture of the dialogue processing system 100 provided in an embodiment of the present application;
[0049] Figure 2A 6 is a schematic structural diagram of an electronic device 600 provided in an embodiment of the present application;
[0050] Figure 2B 7 is a schematic structural diagram of an electronic device 700 provided in an embodiment of the present application;
[0051] Figure 3 This is a flow diagram of the conversation processing method provided in the embodiment of the present application. Figure 1 ;
[0052] Figure 4 This is a second flow chart of the conversation processing method provided in an embodiment of the present application;
[0053] Figure 5 This is a timing flow diagram of the conversation processing method provided in an embodiment of the present application;
[0054] Figure 6 This is a diagram of the dialogue processing architecture provided by an embodiment of the present application;
[0055] Figure 7 This is a schematic diagram of Meng UU receiving content pushed by a server and presenting it, provided by an embodiment of the present application;
[0056] Figure 8 This is a schematic diagram of a terminal device receiving content pushed by a server and presenting it, provided in an embodiment of the present application;
[0057] Figure 9 This is a schematic diagram of the interface display of Meng UU in different states provided by the embodiment of the present application;
[0058] Figure 10 This is a schematic diagram of message push by Meng UU provided in an embodiment of the present application. DETAILED DESCRIPTION
[0059] In order to make the purpose, technical solutions and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limiting this application. All other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.
[0060] In the following description, reference is made to “some embodiments”, which describes a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0061] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program that has a predetermined function and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories) or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the function of the module or unit.
[0062] In the following description, the terms "first\second\..." are merely used to distinguish similar objects and do not represent a specific ordering of the objects. It is understandable that "first\second\..." can be interchanged with a specific order or sequence where permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.
[0063] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application pertains. The terms used herein are for the purpose of describing the embodiments of this application only and are not intended to limit this application.
[0064] In the related technologies, currently, the interactive mode of the game system is mainly based on graphical user interface (GUI) graphic interaction, including all current mainstream personal computer (PC) games, mobile games and console games; with the development and application of natural language processing (NLP) and LLM, some smart hardware will also realize conversational game play through VUI voice interaction, such as many smart speakers or artificial intelligence (AI) dialogue toys built-in game forms such as idiom chain, werewolf killing script killing, role-playing, etc.
[0065] The GUI game interaction method can more fully and vividly display the game screen and visual effects, and better spread the game content within the community through sharing of pages / pictures, thereby achieving better spreadability, stickiness and topicality. Therefore, compared with the abstract and difficult-to-spread VUI interaction, the game system will adopt GUI interaction more often.
[0066] For AI smart hardware products that use voice dialogue as their primary mode of use, they need to break out of the gameplay limitations of VUI interaction and combine it with the most mainstream game interaction methods of GUI. This way, the games of voice interactive smart hardware products are no longer limited to the primary stage such as idiom chain games, but can be expanded to a richer variety similar to mobile games and computer games.
[0067] In view of this, the embodiments of the present application provide a conversation processing method, device, electronic device, and computer-readable storage medium that can create a dynamic narrative space through LLM-driven generative content and feed back the LLM-driven dynamic state analysis results to the intelligent voice device so that the intelligent voice device can present the user with generated chat feedback. The electronic device provided in the embodiments of the present application can be implemented as a server or as an intelligent voice device. The following is an example of the conversation processing method provided by the embodiment of the present application being implemented by a server.
[0068] For example, see Figure 1 , Figure 1 This is a schematic diagram of the architecture of the dialogue processing system 100 provided in an embodiment of the present application, which is used to support a dialogue processing application, such as Figure 1 As shown, the dialogue processing system 100 includes: a server 200, a network 300, an intelligent voice device 400 (such as various types of robots, voice assistants), and a terminal device 500. The intelligent voice device 400 and the terminal device 500 are respectively connected to the server 200 through the network 300, wherein the network 300 can be a local area network or a wide area network, or a combination of the two; the terminal device 500 is a terminal associated with the user, and a client runs on the intelligent voice device 400 and the terminal device 500. The client can be various types of clients, such as a client related to real-time video display or a browser.
[0069] It should be noted that the technical solution provided in this application can be applied to various scenarios, such as VUI intelligent interaction, smart speaker interaction and other AI voice dialogue scenarios.
[0070] For example, Figure 1 The server 200 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms. The intelligent voice device 400 can be various types of robots (such as remote-controlled robots, household intelligent robots, etc.) and smart speakers. The terminal device 500 can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart watch, a car terminal, etc., but is not limited to this. The intelligent voice device 400, the terminal device 500 and the server 200 can be directly or indirectly connected by wired or wireless communication, which is not limited in the embodiments of the present application.
[0071] The following continues to describe the structure of the electronic device provided in the embodiment of the present application. Take the electronic device as an example, see Figure 2A , Figure 2A is a structural diagram of an electronic device 600 provided in an embodiment of the present application, Figure 2A The electronic device 600 shown includes at least one processor 610, a memory 640, and at least one network interface 620. The various components in the electronic device 600 are coupled together via a bus system 630. It will be appreciated that the bus system 630 is used to enable connectivity and communication between these components. In addition to a data bus, the bus system 630 also includes a power bus, a control bus, and a status signal bus. However, for clarity, in FIG. 2 , all of these buses are labeled as the bus system 630.
[0072] The processor 610 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc., where the general-purpose processor can be a microprocessor or any conventional processor, etc.
[0073] The memory 640 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memory, hard drives, optical drives, etc. The memory 640 may optionally include one or more storage devices physically located away from the processor 610.
[0074] The memory 640 includes volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory may be a read-only memory (ROM), and the volatile memory may be a random access memory (RAM). The memory 640 described in the embodiments of the present application is intended to include any suitable type of memory.
[0075] In some embodiments, the memory 640 can store data to support various operations, examples of which include programs, modules, and data structures, or a subset or superset thereof, as exemplified below.
[0076] Operating system 641, including system programs for processing various basic system services and performing hardware-related tasks, such as the framework layer, core library layer, and driver layer, which are used to implement various basic services and process hardware-based tasks;
[0077] A network communication module 642 for reaching other computing devices via one or more (wired or wireless) network interfaces 620 , exemplary network interfaces 620 including Bluetooth, Wireless LAN (WiFi), and Universal Serial Bus (USB);
[0078] In some embodiments, the apparatus provided in the embodiments of the present application may be implemented in software. Figure 2A The dialog processing device 643 stored in the memory 640 is shown. It can be software in the form of a program or plug-in, and includes the following software modules: a first receiving module 6431, a first sending module 6432, a first obtaining module 6433, and a first processing module 6434. These modules are logical and can be arbitrarily combined or further divided according to the functions implemented. It should be noted that in Figure 2A For the sake of convenience, all the above modules are shown at once, but it should not be regarded as excluding the implementation of the dialogue processing device 643 that can only include the first receiving module 6431, the first sending module 6432, the first acquisition module 6433 and the first processing module 6434. The functions of each module will be explained below.
[0079] The following continues to describe the structure of the electronic device for implementing the dialogue processing method provided in the embodiment of the present application, taking the electronic device 700 as an intelligent voice device as an example, see Figure 2B , Figure 2B is a structural diagram of an electronic device 700 provided in an embodiment of the present application, Figure 2B The electronic device 700 shown includes a conversation processing device 755, which can be software in the form of a program or plug-in, and includes the following software modules: a second acquisition module 7551, a second sending module 7552, and a second receiving module 7553. These modules are logical and can be arbitrarily combined or further divided according to the functions implemented. It should be noted that in Figure 2B For the sake of convenience, all the above modules are shown at once, but it should not be regarded as excluding the implementation of the dialogue processing device 755 that can only include the second acquisition module 7551, the second sending module 7552 and the second receiving module 7553. The functions of each module will be explained below.
[0080] It should be noted that Figure 2B The electronic device 700 shown includes a processor 710, a network interface 720, a bus system 740, an operating system 751, and a network communication module 752. Figure 2A The corresponding modules included in have the same structure and the same function, and the embodiments of the present application will not be repeated here.
[0081] also, Figure 2B The electronic device 700 shown also includes a user interface 730. The user interface 730 includes one or more output devices 731 that enable the presentation of media content, including one or more speakers and / or one or more visual display screens. The user interface 730 also includes one or more input devices 732, including user interface components that facilitate user input, such as a keyboard, mouse, microphone, touch screen display, camera, other input buttons and controls.
[0082] Figure 2B The memory 750 in the illustrated electronic device 700 further includes a presentation module 753 (eg, a display screen, etc.) and an input processing module 754 .
[0083] a presentation module 753 for enabling output of information via one or more output devices 731 (e.g., speakers, etc.) associated with the user interface 730 (e.g., a user interface for operating peripheral devices and displaying content and information);
[0084] The input processing module 754 is configured to detect one or more user inputs or interactions from one of the one or more input devices 732 and to translate the detected inputs or interactions.
[0085] The conversation processing method provided in the embodiment of the present application will be described in detail below in combination with the exemplary application and implementation of the server provided in the embodiment of the present application.
[0086] See also Figure 3 , Figure 3 This is a flow chart of the conversation processing method provided by the embodiment of the present application. The above conversation processing method is applied to the server and is combined with Figure 3 The steps shown are explained.
[0087] In step 101, first voice dialogue information of a user sent by an intelligent voice device is received, wherein the first voice dialogue information is used to instruct a virtual character to perform a target task initiated by the user.
[0088] Here, the user interacts with an intelligent voice device (such as Meng UU, an intelligent voice dialogue device) through a client, and Meng UU communicates with a server, which is used to execute the dialogue processing method. Meng UU deploys a virtual character, and the user can interact with the virtual character through Meng UU.
[0089] Here, the target task includes but is not limited to one or more of the following: activity events, destinations, location reporting, and information presentation. Activity events include but are not limited to travel events, such as "Let's go on a trip", "End the trip and go home", and "End the trip and go home after passing through another scenic spot". The destination includes the destination that the activity event indicates the virtual character is to reach. Location reporting includes real-time location reporting, such as voice location reporting instructions such as "Where are you?" and "Where is the next stop?" Information presentation includes the presentation of relevant information during the virtual character's activities and the presentation of relevant information before and / or after the virtual character's previous activities.
[0090] In step 102, the initial activity state of the virtual character when the user initiates the target task is obtained.
[0091] Here, when the user initiates the target task, that is, the server receives the first voice dialogue information sent by the intelligent voice device, it checks the initial activity state (that is, the current state) of the virtual character.
[0092] In step 103, the target activity state of the virtual character is determined according to the target task and the initial activity state.
[0093] Here, the user-initiated target task determines whether the initial active state requires a state machine change. Finally, if a state change is necessary, the state to switch to (the target active state) is determined. Otherwise, the initial active state remains. This allows LLM to generate voice chat feedback based on dynamic states, breaking through the mechanical response limitations of traditional voice assistants.
[0094] In step 104, first dialogue reply information of the target task is sent to the intelligent voice device, wherein the first dialogue reply information includes information obtained by the server through analyzing at least the target activity state of the virtual character through the large language model.
[0095] Here, in the voice interaction scenario, the server analyzes at least the target activity state of the virtual character through a large language model to obtain a first dialogue reply message. The first dialogue reply message is the result of a dynamic state analysis driven by LLM, which enables the virtual character to have the characteristics of a continuously evolving "digital life", opening up a new paradigm for emotional human-computer interaction. In this application, the intelligent voice device receives the first voice dialogue message initiated by the user and forwards it to the server. The server creates a dynamic narrative space through the generative content of LLM for the target task initiated by the user in the first voice dialogue message, and feeds back the dynamic state analysis result driven by LLM to the intelligent voice device, so that the intelligent voice device can present the generated chat feedback to the user. In this way, in products with AI voice interaction intelligent hardware as the main body, it breaks through the menu-based interaction limitations of traditional VUI and the mechanical response limitations of traditional voice assistants, achieving an emotional immersion experience that goes beyond GUI graphic interaction products, and realizing a deep gaming experience through multimodal perception and generative narrative.
[0096] In some embodiments, the above-mentioned step 103 can be implemented in the following manner: when the activity state indicated by the state identifier contained in the target task is different from the initial activity state, and the state transition condition is met between the initial activity state and the activity event contained in the target task, the activity state of the virtual character is controlled to be converted from the initial activity state to the activity state indicated by the state identifier, wherein the target activity state includes the activity state indicated by the state identifier.
[0097] Here, the state transition condition is satisfied between the initial activity state and the activity event included in the target task, which means that if both the event triggering and the condition check are detected to be satisfied, the state transition is executed.
[0098] The activity states of the virtual character include: on-the-go state, on-site state, and at-home state.
[0099] The "in-transit" state indicates that the avatar is performing activities according to the activity event but has not yet reached the destination included in the target task. For example, the avatar selects a destination to travel and has left the body but has not yet reached the destination.
[0100] The on-site state indicates that the avatar has arrived at the destination. For example, the avatar has arrived at the selected destination and is staying there to play.
[0101] The home state indicates that the avatar is at the departure location. For example, the state where Meng UU's inspiration does not leave the body and does not go out for travel is the home state of Meng UU.
[0102] Example 1: When the activity state (such as the in-transit state) indicated by the status identifier contained in the target task (such as let's go traveling) is different from the initial activity state (at home state), and the state transition conditions are met between the initial activity state (at home state) and the activity event contained in the target task (such as going traveling), the activity state of the virtual character is controlled to be converted from the at-home state to the in-transit state indicated by the status identifier.
[0103] Example 2: When the activity state (such as the at-home state) indicated by the state identifier contained in the target task (such as returning home after a trip) is different from the initial activity state (in transit / at-home state), and the state transition conditions are met between the initial activity state (in transit / at-home state) and the activity event contained in the target task (such as the end of a trip), the activity state of the virtual character is controlled to be converted from the in transit / at-home state to the at-home state indicated by the state identifier.
[0104] Example 3: When the activity state indicated by the state identifier included in the target task is the same as the initial activity state, the activity state of the controlled virtual character remains unchanged.
[0105] Here, the state machine changes are controlled according to the current state. This condition-driven state migration mechanism not only ensures the logical rigor of the game world, but also creates a deep strategy space through structured event relationships. Based on this, the dialogue processing method provided in this application can be explored and evolved towards more mainstream games.
[0106] In some embodiments, after the activity state of the virtual character is controlled to be converted from the initial activity state to the activity state indicated by the state identifier, one of the following may be further performed:
[0107] When the target activity state is the en-route state, the activity state of the virtual character is controlled to switch from the en-route state to the at-location state after a first duration, wherein the en-route state indicates that the virtual character is performing activities according to the activity event but has not yet reached the destination included in the target task, and the at-location state indicates that the virtual character has reached the destination;
[0108] When the target activity state is the on-site state, the activity state of the virtual character is controlled to be converted from the on-site state to the on-the-go state after the second time period.
[0109] When the target activity state is the local state, the activity state of the virtual character is controlled to be converted from the local state to the home state after a third time period, wherein the home state indicates that the virtual character is in a state of being at the departure place.
[0110] Here, the state machine changes are controlled according to the timer algorithm. This state migration mechanism based on strict timing control ensures timing accuracy and improves synchronization and coordination capabilities. It should be noted that the control of state machine changes in this application can first control the state machine changes according to the current state, and further use the built-in timer algorithm to continue to control the state machine changes. In this way, the optimal balance between condition-driven and timing-driven is achieved.
[0111] In some embodiments, before executing the above step 104, the following processing may be performed:
[0112] Obtaining real-time activity status and real-time location information of the virtual character, wherein the real-time activity status includes en route status or on-site status;
[0113] The real-time activity state and the real-time location information are analyzed by a large language model to generate first dialogue reply information, wherein the real-time activity state includes the target activity state and / or the activity state after the target activity state is converted.
[0114] Here, real-time location information includes the real-time location information of the virtual character when performing the activity event, including but not limited to latitude and longitude. The real-time activity status and real-time location information are used by the LLM to determine the location of the virtual character (or its approximate location), thereby generating the first dialogue reply information (one or more of the following: news, attraction recommendations, postcard text, etc., related to the determined location).
[0115] Here, the current position longitude = (destination longitude - departure point longitude) × departure time / t + departure point longitude; the current position latitude = (destination latitude - departure point latitude) × departure time / t + departure point latitude.
[0116] In some embodiments, the first conversation reply message includes postcard text, wherein the postcard text includes one or more of the following: text corresponding to the real-time activity status and / or real-time location information, and a postcard picture corresponding to the real-time location information.
[0117] Here, the postcard text can be generated by the server through LLM based on the set "Postcard Summary Table" by extracting key content; the postcard summary table includes but is not limited to: postcard name, postcard description text, postcard picture, and the corresponding destination of the picture.
[0118] In this way, based on the structured constraints of the "Postcard List" + LLM generative innovation, multimodal feature fusion is achieved, which improves the personalized customization effect of postcards.
[0119] In some embodiments, the above-mentioned step 104 can be implemented in the following manner: in response to the virtual character's activity status being in a transit state, and the virtual character leaving the scenic spot along the way and heading for the destination included in the target task, or in response to the virtual character's activity status being in a transit state, and the virtual character leaving the scenic spot along the way and heading for the departure point, sending a postcard text to the intelligent voice device.
[0120] In some embodiments, the following processing may also be performed: sending a postcard text to a terminal device associated with the intelligent voice device, so that the terminal device displays the postcard text.
[0121] Here, the user interacts with the intelligent voice device (such as MengUU, an intelligent voice dialogue device) through the client, and MengUU is connected to the server and the terminal device. In the case where the service generates a postcard text, the postcard text can be sent to MengUU and the terminal device.
[0122] In some embodiments, after executing step 104 above, the following processing may be performed:
[0123] Receiving the second voice conversation information of the user sent by the intelligent voice device;
[0124] When the acquired real-time location information of the virtual character is within the location information included in the scenic spot list, and the key information of the user's intention included in the second voice dialogue message is the same as the key information corresponding to the location information included in the scenic spot list, a second dialogue reply message is generated using the large language model based on the content indexed by the key information corresponding to the location information included in the scenic spot list;
[0125] Send a second conversation reply message to the smart voice device.
[0126] For example, in a voice interaction scenario, after the server receives the user's first voice conversation message from the intelligent voice device, the user may chat with MengUU about travel-related experiences. When the user and MengUU voice chat about content that triggers travel-related intentions, the server extracts key content from the "Attractions List" / "Attractions List" and generates real-time chat feedback through LLM. The attraction list includes but is not limited to: attractions and the following information corresponding to the attraction: affiliated destination, special features, and special events.
[0127] This application builds a structured knowledge constraint + generative creative solution based on the attraction list and LLM, generates real-time travel dialogue feedback, and improves the real-time interactive experience.
[0128] See also Figure 4 , Figure 4This is a flow chart of the dialogue processing method provided by the embodiment of the present application. The above dialogue processing method is applied to intelligent voice devices and will be combined with Figure 4 The steps shown are explained.
[0129] In step 201, first voice dialogue information of a user is obtained, wherein the first voice dialogue information is used to instruct a virtual character to perform a target task initiated by the user.
[0130] Here, the user interacts with an intelligent voice device (such as MengUU, an intelligent voice dialogue device) through a client, which is connected to a server and is used to execute the dialogue processing method. MengUU deploys a virtual character, and the user can interact with the virtual character through MengUU.
[0131] In step 202, first voice dialogue information is sent to a server.
[0132] Here, after receiving the first voice dialogue information initiated by the user, Meng UU can send the obtained first voice dialogue information to the server, and the server analyzes at least the target activity state of the virtual character through the large language model to obtain the first dialogue reply information.
[0133] In step 203, first dialogue reply information of the target task sent by the server is received, wherein the first dialogue reply information includes information obtained by the server through analyzing at least the target activity state of the virtual character through the large language model.
[0134] Here, the server feeds back the first dialogue reply information to Meng UU. The first dialogue reply information is the result of the dynamic state analysis driven by LLM, which enables the virtual character to have the characteristics of a continuously evolving "digital life", opening up a new paradigm for emotional human-computer interaction. In this application, the intelligent voice device receives the first voice dialogue information initiated by the user and forwards it to the server. The server creates a dynamic narrative space through the generative content of LLM for the target task initiated by the user in the first voice dialogue information, and feeds back the dynamic state analysis result driven by LLM to the intelligent voice device, so that the intelligent voice device can present the generated chat feedback to the user. In this way, in products with AI voice interactive intelligent hardware as the main body, it breaks through the menu-based interaction limitations of traditional VUI and the mechanical response limitations of traditional voice assistants, and realizes an emotional immersion experience that goes beyond the GUI graphic interaction products, and realizes a deep gaming experience through multimodal perception and generative narrative.
[0135] In some embodiments, the first conversation reply message includes postcard text, wherein the postcard text includes one or more of the following: text corresponding to the real-time activity status and / or real-time location information of the virtual character, and a postcard picture corresponding to the real-time location information.
[0136] In some embodiments, after executing step 203 above, the following processing may also be performed:
[0137] Obtaining the user's third voice conversation information, where the third voice conversation information is used to display the postcard text;
[0138] The postcard image included in the postcard text is displayed on the human-computer interaction interface of the intelligent voice device; and the text is read through speech synthesis.
[0139] Here, when the intelligent voice device displays the information fed back by the server, the display of the information can be triggered by the user's voice.
[0140] For example, the text of the postcard can be converted into voice output through a voice broadcast controller (Text-to-Speech, TTS), and the picture of the postcard can be displayed on the human-computer interaction interface of the intelligent voice device.
[0141] In some embodiments, after executing step 203 above, the following processing may also be performed:
[0142] Obtaining the user's second voice conversation information;
[0143] Sending the second voice dialogue information to the server;
[0144] Receive a second dialogue reply message sent by the server, wherein the second dialogue reply message includes information generated by the server through a large language model based on the content of the key information index corresponding to the location information contained in the scenic spot list, and the key information corresponding to the location information contained in the scenic spot list is the same as the key information of the user's intention contained in the second voice dialogue message.
[0145] Here, in the voice interaction scenario, after Meng UU sends the user's first voice dialogue information to the server, the user may chat with Meng UU about travel-related experiences. At this time, Meng UU obtains the user's second voice dialogue information and sends the second voice dialogue information to the server. When the user and Meng UU chat about content that triggers travel-related intentions, the server extracts the key content based on the "General List of Attractions" / "List of Attractions" and generates real-time chat feedback through LLM. At this time, Meng UU receives the second dialogue reply information fed back by the server. This application constructs structured knowledge constraints + generative creative solutions based on the attraction list and LLM to generate real-time travel dialogue feedback and enhance the real-time interactive experience.
[0146] The following describes an exemplary application of the embodiment of the present application in a practical application scenario, which describes the specific implementation process of the dialogue processing method in a voice interaction scenario.
[0147] In some embodiments, see Figure 5 , Figure 5 This is a timing flow chart of the dialogue processing method provided by the embodiment of the present application. Figure 5 Describe the specific implementation process.
[0148] Phase 1: MengUU receives the first voice dialogue information input by the user and sends the first voice dialogue information to the server. Here, the first voice dialogue information includes the user's voice command.
[0149] For example, the first voice dialogue information includes but is not limited to: "Let's go travel", "Where are you traveling", "Return home after the trip", "Other gameplay dialogues, etc."
[0150] Phase 2: The server receives the first voice dialogue information, generates the first dialogue reply information through LLM, and feeds back the first dialogue reply information to Meng UU.
[0151] Exemplarily, after receiving the first voice dialogue information, the server checks the current status of the virtual character (eg, at home, on the way, or at the location).
[0152] Here, when the target task carried by the first voice dialogue information (such as let's go out for a trip) and Meng UU's current state is at home, at this time, Meng UU hears "Let's go out for a trip" when at home, which meets the state transition conditions. Then the activity state of Meng UU is controlled to be converted from the at-home state to the on-the-way state, and the game state machine runs to complete the state switching.
[0153] Here, when the target task carried by the first voice dialogue information (such as ending the trip and going home) and Meng UU's current state is in transit / at home, at this time, when Meng UU is in transit / at home, it hears "end the trip and go home" which meets the state transition conditions, then the activity state of Meng UU is controlled to be converted from in transit / at home to at home state, the game state machine runs, and the state switch is completed.
[0154] Here, when the target task carried by the first voice dialogue information (such as let's go out for a trip) and Meng UU's current state is in transit, at this time, when Meng UU hears "Let's go out for a trip" while in transit, the state transition condition is not met, and the activity state of Meng UU is controlled to remain unchanged.
[0155] LLM analyzes real-time activity status and real-time location information to generate a first conversation response (one or more of: news, attraction recommendations, postcard text, etc., related to the determined location). This first conversation response is fed back to MoeUU. After receiving the first conversation response, MoeUU responds to the user's instructions via TTS and displays a status screen on the screen.
[0156] In some embodiments, the state machine changes may continue to be controlled according to a timer algorithm.
[0157] The time t1 that Meng UU stays on the ground satisfies the following formula: minimum stay time < t1 < maximum stay time.
[0158] The duration t of the in-transit state satisfies the following formula: t = straight-line distance between the starting point and the destination / current speed. The longitude and latitude of the starting point and the destination can be extracted from the set "Destination Master Table" to calculate the distance between them.
[0159] While en route, users might chat with MengUU about travel experiences. When a voice chat with MengUU triggers a travel intent, MengUU needs to know their approximate location. The current location needs to be obtained as longitude and latitude, allowing LLM to determine the approximate location and generate relevant responses.
[0160] In some embodiments, the parameters involved in the travel process of Meng UU can be extracted from the set "Variable Summary Table". For example, the variable summary table is as follows:
[0161]
[0162]
[0163] Variable Summary Table
[0164] This application adopts a table configuration method to coordinate the extraction of various data. The parameters of the above table configuration can be flexibly set according to actual needs, and this application does not specifically limit it. In this way, not only the content of each attribute parameter can be flexibly configured, but also the data acquisition efficiency is improved.
[0165] In some embodiments, the destination selection formula is: range radius r = current vitality value × vitality value range.
[0166] In some embodiments, the destinations involved in the travel process of Meng UU can be extracted from the set "destination summary table". The destination summary table includes but is not limited to: a place and the following information corresponding to the place: location (latitude and longitude), level requirements, affiliated unit, and weight.
[0167] Here, vitality value can be understood as an attribute that affects the distance of Meng UU's inspirational flight. Each point of vitality value represents the fixed distance that Meng UU can fly.
[0168] In some embodiments, the vitality value consumption formula is: the timing starts from the departure of Meng UU, 1 vitality value is consumed every n minutes, where n = vitality value range / current speed, and the timing stops when reaching the destination.
[0169] In some embodiments, the current speed formula of Meng UU in the on-the-go state is: v=initial speed×(1-fatigue×fatigue decay).
[0170] Here, the initial speed can be understood as a hidden attribute, which is a fixed value and affects the length of time Meng UU is in the on-the-go state after departure.
[0171] Here, fatigue can be understood as a parameter that affects the actual speed of Meng UU's flight. The actual speed will decay based on the initial speed as Meng UU's fatigue increases.
[0172] For the control of state machine changes, this application can first control the state machine changes according to the current state, and further use the built-in timer algorithm to continue to control the state machine changes. In this way, the optimal balance between condition-driven and timing-driven is achieved.
[0173] Phase 3: The server indexes the corresponding postcard image, generates postcard text through LLM, obtains postcard text, sends the postcard text to Meng UU, and sends the postcard text to the terminal device.
[0174] Here, the server sends the postcard text to MengUU. After receiving the postcard text, MengUU announces the postcard text copy via TTS and displays the postcard image on the screen. The server sends the postcard text to the terminal device. After receiving the postcard text, the terminal device displays the postcard image and copy on the screen.
[0175] In a realizable dialog processing architecture, see Figure 6 , Figure 6 This is a schematic diagram of the dialogue processing architecture provided by the embodiment of the present application. Figure 6 Describe the specific implementation process.
[0176] Meng UU side: serves as the interactive entry point for voice commands, supports game screen presentation, and also serves as the voice outlet for LLM-generated content.
[0177] Server side: supports game system algorithm calculation, game state machine operation, game interactive content generation and feedback.
[0178] Terminal devices such as mobile phones: support auxiliary presentation of game images.
[0179] In a realistic scenario, Meng UU receives user voice commands and sends them to the server. This in turn receives the postcard text from the server, enabling the postcard image to be displayed in a GUI, with the postcard text read out via TTS. The server can also send the postcard text back to the phone, allowing the user to see the game screen rendered on their phone.
[0180] In a realistic scenario, see Figure 7 , Figure 7This is a schematic diagram of the embodiment of the present application providing a method for Meng UU to receive and present content pushed by the server. Figure 7 Describe the specific implementation process.
[0181] Meng UU receives the postcard text pushed by the server, and the user triggers the display of the postcard text through voice commands (such as "open push content"). At this time, Meng UU displays the postcard picture on the screen and announces the postcard text through TTS.
[0182] In a realistic scenario, see Figure 8 , Figure 8 This is a schematic diagram of a terminal device receiving and presenting content pushed by a server provided in an embodiment of the present application. Figure 8 Describe the specific implementation process.
[0183] Users can see the auxiliary presentation effect of the game screen on the mobile phone. Among them, the first area on the mobile phone interface displays: postcard pictures; and the second area on the mobile phone interface displays: postcard copy generated by LLM based on the key information of the game system postcard configuration table.
[0184] In a realistic scenario, see Figure 9 , Figure 9 This is a schematic diagram of the interface display of the Meng UU in different states provided by the embodiment of the present application. Figure 9 Describe the specific implementation process.
[0185] For example, when in transit, the user may chat with Meng UU about travel-related experiences. When the user and Meng UU chat about content that triggers travel-related intentions, Meng UU needs to know the approximate location of the user at that time. For example, the server feedbacks Meng UU's activity status and approximate location. Figure 9 Figure A shows MoeUU as the in-transit interface. MoeUU can play TTS audio (the server generates chat feedback through LLM, including but not limited to approximate location and in-transit status). Three seconds after the audio ends, it will jump to the double-eye interface (MoeUU's default display interface), completing the travel page jump.
[0186] For example, when in a local state, the user may chat with Meng UU about travel-related experiences. When the user and Meng UU have a voice chat about content that triggers travel-related intentions, the game system will extract the key content based on the set "Scenic Spots List" and generate chat feedback through LLM. Figure 9 Figure B shows the MengUU app's local status interface. MengUU can play TTS audio (the server generates chat feedback through LLM, including but not limited to the destination and local status). Three seconds after the audio ends, it will jump to the "eyes" interface (the default display interface for MengUU), redirecting to the travel page.
[0187] For example, when the in-transit / at-home state changes to the home state, the user asks Meng UU where it is currently. Figure 9 Figure C shows the home status interface of Meng UU. Meng UU can play TTS voice (including but not limited to home status). 3 seconds after the playback ends, it jumps to the double-eye interface (the default display interface of Meng UU).
[0188] It should be noted that all information on the interface is provided by the backend server and is information that already exists in the backend. Here, the interface display content is explained based on whether there are unread messages when in the local state. Figure 9 Figure D shows the interface display when there are unread messages. Figure 9 Figure E shows the interface display when there are no unread messages. Whether there are unread messages can be indicated by the message / gift icon on the interface.
[0189] In a realistic scenario, see Figure 10 , Figure 10 This is a schematic diagram of the message push provided by the embodiment of this application. Figure 10 Describe the specific implementation process.
[0190] Postcard message push is an important interactive feature of MengUU. This application has arranged a series of emoji animations. Other push notifications are sent silently in the background, and users can read them by short pressing.
[0191] MengUU's working status includes but is not limited to: active, standby, and dormant.
[0192] At this time, a new message enters the active state; determine whether it is a postcard message. If so, enter the reminder interaction process; if not, wait in standby mode for the user to read the message.
[0193] Reminder interaction process: Play the messageB emoticon (no sound) in a loop, vibrate for 0.5 seconds, and then enter the standby state after 5 minutes. In the standby state, it continues to play messageB.Bin (the emoticon mentioned above) in a loop, and then enters the sleep state. In the sleep state, a static wallpaper of messageB is displayed. After each playback / display in the reminder interaction process, the device enters the waiting for user action phase.
[0194] User operations include but are not limited to operations on the following components of Meng UU: AI dialogue button, intercom button, power button, and inertial measurement unit (IMU).
[0195] A short press of the AI dialogue button triggers a postcard check. If so, the postcard display process begins; if not, the payment reminder message check process begins. Here, after receiving a message, only the first AI dialogue short press is "Read Message." The rest of the user interaction logic remains the same.
[0196] Postcard display process: Play the emoji messageA.bin (the emoji of an opened letter). It contains the local TTS: "During the time we haven't seen each other, I prepared a small gift for you~" and the audio [postcard.mp3]. Then, it vibrates for 0.5 seconds, and the screen displays the postcard image pushed by the backend, with the background TTS voice. If the user does not take any action, the screen automatically switches to the next postcard; or, the user can press the AI dialogue button "Cut Song" to play the next postcard. If there are multiple postcards, the system will play the next postcard after each one finishes playing. At this point, if the user does not take any action, the screen automatically switches to the next postcard; or, the user can press the AI dialogue button "Cut Song" to play the next postcard.
[0197] Payment reminder message check process: If it is a payment reminder, it reminds you that your Mengzi Pass needs to be renewed. Expiration date: YYYY-MM-DD. After playing, the payment message inventory count decreases by 1. The app download QR code is displayed. You can also play the audio [renew.mp3] at this time, for example, "Without a pass, Mengzi can't travel in the human world. Please scan the QR code below and help me renew my pass, Master!"
[0198] After the above message is played, the last page will still be displayed for 5 minutes before becoming active; then, when entering the standby state, the default standby effect wallpaper will be restored.
[0199] The message push process of Meng UU provided in this application realizes dynamic information transmission, converts abstract notifications into concrete emotional expressions (such as different expressions), and reduces the time it takes for users to understand the push intention. This dynamic feedback + emotional expression method achieves a balance between functionality and emotion, and improves user interaction perception efficiency.
[0200] The dialogue processing method provided in this application, through VUI voice interaction and combined with LLM generative content, realizes gameplay similar to that of more complex traditional graphical interface games such as "Travel Frog" on products with AI voice interactive intelligent hardware as the main body. As an intelligent hardware product based on AI voice dialogue, Meng UU is not limited to the framework of traditional voice intelligent hardware that can only perform VUI interactive games such as idiom chain games, but involves the gameplay of mainstream mobile games. Through state machines, algorithm formulas, LLM prompt (promt) control and various device-side displays and supporting applications (Application, App), it fully explores the scalability and possibilities of large-scale model voice intelligent hardware in game interaction and game systems, and can explore and evolve towards more mainstream game types in the future based on this.
[0201] The following continues to describe the exemplary structure of the dialogue processing device 643 provided in the embodiment of the present application implemented as a software module. In some embodiments, such as Figure 2A As shown, the software modules stored in the dialogue processing device 643 of the memory 640 may include: a first receiving module 6431, a first sending module 6432, a first obtaining module 6433, and a first processing module 6434. The first receiving module 6431 is configured to receive a first voice dialogue message from a user sent by an intelligent voice device, wherein the first voice dialogue message is used to instruct a virtual character to perform a target task initiated by the user; the first obtaining module 6433 is configured to obtain the initial activity state of the virtual character when the user initiates the target task; the first processing module 6434 is configured to determine the target activity state of the virtual character based on the target task and the initial activity state; and the first sending module 6432 is configured to send a first dialogue reply message for the target task to the intelligent voice device, wherein the first dialogue reply message includes information obtained by the server through analysis of at least the target activity state using a large language model.
[0202] In some embodiments, the first processing module 6434 is used to control the activity state of the virtual character to be converted from the initial activity state to the activity state indicated by the state identifier when the activity state indicated by the state identifier included in the target task is different from the initial activity state and the state migration condition is met between the initial activity state and the activity event included in the target task, wherein the target activity state includes the activity state indicated by the state identifier.
[0203] In some embodiments, the first processing module 6434 is configured to perform one of the following:
[0204] When the target activity state is the en-route state, the activity state of the virtual character is controlled to switch from the en-route state to the at-location state after a first duration, wherein the en-route state indicates that the virtual character is performing activities according to the activity event but has not yet reached the destination included in the target task, and the at-location state indicates that the virtual character has reached the destination;
[0205] When the target activity state is the on-site state, the activity state of the virtual character is controlled to be converted from the on-site state to the on-the-go state after the second time period.
[0206] When the target activity state is the local state, the activity state of the virtual character is controlled to be converted from the local state to the home state after a third time period, wherein the home state indicates that the virtual character is in a state of being at the departure place.
[0207] In some embodiments, the first acquisition module 6433 is used to acquire the real-time activity status and real-time location information of the virtual character, wherein the real-time activity status includes the en route status or the on-site status;
[0208] The first processing module 6434 is configured to analyze the real-time activity state and the real-time location information using a large language model to generate a first dialogue reply message, wherein the real-time activity state includes the target activity state and / or the activity state after the target activity state is converted.
[0209] In some embodiments, the first conversation reply message includes postcard text, wherein the postcard text includes one or more of the following: text corresponding to the real-time activity status and / or real-time location information, and a postcard picture corresponding to the real-time location information.
[0210] In some embodiments, the first sending module 6432 is used to send a postcard text to the intelligent voice device in response to the virtual character's activity status being in-transit and the virtual character leaving the scenic spot along the way and heading for the destination included in the target task, or in response to the virtual character's activity status being in-transit and the virtual character leaving the scenic spot along the way and heading for the departure point.
[0211] In some embodiments, the first sending module 6432 is used to send the postcard text to the terminal device associated with the intelligent voice device, so that the terminal device displays the postcard text.
[0212] In some embodiments, a first receiving module 6431 is used to receive a second voice dialogue message of a user sent by an intelligent voice device; a first processing module 6434 is used to generate a second dialogue reply message through a large language model based on the content indexed by the key information corresponding to the location information contained in the scenic spot list when the acquired real-time location information of the virtual character is in the location information contained in the scenic spot list and the key information of the user's intention contained in the second voice dialogue message is the same as the key information corresponding to the location information contained in the scenic spot list; a first sending module 6432 is used to send the second dialogue reply message to the intelligent voice device.
[0213] The following continues to describe the exemplary structure of the dialogue processing device 755 provided in the embodiment of the present application implemented as a software module. In some embodiments, such as Figure 2B As shown, the software modules stored in the dialogue processing device 755 of the memory 750 may include: a second acquisition module 7551 , a second sending module 7552 and a second receiving module 7553 .
[0214] The second acquisition module 7551 is used to acquire the user's first voice dialogue information, wherein the first voice dialogue information is used to instruct the virtual character to perform the target task initiated by the user;
[0215] The second sending module 7552 is used to send the first voice dialogue information to the server;
[0216] The second receiving module 7553 is used to receive the first dialogue reply information of the target task sent by the server, wherein the first dialogue reply information includes information obtained by the server through analyzing at least the target activity state of the virtual character through the large language model.
[0217] In some embodiments, the first conversation reply message includes postcard text, wherein the postcard text includes one or more of the following: text corresponding to the real-time activity status and / or real-time location information of the virtual character, and a postcard picture corresponding to the real-time location information.
[0218] In some embodiments, the second acquisition module 7551 is used to acquire the user's third voice conversation information, and the third voice conversation information is used to display the postcard text;
[0219] The device also includes: a presentation module 753 for displaying the postcard picture included in the postcard text on the human-computer interaction interface of the intelligent voice device, and an output device 731 for broadcasting the text through voice synthesis.
[0220] In some embodiments, the second acquisition module 7551 is used to obtain the user's second voice dialogue information; the second sending module 7552 is used to send the second voice dialogue information to the server; the second receiving module 7553 is used to receive the second dialogue reply information sent by the server, wherein the second dialogue reply information includes information generated by the server through a large language model based on the content indexed by the key information corresponding to the location information contained in the scenic spot list, and the key information corresponding to the location information contained in the scenic spot list is the same as the key information of the user's intention contained in the second voice dialogue information.
[0221] It should be noted that the description of the device of the embodiment of the present application is similar to the description of the method embodiment above, and has similar beneficial effects as the method embodiment, so it will not be repeated here. Figure 3 ,or Figure 4 The present invention should be understood by referring to the description of any one of the accompanying drawings.
[0222] An embodiment of the present application provides a computer program product, which includes a computer program or computer-executable instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer-executable instructions from the computer-readable storage medium and executes the computer-executable instructions, causing the electronic device to perform the conversation processing method described above in the embodiment of the present application.
[0223] The embodiment of the present application provides a computer-readable storage medium in which computer-executable instructions or computer programs are stored. When the computer-executable instructions or computer programs are executed by a processor, the processor will execute the conversation processing method provided by the embodiment of the present application, for example, Figure 3 ,or Figure 4 The dialogue processing method shown.
[0224] In some embodiments, the computer-readable storage medium may be a ferroelectric random access memory (FRAM), ROM, programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash memory, magnetic surface memory, optical disk, or compact disc read-only memory (CD-ROM); or various devices including one or any combination of the above memories.
[0225] In some embodiments, computer-executable instructions may be in the form of a program, software, software module, script, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment.
[0226] As an example, computer-executable instructions may, but need not, correspond to a file in a file system, may be stored as part of a file that stores other programs or data, such as in one or more scripts in a HyperText Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple coordinating files (e.g., files storing one or more modules, subroutines, or code portions).
[0227] By way of example, computer-executable instructions may be deployed to be executed on one electronic device, or on multiple electronic devices located at one site, or on multiple electronic devices distributed across multiple sites and interconnected by a communication network.
[0228] The above are merely examples of the present application and are not intended to limit the scope of protection of the present application. Any modifications, equivalent replacements, and improvements made within the spirit and scope of the present application are included in the scope of protection of the present application.
Claims
1. A method for processing a conversation, characterized in that: The method is applied to a server and includes: Receiving first voice dialogue information of a user sent by an intelligent voice device, wherein the first voice dialogue information is used to instruct a virtual character to perform a target task initiated by the user; Acquiring the initial activity state of the virtual character when the user initiates the target task; Determining a target activity state of the virtual character according to the target task and the initial activity state; Sending first dialogue reply information of the target task to the intelligent voice device, wherein the first dialogue reply information includes information obtained by the server analyzing at least the target activity state through a large language model.
2. The method according to claim 1, characterized in that The determining the target activity state according to the target task and the initial activity state includes: When the activity state indicated by the state identifier contained in the target task is different from the initial activity state, and the state transition condition is met between the initial activity state and the activity event contained in the target task, the activity state of the virtual character is controlled to be converted from the initial activity state to the activity state indicated by the state identifier, wherein the target activity state includes the activity state indicated by the state identifier.
3. The method according to claim 2, characterized in that After the controlling the activity state of the virtual character is converted from the initial activity state to the activity state indicated by the state identifier, the method includes one of the following: When the target activity state is an in-transit state, controlling the activity state of the virtual character to transition from the in-transit state to the at-location state after a first duration, wherein the in-transit state indicates a state in which the virtual character performs activities according to the activity event but has not yet reached the destination included in the target task, and the at-location state indicates a state in which the virtual character has reached the destination; When the target activity state is the on-site state, controlling the activity state of the virtual character to switch from the on-site state to the on-the-go state after a second time period; When the target activity state is the on-site state, the activity state of the virtual character is controlled to be converted from the on-site state to the at-home state after a third time period, wherein the at-home state indicates that the virtual character is at the departure place.
4. The method according to claim 1 or 3, characterized in that Before sending the first dialogue reply information of the target task to the intelligent voice device, the method includes: Acquiring real-time activity status and real-time location information of the virtual character, wherein the real-time activity status includes a transit status or a local status; The real-time activity state and the real-time location information are analyzed by the large language model to generate the first dialogue reply information, wherein the real-time activity state includes the target activity state and / or the activity state after the target activity state is converted.
5. The method according to claim 4, characterized in that The first conversation reply message includes a postcard text, wherein the postcard text includes one or more of the following: text corresponding to the real-time activity status and / or the real-time location information, and a postcard picture corresponding to the real-time location information.
6. The method according to claim 5, characterized in that The sending the first dialogue reply information of the target task to the intelligent voice device includes: In response to the virtual character's activity status being in-transit and the virtual character leaving the scenic spot along the way and heading towards the destination included in the target task, or in response to the virtual character's activity status being in-transit and the virtual character leaving the scenic spot along the way and heading towards the departure point, the postcard text is sent to the intelligent voice device.
7. The method according to claim 1, characterized in that After sending the first dialogue reply information of the target task to the intelligent voice device, the method includes: Receiving second voice conversation information of the user sent by the intelligent voice device; When the acquired real-time location information of the virtual character is included in the location information included in the scenic spot list, and the key information of the user's intention included in the second voice dialogue message is the same as the key information corresponding to the location information included in the scenic spot list, a second dialogue reply message is generated by the large language model based on the content indexed by the key information corresponding to the location information included in the scenic spot list; The second dialogue reply information is sent to the intelligent voice device.
8. A method for processing a conversation, characterized in that: The method is applied to an intelligent voice device, and the method includes: Acquiring first voice dialogue information of a user, wherein the first voice dialogue information is used to instruct a virtual character to perform a target task initiated by the user; Sending the first voice dialogue information to the server; Receive first dialogue reply information of the target task sent by the server, wherein the first dialogue reply information includes information obtained by the server analyzing at least the target activity state of the virtual character through a large language model.
9. The method according to claim 8, characterized in that The first dialogue reply message includes a postcard text, wherein the postcard text includes one or more of the following: text corresponding to the real-time activity status and / or real-time location information of the virtual character, and a postcard picture corresponding to the real-time location information.
10. The method according to claim 9, characterized in that After receiving the first dialogue reply information of the target task sent by the server, the method includes: Acquiring third voice conversation information of the user, wherein the third voice conversation information is used to display the postcard text; Displaying the postcard picture included in the postcard text on the human-computer interaction interface of the intelligent voice device; The text is read out through speech synthesis.
11. The method according to claim 10, characterized in that After receiving the first dialogue reply information of the target task sent by the server, the method includes: Acquiring second voice conversation information of the user; sending the second voice dialogue information to the server; Receive a second dialogue reply message sent by the server, wherein the second dialogue reply message includes information generated by the server through the large language model based on the content of the key information index corresponding to the location information contained in the scenic spot list, and the key information corresponding to the location information contained in the scenic spot list is the same as the key information of the user intention contained in the second voice dialogue message.
12. A dialogue processing device, characterized in that: The device is applied to a server, and includes: A first receiving module is configured to receive first voice dialogue information of a user sent by an intelligent voice device, wherein the first voice dialogue information is used to instruct the virtual character to perform a target task initiated by the user; A first acquisition module is used to acquire the initial activity state of the virtual character when the user initiates the target task; A first processing module, configured to determine a target activity state of the virtual character according to the target task and the initial activity state; A first sending module is used to send first dialogue reply information of the target task to the intelligent voice device, wherein the first dialogue reply information includes information obtained by the server through analyzing at least the target activity state through a large language model.
13. A dialogue processing device, characterized in that: The device is applied to an intelligent voice device, and the device includes: A second acquisition module is configured to acquire first voice dialogue information of a user, wherein the first voice dialogue information is used to instruct the virtual character to perform a target task initiated by the user; A second sending module, configured to send the first voice dialogue information to a server; The second receiving module is used to receive the first dialogue reply information of the target task sent by the server, wherein the first dialogue reply information includes information obtained by the server through analyzing at least the target activity state of the virtual character through a large language model.
14. An electronic device, characterized in that: include: a memory for storing computer-executable instructions or computer programs; The processor is configured to implement the dialogue processing method according to any one of claims 1 to 11 when executing the computer-executable instructions or computer programs stored in the memory.
15. A computer-readable storage medium storing computer-executable instructions or a computer program, characterized in that: When the computer-executable instructions or computer program are executed by a processor, the dialogue processing method according to any one of claims 1 to 11 is implemented.