Multi-modal data conversation robot interaction method, configuration method, terminal, medium and product
Through the multimodal data dialogue robot interaction method, the problem of single interaction type of existing dialogue robots is solved, and the interactive support of multiple categories of data is realized, which improves communication efficiency and user experience.
Patent Information
- Application Number
- CN202510484155.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-04-17
AI Technical Summary
Existing dialogue robots only support plain text interactions and are unable to send multimedia content, resulting in low communication efficiency and poor user experience. The lack of multi-category support makes the dialogue process monotonous and tedious, making it difficult to meet the diverse expression needs of users.
The multimodal data dialogue robot interaction method is adopted, and through the collaborative work of the interaction layer, processing layer and model layer, it supports the interaction of various data such as text, images, video, audio, etc., and uses streaming technology to improve the interaction fluency, introduce pre-reply information and guide information to optimize the interaction process.
It has achieved the expansion of multimodal data interaction capabilities, improved user participation and interactive experience, reduced interaction lag, and met diversified interaction needs.
Smart Images

Figure CN120336486A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and particularly to a multi-modal data dialogue robot interaction method, configuration method, terminal, medium and product. Background Art
[0002] At present, there are great limitations in the function implementation of dialogue robots, and only support the pure text message interaction mode. This single function design exposes multiple problems in actual application scenarios: First, users can only interact through text input during use, and cannot send multimedia content such as pictures, videos, and audio, which seriously affects the communication efficiency in scenarios that require visual assistance or voice interaction; Second, the lack of multi-category support makes the dialogue process appear monotonous and boring, reducing user participation and interaction experience; Moreover, the pure text interaction mode is difficult to meet the diverse expression needs of users, especially when it is necessary to quickly convey complex information or emotions, the limitations of text are particularly obvious. If a multi-modal model processing architecture is introduced, it will affect the processing speed of interactive text during the robot interaction process, greatly reducing the smoothness of human-computer interaction. This functional deficiency directly leads to a decline in user experience, specifically manifested in multiple dimensions such as reduced user participation, increased dialogue interruption rate, and decreased user satisfaction. Summary of the Invention
[0003] In view of the above-mentioned disadvantages of the prior art, the purpose of this application is to provide a multi-modal data dialogue robot interaction method, configuration method, terminal, medium and product, which is used to solve the problem of the single interaction type of existing dialogue robots.
[0004] To achieve the above purpose and other related purposes, the first aspect of this application provides a multi-modal data dialogue robot interaction method, which is applied to a dialogue robot including an interaction layer, a processing layer and a model layer. The method includes: the interaction layer responds to the user dialogue information input at the current dialogue node and establishes a session connection with the processing layer; the processing layer completes the initialization configuration with the model layer and sends the user dialogue information to the model layer; the model layer based on the prompt information preset for the current dialogue node and the user dialogue information, returns interaction information and multi-modal extension information to the processing layer; the processing layer respectively assembles the formats of the interaction information and the multi-modal extension information, and then streams and sends them to the interaction layer.
[0005] In some embodiments of the first aspect of this application, the model layer returns interaction information to the processing layer, including: after the user dialogue information triggers the function function of the model layer, the function function configures the corresponding pre-reply information; the model layer sends the pre-reply information to the processing layer as interaction information.
[0006] In some embodiments of the first aspect of the present application, the multimodal extension information carries a corresponding data assembly structure; after the processing layer respectively performs format assembly on the interaction information and the multimodal extension information, it is streamed to the interaction layer, including: the processing layer performs format assembly on the multimodal extension information according to the data assembly structure; after the processing layer sends the interaction information in text format, it streams the assembled multimodal extension information.
[0007] In some embodiments of the first aspect of the present application, the multimodal extension information includes a preset variable configuration slot; after the model layer assigns values to variables based on the API interfaces pre-configured in the variable configuration slot, it returns interaction information and the multimodal extension information to the processing layer.
[0008] In some embodiments of the first aspect of the present application, the model layer includes a pre-configured node knowledge base; the model layer returns interaction information and multimodal extension information to the processing layer based on the node knowledge base, the prompt information preset for the current dialogue node, and the user dialogue information.
[0009] In some embodiments of the first aspect of the present application, before the processing layer sends the user dialogue information to the model layer, the model layer returns corresponding guiding information to the processing layer based on the prompt information; the processing layer sends the guiding information to the interaction layer.
[0010] To achieve the above and other related purposes, the second aspect of the present application provides a configuration method for a multimodal data dialogue robot, which is applied to a dialogue robot including an interaction layer, a processing layer, and a model layer, including: the dialogue robot receives the prompt information configured at the current dialogue node; the interaction layer responds to the user dialogue information input at the current dialogue node and establishes a session connection with the processing layer; the processing layer completes the initialization configuration with the model layer and sends the user dialogue information to the model layer; the model layer returns interaction information and multimodal extension information to the processing layer based on the prompt information and the user dialogue information; the processing layer respectively performs format assembly on the interaction information and the multimodal extension information and then streams them to the interaction layer.
[0011] In some embodiments of the second aspect of the present application, the configuration method for the multimodal data dialogue robot further includes: the dialogue robot receives the function function configuration information of the current dialogue node, and the function function configuration information includes pre-reply information; the model layer sends the pre-reply information to the processing layer as interaction information.
[0012] In some embodiments of the second aspect of the present application, the functional function configuration information further includes a data assembly structure. After the processing layer performs format assembly on the interaction information and the multimodal extension information respectively, it is streamed to the interaction layer, including: the processing layer performs format assembly on the multimodal extension information according to the data assembly structure; after the processing layer sends the interaction information in text format, it streams the assembled multimodal extension information.
[0013] In some embodiments of the second aspect of the present application, the functional function configuration information further includes a preset variable configuration slot; after the model layer performs variable assignment based on the API interfaces pre-configured in the variable configuration slot, it returns interaction information and the multimodal extension information to the processing layer.
[0014] In some embodiments of the second aspect of the present application, the functional function configuration information further includes preset guiding information; before the processing layer sends the user dialogue information to the model layer, the model layer returns corresponding guiding information to the processing layer based on the prompt information; the processing layer sends the guiding information to the interaction layer.
[0015] To achieve the above object and other related objects, a third aspect of the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the multimodal data dialogue robot interaction method. To achieve the above object and other related objects,
[0016] A fourth aspect of the present application provides a computer program product, which includes computer program code. When the computer program code runs on a computer, it enables the computer to implement the multimodal data dialogue robot interaction method.
[0017] To achieve the above object and other related objects, a fifth aspect of the present application provides an electronic terminal, including a memory, a processor, and a computer program stored on the memory; the processor executes the computer program to implement the multimodal data dialogue robot interaction method.
[0018] As described above, the multimodal data dialogue robot interaction method, configuration method, terminal, medium, and product of the present application have the following beneficial effects:
[0019] The present application expands the boundary of the multimodal data interaction ability of the dialogue robot, can support the interaction of various types of data, and meets the requirements of different interaction scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 It shows a schematic flowchart of the multimodal data dialogue robot interaction method in an embodiment of the present application.
[0021] Figure 2 Shown is a schematic diagram of a multi-modal data dialogue robot in an embodiment of the present application.
[0022] Figure 3 Shown is a schematic diagram of the working process of a multi-modal data dialogue robot in an embodiment of the present application.
[0023] Figure 4 Shown is a schematic diagram of the structure of an electronic terminal in an embodiment of the present application. Detailed implementation manners
[0024] The following uses specific specific examples to illustrate the implementation manners of the present application. Those skilled in the art can easily understand other advantages and effects of the present application from the content disclosed in this specification. The present application can also be implemented or applied through other different specific implementation manners. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present application. It should be noted that, without conflict, the following embodiments and the features in the embodiments can be combined with each other.
[0025] In the embodiments of the present application, terms such as "first" and "second" are used to distinguish the same items or similar items with basically the same functions and effects. Those skilled in the art can understand that the terms "first", "second", etc. do not limit the quantity and execution order, and the terms "first", "second", etc. do not necessarily limit being different.
[0026] It should be noted that in the embodiments of the present application, words such as "exemplary" or "for example" represent examples, illustrations or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Exactly speaking, using words such as "exemplary" or "for example" aims to present relevant concepts in a specific manner.
[0027] In the embodiments of the present application, "at least one" means one or more, and "a plurality" means two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone, where A and B can be singular or plural. The character " / " generally represents an "or" relationship between the front and back associated objects. "At least one (item)" or its similar expression below refers to any combination of these items, including any combination of single item (item) or plural items (items). For example, at least one (item) of a, b or c can represent: a, b, c, a - b, a - c, b - c or a - b - c, where a, b, c can be single or multiple.
[0028] Before further elaborating on the present invention, the nouns and terms involved in the embodiments of the present invention are explained. The nouns and terms involved in the embodiments of the present invention are applicable to the following explanations:
[0029] <1> Streaming Transmission is a data transmission technology that allows data to be continuously sent and received while it is being generated, without waiting for the entire file to be fully downloaded. The characteristics of streaming transmission include: Real-time: Data is transmitted and consumed while it is being generated, with low latency; Continuity: Data is continuously transmitted in the form of a stream, rather than being sent all at once; On-demand loading: Users can start using the data without waiting for the entire file to be downloaded; Resource saving: Reduces the need for local storage, and the data can be discarded after being played or processed;
[0030] <2> Chatbot Service is an application service based on artificial intelligence technology. Through technologies such as Natural Language Processing (NLP) and Machine Learning, a computer program can conduct human-like dialogue interactions with users.
[0031] <3> Dialog Management Call: It is a method used in a dialogue robot to call the dialogue management module. Dialogue management is the core part of a dialogue system, responsible for controlling the dialogue flow, understanding the user's intention, managing the dialogue state, and generating appropriate responses;
[0032] <4> Assembly refers to combining and arranging information of different types or sources according to certain format requirements to make it conform to a specific structure or standard;
[0033] <5> A Q&A mode database is a database system specifically used for storing and managing question and answer pairs (Question and Answer). It is mainly used for quickly retrieving and providing content related to the user's questions.
[0034] <6> API interface (Application Programming Interface) is a way for computer software to interact and communicate. It allows different software programs or systems to share data, functions, and services.
[0035] To facilitate the understanding of the embodiments of the present application, first, in combination with Figure 1 Detailed description. Figure 1Shows a schematic flowchart of a multi-modal data dialogue robot interaction method in an embodiment of the present invention. The multi-modal data dialogue robot interaction method in this embodiment is applied to a dialogue robot (Large Language Model) including an interaction layer, a processing layer, and a model layer. It should be understood that the Large Language Model includes, but is not limited to, large language models such as Llamabot, GPT, BERT, etc. To intuitively and clearly show the interaction method between the Large Language Model and the other functional levels of the dialogue robot, Llamabot is used as an example in the specification. Those skilled in the art know that the technical content in this application is not limited to using Llamabot as the model layer of the dialogue robot. The multi-modal data dialogue robot interaction method mainly includes the following steps:
[0036] Step S11: The interaction layer responds to the user dialogue information input at the current dialogue node and establishes a session connection with the processing layer.
[0037] Specifically, as Figure 2 and Figure 3 shown, the interaction layer is Chatbot Service, and the processing layer is DMCall. The interaction layer includes a message layer. The model layer mainly completes data interaction with the interaction layer and the large model through the Llamabot interface. The interaction layer interacts with the user through the interaction messages in the application layer and the message layer. When a user initiates a dialogue, the interaction layer starts and obtains the ID of the dialogue robot. Then the interaction layer establishes a session connection with the session layer. A session connection means that both communication parties negotiate and determine communication parameters, rules, etc. through a series of interaction processes, so as to form a session that can perform data exchange.
[0038] In one embodiment, as Figure 2 shown, the type of user dialogue information output or input through the message layer can be any one or more of the following: Text, List, Button, Image, Files, Video, Audio, Interactive, Location, Contacts.
[0039] Step S12: The processing layer completes the initialization configuration with the model layer and sends the user dialogue information to the model layer.
[0040] Specifically, as Figure 3 shown, after the session between the processing layer and the interaction layer is established, the processing layer and the model layer perform initialization configuration. The initialization configuration includes, but is not limited to, data format determination, interface configuration, model configuration, etc.
[0041] Step S13: Based on the preset prompt information and the user's conversation information of the current conversation node, the model layer sequentially returns interaction information and multimodal extension information to the processing layer.
[0042] Step S14: After respectively performing format assembly on the interaction information and the multimodal extension information, the processing layer streams and sends them to the interaction layer.
[0043] In the overall process of the dialogue robot interacting with the user, the customer configures the prompt information of the dialogue robot, introducing the generation of modal information such as text, list, button, image, file, video, audio, interactive, location, and contacts by the model layer based on the pre-configured prompt information. To adapt to the interaction fluency of the dialogue robot, interaction information and multimodal extension information are generated in a streaming manner and sent to the processing layer. The information after format assembly executed by the processing layer is fed back to the user through the interaction layer, which not only avoids the problem that the interaction information is not user-friendly and inaccurate when the processing layer selects the interaction content, but also solves the problem that the model layer is difficult to directly generate an interaction format that conforms to the current scenario, resulting in the inability to provide intuitive interaction information to the user.
[0044] For example, when the dialogue robot provides list-type function services such as a service list or an option list to the user, and the description of the menu is expected to display different contents according to different scenarios and the number of questions asked, the Llamabot interface will return text-type interaction information and configuration information of the list-type modality to the processing layer. After the processing layer assembles the text-type interaction information and the configuration information of the list-type modality into list-type interaction information, the interaction layer feeds it back to the user through the corresponding interaction applications (such as WhatsApp, Facebook, Instagram, Line, Viber, Telegram) in the application layer.
[0045] In one embodiment, the multimodal extension information carries a corresponding data assembly structure; after respectively performing format assembly on the interaction information and the multimodal extension information, the processing layer streams and sends them to the interaction layer, including: the processing layer performs format assembly on the multimodal extension information according to the data assembly structure; after sending the interaction information in text format, the processing layer streams and sends the assembled multimodal extension information.
[0046] Specifically, such as Figure 2As shown, the function function of the model layer is preset with data assembly structures including List, button, image, text, flies, etc. After the model layer generates multimodal extended information, the function function will configure the corresponding data assembly structure for it, and return the multimodal extended information and the corresponding data assembly structure to the processing layer. The processing layer will assemble the multimodal extended information according to the data assembly structure. After the processing layer streams the interactive information to the interactive layer in text format, the processing layer streams the assembled multimodal extended information to the interactive layer. Then, the interactive layer sends the interactive information and the multimodal extended information to the user.
[0047] For example, if the data assembly structure is "markdown", the multimodal extended information is assembled into information in text format; if the data assembly structure is "", the multimodal extended information is assembled into information in image format. Figure 3 As shown, after the initialization configuration is completed between the processing layer and the model layer, before the processing layer sends the user dialogue information to the model layer, the model layer returns the corresponding guidance information to the processing layer based on the preset prompt information of the current dialogue node; the processing layer sends the guidance information to the interaction layer.
[0048] Specifically, the model layer has multiple pre-set guidance speech templates. The preset prompt information of the current dialogue node is used to clarify the dialogue intention or scenario of the current dialogue node. The model layer matches the corresponding guidance speech template from the preset multiple guidance speech templates according to the preset prompt information of the current dialogue node, and streams the matched guidance speech template to the interaction layer as guidance information.
[0049] It should be understood that guidance information, also known as entry language, refers to the opening remarks or guiding words when users interact with robots, which are intended to attract users' attention, clarify the functions of the robot and guide users to the next step.
[0050] During the conversational interaction between the conversational robot and the user, if the streaming transmission of multimodal interactive information is involved, it may be affected by the data processing speed of the model layer and the processing layer, resulting in interaction jams. Therefore, in this embodiment, guidance information is introduced as the interactive content of the first round of interaction with the user. During the first round of interaction through the guidance information, the model layer configuration and other steps are completed to improve the interaction fluency of the conversational robot. At the same time, at some interaction nodes, the conversational robot needs to clarify the corresponding information to the user in the first round of interaction with the user (such as some nodes need to first confirm the identity information, address information, financial information, etc. with the user). Therefore, introducing preset guidance information at such nodes can reduce the configuration requirements of this part of the prompt information and the number of interaction rounds.
[0051] In one embodiment, the model layer returns interaction information to the processing layer, including: after the user dialogue information triggers the function of the model layer, the function function configures the corresponding pre-reply information; the model layer sends the pre-reply information as interaction information to the processing layer.
[0052] Specifically, in order to avoid the uncertainty of the output of LLM when the function function is triggered, so that LLM can provide certain output content when executing the current function function, pre-reply information is configured in the function function. After the processing layer sends the user dialogue information to the model layer, the function function in the model layer is triggered. The model layer ignores the interactive information generated by itself, and the function function selects a corresponding pre-reply information from the pre-configured pre-reply information list, and the selected pre-reply information is sent to the processing layer as interactive information. The interactive information is generally text language, which is used to respond to questions, prevent users from waiting too long, and enhance interactivity. For example, "Regarding this issue, I suggest you try..." It should be noted that the interactive information generated by the model layer itself is uncertain, but the configuration of the function function can output certain interactive information.
[0053] It should be understood that the function function is a configuration item in the LLM-BOT (model interface) of the model layer for each dialogue node; wherein the configuration information of the function function includes name, description, function category (mainly including API interface call function), API interface information, function request information, function response information, pre-reply information and data assembly structure. In one embodiment, variable configuration slots may be preset in the interaction information. After the model layer fills the corresponding variables into the variable configuration slots of the interaction information, it returns the filled interaction information to the processing layer.
[0054] In one embodiment, a large language model is built into the model layer. The model layer generates corresponding multimodal extension information (additional data) according to the prompt information and the user dialogue information. The multimodal extension information is one-end json data. The type of the multimodal extension information can be any one or more of the following: text, list, button, image, file, video, audio, interactive, location, and contact.
[0055] In one embodiment, the multimodal extension information includes a preset variable configuration slot; after the model layer assigns variables based on the preconfigured API interface in the variable configuration slot, it returns the interaction information and the multimodal extension information to the processing layer.
[0056] Specifically, the variable configuration slot is used to fill the result returned by the API interface. Figure 2As shown, for the functions of the APIStore in the model layer, after the API is called and the result is returned, the obtained output parameters can be post - processed through the arguments of the Response converter and assigned to the specified variables, and the variables are configured in the Response data.
[0057] For example, in the logistics scenario, the variable is waybill_number. After the API is called, the value of waybill_number obtained is XX123456789. The variable waybill_number is configured on the button in the multi - modal extension information. When the model layer returns the multi - modal extension information, XX123456789 is assigned to the slot corresponding to the variable waybill_number in the multi - modal extension information. That is, the button shows "XX123456789".
[0058] In one embodiment, a pre - configured node knowledge base (a database in Q&A mode) can be introduced into the model layer. The node knowledge base can be configured in a direct mode (original) and a model - generated mode (llm_generated). When the node knowledge base is configured in the direct mode, the corresponding answer information is selected from the node knowledge base and transmitted to the user; when the node knowledge base is configured in the model - generated mode, the node knowledge base and the user's dialogue information are input into the large - language model built in the model layer together to generate more accurate complex information.
[0059] In one embodiment, as Figure 2 shown, after the user receives the interaction information and the multi - modal extension information, the user can make a selection or reply. After the interaction layer receives the selection and reply information input by the user, it sends it to the processing layer. The processing layer block sends a request corresponding to the selection and reply information to the model layer and performs corresponding result analysis (DialogueAnalysis). The dialogue analysis in the model layer refers to extracting key information from the dialogue, mining potential causes, and formulating solutions to optimize the dialogue system or improve the user experience. After the result analysis is completed, the model layer ends its work and transmits the execution result information to the processing layer. After the processing layer receives the execution result information, it notifies the interaction layer that the dialogue ends. The interaction layer sends a dialogue evaluation to the user, and after the user completes the dialogue evaluation, all dialogue processes end.
[0060] Similar to the above - mentioned embodiment, the present invention also provides a configuration method for a multi - modal data dialogue robot, which is applied to a dialogue robot including an interaction layer, a processing layer, and a model layer, and includes:
[0061] The dialogue robot receives the prompt information configured at the current dialogue node;
[0062] The interaction layer responds to the user dialogue information input at the current dialogue node and establishes a session connection with the processing layer;
[0063] Initial configuration is completed between the processing layer and the model layer, and the user dialogue information is sent to the model layer;
[0064] Based on the prompt information and the user dialogue information, the model layer returns interaction information and multimodal extension information to the processing layer;
[0065] After the processing layer assembles the formats of the interaction information and the multimodal extension information respectively, it is sent to the interaction layer in a streaming manner.
[0066] In one embodiment, the configuration method of the multimodal data dialogue robot further includes: the dialogue robot receives the functional function configuration information of the current dialogue node, and the functional function configuration information includes pre-reply information; the model layer sends the pre-reply information to the processing layer as interaction information.
[0067] In one embodiment, the functional function configuration information further includes a data assembly structure. After the processing layer assembles the formats of the interaction information and the multimodal extension information respectively and sends them to the interaction layer in a streaming manner, it includes: the processing layer assembles the format of the multimodal extension information according to the data assembly structure; after the processing layer sends the interaction information in text format, it streams and sends the assembled multimodal extension information.
[0068] In one embodiment, the functional function configuration information further includes a preset variable configuration slot; after the model layer assigns variables based on the API interfaces pre-configured in the variable configuration slot, it returns interaction information and multimodal extension information to the processing layer.
[0069] In one embodiment, the functional function configuration information further includes preset guiding information; before the processing layer sends the user dialogue information to the model layer, the model layer returns the corresponding guiding information to the processing layer based on the prompt information; the processing layer sends the guiding information to the interaction layer.
[0070] It should be noted that the prompt information, the functional function configuration information (including: pre-reply information, data assembly structure, variable configuration slot, API interface) and the guiding information of the current dialogue node can be configured with reference to the above embodiments, and will not be elaborated here.
[0071] Figure 4 It is a schematic block diagram of an electronic terminal provided by an embodiment of the present application. As Figure 4As shown, the electronic terminal includes: at least one processor 401, a memory 402, at least one network interface 403, and a user interface 405. Each component in the device is coupled together through a bus system 404. It can be understood that the bus system 404 is used to implement the connection and communication between these components. In addition to the data bus, the bus system 404 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clear illustration, in Figure 4 all kinds of buses are labeled as the bus system.
[0072] Among them, the user interface 405 may include a display, a keyboard, a mouse, a trackball, a click gun, a key, a button, a touchpad, or a touch screen, etc.
[0073] It can be understood that the memory 402 may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM, Read Only Memory), a programmable read-only memory (PROM, Programmable Read-Only Memory), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static random access memory (SRAM, Static Random Access Memory), synchronous static random access memory (SSRAM, Synchronous Static Random Access Memory). The memory described in the embodiments of the present invention is intended to include but not limited to these and any other suitable categories of memory.
[0074] The memory 402 in the embodiments of the present invention is used to store various categories of data to support the operation of the electronic terminal 400. Examples of these data include: any executable program for operating on the electronic terminal 400, such as an operating system 4021 and application programs 4022; the operating system 4021 contains various system programs, such as a framework layer, a core library layer, a driver layer, etc., for implementing various basic services and processing hardware-based tasks. The application programs 4022 may include various application programs, such as a media player (Media Player), a browser (Browser), etc., for implementing various application services. Implementing the multi-modal data dialogue robot interaction method provided by the embodiments of the present invention may be included in the application programs 4022.
[0075] The method disclosed in the embodiments of the present invention above can be applied to the processor 401 or implemented by the processor 401. The processor 401 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit in the hardware of the processor 401 or the instructions in the form of software. The above-mentioned processor 401 may be a general-purpose processor, a digital signal processor (DSP, Digital Signal Processor), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The processor 401 can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present invention. The general-purpose processor 401 may be a microprocessor or any conventional processor, etc. Combining the steps of the accessory optimization method provided in the embodiments of the present invention can be directly embodied as being executed and completed by the hardware decoding processor, or executed and completed by a combination of the hardware and software modules in the decoding processor. The software module may be located in the storage medium, and this storage medium is located in the memory. The processor reads the information in the memory and combines its hardware to complete the steps of the foregoing method.
[0076] In an exemplary embodiment, the electronic terminal 400 may be an application-specific integrated circuit (ASIC, Application Specific Integrated Circuit), a DSP, a programmable logic device (PLD, ProgrammableLogic Device), a complex programmable logic device (CPLD, Complex Programmable Logic Device), for executing the foregoing method.
[0077] According to the method provided in the embodiments of the present application, the present application also provides a computer program product, which includes: computer program code, when the computer program code runs on a computer, it causes the computer to execute Figure 1 the multi-modal data dialogue robot interaction method in the illustrated embodiment.
[0078] According to the method provided in the embodiments of the present application, the present application also provides a computer-readable storage medium, which stores program code, when the program code runs on a computer, it causes the computer to execute Figure 1 the multi-modal data dialogue robot interaction method in the illustrated embodiment.
[0079] As used in this specification, the terms "component", "module", "system", etc. are used to denote computer-related entities, hardware, firmware, combinations of hardware and software, software, or software in execution. For example, a component can be, but is not limited to, a process running on a processor, a processor, an object, an executable, an execution thread, a program, and / or a computer. By way of illustration, both an application running on a computing device and the computing device can be components. One or more components can reside in a process and / or an execution thread, and a component can be located on one computer and / or distributed between two or more computers. In addition, these components can execute from various computer-readable media storing various data structures. A component can communicate, for example, through a signal with other systems (e.g., the Internet) having one or more data packets (e.g., data from two components interacting with another component in a local system, a distributed system, and / or a network) through local and / or remote processes.
[0080] Those of ordinary skill in the art will appreciate that the various illustrative logical blocks and steps described in connection with the embodiments disclosed herein can be implemented in electronic hardware, or in a combination of computer software and electronic hardware. Whether such functions are implemented in hardware or software depends upon the particular application and design constraints of the technical solution. Skilled artisans may implement the described functions in different ways for each particular application, but such implementation should not be considered to exceed the scope of this application.
[0081] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, apparatuses, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be described herein again.
[0082] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For example, the division of units is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the couplings or direct couplings or communication connections shown or discussed with each other can be through some interfaces, and the indirect couplings or communication connections of devices or units can be in electrical, mechanical, or other forms.
[0083] The unit described as a separation component may or may not be physically separated. The component displayed as a unit may or may not be a physical unit, that is, it may be located in one place or distributed across multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0084] In addition, in each embodiment of the present application, each functional unit may be integrated in a processing unit, or each unit may exist physically alone, or two or more units may be integrated in one unit.
[0085] In the above embodiments, the functions of each functional unit may be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it may be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions (programs). When the computer program instructions (programs) are loaded and executed on a computer, the processes or functions according to the embodiments of the present application are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium may be any available medium that the computer can access or a data storage device such as a server or data center that includes one or more available media integrated. The available medium may be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a high-definition digital video disc (DVD)), or a semiconductor medium (for example, a solid state disk (SSD), etc.).
[0086] If a function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods of various embodiments of this application. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs.
[0087] As described above, the above is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed by this application can easily think of changes or substitutions, which should all be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.
[0088] In summary, this application provides a multi-modal data dialogue robot interaction method, configuration method, terminal, medium, and product. This method is applied to a dialogue robot including an interaction layer, a processing layer, and a model layer, and includes: the interaction layer responds to the user dialogue information input at the current dialogue node and establishes a session connection with the processing layer; the processing layer completes initialization configuration with the model layer and sends the user dialogue information to the model layer; the model layer, based on the prompt information preset for the current dialogue node and the user dialogue information, returns interaction information and multi-modal extension information to the processing layer; the processing layer respectively assembles the formats of the interaction information and the multi-modal extension information and then streams them to the interaction layer. This application expands the boundary of the multi-modal data interaction ability of the dialogue robot, can support the interaction of various types of data, and meets the requirements of different interaction scenarios. Therefore, this application effectively overcomes various disadvantages in the prior art and has high industrial utilization value.
[0089] The above embodiments only illustrate the principles and effects of this application by way of example, rather than limiting this application. Any person familiar with this technology can modify or change the above embodiments without departing from the spirit and scope of this application. Therefore, all equivalent modifications or changes completed by those with ordinary knowledge in the technical field without departing from the spirit and technical idea disclosed by this application should still be covered by the claims of this application.
Claims
1. A multimodal data dialogue robot interaction method, characterized in that, Applied to a dialogue robot including an interaction layer, a processing layer, and a model layer, the method includes: The interaction layer responds to the user dialogue information input at the current dialogue node and establishes a session connection with the processing layer; Initial configuration is completed between the processing layer and the model layer, and the user dialogue information is sent to the model layer; Based on the prompt information preset for the current dialogue node and the user dialogue information, the model layer returns interaction information and multimodal extension information to the processing layer; After the processing layer respectively assembles the formats of the interaction information and the multimodal extension information, it is streamed and sent to the interaction layer.
2. The multimodal data dialogue robot interaction method according to claim 1, wherein The model layer returns interaction information to the processing layer, including: After the user dialogue information triggers the function function of the model layer, the function function configures the corresponding pre-reply information; The model layer sends the pre-reply information as interaction information to the processing layer.
3. The multimodal data dialogue robot interaction method according to claim 1, wherein The multimodal extension information carries a corresponding data assembly structure; after the processing layer respectively assembles the formats of the interaction information and the multimodal extension information and streams them to the interaction layer, it includes: The processing layer assembles the format of the multimodal extension information according to the data assembly structure; After the processing layer sends the interaction information in text format, it streams and sends the assembled multimodal extension information.
4. The multimodal data dialogue robot interaction method according to claim 1, wherein, The multimodal extension information includes preset variable configuration slots; after the model layer assigns values to variables based on the API interfaces pre-configured in the variable configuration slots, it returns interaction information and the multimodal extension information to the processing layer.
5. The multimodal data dialogue robot interaction method according to claim 1, wherein The model layer includes a pre-configured node knowledge base; Based on the node knowledge base, the prompt information preset for the current dialogue node, and the user dialogue information, the model layer returns interaction information and multimodal extension information to the processing layer.
6. The multimodal data dialogue robot interaction method according to claim 1, wherein Before the processing layer sends the user dialogue information to the model layer, the model layer returns the corresponding guiding information to the processing layer based on the prompt information; the processing layer sends the guiding information to the interaction layer.
7. A configuration method for a multi-modal data dialogue robot, characterized in that, Applied to a dialogue robot including an interaction layer, a processing layer, and a model layer, it includes: The dialogue robot receives the prompt information configured at the current dialogue node; The interaction layer responds to the user dialogue information input at the current dialogue node and establishes a session connection with the processing layer; Initial configuration is completed between the processing layer and the model layer, and the user dialogue information is sent to the model layer; Based on the prompt information and the user dialogue information, the model layer returns interaction information and multimodal extension information to the processing layer; After the processing layer respectively assembles the formats of the interaction information and the multimodal extension information, it is streamed and sent to the interaction layer.
8. The configuration method of the multi-modal data dialogue robot according to claim 7, characterized in that, It also includes: The dialogue robot receives the function function configuration information of the current dialogue node, and the function function configuration information includes pre-reply information; The model layer sends the pre-reply information as interaction information to the processing layer.
9. The configuration method of the multi-modal data dialogue robot according to claim 8, wherein The function function configuration information further includes a data assembly structure. After the processing layer performs format assembly on the interaction information and the multimodal extension information respectively, it is sent to the interaction layer in a streaming manner, including: The processing layer performs format assembly on the multimodal extension information according to the data assembly structure; after the processing layer sends the interaction information in text format, it sends the assembled multimodal extension information in a streaming manner.
10. The configuration method of the multi-modal data dialogue robot according to claim 8, characterized in that, The function function configuration information further includes a preset variable configuration slot; after the model layer assigns values to variables based on the API interfaces pre-configured in the variable configuration slot, it returns the interaction information and the multimodal extension information to the processing layer.
11. The configuration method of the multimodal data dialogue robot according to claim 8, characterized in that, The function function configuration information further includes preset guiding information; Before the processing layer sends the user dialogue information to the model layer, the model layer returns corresponding guiding information to the processing layer based on the prompt information; The processing layer sends the guiding information to the interaction layer.
12. A computer program product, characterized in that, The computer program product includes computer program code, and when the computer program code runs on a computer, the computer implements the method according to any one of claims 1 to 6.
13. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method according to any one of claims 1 to 6.
14. An electronic terminal, comprising a memory, a processor, and a computer program stored on the memory, characterized in that, The processor executes the computer program to implement the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Customer service multi-round dialogue intention and slot position identification method and system
CN114186042A
Multi-modal task processing and dialogue task processing method, system and equipment
CN117033585A
Voice interaction method and system based on response words, medium, product and terminal
CN118918896A
Voice robot interaction method and device based on integrated ASR, medium, program product and terminal
CN119091883A
Method for QA with multi-modal information
US11720750B1