Method and apparatus for providing conversational agent using call sequence
By generating a single input-output process for the call sequence, the response generation of the conversational agent is optimized, solving the problems of low efficiency and insufficient accuracy in the existing technology, and achieving cost reduction and improved interactive experience.
Patent Information
- Application Number
- CN202511169809.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-08-20
- Filing Date
- 2025-08-20
- Publication Date
- 2026-03-03
AI Technical Summary
In existing technologies, task-oriented dialogue systems require multiple executions of the language model's input-output process when generating responses, resulting in low efficiency, reliance on local optimization, and a low likelihood that the response will meet the user's purpose.
By generating a call sequence using a pre-learned first language model, obtaining information of interest using this sequence, and generating output text using a pre-learned second language model, a single input-output process is achieved, optimizing the type and structure of the call sequence.
It reduces the cost of using language models, shortens response time, improves the interactive experience of conversational agents, and enhances response accuracy through global optimization.
Smart Images

Figure CN121597790A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to a method and apparatus for providing an interactive agent that utilizes a sequence of calls. Background Technology
[0002] In recent years, the automotive industry has developed rapidly, and automobiles have evolved from simple means of transportation into platforms encompassing a variety of digital functions. In particular, automotive infotainment systems have evolved from simple radios and cassette players into systems capable of providing multimedia, navigation, internet-based services, and smartphone integration. Such systems have become essential for enhancing driver convenience and safety.
[0003] Furthermore, with the development of Natural Language Processing (NLP) technology, services enabling natural dialogue between human users and AI agents are becoming available. These conversational AI services are being integrated into various technology sectors, including the automotive industry, in the form of chatbots or voice recognition assistants.
[0004] In particular, the importance of task-oriented dialogue systems, which utilize AI agents to meet specific user needs, is increasingly evident. AI agents process user input through large language models (LLMs) to generate responses that align with the user's objectives. However, when the internal knowledge of the language model alone is insufficient to meet the user's needs, it becomes necessary to gather information from outside the language model through application programming interfaces (APIs).
[0005] In existing technologies, to handle complex user input, the language model's input-output process is executed multiple times, generating and executing the necessary calls in each process to collect information. However, according to existing technologies, since call generation requires multiple input-output processes of the language model, it suffers from inefficiency in terms of execution time and cost. Furthermore, the calls required to generate a response in existing technologies are not generated all at once but sequentially, thus relying on local optimization rather than global optimization, resulting in a relatively low probability that the response will match the user's intended purpose.
[0006] The aforementioned background technology refers to technical information that the inventors acquired or obtained during the derivation of this invention, and is not necessarily publicly known technology disclosed to the general public before the application for this invention. Summary of the Invention
[0007] The problem the invention aims to solve
[0008] This disclosure provides a method and apparatus for providing a conversational agent utilizing call sequences. It should be noted that the technical problems to be solved by this disclosure are not limited to those mentioned above. Other technical problems and advantages of this disclosure not mentioned above can be understood through the following description and will become clearer through the embodiments of this disclosure. Furthermore, it is understood that the technical problems to be solved and the advantages of this disclosure can all be achieved by the means and combinations thereof described in the claims.
[0009] means for solving problems
[0010] As a technical means to solve the above-mentioned technical problems, the first aspect of this disclosure provides a method for providing a conversational agent using a call sequence, comprising: generating a call sequence based on input text of a vehicle passenger using a pre-learned first language model; obtaining information of interest by executing the call sequence; and generating output text corresponding to the input text using a pre-learned second language model and based on the information of interest, wherein the call sequence includes multiple calls.
[0011] A second aspect of this disclosure provides an apparatus for providing a conversational agent utilizing a call sequence, comprising: a communication module for performing communication; a memory storing at least one program; and a processor that operates by executing the at least one program, the processor being configured to: generate a call sequence using a pre-learned first language model and based on input text from vehicle occupants; control the communication module to acquire information of interest by executing the call sequence; and generate output text corresponding to the input text using a pre-learned second language model and based on the information of interest, the call sequence comprising a plurality of calls.
[0012] A third aspect of this disclosure provides a computer-readable recording medium having a program recorded thereon for performing the methods described in the first aspect of this disclosure on a computer.
[0013] Other aspects, features, and advantages besides those described above will become clear from the following drawings, claims, and detailed description of the invention.
[0014] Invention Effects
[0015] According to the above-described problem-solving methods of this disclosure, by implementing the generation of the call sequence utilized by the response generation of the conversational agent as a single input-output flow of the language model, the cost of using the language model can be reduced and the execution time of the conversational agent required for the response can be shortened.
[0016] Furthermore, according to the problem-solving method of this disclosure, by implementing the generation of the call sequence as a single input-output process of the language model, it is possible to optimize the type of calls constituting the call sequence and the structure of the call sequence based on global optimization.
[0017] Furthermore, according to the problem-solving methods disclosed herein, the process of generating a conversational agent's response can be intuitively conveyed to the user by visualizing the call sequence utilized in generating the response of the conversational agent.
[0018] Furthermore, according to the problem-solving method disclosed herein, by updating the call sequence utilized in generating responses to the conversational agent based on user input, the user can manipulate the call sequence to obtain the desired response, thereby improving the interactive experience with the conversational agent.
[0019] The effects of the embodiments disclosed herein are not limited to those mentioned above, and those skilled in the art can clearly understand other effects not mentioned from the description in this specification. Attached Figure Description
[0020] Figure 1 This is a simplified configuration diagram of the system, including the generating apparatus.
[0021] Figure 2 This is an example of a generation device that provides a method for running a conversational agent that utilizes a sequence of calls.
[0022] Figure 3 This is a schematic diagram illustrating a generation apparatus including a first generation unit, an execution unit, and a second generation unit.
[0023] Figure 4 This is a diagram used to briefly illustrate the process of generating a call sequence based on input text.
[0024] Figure 5 This is a schematic diagram illustrating the process of generating a first input prompt based on the input text as input to the first language model.
[0025] Figure 6 This is a schematic diagram illustrating the process of generating a call sequence based on the output of a first language model.
[0026] Figure 7 This is a diagram illustrating the process of structuring the output of a first language model.
[0027] Figure 8 It is a diagram illustrating the various data used to generate the call sequence.
[0028] Figure 9This is a diagram illustrating the process of generating output text using a second language model.
[0029] Figure 10 This is an example of a method that the generator runs to visualize the call sequence of a conversational agent.
[0030] Figure 11 This is a diagram illustrating the execution of units included in a call sequence.
[0031] Figure 12 This is a diagram illustrating the visual elements corresponding to the unit execution.
[0032] Figure 13 This is a schematic diagram used to illustrate the hierarchical structure of the interface that displays the call sequence.
[0033] Figure 14 This is a diagram illustrating the process of user interaction with the interface that displays the call sequence.
[0034] Figure 15 This is a block diagram of an apparatus according to one embodiment. Detailed Implementation
[0035] The advantages and features of this disclosure, as well as methods for implementing them, will become apparent from the accompanying drawings and detailed description of the embodiments. However, it should be understood that this disclosure is not limited to the embodiments described below, but can be implemented in many different forms and includes all variations, equivalents, or substitutions falling within the spirit and technical scope of this disclosure. The embodiments provided below are intended to complete this disclosure, and this disclosure is provided to fully inform those skilled in the art of the scope of the invention. In describing this disclosure, detailed descriptions of specific known techniques will be omitted if it is determined that such descriptions might obscure the gist of the invention.
[0036] The terminology used in this specification is for describing particular embodiments only and is not intended to limit this disclosure. Unless otherwise defined, all terms used in this specification have the same meaning as commonly understood by those skilled in the art.
[0037] In this specification, unless the context clearly indicates otherwise, singular expressions include plural expressions. Furthermore, terms such as "comprising" or "having" should be understood to specify the presence of features, numbers, steps, actions, elements, components, or combinations thereof described in the specification, without precluding the presence or additional possibilities of one or more other features, numbers, steps, actions, elements, components, or combinations thereof.
[0038] While ordinal terms such as "first" or "second" used in this specification may be used to describe multiple constituent elements, the constituent elements should not be limited by these terms. The terms are used only to distinguish one constituent element from another.
[0039] The phrases "in one embodiment," "according to one embodiment," "with respect to one embodiment," or "implementation based on one embodiment" used in this specification do not necessarily refer to the same embodiment. Furthermore, throughout this specification, the term "embodiment" is merely an arbitrary division for the purpose of illustrating this disclosure, and the various embodiments are not necessarily mutually exclusive. For example, the configuration mentioned to illustrate one embodiment can also be applied and / or implemented in other embodiments, and can be applied and / or implemented after modifications without departing from the scope of this disclosure.
[0040] Some embodiments of this disclosure can be represented by functional block configurations and various processing steps. Some or all of these functional blocks can be implemented in any number of hardware and / or software configurations that perform a specific function. For example, the functional modules of this disclosure can be implemented by more than one microprocessor, or by circuitry for implementing a predetermined function.
[0041] For example, the functional modules of this disclosure can be implemented using various programming languages or scripting languages. Functional blocks can be implemented by algorithms running on more than one processor. Furthermore, this disclosure can employ conventional techniques for electronic environment setup, signal processing, and / or data processing, etc. Terms such as “mechanism,” “element,” “device,” and “configuration” are used extensively and are not limited to mechanical and physical configurations. Additionally, terms such as “part” and “module” refer to a unit that performs at least one function or action, which can be implemented in hardware or software, or a combination of hardware and software.
[0042] Furthermore, the connecting lines or connecting members between the components shown in the figures are merely illustrative of functional connections and / or physical or electrical connections. In actual devices, the connections between multiple components can be represented by various alternative or additional functional connections, physical connections, or electrical connections.
[0043] Furthermore, the dimensions or proportions of some components in the figures may be exaggerated appropriately for illustrative purposes. Also, components shown in some figures may not be shown in others.
[0044] In the following description, "vehicle" can refer to any type of means of transport that has a power unit and is used to transport people or goods, including cars, buses, motorcycles, electric scooters or trucks.
[0045] This disclosure will be described in detail below with reference to the accompanying drawings.
[0046] Figure 1 This is a simplified configuration diagram of the system, including the generating apparatus.
[0047] Reference Figure 1 System 10 may include generation device 100. Generation device 100 of this disclosure refers to an electronic device for providing a conversational agent to a user. In one embodiment, generation device 100 may include means for providing a conversational agent utilizing a sequence of calls, and may include means for visualizing the sequence of calls of the conversational agent. In this case, the conversational agent may include a conversational AI agent utilized by a conversational AI service.
[0048] In this disclosure, a conversational AI agent refers to an interactive interface that uses an AI model to provide users with conversational AI services. Here, conversational AI services refer to AI-based services that enable machines and users to communicate in natural language. Conversational AI services can be implemented through chatbots, virtual assistants, or customer support systems that answer user questions or process user commands.
[0049] In one embodiment, providing a conversational agent may include providing an interactive interface to a user with conversational AI services, thereby providing a conversational agent's response to the user's input.
[0050] According to one embodiment, system 10 may include a vehicle system, and the generation device 100 may be implemented as a component of the vehicle system. The vehicle system may be implemented as at least one electronic device for providing users of the vehicle with various functions and / or information, such as conversational artificial intelligence services.
[0051] In one embodiment, the generation device 100 can acquire input from a user riding in the vehicle (e.g., voice or text input) and generate a response from a conversational agent based on the acquired input. The generation device 100 can display a vehicle-use interface including the response via a display device (not shown) constituting the vehicle system, thereby providing the user with the response of the conversational AI agent. In this case, the vehicle-use interface may include a graphical user interface (GUI).
[0052] On the other hand, the specific process by which the generating device 100 generates dialogue information will be discussed in subsequent references. Figures 2 to 9 Please provide further explanation.
[0053] According to one embodiment, the display device is a means for displaying to a user the interaction between a user and a conversational agent. In one embodiment, the display device may include means for visually displaying the responses of the conversational agent generated by the generation means 100. According to one embodiment, the display device is positioned in a location accessible to the user, such as around the driver's seat of a vehicle, and can visually display the interaction between the user and the conversational agent.
[0054] For example, the display device may include, but is not limited to, a central information display, an instrument panel display, and / or a head-up display (HUD) mounted on the vehicle.
[0055] On the other hand, the generation apparatus 100 according to one embodiment can be implemented as a device mounted inside a vehicle, a server device for managing conversational artificial intelligence services outside the vehicle, a user-portable device, or a combination thereof, to provide a conversational agent.
[0056] For example, the generating device 100 may be a computing device mounted on a vehicle, a server device that provides or manages vehicle software, a user's smartphone, tablet computer, global positioning system (GPS) device, other mobile computing device or non-mobile computing device, etc., but is not limited to these.
[0057] In one embodiment, the generation device 100 can acquire user input and generate a response based on the user input. For example, the generation device 100 can utilize an artificial intelligence model and generate a response corresponding to the user input. During the response generation process, the generation device 100 can utilize information accessible within the vehicle system and / or external information.
[0058] In one embodiment, system 10 may further include an external device 200. The external device 200 of this disclosure refers to a device that provides external information when the generating device 100 is unable to generate a response to user input based solely on information accessible within the vehicle system.
[0059] In one embodiment, external information may include various search results, real-time traffic flow information, specific location information, and / or weather information, but is not limited to these.
[0060] In one embodiment, the generating device 100 can send and receive information by utilizing an external device 200 and a network. Furthermore, components of the vehicle system including the generating device 100 can communicate with each other via a network to send and receive information.
[0061] At this point, "network" refers to a broad data communication network that enables smooth communication between different entities, and can include wired internet, wireless internet, and mobile wireless communication networks. For example, a network can include a Local Area Network (LAN), a Wide Area Network (WAN), a Value Added Network (VAN), a mobile radio communication network, a satellite communication network, and combinations thereof.
[0062] Wired communication can include Ethernet and fiber optic networks. Wireless communication can include, but is not limited to, wireless LAN (Wi-Fi), Bluetooth, Bluetooth Low Energy, ZigBee, Wi-Fi Direct (WFD), ultra-wideband (UWB), infrared data association (IrDA), and near field communication (NFC).
[0063] For example, the generating device 100 can exchange information with the external device 200 via wireless communication, and the generating device 100 and other components of the vehicle system such as the display device can send and receive information via wired communication, but are not limited thereto.
[0064] In one embodiment, the generating device 100 may transmit the response of a conversational agent and / or a vehicle interface including the response to a display device via a network to a display device, which may display the data obtained from the generating device 100.
[0065] Furthermore, in one embodiment, the generating device 100 can acquire various external information from the external device 200 to satisfy the user's purpose predicted from the user input by performing communication using a network.
[0066] Figure 2 This is an example of a generation device that provides a method for running a conversational agent that utilizes a sequence of calls.
[0067] Reference Figure 2 In step 210, the generation device 100 can generate a call sequence using a pre-learned first language model and based on the input text of the vehicle occupants. At this time, the call sequence may include multiple calls.
[0068] In one embodiment, the call sequence may include multiple calls with a nested structure.
[0069] In one embodiment, the generating device 100 can generate a call sequence based on the input text and the passenger's conversation history.
[0070] In one embodiment, the generation device 100 may preprocess the input text based on at least one of entity retrieval, dialogue example retrieval, and prompt template application.
[0071] For example, the generation device 100 can utilize a pre-generated entity database to generate a first retrieval result that includes entities corresponding to at least one string constituting the input text. Furthermore, the generation device 100 can utilize a pre-generated dialogue example database to generate a second retrieval result that includes at least one dialogue example having a similarity of at least a threshold to the input text.
[0072] Subsequently, the generation device 100 can identify at least one call used in at least one dialogue example included in the second search result as the target call. Furthermore, the generation device 100 can apply a pre-generated first prompt template to the input text, the first search result, and the target call to generate a first input prompt for the first language model.
[0073] In one embodiment, the generation apparatus 100 may post-process the output of the first language model based on at least one of parsing and slot normalization.
[0074] For example, the generation device 100 can generate a structured output representing the structure of a string by parsing the output of a first language model expressed as a string. Subsequently, the generation device 100 can use a pre-generated slot normalization database to convert at least one string included in the structured output into a normalized expression, thereby generating a call sequence.
[0075] In step 220, the generating device 100 can obtain information of interest by executing a call sequence.
[0076] In one embodiment, the generating device 100 may execute at least a portion of a plurality of calls in a preset order, or execute at least a portion of a plurality of calls in parallel, thereby obtaining information of interest.
[0077] For example, the generating device 100 can execute multiple calls sequentially based on a depth-first search algorithm to obtain information of interest.
[0078] In step 230, the generation device 100 can use a pre-learned second language model and generate output text corresponding to the input text based on information of interest.
[0079] In one embodiment, the call sequence may include all calls used to generate the output text.
[0080] In one embodiment, the generating device 100 can generate output text based on input text and information of interest.
[0081] In one embodiment, the generation device 100 can update the dialogue history by adding the input text to the dialogue history. Then, the generation device 100 can generate output text based on the updated dialogue history and information of interest.
[0082] In one embodiment, the generation device 100 can generate a second input prompt for a second language model by applying a pre-generated second prompt template to the updated dialogue history and information of interest.
[0083] Figure 3 This is a schematic diagram illustrating a generation apparatus including a first generation unit, an execution unit, and a second generation unit.
[0084] Reference Figure 3 The generation apparatus 100 may include a memory 101, a first generation unit 110, an execution unit 120, and a second generation unit 130.
[0085] In one embodiment, the generation device 100 can obtain input text from users such as vehicle occupants. For example, the generation device 100 can obtain user input such as voice speech and / or text input from users through input interfaces (such as microphones, keyboards, and / or touchscreens) provided in the vehicle.
[0086] In one embodiment, the generation device 100 can directly acquire input text in the form of text input. In another embodiment, the generation device 100 can utilize a speech recognition model to convert speech signals corresponding to spoken utterances into text, thereby generating and acquiring input text.
[0087] In one embodiment, the first generation unit 110 can generate a call sequence based on the input text. The execution unit 120 can obtain information of interest by executing the call sequence generated by the first generation unit 110. The second generation unit 130 can generate output text corresponding to the input text based on the information of interest.
[0088] In one embodiment, the first generation unit 110 may utilize a pre-learned first language model and generate a call sequence based on the user's input text. In this case, the first language model may include a pre-learned language model for performing Natural Language Processing (NLP) tasks. The first language model can be implemented using various language models, such as, but not limited to, Bidirectional Encoder Representations from Transformers (BERT), Generative Pre-trained Transformer (GPT), Transformer, Long Short-Term Memory (LSTM) networks, and XLNet. According to one embodiment, the first language model may include a large language model (LLM) learned based on a large-scale text dataset.
[0089] In this disclosure, a call sequence refers to a set of instructions consisting of at least one call. In one embodiment, a call sequence may include multiple calls. On the other hand, in order for the generation device 100 to generate a response corresponding to the input text, a call according to one embodiment may include retrieving various data from the memory 101 of the generation device 100, other devices within the vehicle system, or external devices 200, or requesting a predetermined operation from the memory 101, other devices within the vehicle system, or external devices 200.
[0090] As an example, a call may include a call to an API endpoint of external device 200 for retrieving data from external device 200. As another example, a call may include a system instruction call for controlling at least a portion of the hardware or software configuration of a vehicle system including generation device 100. As yet another example, a call may include, but is not limited to, a database query call for executing a predetermined query to query or modify information from a database accessible to generation device 100.
[0091] In one embodiment, a call may include a job request for various purposes such as retrieving a specific search term, confirming whether a specific restaurant sells a specific menu item, making a phone call to a specific person, setting an alarm clock, opening or closing a vehicle window, or turning a vehicle's air conditioning on or off.
[0092] In one embodiment, the execution unit 120 can acquire information of interest by executing a sequence of calls. In this disclosure, information of interest refers to information required to generate a response from a conversational agent corresponding to the input text. In one embodiment, the information of interest may include information of a specific type determined by the first generation unit 110 based on the input text.
[0093] Furthermore, in one embodiment, the information of interest may include external information obtained from external device 200. According to one embodiment, the external information obtained from external device 200 may include, but is not limited to, structured text (e.g., JSON or XML).
[0094] For example, when a passenger in a vehicle inputs text that says "Tell me the current weather here", the first generation unit 110 can determine the information of interest required to generate a response as the vehicle's location information and the current weather information at that location, and generate a call sequence for obtaining the vehicle's location information and the current weather information at that location.
[0095] As an example, in order to obtain the vehicle's location information, the first generation unit 110 may generate a call sequence including a first call for requesting the vehicle's location information from a location sensor inside the vehicle system or a first external device that provides the vehicle's location information. Furthermore, in order to obtain the current weather information for that location, the first generation unit 110 may generate a call sequence including a second call for requesting weather information corresponding to the vehicle's current location from a second external device that provides the weather information.
[0096] Subsequently, the execution unit 120 can obtain the vehicle's location information and the current weather information at that location as information of interest by executing the call sequence generated by the first generation unit 110.
[0097] In one embodiment, the execution unit 120 may execute at least a portion of multiple calls in a preset order, or execute at least a portion of multiple calls in parallel, thereby obtaining information of interest.
[0098] In one embodiment, when there is a data dependency between the predetermined calls constituting multiple calls, the execution unit 120 may execute the predetermined calls in a preset order so as to utilize the results of the previous calls in subsequent calls to obtain information of interest.
[0099] For example, when a passenger in a vehicle inputs the text "Tell me the current weather here", the vehicle's current location information needs to be obtained first, and then the weather information needs to be queried based on the determined location. Therefore, the execution unit 120 can first execute the call related to the location information query, and then execute the call related to the weather information query.
[0100] In one embodiment, the execution unit 120 can execute multiple calls sequentially based on a depth-first search algorithm to obtain information of interest. Depth-first search refers to, for a tree-like data structure as the search object, starting from the node where the search begins, searching as deeply as possible to find child nodes until no child nodes remain. When no child nodes exist, the process returns to the previous step and performs a search along another path, thereby searching all nodes.
[0101] That is, in one embodiment, the execution unit 120 can sequentially execute all calls constituting the call sequence in a depth-first search order within a call sequence with a tree structure, thereby obtaining information of interest.
[0102] For example, when a passenger's input text indicates "Tell me the current weather here," the starting node can correspond to a vehicle location query, and the child nodes of the starting node can correspond to location-based weather information queries. The execution unit 120 can execute the calls related to the vehicle location query first, and then execute the calls related to the location-based weather information query, following the search order from the starting node to the child nodes.
[0103] In another embodiment, when there is no data dependency between the predetermined calls constituting multiple calls, the execution unit 120 executes the predetermined calls in parallel for efficient operation, thereby acquiring information of interest.
[0104] For example, when a passenger in a vehicle inputs the text "Please tell me today's weather and news headlines", since there is no data dependency between the weather information query and the news information query, the execution unit 120 can execute the calls related to the weather information query and the news information query in parallel without considering the order.
[0105] In one embodiment, the second generation unit 130 may utilize a pre-learned second language model and generate output text corresponding to the input text based on information of interest. In this case, the second language model may include, but is not limited to, the same language model as the first language model.
[0106] In one embodiment, the second language model may include a language model pre-learned for performing natural language processing tasks. The second language model can be implemented using various language models, such as, but not limited to, Bidirectional Encoder Representations from Transformers (BERT), Generative Pre-trained Transformer (GPT), Transformer, Long Short-Term Memory (LSTM) networks, and XLNet. According to one embodiment, the second language model may include a large language model (LLM) learned based on a large-scale text dataset.
[0107] On the other hand, the call sequence can include all calls made by the second language model to generate the output text.
[0108] In one embodiment, the first generation unit 110 can utilize a single input-output flow of the first language model to generate a call sequence including all calls required for generating the output text, without utilizing multiple repetitions of the first language model's input-output flow. Then, the execution unit 120 obtains information of interest by executing all calls included in the call sequence, and the second generation unit 130 can generate the output text based on the information of interest.
[0109] According to one embodiment of this disclosure, compared to generating the next call to be executed based on information obtained after executing a call, the time required to generate output text from input text can be shortened by reducing the number of input-output processes of the language model used to generate the call, and unnecessary waste of computing resources can be prevented.
[0110] Furthermore, generating the next call based on information obtained after executing one call, which relies on local optimum in the process of generating each call, may result in relatively low accuracy of the output text. Conversely, according to the above embodiment, a call sequence based on global optimum can be generated through a single input-output flow of the language model, thereby improving the accuracy of the output text.
[0111] On the other hand, the memory 101 is hardware used to store various data processed in the generation apparatus 100. The memory 101 can store program code for implementing the first language model or the second language model, as well as various programs for the operation, processing, and control of the first generation unit 110, the execution unit 120, and the second generation unit 130. In addition, the memory 101 can also store various data used or generated by the generation apparatus 100, such as input text, call sequences, and output text.
[0112] In one embodiment, the generation device 100 may generate a dialogue history based on input text. According to one embodiment, the dialogue history may include at least one input text and a response corresponding to each input text. For example, when the generation device 100 generates a first response corresponding to a first input text, the acquired first input text and the generated first response may be stored in memory 101.
[0113] In one embodiment, the first generation unit 110 can generate a call sequence based on the input text and the passenger's dialogue history. For example, the first generation unit 110 can generate a call sequence based on the second input text and the dialogue history consisting of the first input text and the first response. That is, by referring not only to the input text but also to the dialogue history, the first generation unit 110 can consider the dialogue pattern with the user to generate a call sequence suitable for the user.
[0114] For example, the dialogue history may include multiple dialogue contents, such as input text indicating "Tell me the weather," a response including weather information based on the current location as a response to the input text, and input text indicating "Not the current location, please tell me the weather at my workplace." The first generation unit 110 can generate a call sequence by reflecting the previous dialogue context and dialogue patterns that appear in the dialogue history, thereby generating a call sequence for querying weather information based on the workplace for the input text indicating "Tell me the weather."
[0115] In one embodiment, the generation device 100 can update the dialogue history by adding input text to the dialogue history. For example, when the memory 101 stores a dialogue history including a first input text and a first response, the generation device 100 can update the dialogue history by obtaining a second input text and adding a second input text to the dialogue history.
[0116] Figure 4 This is a diagram used to briefly illustrate the process of generating a call sequence based on input text.
[0117] Reference Figure 4 The generating device 100 can generate a calling sequence 400 based on the input text 300. (See reference...) Figure 3The first generation unit 110 of the generation device 100 can generate a call sequence 400 based on the input text 300.
[0118] In one embodiment, the generation device 100 may utilize a pre-learned first language model and generate a call sequence 400 including at least one call based on the input text 300 of the vehicle occupants.
[0119] For example, the generation device 100 can construct input data for the first language model based on the input text 300. Then, the generation device 100 can input the input data into the first language model and obtain the output of the first language model. Afterwards, the generation device 100 can generate a calling sequence 400 based on the output of the first language model.
[0120] On the other hand, the specific process of constructing the input data for the first language model and the specific process of generating the call sequence 400 based on the output of the first language model will refer to Figures 5 to 7 This will be explained in detail later.
[0121] Figure 4 Examples of input text and examples of the output of a first language model corresponding to the input text are shown. As an example, the generation device 100 may, in response to acquiring a first input text 301, input first input data for a first language model based on the first input text 301, and acquire a first output 302 as the output of the first language model.
[0122] As another example, the generation device 100 may respond to acquiring a second input text 303, input second input data for a first language model based on the second input text 303, and acquire a second output 304 as the output of the first language model.
[0123] In one embodiment, the call sequence may include multiple calls with a nested structure. A nested structure refers to a structure where one call includes another.
[0124] For example, when the first input text 301 means "Please help me find a high-speed charging station near Seoul Station," the output of the first language model corresponding to the first input text 301 (i.e., the first output 302) can represent a nested structure of calls related to charging station retrieval, including calls related to location retrieval. In this case, the calls related to charging station retrieval can use charging speed and proximity location conditions as parameters for the retrieval criteria, while the calls related to location retrieval can be understood as calls that are executed first, prior to the calls related to charging station retrieval, in order to set proximity location conditions.
[0125] The generation device 100 can use the first output 302 of the first language model that exhibits a nested structure to generate a call sequence 400 that reflects the nested structure exhibited in the first output 302.
[0126] Figure 5 This is a schematic diagram illustrating the process of generating a first input prompt based on the input text as input to the first language model.
[0127] Reference Figure 5 The generation device 100 can preprocess the input text 300 based on at least one of entity retrieval, dialogue example retrieval, and prompt template application. That is, the generation device 100 can construct input data for the first language model 340 by preprocessing the input text 300.
[0128] In one embodiment, the generating apparatus 100 may use a pre-generated entity database 310 to generate a first search result 311 that includes entities corresponding to at least one string constituting the input text 300.
[0129] In one embodiment, entity database 310 may represent a database that includes real-world names and their attributes that a user may mention as entity information. Entity database 310 may include, but is not limited to, area names or location names, etc.
[0130] For example, when the input text 300 includes "Seoul", the generating device 100 can identify "a city in South Korea" as an entity corresponding to the string "Seoul" by searching the input text 300 in the entity database 310, and generate a first search result 311 based on the identification result.
[0131] In one embodiment, the first search result 311 may include the search results of the entity database 310 generation device 100. By utilizing the search of the entity database 310, the intent and context of the input text 300 can be more accurately identified, thereby making the call sequence and response more likely to match the user's intent.
[0132] In one embodiment, the generating device 100 may utilize a pre-generated dialogue example database 320 to generate a second retrieval result 321 that includes at least one dialogue example having a similarity of more than a threshold to the input text 300.
[0133] In another embodiment, the generation device 100 may utilize the dialogue example database 320 to generate a second search result 321 including a preset number of dialogue examples filtered based on their similarity to the input text 300.
[0134] In one embodiment, the dialogue example database 320 may represent a database including examples of dialogues that may occur between a user and an agent (i.e., dialogue examples). According to one embodiment, a dialogue example may include input text and examples of call sequences corresponding to the input text. Furthermore, a dialogue example may also include a response generated based on the call sequence, and examples of additional input text.
[0135] In one embodiment, the generation device 100 can retrieve at least one dialogue example that has a similarity of more than a threshold to the input text 300 by performing a vector similarity search on the dialogue example database 320. In another embodiment, the generation device 100 can retrieve a preset number of dialogue examples filtered based on their similarity to the input text 300 by performing a vector similarity search on the dialogue example database 320.
[0136] In one embodiment, each dialogue example can be converted into an embedding vector through sentence embedding and stored in the dialogue example database 320. That is, the dialogue example database 320 can map and store each dialogue example with its corresponding embedding vector. Here, sentence embedding refers to expressing the semantics of a sentence as a numerical embedding vector.
[0137] Subsequently, the generation device 100 can generate an embedding vector corresponding to the input text 300 through sentence embedding. The generation device 100 retrieves embedding vectors that have a similarity of more than a threshold with the embedding vector corresponding to the input text 300, or retrieves a predetermined number of dialogue examples based on high similarity with the embedding vector corresponding to the input text 300, thereby retrieving at least one dialogue example from the dialogue example database 320. The generation device 100 can generate a second search result 321 that includes the retrieved dialogue examples.
[0138] At this point, the similarity between embedded vectors can be calculated using cosine similarity, and the similarity above the threshold can be determined based on any value in the range of 0.7 to 0.9. Alternatively, similarity can also be calculated using Euclidean distance, but is not limited to these methods.
[0139] In one embodiment, the second search result 321 may include search results for the generation device 100 of the dialogue example database 320. By utilizing the dialogue example database 320, the generation device 100 can retrieve dialogue examples relevant to the context of the input text 300, and can use the retrieved dialogue examples to generate call sequences and / or responses suitable for the dialogue context.
[0140] In one embodiment, the generation device 100 may identify at least one call used by at least one dialogue example included in the second search result 321 as the target call 322. For example, the generation device 100 may identify all calls used by all dialogue examples included in the second search result 321 as the target call 322.
[0141] In another example, the generation device 100 may identify at least one call selected based on usage frequency from all calls used by all dialogue examples included in the second retrieval result 321 as the target call 322. For example, when the second retrieval result 321 includes multiple dialogue examples, the remaining calls other than those used only in one dialogue example may be identified as the target call 322.
[0142] In one embodiment, the generation device 100 may apply a pre-generated first prompt template 330 to the input text 300, the first search result 311, and the target call 322 to generate a first input prompt 331 for the first language model 340.
[0143] At this point, the first prompt template 330 is a template used to generate input data to be provided to the first language model 340, and can be implemented as a pre-designed structured document or data structure for constituting prompts for the first language model 340. In one embodiment, the first prompt template 330 may be defined as JSON, YAML, or other structured data formats, but is not limited thereto.
[0144] In one embodiment, the generation device 100 may use a first prompt template 330 to combine the elements of the input data constituting the first language model 340, thereby generating input data in a format that the language model can understand. That is, according to one embodiment, the first input prompt 331 may represent the input data of the first language model 340 generated by the generation device 100.
[0145] The first prompt template 330 may include slots for various elements constituting the first input prompt 331, such as the input text 300, the first search result 311, and / or the target call 322. The generation device 100 can generate the first input prompt 331 by inserting the various elements into the slots of the first prompt template 330.
[0146] On the other hand, the first prompt template 330 may also include instructions for generating calls that satisfy the user's purpose predicted from the input text 300. Furthermore, the first input prompt 331 may also include a description for each target call 322.
[0147] On the other hand, with Figure 5 Unlike the previous method, the generation device 100 may directly generate a first input prompt 331 for the first language model 340 without performing preprocessing including entity retrieval, dialogue example retrieval, or prompt template application. For example, the generation device 100 may generate a first input prompt 331 consisting of input text 300 and an instruction indicating "generate the call required to generate a response using the input text".
[0148] In another example, the generation device 100 can generate a first input prompt 331 by performing preprocessing consisting of entity retrieval and dialogue example retrieval, or performing preprocessing consisting of entity retrieval and prompt template application, or performing preprocessing consisting of dialogue example retrieval and prompt template application.
[0149] Figure 6 This is a schematic diagram illustrating the process of generating a call sequence based on the output of a first language model.
[0150] Reference Figure 6 The generation device 100 can post-process the output 341 of the first language model 340 based on at least one of parsing and slot normalization. That is, the generation device 100 can generate the call sequence 400 by post-processing the output 341 of the first language model 340.
[0151] In one embodiment, the generating device 100 may input a first input prompt 331 into the first language model 340 and obtain the output 341 of the first language model 340. Figure 4 The first output 302 and the second output 304 shown are examples of the output 341 of the first language model 340.
[0152] return Figure 6 The generation device 100 can generate a structured output 342 representing the structure of a string by parsing the output 341 of the first language model 340 expressed in string form. The specific process by which the generation device 100 generates the structured output 342 from the output 341 will be discussed in detail below. Figure 7 Please provide an explanation.
[0153] In one embodiment, the generation apparatus 100 may use a pre-generated slot normalization database 350 to transform at least one string included in the structured output 342 into a normalized expression, thereby generating a call sequence 400.
[0154] Slot normalization can be understood as processing synonyms and / or near-synonyms through a unified string representation. Since users can refer to the same concept in multiple ways, the output 341 and / or structured output 342 of the first language model 340 may directly include multiple terms mentioned by the user through the input text 300.
[0155] The slot normalization database 350 is a database used to convert multiple expressions into a unified format according to prescribed rules. For example, expressions such as "gas station", "place to refuel", and "refueling point" can all be normalized to a single format "gas_station". In this case, the slot normalization database 350 can map and store various expressions that can be understood as the same as "gas_station" under the "gas_station" entry.
[0156] When either of the expressions is found in output 341 or structured output 342, the generation device 100 can change the expression to "gas_station" to perform slot normalization on output 341 or structured output 342.
[0157] That is, the generation device 100 can replace at least one unnormalized string expression with a normalized expression by referring to the slot normalization database 350. In this way, the generation device 100 constructs the call sequence 400 with normalized expressions, thereby reducing ambiguity caused by multiple expressions and improving the accuracy of data processing, and improving the consistency and clarity of the response generated based on the call sequence 400.
[0158] On the other hand, with Figure 6 Unlike the previous method, the generation device 100 can generate the call sequence 400 directly from the output 341 itself without performing additional post-processing, or it can generate the call sequence 400 by performing either parsing or slot normalization post-processing on the output 341.
[0159] Figure 7 This is a diagram illustrating the process of structuring the output of a first language model.
[0160] Reference Figure 7The generation device 100 can generate a structured output 342 representing the string structure by parsing the output 341 of the first language model 340 expressed in string form.
[0161] In one embodiment, the output 341 of the first language model 340 may include a string expressed in text format to represent at least one call as an execution object. For example, the output 341 of the first language model 340 corresponding to the input text "Please help me find a high-speed charging station near Seoul Station" may be a string expressed as "search_ev_charging_station(charge_speed = "high-speed", area=search_place(name = "Seoul Station"))".
[0162] In one embodiment, the generation device 100 can generate a structured output 342 representing the structure of a string by parsing the output 341. Here, the structure representing the string refers to the relationship between the various constituent elements expressed in the string, such as a call or parameters.
[0163] In one embodiment, the generating device 100 can perform parsing on the output 341, identify the structure of the output 341 by analyzing the string data constituting the output 341, and split the constituent elements included in the output 341 according to the identified structure.
[0164] For example, output 341 can be expressed as "search_ev_charging_station(charge_speed = "high-speed", area = search_place(name = "Seoul Station"))". The generating device 100 can parse "search_ev_charging_station" into a first call function representing the retrieval of electric vehicle charging stations.
[0165] Furthermore, the generation device 100 can define "charge_speed" and "area" as the first and second parameters of the first calling function, respectively. Additionally, the generation device 100 can define "search_place" as a second calling function used as a second parameter factor. Furthermore, the generation device 100 can define "name" as a parameter of the second calling function. In this case, the second calling function can be understood as nested within the first calling function through the second parameter.
[0166] Furthermore, the generation device 100 can define "high speed" as a factor value of the first parameter of the first calling function, and define "Seoul Station" as a factor value of the parameter of the second calling function.
[0167] That is, in one embodiment, as shown in the example above, the generation device 100 can confirm the structure of the output 341 expressed in string form, and define the constituent elements of the output 341 according to the confirmed structure, thereby generating a structured output 342. At this time, the structured output 342 can be expressed in a data format such as JSON or XML, and can be expressed in a tree structure with the constituent elements of the output 341, such as calls and parameters, as nodes, but is not limited thereto.
[0168] Figure 8 It is a diagram illustrating the various data used to generate the call sequence.
[0169] Reference Figure 8 In the process of generating the call sequence 400 based on the input text 300, at least one of the following can be used: entity database 310, dialogue example database 320, first prompt template 330, slot normalization database 350, and description 360.
[0170] Figure 8 The document shows entity information 315 constituting entity database 310, dialogue example 325 constituting dialogue example database 320, implementation example 335 of first prompt template 330, normalized information 355 constituting slot normalized database 350, and implementation example 365 of description 360.
[0171] In one embodiment, the entity database 310 may consist of pre-collected entity information 315. Entity information 315 may consist of specific entries and the various strings that the entries can represent. For example, entity information 315 may include the entry "Korean festivals," and the strings that "Korean festivals" can represent may include "Spring Festival" and "Chuseok," etc.
[0172] When the input text 300 includes "Spring Festival" or "Chuseok", the generation device 100 can refer to the entity database 310 to generate a first search result 311 including "Korean festivals" as an attribute mapped to the string.
[0173] Furthermore, in one embodiment, the dialogue example interface 320 may consist of dialogue examples 325 from a variety of pre-collected dialogue scenarios. For example, dialogue examples 325 may include examples of input text and examples of the output of a first language model 340 thereof. On the other hand, dialogue examples 325 may also include response examples generated from the output of the first language model 340.
[0174] The generation device 100 can retrieve dialogue examples similar to the input text 300 from the dialogue example database 320 through similarity-based retrieval, thereby generating a second retrieval result 321 that includes at least one dialogue example.
[0175] Furthermore, in one embodiment, the first prompt template 330 can be used as input data to constitute the first language model 340.
[0176] For example, the first prompt template 330 may consist of an instruction indicating "generating the call required to generate a response using the input text" and a slot for the input text 300. The generation device 100 can generate the first input prompt 331 by inserting the input text 300 into the slot for the input text 300.
[0177] In another example, the first prompt template 330 may also include at least one of a reasoning method and an output format. In this case, the reasoning method may define the algorithm or rules used when applying the first language model 340 in generating the output, and the output format may specify the form of expression of the output 341.
[0178] As another example, the first prompt template 330 may also include slots for at least one of various reference data, such as a first search result 311, a second search result 321, a target call 322, and a description of the target call 322. The generation device 100 can generate the first input prompt 331 by inserting reference data into each slot.
[0179] Furthermore, in one embodiment, the slot normalization database 350 may be composed of pre-collected normalization information 355. The normalization information 355 may consist of a normalized expression and various strings that the expression can represent. For example, the normalization information 355 may include the entry "@ac_3" as a normalized expression for three-phase alternating current, and the strings that "@ac_3" can represent may include "ac3 phase", "7 pins", and "ac3 phase 7 pins", etc.
[0180] When the generation device 100 includes "ac3 phase", "7 pins" or "ac3 phase 7 pins" in the output 341 or structured output 342, it refers to the slot normalization database 350 and changes the string to "@ac_3", thereby generating the call sequence 400.
[0181] Furthermore, in one embodiment, description 360 may include comprehensive information about each of the multiple calls, such as the definition, purpose, function, calling method, types of available parameters, format and / or content of the returned response data.
[0182] For example, description 360 may include the names of the called functions and descriptive information about the purpose and functionality of each called function. Furthermore, description 360 may include the name of each parameter that can be used as a factor in the called functions, along with a description of the function and usage of each parameter.
[0183] In one embodiment, the generation device 100 may insert a description of all calls that can be generated from the first language model 340 into the first input prompt 331 to form a call sequence 400. The first language model 340 may generate an output 341 consisting of at least one call that conforms to the purpose of the input text 300, referring to the input text 300 and the description of all calls.
[0184] In another embodiment, the generation device 100 may insert the description of the target call 322 filtered from the second search result 321 as a reference result of the dialogue example database 320 into the first input prompt 331. The first language model 340 may generate an output 341 consisting of at least one call that conforms to the purpose of the input text 300, referring to the input text 300 and the description of the filtered target call 322.
[0185] Figure 9 This is a diagram illustrating the process of generating output text using a second language model.
[0186] Reference Figure 9 The generation device 100 can utilize a pre-learned second language model 520 and generate output text 600 corresponding to the input text 300 based on information of interest 500. At this time, the information of interest 500 can be referenced... Figure 3 The process of acquiring information of interest described above is the same as that of acquiring information of interest by the generating device 100.
[0187] For example, the generation device 100 can acquire information of interest 500 by executing at least one call constituting the call sequence 400. As an example, in the first and second calls constituting the call sequence 400, the generation device 100 can acquire first external information from a first external device by executing the first call, and acquire second external information from a second external device by executing the second call. In this case, the information of interest may include both the first and second external information.
[0188] For reference Figure 3 The call sequence 400 for obtaining the information of interest 500 may include all calls for generating the output text 600. In this way, the generation device 100 does not need to repeat the input-output process of the first language model 340 multiple times, but only needs to utilize the single input-output process of the first language model 340 to generate the call sequence 400 that includes all calls required for generating the output text 600.
[0189] Subsequently, the generating device 100 can obtain all the information of interest 500 required for generating the output text 600 by executing all the calls included in the call sequence 400, and generate the output text 600 based on the obtained information of interest 500.
[0190] In one embodiment, the generation device 100 can generate output text based on input text 300 and information of interest 500. For example, the generation device 100 can construct input data for a second language model 520 based on input text 300 and information of interest 500, and generate output text 600 as the output of the second language model 520 by inputting the input data into the second language model 520.
[0191] In one embodiment, the generation device 100 may generate a second input prompt 511 based on the input text 300 and the information of interest 500. Furthermore, in one embodiment, the generation device 100 may generate the second input prompt 511 based on the input text 300, the calling sequence 400, and the information of interest 500. In this case, the second input prompt 511 can be understood as input data for the second language model 520.
[0192] In one embodiment, the generation device 100 can generate a second input prompt 511 for a second language model 520 by applying a pre-generated second prompt template 510 to at least one of the input text 300, the call sequence 400, and the information of interest 500.
[0193] The second prompt template 510 is a template for generating a second input prompt 511 to be provided to the second language model 520. It can be implemented as a pre-designed structured document or data structure used to constitute the input data of the second language model 520. In one embodiment, the second prompt template 510 can be defined as JSON, YAML, or other structured data formats, but is not limited thereto.
[0194] In one embodiment, the generation device 100 may use the second prompt template 510 to combine the elements used to form the second input prompt 511, thereby generating input data in a format that the language model can understand.
[0195] The second prompt template 510 may include slots for elements constituting the second input prompt 511, such as the input text 300, the call sequence 400, and the information of interest 500. The generation device 100 can generate the second input prompt 511 by inserting the elements into the slots of the second prompt template 510. On the other hand, the second prompt template 510 may also include instructions for generating an answer to the input text 300 using the information of interest 500.
[0196] In one embodiment, the second prompt template 510 may consist of an instruction indicating "generate a natural language response corresponding to the input text using information of interest", a slot for the input text 300, and a slot for the information of interest 500. The generation device 100 generates the second input prompt 511 by inserting the input text 300 and the information of interest 500 into the respective slots.
[0197] In another example, the second prompt template 510 may also include at least one of a reasoning method and an output format. In this case, the reasoning method may define the algorithm or rules used by the second language model 520 to generate the output text 600, and the output format may specify the expression of the output text 600. For example, the output format may include whether to use honorifics and a maximum character limit, but is not limited to these.
[0198] As another example, the second prompt template 510 may also include slots for at least one of various reference data such as the call sequence 400. The generation device 100 can generate a second input prompt 511 by inserting the reference data such as the call sequence 400 into the slots.
[0199] The generation device 100 generates a second input prompt 511 that reflects the call sequence 400, so that the second language model 520 can reflect the process of obtaining the information of interest 500 from the input text 300 to the call sequence 400 and the information of interest 500 in the generation process of the output text 600, thereby improving the correlation between the input text 300 and the output text 600.
[0200] In one embodiment, the generation device 100 can update the dialogue history by adding the input text 300 to the dialogue history. Then, the generation device 100 can generate a call sequence 400 based on the input text 300, obtain information of interest 500 based on the call sequence 400, and generate output text 600 based on the updated dialogue history and the information of interest 500. At this point, the dialogue history can represent a reference... Figure 3 The aforementioned dialogue history.
[0201] In one embodiment, the generating device 100 may generate a second input prompt 511 based on the dialogue history and interest information 500 with the input text 300 added.
[0202] As an example, the generation device 100 can generate a second input prompt 511 for a second language model 520 by applying a pre-generated second prompt template 510 to the updated dialogue history and information of interest 500. According to one embodiment, the second prompt template 510 may include slots containing input text 300 for each dialogue history and information of interest 500. The generation device 100 can generate the second input prompt 511 by inserting elements into the slots of the second prompt template 510.
[0203] By generating a second input cue 511 for the second language model 520 based on dialogue history and information of interest 500, the second language model 520 can generate output text 600 by reflecting previous dialogue context and dialogue patterns that appear in the dialogue history.
[0204] In one embodiment, the generating device 100 may provide the generated output text 600 to a user via a display device. For example, the generating device 100 may transmit the generated output text 600 to the display device, or generate an interface (e.g., a GUI) that includes the output text and transmit it to the display device, thereby providing the output text 600 via the display device.
[0205] As an example, the vehicle system may include a generation device 100 and a display device. The generation device 100 may generate output text 600 based on input text 300 from a user riding in the vehicle and generate a vehicle interface including the output text 600, which is then transmitted to the display device, thereby providing the output text 600 to the user riding in the vehicle.
[0206] Figure 10 This is an example of a method that the generator runs to visualize the call sequence of a conversational agent.
[0207] Reference Figure 10 In step 1010, the generating device 100 can use a pre-learned language model and generate a first output text corresponding to the input text based on the input text of the vehicle occupants.
[0208] In step 1020, the generating device 100 can generate an interface that displays the calling sequence used to generate the first output text. At this time, the interface can display each of the multiple unit executions included in the calling sequence.
[0209] In one embodiment, each of the multiple unit executions may represent a call or a parameter.
[0210] An interface according to one embodiment may display visual elements corresponding to each of the plurality of unit executions, and the visual elements according to one embodiment may include at least one of icons and text.
[0211] According to one embodiment, the interface can utilize multiple visual elements to display the call sequence as a tree structure, and each of the unit executions according to one embodiment can be a node in the tree structure.
[0212] In one embodiment, the generating apparatus 100 can generate an interface for displaying upper-level nodes in the execution of multiple units, and can modify the interface to further display lower-level nodes relative to the upper-level nodes based on received extended input for the upper-level nodes.
[0213] In one embodiment, each of the plurality of visual elements can be determined based on at least one of the description of each node and the factor value of each node.
[0214] In one embodiment, a predetermined unit execution included in a plurality of unit executions can represent information access of a preset type. The generation apparatus 100 can generate an interface in which visual effects corresponding to the preset type of information access are further displayed on the visual elements representing the predetermined unit executions.
[0215] In one embodiment, the generating apparatus 100 can generate an interface for displaying the execution of a predetermined unit included in a call sequence, and can modify the interface to display the details of the execution of the predetermined unit based on input provided by receiving detailed information for the predetermined node.
[0216] At this time, the detailed information may include at least one of the following: information related to the description of the execution of the predetermined unit, information related to the factor value of the execution of the predetermined unit, and information related to the execution result of the execution of the predetermined unit.
[0217] In one embodiment, the generation device 100 can update the call sequence based on user input. The generation device 100 can then generate a second output text corresponding to the input text based on the updated call sequence, and modify the interface to display the second output text and the updated call sequence.
[0218] In one embodiment, the generation apparatus 100 can generate an interface for displaying a first visual element of a predetermined unit execution included in a call sequence. At this time, the generation apparatus 100 can update the call sequence by changing the factor value of the predetermined unit execution in the call sequence based on received change input for the predetermined unit execution. Subsequently, the generation apparatus 100 can modify the interface to display a second visual element indicating that the factor value of the predetermined unit execution has been changed. The second visual element may be different from the first visual element.
[0219] Furthermore, in one embodiment, the generation apparatus 100 may provide at least one change suggestion for the predetermined unit based on receiving a first change input for execution of the predetermined unit. In this case, the generation apparatus 100 may change the factor value of the predetermined unit execution based on receiving a second change input for selecting any one of the at least one change suggestion, thereby updating the call sequence.
[0220] Furthermore, in one embodiment, the generation device 100 can generate an interface for displaying predetermined unit executions included in the call sequence. In this case, the generation device 100 can delete a predetermined unit execution from the call sequence based on receiving a deletion input for that predetermined unit execution, thereby updating the call sequence.
[0221] Furthermore, in one embodiment, the generation device 100 may add predetermined unit execution to the call sequence based on receiving an addition input for predetermined unit execution, thereby updating the call sequence.
[0222] Figure 11 This is a diagram illustrating the execution of units included in a call sequence.
[0223] In one embodiment, the generation device 100 may utilize a pre-learned language model and generate a first output text corresponding to the user's input text 300. In one embodiment, the user may include vehicle occupants.
[0224] On the other hand, the process by which the generating device 100 generates the first output text can be compared with the reference. Figures 3 to 9 The process by which the generating device 100 generates output text 600 based on input text 300 is the same. Furthermore, the language model used to generate the first output text may include at least one of a first language model 340 and a second language model 520.
[0225] For example, the generation device 100 can generate a call sequence 400 based on the input text 300, obtain information of interest 500 based on the call sequence 400, and generate a first output text based on the information of interest 500.
[0226] In one embodiment, the generating device 100 can generate an interface displaying a call sequence 400 for generating first output text. The interface can then display each of the plurality of unit executions 1100 included in the call sequence 400. In one embodiment, each of the plurality of unit executions 1100 can represent a call or a parameter. The call sequence 400 can then be compared with a reference... Figures 3 to 9 The call sequence 400 is the same as described above.
[0227] Each of the multiple unit executions 1100 can represent a call or a parameter. In one embodiment, a call executed as a unit may include the generation device 100 retrieving various data from the memory of the generation device 100 or the external device 200 in order to generate a response corresponding to the input text 300, or requesting a specific job from the memory of the generation device 100 or the external device 200. In one embodiment, a call may represent a function call corresponding to a function.
[0228] As an example, a call may include invoking an API endpoint of external device 200 to retrieve data from external device 200. As another example, a call may include a system instruction call for controlling at least a portion of the hardware or software configuration of a predetermined system (e.g., a vehicle system) including generation device 100. As yet another example, a call may include, but is not limited to, a database query call to execute a predetermined query to query or modify information from a database accessible to generation device 100.
[0229] On the other hand, in one embodiment, the parameters executed as a unit may include variables or constants that serve as input values for the call (implemented by calling a function, etc.).
[0230] As an example, parameters may include those used to transmit conditions for filtering specific data when calling the API endpoint of external device 200. As another example, parameters may include those used to adjust the operation of a specific algorithm executed in a predetermined system including generation device 100. As yet another example, parameters may include those used to specify retrieval conditions when making a database query call.
[0231] In one embodiment, the generating device 100 can classify the call sequence 400 into calls, parameters, and factor values of the parameters, and select multiple units as display objects from the calls and parameters according to the classification results to perform 1100.
[0232] For example, the generation device 100 can select all calls and all parameters as multiple units to execute 1100, or it can select multiple units to execute 1100 according to a preset benchmark. In this case, the preset benchmark can be appropriately set within an intuitive and visual range to achieve the importance of each call, the importance of each parameter, and the maximum number of selections.
[0233] In one embodiment, the generation device 100 can generate an interface displaying the execution 1100 of multiple units included in the call sequence 400 for generating the first output text, thereby visualizing the generation process of the first output text. That is, the generation device 100 can visualize the call sequence 400 used in the process of the conversational agent generating a response to user input.
[0234] Reference Figure 11 The generation device 100 can generate output 341 using a pre-learned first language model 340, and generate a call sequence 400 based on the output 341 of the first language model 340.
[0235] on the other hand, Figure 11 For ease of explanation, call sequence 400 is shown as a tree structure, but call sequence 400 can be implemented using specific text formats such as JSON or XML. Furthermore, according to one embodiment, call sequence 400 can be identical to output 341, or generated by performing post-processing such as structuring and / or slot normalization on output 341.
[0236] The generating device 100 can select from the calls and parameters included in the call sequence 400 a plurality of units as visualization objects, namely, a first call expressed as "search_ev_charging_station" indicating a search for electric vehicle charging stations, a first parameter expressed as "charge_speed" indicating a condition regarding charging speed, a second parameter expressed as "area" indicating a search reference location condition, and a third parameter expressed as "name" indicating a location search term condition for the second parameter.
[0237] Figure 12 This is a diagram illustrating the visual elements corresponding to the unit execution.
[0238] Reference Figure 12 The generating apparatus 100 can generate an interface for displaying visual elements 1200 corresponding to each of the plurality of unit executions 1100. That is, the interface according to one embodiment can display visual elements 1200 corresponding to each of the plurality of unit executions 1100, and the visual elements 1200 according to one embodiment may include at least one of icons 1210 and text 1220.
[0239] In one embodiment, the call sequence 400 can represent a hierarchical structure including parallel and / or nested structures. The generation apparatus 100, by analyzing the structure of the call sequence 400, can interpret it as a tree structure composed of multiple nodes. In this case, the multiple nodes constituting the tree structure can each correspond to a specific unit and be executed.
[0240] For example, refer to Figure 11 When the first call described above is the topmost node, the first parameter and the second parameter can be understood as the lower-level nodes of the first call that have a parallel structure with each other, and the third parameter can be understood as the lower-level node of the second parameter.
[0241] In one embodiment, the generating apparatus 100 can generate an interface that displays a tree structure with each of the plurality of unit executions 1100 as a node. That is, according to one embodiment, the interface can utilize a plurality of visual elements 1200 to display the call sequence 400 in a tree structure, where each of the unit executions in one embodiment can be a node in the tree structure.
[0242] In one embodiment, each of the plurality of visual elements 1200 can be determined based on at least one of a description of each node and a factor value of each node. For example, an icon 1210 corresponding to a node can be determined based on at least one of the node's description and the node's factor value. Furthermore, for example, text 1220 corresponding to a node can be determined based on at least one of the node's description and the node's factor value.
[0243] The following is for reference Figure 12 The example described later is an example of visualizing a call sequence 400 with a first call, a first parameter, a second parameter, and a third parameter as nodes.
[0244] At this point, call sequence 400 is generated during the process of generating a response from the conversational agent based on the input text 300, which represents "Please help me find a high-speed charging station near Seoul Station". Furthermore, the first call, first parameter, second parameter, and third parameter correspond to nodes one through four.
[0245] The description corresponding to the first node can include "electric vehicle charging station search," and the factor values of the first node can include the second to fourth nodes. In this case, the icon 1210 corresponding to the first node can be implemented as an icon that intuitively represents the electric vehicle charging station search. Furthermore, the text 1220 corresponding to the first node can be implemented as "charging station search," etc., which intuitively represents the function of searching for electric vehicle charging stations.
[0246] The generating device 100 can intuitively represent the electric vehicle charging station retrieval performed in response generation based on the input text 300 by generating an interface that displays the visual element 1200 corresponding to the first node.
[0247] Furthermore, the description corresponding to the second node may include "charging speed," and the factor value of the second node may include "high speed" or "standard," etc. In this case, the icon 1210 corresponding to the second node can be implemented as an icon that visually represents high-speed charging by further reflecting the factor value in the description of the second node. Furthermore, the text 1220 corresponding to the second node can be implemented as visually representing "high-speed charging" as a charging station search condition, etc.
[0248] The generation device 100 generates an interface that displays the visual element 1200 corresponding to the second node. The response generation corresponding to the input text 300 takes into account the charging speed condition. In addition, it can intuitively show that the condition is set to high-speed charging.
[0249] Furthermore, the description corresponding to the third node may include "search benchmark location," and the factor value of the third node may include the fourth node. In this case, the icon 1210 corresponding to the third node can be implemented as an icon that visually represents setting the surrounding area of a specific location as the search target. Furthermore, the text 1220 corresponding to the third node can be implemented as "location search," etc., visually representing the location conditions for searching charging stations.
[0250] The generating device 100 can intuitively display the location conditions of charging station retrieval in the response generation corresponding to the input text 300 by generating an interface that displays the visual element 1200 corresponding to the third node.
[0251] Furthermore, the description corresponding to the fourth node may include "location search terms," and the factor values of the fourth node may include location names such as "Seoul Station." In this case, the icon 1210 corresponding to the fourth node can be implemented as an icon that visually represents the search term conditions. Furthermore, the text 1220 corresponding to the fourth node can be implemented as "location search terms," etc., representing search terms used to set location conditions for charging station searches.
[0252] The generating device 100 can intuitively display the search location conditions that reflect the application of the predetermined location search terms in the response generation corresponding to the input text 300 by generating an interface that displays the visual element 1200 corresponding to the fourth node.
[0253] On the other hand, in another embodiment, for input text 300 that means “Please recommend a place to eat when charging at the next service area based on my itinerary”, the generating device 100 can generate a call sequence 400 that includes the user’s schedule search, the search for places of interest on the navigation route, the vehicle battery remaining amount query, and the restaurant search.
[0254] At this time, the generating device 100 can obtain the information of interest 500 by executing the call sequence 400 and generate the first output text "In order not to delay your trip, would you like to order some simple French fries? The charging parking area is area A".
[0255] The generation device 100 can determine the execution of each unit of the call sequence 400 constituting the generation of the first output text (i.e., user's schedule retrieval, location of interest retrieval on navigation path, charging station query within location of interest, and restaurant query within location of interest) as multiple units of the visualization object 1100.
[0256] The generating device 100 can intuitively display, by generating an interface that displays a visual element 1200 corresponding to each of the multiple unit executions 1100, taking into account the user's schedule, confirming that the rest stop to be visited is a place of interest on the path, confirming whether charging is available at the rest stop, and confirming that the available restaurant at the rest stop is open.
[0257] According to one embodiment, the generation device 100 can transmit to the user the specific process of response generation, such as what information was used and what conditions were set in the response of the conversational agent, by visualizing the execution 1100 of multiple units.
[0258] The generation device 100 can visually confirm the process by which the response of the conversational agent is generated by generating an interface that displays the visual elements 1200 of each of the multiple unit executions 1100, thereby improving the interpretability of the operation for the conversational agent for the user.
[0259] In this way, the problem of impaired interpretability of responses caused by the use of conversational agents as black boxes that cannot explain their operation in traditional technologies, as well as the problem of hindering error repair attempts, can be improved.
[0260] On the other hand, in one embodiment, the generation device 100 can generate an interface for displaying visual elements 1200 corresponding to each of the plurality of unit executions 1100. For example, the generation device 100 can arrange and display the interface corresponding to each of the plurality of unit executions 1100 in a predetermined direction around the display position of the first output text. In this way, the user can intuitively confirm the various unit executions used to generate the first output text.
[0261] Figure 13 This is a schematic diagram used to illustrate the hierarchical structure of the interface that displays the call sequence.
[0262] Reference Figure 13 The call sequence 400 can represent a hierarchical structure. For example, the call sequence 400 can represent a tree structure. In one embodiment, the generation device 100 can visualize the hierarchical structure of the call sequence 400 by generating an interface that displays the structured visual elements 1300.
[0263] In one embodiment, the plurality of unit executions 1100 constituting the call sequence 400 may include at least one unit execution corresponding to an upper node 1311 and a unit execution corresponding to a lower node 1321 for the upper node. In this case, the generation device 100 may generate an interface displaying the upper node 1311 in the plurality of unit executions 1100, and, based on received extended input regarding the upper node 1311, modify the interface to further display the lower node 1321 for the upper node 1311.
[0264] For example, the generation device 100 includes a plurality of units executing a first visual element 1310 corresponding to an upper node 1311 in 1100, and can generate an interface that displays a structured visual element 1300 that does not include a second visual element 1320 corresponding to a lower node 1320.
[0265] Subsequently, the generation device 100 can receive extended input for the upper-level node 1311. For example, the generation device 100 can receive extended input in a preset manner, such as selecting a first visual element 1310 corresponding to the upper-level node 1311 through user gestures such as clicking or touching.
[0266] As an example, the structured visual element 1300 can display folded icons in parallel within the first visual element 1310 corresponding to the upper-level node 1311. The generation device 100 can receive extended input from a user who selects a folded icon. At this time, the generation device 100 can change the selected folded icon to an expanded icon based on the received extended input from the user.
[0267] Subsequently, in response to receiving extended input, the generating device 100 generates an interface displaying a structured visual element 1300 including a first visual element 1310 and a second visual element 1320 corresponding to the lower node 1321.
[0268] As an example, the factor value of a parameter expressed as "area" can include a parameter expressed as "name". In this case, the parameter expressed as "area" can be executed by the unit corresponding to the upper-level node 1311 of the parameter expressed as "name", and the parameter expressed as "name" can be executed by the unit corresponding to the lower-level node 1321 of the parameter expressed as "area". On the other hand, the parameter expressed as "area" can represent the retrieval benchmark location condition, and the parameter expressed as "name" can represent the location search term.
[0269] First, the generating device 100 can generate an interface that displays a structured visual element 1300 that includes a first visual element 1310 for representing a retrieval reference location but does not include a second visual element 1320 for representing a location retrieval term.
[0270] Subsequently, based on the extended input received from the user for the first visual element 1310 corresponding to the retrieval reference location, the generating device 100 can generate an interface displaying a structured visual element 1300, which includes a second visual element 1320 representing a location search term as a lower-level node of the retrieval reference location and the first visual element 1310.
[0271] In this way, users can first confirm the simplified hierarchical structure of multiple units, and expand the upper-level nodes that need to be specifically confirmed as needed, thereby confirming the expanded hierarchical structure, and then confirming the lower-level nodes.
[0272] Figure 14 This is a diagram illustrating the process of user interaction with the interface that displays the call sequence.
[0273] Reference Figure 14 It can be confirmed that, as examples of interfaces generated by the generating device 100, there are a first interface 1410 corresponding to the first output text 1411 and a second interface 1420 corresponding to the second output text 1421.
[0274] In the embodiments described later, the interface generated by the generating device 100 may represent the first interface 1410.
[0275] In one embodiment, the execution of predetermined units included in the plurality of unit executions 1100 can represent information access of a preset type. For example, information access of a preset type may include personal information query based on query object information, or query of a specific database based on query object database, etc.
[0276] Another example is that preset types of information access can include location-based access such as current or historical location queries, financial information access such as e-commerce records, medical information access such as medication records, or social media information access such as user-posted posts. These can be classified into specific types of information access.
[0277] On the other hand, when performing a pre-defined information access, a user can request that the results of that type of information access be displayed to them. In one embodiment, the entity providing the conversational AI service or the user can pre-set the type of information access to be displayed to the user.
[0278] When multiple units execute 1100 including the execution of a predetermined unit representing information access of a preset type, the generation device 100 can generate an interface on which the visual element 1412 representing the execution of the predetermined unit also displays a visual effect 1413 corresponding to the information access of the preset type.
[0279] At this time, the visual effect 1413 can be set differently depending on the type of information accessed. According to one embodiment, the visual effect 1413 may include additional objects with specific size, shape, color and / or animation effects displayed around the visual element 1412.
[0280] For example, the first visual effect corresponding to the first type of information access may include displaying a border of a first color around the visual element 1412, and the second visual effect corresponding to the second type of information access may include displaying a border of a second color around the visual element 1412. In this case, the first type may correspond to personal information queries, and the second type may correspond to queries targeting a specific database, but is not limited to these.
[0281] In one embodiment, the generating apparatus 100 can generate an interface for displaying the execution of a predetermined unit included in a call sequence, and can modify the interface to display the detailed information 1414 of the predetermined unit execution based on input received for detailed information on the execution of the predetermined unit.
[0282] Providing detailed information input according to one embodiment may include user input in a preset manner for a predetermined node corresponding to the execution of a predetermined unit. In this case, the manner of providing detailed information input may be set to a short click or short touch on the visual element 1412, but is not limited to this.
[0283] According to one embodiment, the details 1414 may include at least one of information related to a description of the predetermined unit's execution, information related to factor values of the predetermined unit's execution, and information related to the execution result of the predetermined unit's execution.
[0284] Information related to the description of predetermined unit execution according to one embodiment may include a description of the purpose and / or function of the unit execution. Furthermore, information related to factor values of predetermined unit execution according to one embodiment may include parameters included in the unit execution or descriptions related to parameter values. Additionally, information related to the execution result of predetermined unit execution according to one embodiment may include descriptions related to at least a portion of the information obtained as the result of the executed unit execution.
[0285] For example, according to one embodiment, details 1414 may include retrieval provider information as information related to a description or factor value performed by the predetermined unit. The retrieval provider information may include information about a search provider that provides search services using a search engine, such as a web search provider, a specific database retrieval provider, or a geographic retrieval provider.
[0286] The generating device 100 indicates, through detailed information 1414, which retrieval providers are used in the retrieval unit, thereby enabling the user to identify the source of the information.
[0287] As an example of the detailed information provided, a scheduled unit execution could represent a call to query a user's photo library database for a photo of the user with their passport. In this case, the call description could include "personal photo library query," the unit execution factor values could include the query criteria for the photo (e.g., a photo with an ID card), and the execution result of the unit execution could include relevant information about the retrieved photo (e.g., the date the photo was taken) and the passport number.
[0288] At this time, in response to receiving input of detailed information from the visual element 1412 that is selected to represent the user’s photo library for querying photos of the user’s passport, the generating device 100 changes the interface to display detailed information 1414 to show that the passport number was extracted from the photos included in the personal photo library, and the date the photo used to extract the passport number was taken, etc.
[0289] On the other hand, in the embodiments described later, the interface generated by the generating device 100 can represent the second interface 1420.
[0290] In one embodiment, the generation device 100 can update the call sequence 400 based on user input. The generation device 100 can generate a second output text corresponding to the input text 300 based on the updated call sequence 1430, and can change the interface to display the second output text and the updated call sequence 1430.
[0291] At this point, the second information of interest used to generate the second output text can be different from the first information of interest used to generate the first output text. Furthermore, the second output text can be different from the first output text.
[0292] On the other hand, the update may include at least one of addition, deletion, and modification performed by a predetermined unit. That is, the generation device 100 may perform at least one unit that constitutes the call sequence 400 based on user input to change and / or delete, or add at least one unit to the call sequence 400 to generate the updated call sequence 1430.
[0293] At this point, user input methods corresponding to the addition, deletion, and changes performed by the predetermined unit can be pre-set. For example, the user input method corresponding to any of the addition, deletion, and changes performed by the predetermined unit can be appropriately set from various input methods such as click, long press, double click, triple click, swipe, two-finger zoom, single click, double click, and drag, based on user experience.
[0294] In another example, the generation device 100 may generate an interface representing menu areas such as "add," "delete," and "change" in response to receiving user input that selects a predetermined unit to be executed (e.g., user input that selects a visual element corresponding to the predetermined unit execution). Subsequently, the generation device 100 can obtain user input corresponding to any of the addition, deletion, and change performed by the predetermined unit by receiving user input that selects any one of "add," "delete," or "change" displayed in the menu area.
[0295] On the other hand, after generating the updated call sequence 1430, the generating device 100 can obtain information of interest 500 by executing the updated call sequence 1430, and generate a second output text based on the obtained information of interest 500. At this time, since the updated call sequence 1430 may be different from the call sequence 400 before the update, the information of interest 500 obtained by executing the call sequence 400 before the update may be different from the information of interest 500 obtained by executing the updated call sequence 1430.
[0296] In one embodiment, the call sequence 1430 may be generated by a single input-output flow of the first language model 340. In this case, the generation device 100 can add, delete, and / or modify specific units based on user input to generate an updated call sequence 1430. Therefore, it is not necessary to utilize the additional input-output flow of the first language model 340, but can generate a second output text based solely on the second language model 520 and the call sequence 1430.
[0297] In this way, by increasing the operability in the response generation process of the conversational agent, users can directly manipulate the calls used to generate answers to obtain the expected response, thereby improving the user experience of interacting with the conversational agent.
[0298] In one embodiment, the generating apparatus 100 can generate an interface for displaying the execution of a first visual element 1422 of a predetermined unit included in the calling sequence 400. In this case, the generating apparatus 100 can update the calling sequence 400 by changing the factor value of the predetermined unit execution in the calling sequence 400 based on received change input for the predetermined unit execution.
[0299] Subsequently, the generating device 100 can modify the interface to display a second visual element executed by a predetermined unit to indicate that the factor value has been changed. At this time, the second visual element may be different from the first visual element 1422.
[0300] For example, the unit execution corresponding to the first visual element 1422 may include "high-speed charging" as a factor value for the search condition of charging station retrieval, while the unit execution corresponding to the second visual element (i.e., the modified unit execution) may include "slow charging" as a factor value. In this case, the first visual element 1422 may include an icon representing high-speed charging, and the second visual element may include an icon representing slow charging.
[0301] According to one embodiment, the method of changing the input can be set to a specific gesture such as a long press on the first visual element 1422, or it can be set to a voice input requesting a change, but is not limited to these. For example, the change input could be a user's voice statement saying, "Don't look for a high-speed charging station, help me find a slow charging station." Another example is a user's voice statement saying, "Don't look for Seoul Station, help me find one near Gangnam Station."
[0302] Furthermore, in one embodiment, the generation device 100 may provide at least one change suggestion 1423 for the predetermined unit based on receiving a first change input for execution on the predetermined unit. In this case, the generation device 100 may change the factor value of the predetermined unit execution based on receiving a second change input that selects any one of the at least one change suggestion 1423, thereby updating the call sequence 400.
[0303] On the other hand, according to one embodiment, the first change input method can be set to perform a specific gesture such as long press on a visual element corresponding to a predetermined unit, or it can be set to request a change via voice input, but is not limited thereto.
[0304] Furthermore, according to one embodiment, the second change input method can be set to select any one of the suggested change candidates by clicking, touching, and / or voice input, but is not limited thereto.
[0305] For example, the invocation sequence 400 may include the execution of a predetermined unit containing a first factor value. The generation device 100 receives a first change input from a user selecting a visual element 1422 corresponding to the predetermined unit execution, and in response to receiving the first change input, can generate an interface displaying the first factor value and a second factor value as a change suggestion 1423 for the predetermined unit execution. As an example, the first factor value may represent "high-speed charging" related to the charging station search criteria, and the second factor value may represent "slow charging".
[0306] The generation device 100 can receive a second change input from the user selecting a second factor value, and in response to receiving the second change input, change the first factor value of the call sequence 400 to the second factor value, thereby generating an updated call sequence 1430. The generation device 100 can obtain information of interest 500 by executing the updated call sequence 1430, and generate a second output text based on the information of interest 500, thereby reflecting the change in unit execution to provide a response that is more in line with the user's intent.
[0307] Then, the generating device 100 can change the interface that displays visual element 1422 (representing the first factor value) to the interface that displays visual element (representing the second factor value).
[0308] Furthermore, in one embodiment, the generation device 100 can generate an interface for displaying the execution of predetermined units included in the call sequence 400. In this case, the generation device 100 can delete the predetermined unit execution from the call sequence 400 based on receiving a delete input for the predetermined unit execution, thereby updating the call sequence.
[0309] According to one embodiment, the method for deleting input can be set to perform a gesture such as double-clicking on a visual element corresponding to a predetermined unit, or it can be set to request deletion via voice input, but is not limited to these. For example, the deletion input could be a user's voice statement indicating "high speed or low speed doesn't matter".
[0310] In one embodiment, in response to receiving a deletion input for a predetermined unit, the generation device 100 can delete the predetermined unit execution from the call sequence 400 to generate an updated call sequence 1430. The generation device 100 can obtain information of interest 500 by executing the updated call sequence 1430 and generate a second output text based on the information of interest 500, thereby reflecting the deletion of the unit execution to provide a response that is more in line with the user's intent.
[0311] Furthermore, in one embodiment, the generation device 100 may add predetermined unit execution to the call sequence based on receiving an addition input for predetermined unit execution, thereby updating the call sequence.
[0312] According to one embodiment, the method of adding input can be set to clicking or touching a specific area of the interface displaying the call sequence 400 (such as the display area of the first output text), or to adding voice input executed by the request unit, but is not limited to these. For example, the added input could be a user's voice statement expressing "searching for charging stations that provide food".
[0313] In one embodiment, the generation device 100 may receive an addition input for a predetermined unit execution and, in response to receiving the addition input, add the predetermined unit execution to the call sequence 400 to generate an updated call sequence 1430. The generation device 100 can obtain information of interest 500 by executing the updated call sequence 1430 and generate a second output text based on the information of interest 500, thereby reflecting the addition of the unit execution to provide a response that is more in line with the user's intent.
[0314] Figure 15 This is a block diagram of an apparatus according to one embodiment. Figure 15 The device 1500 shown can correspond to Figure 1 The generating apparatus 100 shown.
[0315] Reference Figure 15 The device 1500 may include a communication module 1510, a processor 1520, and a memory 1530. Figure 15 The apparatus 1500 is shown only for the components relevant to the embodiments. Therefore, those skilled in the art will understand that the apparatus 1500, in addition to... Figure 15 In addition to the constituent elements shown, other general constituent elements may also be included.
[0316] The communication module 1510 may include at least one component that enables the device 1500 to communicate with other devices via wired or wireless communication. For example, the communication module 1510 may include a wired communication unit for implementing Ethernet, serial communication, or optical communication, and / or a wireless communication unit for implementing Wi-Fi, Bluetooth, or cellular network-based communication.
[0317] The device 1500 performs wired and / or wireless communication using the communication module 1510, thereby exchanging information with other devices constituting the system 10.
[0318] The processor 1520 controls the overall operation of the device 1500. For example, by executing a program stored in the memory 1530, the processor 1520 can control the communication module 1510, the memory 1530, the input unit (not shown), and / or the output unit (not shown). The processor 1520 can control the operation of the device 1500 by executing a program stored in the memory 1530.
[0319] The processor 1520 can be referenced. Figures 1 to 14The processor 1520 controls at least a portion of the operation of the aforementioned device 1500. For example, the processor 1520 can generate a call sequence based on the input text of vehicle occupants using a pre-learned first language model, and can control the communication module 1510 to execute the call sequence to acquire information of interest, and can generate output text corresponding to the input text based on the pre-learned second language model and the information of interest.
[0320] In another example, processor 1520 uses a pre-learned language model and generates a first output text corresponding to the input text based on the input text of the vehicle occupants, and generates an interface that displays the call sequence used to generate the first output text.
[0321] On the other hand, specific examples and references of processor 1520 operation Figures 1 to 14 The same applies. Therefore, a detailed description of the operation of processor 1520 is omitted below.
[0322] The processor 1520 may be implemented using at least one of application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, micro-controllers, microprocessors, and other electrical units for performing functions.
[0323] On the other hand, such as Figure 3 As shown, the first generation unit 110, the execution unit 120, and the second generation unit 130 constitute at least a part of the functional modules of the generation device 100, and their operation can be realized by the processor 1520 performing arithmetic processing corresponding to each functional module.
[0324] The memory 1530, as hardware within the device 1500 for storing various processed data, can store various programs for the operation, processing, and control of the processor 1520. On the other hand, Figure 15 The memory 1530 shown can be used with Figure 3 The memory 101 shown is the same.
[0325] The memory 1530 may include random access memory (RAM) (e.g., dynamic random access memory (DRAM), static random access memory (SRAM), etc.), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), CD-ROM, Blu-ray disc or other optical disc storage, hard disk drive (HDD), solid state drive (SSD), or flash memory, etc.
[0326] In one embodiment, device 1500 can be a mobile electronic device. For example, device 1500 can be implemented as a smartphone, tablet computer, personal computer, smart TV, personal digital assistant (PDA), laptop computer, media player, device with camera, and other mobile electronic devices. Furthermore, device 1500 can be implemented as a wearable device with communication and data processing capabilities, such as a watch, glasses, hairband, and ring.
[0327] In another embodiment, device 1500 may be an electronic device embedded in a vehicle. For example, device 1500 may be an electronic device embedded in the vehicle during the vehicle manufacturing process, or an electronic device integrated with the vehicle after the manufacturing process through tuning.
[0328] In yet another embodiment, device 1500 may be a server located outside the vehicle. The server can communicate via a network and is implemented as at least one computing device to provide commands, code, files, content, services, etc.
[0329] In one embodiment, the process executed on device 1500 may be executed by at least some of the following: mobile electronic devices, electronic devices embedded in a vehicle, and servers located outside the vehicle.
[0330] On the other hand, embodiments of this disclosure can be implemented in the form of a computer program executable on a computer through various components, and such a computer program can be recorded on a computer-readable medium. In this case, the medium may include, but is not limited to: magnetic media such as hard disks, floppy disks, and magnetic tapes; optical recording media such as CD-ROMs and DVDs; magneto-optical media such as floppy disks; and hardware devices specifically used for storing and executing program instructions such as ROMs, RAMs, and flash memory.
[0331] On the other hand, the computer program may be a program specifically designed and configured for this disclosure, or it may be a program known and available to those skilled in the art of computer software. Examples of computer programs may include not only machine language code generated by a compiler, but also high-level language code that can be executed by a computer using an interpreter or the like.
[0332] According to one embodiment, methods according to various embodiments of this disclosure may be provided in a computer program product. The computer program product, as a commodity, can be traded between a seller and a buyer. The computer program product may be distributed (e.g., downloaded or uploaded) in the form of a device-readable storage medium (e.g., a CD-ROM, compact disc read-only memory) or through an application store (e.g., the Play Store™) or directly or online between two user devices. In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily created in a machine-readable storage medium, such as a manufacturer's server, an application store's server, or a relay service storage.
[0333] The steps constituting the method according to this disclosure may be performed in any suitable order unless the order is explicitly stated or otherwise. This disclosure is not necessarily limited to the order in which the steps are described above. All examples or exemplary terms used in this disclosure (e.g., etc.) are merely for the purpose of detailing the disclosure, and the scope of this disclosure is not limited by the examples or exemplary terms above unless limited by the claims. Furthermore, those skilled in the art will understand that various modifications, combinations, and variations can be made within the scope of the appended claims or their equivalents, depending on design conditions and factors.
[0334] Therefore, the ideas of this disclosure should not be limited to the above embodiments, and the scope of this disclosure includes not only the appended claims, but also all scopes that are equivalent to or modified by those claims.
Claims
1. A method for providing a dialogic proxy utilizing a sequence of calls, characterized in that, include: The steps involve using a pre-learned first language model and generating a call sequence based on the input text from vehicle occupants. The steps of obtaining information of interest by executing the call sequence, and The steps of using a pre-learned second language model and generating output text corresponding to the input text based on the information of interest; The call sequence includes multiple calls.
2. The method for providing a dialogic proxy utilizing a call sequence according to claim 1, characterized in that, The call sequence includes all calls used to generate the output text.
3. The method for providing a dialogic proxy utilizing a call sequence according to claim 1, characterized in that, The call sequence includes multiple calls with a nested structure.
4. The method for providing a conversational proxy utilizing a call sequence according to claim 1, characterized in that, The step of generating the call sequence further includes: The step of preprocessing the input text based on at least one of entity retrieval, dialogue example retrieval, and prompt template application.
5. The method for providing a conversational proxy utilizing a call sequence according to claim 4, characterized in that, The preprocessing steps include: The step of generating a first retrieval result comprising entities corresponding to at least one string constituting the input text, using a pre-generated entity database; The step of generating a second retrieval result, which includes at least one dialogue example that has a similarity of more than a threshold to the input text, using a pre-generated database of dialogue examples; The step of identifying at least one call used in at least one dialogue example included in the second search results as the target call; and The step of generating a first input prompt for the first language model by applying a pre-generated first prompt template to the input text, the first search result, and the target.
6. The method for providing a conversational proxy utilizing a call sequence according to claim 1, characterized in that, The steps for generating the call sequence include: The step of post-processing the output of the first language model based on at least one of parsing and slot normalization.
7. The method for providing a conversational proxy utilizing a call sequence according to claim 6, characterized in that, The post-processing steps include: The steps include: generating a structured output representing the structure of the string by parsing the output of the first language model expressed as a string; and... The step of generating the call sequence is to transform at least one string included in the structured output into a normalized expression using a pre-generated slot normalization database.
8. The method for providing a conversational proxy utilizing a call sequence according to claim 1, characterized in that, In the step of obtaining the information of interest The information of interest is obtained by executing at least a portion of the multiple calls in a preset order or by executing at least a portion of the multiple calls in parallel.
9. The method for providing a conversational proxy utilizing a call sequence according to claim 8, characterized in that, In the step of obtaining the information of interest Based on the depth-first search algorithm, multiple calls are executed sequentially to obtain the information of interest.
10. The method for providing a dialogic proxy utilizing a call sequence according to claim 1, characterized in that, In the step of generating the call sequence The call sequence is generated based on the input text and the passenger's conversation history.
11. The method for providing a dialogic proxy utilizing a call sequence according to claim 10, characterized in that, In the step of generating the output text The output text is generated based on the input text and the information of interest.
12. The method for providing a dialogic proxy utilizing a call sequence according to claim 11, characterized in that, The steps for generating the call sequence include: The step of updating the conversation history by adding the input text to the conversation history. In the step of generating the output text The output text is generated based on the updated dialogue history and the information of interest.
13. The method for providing a dialogic proxy utilizing a call sequence according to claim 12, characterized in that, The steps for generating the output text also include: The step of generating a second input prompt for the second language model by applying a pre-generated second prompt template to the updated dialogue history and the information of interest.
14. An apparatus for providing a dialogic agent utilizing a sequence of calls, characterized in that, include: The communication module performs communication. The memory stores at least one program, and The processor runs by executing the at least one program; The processor is configured to: The system utilizes a pre-learned first language model and generates call sequences based on the input text from vehicle occupants. Control the communication module to obtain information of interest by executing the call sequence. The output text corresponding to the input text is generated using a pre-learned second language model and based on the information of interest. The call sequence includes multiple calls.
15. A computer-readable recording medium, characterized in that, The computer-readable recording medium contains a program for performing the method according to claim 1 on a computer.