Method and apparatus for generating prompt template corresponding to input text

KR1020260138908APending Publication Date: 2026-09-2142DOT INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
KR1020250032280
Authority / Receiving Office
KR · KR
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-03-12
Publication Date
2026-09-21

Smart Images

  • Figure PAT00016_ABST
    Figure PAT00016_ABST
Patent Text Reader

Abstract

The present disclosure relates to a method and apparatus for generating a prompt template corresponding to input text. According to one embodiment of the present disclosure, a method for generating a prompt template corresponding to input text may be provided, comprising: generating a text embedding based on input text of a vehicle occupant; obtaining a search result by performing a search based on the text embedding on at least a portion of a previously generated dialogue example database; selecting at least one response guideline from a plurality of previously generated response guidelines based on the search result; and generating a prompt template corresponding to the input text based on the at least one response guideline.
Need to check novelty before this filing date? Find Prior Art

Description

Technology Field

[0001] The present disclosure relates to a method and apparatus for generating a prompt template corresponding to input text. Background Technology

[0002] The automotive industry has been developing rapidly in recent years, and vehicles are evolving beyond mere means of transportation into platforms that incorporate various digital functions. In particular, in-vehicle infotainment systems have evolved from simple radios and cassette players into systems offering diverse capabilities, including multimedia, navigation, internet-based services, and smartphone connectivity. Such systems are becoming essential elements for enhancing driver convenience and safety.

[0003] In addition, advancements in natural language processing (NLP) technology have made it possible to provide services that offer natural conversations between human users and artificial intelligence agents. These conversational AI services are being integrated into various technological fields, including the automotive industry, in the form of chatbots or voice recognition assistants.

[0004] In particular, the importance of task-oriented dialogue systems, which utilize artificial intelligence agents to satisfy specific user needs, is emerging. AI agents process user input using Large Language Models (LLMs) to generate responses that meet the user's objectives. However, the performance of a language model is significantly affected by how its prompts are configured.

[0005] Conventional technologies have faced difficulties in generating responses optimized for various situations due to fixed or simple prompt designs. Particularly in goal-oriented dialogue systems, inefficient prompt designs can lead to significant performance degradation by introducing uncertainty into the language model. Therefore, there is a need for research and development of prompt designs capable of optimizing for diverse scenarios.

[0006] The aforementioned background technology is technical information that the inventor possessed for the derivation of the present invention or acquired during the process of deriving the present invention, and it cannot be considered as prior art disclosed to the general public prior to the filing of the present invention. The problem to be solved

[0007] The present disclosure provides a method and apparatus for generating a prompt template corresponding to input text. The problems to be solved by the present disclosure are not limited to those mentioned above, and other problems and advantages of the present disclosure not mentioned can be understood from the following description and will be more clearly understood from the embodiments of the present disclosure. Furthermore, it will be understood that the problems and advantages to be solved by the present disclosure can be realized by the means and combinations thereof set forth in the claims. means of solving the problem

[0008] As a technical means for achieving the technical problem described above, the first aspect of the present disclosure may provide a method for generating a prompt template corresponding to an input text, comprising: generating a text embedding based on input text of a vehicle occupant; obtaining a search result by performing a search based on the text embedding on at least a portion of a previously generated dialogue example database; selecting at least one response guideline from a plurality of previously generated response guidelines based on the search result; and generating a prompt template corresponding to the input text based on the at least one response guideline.

[0009] A second aspect of the present disclosure may provide an apparatus for generating a prompt template corresponding to an input text, comprising: a memory in which at least one program is stored; and a processor that operates by executing said at least one program, wherein the processor generates a text embedding based on input text of a vehicle occupant, obtains a search result by performing a search based on said text embedding on at least a portion of a previously generated dialogue example database, selects at least one response guideline from a plurality of previously generated response guidelines based on said search result, and generates a prompt template corresponding to said input text based on said at least one response guideline.

[0010] A third aspect of the present disclosure may provide a computer-readable recording medium having a program for executing the method of the first aspect of the present disclosure on a computer.

[0011] Other aspects, features, and advantages other than those described above will become clear from the following drawings, claims, and detailed description of the invention. Effects of the invention

[0012] According to the problem-solving means of the present disclosure described above, by optimizing the prompt of a language model according to various situations, a response that matches the user's intention with high accuracy can be provided.

[0013] In addition, according to the means for solving the problem of the present disclosure, response speed and cost efficiency can be maximized by minimizing the number of times input and output of the language model are performed through prompt optimization.

[0014] The effects of the embodiments of the present disclosure are not limited to the effects mentioned above, and other unmentioned effects will be clearly understood by those skilled in the art from the description in this specification. Brief explanation of the drawing

[0015] The following drawings attached to this specification are intended to illustrate at least one embodiment according to the present disclosure and serve to further enhance understanding of the technical concept of the present disclosure together with the detailed description of the invention set forth below; therefore, the present disclosure should not be interpreted as being limited only to the matters described in the drawings. FIGS. 1 and 2 are exemplary drawings for explaining an environment in which an interactive artificial intelligence service is provided. FIG. 3 is a block diagram of an electronic device according to one embodiment. Figure 4 is an example of how an electronic device operates. FIG. 5 is an exemplary drawing for explaining a conversation method according to one embodiment. FIG. 6 is an exemplary drawing for illustrating an electronic device including a first generating unit, an execution unit, and a second generating unit. FIG. 7 is an example of a method for configuring the input of a language model according to one embodiment. Figure 8 is an exemplary diagram illustrating a conversation example database. FIG. 9 is an example of a method for generating a prompt template corresponding to input text according to one embodiment. Figure 10 is an exemplary diagram illustrating the process of filtering multiple conversation rules. FIGS. 11 and 12 are exemplary drawings for explaining the process of generating a prompt template corresponding to input text based on at least one conversation rule. FIG. 13 is an exemplary diagram illustrating an input prompt including a situational prompt and a fixed prompt. FIG. 14 is an exemplary diagram illustrating the process of generating an input prompt by applying a prompt template and generating output text based on the input prompt. FIG. 15 is an exemplary drawing for illustrating a conversation history according to one embodiment. FIG. 16 is an example of a method for generating a prompt template corresponding to input text according to one embodiment. Figure 17 is an exemplary diagram illustrating the process of performing a search based on text embeddings. FIG. 18 is an exemplary diagram illustrating the process of obtaining search results. FIG. 19 is an exemplary diagram illustrating the process of selecting at least one response guideline from a plurality of previously generated response guidelines based on search results. FIG. 20 is an exemplary drawing for explaining a method of using search results according to one embodiment. FIG. 21 is an exemplary drawing for explaining the process of generating output text according to one embodiment. Specific details for implementing the invention

[0016] The advantages and features of the present disclosure and the methods for achieving them will become clear by referring to the embodiments described in detail together with the accompanying drawings. However, the present disclosure is not limited to the embodiments presented below, but can be implemented in various different forms and should be understood to include all modifications, equivalents, and substitutions that fall within the spirit and scope of the present disclosure. The embodiments presented below are provided to make the present disclosure complete and to fully inform those skilled in the art of the scope of the invention. In describing the present disclosure, detailed descriptions of related prior art are omitted if it is determined that such detailed descriptions may obscure the essence of the present disclosure.

[0017] The terms used herein are used merely to describe specific embodiments and are not intended to limit the disclosure. Unless otherwise defined, all terms used herein have the same meaning as generally understood by those skilled in the art to which this disclosure pertains.

[0018] In this specification, singular expressions include plural expressions unless the context clearly indicates otherwise. Furthermore, terms such as "comprising" or "having" are intended to specify the existence of the features, numbers, steps, actions, components, parts, or combinations thereof described in the specification, and should be understood as not precluding the existence or addition of one or more other features, numbers, steps, actions, components, parts, or combinations thereof.

[0019] Additionally, terms including ordinal numbers, such as "first" or "second" as used herein, may be used to describe various components, but the components should not be limited by the terms. The terms are used solely for the purpose of distinguishing one component from another.

[0020] Phrases such as "in one embodiment," "according to one embodiment," "related to one embodiment," or "according to an implementation of one embodiment" in this specification do not necessarily refer to the same embodiment. Furthermore, throughout this specification, "examples" are arbitrary distinctions to facilitate the description of the present disclosure, and each embodiment does not need to be mutually exclusive. For example, configurations mentioned for the description of one embodiment may be applied and / or implemented in other embodiments, and may be modified and applied and / or implemented to the extent that they do not depart from the scope of the present disclosure.

[0021] Some embodiments of the present disclosure may be represented by functional block configurations and various processing steps. Some or all of these functional blocks may be implemented by various numbers of hardware and / or software configurations that perform specific functions. For example, the functional blocks of the present disclosure may be implemented by one or more microprocessors or by circuit configurations for a specific function.

[0022] For example, the functional blocks of the present disclosure may be implemented in various programming or scripting languages. The functional blocks may be implemented as algorithms executed on one or more processors. Additionally, the present disclosure may employ prior art for electronic configuration, signal processing, and / or data processing, etc. Terms such as "mechanism," "element," "means," and "configuration" may be used broadly and are not limited to mechanical and physical configurations. Furthermore, terms such as "-part," "-module," etc. refer to a unit that processes at least one function or operation, which may be implemented in hardware or software, or as a combination of hardware and software.

[0023] Furthermore, the connecting lines or connecting members between the components depicted in the drawings are merely illustrative of functional connections and / or physical or circuit connections. In the actual device, connections between components may be represented by various alternative or added functional connections, physical connections, or circuit connections.

[0024] In addition, some components in the drawings may be depicted with their size or proportions slightly exaggerated. Also, components depicted in one drawing may not be depicted in another drawing.

[0025] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the attached drawings.

[0026] FIGS. 1 and 2 are exemplary drawings for explaining an environment in which an interactive artificial intelligence service is provided.

[0027] Referring to FIG. 1, a conversational system (1) according to one embodiment may include a user (10) receiving a conversational artificial intelligence service and a conversational means (20) used to provide a conversational artificial intelligence service to the user (10).

[0028] In the present disclosure, the conversation means (20) refers to at least one device implementing a conversational artificial intelligence agent. At this time, the conversational artificial intelligence agent may include an interaction interface for providing a conversational artificial intelligence service to a user (10) using an artificial intelligence model.

[0029] Conversational AI service refers to various AI-based services that enable a machine and a user (10) to communicate in natural language, and conversational AI service can be implemented as a chatbot service that answers questions from the user (10) or processes commands from the user (10), a virtual assistant service, or a customer support system.

[0030] In one embodiment, providing a conversational artificial intelligence agent may include providing a conversational artificial intelligence service interaction interface to a user (10), thereby providing a response from the conversational artificial intelligence agent to the user (10) regarding the user's (10) input. In this case, the response may be implemented by providing output text or a voice signal corresponding to the output text, but is not limited thereto, and may comprehensively represent the processing results of the conversational artificial intelligence agent regarding the user's (10) input, such as displaying search results or a series of commands for physically operating a mechanical device.

[0031] That is, the conversation means (20) may refer to any means of providing a conversational artificial intelligence service to the user (10) by interacting with the user (10) on the side of the user (10), and the conversation means (20) according to one embodiment may be implemented including a display and / or a speaker, etc., that provides at least some elements of the conversational artificial intelligence service audiovisually. In addition, the conversation means (20) may be implemented including, but is not limited to, a touchscreen, a microphone, and / or a keypad, etc., that acquire user input.

[0032] In one embodiment, the conversation system (1) may include a transportation system using a vehicle. For example, the user (10) may include a passenger on board the vehicle. As an example, the user (10) may include a driver operating the vehicle, but is not limited thereto.

[0033] In this case, a vehicle may refer to any type of means of transportation used to move people or goods using an engine, such as a car, bus, motorcycle, kickboard, or truck.

[0034] Additionally, the vehicle may include autonomous vehicles. In this case, the term "autonomous vehicle" may comprehensively refer to a vehicle in which at least a part of the driving functions are automated, rather than a vehicle in which the entire driving function is automated. As an example, an autonomous vehicle may include a vehicle equipped with an Advanced Driver Assistance System (ADAS), such as a hazardous situation detection system.

[0035] A conversation means (20) according to one embodiment may constitute at least a part of a vehicle. For example, the conversation means (20) may include a vehicle display that displays a vehicle interface containing a response from a conversational artificial intelligence agent. In this case, the vehicle interface may include a graphical user interface (GUI).

[0036] Referring to FIG. 2, the conversational system (1) may include an electronic device (30). In the present disclosure, the electronic device (30) refers to at least one device that generates a response of a conversational artificial intelligence agent corresponding to the input based on the input of a user (10).

[0037] In one embodiment, the electronic device (30) can obtain input from the user (10) through the conversation means (20). For example, the conversation means (20) may obtain the user's (10) utterance as a voice signal, and the electronic device (30) may obtain the voice signal from the conversation means (20) and generate input text from the voice signal. At this time, the electronic device (30) may generate text based on the voice signal using ASR (Automatic Speech Recognition), but is not limited thereto.

[0038] In another example, the conversation means (20) can obtain input text from a user (10) who writes text directly, and the electronic device (30) can obtain input text from the user (10) from the conversation means (20).

[0039] When the conversation system (1) is implemented as a transportation system using a vehicle, the electronic device (30) according to one embodiment may be implemented as a device mounted inside the vehicle, but is not limited thereto. For example, the electronic device (30) may be implemented as a server device of an entity supplying or managing vehicle software, a smartphone, tablet PC, GPS (global positioning system) device, other mobile or non-mobile computing device, etc. of a user (10), but is not limited thereto.

[0040] Meanwhile, the electronic device (30) according to one embodiment may be implemented as a plurality of electronic devices, for example, the electronic device (30) may be implemented as a combination of at least some of a device mounted inside a vehicle to provide a conversational artificial intelligence agent, a server device that manages a conversational artificial intelligence service outside the vehicle, and a device that is portable by a user (10).

[0041] In one embodiment, the electronic device (30) can use a previously trained artificial intelligence model to generate a response from an interactive artificial intelligence agent corresponding to the input text based on the input text of the user (10).

[0042] Meanwhile, the conversation system (1) may further include at least one external device (40). In one embodiment, the electronic device (30) may use at least one external device (40) in the process of generating a response based on the input text of the user (10).

[0043] In one embodiment, the external device (40) may include an external server that supports various natural language processing using a language model that is not directly stored by the electronic device (30). For example, the external device (40) may include an external server that hosts a Large Language Model (LLM) and can return a response based on data obtained from the electronic device (30).

[0044] Additionally, in one embodiment, the external device (40) may include a device that provides external information when the electronic device (30) cannot generate an appropriate response to the input text using only information accessible within a local environment (e.g., a single vehicle system). For example, the external information may include, but is not limited to, various search results, information regarding real-time traffic flow, information regarding specific locations and / or weather information.

[0045] In one embodiment, the electronic device (30) may use a plurality of external devices (40) to generate a response. As an example, the electronic device (30) may perform an API call to a first external device (41) for natural language processing using a large-scale language model, and may perform an API call to a second external device (42) for obtaining external information such as various search results.

[0046] In one embodiment, the electronic device (30) can exchange information by communicating with at least one external device (40) using a network. Additionally, components of a local environment including the electronic device (30) can exchange information by communicating with each other using a network.

[0047] In this case, the network is a comprehensive data communication network that enables different entities to communicate smoothly with each other, and may include wired internet, wireless internet, and mobile wireless communication networks. For example, the network may include a Local Area Network (LAN), a Wide Area Network (WAN), a Value Added Network (VAN), a mobile radio communication network, a satellite communication network, and combinations thereof.

[0048] Wired communication may include Ethernet and fiber optic networks. Additionally, wireless communication may include, for example, Wi-Fi, Bluetooth, Bluetooth Low Energy, ZigBee, Wi-Fi Direct (WFD), ultra-wideband (UWB), infrared data association (IrDA), and near field communication (NFC), but is not limited thereto.

[0049] For example, the electronic device (30) may exchange information with an external device (40) using wireless communication, and components of the local environment, such as the electronic device (30) and the communication means (20), may exchange information using wired communication, but are not limited thereto.

[0050] In one embodiment, the electronic device (30) can transmit the response of the conversational artificial intelligence agent to the conversation means (20) by performing communication using a network, and the conversation means (20) can provide the data obtained from the electronic device (30) to the user (10).

[0051] In addition, in one embodiment, the electronic device (30) can obtain various external information from an external device (40) to satisfy the user's (10) objectives predicted from the user's (10) input text by performing communication using a network.

[0052] FIG. 3 is a block diagram of an electronic device according to one embodiment.

[0053] Referring to FIG. 3, the electronic device (30) may include an input / output interface (31), a communication interface (32), a memory (33), and a processor (34). Only the components related to the embodiments of the present disclosure are shown in the electronic device (30) of FIG. 3. Therefore, a person skilled in the art will understand that the electronic device (30) may include other general-purpose components in addition to the components shown in FIG. 3.

[0054] The input / output interface (31) may include at least one component used by the electronic device (30) to acquire input text and provide a response. In one embodiment, when the electronic device (30) is implemented in conjunction with the conversation means (20), the input / output interface (31) may be understood as being identical to the conversation means (20).

[0055] For example, the input / output interface (31) may include an input interface such as a keypad, microphone, and / or camera, and an output interface such as a display, speaker, and / or vibration module, but is not limited thereto, and may be implemented in a form where the input interface and the output interface are integrated, such as a touchscreen.

[0056] Additionally, if the communication means (20) and the electronic device (30) are implemented as separate devices, the input / output interface (31) may represent a communication channel or wired / wireless connection element for exchanging data with the communication means (20).

[0057] The communication interface (32) may include at least one component that enables the electronic device (30) to perform wired / wireless communication with another device in the process of generating a response corresponding to the input text. For example, the communication interface (32) may include a wired communication unit for implementing Ethernet, serial communication, or optical communication and / or a wireless communication unit for implementing Wi-Fi, Bluetooth, or cellular network-based communication.

[0058] The electronic device (30) can exchange information with other devices constituting the conversation system (1), such as an external device (40), by performing wired communication and / or wireless communication using the communication interface (32).

[0059] The memory (33) is hardware that stores various data processed within the electronic device (30) and can store programs for various operations, processing, and control of the processor (34). Additionally, the memory (33) can store at least temporarily the various data mentioned in the present disclosure as data processing targets of the processor (34).

[0060] Memory (33) may include, but is not limited to, RAM (Random Access Memory), ROM (Read-Only Memory), EEPROM (Electrically Erasable Programmable Read-Only Memory), CD-ROM, Blu-ray or other optical disc storage, HDD (Hard Disk Drive), SSD (Solid State Drive), or flash memory such as DRAM (Dynamic Random Access Memory), SRAM (Static Random Access Memory).

[0061] The processor (34) controls the overall operation of the electronic device (30). For example, the processor (34) can control the input / output interface (31), the communication interface (32), and / or the memory (33) overall by executing programs stored in the memory (33). The processor (34) can control the operation of the electronic device (30) by executing programs stored in the memory (33).

[0062] The processor (34) can control at least some of the operations of the electronic device (30) mentioned in the present disclosure. For example, the processor (34) can generate a text embedding based on input text from a vehicle occupant, obtain search results by performing a search based on the text embedding on at least some of a previously generated dialogue example database, select at least one response guideline from a plurality of previously generated response guidelines based on the search results, and generate a prompt template corresponding to the input text based on the at least one response guideline.

[0063] In another example, the processor (34) can determine the domain of the input text based on the input text of the vehicle occupant, and, based on a filtering condition including the domain, perform filtering on a dialogue example database composed of dialogue example data grouped with a plurality of dialogue rules as items, thereby selecting at least one dialogue rule and generating a prompt template corresponding to the input text based on the at least one dialogue rule.

[0064] Meanwhile, specific examples of the operation of the processor (34) can be understood as controlling the operation of the electronic device (30) described later with reference to FIGS. 4 to 21.

[0065] The processor (34) can be implemented using at least one of ASICs (Application Specific Integrated Circuits), DSPs (Digital Signal Processors), DSPDs (Digital Signal Processing Devices), PLDs (Programmable Logic Devices), FPGAs (Field Programmable Gate Arrays), controllers, microcontrollers, microprocessors, and other electrical units for performing functions.

[0066] In one embodiment, the electronic device (30) may be a mobile electronic device. For example, the electronic device (30) may be implemented as a smartphone, tablet PC, PC, smart TV, PDA (personal digital assistant), laptop, media player, camera-equipped device, and other mobile electronic devices. Additionally, the electronic device (30) may be implemented as a wearable device such as a watch, glasses, hair band, and ring equipped with communication functions and data processing functions.

[0067] In another embodiment, the electronic device (30) may be an electronic device embedded in a vehicle. For example, the electronic device (30) may be an electronic device inserted into the vehicle during the vehicle production process or coupled to the vehicle through tuning after the production process.

[0068] In another embodiment, the electronic device (30) may be a server located outside the vehicle. The server may be implemented as at least one computing device that communicates via a network to provide commands, code, files, content, services, etc.

[0069] In one embodiment, the process performed in the electronic device (30) may be performed by at least some of the following: a mobile electronic device, an electronic device embedded in a vehicle, and a server located outside the vehicle.

[0070] FIG. 4 is an example of how an electronic device operates. Specifically, FIG. 4 is an example of how an electronic device (30) operates when a conversation system (1) is applied to a transportation system using a vehicle.

[0071] Referring to FIG. 4, in step 410, the electronic device (30) can obtain input text from a vehicle occupant.

[0072] In one embodiment, the electronic device (30) can acquire input text from a vehicle occupant through an input interface. For example, the electronic device (30) can acquire input text from a vehicle occupant by acquiring user input in the form of a voice signal or text through various forms of input means, such as a microphone, a camera, a keypad, and / or a touchscreen.

[0073] Meanwhile, the electronic device (30) may further acquire vehicle status regarding the vehicle being used by the vehicle occupant. In the present disclosure, vehicle status refers to various data regarding the vehicle, and the electronic device (30) may include driving data (e.g., vehicle speed, fuel status, or set driving route, etc.) acquired through at least one sensor provided inside the vehicle or stored in a memory device provided in the vehicle, and / or surrounding environment data (e.g., images of the vehicle surroundings, etc.) acquired through at least one sensor provided outside the vehicle.

[0074] The electronic device (30) may use the vehicle state in the process of generating a response corresponding to the input text. For example, the electronic device (30) may use the vehicle state to interpret the requirements of the vehicle occupant more accurately or to generate a response customized to the vehicle situation.

[0075] In step 420, the electronic device (30) can construct the input of the language model based on the input text.

[0076] In one embodiment, the language model may include a pre-trained language model to perform Natural Language Processing (NLP) tasks. The language model may be implemented as various language models based on BERT, GPT, Transformer, LSTM, and / or XLNet, but is not limited thereto. The language model according to one embodiment may include a large-scale language model (LLM) trained on a large text dataset.

[0077] The electronic device (30) may use the input text itself as the input to the language model, but may also construct the input to the language model by applying a pre-generated prompt template to the input text to generate an input prompt.

[0078] In one embodiment, the electronic device (30) may use the current situation in which the vehicle is placed as the context of the input text by using the vehicle state in the process of generating an input prompt. For example, the electronic device (30) may use the vehicle state in the process of generating a prompt template, or may use the vehicle state as the target of application of the prompt template together with the input text.

[0079] In step 430, the electronic device (30) can use a language model to generate an output of a language model corresponding to the input.

[0080] In one embodiment, the electronic device (30) can generate an output of a language model corresponding to an input based on the input of the language model configured in step 420, using a language model previously stored in memory (33).

[0081] In another embodiment, the electronic device (30) can obtain the output of a language model corresponding to the input based on the input of the language model configured in step 420 by calling a language model API hosted through an external device (40).

[0082] In step 440, the electronic device (30) can provide a response corresponding to the input text based on the output.

[0083] In one embodiment, the output of the language model may include output text as a response from a conversational artificial intelligence agent to input text. For example, providing a response corresponding to the input text may include displaying a vehicle interface containing the output text on a vehicle display. In another example, providing a response corresponding to the input text may include converting the output text into a speech signal using Text-to-Speech (TTS) or the like, and playing the converted speech signal through a speaker provided in the vehicle.

[0084] Additionally, in one embodiment, the output of the language model may include a command. For example, the command may include a command that causes the electronic device (30) to generate a signal to control a specific function within the vehicle (e.g., air conditioning or navigation). In another example, the command may include a command that causes the electronic device (30) to provide specific information (e.g., restaurant search results) obtained from an external device (40) in an audiovisual manner. In this case, providing a response corresponding to the input text may include executing a command included in the output of the language model.

[0085] Through this, the conversational artificial intelligence agent provided via the conversation means (20) and the user (10) can interact in natural language, and the user (10) can easily receive desired information and conveniently use desired vehicle functions.

[0086] FIG. 5 is an exemplary drawing for explaining a conversation method according to one embodiment. The vehicle display (50) shown in FIG. 5 can be understood as an example of the conversation means (20) shown in FIG. 1.

[0087] A vehicle display (50) according to one embodiment is provided in a location where a user (10) can see it, such as around the driver's seat of a vehicle, and can visually display the interaction between the user (10) and an interactive artificial intelligence agent. For example, the vehicle display (50) may include, but is not limited to, a central information display, a cluster display, and / or a head-up display mounted in the vehicle.

[0088] A vehicle interface (51) may be displayed on a vehicle display (50). In one embodiment, the vehicle interface (51) is an interface for providing various vehicle functions to a user (10), and may include an interface for providing various vehicle functions such as functions related to vehicle control and infotainment functions. The infotainment functions according to one embodiment may include, but are not limited to, conversational artificial intelligence services, navigation, calling, messaging, social media, in-vehicle audio and video functions.

[0089] In addition, in one embodiment, the vehicle interface (51) may include various interfaces used to provide various information, such as fuel level, battery level, and safety of operation.

[0090] In one embodiment, the vehicle display (50) may be implemented as an input / output interface (31) of an electronic device (30), or as a separate device distinct from the electronic device (30) so as to be able to exchange data through the input / output interface (31). The electronic device (30) generates various data necessary for the vehicle interface (51) to be displayed on the vehicle display (50), and can display the vehicle interface (51) on the vehicle display (50) based on the generated data.

[0091] In one embodiment, the vehicle interface (51) may display at least a portion of the conversation status and / or conversation content. For example, the vehicle interface (51) may include a first area (52) that displays the conversation status and / or a second area (53) that displays the conversation content.

[0092] In one embodiment, the first region (52) may include text and / or an icon representing the state of the conversational AI agent. For example, the state of the conversational AI agent may include a waiting state, a state of receiving user input, a state of generating a response, and a state of playing output text voice. Meanwhile, the icon may be used as a graphic element representing the state of the conversational AI agent and as a means to intuitively display the current state of the conversational AI agent through various visual representations such as shape, color, and / or animation effects.

[0093] Additionally, in one embodiment, the second area (53) may include the input text of the user (10) obtained by the electronic device (30) and / or the output text of the conversational artificial intelligence agent included in the response corresponding to the input text. Further information requested by the user (10) may be displayed in the second area (53). For example, search results, etc., in response to the user's (10) request may be displayed in the second area (53).

[0094] In one embodiment, the display state of at least part of the vehicle interface (51), such as the first area (52) and / or the second area (53), can be switched by interaction between the user (10) and the conversation means (20).

[0095] For example, the state of the first area (52) can be switched from a state indicating a waiting state to a state indicating a user input reception state while acquiring the user's (10) speech or text input.

[0096] As another example, the state of the second region (53) may be switched to a state that further displays the input text of the user (10) recognized at the time the input text of the user (10) is acquired, and subsequently, based on the fact that the electronic device (30) has completed the generation of a response, it may be switched to a state that further displays the output text corresponding to the input text and / or the information requested by the user (10).

[0097] In one embodiment, the conversation state and / or conversation content may not be included in the vehicle interface (51), and although FIG. 5 is illustrated as having a first area (52) and a second area (53) separated, the area displaying the conversation state and the area containing the conversation content may be set to be adjacent or integrated.

[0098] FIG. 6 is an exemplary drawing for illustrating an electronic device including a first generating unit, an execution unit, and a second generating unit. Each of the electronic device (610) and external device (620) shown in FIG. 6 may correspond to the electronic device (30) or external device (40) shown in FIG. 2.

[0099] Meanwhile, the operation of at least some of the functional modules constituting the electronic device (610), such as the first generation unit (611), the execution unit (612), and the second generation unit (613) shown in FIG. 6, can be implemented by the processor (34) performing computational processing corresponding to each functional module.

[0100] Referring to FIG. 6, the electronic device (610) may include a first generating unit (611), an execution unit (612), a second generating unit (613), and a memory (614).

[0101] The first generation unit (611) can generate a call sequence based on the input text obtained by the electronic device (610). In one embodiment, the first generation unit (611) can generate a call sequence based on the user's input text using a previously trained first language model.

[0102] In the present disclosure, a call sequence refers to a series of commands consisting of at least one call. In one embodiment, the call sequence may include a plurality of calls. Meanwhile, a call according to one embodiment may include the electronic device (610) obtaining various data from the memory (614) of the electronic device (610), another device in the vehicle system, or an external device (620), or requesting a specific operation from the memory (614), another device in the vehicle, or an external device (620) in order for the electronic device (610) to generate a response corresponding to an input text.

[0103] As one example, the call may include calling an API endpoint of an external device (620) to obtain data from the external device (620). As another example, the call may include a system command call to control at least some of the hardware or software configurations of a vehicle including an electronic device (610). As yet another example, the call may include, but is not limited to, a database query call to execute a predetermined query to look up or modify information from a database accessible to the electronic device (610).

[0104] In one embodiment, the call may include a request for a task to achieve various purposes, such as searching for a specific search term, checking whether a specific restaurant sells a specific menu, calling a specific person, scheduling an alarm, opening and closing the windows of a vehicle, or turning the air conditioner of a vehicle on and off.

[0105] The execution unit (612) can obtain a call execution result by executing a call sequence generated from the first generation unit (611). In one embodiment, the call execution result may include various information (e.g., search results, etc.) obtained to generate a response from an interactive agent corresponding to the input text. In addition, in one embodiment, the call execution result may include that a specific request included in the call has been completed (e.g., that a message transmission has been completed).

[0106] Additionally, in one embodiment, the result of the call execution may include external information obtained from an external device (620). According to one embodiment, the external information obtained from the external device (620) may include text in a structured format (e.g., JSON or XML, etc.), but is not limited thereto.

[0107] For example, if the input text of a vehicle occupant indicates 'tell me the current weather here,' the first generation unit (611) determines the call execution result required to generate a response as the vehicle's location information and the current weather information of that location, and can generate a call sequence to obtain the vehicle's location information and the current weather information of that location.

[0108] As an example, the first generating unit (611) may generate a call sequence including a first call to request the vehicle's location information from a location sensor inside the vehicle system or a first external device providing the vehicle's location information in order to obtain the vehicle's location information. Additionally, the first generating unit (611) may generate a call sequence further including a second call to request weather information corresponding to the vehicle's current location from a second external device providing weather information in order to obtain the current weather information of the location.

[0109] Afterwards, the execution unit (612) can obtain the vehicle's location information and the current weather information of the location as a result of the call execution by executing the call sequence generated from the first generation unit (611).

[0110] The second generation unit (613) can generate a response corresponding to the input text based on at least a portion of the input text and the call execution result. In one embodiment, the second generation unit (613) can generate an output text corresponding to the input text based on at least a portion of the input text and the call execution result using a previously trained second language model. At this time, the second language model may include the same language model as the first language model, but is not limited thereto.

[0111] In one embodiment, the second generation unit (613) may generate a response corresponding to the input text by configuring the input of the second language model and generating the output of the second language model corresponding to the input using the second language model. For example, the specific process of generating the response may include the same process as steps 420 to 440 illustrated in FIG. 4.

[0112] In one embodiment, the second generation unit (613) may use the call execution result in the process of generating a response corresponding to the input text. For example, the second generation unit (613) may use the call execution result to interpret the requirements of the vehicle occupant more accurately or to generate a response customized to the vehicle situation.

[0113] In one embodiment, the second generation unit (613) may use the call execution result as the context of the input text by using the call execution result in the process of generating an input prompt. For example, the second generation unit (613) may use the call execution result in the process of generating a prompt template, or may use the call execution result as the target of application of the prompt template together with the input text.

[0114] Meanwhile, the memory (614) can store a program used for the operation, processing, and control of the first generation unit (611), the execution unit (612), and the second generation unit (613). Additionally, the memory (614) can store various data that the electronic device (610) uses or generates, such as input text, call sequences, and output text.

[0115] In one embodiment, the electronic device (610) may generate a conversation history based on input text. The conversation history according to one embodiment may include at least one input text and a response corresponding to each input text. For example, when the electronic device (610) generates a first response corresponding to a first input text, it may store the acquired first input text and the generated first response in memory (614).

[0116] In one embodiment, the electronic device (610) can update the conversation history by adding input text and / or a response to the conversation history. For example, if a conversation history including a first input text and a first response is stored in memory (614), the electronic device (610) can update the conversation history by adding a second input text to the conversation history based on having obtained a second input text.

[0117] FIG. 7 is an example of a method for configuring the input of a language model according to one embodiment.

[0118] Referring to FIG. 7, in step 710, the electronic device (30) can generate a call sequence based on input text from a vehicle occupant. For example, in response to obtaining input text from a vehicle occupant indicating ‘guide me to a nearby restaurant,’ the electronic device (30) can generate a call sequence including a search with ‘restaurant’ as the search term.

[0119] In step 720, the electronic device (30) can obtain a result of the call execution by executing a call sequence. For example, the electronic device (30) can obtain information such as the place name, address, and electric vehicle charging availability for a plurality of restaurants as a result of the search execution.

[0120] In step 730, the electronic device (30) can construct the input of the language model based on at least some of the results of the call execution and the input text. For example, the electronic device (30) can construct the input of the language model by constructing an input prompt that includes the input text 'Guide me to a nearby restaurant' and information obtained as the result of the search execution.

[0121] Meanwhile, properly configuring the input of a language model is essential for generating an appropriate response, and depending on how the input is configured, a satisfactory response may be generated even when using a language model with few parameters, or an unsatisfactory response may be generated even when using a language model with many parameters.

[0122] For example, if information on hundreds of restaurants is obtained as a result of a search, including the entire result in the input of the language model may generate inaccurate responses or consume unnecessary computational resources.

[0123] Furthermore, if no information about restaurants is obtained as a result of the search, constructing the input to the language model without including contextual prompts indicating such specific situations may lead to hallucinations, such as including distorted content in the response.

[0124] Accordingly, a specific process for generating a prompt template corresponding to the input text to properly configure the input of the language model will be described later with reference to FIGS. 9 and FIGS. 16, etc.

[0125] Figure 8 is an exemplary diagram illustrating a conversation example database.

[0126] The electronic device (30) may use a dialogue example database (810) in the process of generating a response corresponding to an input text. In one embodiment, the electronic device (30) may use the dialogue example database (810) to generate an input prompt of a language model based on the input text. Hereinafter, an example of how the dialogue example database (810) stores data will be described in detail with reference to FIG. 8.

[0127] Referring to FIG. 8, the conversation example database (810) may contain conversation example data. In the present disclosure, conversation example data may represent a set of data records that combine a scenario represented by a specific conversation example and various data corresponding to that scenario.

[0128] In one embodiment, the conversation example database (810) may be composed of conversation example data grouped by a plurality of dialogue rules as items. For example, the first group (821) may be composed of conversation example data grouped under the first dialogue rule. That is, each of the conversation example data constituting the first group (821) may represent conversation example data corresponding to each of a plurality of scenarios that follow a single dialogue rule.

[0129] As an example, the first conversation rule may represent a 'guidance on the processing result of a vehicle control request'. For example, the first conversation example data included in the first group (821) may represent a data record corresponding to an 'air conditioner setting change scenario', and the second conversation example data may represent a data record corresponding to a 'window opening / closing scenario', etc. Additionally, the second conversation rule may represent a 'guidance on the processing result of a location-based search'. For example, the conversation example data included in the second group (822) may represent a data record corresponding to a 'restaurant search scenario', etc.

[0130] As another example, the first and second conversation rules may be configured with greater detail. For instance, the first conversation rule may represent a "conversation in a situation where no location search results exist for the request," and the second conversation rule may represent a "conversation in a situation where there are an excessive number of location search results for the request."

[0131] In one embodiment, each of the plurality of conversation rules may correspond one-to-one with each of the plurality of response guidelines. In this disclosure, response guidelines refer to various data that guide a language model to generate systematic responses.

[0132] A response guideline according to one embodiment may be composed of various contents for generating a response appropriate to the situation, such as guiding the context of the conversation, specifying the format of the output, or guiding necessary additional queries. In one embodiment, the response guideline may include a description that guides the context of the conversation.

[0133] Additionally, in one embodiment, the response instructions may include a command regarding how to utilize the output of the language model. For example, the response instructions may specify the format of the output, such as whether to provide the output text of the language model to the user only as synthesized speech or text displayed on a vehicle interface, or whether to display search results resulting from an API call in the form of a list on the vehicle interface.

[0134] For example, the first conversation rule may correspond to the first response guideline (831), and the second conversation rule may correspond to the second response guideline (832). As an example, the first conversation rule may represent a 'conversation in a situation where there are no location search results for the request,' and the second conversation rule may represent a 'conversation in a situation where there are too many location search results for the request.'

[0135] At this time, the first response guideline (831) may include a description including ‘no place search results exist’ and ‘if necessary, request a change of search terms and inform that a waypoint cannot be set,’ and the second response guideline (832) may include a description including ‘multiple place search results exist’ and ‘if necessary, request a place selection so that a waypoint can be specified.’

[0136] Meanwhile, conversation example data may represent conversation example data corresponding to each of a plurality of conversation examples. In the present disclosure, a conversation example is an exemplary conversation pre-generated to reflect a specific scenario.

[0137] In one embodiment, the dialogue example may consist of a single example of input text, or may consist of an example of input text and an example of corresponding output text. Meanwhile, the dialogue example may further include additional examples of input text and additional examples of corresponding output text.

[0138] Below, the types of data items that may be included in the conversation example data (840) corresponding to one conversation example (843) are described in detail.

[0139] In one embodiment, conversation example data (840) corresponding to a conversation example (843) may include a conversation example embedding (841) and a response guideline index (842). In this case, the conversation example embedding (841) may be generated based on the conversation example (843).

[0140] For example, a conversation example (843) can be converted into a conversation example embedding (841) through sentence embedding. In this case, sentence embedding means expressing the meaning of a sentence in a pre-set format, such as the format of a numerical embedding vector.

[0141] Meanwhile, the response guideline index (842) may indicate an index that indicates a response guideline corresponding to a higher conversation rule of the conversation example data (840) corresponding to the conversation example (843). For example, if the conversation example data (840) is grouped under the first conversation rule, the response guideline index (842) included in the conversation example data (840) may indicate the first response guideline (831) corresponding to the first conversation rule.

[0142] In one embodiment, the conversation example data (840) may include only a conversation example index indicating the conversation example data (840) instead of including a response guideline index (842). In this case, the conversation example index may be understood as an identifier of the conversation example data (840) and may be mapped to a response guideline index (842) indicating a response guideline corresponding to a corresponding conversation rule through a higher conversation rule of the conversation example data (840). That is, the conversation example data (840) may include a response guideline index (842) or include certain data mapped to the response guideline index (842).

[0143] In one embodiment, the conversation example data (840) may further include a conversation example (843) and a response guideline (844) corresponding to the conversation example (843). In this case, the conversation example (843) represents a conversation example used to generate the conversation example embedding (841), and the response guideline (844) represents a response guideline indicated by a response guideline index (842).

[0144] Meanwhile, the conversation example (843) and response guideline (844) may require relatively large memory resources compared to the conversation example embedding (841) and response guideline index (842). In one embodiment, the conversation example (843) and response guideline (844) may be stored separately from the conversation example data (840), and the conversation example data (840) may include only the conversation example index mapped to the conversation example (843) and response guideline (844).

[0145] FIG. 9 is an example of a method for generating a prompt template corresponding to input text according to one embodiment.

[0146] Referring to FIG. 9, in step 910, the electronic device (30) can determine the domain of the input text based on the input text of the vehicle occupant.

[0147] In step 920, the electronic device (30) can select at least one dialogue rule by performing filtering on a dialogue example database composed of dialogue example data grouped with a plurality of dialogue rules as items, based on a filtering condition including a domain.

[0148] In one embodiment, the filtering condition may further include at least one of the call execution result and the vehicle status, and the call execution result may be generated based on the input text.

[0149] In one embodiment, the electronic device (30) can select at least one conversation rule by performing primary filtering on a plurality of conversation rules based on a domain and performing secondary filtering on the result of the primary filtering based on at least one of the call execution result and the vehicle state.

[0150] In step 930, the electronic device (30) can generate a prompt template corresponding to the input text based on at least one conversation rule.

[0151] In one embodiment, each of the plurality of conversation rules may correspond one-to-one with each of the plurality of response guidelines. In this case, the electronic device (30) may generate a prompt template based on at least one response guideline among the response guidelines corresponding to each of at least one conversation rule.

[0152] In one embodiment, the electronic device (30) can generate an input prompt for a pre-trained language model by applying a prompt template to at least one of a vehicle state, a call execution result, a pre-generated conversation history, and an input text. Subsequently, the electronic device (30) can generate an output text based on the input prompt using the language model.

[0153] In one embodiment, the input prompt may include a contextual prompt generated based on at least one of the vehicle status, call execution result, conversation history, and input text, and a pre-set fixed prompt.

[0154] In one embodiment, the prompt template may apply to at least a portion of the total call execution results regarding the call execution results.

[0155] For example, the result of a call execution may include multiple metadata items, and the prompt template may apply to at least some of the multiple metadata items.

[0156] In one embodiment, the conversation history may consist of a number of unit input / output texts that does not exceed a preset number among the entire conversation history.

[0157] In one embodiment, the first memory storing the entire conversation history may be distinguished from the second memory storing the conversation history composed of a number of unit input / output texts that does not exceed a preset number.

[0158] Figure 10 is an exemplary diagram illustrating the process of filtering multiple conversation rules.

[0159] Referring to FIG. 10, the electronic device (30) can derive a filtered conversation rule (1030) by performing filtering on a plurality of conversation rules (1010). In one embodiment, the electronic device (30) can select a filtered conversation rule (1030) by performing filtering on a conversation example database composed of conversation example data grouped with a plurality of conversation rules (1010) as items, based on a filtering condition (1020). At this time, the filtered conversation rule may be composed of at least one conversation rule.

[0160] In one embodiment, the filtering condition (1020) may include a domain (1021). For example, the electronic device (30) may determine a domain (1021) corresponding to an input text and select a filtered conversation rule (1030) by filtering a plurality of conversation rules (1010) based on the filtering condition (1020) that includes the domain (1021).

[0161] In the present disclosure, a domain (1021) refers to a problem area to be solved through natural language processing. In one embodiment, the domain may include at least one of various areas where natural language processing can be usefully utilized, such as navigation, map services, reservation services, medical record analysis, medical conversation analysis, legal document analysis, stock market analysis and / or real estate information analysis, but is not limited thereto.

[0162] In one embodiment, the electronic device (30) can determine the domain (1021) of the input text based on the input text of the vehicle occupant.

[0163] For example, the electronic device (30) can determine the domain (1021) of the input text based on the output of the first language model used by the first generation unit (611) described above with reference to FIG. 6. As an example, the first language model may include a language model that returns the domain (1021) of the input text along with a call sequence based on a predetermined classification method, such as zero-shot or few-shot classification or embedding-based classification, and the electronic device (30) can determine the domain (1021) of the input text with the output of the first language model.

[0164] As another example, the electronic device (30) can determine a domain (1021) based on a call sequence included in the output of the first language model. For example, if the output of the first language model includes an API call representing a place search, the electronic device (30) can determine the domain (1021) of the input text corresponding to the API call as navigation and / or map search, etc. Meanwhile, a domain (1021) corresponding to each individual call including a call sequence may be pre-set, and the electronic device (30) can determine the domain (1021) of the input text based on the call sequence using a pre-set correspondence.

[0165] If the filtering condition (1020) includes a domain (1021), the electronic device (30) can filter a plurality of conversation rules (1010) based on the domain (1021). For example, if the domain (1021) of the input text is navigation, the filtered conversation rules (1030) may not include conversation rules among the plurality of conversation rules (1010) that are unrelated to tasks related to navigation.

[0166] In one embodiment, each of the plurality of conversation rules (1010) may be pre-mapped to associated domains, and the electronic device (30) may select a filtered conversation rule (1030) by removing the remaining conversation rules, excluding the conversation rule mapped to the domain (1021) corresponding to the input text among the plurality of conversation rules (1010).

[0167] Meanwhile, in one embodiment, the filtering condition may further include at least one of a call execution result (1022) and a vehicle state (1023). In this case, the call execution result (1022) may be generated based on input text. For example, the electronic device (30) may obtain the call execution result (1022) as the result of executing the call sequence described above with reference to FIG. 6, etc.

[0168] Additionally, the vehicle status (1023) refers to various data regarding the vehicle being used by the vehicle occupant as described above with reference to FIG. 3, etc. The electronic device (30) can obtain the vehicle status (1023) by obtaining driving data using at least one sensor provided inside the vehicle or a memory device of the vehicle, or by obtaining surrounding environment data through at least one sensor provided outside the vehicle.

[0169] In one embodiment, the filtering condition (1020) includes a call execution result (1022) and / or a vehicle state (1023), which may indicate that a predetermined condition is set based on the call execution result (1022) and / or the vehicle state (1023), and that at least some of the plurality of conversation rules (1010) may be filtered depending on whether they satisfy the predetermined condition.

[0170] For example, the input text may represent 'Set A as a waypoint'. In this case, the call execution result (1022) may include a place search result with A as the search term, and the vehicle status (1023) may include data regarding the currently set destination, etc. The electronic device (30) may filter multiple conversation rules (1010) according to whether pre-set conditions regarding the call execution result (1022) and / or vehicle status (1023) are satisfied, such as a condition identifying whether there is no waypoint search result, a condition identifying whether a waypoint cannot be set because a destination is not set, a condition identifying whether there are multiple waypoint search results, and a condition identifying whether the searched waypoint is the same as the currently set waypoint.

[0171] For example, when considering the call execution result (1022) and / or vehicle status (1023), if the searched waypoint is identified as being the same as the currently set waypoint, the electronic device (30) may select only at least one conversation rule that is pre-mapped to be relevant to the situation (e.g., a conversation rule indicating a scenario asking whether to enter another waypoint or a scenario indicating that it is already set) as the filtered conversation rule (1030).

[0172] In one embodiment, the electronic device (30) may perform primary filtering based on at least some of the filtering conditions (1020) and perform secondary filtering based on the remainder of the filtering conditions (1020). For example, the electronic device (30) may select a filtered conversation rule (1030) by performing primary filtering on a plurality of conversation rules (1010) based on a domain (1021) and performing secondary filtering on the result of the primary filtering based on at least one of a call execution result (1022) and a vehicle state (1023).

[0173] For example, if the input text represents 'Set A as a waypoint', the domain (1021) corresponding to the input text can be determined as navigation. At this time, the result of the first filtering based on the domain (1021) may be a part of the multiple conversation rules (1010) mapped to navigation. Subsequently, the electronic device (30) can perform a second filtering on a part of the multiple conversation rules (1010) mapped to navigation.

[0174] For example, when considering the call execution result (1022) and / or vehicle status (1023), if the searched waypoint is identified as being the same as the currently set waypoint, the electronic device (30) may select only at least one conversation rule that is pre-mapped to be related to the situation identified among the results of the first filtering as the filtered conversation rule (1030).

[0175] Meanwhile, filtering based on a domain (1021) may require relatively fewer computational resources compared to filtering using pre-set conditions regarding the call execution result (1022) and vehicle status (1023), and the conditions regarding the call execution result (1022) and vehicle status (1023) for filtering may be set differently for each domain (1021). Accordingly, by sequentially performing first filtering based on the domain (1021) and second filtering based on the call execution result (1022) and / or vehicle status (1023), filtered conversation rules (1030) can be derived quickly and efficiently using relatively fewer computational resources.

[0176] FIGS. 11 and 12 are exemplary drawings for explaining the process of generating a prompt template corresponding to input text based on at least one conversation rule.

[0177] The electronic device (30) can generate a prompt template corresponding to the input text based on at least one conversation rule. In this case, the at least one conversation rule may represent at least some of the filtered conversation rules (1030) shown in FIG. 10.

[0178] In the present disclosure, a prompt template is a template used to generate an input prompt for a language model and may be implemented as a pre-designed document or data structure in a structured format to configure the input prompt for the language model. In one embodiment, the prompt template may be defined in JSON, YAML, or other structured data formats, but is not limited thereto.

[0179] As described above with reference to FIG. 8, each of the multiple conversation rules can correspond one-to-one with each of the multiple response guidelines. With reference to FIG. 11, the first conversation rule (1111) corresponds to the first response guideline (1121), the second conversation rule (1112) corresponds to the second response guideline (1122), the third conversation rule (1113) corresponds to the third response guideline (1123), and the fourth conversation rule (1114) corresponds to the fourth response guideline (1124).

[0180] For example, the first conversation rule (1111) and the second conversation rule (1112) may represent the filtered conversation rule (1030) described above with reference to FIG. 10, and the third conversation rule (1113) and the fourth conversation rule (1114) may represent the remaining conversation rules among the plurality of conversation rules (1010), excluding the filtered conversation rule (1030).

[0181] At this time, the first response guideline (1121) and the second response guideline (1122) may be classified as a candidate group (1130) that can be used when the electronic device (30) generates a prompt template, and the third response guideline (1123) and the fourth response guideline (1124) may be classified as an unused group (1140) that is not used when the electronic device (30) generates a prompt template.

[0182] In one embodiment, the electronic device (30) may generate a prompt template based on at least one of the response guidelines included in the candidate group (1130). For example, as illustrated in FIG. 11, if the candidate group (1130) includes a first response guideline (1121) and a second response guideline (1122), the electronic device (30) may generate a prompt template based on the first response guideline (1121), generate a prompt template based on the second response guideline (1122), or generate a prompt template based on both the first response guideline (1121) and the second response guideline (1122).

[0183] Referring to FIG. 12, a first template (1210) and a second template (1220) can be seen, in which at least a portion of a prompt template according to one embodiment is shown.

[0184] The first template (1210) is an example of a prompt template generated based on response instructions corresponding to a conversation rule indicating 'a conversation in a situation where no place search result exists according to the request'.

[0185] For example, the first template (1210) may include a template element (1211) that guides the generation of a situation-appropriate response by including a description included in the response instructions, a template element (1212) that guides the conversation context through the conversation history, and a template element (1213) that provides input text from the vehicle occupant to be processed. At this time, the response instructions may include information indicating how to arrange the situation-specific description and specific template elements, and the electronic device (30) may generate a prompt template that includes the situation-specific description and arranges the template elements in a predetermined manner by using the response instructions that are ultimately used.

[0186] Meanwhile, the second template (1220) is an example of a prompt template generated based on response guidelines corresponding to a conversation rule indicating 'a conversation in a situation where there are too many location search results in response to a request'.

[0187] For example, the second template (1220) may include a template element (1221) that guides the generation of a situation-appropriate response by including a description included in the response instructions, a template element (1222) that indicates which data to use from the call execution results, a template element (1223) that guides the conversation context through the conversation history, and a template element (1224) that provides the input text of the vehicle occupant to be processed.

[0188] As such, when using prompt templates based on different response guidelines, the types and arrangement of contextual descriptions and / or template elements constituting the prompt templates may differ.

[0189] In one embodiment, the electronic device (30) can include only data that is relatively important for generating a response in the input prompt by limiting the application targets of the prompt template using response guidelines.

[0190] For example, a predetermined response guideline may include a setting that does not include the vehicle state as a target for application to the prompt template, such as in the first template (1210) and the second template (1220), or does not include the call execution result as a target for application to the prompt template, such as in the first template (1210).

[0191] In one embodiment, a predetermined response guideline may include a setting such that a prompt template includes at least a portion of the total call execution results as the target of application regarding the call execution results. For example, a prompt template generated based on a predetermined response guideline may have a portion of the total call execution results as the target of application regarding the call execution results.

[0192] For example, in the process of generating the second template (1220), the data obtained by the electronic device (30) as a result of the call execution may include data regarding dozens of places. At this time, a predetermined response guideline may include a setting in such a scenario where the prompt template, as the target of application, includes only a preset number of place search results (e.g., three, etc.) based on distance, rather than the entire search result.

[0193] In addition, in one embodiment, the result of the call execution may include a plurality of metadata items, and the prompt template may have the application target regarding the plurality of metadata items as at least some of the plurality of metadata items.

[0194] For example, in the process of generating the second template (1220), the data obtained by the electronic device (30) regarding each location as a result of the call execution may consist of multiple metadata items such as the location name, address, availability of parking, and availability of electric vehicle charging. At this time, a predetermined response guideline may include a setting such that, in such a scenario, the prompt template includes only items of a pre-set type (e.g., location name and address, etc.) as the target of application, rather than all of the multiple metadata items.

[0195] By doing so, the electronic device (30) can either not include relatively unimportant data in the input prompt or include relatively important types of data in the input prompt depending on the situation, thereby preventing performance degradation of the language model.

[0196] In addition, since the type of data that needs to be considered primarily in a specific environment, such as a transportation environment using a vehicle, may differ from the type of data that needs to be considered primarily in another environment, the electronic device (30) can generate an input prompt suitable for the individual environment by generating a prompt template that applies appropriate data items based on pre-generated response guidelines.

[0197] Meanwhile, after the electronic device (30) obtains the call execution result, it may store at least a portion of the call execution result in memory (33). Through this, the electronic device (30) may use at least a portion of the call execution result as an application target to generate a prompt template or to generate an input prompt by applying the prompt template.

[0198] For example, the second template (1220) may not include the possibility of parking by location as a target for application. In this case, after the first response corresponding to the first input text, the vehicle occupant may provide a second input text indicating "tell me where I can park among the presented locations," and the electronic device (30) may require information regarding the possibility of parking by location during the process of generating the second response corresponding to the second input text.

[0199] Accordingly, the electronic device (30) stores at least a portion of the call execution results in memory (33), and can then use the stored call execution results in the process of generating an input prompt.

[0200] In one embodiment, the electronic device (30) can initialize a stored call execution result based on the fact that the domain determined in step 910 shown in FIG. 9 is different from the previous domain corresponding to the previous input text.

[0201] For example, the electronic device (30) may retain a previously obtained call execution result so that it can be used later based on the domain of the input text being maintained. In another example, the electronic device (30) may retain a previously obtained call execution result so that it can be used later based on the domain of the input text being maintained and the command corresponding to the user's input text not yet being executed.

[0202] Additionally, the electronic device (30) can secure memory resources by deleting previously obtained call execution results based on the fact that the domain of the input text has changed (e.g., from navigation to vehicle control, etc.). As another example, the electronic device (30) can delete previously obtained call execution results based on the fact that the previous domain is not identified in the conversation for a preset number of conversations (e.g., three pairs of input / output processes, etc.) after the domain of the input text has changed.

[0203] Additionally, in one embodiment, the electronic device (30) may retain a previously stored call execution result so that the previously obtained call execution result can be used later, based on the fact that the input text does not correspond to any domain or does not represent a purposeful conversation (TOD).

[0204] Meanwhile, although not shown in FIG. 12, the prompt template may include the vehicle state as a target for application. The electronic device (30) may store the vehicle state, which is updated in real time, in memory (33), and may use the stored vehicle state to generate a prompt template or to generate an input prompt by applying the prompt template.

[0205] FIG. 13 is an exemplary diagram illustrating an input prompt including a situational prompt and a fixed prompt.

[0206] Referring to FIG. 13, the input prompt may include a fixed prompt (1310) and a contextual prompt (1320). In this case, the fixed prompt (1310) may include a preset prompt content regardless of the response instructions, and the contextual prompt (1320) may include a prompt content generated by applying a prompt template generated based on the response instructions to at least one of the vehicle status, call execution result, conversation history, and input text.

[0207] For example, the fixed prompt (1310) may include a description indicating the persona, tone and manner, and language settings of the language model. Additionally, for example, the fixed prompt (1310) may include a description indicating a defense element to prevent harmful manipulations such as tampering and prompt injection attacks.

[0208] Meanwhile, the situational prompt (1320) can be generated by applying a generated prompt template, created according to the process of generating a prompt template suitable for the situation described above with reference to FIG. 12, etc., to at least one of the vehicle status, call execution result, conversation history, and input text. The specific process of applying the generated prompt template will be described later with reference to FIG. 14, etc.

[0209] By generating an input prompt that includes a fixed prompt (1310) and a situational prompt (1320), the electronic device (30) can generate an input prompt that basically includes both the considerations that a language model must consider during the output generation process and considerations specific to the current situation, so that it can provide a response that matches the vehicle occupant's intention with high accuracy.

[0210] FIG. 14 is an exemplary diagram illustrating the process of generating an input prompt by applying a prompt template and generating output text based on the input prompt.

[0211] Referring to FIG. 14, the electronic device (30) can generate an input prompt (1430) in a format that can be processed by a natural language processing language model (1440) by combining each element constituting the input prompt (1430) of the language model (1440) using a prompt template (1420). That is, the input prompt (1430) according to one embodiment can represent an input to the language model (1440) generated by the electronic device (30).

[0212] For example, the electronic device (30) may include slots for each element that constitute the input prompt (1430). The electronic device (30) may generate the input prompt (1430) by inserting each element into each slot of the prompt template (1420).

[0213] In one embodiment, the prompt template (1420) may include a slot corresponding to the input text (1411), and the electronic device (30) may generate an input prompt (1430) by inserting at least a portion of the input text (1411) into the slot.

[0214] Additionally, in one embodiment, the prompt template (1420) may include slots corresponding to each of the call execution result (1412), vehicle status (1413), and conversation history (1414), and the electronic device (30) may generate an input prompt (1430) by inserting elements corresponding to the slots.

[0215] Referring to the second template (1220) illustrated in FIG. 12, a prompt template (1420) according to one embodiment may include a slot corresponding to a 'place name' and a slot corresponding to an 'address' among the call execution results (1412), and may include a slot corresponding to a conversation history (1414) and a slot corresponding to an input text (1411). The electronic device (30) may apply the prompt template (1420) by inputting data of a type for which a corresponding slot exists into each slot, and data of a type for which a corresponding slot does not exist may be understood as not being subject to the application of the prompt template (1420).

[0216] Subsequently, the electronic device (30) can obtain the output of the language model (1440) based on the input prompt (1430) using the previously learned language model (1440). In one embodiment, the output of the language model (1440) may include an output text (1450) as a response from a conversational artificial intelligence agent to the input text (1411). Additionally, in one embodiment, the output of the language model may include a command to generate a signal for controlling a specific function within the vehicle.

[0217] Meanwhile, the specific process for obtaining the output of the language model (1440) may include the same process as step 430 described above with reference to FIG. 4.

[0218] By comprehensively considering the user's situation and configuring an input prompt (1430) suitable for situational response, a response matching the user's intention can be provided through a single input / output process of the language model (1440), thereby reducing the cost associated with using the language model (1440) and shortening the execution time of the conversational artificial intelligence agent required for the response.

[0219] FIG. 15 is an exemplary drawing for illustrating a conversation history according to one embodiment.

[0220] Referring to FIG. 15, the accumulated input / output text can constitute the entire conversation history (1500). At this time, the input / output text may represent an input text obtained once from a vehicle occupant for the purpose of generating a response, a response corresponding to the input text, or an output text corresponding to the input text. That is, the conversation history may include at least one of the input text of the vehicle occupant and the output text of the conversational artificial intelligence agent, and may also include a response of the conversational artificial intelligence agent that includes additional commands for operating functions within the vehicle in addition to the output text.

[0221] Meanwhile, a method for managing conversation history according to one embodiment may include storing and / or updating the conversation history described above with reference to FIG. 6, etc.

[0222] Additionally, as described above with reference to FIG. 12 and FIG. 14, etc., the electronic device (30) may use a conversation history in the process of generating an input prompt. In one embodiment, the conversation history used by the electronic device (30) in the process of generating an input prompt may consist of a number of unit input / output texts that does not exceed a preset number among the total conversation history (1500) in which all input / output texts are accumulated. For example, the preset number may be set to 3 to 10, but is not limited thereto.

[0223] For example, the number of pre-set items may be set to 6. At this time, the conversation history used in the process of generating a response corresponding to the fourth input text (1541) may consist of the first input text (1511), the first output text (1512), the second input text (1521), the second output text (1522), the third input text (1531), and the third output text (1532).

[0224] Meanwhile, the electronic device (30) may store the conversation history used in the process of generating a response in memory (33). In one embodiment, the electronic device (30) may store the output text included in the conversation history together with the call execution result corresponding to the output text. For example, the call execution result obtained to generate the first output text (1512) may be stored together with the first output text (1512).

[0225] In one embodiment, the first memory storing the entire conversation history (1500) may be distinguished from the second memory storing the conversation history composed of a number of unit input / output texts that does not exceed a preset number.

[0226] For example, the first memory may include memory constituting a server device that provides conversational artificial intelligence services, and the second memory may include memory provided in the vehicle. As another example, the first memory and the second memory may be implemented as distinct memories constituting the server device or as distinct memories provided in the vehicle. Through this, the entire conversation history (1500) can be safely stored, and some conversation history necessary for understanding the context can be stored in a local environment easily accessible to the electronic device (30), thereby allowing for efficient distribution of memory resources.

[0227] In one embodiment, the electronic device (30) may initialize the conversation history based on the fact that the domain corresponding to the current input text is different from the previous domain corresponding to the previous input text. At this time, initializing the conversation history does not mean deleting the entire conversation history (1500), but means not using at least some of the conversation history in the process of generating a prompt template or input prompt.

[0228] In one embodiment, the electronic device (30) may initialize the conversation history stored in the second memory based on the domain being changed. As an example, the electronic device (30) may maintain the conversation history stored in the second memory based on the domain of the input text being maintained, or based on the domain of the input text being maintained and the command corresponding to the user's input text not yet being executed. Additionally, the electronic device (30) may secure memory resources by deleting the conversation history stored in the second memory based on the domain of the input text being changed (e.g., changing from navigation to vehicle control, etc.).

[0229] Meanwhile, when maintaining a conversation history, the electronic device (30) may maintain call execution results corresponding to each input / output text constituting the maintained conversation history. Additionally, when initializing the conversation history, the electronic device (30) may initialize call execution results corresponding to each initialized input / output text.

[0230] Through this, the electronic device (30) uses data elements that are precisely adjusted according to the situation during the process of generating a situational prompt template, and since the target of application of the prompt template can be precisely adjusted according to the situation, it can provide a response that matches the user's intention with only one language model call, reduce the cost associated with the language model call, and shorten the execution time of the conversational artificial intelligence agent required for the response.

[0231] Additionally, in one embodiment, the electronic device (30) may use dialogue parameters in the process of generating an input prompt for a language model. In this case, the dialogue parameters may include vehicle status and call execution results. In one embodiment, the vehicle status may represent the real-time status of the vehicle corresponding to the time of processing the input text, and the call execution result may include at least some of the call execution results obtained for the entire dialogue history (1500).

[0232] The result of a call execution used to generate an input prompt for a language model according to one embodiment may be stored in the second memory described above together with at least a portion of the entire conversation history (1500), but is not limited thereto.

[0233] In one embodiment, the electronic device (30) may determine whether to retain a call execution result along with determining whether to retain a conversation history as previously mentioned, or may specify a call execution result to be applied using a slot of a prompt template, but may also change a call execution result available for generating an input prompt by initializing a specific call execution result.

[0234] In one embodiment, the electronic device (30) may initialize the call execution result based on whether a final response corresponding to the user's previous input text has been provided. In this case, initializing the call execution result means not using at least some of the call execution results among the entire call execution results during the process of generating a prompt template or an input prompt. In this case, the previous input text is not limited to only the input text immediately preceding the current input text.

[0235] For example, if the first input text (1511) is a request from a user requesting weather information, and the first output text (1512) and the second output text (1522) failed to complete the weather information for various reasons, the electronic device (30) can retain the results of the call execution, such as weather information obtained during the process of generating the first output text (1512) and the second output text (1522).

[0236] Additionally, for example, if the first input text (1511) is a request from a user asking for weather information, the third output text (1513) is a text providing weather information requested in the first input text (1511), or if weather information is displayed on a vehicle display along with the provision of the third output text (1513), the electronic device (30) can initialize the call execution result, such as weather information obtained during the process of generating the first output text (1512) and the second output text (1522), based on the provision of a final response corresponding to the user's first input text (1511).

[0237] Meanwhile, the result of the call execution may be initialized based on factors such as the input / output text accumulating beyond a preset number or the domain of the input text changing (e.g., changing from weather guidance to navigation), even though a final response corresponding to the user's previous input text has not been provided.

[0238] Additionally, in one embodiment, the electronic device (30) can initialize the call execution result based on the fact that the current domain corresponding to the current input text is different from the previous domain corresponding to the previous input text.

[0239] For example, the electronic device (30) may initialize the call execution result stored in the second memory based on the domain being changed. As an example, the electronic device (30) may maintain the call execution result stored in the second memory based on the domain of the input text being maintained, or based on the domain of the input text being maintained and the command corresponding to the user's input text not yet being executed. Additionally, the electronic device (30) may secure memory resources by deleting the call execution result stored in the second memory based on the domain of the input text being changed (e.g., changing from weather guidance to navigation, etc.).

[0240] Through this, the computational resources, memory resources, and costs required to generate input prompts can be optimized, and the accuracy of the response can be improved.

[0241] FIG. 16 is an example of a method for generating a prompt template corresponding to input text according to one embodiment.

[0242] Referring to FIG. 16, in step 1610, the electronic device (30) can generate a text embedding based on the input text of the vehicle occupant.

[0243] In one embodiment, the electronic device (30) can generate a text embedding based on at least one of a previously generated conversation history and previously generated conversation parameters and an input text.

[0244] In step 1620, the electronic device (30) can obtain search results by performing a search based on text embeddings on at least a portion of the previously generated dialogue example database.

[0245] A conversation example database may be composed of conversation example data corresponding to each of a plurality of conversation examples. In this case, each of the conversation example data corresponding to each of the plurality of conversation examples may include a conversation example embedding generated based on the conversation example, and a response guideline index representing a response guideline corresponding to the conversation example.

[0246] In one embodiment, the conversation example database may be composed of conversation example data grouped with a plurality of dialogue rules as items. In this case, each of the plurality of response guidelines may correspond one-to-one with each of the plurality of dialogue rules.

[0247] In one embodiment, a search based on text embeddings may include a search based on embedding similarity based on text embeddings and conversation example embeddings included in each of the conversation example data corresponding to each of the plurality of conversation examples.

[0248] In step 1630, the electronic device (30) can select at least one response guideline from a plurality of previously generated response guidelines based on the search results.

[0249] In one embodiment, the electronic device (30) can obtain a search result including a distribution of response guideline indices for conversation example data corresponding to each of a preset number of conversation examples by performing an embedding similarity-based search. In one embodiment, the selected at least one response guideline may be composed of at least one of the response guidelines corresponding to each of a preset number of conversation examples.

[0250] In one embodiment, each conversation example data corresponding to each of a plurality of conversation examples may further include a language type corresponding to the conversation example. In one embodiment, the electronic device (30) can obtain a search result including a distribution of language types for conversation example data corresponding to each of a preset number of conversation examples by performing an embedding similarity-based search.

[0251] In step 1640, the electronic device (30) can generate a prompt template corresponding to the input text based on at least one response instruction.

[0252] In one embodiment, the electronic device (30) can generate a prompt template to have a single task format based on the selection of a single response guideline as at least one response guideline, and can generate a prompt template to have multiple task formats based on the selection of multiple response guidelines as at least one response guideline.

[0253] In one embodiment, the electronic device (30) can generate an input prompt for a pre-trained language model by applying a prompt template to at least one of a vehicle state, a call execution result, a pre-generated conversation history, and an input text. Subsequently, the electronic device (30) can generate an output text based on the input prompt using the language model. At this time, the call execution result may be generated based on the input text.

[0254] In one embodiment, the electronic device (30) can select a language model used to generate output text from among a plurality of previously trained language models based on search results.

[0255] Figure 17 is an exemplary diagram illustrating the process of performing a search based on text embeddings.

[0256] Referring to FIG. 17, the electronic device (30) can generate a text embedding (1730) based on the input text (1710) of a vehicle occupant. In one embodiment, the electronic device (30) can generate a text embedding (1730) corresponding to the input text (1710) using an embedding generation model (1720).

[0257] In the present disclosure, the embedding generation model (1720) is a learning model that converts text expressed in natural language into a pre-set format such as a vector, and means a learning model that numerically expresses the contextual features of the text. For example, the embedding generation model (1720) can be used to generate text embeddings (1730) by dividing text, such as input text (1710), into tokens and reflecting the context and / or meaning of each token based on an artificial neural network, etc., and mapping the input text (1710) into a multidimensional space.

[0258] In one embodiment, the embedding generation model (1720) may include a learning model based on BERT, GPT, or XLNet, etc. For example, the embedding generation model (1720) may be implemented integrally with the first language model used by the first generation unit (611) described above with reference to FIG. 6 in the process of generating a call sequence, and the electronic device (30) may obtain a text embedding (1730) corresponding to the input text (1710) as the output of the first language model in the process of generating a call sequence.

[0259] In one embodiment, the electronic device (30) may generate a text embedding (1730) based on at least one of a previously generated conversation history and previously generated conversation parameters in addition to the input text (1710). In this case, the conversation parameters refer to various information that can provide conversation context to the input text (1710) in addition to the conversation history.

[0260] For example, the conversation history may be configured to be identical to the conversation history used in the process of generating the prompt template described above with reference to FIG. 15, etc. Additionally, for example, the conversation parameters may include at least some of the call execution results and / or vehicle states available in the process of generating the prompt template described above with reference to FIG. 12, etc.

[0261] Through this, the electronic device (30) can reflect information about the conversational context and situation, which is difficult to grasp from the input text (1710) alone, into the text embedding (1730), thereby generating a more accurate text embedding (1730).

[0262] Meanwhile, as described above with reference to FIG. 8, the conversation example database may be composed of conversation example data corresponding to each of a plurality of conversation examples. The conversation example database illustrated in FIG. 17 is an example of a conversation example database composed of conversation example data grouped with a plurality of conversation rules as items. For example, the first group (1741) may be composed of conversation example data grouped under the first conversation rule, and the second group (1742) may be composed of conversation example data grouped under the second conversation rule.

[0263] Additionally, as described above with reference to FIG. 8, conversation example data (1750) corresponding to one conversation example may include a conversation example embedding (1751) that is generated in advance based on the conversation example. At this time, the process of generating the conversation example embedding (1751) based on the conversation example may be the same as the process of generating the text embedding (1730) based on the input text (1710).

[0264] For example, a conversation example embedding (1751) can be generated by using an embedding generation model (1720) to generate a conversation example embedding (1751) corresponding to a conversation example, but the specific method of generating the conversation example embedding (1751) is not limited to this.

[0265] Meanwhile, the electronic device (30) can obtain search results by performing a search based on text embeddings (1730) on at least a portion of a previously generated conversation example database. At this time, at least a portion of the conversation example database to be searched may include all conversation example data under the filtered conversation rule (1030) described above with reference to FIG. 10, etc., that is, all conversation example data included in the candidate group (1130) described above with reference to FIG. 11, etc.

[0266] Through this, the scope of the search based on the text embedding (1730) can be reduced, so that the search based on the text embedding (1730) can be performed with relatively few computational resources.

[0267] In one embodiment, a search based on text embeddings (1730) may include a search based on embedding similarity based on conversation example embeddings (1751) and text embeddings (1730) included in each of the search targets. In this case, the search targets refer to conversation example data corresponding to each of the plurality of conversation examples, and the similarity between the conversation example embeddings (1751) and text embeddings (1730) may be calculated using cosine similarity or Euclidean distance, but is not limited thereto.

[0268] Hereinafter, with reference to FIG. 18 and the like, the process of obtaining search results of a search based on text embeddings (1730) and the process of using the results will be explained in detail.

[0269] FIG. 18 is an exemplary diagram illustrating the process of obtaining search results.

[0270] Referring to FIG. 18, the electronic device (30) can determine a similar data set (1810) composed of conversation example data corresponding to each conversation example having a high similarity to the text embedding (1730) by performing a search based on the text embedding (1730).

[0271] In one embodiment, the electronic device (30) can determine a similar data set (1810) by performing an embedding similarity search for a search target and determining a preset number of similar data selected based on similarity with the input text (1710). In another embodiment, the electronic device (30) can determine a similar data set (1810) by performing an embedding similarity search for a search target and determining similar data greater than or equal to a preset threshold with respect to the input text (1710).

[0272] In one embodiment, the electronic device (30) can obtain a search result including a distribution of response guideline indices for a similar data set (1810) by performing an embedding similarity-based search. That is, the electronic device (30) can obtain a search result including a distribution of response guideline indices for similar data included in the similar data set (1810).

[0273] Graph (1820) is an example of the distribution of response guideline indices. For example, all similar data included in the similar data set (1810) may include a first response guideline index, a second response guideline index, or a third response guideline index. In this case, each response guideline index may represent a response guideline index (842) stored together with the conversation example embedding (841) described above with reference to FIG. 8, etc.

[0274] For example, the first similar data (1811) may correspond to any one of the first response guideline index, the second response guideline index, and the third response guideline index, and the second similar data (1812) may also correspond to any one of the first response guideline index, the second response guideline index, and the third response guideline index, and the response guideline index corresponding to the first similar data (1811) may be different from the response guideline index corresponding to the second similar data (1812).

[0275] Meanwhile, similar data can represent conversation example data grouped by a higher conversation rule, and since each response guideline index can correspond to one conversation rule, the distribution of the response guideline index can be understood as a distribution indicating which conversation rule's sub-conversation example data each similar data included in the similar data set (1810) is.

[0276] In one embodiment, the electronic device (30) may select at least one response guideline from a plurality of previously generated response guidelines based on search results. At this time, the selected at least one response guideline may be composed of at least one of the response guidelines corresponding to each of the similar data.

[0277] For example, if all similar data included in the similar data set (1810) corresponds to the first response guideline index, the second response guideline index, or the third response guideline index, the selected at least one response guideline may be composed of at least one of the first response guideline indicated by the first response guideline index, the second response guideline indicated by the second response guideline index, and the third response guideline indicated by the third response guideline index.

[0278] At this time, at least one selected response guideline can be used to generate a prompt template. That is, the range of response guidelines to be used to generate a prompt template can be primarily narrowed based on embedding similarity, thereby increasing the likelihood of generating an appropriate prompt template.

[0279] FIG. 19 is an exemplary diagram illustrating the process of selecting at least one response guideline from a plurality of previously generated response guidelines based on search results.

[0280] In one embodiment, the electronic device (30) may select at least one response guideline among the response guidelines corresponding to each of the similar data based on the distribution of the response guideline index included in the search results. For example, the electronic device (30) may select a response guideline corresponding to a response guideline index that shows a relatively high proportion.

[0281] Referring to FIG. 19, graphs (1910) and (1920) are examples of the distribution of response guideline indices appearing in the similar data set (1810) shown in FIG. 18.

[0282] Graph (1910) shows a distribution of response guideline indices in which one response guideline index shows a higher proportion than other response guideline indices, and graph (1920) shows a distribution of response guideline indices in which multiple response guideline indices show a higher proportion than other response guideline indices.

[0283] In one embodiment, the electronic device (30) can calculate the ratio of each response guideline index. As an example, the ratio of the first response guideline index can be calculated as 80%, the ratio of the second response guideline index as 10%, and the ratio of the third response guideline index as 10%. At this time, the electronic device (30) can select a first response guideline corresponding to the first response guideline index.

[0284] As another example, the ratio of the first response guideline index can be calculated as 50%, the ratio of the second response guideline index as 40%, and the ratio of the third response guideline index as 10%. At this time, the electronic device (30) may select only the first response guideline corresponding to the first response guideline index, or select the second response guideline corresponding to both the first response guideline and the second response guideline index.

[0285] In one embodiment, the electronic device (30) may select a response guideline corresponding to a response guideline index having a maximum value, and a response guideline corresponding to a response guideline index having a difference from the maximum value within a threshold value. In another embodiment, the electronic device (30) may select a response guideline corresponding to a response guideline index having a maximum value, and a response guideline corresponding to a response guideline index having a ratio to the maximum value within a predetermined range.

[0286] Afterward, the electronic device (30) can generate a prompt template corresponding to the input text based on at least one selected response guideline. At this time, the specific process of generating the prompt template based on the response guideline may be the same as described above with reference to FIG. 12, etc.

[0287] Meanwhile, the electronic device (30) may select a single response guideline based on the search results, but may also select multiple response guidelines.

[0288] In one embodiment, the electronic device (30) can generate a prompt template to have a single task format based on the selection of a single response guideline as at least one response guideline, and can generate a prompt template to have multiple task formats based on the selection of multiple response guidelines as at least one response guideline.

[0289] In the present disclosure, a single task format refers to a format that a prompt template may have when the task currently being pursued by a conversational AI agent is predicted to be a single task, and a multiple task format refers to a format that a prompt template may have when the task currently being pursued by a conversational AI agent is predicted to be at least one of a plurality of tasks.

[0290] In one embodiment, a single task format may include a format containing only a description of a single selected response instruction. For example, the first template (1210) and the second template (1220) illustrated in FIG. 12 may each be understood as examples of prompt templates having a single task format.

[0291] Meanwhile, the multi-task format may include all descriptions of multiple selected response instructions, and furthermore, may include additional descriptions instructing to provide an appropriate response by appropriately grasping the meaning of input text predicted to be related to various scenarios.

[0292] For example, if the search result has a distribution of response guideline indices such as graph (1910), the electronic device (30) can prevent unnecessary data from being included in the input prompt and ensure that relatively important data is included in the input prompt by generating a description of the first response guideline corresponding to the first response guideline index and a prompt template having a structure and slots according to the first response guideline.

[0293] As another example, if the search results have a distribution of response guideline indices such as graph (1920), the electronic device (30) may generate a description of a first response guideline corresponding to the first response guideline index, a description of a second response guideline corresponding to the second response guideline index, and a prompt template having a structure and slots according to the first response guideline and the second response guideline. In this case, the prompt template may further include additional descriptions instructing to provide an appropriate response by appropriately identifying the meaning of input text predicted to be related to both the first response guideline and the second response guideline.

[0294] Through this, inefficient prompt design can be prevented by optimizing the length of the input prompt, and uncertainty regarding the input prompt can be minimized when the intent of the input text is clear. In addition, in relatively complex situations, such as when the intent of the input text is unclear, performance degradation of the language model can be prevented by selecting and minimizing multiple tasks predicted to be the objective.

[0295] FIG. 20 is an exemplary drawing for explaining a method of using search results according to one embodiment.

[0296] Referring to FIG. 20, the electronic device (30) can determine a similar data set (2010) composed of conversation example data corresponding to each conversation example having high similarity to the text embeddings by performing a search based on text embeddings. At this time, the specific process for determining the similar data set (2010) may be the same as the specific process for determining the similar data set (1810) described above with reference to FIG. 18.

[0297] In one embodiment, each of the similar data may further include various variables specific to the similar data. For example, the types of variables specific to the similar data may include, but are not limited to, the type of language corresponding to the conversation example and the type of task required by the conversation example (e.g., classification task or technology-related task, etc.).

[0298] Graph (2020) is an example of the distribution of a specific variable appearing in a similar dataset (2010). Variables 1, 2, and 3 may represent types of languages ​​such as English and Korean, or types of tasks such as classification tasks and technology-related tasks.

[0299] As an example, the electronic device (30) can obtain a search result including a distribution of language types for a similar data set (2010) by performing an embedding similarity-based search.

[0300] Response instructions may be written in advance in each language or written in a specific language and translated into another language in real time. In one embodiment, the electronic device (30) may generate a response in a language corresponding to the input text even when the input text is input in a language different from the pre-set language, by selecting a response instruction written in a language having a maximum value based on the distribution of language types for a similar data set (2010) or by translating a response instruction written in a specific language into a language having a maximum value.

[0301] In addition, in one embodiment, the electronic device (30) can select a language model used to generate output text from among a plurality of previously trained language models based on search results.

[0302] For example, a plurality of language models may include a first language model trained with relatively large weights for processing complex tasks and a second language model trained with relatively small weights for processing simple tasks. In this case, the electronic device (30) may select the first language model or the second language model as the language model used to generate output text based on search results.

[0303] As an example, the electronic device (30) can generate a prompt template having a single task format or a multiple task format based on search results, select a second language model based on the generated prompt template having a single task format, and select a first language model based on the generated prompt template having a multiple task format.

[0304] As another example, multiple language models may include a first language model specialized for English text processing and a second language model specialized for Korean text processing. In this case, the electronic device (30) may select the first language model or the second language model as the language model used to generate output text based on the search results.

[0305] As an example, the electronic device (30) can select a language having a maximum value in the distribution of language types based on the search results, and if the language having the maximum value is English, it can select a first language model, and if the language having the maximum value is Korean, it can select a second language model.

[0306] Through this, along with the generation of a suitable input prompt, a language model capable of appropriately processing the input prompt using information identified during the input prompt generation process can be selected, thereby generating a response optimized for various situations.

[0307] FIG. 21 is an exemplary drawing for explaining the process of generating output text according to one embodiment.

[0308] Referring to FIG. 21, in step 2110, the electronic device (30) can obtain search results by performing a search based on text embeddings on at least a portion of a previously generated conversation example database. At this time, the specific process of obtaining search results may include the process of obtaining search results described above with reference to FIG. 17 and FIG. 18, etc.

[0309] For example, search results may include the distribution of various variables by conversation example data, such as response guideline indices.

[0310] In step 2120, the electronic device (30) may select at least one response guideline from a plurality of previously generated response guidelines based on the search results. At this time, the specific process of selecting the response guideline may include the process of selecting the response guideline described above with reference to FIG. 19, etc.

[0311] At this time, the range of response guidelines can be limited to response guidelines corresponding to each conversation example data included in the search results obtained in step 2110, and accordingly, the suitability of the prompt template for generating an appropriate response can be improved.

[0312] In step 2130, the electronic device (30) may select a language model to be used for generating output text from among a plurality of previously trained language models based on search results. At this time, the specific process of selecting the language model may include the process of selecting the language model described above with reference to FIG. 20, etc.

[0313] For example, the electronic device (30) can select the language model most suitable for processing the input prompt by determining the language and / or type of task most suitable for the input text from the search results.

[0314] In step 2140, the electronic device (30) can generate output text based on an input prompt generated based on at least one selected response instruction using a selected language model.

[0315] In one embodiment, the electronic device (30) can generate a prompt template based on selected response instructions and generate an input prompt by applying the generated prompt template to at least one of the vehicle status, call execution result, previously generated conversation history, and input text. For example, the electronic device (30) can generate an input prompt by inserting at least some of the vehicle status, call execution result, previously generated conversation history, and input text into each slot of the prompt template.

[0316] Subsequently, the electronic device (30) can generate output text based on an input prompt using a selected language model. At this time, generating output text using the selected language model may include generating a response by loading the weights of the corresponding language model stored in memory (33) or generating a response by calling a language model API hosted through an external server, etc.

[0317] Meanwhile, embodiments according to the present disclosure may be implemented in the form of a computer program that can be executed through various components on a computer, and such a computer program may be recorded on a computer-readable medium. In this case, the medium may include, but is not limited to, magnetic media such as hard disks, floppy disks, and magnetic tapes, optical recording media such as CD-ROMs and DVDs, magneto-optical media such as floptical disks, and hardware devices specifically configured to store and execute program instructions such as ROM, RAM, and flash memory.

[0318] Meanwhile, the above computer program may be one specifically designed and configured for the present disclosure or one known and available to those skilled in the art of computer software. Examples of computer programs may include machine code, such as that produced by a compiler, as well as high-level language code that can be executed by a computer using an interpreter, etc.

[0319] According to one embodiment, the method according to various embodiments of the present disclosure may be provided by being included in a computer program product. The computer program product may be traded between a seller and a buyer as a product. The computer program product may be distributed in the form of a device-readable storage medium (e.g., compact disc read-only memory (CD-ROM)), or distributed online (e.g., download or upload) through an application store (e.g., Play Store™) or directly between two user devices. In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily created in a device-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or a relay server.

[0320] Unless explicitly stated otherwise, the steps constituting the method according to the present disclosure may be performed in a suitable order. The present disclosure is not necessarily limited by the order in which the steps are described. The use of any examples or exemplary terms (e.g., etc.) in the present disclosure is merely for the purpose of describing the present disclosure in detail and, unless limited by the claims, the scope of the present disclosure is not limited by such examples or exemplary terms. Furthermore, those skilled in the art will understand that various modifications, combinations, and changes may be made according to design conditions and factors within the scope of the claims or equivalents to which they are added.

[0321] Accordingly, the scope of the present disclosure should not be limited to the embodiments described above, and all scopes equivalent to or equivalently modified from the claims set forth below, as well as the claims set forth below, shall be considered to fall within the scope of the scope of the present disclosure.

Claims

Claim 1 A method for generating a prompt template corresponding to an input text, comprising: generating a text embedding based on input text from a vehicle occupant; obtaining search results by performing a search based on the text embedding on at least a portion of a previously generated dialogue example database; selecting at least one response guideline from a plurality of previously generated response guidelines based on the search results; and generating a prompt template corresponding to the input text based on the at least one response guideline. Claim 2 A method according to claim 1, wherein the conversation example database is composed of conversation example data corresponding to each of a plurality of conversation examples, and each of the conversation example data corresponding to each of the plurality of conversation examples includes a conversation example embedding generated based on the conversation example, and a response guideline index representing a response guideline corresponding to the conversation example. Claim 3 In claim 2, the conversation example database is composed of conversation example data grouped into items of a plurality of conversation rules, and each of the plurality of response guidelines corresponds one-to-one with each of the plurality of conversation rules. Claim 4 In claim 2, the search comprises a search based on embedding similarity based on a conversation example embedding included in each conversation example data corresponding to each of the plurality of conversation examples and the text embedding. Claim 5 In claim 4, the step of obtaining the search result comprises: obtaining the search result including the distribution of response guideline indices for conversation example data corresponding to each of a preset number of conversation examples by performing the embedding similarity-based search. Claim 6 In claim 5, the method wherein the at least one response guideline is composed of at least one of the response guidelines corresponding to each of the preset number of conversation examples. Claim 7 In claim 4, each conversation example data corresponding to each of the plurality of conversation examples further includes a language type corresponding to the conversation example, and the step of obtaining the search result includes the step of obtaining the search result including a distribution of language types for conversation example data corresponding to each of a preset number of conversation examples by performing the embedding similarity-based search. Claim 8 A method according to claim 1, wherein the step of generating the prompt template comprises: generating the prompt template to have a single task format based on the selection of a single response guideline as the at least one response guideline, and generating the prompt template to have multiple task formats based on the selection of multiple response guidelines as the at least one response guideline. Claim 9 The method according to claim 1, wherein the step of generating the prompt template further comprises: generating an input prompt for a pre-trained language model by applying the prompt template to at least one of a vehicle state, a call execution result, a pre-generated conversation history, and the input text; and generating an output text based on the input prompt using the language model; wherein the call execution result is generated based on the input text. Claim 10 In claim 9, the step of generating the prompt template further comprises the step of selecting the language model used to generate the output text from among a plurality of language models that have been trained based on the search results. Claim 11 A method according to claim 1, wherein the step of generating the text embedding comprises: generating the text embedding based on at least one of a previously generated conversation history and previously generated conversation parameters and the input text. Claim 12 A device for generating a prompt template corresponding to an input text, comprising: a memory in which at least one program is stored; and a processor that operates by executing the at least one program; wherein the processor generates a text embedding based on input text of a vehicle occupant, obtains a search result by performing a search based on the text embedding on at least a portion of a previously generated dialogue example database, selects at least one response guideline from a plurality of previously generated response guidelines based on the search result, and generates a prompt template corresponding to the input text based on the at least one response guideline. Claim 13 A computer-readable recording medium storing a program for executing the method according to claim 1 on a computer.