Method and apparatus for generating multi-intent utterance dataset

By using various conjunctions and complex patterns in the generation of multi-intention discourse datasets, the problem of insufficient diversity of conjunction words in the existing datasets is solved, and more complex and diverse intent datasets are generated, which improves the ability of dialogue systems to understand user intentions.

CN120162633APending Publication Date: 2025-06-17HYUNDAI MOTOR CO LTD +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411789815.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-06-03
Filing Date
2024-12-06
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

The existing multi-intention discourse datasets have difficulty in accurately understanding multiple intents of users due to insufficient diversity of connective words.

Method used

By connecting discourses using various conjunctions and complex patterns, a more complex and diverse intent dataset than existing multi-intent datasets are generated. Specific methods include collecting a single-intention dataset, preprocessing it, selecting a single-intention utterance to be merged, and generating multi-intention utterances through different merging methods.

Benefits of technology

The generated multi-intention dataset can more accurately reflect the diversity and complexity of real-world dialogue, helping task-oriented dialogue systems better understand and grasp multiple user intentions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120162633A_ABST
    Figure CN120162633A_ABST
Patent Text Reader

Abstract

A computationally implemented method for generating a multi-intent dataset includes: collecting a single-intent dataset; preprocessing the collected single-intention data set while keeping the meaning and the structure of the utterance in the collected single-intention data set; selecting a plurality of single-intention utterances to be merged from the preprocessed single-intention data set; and merging the selected plurality of single-intent utterances into one multi-intent utterance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a method and apparatus for generating a multi-intent utterance dataset. Background Art

[0002] The following description merely provides background information related to the present embodiment and does not constitute related art.

[0003] A task-oriented dialogue system (TOD) is an artificial intelligence system that understands specific goals or needs by analyzing user utterances and generates responses to specific goals or needs. Task-oriented dialogue systems can be used in various situations, such as online shopping, reservation systems, customer service, etc. Task-oriented dialogue systems need to be able to accurately understand user needs and effectively handle user needs.

[0004] "Intent" refers to the purpose or goal that a user intends to achieve by interacting with a dialogue system (such as a chatbot). For example, an intent can be the underlying meaning or need behind a user asking a question. A multi-intent utterance is when a user expresses two or more different needs or intentions in a single utterance. An example of such an utterance is a sentence like "Book a flight for tomorrow morning and recommend nearby hotels." Users often speak with multiple purposes. However, task-oriented dialogue systems traditionally prefer to interpret user utterances as being related to a single purpose.

[0005] According to a study published in 2019, In the TOD dataset created by (a company engaged in e-commerce and artificial intelligence), more than half of the utterances were reported as multi-intent utterances. When designing and developing task-oriented dialogue systems, it is becoming increasingly important to grasp the various intentions and needs of users.

[0006] Despite great interest in multi-intent utterances, resources to support this research are very limited. MixATIS and MixSNIPS are currently widely used datasets of multi-intent utterances, generated by merging two or more single-intent utterances. MixATIS and MixSNIPS always include one of "and", "and then", and "and" to merge single-intent utterances. However, MixATIS and MixSNIPS have been criticized for the insufficient diversity of connectives used to generate the datasets. MixATIS and MixSNIPS exploit simple merging patterns as they only use AND variants. The simple pattern allows multi-intent detection models to learn too easily to identify the number of intents. For example, a multi-intent detection model can easily identify the number of intents in an utterance by counting the occurrences of the conjunction "and" or discerning the presence of ", (comma)". Existing research has not paid enough attention to this issue and has relied solely on the datasets. Summary of the invention

[0007] The main aspects of the present disclosure are directed to providing a method and apparatus for generating a dataset characterized by more complex and diverse intents than existing multi-intent datasets by connecting utterances using various conjunctions and complex patterns.

[0008] Aspects of the present disclosure are not limited to the above contents, and other aspects not mentioned herein should be clearly understood by those of ordinary skill in the art from the following description.

[0009] According to one aspect of the present disclosure, a computationally achievable method for generating a multi-intent dataset includes collecting a single-intent dataset. The method also includes preprocessing the collected single-intent dataset while retaining the meaning and structure of the utterances in the collected single-intent dataset. The method also includes selecting a plurality of single-intent utterances to be merged from the preprocessed single-intent dataset. The method also includes merging the plurality of selected single-intent utterances into a multi-intent utterance.

[0010] According to another aspect of the present disclosure, a device for generating a multi-intent dataset includes a memory configured to store one or more instructions and at least one processor configured to execute the one or more instructions stored in the memory. By executing one or more instructions, the at least one processor is configured to collect a single-intent dataset. The at least one processor is also configured to preprocess the collected single-intent dataset while retaining the meaning and structure of the utterances in the collected single-intent dataset. The at least one processor is also configured to select multiple single-intent utterances to be merged from the preprocessed single-intent dataset. The at least one processor is also configured to merge multiple selected single-intent utterances into one multi-intent utterance.

[0011] According to an embodiment of the present disclosure, a multi-intent dataset can be generated using complex patterns and various conjunctions. Therefore, a multi-intent dataset can be obtained using different conjunctions without resorting to the simple conjunction rules used in MixATIS and MixSNIPS.

[0012] According to an embodiment of the present disclosure, a multi-intent dataset can be generated to reflect various situations or contexts. Therefore, the diversity and complexity of real-world conversations between people can be reflected in the multi-intent dataset. In addition, using the generated multi-intent dataset allows a task-oriented dialogue system to more accurately grasp and understand multiple intents from the user's utterances.

[0013] The effects of the present disclosure are not limited to the above-mentioned contents, and other effects not mentioned herein should be clearly understood by those of ordinary skill in the art from the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] Figure 1Schematically illustrates a block diagram of an apparatus for generating a multi-intent dataset according to an embodiment of the present disclosure.

[0015] Figure 2 is a view illustrating a process for generating a multi-intent dataset by merging single-intent datasets.

[0016] Figure 3 is a flowchart illustrating a method for generating a multi-intent dataset according to an embodiment of the present disclosure.

[0017] Figure 4 is a block diagram schematically showing a computing device that can be used to implement the method or device according to the present disclosure. DETAILED DESCRIPTION

[0018] Hereinafter, some embodiments of the present disclosure should be described in detail with reference to the accompanying drawings. In the following description, similar reference numerals represent the same or similar elements, although these elements are shown in different drawings. In addition, in the following description of some embodiments, for the purpose of clarity and brevity, detailed descriptions of known functions and configurations incorporated therein are omitted.

[0019] In addition, various terms such as first, second, A, B, (a), (b), etc. are only used to distinguish one component from another component, and are not intended to imply or suggest the substance, order or sequence of the components. Throughout this disclosure, when a part "includes" or "comprises" a component, the part is meant to further include other components, and other components are not excluded unless otherwise specifically stated. Terms such as "unit", "module", etc. refer to one or more units for processing at least one function or operation, and the unit can be implemented by hardware, software or a combination thereof.

[0020] The following detailed description, together with the appended drawings, are intended to describe embodiments of the present disclosure and are not intended to represent the only embodiments in which the present disclosure may be practiced.

[0021] Figure 1 Schematically illustrates a block diagram of an apparatus for generating a multi-intent dataset according to an embodiment of the present disclosure.

[0022] The apparatus 10 for generating a multi-intent dataset according to an embodiment of the present disclosure may include all or part of a data preprocessing module 100, a data selection module 120, a data merging module 140 or a data checking module 160. It should be noted that Figure 1 All blocks shown are essential components, and in other embodiments, some of the included blocks may be added, modified, or deleted. Figure 1The components shown show functionally different elements, and at least one component may be implemented in a manner of being integrated together in an actual physical environment.

[0023] The data preprocessing module 100 preprocesses the data for merging single-intent utterances while retaining their original meaning or structure. The data selection module 120 selects utterances to be merged from the single-intent dataset. The data merging module 140 generates a multi-intent dataset by merging the selected utterances. The data checking module 160 checks the generated multi-intent dataset by using a generative model.

[0024] Figure 2 is a view illustrating a process for generating a multi-intent dataset by merging single-intent datasets.

[0025] Methods for merging single-intent datasets include "explicit concatenation" and "implicit concatenation". Explicit concatenation is a concatenation method in which connectives are explicitly used to connect utterances. Explicit concatenation includes "AND variant method" and "various connective methods".

[0026] The AND variant method is a method to merge two or more single-intent utterances by using one or more of “and”, “and then”, “and”, or “, (comma)”. It is the same method used to generate MiXATIS and MixSNIPS, which are currently widely used multi-intent datasets.

[0027] The various conjunction methods are methods of combining two or more single-intent utterances by using one or more of the following: "and", "and then", "also", ", (comma)", "; (semicolon)", "or", "before", "after", "also", "finally". Single-intent utterances can be combined by using various conjunctions without relying entirely on the AND variant method.

[0028] Implicit linking is a method of merging in which no explicit connectives are used to link utterances. Implicit linking includes the "conjunction removal method", "gerund phrase method", "ellipse method" and "coreference method".

[0029] The conjunction removal method is a method for merging single-intent utterances by removing conjunctions. With this method, single-intent utterances can be merged easily and simply. The conjunction removal method effectively reflects the intuition that speakers tend to use shorter utterances.

[0030] The gerund phrase method is a method of merging sentences by transforming specific utterances into gerund phrases (-ing form of verbs). It emphasizes the concurrence of multiple sentences and allows the merging of single-intention utterances. The gerund phrase method is applicable to utterances that meet certain conditions. The utterances that meet certain conditions can be sentences starting with a verb, or can be some interrogative sentences that can be naturally converted into participle structures.

[0031] The ellipsis method is a method to merge single-intent utterances by arbitrarily eliminating redundant expressions in multiple sentences.

[0032] The coreference method is a method to merge single-intent utterances by eliminating redundant expressions in multiple sentences or replacing redundant expressions with pronouns.

[0033] Table 1 shows the concatenation results of two single-intent utterances by various merging methods.

[0034] Table 1:

[0035]

[0036] Single intent utterance 1 is "Play my 88 key playlist", and the intent of single intent utterance 1 is "PlayMusic". Single intent utterance 2 is "Add another song to my 88 key playlist", and the intent of single intent utterance 2 is "AddToPlaylist".

[0037] By combining single-intent utterance 1 and single-intent utterance 2 using the AND variation method, “Play my 88-key playlist and add another song to my 88-key playlist” can be formed. Single-intent utterances are combined by adding “and” between the sentences.

[0038] “Add another song to my 88-key playlist” can be formed by merging single-intention utterance 1 and single-intention utterance 2 using the gerund phrase method.

[0039] “Play my 88-key playlist, add another song to my 88-key playlist” can be formed by merging single-intent utterance 1 and single-intent utterance 2 using the conjunction removal method.

[0040] “Play my 88-key playlist and add another song” can be formed by merging single-intent utterance 1 and single-intent utterance 2 using the ellipsis method.

[0041] “Play my 88-key playlist and add another song to it” can be formed by merging single-intention utterance 1 and single-intention utterance 2 using the coreference method. Single-intention utterance 1 and single-intention utterance 2 are merged by eliminating the redundant expression “my 88-key playlist” and replacing it with the pronoun “it”.

[0042] Figure 3 is a flowchart illustrating a method for generating a multi-intent dataset according to an embodiment of the present disclosure.

[0043] In step S300, the apparatus 10 for generating a multi-intent dataset collects a single-intent dataset for merging into the multi-intent dataset. Although the single-intent dataset used in one embodiment of the present disclosure is in English, the dataset is not limited thereto. Intent needs to be clearly defined in the single-intent dataset.

[0044] In the embodiments of the present disclosure, single-intent datasets ATIS, SNIPS, Banking77, and CLINC150 are used. ATIS (Air Travel Information System) is a natural language processing dataset focused on air travel information. SNIPS is a speech recognition training dataset covering various fields (weather, music, etc.). Banking77 is a dataset focused on bank-related queries. CLINC150 is a dataset dealing with fields such as banking, travel, restaurants, etc. ATIS, SNIPS, Banking77, and CLINC150 can be used for data merging.

[0045] In step S302, the device 10 for generating a multi-intent dataset preprocesses the single-intent dataset. In the preprocessing step, the original meaning and structure of the utterances in the single-intent dataset need to be retained. For example, during the preprocessing process, the uppercase letters in the dataset are converted to lowercase letters. Then the punctuation marks ".", "?", "!" are removed.

[0046] In step S304, the apparatus 10 for generating a multi-intent dataset selects a single-intent utterance to be merged. Methods for selecting utterances to be merged from a single-intent dataset include "random selection" and "selection based on cosine similarity". Random selection is a method for randomly selecting utterances with different intents. Because the utterances are randomly selected, the selected utterances may not be similar in words or structures. Therefore, in the case of random selection, it may be difficult to perform the omission method and the coreference method, both of which are implicit connection methods.

[0047] Cosine similarity based selection is a method for calculating cosine similarity between utterances and selecting utterances of a specific value or higher. Cosine similarity is a method for measuring the similarity between two vectors by calculating the angle between the two vectors. The cosine value is calculated by using the dot product of the two vectors and the magnitude of the vectors. The closer the cosine value is to 1, the more similar the two vectors are in direction. Cosine similarity is a measure of how similar given words are in structure or format. The closer the cosine similarity between words is to 1, the more similar the words are in structure or format. Cosine similarity can be applied after the text is converted into a vector.

[0048] In order to adopt the omission method or the coreference method, redundant expressions are required in the sentence. Therefore, in the present disclosure, utterances are selected based on cosine similarity in order to adopt the omission method or the coreference method.

[0049] In the present disclosure, two or more utterances having a cosine similarity of a specific value or higher are selected. For example, in the present disclosure, utterances having a cosine similarity of 0.7 or higher are basically selected. As another example, in the present disclosure, utterances having a minimum cosine similarity of 0.5 are selected depending on the data set. In the present disclosure, the specific value is not fixed, and the present disclosure is not limited to the above examples.

[0050] When selecting three utterances, cosine similarity is measured for each case of two utterances, and when all utterances have a cosine similarity of a certain value or higher, the utterance is selected. Using selection based on cosine similarity, utterances with similar structures or containing redundant words can be retrieved, and implicit connections may be easier than when utterances are selected randomly.

[0051] In step S306, the device 10 for generating a multi-intent dataset generates a multi-intent utterance by merging two or more single-intent utterances. The merging methods include "manual rule-based concatenation" and "concatenation using a generative artificial intelligence model". Manual rule-based concatenation is a method for merging single-intent utterances based on various manually defined rules. Manual rule-based concatenation includes AND variant methods, various conjunction methods, conjunction removal methods, and gerund phrase methods. Concatenation using a generative artificial intelligence model is a method for merging two or more single-intent utterances by using a pre-trained generative artificial intelligence language model. Concatenation using a generative artificial intelligence model includes AND variant methods, various conjunction methods, conjunction removal methods, gerund phrase methods, omission methods, and co-reference methods. In the present disclosure, a prior art generative artificial intelligence model capable of both explicit and implicit concatenation is used to generate a complex multi-intent dataset. For example, ChatGPT can be used for various natural language processing tasks, such as summarization, providing intelligence, and translation, and shows high performance.

[0052] Depending on the utterance selection method, a multi-intent dataset can be generated from randomly selected utterances by using manual rule-based concatenation, and a multi-intent dataset can be generated from utterances selected based on cosine similarity by using a concatenation method using a generative artificial intelligence model. Two or more pre-selected utterances and augmented prompts are fed into a generative artificial intelligence model to perform explicit or implicit concatenation and generate a multi-intent dataset.

[0053] Table 2 depicts an example of providing hints to a generative artificial intelligence model in the present disclosure.

[0054] Table 2:

[0055]

[0056] The prompt format fed to the generative artificial intelligence model is not limited to the specific format or the format shown in Table 1. For example, a prompt to add or remove instructions can be fed. As another example, a prompt to add, remove or modify an example can be fed.

[0057] Step S308 is a step of checking the multi-intent dataset generated by the generative artificial intelligence model.

[0058] Generative AI models do not always guarantee correct answers. Even when examples and instructions are clearly presented, the model may generate unwanted results. For example, when generating multi-intent utterances by merging single-intent utterances, the generative AI model may distort the intent or even partially remove the intent. In addition, it may fail to merge and, as a result, generate incorrect sentences. This may make users doubt the credibility of the results of the generative AI model. Therefore, steps are needed to check the multi-intent datasets generated by the generative AI model.

[0059] The evaluation metrics used to determine whether the generative AI model correctly performs explicit concatenation include word frequency measure, linker frequency measure, and pronoun frequency measure. The evaluation metrics provide insights into estimating the degree of language transformation caused by concatenation. All metrics have values ​​of 0 or 1.

[0060] Mathematical formula 1 is a formula representing word frequency measurement.

[0061] Mathematical formula 1:

[0062]

[0063] If the number of words in the post-concatenation utterance is less than or equal to the total number of words in the pre-concatenation utterance, let the metric be 1, otherwise 0. A metric value of 1 may indicate that the ellipsis or coreference method was used in the explicit concatenation method, or that no words were added during the concatenation process.

[0064] Mathematical formula 2 is a formula representing the link word frequency measure.

[0065] Mathematical formula 2:

[0066]

[0067] If the number of linkers in the post-join utterance is less than or equal to the total number of linkers in the pre-join utterance, let the measure be 1; otherwise, let it be 0. A measure value of 1 may indicate that no explicit linkers were used when merging utterances.

[0068] Mathematical formula 3 is a format for expressing pronoun frequency measurement.

[0069] Mathematical formula 3:

[0070]

[0071] If the number of pronouns in the post-concatenation utterance is greater than the total number of pronouns in the pre-concatenation utterance, let the measure be 1; otherwise, let it be 0. A measure value of 1 may indicate that redundant expressions are replaced by pronouns when merging two or more utterances.

[0072] Therefore, it can be assumed that if the metric is 1, the utterance has been generated as intended by the user. However, there may be cases where there are significant differences before and after concatenation. For example, if the number of linked words in the original utterance increases from 2 to 6 after concatenation, it may indicate an unnecessary paraphrase that is different from the original utterance, while a significant decrease may indicate that the utterance was ignored during concatenation and the original intent of the utterance could not be preserved. In this case, concatenation is considered unsuccessful and the concatenated results may be excluded from the final dataset.

[0073] In step 308, there are three checking stages. First, only sentences with a metric value of 1 are selected. If the number of words in the utterance after concatenation is less than or equal to the total number of words in the utterance before concatenation, the value of the word frequency metric is 1. If the number of linking words in the utterance after concatenation is less than or equal to the total number of linking words in the utterance before concatenation, the value of the linking word frequency metric is 1. If the number of pronouns in the utterance after concatenation is greater than the total number of pronouns in the utterance before concatenation, the value of the pronoun frequency metric is 1. In the next step, TFMN, which has cutting-edge capabilities in multi-intent detection, is used to filter out sentences characterized by failed intent detection. Finally, experts who understand artificial intelligence will check the results and remove sentences with damaged intent.

[0074] In step S310, an evaluation is made on how to generate a multi-intent dataset using an artificial intelligence model and a generated dataset. The model is trained on a dataset generated in a traditional method and evaluated using a multi-intent dataset generated according to the present disclosure. The datasets generated in the traditional method refer to MixSNIPS and MixATIS. The datasets generated by the method of creating MixSNIPS and MixATIS may include MixBanking77 and MixCLINC150. The datasets generated by the dataset generation method of the present disclosure may include BlendSNIPS, BlendATIS, BlendBanking77, and BlendCLINC150.

[0075] The models used in supervised learning are TFMN and SLIM. The TFMN model predicts the number of intents in a multi-intent utterance. Subsequently, the TFMN model generates the most likely intent. If the output generated by the neural network for an intent exceeds a set threshold, the SLIM model selects that intent. The neural network predicts the probability of each intent and then passes the predicted probability through an activation function. Afterwards, if the probability of an intent is greater than or equal to the set threshold, the intent is considered to exist and is generated as the final result. For example, the set threshold can be 0.5.

[0076] In unsupervised learning, evaluation is performed via ChatGPT.

[0077] Accuracy is used to evaluate the multi-intent detection performance. Table 3 shows the evaluation results.

[0078] Table 3:

[0079]

[0080] These models show reasonable performance when trained and evaluated on datasets generated in traditional methods through supervised learning. On the other hand, when trained on datasets generated in traditional methods and evaluated on datasets generated according to the present disclosure, these models show significant performance degradation, some of which are reduced by 40%. This result shows that traditional dataset generation methods lack the complexity required to fully evaluate multi-intent detection capabilities.

[0081] Although replacing the training data with the datasets generated according to the present disclosure does lead to some performance recovery, they are less accurate than the datasets generated in the traditional method. This means that the data generation method according to the present disclosure is inherently more complex.

[0082] In addition, it can be found that the performance of unsupervised learning is poor. This indicates that the generative AI model has not been adapted to the multi-intent detection task.

[0083] In step S312, a final dataset is generated by including both multi-intent utterances generated by manual rule-based concatenation and checked multi-intent utterances generated by concatenation using a generative artificial intelligence model.

[0084] Figure 4 is a block diagram schematically showing a computing device that can be used to implement the method or device according to the present disclosure.

[0085] The computing device 40 may include some or all of the memory 400, the processor 420, the storage device 440, the input / output interface 460, or the communication interface 480. The computing device 40 may structurally and / or functionally include at least some of the data preprocessing module 100, the data selection module 120, the data merging module 140, or the data checking module 160. The computing device 40 may be a fixed computing device such as a desktop computer, a server, an AI accelerator, etc., or may be a portable computing device such as a laptop computer, a smart phone, etc.

[0086] The memory 400 may store a program that allows the processor 420 to perform methods or operations according to various embodiments of the present disclosure. For example, the program may include a plurality of instructions that can be executed by the processor 420, and when the processor 420 executes the plurality of instructions, the program may be executed. Figure 3 The method shown.

[0087] The memory 400 may be a single memory or multiple memories. In this case, the information required to perform the methods or operations according to various embodiments of the present disclosure may be stored in a single memory or in multiple memories in a distributed manner. If the memory 400 includes multiple memories, the multiple memories may be physically separated.

[0088] The memory 400 may include at least one of a volatile memory or a nonvolatile memory. The volatile memory includes an SRAM (Static Random Access Memory) or a DRAM (Dynamic Random Access Memory), and the nonvolatile memory includes a flash memory.

[0089] The processor 420 may include at least one core for executing at least one instruction. The processor 420 may execute instructions stored in the memory 400. The processor 420 may be a single processor or a plurality of processors.

[0090] Storage device 440 retains stored data even if power is cut off to computing device 40. For example, storage device 440 may include non-volatile memory and / or include a storage medium such as a tape, optical disk, or magnetic disk.

[0091] The program stored in the storage device 440 may be loaded onto the memory 400 before being executed by the processor 420. The storage device 440 may store a file made using a program language, and may load a program created from the file by a compiler or the like onto the memory 400. The storage device 440 may store data to be processed by the processor 420 and / or data processed by the processor 420.

[0092] The input / output interface 460 may include input devices such as a keyboard, a mouse, etc., and may include output devices such as a display device, a printer, etc. The user may trigger the processor 420 to execute a program and / or check processing results from the processor 420 via the input / output interface.

[0093] The communication interface 480 may provide access to an external network. For example, the computing device 40 may communicate with other devices (eg, the data preprocessing module 100 , the data selection module 120 , the data merging module 140 , or the data checking module 160 ) via the communication interface 480 .

[0094] Each component of the device or method according to the present disclosure can be implemented as hardware or software, or can be implemented as a combination of hardware and software. In addition, the function of each component can be implemented as software, and a microprocessor can be implemented to execute the function of the software corresponding to each component.

[0095] Various embodiments of the systems and techniques described herein may be implemented with digital electronic circuitry, integrated circuits, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), computer hardware, firmware, software, and / or combinations thereof. Various embodiments may include implementations with one or more computer programs executable on a programmable system. The programmable system includes at least one programmable processor, which may be a special purpose processor or a general purpose processor, coupled to receive data and instructions from, and transmit data and instructions to, a storage system, at least one input device, and at least one output device. A computer program (also referred to as a program, software, software application, or code) includes instructions for a programmable processor and is stored in a "computer-readable recording medium."

[0096] The computer-readable recording medium may include all types of storage devices on which computer-readable data may be stored. The computer-readable recording medium may be a non-volatile or non-temporary medium, such as a read-only memory (ROM), a random access memory (RAM), a compact disc ROM (CD-ROM), a magnetic tape, a floppy disk, or an optical data storage device. In addition, the computer-readable recording medium may further include a temporary medium such as a data transmission medium. In addition, the computer-readable recording medium may be distributed on a computer system connected via a network, and the computer-readable program code may be stored and executed in a distributed manner.

[0097] Although the operations shown in the flowchart / timing diagram in the present disclosure are performed sequentially, this is merely a description of the technical idea of ​​one embodiment of the present disclosure. In other words, a person skilled in the art of one embodiment of the present disclosure can understand that various modifications and changes can be made without departing from the present disclosure, that is, the order shown in the flowchart / timing diagram can be changed, and one or more operations in the operation can be performed in parallel. Therefore, the flowchart / timing diagram is not limited to the time order.

[0098] Although the embodiments of the present disclosure are described for the purpose of illustration, it will be understood by those skilled in the art that various modifications, additions and substitutions may be made without departing from the spirit and scope of the present disclosure. Therefore, for the sake of simplicity and clarity, the embodiments of the present disclosure have been described. The scope of the technical ideas of the present embodiments is not limited by the description. Therefore, it will be understood by those skilled in the art that the scope of the present disclosure should not be limited by the embodiments explicitly described above, but should be limited by the claims and their equivalents.

Claims

1. A computationally implementable method for generating a multi-intent dataset, the method comprising the following steps: Collect single intent dataset; Preprocessing the collected single-intent dataset while preserving the meaning and structure of the utterances in the collected single-intent dataset; Select multiple single-intent utterances to be merged from the preprocessed single-intent dataset; as well as The selected multiple single-intent utterances are merged into a multi-intent utterance.

2. The method according to claim 1, wherein: The steps to select multiple single-intent utterances to merge include: Randomly select utterances with different intentions.

3. The method according to claim 1, wherein: The steps to select multiple single-intent utterances to merge include: Based on cosine similarity, multiple single-intent utterances that are similar in sentence structure or format are selected.

4. The method according to claim 1, further comprising the steps of: The multi-intent dataset generated by the generative artificial intelligence model is examined by using a frequency metric, which is an evaluation scale configured to determine whether the merging was done correctly.

5. The method according to claim 2, further comprising the steps of: The multi-intent dataset is evaluated by using an artificial intelligence model and the generated dataset.

6. The method according to claim 1, wherein: The step of merging the selected multiple single-intent utterances into a multi-intent utterance comprises: The selected multiple single-intent utterances are merged into the one multi-intent utterance by using at least one of ";", "or", "before", "after", "in addition" or "last".

7. The method according to claim 1, wherein: The step of merging the selected multiple single-intent utterances into a multi-intent utterance comprises: The selected multiple single-intent utterances are merged into the one multi-intent utterance by removing conjunctions.

8. The method according to claim 1, wherein: The step of merging the selected multiple single-intent utterances into a multi-intent utterance comprises: The selected plurality of single-intent utterances are merged into the one multi-intent utterance by transforming a specific utterance into a gerund phrase.

9. The method according to claim 1, wherein: The step of merging the selected multiple single-intent utterances into a multi-intent utterance comprises: The selected multiple single-intent utterances are merged into the one multi-intent utterance by arbitrarily eliminating redundant expressions in the multiple sentences.

10. The method according to claim 1, wherein: The step of merging the selected multiple single-intent utterances into a multi-intent utterance comprises: The selected multiple single-intent utterances are merged into the one multi-intent utterance by eliminating redundant expressions in the multiple sentences and replacing the redundant expressions with pronouns.

11. A device for generating a multi-intent dataset, the device comprising: A memory configured to: store one or more instructions; and at least one processor configured to: execute one or more instructions stored in the memory, Wherein, by executing the one or more instructions, the at least one processor is configured to: Collect single intent dataset; Preprocessing the collected single-intent dataset while preserving the meaning and structure of the utterances in the collected single-intent dataset; selecting multiple single-intent utterances to be merged from the preprocessed single-intent dataset; and The selected multiple single-intent utterances are merged into a multi-intent utterance.

12. The device according to claim 11, wherein When selecting multiple single-intent utterances to be merged, the at least one processor is configured to: Randomly select utterances with different intentions.

13. The device according to claim 11, wherein: When selecting multiple single-intent utterances to be merged, the at least one processor is configured to: Based on cosine similarity, multiple single-intent utterances that are similar in sentence structure or format are selected.

14. The device according to claim 11, wherein: The at least one processor is configured to: Examining multi-intent datasets generated by generative AI models by using frequency metrics, and The frequency metric is an evaluation metric configured to determine whether the merging is completed correctly.

15. The device according to claim 12, wherein: The at least one processor is further configured to: The multi-intent dataset is evaluated by using an artificial intelligence model and the generated dataset.

16. The device according to claim 11, wherein When the selection is to merge the multiple single-intent utterances into one multi-intent utterance, the at least one processor is configured to: The selected multiple single-intent utterances are merged into the one multi-intent utterance by using at least one of ";", "or", "before", "after", "in addition" or "last".

17. The device according to claim 11, wherein: When merging the selected multiple single-intent utterances into one multi-intent utterance, the at least one processor is configured to: The selected multiple single-intent utterances are merged into the one multi-intent utterance by removing conjunctions.

18. The device according to claim 11, wherein When merging the selected multiple single-intent utterances into one multi-intent utterance, the at least one processor is configured to: The selected plurality of single-intent utterances are merged into the one multi-intent utterance by transforming a specific utterance into a gerund phrase.

19. The device according to claim 11, wherein: When merging the selected multiple single-intent utterances into one multi-intent utterance, the at least one processor is configured to: The selected multiple single-intent utterances are merged into the one multi-intent utterance by arbitrarily eliminating redundant expressions in the multiple sentences.

20. The device according to claim 11, wherein When merging the selected multiple single-intent utterances into one multi-intent utterance, the at least one processor is configured to: The selected multiple single-intent utterances are merged into the one multi-intent utterance by eliminating redundant expressions in the multiple sentences and replacing the redundant expressions with pronouns.