Machine learning apparatus and method for large language model on civil service
Patent Information
- Application Number
- KR1020230118915
- Authority / Receiving Office
- KR · KR
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2023-09-07
- Publication Date
- 2026-09-02
- Estimated Expiration
- 2043-09-07
Smart Images

Figure R1020230118915_ABST
Abstract
Description
Technology Field
[0001] The embodiments disclosed in this document relate to an apparatus and method for performing machine learning of a large language model. Background Technology
[0002] Machine learning is the process of optimizing model parameters, and its development trends are shifting from statistical-based models to deep learning, from RNNs (recurrent neural networks) to transfer learning models, and from small transfer learning models to large language model inference.
[0003] As deep learning models have become larger, the performance of unsupervised pattern learning has improved, and as a result, performance capable of commercial use can be achieved even in the field of text generation where existing models did not perform well.
[0004] However, as large language model inference has become the focus of research, the examination of specific domains and tasks has relatively decreased. Since large language models must undergo extensive training to perform tasks for specific domains or tasks at the level of a concrete model, there is a problem in that the training cost can be very high, and even if such high-cost training is performed, it is difficult to guarantee superior performance. The problem to be solved
[0005] When applying large language models to civil complaint handling in public institutions, training methods capable of guaranteeing superior performance may be required, as the nature of these tasks demands a very high level of reliability. Furthermore, to ensure the security of sensitive data, there may be a high need for installable models capable of local execution, allowing for data processing without external network transmission.
[0006] The embodiments of the present invention are intended to provide a method for efficiently performing the training of a language model specialized for civil complaint resolution or consultation by utilizing the text generation capabilities of a large deep learning model while limiting the domain and tasks during the training process. means of solving the problem
[0007] A learning device for a large language model for civil complaint handling according to one embodiment disclosed in this document comprises a communication circuit configured to communicate with the outside, a memory storing a language model and a regression model, and a processor electrically connected to the communication circuit and the memory. The processor obtains a query related to a civil complaint to a public institution and an answer corresponding to the query using the communication circuit, obtains meta-information including at least some of the institution information and regional information associated with the query and answer using the communication circuit, performs unsupervised learning of the language model by inputting an answer to the language model to learn the style of the result text generated by the language model, performs supervised learning of the language model by fine-tuning the language model so that an answer is output when the query and meta-information are input to the language model after the unsupervised learning is completed, performs reinforcement learning of the language model using evaluation scores for the result text generated by the language model and the result text output by the regression model when the query and meta-information are input, and performs machine learning for the regression model using the result text, the result score calculated by the regression model upon input of the result text, and an externally input score for the result text. there is.
[0008] According to one embodiment, the processor may perform data augmentation for a query based on a generative model, back translation, or named entity substitution, and perform preprocessing for the answer to delete repetitive phrases and de-identify text.
[0009] According to one embodiment, the processor inputs an array of tokens corresponding to the result text into a regression model, calculates a result score based on output values calculated by a token order-based model composed of a neural network (NN) included in the regression model and a token frequency-based model composed of a support vector machine (SVM), respectively, calculates a performance indicator for the regression model by comparing the result score with an externally input score, and outputs an evaluation score based on the result score, the externally input score, and the performance indicator.
[0010] According to one embodiment, the processor may set a first weight for the result score to the maximum and a second weight for the externally input score to the minimum when the performance indicator is greater than or equal to the maximum reference value, set the first weight to the minimum and the second weight to the maximum when the performance indicator is less than the minimum reference value, set the first weight to the minimum and the second weight to the maximum when the performance indicator is less than the maximum reference value and greater than or equal to the minimum reference value, set the first weight to be proportional to the performance indicator and the second weight to be inversely proportional to the performance indicator, and set a regression model such that when the case where the performance indicator is greater than or equal to the maximum reference value accumulates above a specified threshold, the evaluation score is output as the same as the result score.
[0011] A method for learning a large language model for civil complaint handling according to one embodiment disclosed in this document may include: a step of obtaining a question related to a civil complaint to a public institution and an answer corresponding to the question; a step of obtaining meta-information including at least some of the institution information and regional information associated with the question and answer; a step of performing unsupervised learning of a language model by inputting an answer into the language model to learn the writing style of the result text generated by the language model; a step of performing supervised learning of a language model to fine-tune the language model so that an answer is output when the question and meta-information are input into the language model after unsupervised learning is completed; a step of performing reinforcement learning of a language model using evaluation scores for the result text generated by the language model when the question and meta-information are input and the result text output by a regression model after supervised learning is completed; and a step of performing machine learning on a regression model using the result text, the result score calculated by the regression model when the result text is input, and an externally input score for the result text. Effects of the invention
[0012] According to the embodiments disclosed in this document, the performance and reliability of the generated responses can be improved by utilizing unsupervised learning, supervised learning, and reinforcement learning together when training a large language model for resolving civil complaints of public institutions.
[0013] In addition, by training a regression model that generates evaluation scores input during reinforcement learning, reinforcement learning can be performed efficiently without score input by an evaluator.
[0014] In addition, various effects that can be identified directly or indirectly through this document may be provided. Brief explanation of the drawing
[0015] FIG. 1 illustrates the operating environment of a large language model for civil complaint handling according to one embodiment. FIG. 2 is a block diagram illustrating the configuration of a machine learning device for a large language model for civil complaint processing according to one embodiment. FIG. 3 is a diagram illustrating exemplary data augmentation and maintenance operations of a machine learning device of a large language model for civil complaint handling according to one embodiment. FIG. 4 is a diagram illustrating the unsupervised learning operation of an exemplary language model of a machine learning device for a large language model for civil complaint work according to one embodiment. FIG. 5 is a diagram illustrating the supervised learning operation of an exemplary language model of a machine learning device for a large language model for civil complaint handling according to one embodiment. FIG. 6 is a diagram illustrating the reinforcement learning operation of an exemplary language model of a machine learning device for a large language model for civil complaint handling according to one embodiment. FIG. 7 is a diagram illustrating the machine learning operation of an exemplary regression model of a machine learning device for a large language model for civil complaint handling according to one embodiment. FIG. 8 is a flowchart illustrating a machine learning method for a large language model for civil complaint handling according to one embodiment. In relation to the description of the drawings, the same or similar reference numerals may be used for identical or similar components. Specific details for implementing the invention
[0016] Hereinafter, some embodiments of the present invention will be described in detail with reference to the exemplary drawings. However, this is not intended to limit the present invention to specific embodiments, and it should be understood that the invention includes various modifications, equivalents, or substitutions of the embodiments. It should be noted that in assigning reference numerals to the components of each drawing, the same components are given the same reference numeral whenever possible, even if they are shown in different drawings. Furthermore, in describing the embodiments of the present invention, if it is determined that a detailed description of related known components or functions would hinder understanding of the embodiments of the present invention, such detailed description is omitted.
[0018] FIG. 1 illustrates the operating environment of a large language model for civil complaint handling according to one embodiment.
[0019] Referring to FIG. 1, a large language model (LM) for civil complaint processing according to one embodiment may be a model for generating a response corresponding to the content of a civil complaint entered by a complainant. The LLM for civil complaint processing may be installed on a model execution server of an institution's intranet.
[0020] Petitioners may enter a query including the relevant department, the title of the complaint, and the content of the complaint through a personal web browser. The query entered by the petitioner may be stored in the database for complaint processing.
[0021] The database for civil complaint processing can store queries and metadata regarding the queries. The scheduler of the model execution server can send queries to the database, and in response to the queries, the queries and metadata can be input into the LLM for civil complaint operations. The LLM for civil complaint operations can output a draft civil complaint response derived from the queries and metadata. The draft civil complaint response can be updated in the database.
[0022] The complaint handling officer can check processed complaints through the complaint processing UI. The officer can review and modify the draft response. Once the complaint is marked as completed by the officer, the response can be updated in the database and displayed in the complainant's personal web browser.
[0023] A machine learning device for a large language model for civil complaint handling according to one embodiment of the present invention is a device for efficiently performing machine learning of the LLM for civil complaint handling described above.
[0025] FIG. 2 is a block diagram illustrating the configuration of a machine learning device for a large language model for civil complaint processing according to one embodiment.
[0026] Referring to FIG. 2, a machine learning device (200) for a large language model for civil complaint work according to one embodiment may be implemented as one of various types of computing devices. The machine learning device (200) for a large language model for civil complaint work according to one embodiment may include a communication circuit (210), a memory (220), and a processor (230).
[0027] The communication circuit (210) may be an interface that communicates wirelessly or via a wire with an external device (e.g., an external storage device or an external computing device). The external storage device may be, for example, a storage device for providing data related to civil complaints to a machine learning device (200). The external computing device may be, for example, a server of a public institution that processes and stores various data related to civil complaints. The communication circuit (210) may transmit and receive data with the external device.
[0028] The memory (220) may include volatile memory and / or non-volatile memory. The memory (220) may store various data handled by the machine learning device (200) of the large language model for civil complaint work, and may store artificial intelligence models such as language models and regression models. The language model may be, for example, a Polyglot-ko based or KoBART based answer generation model.
[0029] The processor (230) may be electrically connected to the communication circuit (210) and the memory (220). The processor (230) may control the communication circuit (210) and the memory (220) and perform various data processing and operations. The processor (230) may perform the following operations by executing software or instructions stored in the memory (220).
[0030] According to one embodiment, the processor (230) can obtain a query related to a civil complaint regarding a public institution and an answer corresponding to the query using a communication circuit (210). The processor (230) can obtain the query and the answer from an external device (e.g., an external storage device or an external computing device) that stores the query text and the answer text registered on the civil complaint bulletin board. The query may be text entered by a civil complainant, and the answer may be text entered by a civil complaint processing officer. The query and the answer corresponding to the query may be distinguished from other queries and other answers so that they can be matched with each other.
[0031] According to one embodiment, the processor (230) can obtain meta-information including at least some of the agency information and regional information associated with the inquiry and answer using a communication circuit (210). Due to the domain characteristics of civil complaint work, where rules applied vary depending on the region, relevant agency, and department, etc., the efficiency and performance of learning can be improved when collecting meta-information regarding the inquiry and answer. The processor (230) can obtain agency information (including department information within the agency) and regional information, etc., collected from the civil complaint bulletin board or the relevant website. The meta-information can be matched with the previously collected inquiry and answer.
[0032] According to one embodiment, the processor (230) may perform data augmentation for a query. For example, the processor (230) may utilize a generative model or perform back-translation or named entity substitution (e.g., personal information).
[0033] According to one embodiment, the processor (230) may perform preprocessing for the deletion of repetitive phrases and de-identification of text on the answer. The processor (230) may remove repetitive phrases, such as greetings, that are unnecessary for learning from the answer, and may remove names, phone numbers, extension numbers, and specific URLs for de-identification.
[0034] According to one embodiment, the processor (230) can perform unsupervised learning of the language model by inputting a response into the language model to learn the writing style of the result text generated by the language model. The processor (230) can input a response to a civil complaint into the language model to perform the learning. The language model can learn the writing style applied to the input response. Since the response is text input by a civil complaint processing officer of a public institution (e.g., a public official), it may consist of a writing style suitable for the text generated by the language model, and the language model can learn the writing style of the civil complaint processing officer by repeating unsupervised learning.
[0035] According to one embodiment, the processor (230) may perform supervised learning of a language model that fine-tunes the language model so that when a query and meta-information are input to the language model after unsupervised learning is completed, an answer is output. The processor (230) may use a query (including original data and augmented data) and meta-information as inputs to the supervised learning, and an answer as output to the supervised learning. The parameters of the language model may be fine-tuned so that when the corresponding query and meta-information are input, the corresponding answer is output (to simulate the probability of the tokens of the corresponding answer being seen). The language model may be trained to output a corresponding answer when a query is input by repeating the supervised learning.
[0036] According to one embodiment, the processor (230) can perform reinforcement learning of the language model using evaluation scores for result text generated by the language model and result text output by the regression model when a query and meta-information are input after supervised learning is completed. In this reinforcement learning, the query and meta-information may be a state, the result text may be an action, and the evaluation score may be a reward. The processor (230) can repeat reinforcement learning to improve the evaluation score for the result text generated by the language model.
[0037] According to one embodiment, the processor (230) can perform machine learning on the regression model using the result text, the result score calculated by the regression model upon input of the result text, and the externally input score for the result text. The processor (230) can compare the score calculated by the regression model with the score input by the evaluator for the same result text, and perform learning on the regression model so that the score calculated by the regression model becomes closer to the score input by the evaluator.
[0038] As a specific method for performing training of the regression model, the processor (230) may input an array of tokens corresponding to the result text into the regression model. The processor (230) may calculate a result score based on the output values produced by the token order-based model, which consists of a neural network (NN) included in the regression model, and the token frequency-based model, which consists of a support vector machine (SVM), respectively. The processor (230) may calculate a performance indicator for the regression model by comparing the result score with the externally input score. The processor (230) may output an evaluation score based on the result score, the externally input score, and the performance indicator.
[0039] In a specific method for calculating an evaluation score, if the performance indicator is greater than or equal to a maximum threshold value, the processor (230) may set the first weight for the result score to the maximum and the second weight for the externally input score to the minimum. If the performance indicator is less than the minimum threshold value, the processor (230) may set the first weight to the minimum and the second weight to the maximum. If the performance indicator is less than the maximum threshold value and greater than or equal to the minimum threshold value, the processor (230) may set the first weight to be proportional to the performance indicator and the second weight to be inversely proportional to the performance indicator. The processor (230) may set a regression model to output an evaluation score equal to the result score when the number of cases where the performance indicator is greater than or equal to the maximum threshold value accumulates above a specified threshold.
[0041] FIG. 3 is a diagram illustrating exemplary data augmentation and maintenance operations of a machine learning device of a large language model for civil complaint handling according to one embodiment.
[0042] Referring to FIG. 3, the machine learning device may obtain training data (310) from an external device (30) (e.g., an external storage device or an external computing device). The training data (310) may include a petitioner's inquiry (311) and a person in charge's answer (312), and may include information about an institution (313) and a region (314) as meta-information for the inquiry (311) and the answer (312).
[0043] The machine learning device can perform data augmentation on the query (311) in particular among the training data (310). For example, the machine learning device can input the query (311) into a generative model, perform back-translation on the query (311), or substitute entity names (e.g., names, addresses, and phone numbers, etc.) included in the query (311). The generative model may be a model capable of performing a paraphrasing operation separate from the device of the invention, where the term “paraphrasing” may mean changing the expression while maintaining the meaning of the text. Since the query (311) is unstructured text freely entered by a petitioner, data augmentation can enrich the training data (310) and enable the language model to respond effectively to the unstructured data.
[0044] The machine learning device may perform data cleaning and de-identification on the training data (310), particularly on the answers (312). Data cleaning may be performed on the text included in the answers (312), limited to the generation of actual responses to the civil complaint. Additionally, repetitive phrases that are not helpful for machine learning, such as opening greetings, summaries of the questions (311), and closing greetings, may be deleted. Furthermore, the text may be de-identified by deleting names, phone numbers, extension numbers, and specific URLs. Only minimal processing may be performed on the actual answer content, excluding the described parts.
[0046] FIG. 4 is a diagram illustrating the unsupervised learning operation of an exemplary language model of a machine learning device for a large language model for civil complaint work according to one embodiment.
[0047] Referring to FIG. 4, a machine learning device according to one embodiment can perform unsupervised learning by inputting answers (412) from training data (410) into a language model (420). The training data (410) includes data prior to data augmentation (questions and responses of complaints actually received) and may further include data after data augmentation and preprocessing (data cleaning and de-identification) have been completed.
[0048] In unsupervised learning for the language model (420), the language model (420) can learn the patterns of text that the model is to generate. The language model (420) can learn the style of the answer (412) by learning the tendency of tokens to appear in an array of tokens corresponding to the answer (412) (e.g., conditional probability). The language model (420) can output text in a format similar to that of a civil complaint processing officer at a public institution by repeatedly learning the tendency of the token array.
[0050] FIG. 5 is a diagram illustrating the supervised learning operation of an exemplary language model of a machine learning device for a large language model for civil complaint handling according to one embodiment.
[0051] Referring to FIG. 5, a machine learning device according to one embodiment can perform supervised learning by providing training data (410) containing information about a query (411), an answer (412), an organization (413), and a region (414) to a language model (520). The training data (410) may include plain text and a separator that distinguishes between input and output. Here, the input may include the query (411), the organization (413), and the region (414), and the output may include the answer (412). The input and output may be provided to the language model (520) as a pair.
[0052] The parameters of the language model (520) can be tuned so that when an input is provided, the generated result (530) mimics the probability that the tokens of the output (answer (412)) are seen in order to minimize the loss value at the token level. The language model (520) can be repeatedly supervised learning with various inputs and outputs, and as the learning is repeated, the generated result (530) can become closer to the answer (412).
[0054] FIG. 6 is a diagram illustrating the reinforcement learning operation of an exemplary language model of a machine learning device for a large language model for civil complaint handling according to one embodiment.
[0055] Referring to FIG. 6, a machine learning device according to one embodiment can update parameters in a proximal policy optimization (PPO) manner by combining the generated result (630) by a language model (620) and the evaluation score (650) therefrom. For example, the machine learning device can provide information about a query (411), an organization (413), and a region (414) among the training data to the language model (620). The language model (620) can provide a generated result (630) consisting of text based on the query (411), the organization (413), and the region (414). A regression model (640) can output an evaluation score (650) when the generated result (630) is input. The parameters of the language model (620) can be continuously updated so that the evaluation score (650) of the generated result (630) increases, thereby enabling reinforcement learning of the language model (620).
[0057] FIG. 7 is a diagram illustrating the machine learning operation of an exemplary regression model of a machine learning device for a large language model for civil complaint handling according to one embodiment.
[0058] Referring to FIG. 7, a machine learning device according to one embodiment has a performance indicator Regression model to increase this It can be trained. Language model is a token array If input is given, a token array as the generation result. Can output. Regression model The training data is a language model It can accumulate depending on the training process.
[0059] Generation result Regarding this, the external evaluator scores the external input score according to predefined evaluation criteria. regression model It can be entered into. Meanwhile, regression model The generated result If this is entered, the result score Can calculate. Result score It can be calculated as an ensemble (e.g., weighted average) of results (scores) derived from each of the neural network (NN)-based model and the support vector machine (SVM)-based model.
[0060] Regression model It is an external input score in a manner similar to MSE (mean squared error), etc. Result score based on A performance metric indicating the accuracy of It can calculate performance metrics It can be adjusted to become a rational number indicator between 0 and 1.
[0061] Evaluation score External input score when calculating and result score Weights applied to is a performance metric It can be determined by. Performance indicators This maximum threshold If greater than or equal to, weight It can be 1. Performance indicator This minimum threshold If less than, weight It can be 0. Performance indicator This minimum threshold Above and maximum threshold If less than, weight Is It could be.
[0062] Evaluation score is a mathematical formula It can be calculated as. That is, performance indicators The higher this is, the more the regression model Result score calculated by The weight for increases, and performance metrics The lower this is, the higher the external input score entered by the external evaluator The weight for may increase. Performance metrics This maximum threshold If above, evaluation score is a regression model Result score calculated by It can be, performance metrics This minimum threshold If less than, evaluation score is external input score It could be.
[0063] Evaluation score is a language model It can be provided as, and language model is evaluation score It can be learned to increase.
[0065] FIG. 8 is a flowchart illustrating a machine learning method for a large language model for civil complaint handling according to one embodiment.
[0066] In the following, it is assumed that the machine learning device of FIG. 2 performs the process of FIG. 8. Also, in the description of FIG. 8, the operation described as being performed by the machine learning device can be understood as being controlled by the processor (230).
[0067] Referring to FIG. 8, in step 810, the machine learning device can obtain a question related to a complaint to a public institution and an answer corresponding to the question.
[0068] In step 820, the machine learning device can obtain meta-information including at least some of the organization information and region information associated with the query and answer.
[0069] In step 830, the machine learning device can perform unsupervised learning of the language model by inputting an answer to the language model to learn the style of the resulting text generated by the language model.
[0070] In step 840, the machine learning device can perform supervised learning of a language model that fine-tunes the language model so that when a query and meta-information are input into the language model after unsupervised learning is completed, an answer is output.
[0071] In step 850, the machine learning device can perform reinforcement learning of the language model using evaluation scores for the result text generated by the language model and the result text output by the regression model when the query and meta-information are input after supervised learning is completed.
[0072] In step 860, the machine learning device can perform machine learning on the regression model using the result text, the result score calculated by the regression model upon input of the result text, and the externally input score for the result text.
[0074] The embodiments of this document and the terms used therein are not intended to limit the technology described in this document to specific embodiments and should be understood to include various modifications, equivalents, and / or substitutions of said embodiments. In relation to the description of the drawings, similar reference numerals may be used for similar components. A singular expression may include a plural expression unless the context clearly indicates otherwise. In this document, expressions such as "A or B," "at least one of A and / or B," "A, B or C," or "at least one of A, B and / or C" may include all possible combinations of items listed together. Expressions such as "first," "second," "first," or "second" may modify said components regardless of order or importance and are used only to distinguish one component from another and do not limit said components. When it is mentioned that a component is "(functionally or telecommunicationally) connected" or "connected" to another component, said component may be directly connected to said other component or connected through said other component.
[0075] In this document, "adapted to or configured to" may be used interchangeably with, depending on the context, for example, hardware- or software-wise, "suitable for," "capable of," "modified to," "made to," "capable of," or "designed to." In some cases, the expression "device configured to" may mean that the device is "capable of" in conjunction with other devices or components. For example, the phrase "processor configured to perform A, B, and C" may mean a dedicated processor for performing those operations (e.g., an embedded processor) or a general-purpose processor (e.g., a CPU) capable of performing those operations by executing one or more programs stored in a memory device.
[0076] As used in this document, the term “module” includes a unit composed of hardware, software, or firmware, and may be used interchangeably with terms such as logic, logic block, component, or circuit, for example. A “module” may be a component formed as a whole or a minimum unit or part thereof that performs one or more functions. A “module” may be implemented mechanically or electronically and may include, for example, an application-specific integrated circuit (ASIC) chip, field-programmable gate arrays (FPGAs), or programmable logic device, known or under development, that performs certain operations.
[0077] At least a portion of a device (e.g., modules or functions thereof) or a method (e.g., operations) according to one embodiment may be implemented as instructions stored in a computer-readable storage medium in the form of program modules. When said instructions are executed by a processor, the processor may perform a function corresponding to said instructions.
[0078] Each component (e.g., module or program module) according to one embodiment may be composed of a single or multiple entities, and some of the aforementioned sub-components may be omitted or additional sub-components may be included. Generally or additionally, some components (e.g., module or program module) may be integrated into a single entity to perform the functions performed by each of the respective components prior to integration in the same or similar manner. The operations performed by the module, program module, or other components according to one embodiment may be executed sequentially, in parallel, iteratively, or heuristically, or at least some operations may be executed in a different order, omitted, or other operations added.
Claims
Claim 1 A learning device for a large language model for civil complaint handling, comprising: a communication circuit configured to communicate with the outside; a memory in which a language model and a regression model are stored; The system includes a processor electrically connected to the communication circuit and the memory, wherein the processor obtains a query related to a civil complaint regarding a public institution and an answer corresponding to the query using the communication circuit, obtains meta-information including at least some of the institution information and regional information associated with the query and the answer using the communication circuit, performs unsupervised learning of the language model by inputting the answer into the language model to learn the style of the result text generated by the language model, performs supervised learning of the language model by fine-tuning the language model so that when the query and the meta-information are input into the language model, the answer is output, after the unsupervised learning is completed, performs reinforcement learning of the language model using the result text generated by the language model when the query and the meta-information are input and the evaluation score for the result text output by the regression model, and performs machine learning on the regression model using the result text, the result score calculated by the regression model upon input of the result text, and the external input score for the result text input by an external evaluator. A device characterized by performing, wherein the processor inputs an array of tokens corresponding to the result text into the regression model, and calculates the result score based on output values calculated by each of a token order-based model composed of a neural network (NN) included in the regression model and a token frequency-based model composed of a support vector machine (SVM). Claim 2 An apparatus according to claim 1, wherein the processor performs data augmentation for the query based on a generative model, back translation, or named entity substitution, and performs preprocessing for the deletion of repetitive phrases and de-identification of text for the answer. Claim 3 An apparatus according to claim 1, wherein the processor calculates a performance indicator for the regression model by comparing the result score and the external input score, and outputs an evaluation score based on the result score, the external input score, and the performance indicator. Claim 4 In claim 3, the device is characterized in that the processor sets the first weight for the result score to the maximum and the second weight for the external input score to the minimum when the performance indicator is greater than or equal to the maximum reference value, sets the first weight to the minimum and the second weight to the maximum when the performance indicator is less than the minimum reference value, sets the first weight to the minimum and the second weight to the maximum when the performance indicator is less than the maximum reference value and greater than or equal to the minimum reference value, and sets the first weight to be proportional to the performance indicator and the second weight to be inversely proportional to the performance indicator, and sets the regression model to output the evaluation score as equal to the result score when the case where the performance indicator is greater than or equal to the maximum reference value accumulates above a specified reference value. Claim 5 A method for training a large language model for civil complaint handling, comprising: a step of obtaining a query related to a civil complaint concerning a public institution and an answer corresponding to said query; a step of obtaining meta-information including at least some of the institution information and regional information associated with said query and said answer; a step of performing unsupervised learning of said language model by inputting said answer into said language model to learn the style of the result text generated by said language model; a step of performing supervised learning of said language model by fine-tuning said language model so that when said query and said meta-information are input into said language model after said unsupervised learning is completed, said answer is output; and a step of performing reinforcement learning of said language model using an evaluation score for the result text generated by said language model when said query and said meta-information are input and said result text output by a regression model after said supervised learning is completed. A method comprising the step of performing machine learning on a regression model using the result text, a result score calculated by the regression model upon input of the result text, and an external input score for the result text input by an external evaluator, wherein the step of performing machine learning on the regression model comprises the step of inputting an array of tokens corresponding to the result text into the regression model, and the step of calculating the result score based on output values calculated by each of a token order-based model composed of a neural network (NN) and a token frequency-based model composed of a support vector machine (SVM) included in the regression model.
Citation Information
Patent Citations
Big data management system for public complaints services
KR1020160075971A