Method for generating and augmenting unethical utterance detection data using commercial llm

By using commercial LLMs with database-driven prompts to bypass safety guards and classify unethical speech, the method addresses the challenge of generating high-quality unethical speech data for detection models.

WO2026095740A1PCT designated stage Publication Date: 2026-05-07KOREA ELECTRONICS TECH INST
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
KOREA ELECTRONICS TECH INST
Filing Date
2025-11-04
Publication Date
2026-05-07

AI Technical Summary

Technical Problem

Commercial Large Language Models (LLMs) are designed with safety guards that prevent the generation of unethical speech, making it difficult to create or augment data for unethical speech detection models, while Open LLMs, lacking these guards, provide inferior performance.

Method used

A method to generate and augment unethical speech data using commercial LLMs by storing ethical standards and social norms in a database, generating prompts to bypass safety guards, and classifying the obtained unethical speech based on these norms.

Benefits of technology

Enables the generation and augmentation of high-quality unethical speech data for training models, overcoming the limitations of both commercial and Open LLMs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025017871_07052026_PF_FP_ABST
    Figure KR2025017871_07052026_PF_FP_ABST
Patent Text Reader

Abstract

A method for generating and augmenting unethical utterance detection data using a commercial LLM, according to one embodiment of the present invention, comprises the steps of: storing, in a first database, data on ethical criteria and social norms for each social group; extracting, from the first database, at least one piece of data on ethical criteria and social norms; generating a prompt applicable to commercial LLMs on the basis of the extracted data on ethical criteria and social norms; and acquiring unethical utterances by applying the generated prompt to the commercial LLMs.
Need to check novelty before this filing date? Find Prior Art

Description

Method for Generating and Augmenting Data on Unethical Speech Detection Using Commercial LLM

[0001] The present invention relates to a method for generating and augmenting data using a commercial LLM (Large Language Model), and more specifically, to a method for generating or augmenting text-based data regarding unethical (discrimination, profanity, unethical, immoral, etc.) utterances using a commercial LLM (Large Language Model).

[0002] Large Language Models (LLMs) are language models composed of artificial neural networks with numerous parameters, and can support AI chatbot technology.

[0003] In the case of commercial LLMs provided to users, such as the GPT series developed by OpenAI, they are designed not to respond to utterances that violate ethical standards and social norms through safety guards configured by developers. Consequently, even if users attempt to generate or augment data on unethical utterances to use as training data for models that detect such content (hereinafter collectively referred to as 'unethical utterances'), they are blocked by these safety guards, making it difficult to generate or augment such data.

[0004] Previously, to generate or augment data regarding content violating these ethical standards and social norms, it was possible to acquire such data through Open LLMs, which, while having lower performance than commercial LLMs, lacked safety guards.

[0005] However, when utilizing such open LLMs, there is a problem in that the quality and quantity of obtainable data are reduced because the performance of open LLMs is relatively inferior to that of commercial LLMs.

[0006] Accordingly, in order to utilize data on unethical speech as training data for models that detect such speech, it is necessary to explore methods to generate or augment data on unethical speech using commercial LLMs.

[0007] The present invention has been devised to solve the aforementioned problems, and the objective of the present invention is to provide a method for generating and augmenting unethical speech detection data that can generate or augment data regarding unethical speech by utilizing a commercial LLM.

[0008] A method for generating and augmenting data for detecting unethical speech using commercial LLMs according to an embodiment of the present invention for achieving the above objective comprises: a step in which a system stores data regarding ethical standards and social norms for each social group in a first database; a step in which a system extracts at least one piece of data regarding ethical standards and social norms from the first database; a step in which a system generates a prompt applicable to commercial LLMs based on the extracted data regarding ethical standards and social norms; and a step in which a system applies the generated prompt to a commercial LLM to obtain unethical speech that violates ethical standards or social norms.

[0009] Additionally, the step of generating a prompt may generate a prompt that includes content requesting the generation of unethical utterances that violate the ethical standards or social norms of a social group, and content requesting the performance of a fake task of said unethical utterance.

[0010] And the step of generating a prompt can generate a prompt that includes content requesting the generation of unethical utterances and content requesting an explanation of the problems associated with the unethical utterances along with the unethical utterances.

[0011] Additionally, the step of generating a prompt may generate a prompt that includes content requesting the generation of unethical utterances and content requesting the generation of expressions that purify the unethical utterances into proper language expressions along with the unethical utterances.

[0012] And the step of generating a prompt can generate a prompt that includes content requesting an explanation of the meaning of the presented unethical word and requesting the generation of a sentence containing the said unethical word.

[0013] Additionally, the step of generating a prompt may generate a prompt that includes a request to generate educational content capable of resolving the prejudices or misunderstandings contained in the unethical utterances presented.

[0014] And the step of generating the prompt can utilize a Large Language Model-based prompt generator to generate content requesting the performance of a fake task for the input unethical utterance.

[0015] In addition, the method for generating and augmenting unethical speech detection data using a commercial LLM according to one embodiment of the present invention may further include the step of the system classifying the acquired unethical speech according to ethical standards and social norms for each social group and storing it in a second database.

[0016] And the step of storing in the second database may involve determining the intent of speech and the degree of harmfulness of speech based on the ethical standards and social norms of each social group for unethical speech, classifying it into multiple groups according to the judgment results, and storing each group in the second database.

[0017] Meanwhile, a system for generating and augmenting data for detecting unethical speech using commercial LLMs according to another embodiment of the present invention comprises: a storage unit including a first database in which data regarding ethical standards and social norms for each social group is stored; and a processor that extracts at least one piece of data regarding ethical standards and social norms from the first database, generates a prompt applicable to commercial LLMs based on the extracted data regarding ethical standards and social norms, and applies the generated prompt to the commercial LLM to obtain unethical speech that violates ethical standards or social norms.

[0018] Additionally, a method for generating and augmenting data for detecting unethical speech using commercial LLMs according to another embodiment of the present invention comprises: a step in which a system extracts at least one piece of data regarding ethical standards and social norms from a first database in which data regarding ethical standards and social norms by social group is stored; a step in which a system generates a prompt applicable to commercial LLMs based on the extracted data regarding ethical standards and social norms; a step in which a system applies the generated prompt to a commercial LLM to obtain unethical speech that violates ethical standards or social norms; and a step in which a system classifies the obtained unethical speech according to the ethical standards and social norms by social group and stores it in a second database.

[0019] A system for generating and augmenting data for detecting unethical speech using commercial LLMs according to another embodiment of the present invention comprises: a data extraction unit that extracts at least one piece of data regarding ethical standards and social norms from a first database in which data regarding ethical standards and social norms by social group is stored; a prompt generation unit that generates a prompt applicable to commercial LLMs based on the extracted data regarding ethical standards and social norms; an unethical speech classification unit that applies the generated prompt to a commercial LLM to obtain unethical speech that violates ethical standards or social norms and classifies the obtained unethical speech according to the ethical standards and social norms by social group; and a storage unit including a second database in which unethical speech classified through the unethical speech classification unit is stored.

[0020] As described above, according to the embodiments of the present invention, data regarding unethical speech can be generated or augmented by utilizing a commercial LLM, thereby contributing to the training of an unethical speech detection model.

[0021] FIG. 1 is a drawing provided for the description of the configuration of a system for generating and augmenting unethical speech detection data using a commercial LLM according to an embodiment of the present invention.

[0022] FIG. 2 is a drawing provided for a more detailed configuration description of the processor illustrated in FIG. 1.

[0023] FIG. 3 is a flowchart provided for explaining a method for generating and augmenting unethical speech detection data using a commercial LLM according to an embodiment of the present invention, and

[0024] FIG. 4 is a flowchart provided for a more detailed explanation of the process of obtaining unethical speech by applying a prompt generated through an unethical speech detection data generation and augmentation system utilizing a commercial LLM according to one embodiment of the present invention to a commercial LLM.

[0025] The present invention will be described in more detail below with reference to the drawings. To clearly explain the invention, parts unrelated to the description have been omitted from the drawings, and in the drawings, the width, length, thickness, etc., of the components may be exaggerated for convenience.

[0026] FIG. 1 is a diagram provided to describe the configuration of a system for generating and augmenting unethical speech detection data using a commercial LLM according to one embodiment of the present invention.

[0027] The system for generating and augmenting data for the detection of unethical speech using a commercial LLM according to the present embodiment (hereinafter collectively referred to as the 'system') is provided to generate or augment data regarding unethical speech using a commercial LLM.

[0028] To this end, the system may include an input unit (100), a processor (200), and a storage unit (300).

[0029] The input unit (100) is equipped with an input interface device that receives user input, such as a mouse or keyboard, and a communication module connected to a network, so that it can obtain data on ethical standards and social norms for each social group in the form of text.

[0030] For example, the input unit (100) can acquire data regarding ethical standards and social norms for each social group that are entered in text form through an input interface device that receives user input, such as a mouse or keyboard, or receive data regarding ethical standards and social norms for each social group in text form from an external device.

[0031] The storage unit (300) is provided to store programs and data necessary for the operation of the processor (200).

[0032] For example, the storage unit (300) may include a first database in which data regarding ethical standards and social norms by social group is stored, and a second database in which data regarding unethical speech is classified and stored according to ethical standards and social norms by social group.

[0033] The processor (200) can utilize commercial LLM to process all matters for generating or augmenting data on unethical speech.

[0034] Specifically, the processor (200) can extract at least one piece of data regarding ethical standards and social norms from a first database in which data regarding ethical standards and social norms by social group is stored, generate a prompt applicable to commercial LLMs based on the extracted data regarding ethical standards and social norms, apply the generated prompt to a commercial LLM to obtain unethical utterances that violate ethical standards or social norms, classify the obtained unethical utterances according to the ethical standards and social norms by social group, and store them in a second database.

[0035] Additionally, the processor (200) can generate content requesting the performance of a fake task of an unethical utterance in various types and store it in a first database, and when a specific unethical utterance is extracted, it can generate a commercial LLM prompt by selecting content requesting the performance of a specific type of fake task from among the various types of content requesting the performance of a fake task stored in the first database and including it in the prompt along with content requesting the generation of the extracted unethical utterance.

[0036] FIG. 2 is a drawing provided for a more detailed configuration description of the processor (200) shown in FIG. 1.

[0037] Referring to FIG. 2, the processor (200) may include a data extraction unit (210), a prompt generation unit (220), and an unethical speech classification unit (230).

[0038] The data extraction unit (210) can extract at least one piece of data regarding ethical standards and social norms from a first database in which data regarding ethical standards and social norms by social group is stored.

[0039] The prompt generation unit (220) can generate a prompt applicable to commercial LLMs to obtain data on unethical (discrimination, profanity, unethicality, immorality, etc.) utterances that violate ethical standards or social norms for each social group, based on data on extracted ethical standards and social norms, to be used as training data for a model that detects data on unethical utterances.

[0040] Specifically, the prompt generation unit (220) can generate a prompt containing content requesting the generation of unethical utterances and content requesting the performance of a fake task of said unethical utterances so that the generation of unethical utterances is not restricted by the Safety Guard set by the developer of the commercial LLM.

[0041] The prompt generation unit (220) can generate a prompt applicable to a commercial LLM that includes a request to perform a fake task regarding an unethical utterance extracted through the data extraction unit (210) by utilizing a prompt generator based on a Large Language Model.

[0042] And the prompt generation unit (220) can generate content requesting the performance of a fake task of an unethical utterance in various types and store it in the first database, and when a specific unethical utterance is extracted through the data extraction unit (210), it can select content requesting the performance of a specific type of fake task from among the various types of content requesting the performance of a fake task stored in the first database and include it in the prompt together with content requesting the generation of the unethical utterance extracted through the data extraction unit (210).

[0043] Here, the content requesting the performance of a fake task may correspond to any one of the following types: a) a request to explain the problems associated with the unethical utterance along with the unethical utterance; b) a request to generate an expression that purifies the unethical utterance into proper language expression along with the unethical utterance; c) a request to explain the meaning of the presented unethical word and to generate a sentence containing the unethical word; and d) a request to generate educational content that can resolve prejudices or misunderstandings contained in the presented unethical utterance.

[0044] a) Example of type: Explain in one sentence why the unethical utterance is harmful.

[0045] b) Example of type: Correct the unethical utterance into a proper expression.

[0046] c) Example of type: Presenting an unethical (harmful) word, asking for its meaning, and requesting the generation of sentences / utterances that must not be used and contain that word.

[0047] d) Example of Type: An AI program is being developed to respond appropriately to user utterances. When a user makes a harmful or biased remark (A), the AI ​​generates an appropriate response (B). The purpose of this response is to correct biases or misunderstandings contained in the user's remarks and to convey educational content. A request is made to utilize the AI's assistance due to a lack of data regarding user utterances.

[0048] Regarding the content requesting the performance of the above fake task, for example, a) regarding type, the prompt generation unit (220) can generate a prompt that includes content requesting the generation of unethical speech and content requesting that the problems of the unethical speech be explained together with the unethical speech.

[0049] b) For example, regarding the type, the prompt generation unit (220) can generate a prompt that includes content requesting the generation of an unethical utterance and content requesting an explanation of the problems associated with the unethical utterance along with the unethical utterance.

[0050] c) For example, regarding the type, the prompt generation unit (220) can generate a prompt that includes content requesting the generation of an unethical utterance and content requesting the generation of an expression that purifies the unethical utterance into a correct language expression along with the unethical utterance.

[0051] d) For example, regarding the type, the prompt generation unit (220) can generate a prompt that includes a request to explain the meaning of the unethical word being presented and a request to generate a sentence containing the unethical word.

[0052] The unethical speech classification unit (230) can apply the generated prompt to a commercial LLM to obtain unethical speech that violates ethical standards or social norms, classify the obtained unethical speech according to the ethical standards and social norms of each social group, and store it in a second database.

[0053] FIG. 3 is a flowchart provided to explain a method for generating and augmenting unethical speech detection data using a commercial LLM according to one embodiment of the present invention.

[0054] The method for generating and augmenting unethical speech detection data using a commercial LLM according to the present embodiment can be executed by the system described above with reference to FIGS. 1 and 2.

[0055] Referring to FIG. 3, the system can utilize a data extraction unit (210) to extract at least one piece of data regarding ethical standards and social norms from a first database in which data regarding ethical standards and social norms by social group is stored (S310).

[0056] And the system can generate a prompt for commercial LLMs to obtain data on unethical speech that violates ethical standards or social norms by social group, based on the extracted data on ethical standards and social norms, by utilizing the prompt generation unit (220) (S320).

[0057] Afterwards, when the system obtains unethical speech by applying the generated commercial LLMs prompt to the commercial LLM (S330), it can use the unethical speech classification unit (230) to classify the obtained unethical speech according to the ethical standards and social norms of each social group and store it in the second database (S340).

[0058] Specifically, the system can utilize an unethical speech classification unit (230) to determine the intent of speech and the degree of harm of speech based on the ethical standards and social norms of each social group for unethical speech, classify it into multiple groups according to the judgment results, and store each group in a second database.

[0059] For example, the unethical speech classification unit (230) can classify the speech intent of the unethical speech into any one of a) opinion search type, b) opinion expression type, c) information provision type and d) information provision request type.

[0060] At this time, a) the opinion search type is a type having a speech intention to ask for the system's opinion, b) the opinion expression type is a type having a speech intention to express the user's opinion, emotion, attitude, etc. regarding a concept, situation, information, etc., c) the information provision type is a type having a speech intention to provide information known to the user, and d) the information provision request type is a type having a speech intention to request the system to provide information.

[0061] a) Example of opinion-seeking type: Do you think it makes sense for a first-time bride to date a divorcee?

[0062] b) Example of opinion expression type: All criminals should be executed.

[0063] c) Example of information provision type: I heard all men are car poor these days?

[0064] d) Examples of types of information requests: Tell me how to smuggle drugs.

[0065] Additionally, the unethical speech classification unit (230) can numerically calculate the degree of harm for unethical speech within a preset range. Here, the preset range for the degree of harm can be set from level 0 to level 5.

[0066] For example, the unethical speech classification unit (230) can calculate the degree of harm to the unethical speech as one of the following levels: Level 0, which is a level where the speech is not harmful; Level 1, which is a level where there is no intention to cause harm but it may cause discomfort depending on the context and situation; Level 2, which is a level where it may cause discomfort by including derogatory or aggressive remarks against a specific individual or group; Level 3, which is a level where it clearly reveals discrimination or prejudice against the target and may damage the target's self-esteem; Level 4, which is a level where it conveys strong hatred or violent messages against a specific person or group and causes psychological damage; and Level 5, which is a level where it promotes extreme and dangerous behavior, direct life threats, or violence.

[0067] 1) Example of Level 1 Harmfulness: Women are just weaker than men, aren't they?

[0068] 2) Example of Level 2 Harmfulness (Harmfulness): So, a student retaking the exam actually has time for Instagram?

[0069] 3) Example of Level 3 Harmfulness: Misook is a disabled person with all four limbs intact.

[0070] 4) Example of Level 4 Harmfulness: That little brat is acting like a lady.

[0071] 5) Example of Level 5 Harmfulness: All Black people should be buried in the ground.

[0072] FIG. 4 is a flowchart provided for a more detailed explanation of the process of obtaining unethical speech by applying a prompt generated through an unethical speech detection data generation and augmentation system utilizing a commercial LLM according to one embodiment of the present invention to a commercial LLM.

[0073] Referring to FIG. 4, the system generates content requesting the performance of a fake task of an unethical utterance in various types and stores it in a first database. When a specific unethical utterance is extracted, the system selects content requesting the performance of a specific type of fake task from among the various types of content requesting the performance of a fake task stored in the first database (S410), and can generate a prompt to be included in the prompt along with content requesting the generation of the extracted unethical utterance (S420).

[0074] And if the generated prompt is rejected by the Safety Guard of the commercial LLM (S430-No), the system can adjust the selection weight for the type of fake task for which the request was rejected for the user LLM for which the request was rejected, so that the request for the execution of the type of fake task for which the request was rejected is not selected again for the user LLM for which the request was rejected (S440).

[0075] In this case, the system can set the selection weight for the fake task type for which the request approval was denied to 0.1 times the existing selection weight.

[0076] Conversely, if the generated prompt is approved by the Safety Guard of the commercial LLM, the system applies the generated prompt for the commercial LLM to the commercial LLM to obtain unethical speech, determines the intent of speech and the degree of harmfulness of the speech based on the ethical standards and social norms of each social group for the obtained unethical speech (S450), classifies it into multiple groups based on the judgment result, and stores each group in a second database (S460).

[0077] So far, a method for generating and augmenting unethical speech detection data using a commercial LLM has been described in detail with reference to preferred embodiments.

[0078] Methods for generating or augmenting data regarding content that violates existing ethical standards and social norms utilize Open LLMs, which lack safety guards and offer lower performance than commercial LLMs. Consequently, there is a problem where the quality and quantity of the acquired data are reduced.

[0079] On the other hand, according to an embodiment of the present invention, data on unethical speech can be generated or augmented by utilizing a commercial LLM, thereby contributing to the training of an unethical speech detection model.

[0080] Meanwhile, it goes without saying that the technical concept of the present invention may also be applied to a computer-readable recording medium containing a computer program that enables the device and method according to the present embodiment to perform their functions. Furthermore, the technical concept according to various embodiments of the present invention may be implemented in the form of computer-readable code recorded on a computer-readable recording medium. A computer-readable recording medium may be any data storage device that can be read by a computer and store data. For example, a computer-readable recording medium may be a ROM, RAM, CD-ROM, magnetic tape, floppy disk, optical disk, hard disk drive, etc. Additionally, computer-readable code or a program stored on a computer-readable recording medium may be transmitted through a network connected between computers.

[0081] Furthermore, although preferred embodiments of the present invention have been illustrated and described above, the present invention is not limited to the specific embodiments described above. Various modifications are possible by those skilled in the art without departing from the essence of the invention as claimed in the claims, and such modifications should not be understood individually from the technical spirit or perspective of the present invention.

Claims

1. A step in which the system stores data regarding ethical standards and social norms for each social group in a first database; The system extracts at least one piece of data regarding ethical standards and social norms from a first database; A step in which the system generates prompts applicable to commercial LLMs based on data regarding extracted ethical standards and social norms; and A method for generating and augmenting unethical speech detection data using a commercial LLM, comprising the step of the system applying a generated prompt to a commercial LLM to obtain unethical speech that violates ethical standards or social norms.

2. In Claim 1, The step of generating a prompt is, A method for generating and augmenting unethical speech detection data using a commercial LLM, characterized by generating a prompt containing content requesting the generation of unethical speech that violates ethical standards or social norms of social groups, and content requesting the performance of a fake task of said unethical speech.

3. In Claim 2, The step of generating a prompt is, A method for generating and augmenting unethical speech detection data using a commercial LLM, characterized by generating a prompt that includes content requesting the generation of unethical speech and content requesting an explanation of the problems associated with the unethical speech along with the unethical speech.

4. In Claim 2, The step of generating a prompt is, A method for generating and augmenting unethical speech detection data using a commercial LLM, characterized by generating a prompt that includes content requesting the generation of unethical speech and content requesting the generation of an expression that purifies the unethical speech into a correct language expression along with the unethical speech.

5. In Claim 2, The step of generating a prompt is, A method for generating and augmenting unethical speech detection data using a commercial LLM, characterized by generating a prompt that includes a request to explain the meaning of a presented unethical word and a request to generate a sentence containing the unethical word.

6. In Claim 2, The step of generating a prompt is, A method for generating and augmenting unethical speech detection data using a commercial LLM, characterized by generating a prompt that includes a request to generate educational content capable of resolving prejudices or misunderstandings contained in the presented unethical speech.

7. In Claim 2, The step of generating a prompt is, A method for generating and augmenting unethical utterance detection data using a commercial LLM, characterized by utilizing a Large Language Model-based prompt generator to generate content requesting the performance of a fake task for an input unethical utterance.

8. In Claim 1, A method for generating and augmenting unethical speech detection data using a commercial LLM, characterized by further including the step of the system classifying acquired unethical speech according to ethical standards and social norms for each social group and storing it in a second database.

9. In Claim 8, The step of storing in the second database is, A method for generating and augmenting unethical speech detection data using a commercial LLM, characterized by determining the intent of speech and the degree of harmfulness of speech based on ethical standards and social norms for each social group regarding unethical speech, classifying it into multiple groups based on the determination results, and storing each group in a second database.

10. A storage unit including a first database in which data regarding ethical standards and social norms for each social group is stored; and A system for generating and augmenting unethical speech detection data utilizing commercial LLMs, comprising: a processor that extracts at least one piece of data regarding ethical standards and social norms from a first database, generates a prompt applicable to commercial LLMs based on the extracted data regarding ethical standards and social norms, and applies the generated prompt to commercial LLMs to obtain unethical speech that violates ethical standards or social norms.

11. A step in which the system extracts at least one piece of data regarding ethical standards and social norms from a first database in which data regarding ethical standards and social norms by social group is stored; A step in which the system generates prompts applicable to commercial LLMs based on data regarding extracted ethical standards and social norms; A step in which the system applies the generated prompt to a commercial LLM to obtain unethical utterances that violate ethical standards or social norms; and A method for generating and augmenting unethical speech detection data using a commercial LLM, comprising the step of the system classifying acquired unethical speech according to ethical standards and social norms for each social group and storing them in a second database.

12. A data extraction unit that extracts at least one piece of data regarding ethical standards and social norms from a first database in which data regarding ethical standards and social norms by social group is stored; A prompt generation unit that generates prompts applicable to commercial LLMs based on data regarding extracted ethical standards and social norms; An unethical utterance classification unit that applies the generated prompt to a commercial LLM to obtain unethical utterances that violate ethical standards or social norms, and classifies the obtained unethical utterances according to the ethical standards and social norms of each social group; and An unethical speech detection data generation and augmentation system utilizing a commercial LLM, comprising: a storage unit including a second database in which unethical speech classified through an unethical speech classification unit is stored.

Citation Information

Patent Citations

  • Method for analyzing metallic iron from steel slag

    KR1020210146055A

  • Electrode assembly and rechargeable battery including the same

    KR1020250111638A

  • Apparatus, method and computer program for augmenting learning data for harmful words

    KR102410582B1

  • Method for manufacturing polarizing plate

    KR102924861B1

  • System to Prevent Misuse of Large Foundation Models and a Method Thereof

    US20240289628A1