Method for providing ai security governance for external ai service and system therefor

KR103021824B1Active Publication Date: 2026-09-21C A S LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
KR1020260064998
Authority / Receiving Office
KR · KR
Patent Type
Patents
Current Assignee / Owner
Filing Date
2026-04-10
Publication Date
2026-09-21
Estimated Expiration
2046-04-10

Smart Images

  • Figure 112026043896675-PAT00005_ABST
    Figure 112026043896675-PAT00005_ABST
Patent Text Reader

Abstract

A method for providing AI security governance for an external AI (Artificial Intelligence) service and a system for providing such governance are provided. A method for providing AI security governance according to some embodiments may include the steps of obtaining a user prompt for an external AI service; deriving an analysis result related to security risk by analyzing the contextual meaning of the prompt through a language model; calculating the risk level of the prompt based on the analysis result; determining a measure based on the risk level based on a predefined security policy; generating a security prompt by performing a measure on the prompt; and transmitting the security prompt to the external AI service. According to this method, an AI security governance framework can be effectively established in an environment utilizing an external AI service.
Need to check novelty before this filing date? Find Prior Art

Description

Technology Field

[0001] The present disclosure relates to the field of information security and Artificial Intelligence (AI) security technology, and more specifically, to technology for providing AI security governance in an environment utilizing external AI services. Background Technology

[0002] Recently, as the business application of Generative AI (Artificial Intelligence) services expands, particularly among public institutions, financial institutions, and general companies, it is becoming common practice for internal users to compose prompts, input them into external Generative AI services, and utilize the responses in their work.

[0003] However, prompts may contain sensitive information, such as personal and confidential data, and there is a risk of serious security incidents if such information is transmitted externally. Furthermore, due to the nature of generative AI, there is also a possibility that sensitive information could be indirectly exposed externally during the process of handling tasks requested by internal users.

[0004] Furthermore, responses received from external generative AI services may contain malicious code, false information, or content that violates security policies, and security incidents may occur if such responses are provided internally without verification.

[0005] Meanwhile, technologies such as Cross Domain Solutions (CDS) and Data Loss Prevention (DLP) have traditionally been used to prevent information leakage and ensure security. However, since these technologies fail to consider the contextual meaning of data, they have limitations in controlling outbound and inbound data related to externally generated AI services; in particular, they are difficult to effectively prevent indirect information leakage through representation modification or reconstruction. Prior art literature

[0006] Korean Registered Patent No. 10-1859636 (Published May 21, 2018) The problem to be solved

[0007] Various technical problems to be solved through some embodiments of the present disclosure relate to a method for providing AI security governance for external AI (Artificial Intelligence) services and a system for doing so.

[0008] Specifically, the technical problem to be solved through some embodiments of the present disclosure is to provide a method and system capable of effectively controlling security risks that may occur in an environment using external AI services.

[0009] In addition, another technical problem to be solved through some embodiments of the present disclosure is to provide a method and system that can control security risks while simultaneously minimizing the reduction of user convenience regarding external AI services.

[0010] In addition, another technical problem to be solved through some embodiments of the present disclosure is to provide a method and system for effectively establishing an organizational-level AI security governance and auditing framework in an environment utilizing external AI services.

[0011] In addition, another technical problem to be solved through some embodiments of the present disclosure is to provide a method and system capable of establishing a flexible AI security governance and audit framework applicable to various organizational environments.

[0012] The technical problems of the present disclosure are not limited to those mentioned above, and other unmentioned technical problems will be clearly understood by a person skilled in the art of the present disclosure from the description below. means of solving the problem

[0013] A method for providing AI security governance according to some embodiments of the present disclosure for solving the aforementioned technical problem may include, in a method performed by at least one processor, the steps of: obtaining a user prompt for an external AI (Artificial Intelligence) service; deriving an analysis result related to a security risk by analyzing the contextual meaning of the prompt through a language model; calculating the risk level of the prompt based on the analysis result; determining a measure according to the risk level based on a predefined security policy; generating a security prompt by performing the measure on the prompt; and transmitting the security prompt to the external AI service.

[0014] In some embodiments, the prompt is received from the user's terminal located in the internal network, and the at least one processor may be provided in a proxy server positioned between the internal network and the external AI service.

[0015] In some embodiments, the step of deriving the analysis result may include: performing a static check on the prompt; generating an analysis prompt using the result of the static check as context information of the prompt; and providing the analysis prompt to the language model to derive the analysis result.

[0016] In some embodiments, the step of deriving the analysis result may include: obtaining attachment data associated with the prompt; generating an analysis prompt for analyzing the contextual meaning of the prompt; and providing the attachment data together with the analysis prompt to the language model to derive the analysis result.

[0017] In some embodiments, the analysis result is expressed in a structured format including a plurality of items, and the plurality of items may include the risk type and risk assessment indicator of the prompt.

[0018] In some embodiments, the step of calculating the risk of the prompt includes the step of calculating the risk based on a plurality of evaluation factors including the analysis result, and the evaluation factors may further include the user's authority and the type of work requested by the prompt.

[0019] In some embodiments, the step of calculating the risk of the prompt may include: generating an expected response to the prompt through the language model; and calculating the risk based further on the risk analysis result for the expected response.

[0020] In some embodiments, the analysis result includes information regarding the risk type of the prompt, and the step of calculating the risk level of the prompt may include: selecting a weight profile corresponding to the risk type among a plurality of predefined weight profiles—the weight profile includes weights for each of a plurality of evaluation factors, and the evaluation factors include the analysis result—; and calculating the risk level by combining the values ​​of each of the evaluation factors based on the weights.

[0021] In some embodiments, the step of calculating the risk of the prompt includes the step of calculating the risk by synthesizing the values ​​of a plurality of evaluation factors including the analysis result based on weights, and information related to the determined measure is recorded in an audit log, and the AI ​​security governance providing method may further include the step of lowering the weight of one or more evaluation factors related to the false positive when a false positive for the prompt is identified based on the audit log; and the step of raising the weight of one or more evaluation factors related to the false negative when a false negative for the prompt is identified based on the audit log.

[0022] In some embodiments, the step of determining a measure according to the risk level includes: determining a measure to be applied to the prompt as a fully permissible measure when the risk level is less than a first threshold; and determining a measure to be applied to the prompt as a partially permissible measure when the risk level is greater than or equal to the first threshold and less than a second threshold, wherein the second threshold is a value greater than the first threshold, and the partially permissible measure may include processing to de-identify a part of the prompt.

[0023] In some embodiments, the step of determining the action according to the risk level further includes the step of determining the action to be applied to the prompt as a reconfiguration action when the risk level is greater than or equal to the second threshold and less than the third threshold, wherein the third threshold is a value greater than the second threshold, and the reconfiguration action is a measure to generate the security prompt by reconfiguring the prompt through substitution-based de-identification, and when the risk level is greater than or equal to the third threshold, the action to be applied to the prompt may be determined as a transmission blocking action.

[0024] In some embodiments, the analysis result includes information regarding one or more text sections related to the security risk within the prompt, the determined action is a partial allowance action, and the step of generating the security prompt may include the step of generating the security prompt by de-identifying the one or more text sections.

[0025] In some embodiments, the step of generating the security prompt comprises: identifying one or more sensitive entities in the prompt; and generating the security prompt by performing entity-level de-identification on the prompt based on the result of the identification, wherein the generated security prompt may include an abstract information type corresponding to at least some of the one or more sensitive entities.

[0026] In some embodiments, the step of transmitting to the external AI service comprises: generating a verification prompt based on the prompt, the security prompt, and a set of instructions; providing the verification prompt to the language model to verify the security prompt; and transmitting the security prompt to the external AI service based on the result of the verification, wherein the set of instructions may include instructions for verifying the existence of the security risk with respect to the security prompt and instructions for verifying whether the meaning of the security prompt is distorted by comparison with the prompt.

[0027] A method for providing AI security governance according to other embodiments of the present disclosure for solving the aforementioned technical problem may include, in a method performed by at least one processor, the steps of: receiving a response to a user prompt from an external AI (Artificial Intelligence) service—the prompt being transmitted to the external AI service after undergoing a processing process based on a predefined security policy—; deriving an analysis result related to a security risk by analyzing the contextual meaning of the response through a language model—the security risk includes a risk resulting from a violation of the security policy—; calculating the risk level of the response based on the analysis result; determining a measure based on the risk level based on the security policy; generating a security response by performing the measure on the response; and transmitting the security response to the user's terminal.

[0028] In some embodiments, the step of determining the action according to the risk level includes: determining the action to be applied to the response as a fully allowed action when the risk level is less than a first threshold; determining the action to be applied to the response as a partially allowed action when the risk level is greater than or equal to the first threshold and less than a second threshold; determining the action to be applied to the response as a reconfiguration action when the risk level is greater than or equal to the second threshold and less than a third threshold; and determining the action to be applied to the response as an alternative response provision action when the risk level is greater than or equal to the third threshold, wherein the second threshold is a value smaller than the third threshold and larger than the first threshold, the partially allowed action includes processing to de-identify a part of the response, the reconfiguration action is a measure to generate the security response by reconfiguring the representation of the response, and the alternative response provision action may be a measure to generate a new response to replace the response based on the security policy and provide the new response as the security response.

[0029] In some embodiments, the prompt is transmitted to the external AI service through de-identification processing, and the step of transmitting to the user's terminal may include: providing information regarding the de-identification processing and the security response to the language model to generate a restoration response; and transmitting the restoration response to the terminal.

[0030] An AI security governance system according to some embodiments of the present disclosure for solving the technical problem described above comprises: one or more processors; and a memory for storing a computer program executed by said one or more processors, wherein the computer program may include instructions for: obtaining a user prompt for an external AI (Artificial Intelligence) service; deriving an analysis result related to a security risk by analyzing the contextual meaning of said prompt through a language model; calculating a risk level of said prompt based on said analysis result; determining a measure according to said risk level based on a predefined security policy; generating a security prompt by performing said measure on said prompt; and transmitting said security prompt to said external AI service.

[0031] An AI security governance system according to some embodiments of the present disclosure for solving the technical problem described above comprises: one or more processors; and a memory for storing a computer program executed by said one or more processors, wherein the computer program may include instructions for: receiving a response to a user prompt from an external AI (Artificial Intelligence) service—said that the prompt is transmitted to said external AI service after undergoing a processing process based on a predefined security policy—; deriving an analysis result related to a security risk by analyzing the contextual meaning of said response through a language model—said that the security risk includes a risk resulting from a violation of said security policy—; calculating a risk level of said response based on said analysis result; determining a measure according to said risk level based on said security policy; generating a security response by performing said measure on said response; and transmitting said security response to said user terminal.

[0032] A computer program according to some embodiments of the present disclosure for solving the above-described technical problem may be combined with a computer and stored on a non-transitory computer-readable recording medium to execute the AI ​​security governance providing method. Effects of the invention

[0033] According to some embodiments of the present disclosure described above, by analyzing the contextual meaning of a prompt using a language model, even security risks that are difficult to identify through inspection of the surface expression of the prompt alone, such as indirect information leakage and prompt injection, can be effectively detected. Accordingly, security risks that may occur in an environment using external AI (Artificial Intelligence) services can be effectively controlled.

[0034] In addition, by utilizing static inspection results as context information, the accuracy of contextual semantic analysis of prompts can be improved.

[0035] In addition, the risk level of a prompt (or request) can be calculated based on semantic analysis results, static inspection results, etc., and measures based on the risk level can be performed according to predefined security policies. For example, measures such as full allowance, partial allowance, reconfiguration, and transmission blocking can be performed differentially on a prompt depending on the risk level. In such cases, security levels can be maintained while preventing excessive blocking, thereby ensuring both security and user convenience simultaneously.

[0036] In addition, by considering various evaluation factors such as user permissions and the type of task requested by the prompt, in addition to static inspection results and semantic analysis results, the risk level of the prompt can be calculated more accurately.

[0037] Furthermore, by performing anonymization at the level of risk sections or sensitive entities within the prompt, the leakage of sensitive information can be prevented while simultaneously minimizing the loss of context within the prompt. Accordingly, the issue of reduced user convenience regarding external AI services due to security processing can be effectively mitigated.

[0038] Furthermore, by applying risk analysis and differential measures to inbound responses in a manner similar to outbound requests, integrated security management across the entire process of using external AI services becomes possible. Consequently, an AI security governance framework can be established that enables the application and management of consistent security policies within the organizational environment.

[0039] In addition, by recording the processing history of prompts (or requests) and responses in the audit log and transmitting relevant information to administrator terminals when security events occur, an AI governance and audit system can be effectively established in organizational environments utilizing external AI services.

[0040] The effects according to the technical concept of the present disclosure are not limited to those mentioned above, and other unmentioned effects will be clearly understood by a person skilled in the art from the description below. Brief explanation of the drawing

[0041] FIG. 1 is a diagram illustrating an exemplary operating environment of an Artificial Intelligence (AI) security governance system according to some embodiments of the present disclosure. FIG. 2 is a drawing for illustrating an exemplary implementation of an AI security governance system according to some embodiments of the present disclosure. FIGS. 3 and 4 are exemplary block diagrams for explaining the configuration and operation of an AI security governance system according to some embodiments of the present disclosure. FIG. 5 is an exemplary flowchart illustrating an outbound request processing process according to some embodiments of the present disclosure. FIG. 6 is an exemplary drawing for explaining the contextual semantic analysis process of a prompt according to some embodiments of the present disclosure. FIG. 7 is a diagram showing an example of the result of contextual semantic analysis of a prompt that may be referenced in some embodiments of the present disclosure. FIGS. 8 and 9 are exemplary drawings for illustrating a method for calculating risk according to some embodiments of the present disclosure. FIG. 10 is an exemplary drawing for explaining a weight determination method according to some embodiments of the present disclosure. FIG. 11 is an exemplary flowchart illustrating the process of determining and executing risk-based measures according to some embodiments of the present disclosure. FIG. 12 illustrates a process for performing a partial acceptance measure according to some embodiments of the present disclosure. FIG. 13 illustrates a process for performing prompt reconfiguration measures according to some embodiments of the present disclosure. FIG. 14 is an exemplary drawing for illustrating a security prompt verification process according to some embodiments of the present disclosure. FIG. 15 is an exemplary flowchart illustrating an inbound response processing process according to some embodiments of the present disclosure. FIG. 16 is an exemplary drawing for illustrating the process of analyzing the contextual meaning of a response according to some embodiments of the present disclosure. FIG. 17 is an exemplary drawing for explaining a restoration response generation process according to some embodiments of the present disclosure. FIG. 18 illustrates an exemplary computing device capable of implementing an AI security governance system according to some embodiments of the present disclosure. Specific details for implementing the invention

[0042] Hereinafter, various embodiments of the present disclosure will be described in detail with reference to the attached drawings. The advantages and features of the present disclosure and the methods for achieving them will become clear by referring to the embodiments described below in detail together with the attached drawings. However, the technical concept of the present disclosure is not limited to the following embodiments but can be implemented in various different forms. The following embodiments are provided merely to complete the technical concept of the present disclosure and to fully inform those skilled in the art of the scope of the present disclosure, and the technical concept of the present disclosure is defined only by the scope of the claims.

[0043] In describing the various embodiments of the present disclosure, if it is determined that a detailed description of related known configurations or functions could obscure the essence of the present disclosure, such detailed description is omitted.

[0044] Unless otherwise defined, terms used in the following embodiments (including technical and scientific terms) may be used in a meaning commonly understood by those skilled in the art to which this disclosure pertains, but this may vary depending on the intent of those skilled in the art, case law, the emergence of new technology, etc. The terms used in this disclosure are for describing the embodiments and are not intended to limit the scope of this disclosure.

[0045] In the following embodiments, singular expressions include plural concepts unless the context clearly specifies them as singular. Additionally, plural expressions include singular concepts unless the context clearly specifies them as plural.

[0046] In addition, terms such as first, second, A, B, (a), (b), etc. used in the following embodiments are used merely to distinguish one component from another, and the essence, order, or sequence of the said component is not limited by such terms.

[0047] In the following embodiments, the components described using terms such as ~part or unit, module, block, ~or, ~er, etc., and the functional blocks illustrated in the drawings may be implemented in the form of software, hardware, or a combination thereof. Software may include, for example, machine code, firmware, embedded code (or software), application software, or a combination thereof. Additionally, hardware may include, for example, electrical circuits, electronic circuits, processors, computers, integrated circuits, integrated circuit cores, passive components, or a combination thereof. As a more specific example, such as ~part, module, etc., may include components such as software components, object-oriented software components, class components, and task components, as well as processes, functions, attributes, procedures, subroutines, segments of program code, drivers, firmware, microcode, circuits, data, databases, data structures, tables, arrays, and variables.

[0048] Hereinafter, various embodiments of the present disclosure will be described in detail with reference to the attached drawings.

[0049] FIG. 1 is a drawing for illustrating an exemplary operating environment of an AI (Artificial Intelligence) security governance system (10) according to some embodiments of the present disclosure.

[0050] As illustrated in FIG. 1, the AI ​​security governance system (10) is a computing system equipped with various security governance functions for controlling security risks in an environment using one or more external AI services (13-1 to 13-K). The AI ​​security governance system (10) may operate in conjunction with one or more user terminals (11-1 to 11-N), an administrator terminal (12), and / or one or more external AI services (13-1 to 13-K). However, the scope of the present disclosure is not limited thereto, and the AI ​​security governance system (10) may also operate in conjunction with other devices and / or systems (e.g., security monitoring systems, etc.).

[0051] In the following, for the sake of convenience of understanding, the reference number '11' will be used in both cases where any user terminal (e.g., 11-1 or 11-2) is referred to individually and cases where multiple user terminals (e.g., 11-1 to 11-N) are referred to collectively.

[0052] Similarly, the reference number '13' is used in both cases where any external AI service (e.g., 13-1 or 13-2) is referred to individually and multiple external AI services (e.g., 13-1 to 13-K) are referred to collectively.

[0053] Additionally, for the sake of brevity of the present disclosure, the AI ​​security governance system (10) will be abbreviated as 'system' below.

[0054] The system (10) can perform various security governance functions while relaying requests and responses between a user terminal (11) and an external AI service (13). Specifically, the system (10) can receive a request for an external AI service (13) from the user terminal (11) and calculate a risk level by analyzing the prompt and / or attached data included in the request. At this time, the system (10) can use a language model to analyze the contextual meaning of the prompt, etc., and calculate the risk level of the request by reflecting the analysis results. Accordingly, even security risks that are difficult to identify by inspecting only the surface expression of the prompt, such as indirect information leakage or prompt injection, can be effectively detected and controlled. Here, security risk can be understood as a concept that encompasses not only general security risks but also risks arising from violations of security policies (e.g., false information, biased expression, etc.). Then, the system (10) can perform measures according to the risk level (i.e., differential security measures) on the request based on a predefined security policy. Examples of measures to be performed may include, but are not limited to, blocking transmission, allowing transmission (e.g., full allowance, partial allowance, etc.), reconfiguring prompts, and de-identifying prompts. The system (10) may also perform security processing in a similar manner on responses received from an external AI service (13).

[0055] Additionally, the system (10) can record processing history for requests and responses in the form of an audit log and transmit related information to an administrator terminal (12) when a security event occurs. In this case, an AI governance and audit system can be effectively established in an organizational environment using an external AI service (13).

[0056] The detailed operations of the system (10) will be explained in more detail later with reference to the drawings from Fig. 3 onwards.

[0057] The above-described system (10) may be referred to as an ‘AI security management system’, ‘AI security control system’, ‘AI governance system’, ‘AI security audit system’, ‘AI policy management system’, etc. depending on the case, and the term system may be used interchangeably with terms such as ‘device’, ‘server’, ‘platform’, etc.

[0058] The above-described system (10) may be implemented with at least one computing device. For example, all functions of the system (10) may be implemented in a single computing device, or a first function of the system (10) may be implemented in a first computing device and a second function may be implemented in a second computing device. Alternatively, a specific function of the system (10) may be implemented in multiple computing devices.

[0059] A computing device may include any device equipped with computing (or processing) functions, and for an example of such a device, refer to FIG. 18. Since a computing device is a collection of various components (e.g., memory, processor, etc.) that interact, it may be referred to as a 'computing system' depending on the case. Of course, the term computing system may also encompass the concept of a collection of multiple computing devices that interact.

[0060] The external AI service (13) refers to an AI service located in an external network. Such an AI service may be, for example, a generative AI-based service configured to receive a prompt as input and perform a requested task, but the scope of the present disclosure is not limited thereto. As a more specific example, the external AI service (13) may be a service capable of performing various tasks (e.g., summarization, analysis, modification, conversion, reconstruction, generation, etc.) on data such as text, code, images, and audio according to a prompt, based on a generative language model, a generative image model (e.g., text-to-image model, etc.), a generative audio model, etc. The conceptual scope of the external AI service (13) may also include a service in the form of an AI agent based on a generative model.

[0061] A generative language model is a deep learning model equipped with the ability to understand and generate text in which language is expressed, and may be referred to as a 'Large Language Model', 'Generative Text Model', 'Language Model', etc. depending on the case.

[0062] The user terminal (11) is a terminal on the user side that uses the external AI service (13). The user can use the external AI service (13) by entering a prompt through the terminal (11), and in the process, a request including the prompt may be generated. The request from the user terminal (11) is security-processed by the system (10), and depending on the result, transmission to the external AI service (13) may be allowed or blocked. In addition, the response of the external AI service (13) to the request may also be security-processed by the system (10).

[0063] The user may be, for example, a member performing work within an organization, but is not limited thereto. The user terminal (11) may be a mobile device such as a smartphone, or a fixed device such as a desktop.

[0064] The user terminal (11) may be referred to as a 'client terminal', 'service usage terminal', etc. depending on the case, and the term terminal may be used interchangeably with terms such as 'apparatus', 'device', 'system', etc.

[0065] The administrator terminal (12) is a terminal used by an administrator (e.g., security expert, etc.) related to AI security governance. The administrator terminal (12) can receive information related to security events from the system (10), and the administrator can check the information through the terminal (12). In addition, the administrator can perform tasks such as policy setting, security event definition and setting, security event monitoring, and response to security events through the terminal (12).

[0066] As described, the system (10), user terminal (11), administrator terminal (12) and / or external AI service (13) can communicate through a network. Here, the network can be implemented as any type of wired or wireless network, such as a Local Area Network (LAN), a Wide Area Network (WAN), a mobile radio communication network, or Wibro (Wireless Broadband Internet).

[0067] FIG. 2 is a drawing for illustrating an exemplary implementation of a system (10) according to some embodiments of the present disclosure.

[0068] As illustrated in FIG. 2, the system (10) is positioned between the internal network (21) and the external network (22) and can be implemented in the form of a proxy server. Here, being "positioned between" can be understood as a concept that encompasses not only cases where the system (10) is physically located between the internal network (21) and the external network (22) (e.g., boundary area), but also cases where communication between the internal network (21) and the external network (22) is logically configured to pass through the system (10). For example, even if the system (10) is located in the internal network (21), if communication between the user terminal (11) and the external AI service (13) is configured to pass through the system (10), the system (10) can be considered to be positioned between the internal network (21) and the external network (22).

[0069] The system (10) can separate sessions between a user terminal (11) located in an internal network (21) and an external AI service (13) located in an external network (22), and relay requests and responses based on a predefined security policy. At this time, the system (10) can perform appropriate measures based on the security policy, which will be described later.

[0070] Up to now, exemplary operating environments and implementation forms of a system (10) according to some embodiments of the present disclosure have been described with reference to FIGS. 1 and 2. Hereinafter, the system (10) will be described in more detail with reference to FIGS. 3 and 4.

[0071] FIGS. 3 and 4 are exemplary block diagrams for explaining the configuration and operation of a system (10) according to some embodiments of the present disclosure. FIGS. 3 and 4 each illustrate exemplary data flows related to the processing of outbound requests and inbound responses.

[0072] As illustrated in FIG. 3 or FIG. 4, the system (10) may be configured to include a relay unit (31), a static inspection unit (32), a semantic analysis unit (33), a policy execution unit (34), an audit log management unit (35), and a reporting unit (36). However, FIG. 3 illustrates only the components related to the embodiments of the present disclosure. Therefore, a person skilled in the art to which the present disclosure belongs will understand that other general-purpose components (e.g., processor, memory, database, etc.) may be included in addition to the components (31 to 36) illustrated in FIG. 3. Furthermore, the components (31 to 36) of the system (10) illustrated in FIG. 3 represent functionally distinct functional elements; multiple components may be implemented in a form where they are integrated with each other in an actual physical environment, or specific components may be implemented in a form where they are separated into multiple sub-components. Each of the components (31 to 36) of the system (10) will be described below.

[0073] The relay unit (31) is a module that relays requests and responses between a user terminal (11) and an external AI service (13). The relay unit (31) can transmit requests to the external AI service (13) or transmit responses to the user terminal (11) based on a security policy. Here, requests and responses can be understood as encompassing those that are modified or replaced according to security processing.

[0074] The request from the user terminal (11) may include, for example, a prompt, attached data, etc. A prompt is a description indicating the content or instructions of the task the user wishes to request, and may be expressed in natural language form, but is not limited thereto. Attached data is data entered by the user along with the prompt, and may include data of various forms (or file formats) without restriction, such as text documents, image data (e.g., document images, etc.), audio data (e.g., voice recording files, etc.), video data, etc. Such attached data may be used as reference information for performing a specific task by an external AI service (13). Attached data may be provided in the form of a file, but is not limited thereto. Attached data may be entered as data that is not packaged into a file, or may be entered in a form that is referenced from an external storage.

[0075] Additionally, the relay unit (31) can perform session management between the user terminal (11) and the external AI service (13). Specifically, the relay unit (31) manages the entire session between the user terminal (11) and the external AI service (13) by separating it into a first session between the user terminal (11) and the system (10) and a second session between the system (10) and the external AI service (13), and relays requests and responses through the first session and the second session.

[0076] For the operation of the relay unit (31), further refer to the descriptions in the drawings below Fig. 5.

[0077] The static inspection unit (32) is a module that performs static inspection based on predefined inspection rules. The static inspection unit (32) can perform static inspection on the prompt and attached data included in the request of the user terminal (11) and the response of the external AI service (13).

[0078] For example, the static inspection unit (32) can perform static inspection on prompts and / or attached data using inspection rules based on file extensions, MIME (Multipurpose Internet Mail Extensions) types, magic numbers, string patterns, forbidden word dictionaries, sensitive information patterns (e.g., regular expressions, etc.), security level tags, etc. Here, sensitive information can be understood as a concept encompassing personal information and confidential information.

[0079] As another example, the static inspection unit (32) can perform static inspection on the response using inspection rules based on file extensions, MIME types, magic numbers, sensitive information patterns, attack patterns (e.g., script injection patterns, SQL injection patterns, etc.), malicious code patterns, malicious lists (e.g., malicious domain lists, malicious URL lists, etc.), policy violation patterns, etc.

[0080] Static inspection rules can be managed as security policies of the system (10) and can be defined and set according to the organization's operating environment and security requirements.

[0081] For the operation of the static inspection unit (32), further refer to the descriptions in the drawings 5 ​​and below.

[0082] The semantic analysis unit (33) is a module that analyzes contextual semantics related to security risks using a language model. The semantic analysis unit (33) can perform semantic analysis on prompts and attached data included in the request of the user terminal (11) and the response of the external AI service (13).

[0083] The language model may be, for example, a generative language model, but the scope of the present disclosure is not limited thereto. The language model may be, for example, an internal model that operates and runs within an organization without transmitting data externally, or an external model that operates in a secure environment. The language model may be fine-tuned with a dataset of the domain to which the system (10) belongs. The semantic analysis unit (33) may perform semantic analysis using one or more language models.

[0084] For example, the semantic analysis unit (33) can derive a structured analysis result by analyzing the contextual meaning of the prompt through a language model. At this time, the semantic analysis unit (33) may also analyze the prompt based further on static inspection results and may perform semantic analysis on attached data. For details regarding the items of the analysis result, refer to the descriptions in FIGS. 6 and FIGS. 7.

[0085] As another example, the semantic analysis unit (33) can derive a structured analysis result by analyzing the contextual meaning of the response through a language model.

[0086] For the operation of the meaning analysis unit (33), further refer to the descriptions in the drawings below Fig. 5.

[0087] The policy enforcement unit (34) is a module that performs measures to prevent security risks based on a predefined security policy. Specifically, the policy enforcement unit (34) can calculate the risk level of a prompt (or request) based on static inspection results, semantic analysis results, etc., and perform measures according to the risk level based on the security policy. The security policy can be defined to include, for example, a risk level (e.g., risk level, risk range, etc.) and corresponding measures. In addition, the policy enforcement unit (34) can also perform measures based on the security policy for responses received from an external AI service (13).

[0088] For the operation of the policy execution unit (34), please refer further to the descriptions in the drawings below Fig. 5.

[0089] The audit log management unit (35) is a module that records information related to security processing (i.e., processing based on security policies) performed in the system (10) in the audit log and manages it. For example, the audit log management unit (35) can record various information related to the security processing of requests and responses (e.g., static inspection results, semantic analysis results, risk level, applied security policy, measures taken, user identifier, processing time, final processing result, etc.) in the audit log. In addition, the audit log management unit (35) can also record information regarding pre-configured security events in the audit log. Here, security events are related to the occurrence of security policy violations, signs of risk, etc., and may include, for example, prompt injection, detection of sensitive information leakage, detection of evasive information leakage, detection of malicious patterns (e.g., malicious code patterns, attack patterns, etc.), and detection of requests and responses where the risk level exceeds a threshold, but are not limited thereto.

[0090] In addition, the audit log management unit (35) can analyze accumulated audit logs to detect additional security events such as repeated security policy violations and attempts to bypass security policies.

[0091] For the operation of the audit log management unit (35), further refer to the descriptions in the drawings 5 ​​and below.

[0092] The reporting unit (36) is a module that provides information related to security events to the administrator terminal (12) and / or the security control system (not shown). For example, the reporting unit (36) can send an alert to the administrator terminal (12) when a security event occurs, or provide a summary of related information.

[0093] For the operation of the reporting unit (36), further refer to the descriptions in the drawings below Fig. 5.

[0094] For reference, at least some of the components (31 to 36) of the system (10) may form a guardrail module and may be implemented in the form of an agent. For example, the relay unit (31), static inspection unit (32), semantic analysis unit (33), and policy enforcement unit (34) may form a guardrail module, and in some cases, at least one of the audit log management unit (35) and the reporting unit (36) may be further included in the guardrail module.

[0095] Up to now, the configuration and operation of a system (10) according to some embodiments of the present disclosure have been described with reference to FIGS. 3 and 4. Hereinafter, various embodiments related to a method for providing AI security governance will be described with reference to FIGS. 5 and subsequent drawings.

[0096] The method described below may be performed by at least one processor or implemented by said processor. For example, each step / operation of the method described below may be performed by at least one processor equipped in the system (10) described above in the environment illustrated in FIG. 1 or FIG. 2. However, the scope of the present disclosure is not limited thereto, and depending on the implementation method, some steps / operations of the method described below may be performed on other computing devices.

[0097] FIG. 5 is an exemplary flowchart illustrating an outbound request processing process according to some embodiments of the present disclosure. However, this is merely an exemplary embodiment for achieving the purpose of the present disclosure, and it is understood that some steps may be added or deleted as necessary.

[0098] As illustrated in FIG. 5, the processing process according to the embodiments may begin at step S51, which receives a request containing a prompt for an external AI service. For example, the system (10, e.g., relay unit (31)) may receive a request for an external AI service (13) from a user terminal (11). At this time, the request may further include attachment data associated with the prompt.

[0099] In step S52, a static inspection is performed on the prompt. For example, the system (10, e.g., static inspection unit (32)) may perform a static inspection on the prompt based on pre-set inspection rules. If associated attachment data exists, the system (10) may also perform a static inspection on the attachment data. In some cases, the system (10) may convert the attachment data into a text format (e.g., recognizing text from a document image) and apply text-related inspection rules to the conversion result.

[0100] In some embodiments, the external transmission of a prompt may be immediately blocked based on the results of a static inspection. For example, if the static inspection results satisfy conditions indicating a clear security risk, the system (10) may block the external transmission of the corresponding prompt. Examples of such conditions include, but are not limited to, the detection of a certain number of sensitive information patterns (e.g., patterns of social security numbers, account numbers, etc.), the detection of predefined secret keywords (e.g., internal top-secret project names, filenames of major source code, etc.), or the attachment of a file with a predefined extension (e.g., executable files). In the embodiments, the system (10, e.g., audit log management unit (35), reporting unit (36)) may record information related to the blocked prompt (e.g., prompt, static inspection results, related conditions, user identifier, etc.) in the audit log and may send an alert to the administrator terminal (12). Alternatively, the system (10) may send a warning regarding a security policy violation to the user terminal (11).

[0101] In step S53, the contextual meaning of the prompt is analyzed through a language model based on the results of static inspection. This analysis process can be understood as being performed to identify intentions not revealed in the surface expression of the prompt, thereby enabling the effective identification of security risks such as prompt injection and circumvention of information.

[0102] For ease of understanding, the detailed process of step S53 will be explained in detail with reference to Fig. 6.

[0103] FIG. 6 is an exemplary drawing for explaining the contextual semantic analysis process of a prompt according to some embodiments of the present disclosure.

[0104] As illustrated in FIG. 6, the contextual meaning of the prompt (62) can be analyzed through the language model (60). As described above, the language model (60) may be an internal model or an external model operating in a secure environment. The language model (60) may be, for example, fine-tuned with a dataset of a domain associated with the system (10) (e.g., a dataset related to internal organizational business), but is not limited thereto.

[0105] For example, the system (10) can generate an analysis prompt (61) based on a prompt (62), static inspection results (63), a set of instructions (64), etc., and provide it to a language model (60) to analyze the contextual meaning of the prompt (62). As a result, an analysis result (65) related to security risks can be derived (or generated). If there is attached data associated with the prompt (62), the system (10) can provide the attached data to the language model (60) along with the analysis prompt (61). This can be understood as being because the security risk associated with the prompt (62) may vary depending on the content of the attached data. Additionally, the system (10) may perform an independent semantic analysis of the attached data using the language model (60) separately from the prompt (62).

[0106] In some cases, the system (10) may analyze the contextual meaning of the prompt (62) without relying on the static check result (63). That is, the static check result (63) may be omitted from the analysis prompt (61).

[0107] The static inspection result (63) may include, for example, sensitive information detected according to predefined inspection rules, security class tags of assets (e.g., documents, source code, etc.), prohibited keywords, etc., but is not limited thereto. The static inspection result (63) can be used as context information when analyzing the contextual semantics of the prompt (62) to improve the analysis accuracy of the language model (60).

[0108] The set of instructions (64) may include, for example, instructions requesting the analysis of the contextual meaning of the prompt (62), instructions to refer to the static check result (63) during the semantic analysis process, etc.

[0109] In addition, the instruction set (64) and / or analysis prompt (61) may further include various instructions / information / data / conditions / rules / explanations to improve the accuracy of the analysis result (65).

[0110] For example, the analysis prompt (61) may further include definitions of risk types, rules for designating risk ranges, criteria for determining recommended measures, etc. That is, the analysis prompt (61) may further include various descriptions for deriving the values ​​of each item to be included in the analysis result (65).

[0111] As another example, the analysis prompt (61) may further include instructions that define step-by-step procedures for deriving analysis results (65). These instructions may be configured, for example, to identify risk zones in the prompt, evaluate the risk type and risk level for each zone, calculate the risk score of the prompt based on the evaluation results, and determine recommended measures based on the risk score, but are not limited thereto.

[0112] As another example, the analysis prompt (61) may include one or more additional analysis examples. The analysis examples may include, but are not limited to, prompt examples with high difficulty in contextual semantic analysis, and semantic analysis results for said prompt examples (e.g., correct result, incorrect result, reason for error, etc.). Such prompt examples may be prepared in advance, for example, by a security expert, or may be automatically selected based on the reliability of the analysis results produced by the language model (60) (see below). For instance, among the previously stored prompts, those with a reliability below a threshold may be automatically selected, and the semantic analysis results included in the analysis examples may be reviewed by a security expert.

[0113] As another example, the analysis prompt (61) may further include the output format of the analysis result (65) (see FIG. 7) and instructions to enforce it.

[0114] As another example, the analysis prompt (61) may further include instructions to dynamically adjust the number and range of risk intervals based on quantitative risk assessment indicators (e.g., risk score, confidence level, etc.) for the prompt (62).

[0115] As another example, the analysis prompt (61) may be generated based on various combinations of the examples described above.

[0116] The analysis results (65) may include items such as risk type, risk range, type of analysis target, risk assessment indicator, recommended measures, basis for judgment, type of work, etc., but are not limited thereto.

[0117] Risk Type is an item indicating the type of security risk associated with the subject of analysis (e.g., prompt, response, attached data). Types of security risks may be classified, for example, as exposure of personal information, leakage of confidential information, circumvention of information requests, prompt injection, malicious responses, false responses, biased responses, or other responses violating policies, but are not limited thereto.

[0118] The type of analysis target is an item indicating the category of the analysis target. Analysis targets may be classified, for example, as outbound prompts or inbound responses, but are not limited to these.

[0119] A risk section is an item indicating the location and / or scope of a text section related to security risks within the subject of analysis. This location and / or scope information may be expressed in units of documents, paragraphs, sentences, phrases, entities, tokens, and / or characters, but is not limited thereto.

[0120] Risk assessment indicators are items representing quantitative risk assessment results for the subject of analysis. Risk assessment indicators may include, for example, a risk score representing the level of security risk of the subject of analysis, the reliability of the language model (60) for the analysis result (65), but are not limited thereto. The risk score may be expressed in the form of a numerical value, grade, or level.

[0121] Recommended measures are items intended to indicate the types of security measures recommended for the subject of analysis. In addition to type information such as full allow, partial allow, reconfiguration, and providing alternative responses, such items may include information regarding de-identification methods such as removal, masking, pseudonymization, and generalization, but are not limited thereto.

[0122] The basis for judgment is an item to explain the reason why the analysis result (65) was derived.

[0123] The task type is an item to indicate the type of task requested through the prompt (62).

[0124] The analysis result (65) can be output (or generated) in a structured format, for example. An example of such an analysis result (65) is illustrated in FIG. 7. FIG. 7 illustrates an example where the output format of the analysis result (71) is defined in JSON (JavaScript Object Notation) format.

[0125] In some embodiments, multiple items included in the analysis result (65) may be defined as being divided into mandatory items and optional items. For example, items such as recommended measures and grounds for judgment may be defined as optional items, and the remaining items may be defined as mandatory items. In such cases, the system (10) may determine whether there are any missing mandatory items in the analysis result (65), and if there are missing items, it may perform re-analysis through the language model (60). Alternatively, the system (10) may block the external transmission of the prompt (62) without performing re-analysis.

[0126] Up to now, the process of analyzing the contextual meaning of a prompt according to some embodiments of the present disclosure has been described with reference to FIGS. 6 and FIGS. 7.

[0127] Referring again to Fig. 5, the explanation will be provided.

[0128] In step S54, the risk of the prompt is calculated based on static inspection results and semantic analysis results. For example, the system (10, e.g., policy enforcement unit (34)) may calculate the risk of the prompt based on multiple evaluation factors including static inspection results, semantic analysis results, etc. For examples of evaluation factors, refer to Table 1 below. Table 1 below summarizes the evaluation factors used to calculate the risk of the prompt or response. The risk may be calculated using at least some of the evaluation factors listed in Table 1.

[0129] Evaluation factors explanation Score (value) example Static inspection results Indicates the level of security risk based on static inspection results The score increases as the number of violation rules and detections increases. Semantic analysis results Indicates the level of security risk derived through contextual semantic analysis. - The higher the risk score or confidence level, the higher the score. - The score increases if it corresponds to a specific risk type. Asset security rating Indicates the security level of the related asset The higher the security level, the higher the score. User permissions Indicates the user's permission level The lower the authority, the higher the score. Task type Indicates the level of security risk based on the type of operation requested through the prompt. - Scores increase in the following order - Simple Proofreading / Translation < Summarization / Organization < Rewriting for External Submission < Generation of Code / Scripts Related to Internal System Configuration / Security Settings Whether the same or similar requests are repeated Indicates the level of repetition of identical or similar requests The higher the number of repetitions or frequency, the higher the score. Harmfulness of Inbound Responses Indicates the level of security risk based on the maliciousness or inappropriateness of the response content The higher the malignancy and / or inappropriateness, the higher the score

[0130] The specific method for calculating the risk level may vary depending on the example.

[0131] In some embodiments, the risk of the prompt may be calculated by combining the values ​​of each evaluation element based on weights. For example, the system (10) may calculate a score for each evaluation element and calculate the risk of the prompt by aggregating the evaluation scores based on weights. A specific method for determining weights will be described later.

[0132] In some other embodiments, the risk of the prompt can be calculated using a trained machine learning model. In this case, the risk of the prompt can be calculated more precisely, which will be explained in detail shortly with reference to FIGS. 8 and 9.

[0133] In some other embodiments, the risk level of a prompt may be calculated based further on the risk analysis results of the expected response to the prompt (e.g., semantic analysis results, static inspection results, risk level, etc.). For example, the system (10) may generate an expected response to a prompt through a language model (60) and may calculate the risk level of the prompt based further on the risk analysis results of the expected response. Here, the risk analysis results of the expected response may be set to correspond to the maliciousness of the inbound response listed in Table 1. According to these embodiments, it is possible to calculate the risk level by considering in advance not only the prompt itself but also the results that may be generated from the prompt, thereby enabling more proactive and precise security control.

[0134] In some other embodiments, the security level of an asset associated with a prompt (e.g., attachments / data, documents, data, etc. mentioned in the prompt) may be determined in conjunction with an organization’s information asset classification system or related systems, and the risk level of the prompt may be calculated by reflecting this. For example, the system (10) may determine the security level of an asset associated with a prompt by referring to security level information managed by a document management system, a file labeling system (e.g., Data Loss Prevention (DLP) system, Information Rights Management (IRM) system), a metadata management system, or an information asset management system (e.g., Configuration Management Database (CMDB)) established within the organization. In some cases, the system (10) may estimate the security level by recognizing tags or keywords such as "[Confidential]" or "CONFIDENTIAL" included in the prompt. The system (10) may calculate the risk level of the prompt by considering these asset security levels as evaluation factors (see Table 1). According to these embodiments, since the risk level of a prompt is calculated based on an asset security rating organically linked with the organization's internal information security infrastructure, more consistent and precise security controls can be achieved. This method of calculating risk can also be applied to calculating the risk level of a response.

[0135] In some other embodiments, the risk of the prompt may be calculated based on various combinations of the embodiments described above. For example, the system (10) may calculate the final risk of the prompt by combining the first risk according to the first embodiment and the second risk according to the second embodiment.

[0136] Hereinafter, a method for calculating the risk of a prompt using a machine learning model (80) will be explained with reference to FIGS. 8 and FIGS. 9. The method described below can also be used to calculate the risk of a response.

[0137] FIG. 8 is an exemplary drawing for explaining a method for calculating risk according to some embodiments of the present disclosure.

[0138] As illustrated in FIG. 8, the machine learning model (80) can be configured to receive the value of each evaluation element (i.e., evaluation result) as input and output a risk level. The value of the evaluation element can be preprocessed into an appropriate form and input into the machine learning model (80), and any specific preprocessing method is acceptable.

[0139] The machine learning model (80) can be implemented based on various types / forms of models. For example, the machine learning model (80) may be implemented based on a classification model (e.g., when the risk level is expressed in the form of a grade), or it may be implemented based on a regression model. Additionally, the machine learning model (80) may be implemented based on a neural network, or it may be implemented based on a traditional machine learning model.

[0140] The machine learning model (80) can be trained using a supervised learning method. For example, the system (10) can train the machine learning model (80) using a training set labeled with whether a security risk has occurred or a correct risk level. The training set may include a plurality of training samples and corresponding labels, and each training sample may consist of input data as illustrated in FIG. 8. In some cases, the risk level of the response corresponding to the prompt may be used as the correct risk level.

[0141] In some embodiments, as illustrated in FIG. 9, a correct action (96) is determined based on the result of a response risk assessment for a security prompt (e.g., 94), and a risk value corresponding to the correct action (96) may be designated as the correct risk (97) for training a machine learning model (80). For example, the system (10) may generate multiple security prompts (e.g., 94, 95) by performing different actions (e.g., full allow, partial allow, reconstruction, masking, generalization, pseudonymization, etc.) on a prompt sample (92). Then, the system (10) may evaluate the risk of the response corresponding to each security prompt (e.g., 94) and determine the action from which the response with the lowest risk among the multiple actions is derived as the correct action (96). Next, the system (10) can determine a risk level (e.g., a value sampled from a risk range, etc.) corresponding to the correct action (96) based on a security policy and designate it as the correct risk level (97). As illustrated, the correct risk level (97) can be used to update the parameters of the machine learning model (80) based on the difference from the predicted risk level (91). According to these embodiments, the machine learning model (80) can be trained to calculate a risk level suitable for determining the action, and as a result, security risk can be effectively reduced.

[0142] A trained machine learning model (80) can be used to calculate the risk level of a prompt entered by a user. For example, the system (10) can input the value of each evaluation factor into the trained machine learning model (80) to calculate the risk level of the corresponding prompt.

[0143] Up to now, a method for calculating risk according to some embodiments of the present disclosure has been described with reference to FIGS. 8 and 9. As described above, the risk of a prompt can be calculated more precisely by using a machine learning model (80).

[0144] Meanwhile, the specific method for determining the weight of each evaluation element may also vary depending on the embodiment.

[0145] In some embodiments, the weights may be determined by a security expert. For example, the security expert may determine the weight of each evaluation factor by reflecting the organizational environment, the characteristics of the domain in which the system (10) operates, etc.

[0146] In some other embodiments, weight profiles may be defined for each risk type. Here, a weight profile may refer to a set of information including weights for each evaluation element, and in some cases, a risk calculation method may be further included in the weight profile. In such cases, the risk of a corresponding prompt may be calculated using a weight profile corresponding to the risk type of the prompt. For example, the system (10) may select a weight profile corresponding to the risk type information included in the semantic analysis result of the prompt from among a plurality of weight profiles, and calculate the risk of the prompt using this. That is, the system (10) may calculate the risk of the prompt by synthesizing the values ​​of each evaluation element based on the weight profile.

[0147] In some other embodiments, the weights for each evaluation element may be adjusted retrospectively based on the occurrence of false positives and / or false negatives. For example, the system (10) may receive the identification results of false positives and / or false negatives from the administrator's terminal (12) that checked the audit log, or may identify false positives and / or false negatives by analyzing the audit log. If a false positive is identified, the system (10) may lower the weight of one or more evaluation elements associated with the false positive. For example, the system (10) may lower the weight of an evaluation element whose contribution to the prompt risk calculation is above a threshold. On the other hand, if a false negative is identified, the system (10) may raise the weight of one or more evaluation elements associated with the false negative. In this case, the prompt risk can be calculated more accurately, and the occurrence of false positives and false negatives can be effectively reduced.

[0148] In some other embodiments, weights for each evaluation factor may be determined using a trained machine learning model. In this case, the weights for each evaluation factor can be determined precisely, which will be further explained shortly with reference to FIG. 10.

[0149] In some other embodiments, weights for each evaluation factor may be determined based on various combinations of the embodiments described above.

[0150] Below, with reference to FIG. 10, a method for determining weights for each evaluation element using a machine learning model (100) will be explained.

[0151] FIG. 10 is an exemplary drawing for explaining a weight determination method according to some embodiments of the present disclosure.

[0152] As illustrated in FIG. 10, the machine learning model (100) may be configured to receive the value of each evaluation element and output a weight for each evaluation element. As described above, the value of each evaluation element may be preprocessed into an appropriate form and input into the machine learning model (100).

[0153] The machine learning model (100) can be implemented based on various types / forms of models. For example, the machine learning model (100) may be implemented based on a neural network or based on a traditional machine learning model.

[0154] The machine learning model (100) can be trained using a supervised learning method. For example, the system (10) can train the machine learning model (100) using a training set labeled with whether a security risk has occurred or the correct risk level. Specifically, the system (10) can input training samples for a specific prompt into the machine learning model (100) to predict weights for each evaluation factor and calculate the risk level of the corresponding prompt based on the predicted weights. Then, the system (10) can update the parameters of the machine learning model (100) based on a loss representing the difference between the calculated risk level and the correct risk level. The system (10) can repeat the above process for training samples of other prompts.

[0155] A trained machine learning model (100) can be used to calculate the risk level of a prompt entered by a user. For example, the system (10) can determine the weights for each evaluation element by inputting the evaluation results into the trained machine learning model (100) and use this to calculate the risk level of the prompt.

[0156] Up to now, a weight determination method according to some embodiments of the present disclosure has been described with reference to FIG. 10. As described above, by using a machine learning model (100), weights for each evaluation factor can be accurately determined, and as a result, the risk level of the prompt can be calculated more precisely.

[0157] Referring again to Fig. 5, the explanation will be provided.

[0158] In step S55, a measure based on the risk level is determined based on a predefined security policy, and the corresponding measure is performed on the prompt. For example, the security policy may include a set of rules for determining a measure (i.e., security measure) corresponding to the risk level (e.g., risk level, risk range, etc.), and the system (10) may determine the measure to be applied to the prompt by referring to this set of rules. Examples of measures include, but are not limited to, allowing the prompt entirely, allowing it partially, reconfiguring, and blocking transmission (i.e., blocking external transmission). The partial allow measure may include de-identification processing such as masking or removal.

[0159] In some cases, a de-identification method (e.g., masking, removal, generalization, pseudonymization, etc.) may be determined based on security policies and / or risk levels as a component of the measure.

[0160] To facilitate better understanding, the detailed process of step S55 will be explained with reference to FIGS. 11 to 13.

[0161] FIG. 11 is an exemplary flowchart illustrating the process of determining and executing risk-based measures according to some embodiments of the present disclosure. However, this is merely an exemplary embodiment for achieving the purpose of the present disclosure, and it is understood that some steps may be added or deleted as necessary. FIG. 11 illustrates a case where rules are defined so that measures to be applied to a prompt are determined in correspondence with risk ranges.

[0162] In the following description, for the clarity of the present disclosure, different threshold values ​​will be distinguished and referred to using modifiers such as 'first' and 'second'.

[0163] As illustrated in FIG. 11, first, it is determined whether the risk level of the prompt is less than a first threshold (TH1) (S111). If it is less than the first threshold (TH1), step S112 is performed, and if not, step S113 may be performed.

[0164] In step S112, all actions to be applied to the prompt are determined to be allowed actions, and the corresponding actions are performed. For example, the system (10) can set the prompt to a security verification state without any separate modification to the prompt, and since the prompt in the security verification state has undergone a security processing process, it can be considered a security prompt.

[0165] In step S113, it is determined whether the risk level of the prompt is less than a second threshold (TH2). Here, the second threshold (TH2) may be set to a value greater than the first threshold (TH1) and less than the third threshold (TH3). If it is less than the second threshold (TH2), step S114 is performed, and if not, step S115 is performed.

[0166] In step S114, the action to be applied to the prompt is determined to be a partial allowance action, and de-identification, such as masking or removal, is performed on a part of the prompt (i.e., the part that is not allowed). Here, the target for de-identification may be, for example, a risk section derived through contextual semantic analysis of the prompt, but is not limited thereto. The target for de-identification may be determined based on static inspection results and / or sensitive entity detection results, or may be determined as an extended risk section reflecting these results.

[0167] For greater convenience of understanding, step S114 will be further explained with reference to Fig. 12.

[0168] FIG. 12 illustrates a process for performing a partial acceptance measure according to some embodiments of the present disclosure.

[0169] As illustrated in FIG. 12, risk sections (122 to 124) can be identified within the prompt (121) through the language model (60), and a security prompt (125) can be generated by de-identifying each of the risk sections (122 to 124).

[0170] FIG. 12 illustrates an example in which sensitive entities such as username (122), resident registration number (123), and account number (124) are identified as risk sections by a language model (60), and the de-identification method is masking, but the scope of the present disclosure is not limited thereto.

[0171] Up to this point, the process of performing partial allowance measures according to some embodiments of the present disclosure has been described with reference to FIG. 12. As described above, by performing de-identification on risk sections identified based on the contextual meaning of the prompt, security risks such as indirect information leakage and prompt injection can be effectively identified.

[0172] Meanwhile, in some embodiments, entity-level de-identification may be performed on risk sections within the prompt. For example, even if a risk section is identified at the sentence or paragraph level, de-identification may be performed only on entities within that risk section. Additionally, a security prompt may be generated to include abstract information types (e.g., personal identification information, username, etc.) for at least some entities. For example, an abstract information type may be added to the de-identification result of a specific entity (e.g., masking result, removal result, etc.). In such cases, the overall context of the prompt is maintained even in the security prompt, so the problem of reduced user convenience due to security processing can be effectively mitigated.

[0173] Referring again to Fig. 11, the explanation will be provided.

[0174] In step S115, it is determined whether the risk level of the prompt is less than the third threshold (TH3). If it is less than the third threshold (TH3), step S116 is performed, and if not, step S117 may be performed.

[0175] In step S116, the action to be applied to the prompt is determined to be a reconstructive action, and the corresponding action is performed on the prompt. Here, a reconstructive action refers to an action that generates a secure prompt by reconstructing (or modifying) the representation of the prompt through substitution-based de-identification, etc. A reconstructive action can be understood as being distinguished from a partial allowance action in that it performs modification at the representation level rather than masking or removing the unallowed parts.

[0176] For greater convenience of understanding, step S116 will be further explained with reference to Fig. 13.

[0177] FIG. 13 illustrates a process for performing prompt reconfiguration measures according to some embodiments of the present disclosure.

[0178] As illustrated in FIG. 13, sensitive entities (133, 134) within the prompt (131) can first be detected. Specific methods for detecting sensitive entities (e.g., 133) can be designed in various ways.

[0179] Next, a security prompt (135) can be generated by performing substitution-based de-identification for each sensitive entity (133, 134). For example, the system (10) can perform substitution-based de-identification, such as generalization or pseudonymization, for each sensitive entity (133, 134). In this case, the same sensitive entity may be substituted with the same alternative representation to preserve context.

[0180] In some embodiments, an abstract information type (e.g., personal identification information, username, etc.) corresponding to at least some sensitive entities (e.g., 133) may be included in the security prompt (135). For example, the sensitive entity (e.g., 133) may be replaced with the abstract information type, or the abstract information type may be added to the de-identification result of the sensitive entity (e.g., 133) (e.g., pseudonymization result, generalization result, etc.). In such cases, the overall context of the prompt is maintained in the security prompt, thereby effectively mitigating the problem of reduced user convenience due to security processing. FIG. 13 illustrates an example in which pseudonymization is performed using an identifier (e.g., User_A, Project_B, etc.) based on the abstract information type to preserve context.

[0181] Additionally, in some embodiments, a reconstruction measure may be performed using a language model (60). For example, the system (10) may generate a security prompt (e.g., 135) by providing the language model (60) with a prompt (121), a sensitive entity detection result and / or a semantic analysis result (e.g., information regarding a risk zone, risk type, risk assessment indicator, task type, recommended action, etc.). In this case, the set of instructions provided to the language model (60) may include, but is not limited to, instructions to determine the de-identification strength and target by considering the sensitive entity detection result and / or semantic analysis result while maintaining the overall context of the prompt (132), and to perform de-identification and prompt reconstruction using a de-identification method corresponding to the determined strength. According to these embodiments, a reconstruction measure may be performed in which the degree of context loss and the security risk level of the prompt (132) are considered in balance.

[0182] Up to this point, the process of performing prompt reconstruction measures according to some embodiments of the present disclosure has been described with reference to FIG. 13. As described above, by performing reconstruction measures through substitution-based de-identification for prompts with relatively high risk, security risks can be effectively reduced.

[0183] Referring again to Fig. 11, the explanation will be provided.

[0184] In step S117, the action to be applied to the prompt is determined to be a transmission blocking action, and the action is performed. For example, the system (10, e.g., policy enforcement unit (34), audit log management unit (35)) may block the transmission of the prompt and simultaneously record the relevant information in the audit log. If necessary, the system (10, e.g., reporting unit (36)) may also send an alert to the administrator terminal (12). The system (10) may also record the relevant information in the audit log when other actions other than the transmission blocking action are performed.

[0185] In some embodiments, the action to be applied to the prompt may be determined by considering the recommended action information included in the semantic analysis results. For example, the system (10) may apply the recommended action as is, or may apply the stricter of the action based on risk and the recommended action to the prompt.

[0186] Up to now, with reference to FIGS. 11 to 13, the process of determining and performing risk-based measures according to some embodiments of the present disclosure has been described.

[0187] Referring again to Fig. 5, the explanation will be provided.

[0188] In steps S56 and S57, it is determined whether the action performed is a transmission blocking action. If it is not a transmission blocking action, a security prompt is transmitted to an external AI service. For example, the system (10, e.g., relay unit (31)) can transmit the security prompt obtained as a result of performing the action to an external AI service (13). If attachment data exists, the system (10) can transmit the attachment data, on which security processing has been performed, together.

[0189] In some embodiments, as illustrated in FIG. 14, verification is performed on the security prompt (143), and if the verification result (145) is determined to be successful, the security prompt (143) may be transmitted externally. For example, the system (10) can verify the security risk and semantic distortion of the security prompt (143) by generating a verification prompt (141) based on the prompt (142), the security prompt (143), and a set of instructions (144), and providing it to the language model (60). Here, the set of instructions (144) may include instructions to verify the existence of a security risk for the security prompt (143) and instructions to verify semantic distortion by comparing the security prompt (143) with the prompt (142). Additionally, the verification prompt (141) may further include static inspection results, a description (or rule) of the verification criteria or verification method, etc. Next, the system (10) may take appropriate action based on the verification result (145) of the language model (60). The verification result (145) may include, for example, whether the verification was successful, the reason for the verification failure, parts related to the verification failure, etc., but is not limited thereto. If the verification result (145) is determined to be a failure, the system (10) may block the transmission of the security prompt (143) and provide the user with a notification (or warning) about the blocking. Alternatively, the system (10) may regenerate the security prompt (143). FIG. 14 illustrates an example in which the contextual semantic analysis of the prompt (142) and the verification of the security prompt (143) are performed by the same language model, but the scope of the present disclosure is not limited thereto.

[0190] Up to this point, the processing of outbound requests according to several embodiments of the present disclosure has been described with reference to FIGS. 5 through 14. As described above, by analyzing the contextual meaning of a prompt using a language model, even security risks that are difficult to identify by examining only the surface expression of the prompt, such as indirect information leakage and prompt injection, can be effectively detected. Accordingly, security risks that may occur in an environment using external AI services can be effectively controlled.

[0191] In addition, by utilizing static inspection results as context information, the accuracy of contextual semantic analysis of prompts can be improved.

[0192] In addition, the risk level of a prompt (or request) can be calculated based on semantic analysis results, static inspection results, etc., and measures based on the risk level can be performed according to predefined security policies. For example, measures such as full allowance, partial allowance, reconfiguration, and transmission blocking can be performed differentially on a prompt depending on the risk level. In such cases, security levels can be maintained while preventing excessive blocking, thereby ensuring both security and user convenience simultaneously.

[0193] In addition, by considering various evaluation factors such as user permissions and the type of task requested by the prompt, in addition to static inspection results and semantic analysis results, the risk level of the prompt can be calculated more accurately.

[0194] Furthermore, by performing anonymization at the level of risk sections or sensitive entities within the prompt, the leakage of sensitive information can be prevented while simultaneously minimizing the loss of context within the prompt. Accordingly, the issue of reduced user convenience regarding external AI services due to security processing can be effectively mitigated.

[0195] Below, the inbound response processing process will be described with reference to FIGS. 15 and FIGS. 16.

[0196] FIG. 15 is an exemplary flowchart illustrating an inbound response processing process according to some embodiments of the present disclosure. However, this is merely an exemplary embodiment for achieving the purpose of the present disclosure, and it is understood that some steps may be added or deleted as necessary.

[0197] As illustrated in FIG. 15, the processing process according to the embodiments may begin at step S151, which involves receiving a response to a security prompt from an external AI service. For example, the system (10, e.g., relay unit (31)) may receive a response to a previously transmitted security prompt from an external AI service (13).

[0198] In step S152, a static inspection is performed on the response. For example, the system (10, e.g., static inspection unit (32)) may perform an inspection of the format of the response using inspection rules based on file extensions, MIME types, magic numbers, etc. As another example, the system (10) may perform an inspection of malicious elements in the response using inspection rules based on attack patterns (e.g., script injection patterns, SQL injection patterns, etc.), malicious code patterns, malicious lists (e.g., malicious domain lists, malicious URL lists, etc.), policy violation patterns, etc. If necessary, the system (10) may perform additional isolation and / or inspection of inbound responses in the form of files or code by linking with an external malicious code detection engine (e.g., antivirus) or a sandbox system.

[0199] For this step S152, please refer further to the explanation of step S52.

[0200] In step S153, the contextual meaning of the response is analyzed through a language model based on the static test results. Step S153 will be further explained with reference to Fig. 16.

[0201] FIG. 16 is an exemplary drawing for illustrating the process of analyzing the contextual meaning of a response according to some embodiments of the present disclosure.

[0202] As illustrated in FIG. 16, the contextual meaning of the response (162) can be analyzed through the language model (60).

[0203] For example, the system (10, e.g., semantic analysis unit (33)) can generate an analysis prompt (161) based on a response (162), static inspection results (163), a set of instructions (164), etc., and provide it to a language model (60) to analyze the contextual meaning of the response (162). In this case, not only false information, biased expressions, and explicit policy violation expressions, but also indirect policy violation expressions or expressions that may re-expose sensitive information can be accurately identified.

[0204] The static inspection result (163) may include, for example, attack patterns, malicious URLs, malicious domains, policy violation expressions, prohibited keywords, etc. detected according to predefined inspection rules, but is not limited thereto. The static inspection result (163) can be used as context information when analyzing the contextual semantics of the response (162) to improve the analysis accuracy of the language model (60).

[0205] The set of instructions (164) may include, for example, instructions requesting the analysis of the contextual meaning of the response (162), instructions requesting reference to the static test result (163) during the semantic analysis process, etc.

[0206] In addition, the instruction set (164) and / or analysis prompt (161) may further include various instructions / information / data / conditions / rules / descriptions to improve the accuracy of the analysis result (165), for which refer to the description in FIG. 6.

[0207] For details regarding the items and contents of the analysis results (165), please refer to the descriptions in FIGS. 6 and FIGS. 7.

[0208] In some cases, the system (10) may analyze the contextual meaning of the response (162) without relying on the static test result (163).

[0209] Up to now, the process of analyzing the contextual meaning of a response according to some embodiments of the present disclosure has been described with reference to FIG. 16.

[0210] Referring again to Fig. 15, the explanation will be provided.

[0211] For this step S153, please refer further to the explanation of step S53.

[0212] In step S154, the risk of the response is calculated based on static inspection results and semantic analysis results. For example, the system (10, e.g., policy enforcement unit (34)) may calculate the risk of the response based on multiple evaluation factors including static inspection results, semantic analysis results, etc. For examples of evaluation factors, refer to Table 1 described above.

[0213] For reference, the harmfulness of inbound responses listed in Table 1 can be understood as an item determined by combining static inspection results and semantic analysis results, and in some cases, may be excluded when calculating response risk.

[0214] For this step S154, please refer further to the explanation in step S54.

[0215] In step S155, an action based on the risk level is determined based on a predefined security policy, and the corresponding action is performed on the response. Examples of actions performed include, but are not limited to, full allowance, partial allowance, reconfiguration, and providing alternative responses.

[0216] For example, if the risk level is below a first threshold, the system (10, e.g., policy enforcement unit (34)) may determine that the measure to be applied to the response is a fully permissible measure. Additionally, if the risk level is above the first threshold and below the second threshold (wherein the second threshold is greater than the first threshold and less than the third threshold), the system (10) may determine that the measure to be applied to the response is a partially permissible measure. Additionally, if the risk level is above the second threshold and below the third threshold, the system (10) may determine that the measure to be applied to the response is a reconfiguration measure. If the risk level is above the third threshold, the measure to be applied to the response may be determined as an alternative response provision measure.

[0217] Here, each threshold may be distinct from the threshold of FIG. 11 used to perform differential measures for the prompt, but is not limited thereto.

[0218] Similarly to the above, partial allowance measures may include processing to de-identify parts of the response, for which further reference is made to the description in FIG. 12.

[0219] Reconstruction measures refer to measures that generate a secure response by reconstructing the expression of the response, and the reconstruction of the expression of the response may be performed through substitution-based de-identification (see FIG. 13) or in other ways. For example, the system (10) may generate a secure response by modifying policy violation expressions using a language model (60), or it may modify policy violation expressions based on rules (e.g., deleting or replacing with policy compliant expressions).

[0220] The measure to provide an alternative response means a measure to generate a new response to replace the original response based on a security policy and to provide the new response as a security response. For example, the system (10) may block the original response and generate a new response containing a notice or general guide based on a security policy, or provide a new response generated based on a standard template to the user (e.g., generating a new response by reflecting only content that complies with the security policy in the standard template). When the original response is blocked, the system (10, e.g., audit log management unit (35), reporting unit (36)) may record relevant information in the audit log and send an alert to the administrator terminal (12). As described above, the system (10) may also record relevant information in the audit log when other measures are performed.

[0221] For this step S155, please refer further to the explanation of step S55.

[0222] In step S156, a security response obtained as a result of performing the action is transmitted to a user terminal. For example, the system (10, e.g., relay unit (31)) can transmit a security response (e.g., original response of security verification status, partial allowance response, reconfiguration response, alternative response, etc.) to the user terminal (11).

[0223] In some embodiments, as illustrated in FIG. 17, a restoration response (175) may be generated using de-identification processing information (173) and transmitted to a user terminal (11). Here, the de-identification processing information (173) may include mapping information before and after de-identification (e.g., correspondence between original text segments and de-identification results) of a prompt (e.g., partial allowance, reconfiguration action related prompt) corresponding to a security response (172), but is not limited thereto. For example, the system (10) may generate a restoration prompt (171) based on the security response (172), de-identification processing information (173), a set of instructions (174), etc., and provide it to a language model (60) to generate a restoration response (175). The language model (60) can restore at least a portion of the de-identified information based on the de-identification processing information (173), and as a result, the problem of reduced user convenience due to security processing can be further mitigated. In some cases, the system (10) may additionally provide the original text prompt to the language model (60) to improve restoration accuracy.

[0224] For this step S156, please refer further to the explanation in step S57.

[0225] Up to this point, the inbound response processing process according to some embodiments of the present disclosure has been described with reference to FIGS. 15 to 17. As described above, by applying risk analysis and differential measures to inbound responses in a manner similar to that of outbound requests, integrated security management throughout the entire process of using external AI services becomes possible, and accordingly, an AI security governance system capable of applying and managing consistent security policies in an organizational environment can be established.

[0226] Hereinafter, with reference to FIG. 18, an exemplary computing device (180) capable of implementing the system (10) described above will be described.

[0227] FIG. 18 is an exemplary hardware configuration diagram showing a computing device (180).

[0228] As illustrated in FIG. 18, a computing device (180) may include one or more processors (181), a bus (183), a communication interface (184), a memory (182) for loading a computer program (186) executed by the processor (181), and a storage (185) for storing the computer program (186). However, FIG. 18 illustrates only the components related to the embodiments of the present disclosure. Therefore, a person skilled in the art to which the present disclosure belongs will understand that other general-purpose components may be included in addition to the components (181 to 186) illustrated in FIG. 18. That is, the computing device (180) may include various additional components in addition to the components (181 to 186) illustrated in FIG. 18. Furthermore, depending on the case, the computing device (180) may be configured in a form in which some of the components (181 to 186) illustrated in FIG. 18 are omitted. Below, each component of the computing device (180) is described.

[0229] The processor (181) can control the overall operation of each component of the computing device (180). The processor (181) may be configured to include at least one of a CPU (Central Processing Unit), MPU (Micro Processor Unit), MCU (Micro Controller Unit), GPU (Graphic Processing Unit), NPU (Neural Processing Unit), TPU (Tensor Processing Unit), VPU (Vision Processing Unit), APU (Accelerated Processing Unit), or any other type of processor well known in the art of the present disclosure. Additionally, the processor (181) may perform operations for at least one application or program to execute specific operations / steps / methods. The computing device (180) may have one or more processors.

[0230] Next, the memory (182) may store various data, commands and / or information. The memory (182) may load a computer program (186) from storage (185) to execute specific operations / steps / methods. The memory (182) may be implemented as volatile memory such as RAM, but the technical scope of the present disclosure is not limited thereto.

[0231] The bus (183) can provide communication functions between components of the computing device (180). The bus (183) can be implemented as various types of buses, such as an address bus, a data bus, and a control bus.

[0232] The communication interface (184) can support wired and wireless internet communication of the computing device (180). Additionally, the communication interface (184) may support various communication methods other than internet communication. To this end, the communication interface (184) may be configured to include a communication module well known in the art of the present disclosure.

[0233] Storage (185) may store one or more computer programs (186) non-temporarily. Storage (185) may be configured to include non-volatile memory such as ROM (Read Only Memory), EPROM (Erasable Programmable ROM), EEPROM (Electrically Erasable Programmable ROM), flash memory, a hard disk, a removable disk, or any form of computer-readable recording medium well known in the art to which this disclosure belongs.

[0234] A computer program (186) may include instructions that cause a processor (181) to perform specific operations / steps / methods when loaded into memory (182). That is, the processor (181) can perform specific operations / steps / methods by executing the loaded instructions.

[0235] For example, a computer program (186) may include instructions to perform actions such as obtaining a user prompt for an external AI service, deriving an analysis result related to security risks by analyzing the contextual meaning of the prompt through a language model, calculating the risk level of the prompt based on the analysis result, determining an action based on the risk level based on a predefined security policy, generating a security prompt by performing an action on the prompt, and transmitting the security prompt to an external AI service.

[0236] As another example, the computer program (186) may include instructions to perform the following actions: receiving a response to a user prompt from an external AI service; deriving an analysis result related to security risks by analyzing the contextual meaning of the response through a language model; calculating the risk level of the response based on the analysis result; determining a measure based on the risk level based on a security policy; generating a security response by performing a measure on the response; and transmitting the security response to the user's terminal.

[0237] As another example, a computer program (186) may include instructions to perform at least some of the operations / operations / methods described with reference to FIGS. 1 through 17.

[0238] As illustrated, a system (10) according to some embodiments of the present disclosure can be implemented through a computing device (180).

[0239] Meanwhile, in some embodiments, the computing device (180) illustrated in FIG. 18 may refer to a virtual machine implemented based on cloud technology. For example, the computing device (180) may be a virtual machine running on one or more physical servers included in a server farm. In this case, at least some of the processor (181), memory (182), and storage (185) illustrated in FIG. 18 may be virtual hardware, and the communication interface (184) may also be implemented as a virtualized networking element such as a virtual switch.

[0240] Up to now, with reference to FIG. 18, an exemplary computing device (180) capable of implementing a system (10) according to some embodiments of the present disclosure has been described.

[0241] Various embodiments of the present disclosure and effects according to those embodiments have been described with reference to FIGS. 1 through 18. The effects according to the technical concept of the present disclosure are not limited to those mentioned above, and other unmentioned effects will be clearly understood by a person skilled in the art from the description below.

[0242] Furthermore, just because the above embodiments describe a plurality of components being combined into one or operating in combination, the technical concept of the present disclosure is not necessarily limited to these embodiments. That is, within the scope of the purpose of the technical concept of the present disclosure, all such components may be selectively combined into one or more combinations to operate.

[0243] The technical concept of the present disclosure described above may be implemented as computer-readable code on a computer-readable recording medium (e.g., a non-transitory recording medium). A computer program recorded on a computer-readable recording medium may be transmitted to another computing device via a network such as the Internet and installed on said computing device, thereby being used on said computing device.

[0244] Although the operations are depicted in a specific order in the drawings, it should not be understood that the operations must be executed in the specific order depicted or in sequential order, or that all depicted operations must be executed to obtain the desired result. In certain situations, multitasking and parallel processing may be advantageous.

[0245] Although various embodiments of the present disclosure have been described above with reference to the attached drawings, those skilled in the art will understand that the technical concept of the present disclosure may be implemented in other specific forms without altering the technical concept or essential features thereof. Therefore, the embodiments described above should be understood as illustrative in all respects and not restrictive. The scope of protection of the present disclosure shall be interpreted by the claims below, and all technical concepts within the equivalent scope shall be interpreted as being included within the scope of rights of the technical concept defined by the present disclosure. Explanation of the symbols

[0246] 10: AI Security Governance System 11-1, 11-2, 11-N: User terminal 12: Administrator Terminal 13-1, 13-2, 13-K: External AI Services 21: Internal network 22: External network 31: Broadcasting Department 32 Static Inspection Unit 33: Semantic Analysis Department 34: Policy Executive 35: Audit Log Management Department 36: Report Department 60: Language Model 80, 100: Machine learning model 180: Computing device 181: Processor 182: Memory 183: Bus 184: Communication Interface 185: Storage 186: Computer Program

Claims

Claim 1 A method for providing AI security governance, comprising: a step of obtaining a user prompt for an external AI (Artificial Intelligence) service by performing by at least one processor; a step of deriving a first analysis result related to a security risk of the prompt by analyzing the contextual meaning of the prompt through a language model; a step of generating an expected response to the prompt through the language model; a step of calculating a risk level of the prompt based on the first analysis result and a second analysis result related to a security risk of the expected response; a step of determining a measure according to the risk level based on a predefined security policy; a step of generating a security prompt by performing the measure on the prompt; and a step of transmitting the security prompt to the external AI service, wherein the step of transmitting to the external AI service comprises: a step of generating a verification prompt based on the prompt, the security prompt, and a set of instructions, wherein the set of instructions includes instructions that verify whether the meaning of the security prompt is distorted by comparison with the prompt; a step of verifying the security prompt by providing the verification prompt to the language model; and a step of transmitting the security prompt to the external AI service based on the result of the verification. Claim 2 A method for providing AI security governance according to claim 1, wherein the prompt is received from the user's terminal located in the internal network, and the at least one processor is provided in a proxy server positioned between the internal network and the external AI service. Claim 3 A method for providing AI security governance according to claim 1, wherein the step of deriving the first analysis result comprises: a step of performing a static inspection on the prompt based on a preset inspection rule; a step of generating an analysis prompt using the result of the static inspection as context information of the prompt; and a step of providing the analysis prompt to the language model to derive the first analysis result. Claim 4 A method for providing AI security governance according to claim 1, wherein the step of deriving the first analysis result comprises: a step of obtaining attachment data associated with the prompt; a step of generating an analysis prompt for analyzing the contextual meaning of the prompt; and a step of providing the attachment data together with the analysis prompt to the language model to derive the first analysis result. Claim 5 A method for providing AI security governance according to claim 1, wherein the first analysis result is expressed in a structured format including a plurality of items, and the plurality of items include risk types and risk assessment indicators of the prompt. Claim 6 A method for providing AI security governance according to claim 1, wherein the step of calculating the risk of the prompt includes the step of calculating the risk based on a plurality of evaluation factors including the first analysis result and the second analysis result, and the evaluation factors further include the user's authority and the type of work requested by the prompt. Claim 7 delete Claim 8 A method for providing AI security governance according to claim 1, wherein the first analysis result includes information regarding the risk type of the prompt, and the step of calculating the risk level of the prompt comprises: a step of selecting a weight profile corresponding to the risk type among a plurality of predefined weight profiles - wherein the weight profile includes weights for each of a plurality of evaluation factors, and the evaluation factors include the first analysis result and the second analysis result -; and a step of calculating the risk level by synthesizing the values ​​of each of the evaluation factors based on the weights. Claim 9 A method for providing AI security governance according to claim 1, wherein the step of calculating the risk of the prompt comprises: a step of calculating the risk by synthesizing the values ​​of a plurality of evaluation factors, including the first analysis result and the second analysis result, based on weights; wherein information related to the determined measure is recorded in an audit log; and further comprising: a step of lowering the weight of one or more evaluation factors related to the false positive when a false positive for the prompt is identified based on the audit log; and a step of raising the weight of one or more evaluation factors related to the false negative when a false negative for the prompt is identified based on the audit log. Claim 10 A method for providing AI security governance according to claim 1, wherein the step of determining measures according to the risk level includes: a step of determining the measure to be applied to the prompt as a fully allowed measure when the risk level is less than a first threshold; and a step of determining the measure to be applied to the prompt as a partially allowed measure when the risk level is greater than or equal to the first threshold and less than a second threshold, wherein the second threshold is a value greater than the first threshold, and the partially allowed measure includes processing to de-identify a part of the prompt. Claim 11 A method for providing AI security governance according to claim 10, wherein the step of determining measures based on the risk level further includes the step of determining the measure to be applied to the prompt as a reconfiguration measure when the risk level is greater than or equal to the second threshold and less than the third threshold, wherein the third threshold is a value greater than the second threshold, and the reconfiguration measure is a measure to generate the security prompt by reconfiguring the prompt through substitution-based de-identification, and when the risk level is greater than or equal to the third threshold, the measure to be applied to the prompt is determined as a transmission blocking measure. Claim 12 A method for providing AI security governance according to claim 1, wherein the first analysis result includes information regarding one or more text sections related to the security risk of the prompt, the determined measure is a partial permission measure, and the step of generating the security prompt includes the step of generating the security prompt by de-identifying the one or more text sections. Claim 13 A method for providing AI security governance according to claim 1, wherein the step of generating the security prompt comprises: a step of identifying one or more sensitive entities in the prompt; and a step of generating the security prompt by performing entity-unit de-identification on the prompt based on the result of the identification, wherein the generated security prompt includes an abstracted information type corresponding to at least some of the one or more sensitive entities. Claim 14 A method for providing AI security governance according to claim 1, wherein the set of instructions further includes instructions for verifying the existence of a security risk for the security prompt. Claim 15 A method for providing AI security governance according to claim 1, wherein the risk level is a first risk level; receiving a response to the security prompt from the external AI service; deriving a third analysis result related to the security risk of the response by analyzing the contextual meaning of the response through the language model, wherein the security risk of the response includes a risk due to a violation of the security policy; calculating a second risk level of the response based on the third analysis result; determining a measure according to the second risk level based on the security policy; generating a security response by performing the measure determined according to the second risk level on the response; and transmitting the security response to the user's terminal. Claim 16 In claim 15, the step of determining measures according to the second risk level comprises: determining the measure to be applied to the response as a fully permissible measure when the second risk level is less than a first threshold; determining the measure to be applied to the response as a partially permissible measure when the second risk level is greater than or equal to the first threshold and less than a second threshold; determining the measure to be applied to the response as a reconfiguration measure when the second risk level is greater than or equal to the second threshold and less than a third threshold; and determining the measure to be applied to the response as an alternative response provision measure when the second risk level is greater than or equal to the third threshold, wherein the second threshold is a value smaller than the third threshold and larger than the first threshold, the partially permissible measure includes processing to de-identify a part of the response, the reconfiguration measure is a measure to generate the security response by reconfiguring the representation of the response, and the alternative response provision measure is a measure to generate a new response to replace the response based on the security policy and provide the new response as the security response. Claim 17 In claim 15, the security prompt is generated through de-identification processing, and the step of transmitting to the user's terminal comprises: a step of generating a restoration response by providing information regarding the de-identification processing and the security response to the language model; and a step of transmitting the restoration response to the terminal. Claim 18 One or more processors; and memory for storing a computer program executed by said one or more processors, said computer program comprises: an action of obtaining a user prompt for an external AI (Artificial Intelligence) service; an action of deriving a first analysis result related to the security risk of said prompt by analyzing the contextual meaning of said prompt through a language model; an action of generating an expected response to said prompt through said language model; an action of calculating the risk level of said prompt based on the first analysis result and a second analysis result related to the security risk of said expected response; an action of determining an action according to said risk level based on a predefined security policy; an action of generating a security prompt by performing said action on said prompt; and instructions for an action of transmitting said security prompt to said external AI service, said action, said action; and the action of transmitting to said external AI service comprises: an action of generating a verification prompt based on said prompt, said security prompt, and a set of instructions - said set of instructions includes instructions that verify whether the meaning of said security prompt is distorted by comparison with said prompt -; and an action of providing said verification prompt to said language model to verify said security prompt. An AI security governance system comprising the operation of transmitting the security prompt to the external AI service based on the result of the verification above. Claim 19 In claim 18, the above risk level is a first risk level, and the computer program further comprises instructions for: receiving a response to the security prompt from the external AI service; deriving a third analysis result related to the security risk of the response by analyzing the contextual meaning of the response through the language model - the security risk of the response includes a risk due to a violation of the security policy -; calculating a second risk level of the response based on the third analysis result; determining a measure according to the second risk level based on the security policy; generating a security response by performing the measure determined according to the second risk level on the response; and transmitting the security response to the user's terminal. Claim 20 A computer program stored on a non-transitory computer-readable recording medium in combination with a computer to execute a method according to any one of claims 1 through 6 and claims 8 through 16.

Citation Information

Patent Citations

  • Sentence creating system using confidentiality protection prompt

    JP2025073908A

  • Persona-Security Graph-based dynamic security management system and method thereof

    KR102910983B1

  • Dynamic Filtering and Response Method and Apparatus for Preventing Sensitive Data Leakage

    KR102944456B1

  • Security and privacy inspection of bidirectional generative artificial intelligence traffic using a forward proxy

    US12273392B1

  • Validating autonomous artificial intelligence (AI) agents using generative ai

    US20260017525A1