Method and device for preventing semantic data leakage

By combining a generative pre-trained transformer (GPT) and a data leakage prevention (DLP) system, the semantic data leakage of the target context is dynamically customized, which solves the problem of inaccurate detection in the prior art and achieves efficient and reliable data security enhancement.

CN122070542APending Publication Date: 2026-05-19HUAWEI CLOUD COMPUTING TECHNOLOGIES CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HUAWEI CLOUD COMPUTING TECHNOLOGIES CO LTD
Filing Date
2023-10-25
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

Existing DLP systems struggle to accurately detect semantic data leakage when faced with advanced language technologies, leading to the hiding and leakage of sensitive information. Current technologies suffer from inaccurate detection and low efficiency.

Method used

Generative pre-trained transformers (GPT) are used to obtain semantic interpretations of the input text. Combined with a data leakage prevention (DLP) system, hidden data leakage is detected by using a set of words and phrases of interest. The target context is dynamically customized to enhance the detection capability.

Benefits of technology

It improves the accuracy of detecting hidden data leaks, reduces false alarm rates, lowers computational overhead, enhances overall data security, and reduces the risk of semantic leakage attacks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122070542A_ABST
    Figure CN122070542A_ABST
Patent Text Reader

Abstract

In order to prevent semantic data leakage, one or more semantic interpretations of an input text are obtained using a generative pre-trained translator (GPT), where the GPT is to interpret the input text with reference to a target context; and determining whether the input text includes hidden data leaks by checking the semantic interpretation of the input text using a data leak prevention (DLP) system, where the DLP system is used to detect data leaks based on a set of words and phrases of interest. Accordingly, it is possible to prevent hidden data leakage when sensitive information is hidden in text (e.g., by a high-level language technology).
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention generally relates to the field of data security, and more specifically, to methods and apparatus for preventing the leakage of semantic data. Background Technology

[0002] In today's digitally interconnected world, protecting sensitive data has become a top concern for businesses of all sizes. Data leak prevention (DLP) systems act as sentinels, tirelessly scanning and protecting outgoing data traffic to detect sensitive information such as trade secrets, proprietary intellectual property, product specifications, financial records, and personal information of employees and customers. Furthermore, the primary purpose of DLP systems is to identify and prevent unauthorized transmission of sensitive information from an organization's network to external destinations.

[0003] Currently, the text matching techniques used in existing DLP systems can be fooled by attackers, who can employ advanced textual techniques such as rewriting, allegory, and metaphor to hide sensitive information in seemingly neutral text. However, these attempts to prevent data leakage have failed due to various reasons, including inaccurate and inefficient text detection, and the alteration of the precise meaning of data through text shuffling or trimming in formal documents. For example, Trojan horses can successfully bypass DLP checks and leak sensitive information to third parties through semantic recoding. Therefore, there is a technical challenge in detecting semantic leaks to ensure overall data security and confidentiality.

[0004] Therefore, in light of the above discussion, it is necessary to overcome the drawbacks associated with conventional methods and devices for preventing semantic data leakage. Summary of the Invention

[0005] This invention provides a method and apparatus for preventing semantic data leakage. It offers a solution to the existing problem of how to detect semantic leakage to ensure overall data security and confidentiality. The object of this invention is to provide a solution that at least partially solves the problems encountered in the prior art, and to provide an improved method and apparatus for preventing semantic data leakage.

[0006] One or more objects of the present invention are achieved by the solutions provided in the appended independent claims. Advantageous embodiments of the invention are further defined in the dependent claims.

[0007] In one aspect, the present invention provides a method for preventing semantic data leakage. The method includes: obtaining one or more semantic interpretations of input text using a generative pre-trained transformer (GPT), wherein the GPT is used to interpret the input text with reference to a target context; and examining the semantic interpretations of the input text using a data leak prevention (DLP) system to determine whether the input text contains hidden data leakage, wherein the DLP system is used to detect data leakage based on a set of words and phrases of interest.

[0008] Advantageously, methods for preventing semantic data leakage are used to prevent data breaches when any sensitive information (e.g., through advanced language techniques) is hidden within the text. In these methods, interpreting the input text based on the target context enhances the ability to detect hidden data leakage, which traditional keyword matching might fail to detect. Furthermore, by examining the semantic interpretation generated by GPT, understanding the meaning and context of the text significantly reduces false positives and improves the accuracy of detecting hidden data leakage, while simultaneously reducing computational overhead and efficiently and reliably detecting data leakage. Therefore, these methods enhance overall data security and reduce the risk of semantic leakage attacks.

[0009] In one implementation, the method further includes: defining the target context for semantic interpretation based on user input and / or the set of words and phrases of interest; providing a context domain extension to the GPT according to the target context; and updating the context domain extension of the GPT when the target context is updated.

[0010] By defining the target context for semantic interpretation and updating the context domain extension of GPT, semantic interpretation can be dynamically customized based on specific user needs for enhanced data security requirements, thereby achieving accurate, reliable, and context-aware detection of hidden data leaks.

[0011] In another implementation, the method further includes: determining that the DLP system has not detected data leakage in the input text before obtaining the semantic interpretation of the input text.

[0012] In this implementation, detecting data breaches can minimize unnecessary processing and computational overhead for GPT and optimize resource utilization, ensuring that GPT's computational resources are only used on texts that may contain hidden data breaches.

[0013] In another implementation, the method further includes: once the DLP system detects the data leak in any semantic interpretation of the semantic interpretation, reporting to the user and / or control authority that the input text includes the hidden data leak.

[0014] Advantageously, reporting detected hidden data breaches to users and / or control authorities can enhance data security incident response, enabling timely action to mitigate potential risks and protect sensitive information, thereby reducing the impact of data breaches and enhancing overall data security.

[0015] In another aspect, the present invention provides an apparatus for preventing semantic data leakage, the apparatus comprising: a context repository for storing target context for semantic interpretation; and a detection engine for: obtaining one or more semantic interpretations of the input text by loading the input text into a generative pre-trained transformer (GPT), wherein the GPT is used to interpret the input text with reference to the target context; and determining whether the input text contains hidden data leakage by examining the semantic interpretations of the input text using a data leak prevention (DLP) system, wherein the DLP system is used to detect data leakage based on a set of words and phrases of interest.

[0016] The device achieves all the advantages and technical effects of the method described in this invention.

[0017] It should be understood that all of the aforementioned implementation methods can be combined together.

[0018] It should be noted that all devices, elements, circuits, units, and apparatuses described in this application can be implemented in software or hardware elements or any combination thereof. All steps performed by the various entities described in this application, and the functions to be performed by the various entities described, are intended to refer to the respective entities performing the respective steps and functions. Although the specific functions or steps to be performed by external entities are not reflected in the detailed description of the specific elements of the entities performing the specific steps or functions in the following description of specific embodiments, those skilled in the art will understand that these methods and functions can be implemented by the respective software or hardware elements or any combination thereof. It should be understood that the features of the present invention are readily combined in various combinations without departing from the scope of the invention as defined by the appended claims.

[0019] Other aspects, advantages, features and objects of the invention will become apparent from the accompanying drawings and the detailed description of illustrative implementations as explained in conjunction with the following appended claims. Attached Figure Description

[0020] The above-described invention and the following detailed description of illustrative embodiments can be better understood when read in conjunction with the accompanying drawings. Exemplary structures of the invention are shown in the drawings to illustrate the invention. However, the invention is not limited to the specific methods and tools disclosed herein. Furthermore, those skilled in the art will understand that the drawings are not drawn to scale. Where possible, similar elements are represented by the same numbers.

[0021] Embodiments of the present invention are described below by way of example only with reference to the following accompanying drawings, in which: Figure 1 This is a flowchart of a method for preventing semantic data leakage provided in an embodiment of the present invention; Figure 2 This is a block diagram of a device for preventing semantic data leakage provided in an embodiment of the present invention; Figure 3 This is an illustration provided by an embodiment of the present invention, depicting an exemplary scenario for preventing semantic data leakage; Figure 4 This is an illustration of an exemplary scenario depicting an apparatus for preventing semantic data leakage, provided by an embodiment of the present invention; Figure 5 This is an exemplary schematic diagram illustrating the execution sequence for preventing semantic leakage, provided by an embodiment of the present invention.

[0022] In the accompanying diagrams, underlined numbers indicate the item in which the underlined number appears or the item adjacent to the underlined number. Ununderlined numbers relate to the item identified by the line that associates the ununderlined number with the item. When a number is ununderlined and has an associated arrow, the ununderlined number is used to identify the general item that the arrow points to. Detailed Implementation

[0023] The following detailed description illustrates embodiments of the present invention and ways in which these embodiments can be implemented. While some modes of implementing the invention have been disclosed, those skilled in the art will recognize that other embodiments for implementing or practicing the invention may also exist.

[0024] Figure 1 This is a flowchart of a method for preventing semantic data leakage provided in an embodiment of the present invention. (Reference) Figure 1 The diagram shows a flowchart of method 100 used in the apparatus. Method 100 includes steps 102 to 106.

[0025] A method 100 is provided to prevent semantic data leakage. Method 100 is used to reduce the risk of sensitive information being leaked due to semantic operations. Therefore, the overall data security and authenticity of the relevant data are improved.

[0026] During operation, at step 102, method 100 includes: using a generative pre-trained transformer (GPT) to obtain one or more semantic interpretations of the input text, wherein the GPT is used to interpret the input text with reference to a target context. In other words, the GPT is used to analyze and generate one or more semantic interpretations of the input text. Furthermore, the generated one or more semantic interpretations are generated based on an understanding of the text and the target context. The target context can be a predefined target context that helps the GPT to understand the text more deeply by associating the corresponding text with a specific domain or topic.

[0027] Furthermore, at step 104, method 100 includes: determining whether the input text contains a hidden data leak by examining the semantic interpretation of the input text using a data leak prevention (DLP) system, wherein the DLP system is used to detect data leaks based on a set of words and phrases of interest. The DLP system utilizes the set of words and phrases of interest, which can be used to examine the semantic interpretation of the input text to prevent data leaks. Furthermore, the corresponding words and phrases of interest indicate sensitive information or data leaks. Therefore, if any word or phrase from the set of words and phrases of interest is detected in the semantic interpretation, then a data leak is detected in this case. Thus, method 100 is used to assess whether the input text contains a hidden data leak to reduce the risk of any potential data leakage and data threats, where the hidden data leak involves sensitive information that may be intentionally or unintentionally hidden in the text.

[0028] According to one embodiment, method 100 further includes: defining a target context for semantic interpretation based on user input and / or a set of words and phrases of interest; providing a context domain extension to the GPT based on the target context; and updating the context domain extension of the GPT when the target context is updated. The target context for semantic interpretation based on the set of words and phrases of interest includes a meaningful context description (e.g., product version, CPU product name, battery level, etc.) and a list of keywords associated with the corresponding target context. For example, the target context includes: Context descriptor { <id>, <Contextual description>, <Set the interpreter's context directive>, Special keywords and phrases for interpreters and DLP } In one implementation, method 100 further includes defining a target context for semantic interpretation based on user input. In another implementation, method 100 further includes defining a target context for semantic interpretation based on a set of words and phrases of interest. In yet another implementation, method 100 further includes defining a target context for semantic interpretation based on both user input and a set of words and phrases of interest. Furthermore, when the target context is updated based on the topic and domain, the context domain extension is also updated. For example, providing a context domain extension to the GPT based on the target context (e.g., cellular, healthcare, etc.) allows for correct text transformation using domain-specific words, phrases, etc. Therefore, by using a target context for semantic interpretation and updating the context domain extension of the GPT, semantic interpretation can be dynamically customized based on specific user needs for enhanced data security requirements, thereby achieving accurate, reliable, and context-aware detection of hidden data breaches.

[0029] According to one embodiment, method 100 further includes: determining that the DLP system has not detected data leakage in the input text before obtaining a semantic interpretation of the input text. First, a set of words and phrases of interest, such as a list of sensitive words, terms, phrases, etc., is prepared for the DLP system. Then, the GPT is pre-trained by installing relevant context domain extensions to interpret the input text with reference to target contexts. Furthermore, the list of sensitive contexts (i.e., potential targets of data leakage) is updated synchronously. In this case, the input text is transformed using one of the sensitive contexts in the list, and the transformed input text is evaluated using conventional DLP techniques. Then, if a data leakage is detected, the evaluation of the data leakage is stopped, and the corresponding input text is marked as leaky. Furthermore, when all context interpretations successfully pass the DLP system check, the corresponding context is marked as good and passed to the next processing step, such as data transmission or data storage. Moreover, the transformation and analysis of the input text can be completed iteratively or synchronously. Therefore, detecting data leakage can minimize unnecessary processing and computational overhead of the GPT and optimize resource utilization, ensuring that the GPT's computational resources are only used on text that may contain hidden data leakage.

[0030] Furthermore, at step 106, method 100 further includes: once the DLP system detects a data breach in any semantic interpretation of the semantic interpretation, reporting to the user and / or control authority that the input text contains a hidden data breach. In one implementation, method 100 further includes: reporting to the user that the input text contains a hidden data breach. In another implementation, method 100 further includes: once the DLP system detects a data breach in any semantic interpretation of the semantic interpretation, reporting to the control authority that the input text contains a hidden data breach. In yet another implementation, method 100 further includes: once the DLP system detects a data breach in any semantic interpretation of the semantic interpretation, reporting to the user and control authority that the input text contains a hidden data breach. Furthermore, reporting detected hidden data breaches to users and / or control authorities enhances data security incident response, enabling timely action to mitigate potential risks and protect sensitive information, thereby reducing the impact of data breaches and enhancing overall data security.

[0031] Advantageously, Method 100 for preventing semantic data leakage aims to prevent data leakage when any sensitive information (e.g., through advanced language techniques) is hidden in the text. In Method 100, interpreting the input text based on the target context enhances the ability to detect hidden data leakage, which traditional keyword matching might fail to detect. Furthermore, by examining the semantic interpretation generated by GPT, understanding the meaning and context of the text significantly reduces false positives and improves the accuracy of detecting hidden data leakage, while reducing computational overhead and efficiently and reliably detecting data leakage. Therefore, Method 100 enhances overall data security and reduces the risk of semantic leakage attacks.

[0032] Steps 102 to 106 are merely illustrative, and other alternatives may be provided, in which one or more steps are added, one or more steps are deleted, or one or more steps are provided in a different order, without departing from the scope of the claims herein.

[0033] A computer program product including instructions is provided, which, when executed by a computer, include causing the computer to perform a method 100 for preventing semantic data leakage in response to input text. The computer program product is implemented as an algorithm and embedded in software stored in a non-transitory computer-readable storage medium. The instructions are stored on the non-transitory computer-readable storage medium and are executable by one or more processors in a computer system to perform method 100. The non-transitory computer-readable storage medium may include, but is not limited to, electronic storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination of the above. Examples of computer-readable storage media include, but are not limited to, electrically erasable programmable read-only memory (EEPROM), random access memory (RAM), read-only memory (ROM), hard disk drive (HDD), flash memory, secure digital (SD) cards, solid-state drives (SSDs), computer-readable storage media, and / or CPU cache memory.

[0034] Figure 2 This is a block diagram of a device for preventing semantic data leakage provided in an embodiment of the present invention. (Reference) Figure 2 The diagram shows a block diagram 200 of a device 202 for preventing semantic data leakage. Furthermore, device 202 includes a context repository 204, a probe engine 206, and a generative pre-trained transformer (GPT) 210.

[0035] The detection engine 206 is a component responsible for interacting with other components, performing and managing specific tasks, including obtaining semantic interpretations and identifying hidden data leaks. Examples of the detection engine 206 may include, but are not limited to, microcontrollers, microprocessors, central processing units (CPUs), complex instruction set computing (CISC) processors, application-specific integrated circuit (ASIC) processors, reduced instruction set (RISC) processors, very long instruction word (VLIW) processors, data processing units, and other processors or control circuits.

[0036] A device 202 is provided for preventing semantic data leakage. Device 202 reduces the risk of sensitive information being leaked due to semantic operations. Therefore, the overall data security and authenticity are improved.

[0037] Device 202 includes a context repository 204 for storing a target context 208 used for semantic interpretation. The context repository 204 is a storage component within device 202 designed to store the specific target context required for the semantic interpretation of input text 212. Furthermore, the target context 208 stored in the context repository 204 refers to key information used in the semantic interpretation. However, the target context 208 can be updated synchronously according to user needs. Therefore, the context repository 204 is used to ensure that semantic analysis is completed with reference to a specific context.

[0038] Furthermore, the probing engine 206 is used to obtain one or more semantic interpretations of the input text 212 by loading it into GPT 210, wherein GPT 210 is used to interpret the input text 212 with reference to the target context. GPT 210 is a dedicated component or model for natural language processing, analyzing and generating human-like text based on the input text 212 and the target context. Additionally, GPT 210 is used to analyze and generate one or more semantic interpretations of the input text 212. Moreover, the generated one or more semantic interpretations are based on an understanding of the text and the target context. The target context can be a predefined target context that helps GPT 210 to understand the text more deeply by associating the corresponding text with a specific domain or topic.

[0039] Furthermore, the detection engine 206 is used to determine whether the input text 212 contains a hidden data leak by examining the semantic interpretation of the input text 212 using a data leak prevention (DLP) system 214, wherein the DLP system 214 is used to detect data leaks based on a set of words and phrases of interest. The DLP system 214 utilizes the set of words and phrases of interest, which can be used to examine the semantic interpretation of the input text 212 to prevent data leaks. In one implementation, the DLP system 214 refers to a system designed to prevent data leaks by identifying and blocking the disclosure of sensitive information. Furthermore, the corresponding words and phrases of interest indicate sensitive information or data leaks. Therefore, if any word or phrase from the set of words and phrases of interest is detected in the semantic interpretation, then a data leak is detected in this case. Therefore, method 100 is used to assess whether the input text 212 contains a hidden data leak to reduce the risk of any potential data leakage and data threats, where the hidden data leak involves sensitive information that may be intentionally or unintentionally hidden in the text.

[0040] According to one embodiment, the detection engine 206 is further configured to: define a target context for semantic interpretation based on user input and / or a set of words and phrases of interest; provide a context domain extension to the GPT 210 based on the target context; and update the context domain extension of the GPT 210 when the target context is updated. Therefore, by using the target context 208 for semantic interpretation and updating the context domain extension of the GPT 210, semantic interpretation can be dynamically customized based on specific user needs for enhanced data security requirements, thereby achieving accurate, reliable, and context-aware detection of hidden data leaks.

[0041] According to one embodiment, the detection engine 206 is further configured to: determine whether the DLP system 214 has detected a data leak in the input text 212 before acquiring a semantic interpretation of the input text 212. First, a set of words and phrases of interest, such as a list of sensitive words, terms, phrases, etc., is prepared for the DLP system 214. Then, the GPT 210 is pre-trained by installing relevant context domain extensions to interpret the input text 212 with reference to the target context. Furthermore, the list of sensitive contexts (i.e., potential targets of data leaks) is updated synchronously. In this case, the input text 212 is transformed using one of the sensitive contexts in the list, and the transformed input text 212 is evaluated using conventional DLP techniques. Subsequently, if a data leak is detected, the evaluation of the data leak is stopped, and the corresponding input text is marked as leaked. Furthermore, when all context interpretations successfully pass the check of the DLP system 214, the corresponding context is marked as good and passed to the next processing step, such as data transmission or data storage. Moreover, the transformation and analysis of the input text 212 can be completed iteratively or synchronously. Therefore, detecting data breaches can minimize unnecessary processing and computational overhead for GPT 210 and optimize resource utilization, ensuring that GPT 210's computational resources are only used on texts that may contain hidden data breaches.

[0042] According to one embodiment, the detection engine 206 is further configured to: report to the user and / or control authority that the input text 212 contains a hidden data breach once the DLP system 214 detects a data breach in any semantic interpretation of the semantic interpretation. Reporting detected hidden data breaches to the user and / or control authority enhances data security incident response, enabling timely action to mitigate potential risks and protect sensitive information, thereby reducing the impact of data breaches and enhancing overall data security.

[0043] Advantageously, the device 202 for preventing semantic data leakage is used to prevent data leakage when any sensitive information (e.g., through advanced language techniques) is hidden in the text. Interpreting the input text based on the target context 212 enhances the ability to detect hidden data leakage, which traditional keyword matching may fail to detect. Furthermore, by examining the semantic interpretation generated by GPT 210, the meaning and context of the text are understood, which significantly reduces false positives and improves the accuracy of detecting hidden data leakage, while reducing computational overhead and efficiently and reliably detecting data leakage. Therefore, the device 202 is used to enhance overall data security and reduce the risk of semantic leakage attacks.

[0044] Figure 3 This is an illustration provided by an embodiment of the present invention, depicting an exemplary scenario for preventing semantic data leakage. Figure 3 Combination Figure 1 and Figure 2 The elements are shown in the reference. Figure 3 The figure 300 shows a description of the device 202.

[0045] In an exemplary scenario, device 202 is used to recover hidden words, meanings, or messages from text and further convert the content back into an explicit representation so that DLP system 214 can classify the corresponding input text. Device 202 is used to obtain one or more semantic interpretations of input text 212 by loading input text 212 into GPT 210, and also to interpret input text 212 with reference to a target context, as shown in operation 308. Furthermore, device 202 is used to define a target context for semantic interpretation based on input text 212 and / or a set of words and phrases of interest 314. In one example, input text 212 (as shown in Table 318) can be defined as "text 4", "text 5", "text 6" using a context (e.g., "text 1", "text 2", "text 3" with predefined interpretations), as shown in operation 316. Moreover, such semantic interpretations are interpreted by GPT 210 (or a context converter), and device 202 is also used to provide a context domain extension 302 to GPT 210 based on the target context, as shown in operation 304. Subsequently, device 202 is used to update the context domain extension 302 of GPT 210 when the target context is updated. The original content is reinterpreted using different contexts so that sensitive data can be recovered through reinterpretation if one of the different contexts is used to hide sensitive data. Device 202 is used to: determine whether the input text 212 contains hidden data leaks by examining the semantic interpretation of the input text 212 using DLP system 214, wherein DLP system 214 is used to detect data leaks based on a set of words and phrases of interest 314, as shown in operation 312. Furthermore, the reinterpreted content will be examined again by DLP system 214, as shown in operation 310. In this case, when sensitive information reappears, the content can be reported as a data leak, as shown in operation 306. Therefore, GPT 210 can be used in conjunction with DLP system 214 to reduce semantic leak attacks and enhance overall data security during data transmission.

[0046] Figure 4 This is an illustration of an exemplary scenario depicting an apparatus for preventing semantic data leakage, provided by an embodiment of the present invention. Figure 4 Combination Figure 1 , Figure 2 and Figure 3 The elements are shown in the reference. Figure 4 The figure 400 shows a description of the device 202.

[0047] In one implementation scenario, device 202 includes a context repository 204, which is also used to store the target context for semantic interpretation (i.e., Figure 2 The target context 208). Furthermore, the probe engine 206 of device 202 is used to acquire one or more semantic interpretations of the input text 212, as shown in operation 402. Furthermore, at operation 404, the probe engine 206 is used to load the input text 212 into GPT 210, wherein GPT 210 is used to interpret the input text 212 by referencing the target context through setting the context, as shown in operation 406. Furthermore, device 202 is used to: determine whether the input text 212 contains hidden data leaks by examining the semantic interpretations of the input text 212 using DLP system 214, wherein DLP system 214 is used to detect data leaks based on a set of words and phrases of interest, as shown in operation 408. Furthermore, the probe engine 206 is used to: define a target context for semantic interpretation based on user input and / or the set of words and phrases of interest, and at operation 414 to provide a context domain extension 302 to GPT 210 by using a base model 424 according to the target context. At operation 412, context domain extension 302 is installed to enable GPT 210 to define a target context for semantic interpretation based on user input and / or a set of words and phrases of interest. Additionally, at operation 410, device 202 is used to update the context domain extension of GPT 210 when the target context is updated. At operation 420, probe engine 206 is used to determine that all context interpretations have successfully passed the check of DLP system 214; then, in this case, the content can be marked as good and passed. Furthermore, at operation 416, the defined target context for semantic interpretation is provided to probe engine 206 to further reinterpret the corresponding input text by DLP system 214. Furthermore, at operation 418, probe engine 206 is used to determine that DLP system 214 has not detected data leakage in input text 212 before obtaining a semantic interpretation of input text 212. Furthermore, at operation 422, device 202 is configured to: report to the user and / or control authority that the input text 212 contains a hidden data leak once the DLP system 214 detects a data leak in any semantic interpretation of the semantic interpretation. Therefore, device 202 for preventing semantic data leaks can be used to prevent data leaks from occurring when any sensitive information (e.g., through advanced language techniques) is hidden in the text.

[0048] Figure 5 This is an exemplary schematic diagram illustrating the execution sequence for preventing semantic leakage, provided by an embodiment of the present invention. Figure 5 Combination Figure 1 , Figure 2 , Figure 3 and Figure 4 The elements within are described. (See reference.) Figure 5 Figure 500 illustrates a process for preventing semantic leakage. Figure 500 describes what can be achieved by... Figure 2 The device 202 performs operations from 512 to 516.

[0049] At operation 502, the context semantics are represented as ( Figure 1 Method 100 begins by acquiring an input sample (or input text 212), as shown in operation 504. Then, at operation 506, ( Figure 2 The device 202 is used to examine all contexts of the corresponding input sample. At operation 508, the device 202 transforms the content using one of the sensitive contexts from a prepared list of sensitive words, terms, phrases, etc. The transformed content can be evaluated using conventional DLP techniques with the given context. Furthermore, at operation 510, the device 202 checks for data leaks, and if a data leak is detected, for example, at operation 512, the corresponding input sample is marked as having a leak. Additionally, a list of sensitive words, terms, phrases, etc., is prepared for DLP checking and can be updated simultaneously. In this case, at operation 514, the device 202 reports the data leak to the user and / or content provider. In another case, at operation 516, if all context interpretations successfully pass the DLP check, the corresponding input text is marked as good and passed to the next processing step. Furthermore, the transformation and analysis of the input data can be performed iteratively or in parallel, for example, at operation 518. Therefore, the device 202 for preventing semantic data leakage can be used to prevent data leakage when any sensitive information (e.g., through advanced language technology) is hidden in the text.

[0050] Modifications to the embodiments of the invention described above may be made without departing from the scope of the invention as defined by the appended claims. Terms such as "comprising," "including," "incorporated," "having," "is / are," etc., used to describe and claim the invention are intended to be interpreted in a non-exclusive manner, allowing for the presence of items, components, or elements not explicitly described. Singular references should also be interpreted as relating to the plural. The word "exemplary" as used herein means "as an example, instance, or illustration." Any embodiment described as "exemplary" is not necessarily to be construed as more preferred or advantageous than other embodiments, or as excluding combinations of features from other embodiments. The word "optionally" as used herein means "provided in some embodiments and not in others." It should be understood that certain features of the invention described in the context of a single embodiment for clarity may also be provided in combination in a single embodiment. Conversely, various features of the invention described in the context of a single embodiment for brevity may also be provided individually or in any suitable combination or as embodiments of any other described aspect of the invention.< / id>

Claims

1. A method (100) for preventing semantic data leakage, characterized in that, The method (100) includes: One or more semantic interpretations of the input text (212) are obtained using a generative pre-trained transformer (GPT) (210), wherein the GPT (210) is used to interpret the input text (212) with reference to the target context (208). The input text (212) is examined by using a data leak prevention (DLP) system (214) to determine whether the input text (212) contains a hidden data leak, wherein the DLP system (214) is used to detect data leaks based on a set of words and phrases of interest.

2. The method (100) according to claim 1, characterized in that, Also includes: The target context for semantic interpretation is defined based on user input and / or the set of words and phrases of interest (208). Provide the context domain extension to the GPT (210) according to the target context (208). When the target context (208) is updated, the context domain extension of the GPT (210) is updated.

3. The method (100) according to claim 1 or 2, characterized in that, Also includes: Before obtaining the semantic interpretation of the input text (212), it is determined that the DLP system (214) has not detected any data leakage in the input text (212).

4. The method (100) according to any one of claims 1 to 3, characterized in that, Also includes: Once the DLP system (214) detects the data leak in any semantic interpretation of the semantic interpretation, it reports to the user and / or control authority that the input text (212) includes the hidden data leak.

5. A device (202) for preventing semantic data leakage, characterized in that, The device (202) includes: Context repository (204) is used to store the target context for semantic interpretation. Detection engine (206), used for: One or more semantic interpretations of the input text (212) are obtained by loading the input text (212) into a generative pre-trained transformer (GPT) (210), wherein the GPT (210) is used to interpret the input text (212) with reference to the target context (208). The input text (212) is examined by using a data leak prevention (DLP) system (214) to determine whether the input text (212) contains a hidden data leak, wherein the DLP system (214) is used to detect data leaks based on a set of words and phrases of interest.

6. The apparatus (202) according to claim 5, characterized in that, The detection engine (206) is also used for: The target context for semantic interpretation is defined based on user input and / or the set of words and phrases of interest (208). Provide the context domain extension to the GPT (210) according to the target context (208). When the target context (208) is updated, the context domain extension of the GPT (210) is updated.

7. The apparatus (202) according to claim 5 or 6, characterized in that, The detection engine (206) is also used for: Before obtaining the semantic interpretation of the input text (212), it is determined that the DLP system (214) has not detected any data leakage in the input text (212).

8. The apparatus (202) according to any one of claims 5 to 7, characterized in that, The detection engine (206) is also used for: Once the DLP system (214) detects the data leak in any semantic interpretation of the semantic interpretation, it reports to the user and / or control authority that the input text (212) includes the hidden data leak.

9. A computer program product comprising instructions, characterized in that, When the computer program product is executed by a computer, the instructions cause the computer to perform the method (100) for preventing semantic data leakage as described in any one of claims 1 to 4 on the input text (212).

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes instructions that, when executed by a computer, cause the computer to perform the method (100) for preventing semantic data leakage as described in any one of claims 1 to 4 on input text (212).