Multi-modal data loss protection using artificial intelligence
Through the multimodal data loss protection system, using artificial intelligence and machine learning technology, the problems of dictionary dependence and pattern limitation in the existing DLP methods are solved, and efficient sensitive data detection and classification of multiple data formats are realized.
Patent Information
- Application Number
- CN202510040862.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-06-17
- Filing Date
- 2025-01-10
- Publication Date
- 2025-07-11
AI Technical Summary
Existing data loss protection (DLP) methods require pre-entry of dictionary, which may be too strict or omit critical information, making it difficult to effectively detect sensitive data across different data modes.
A multimodal data loss protection system is adopted, and artificial intelligence and machine learning are used to realize sensitive data detection and classification of multiple data formats through large language model (LLM) and zero-sample classifiers, combining optical character recognition (OCR) and computer vision.
It improves the accuracy and recall rate of sensitive data detection, reduces false positives and missed reports, enhances the effectiveness of data loss protection, and adapts to the detection needs of different data modes.
Smart Images

Figure CN120296754A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure generally relates to computer network systems and methods, with a particular focus on protecting sensitive data. More particularly, the present disclosure relates to systems and methods for multi-modal data loss prevention (DLP) using artificial intelligence (such as machine learning). Background Art
[0002] In the fields of computing, networking, information technology (IT), cybersecurity, etc., data loss prevention (DLP), also known as data loss prevention, is simply referred to as data protection (when referring to the protection aspect) or simply data loss (when referring to the detection aspect or the problem itself), and involves protecting all aspects of sensitive enterprise data. Of course, the data of a company or any organization is one of its most important assets, and any loss or theft poses a significant economic risk. Enterprise data can be leaked in different ways, namely through email, webmail, cloud storage, social media, and various other applications, as well as simply by connecting storage devices and copying files. Companies can use predefined templates (such as regular expression matching) to formulate policies for sensitive data formats to avoid data leakage. For example, some existing methods of DLP are described in U.S. Patent No. 11,829,347, commonly assigned and titled "Cloud-based data loss prevention" and issued on November 28, 2023, the entire content of which is incorporated herein by reference. These methods generally require companies to provide a dictionary, and the detection techniques include exact data matching (EDM), in which specific keywords, data categories, etc. are marked, or indexed data matching (IDM), in which content that matches the entire or part of a document in a document repository is marked, and possibly optical character recognition (OCR) to obtain any text from images. Although these methods are effective in protecting data, they require pre-input (i.e., a dictionary), and may be too strict, or may miss some key files due to understanding combinations of various file formats (such as images, text, videos, etc.). DLP needs to detect potential data loss without pre-specifying data, without sharing sensitive data, and be able to detect across different modalities. Summary of the Invention
[0003] The present disclosure relates to systems and methods for multimodal data loss protection (DLP) using artificial intelligence. In particular, the method is multimodal because it can understand or generate information across multiple modalities or data types (e.g., text, image, video, audio, etc.). The present disclosure utilizes artificial intelligence and machine learning, where a trained multimodal system is capable of processing and integrating information from various modalities, namely text, image, sound, video, etc. Advantageously, the trained multimodal system is capable of detecting the data categories being accessed, transmitted, etc., without the need for a pre - dictionary by the company IT. The trained multimodal system can be a tool to enhance data loss prevention capabilities.
[0004] In various embodiments, the present disclosure includes a computer - implemented method for multimodal data loss protection (DLP) having steps, a cloud service configured to implement the steps, a server or any other processing device configured to implement the steps, and a non - transitory computer - readable medium storing instructions that, when executed, cause one or more processors to perform the steps. The steps include receiving an input comprising data in any one of a plurality of formats; processing the input to determine whether the data includes sensitive data; and in response to the input including sensitive data, performing the following steps: processing the input to classify the input into a category of a plurality of categories; and providing an indication of the category of the plurality of categories. The steps can also, in response to the input including non - sensitive data, provide an indication that the data is non - sensitive, thus allowing data in transit or not marking data at rest.
[0005] The steps can also, in response to the input including sensitive data, provide an indication of the category of the plurality of categories and a sub - category associated with the category. The plurality of formats can include text format, image format, audio format, video format, source code, and combinations thereof. Processing the input to determine whether the data includes sensitive data can utilize (1) a large language model (LLM) and embeddings, and (2) a machine learning model configured for classification. Processing the input to classify the input into a category can utilize (1) a large language model (LLM) and (2) a zero - shot classifier.
[0006] Processing the input to determine whether the data includes sensitive data and processing the input to classify the input into categories can both utilize one or more machine learning models trained based on a set of labeled training documents. The steps may further include, before training one or more machine learning models with the set of labeled training documents, identifying any mislabeled documents therein by performing optical character recognition (OCR) and checking for the presence of relevant keywords. The steps may further include, before training one or more machine learning models with the set of labeled training documents, filtering out images in the set of labeled training documents based on file size. The steps may further include, before training one or more machine learning models with the set of labeled training documents, grouping images with minor differences in the training document set based on the hashes of the images.
[0007] The present disclosure relates to systems and methods for data loss protection (DLP) using distilled models. In particular, the method is multimodal as it can understand or generate information across multiple modalities or data types (e.g., text, image, video, audio, etc.). The present disclosure utilizes artificial intelligence and machine learning, where a trained multimodal system is capable of processing and integrating information from various modalities, namely text, image, sound, video, etc. Advantageously, the trained multimodal system is capable of detecting the data categories being accessed, transmitted, etc., without the need for a pre - defined dictionary by the company's IT. Further, the use of the advanced model optimization described herein allows the system to utilize smaller, more efficient models for content / data classification and sensitive data identification.
[0008] In various embodiments, the present disclosure includes a computer - implemented method for multimodal data loss protection (DLP) having steps, a cloud service configured to implement the steps, a server or any other processing device configured to implement the steps, and a non - transitory computer - readable medium storing instructions that, when executed, cause one or more processors to execute the steps. The steps include receiving a plurality of general data predictions from a teacher model; determining one or more advantages of the teacher model based on the received general data predictions; generating a synthetic data set based on the one or more advantages of the teacher model; providing the synthetic data set to the teacher model and receiving a plurality of synthetic data predictions from the teacher model based on the synthetic data set; and performing knowledge distillation on a student model based on the synthetic data predictions received from the teacher model to produce a distilled model.
[0009] The steps may further include: wherein the teacher model and the student model are large language models (LLMs). Before receiving multiple general data predictions from the teacher model, the steps may include providing a general data loss protection (DLP) dataset to the teacher model. The multiple general data predictions and the multiple synthetic data predictions may include content category classification predictions. Determining one or more advantages of the teacher model may include determining one or more categories in which the teacher model performs classification with an accuracy higher than a threshold. Generating a synthetic dataset may include using a large language model (LLM) to generate multiple inputs associated with one or more advantages of the teacher model, wherein the synthetic dataset includes the multiple inputs. These steps may further include using the distilled model to classify inputs to a data loss protection (DLP) system in production. The steps may further include receiving an input including data in any of a plurality of formats; processing the input through the distilled model to classify the input into a category of a plurality of categories; and providing an indication of the category of the plurality of categories. The steps may further include, before processing the input for classification, processing the input to determine whether the data includes sensitive data. The plurality of formats may include text format, image format, audio format, video format, source code, and combinations thereof. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] The present disclosure is illustrated and described herein with reference to various drawings, in which, where appropriate, like reference numerals are used to denote like system components / method steps, and in which:
[0011] Figure 1A is a network diagram of three example network configurations for a user's cybersecurity monitoring.
[0012] Figure 1B is Figure 1A a logical diagram of a cloud operating as a zero-trust platform in
[0013] Figure 2 a block diagram of a server.
[0014] Figure 3 a block diagram of a user device.
[0015] Figure 4 is a schematic diagram of a multimodal DLP system for analyzing different input file formats using various tools.
[0016] Figure 5 is Figure 4 a screenshot of an example output of the multimodal DLP system of
[0017] Figure 6 a flowchart of multimodal DLP with an artificial intelligence process.
[0018] Figure 7Screenshots of three example images that are very similar to each other.
[0019] Figure 8 is Figure 6 A flowchart of a process that is an example implementation of the sensitive data classifier step of a multimodal DLP with an artificial intelligence process, using a combination of an LLM and a zero-shot classifier.
[0020] Figure 9 is Figure 6 A flowchart of process 500 that is an example implementation of the sensitive content recognizer of a multimodal DLP with an artificial intelligence process, using a combination of CLIP embeddings and supervised learning XGboost.
[0021] Figure 10 is using Figure 9 An example table of the classification results of the process.
[0022] Figure 11 is an example table of subcategory results.
[0023] Figure 12 A flowchart of a process for multimodal DLP.
[0024] Figure 13 A graph showing the image size and loading time for various image file types.
[0025] Figure 14 A graph of the trend of multiple estimated loading times and image sizes for various image file types.
[0026] Figure 15 A flowchart of a composite text and image classification architecture.
[0027] Figure 16 A flowchart of a process for inline multimodal DLP.
[0028] Figure 17 A flowchart of a knowledge distillation method.
[0029] Figure 18 A flowchart of implementing knowledge distillation within the present system and method.
[0030] Figure 19 A flowchart showing multiple experiments of models for optimizing content classification using different methods.
[0031] Figure 20 A comparison of the class classification metrics between two models.
[0032] Figure 21 A graph showing the performance of a standard student model and a standard teacher model.
[0033] Figure 22It is a diagram for comparing the performance of student models optimized by various methods.
[0034] Figure 23 It is a flowchart of process 850 for knowledge distillation of LLM for data loss protection (DLP). Detailed implementation
[0035] Similarly, the present disclosure relates to systems and methods for using a distilled model for data loss protection (DLP). Various embodiments include mechanisms for generating a distilled model based on a relatively small student model during the knowledge distillation process. Traditionally, large models are good at performing tasks such as content classification. Although these large models are useful, they cannot be used in production because they require a large amount of computational resources and introduce latency, which has a negative impact on the user experience. The present systems and methods provide mechanisms for training smaller models to perform effectively in a production environment while maintaining prediction accuracy and minimizing computational resource usage.
[0036] §1.0 Examples of Network Security Monitoring and Protection
[0037] Figure 1A It is a network diagram of three example network configurations 100A, 100B, 100C for network security monitoring and protection of user 102. Those skilled in the art will recognize that these are some examples for illustrative purposes, and there may be other network security monitoring methods, and these different methods can be used in combination with each other or alone. Additionally, although shown for a single user 102, actual embodiments will handle a large number of users 102, including multi-tenancy. In this example, user 102 (with a user device 300 as shown, for example) communicates on the Internet 104, including accessing cloud services, software as a service, etc. (each of which can be provided through computational resources, such as using one or more servers 200 as shown). As part of providing network security through these example network configurations 100A, 100B, 100C, a large amount of network security data is obtained. The present disclosure focuses on using this network security data for various purposes. Figure 3 shown in) communicates on the Internet 104, including accessing cloud services, software as a service, etc. (each of which can be provided through computational resources, such as using one or more servers 200 as shown). As part of providing network security through these example network configurations 100A, 100B, 100C, a large amount of network security data is obtained. The present disclosure focuses on using this network security data for various purposes. Figure 2 shown) communicates on the Internet 104, including accessing cloud services, software as a service, etc. (each of which can be provided through computational resources, such as using one or more servers 200 as shown). As part of providing network security through these example network configurations 100A, 100B, 100C, a large amount of network security data is obtained. The present disclosure focuses on using this network security data for various purposes.
[0038] Network configuration 100A includes a server 200 located between a user 102 and the Internet 104. For example, the server 200 can be a proxy, gateway, Secure Web Gateway (SWG), Secure Internet and Web Gateway, Secure Access Service Edge (SASE), Secure Service Edge (SSE), Cloud Application Security Broker (CASB), etc. The server 200 is shown inline with the user 102 and is configured to monitor the user 102. In other embodiments, the server 200 need not be inline. For example, for one or more security purposes, the server 200 can monitor requests from and responses to the user 102, and allow, block, warn, and log these requests and responses. The server 200 can be located on a local network associated with the user 102 or on an external network such as the Internet 104. Network configuration 100B includes an application 110 executed on a user device 300. The application 110 can perform functions similar to those of the server 200, as well as functions coordinated with the server 200. Finally, network configuration 100C includes a cloud service 120 configured to monitor the user 102 and perform security as a service. Of course, various embodiments are contemplated herein, including combinations of network configurations 100A, 100B, 100C.
[0039] Network security monitoring and protection can include firewalls, intrusion detection and prevention, Uniform Resource Locator (URL) filtering, content filtering, bandwidth control, Domain Name System (DNS) filtering, defense against advanced threats (malware, spam, cross-site scripting (XSS), phishing, etc.), data protection, sandboxing, antivirus, and any other security technology. Any one of these functions can be implemented by any of the network configurations 100A, 100B, 100C. A firewall can provide deep packet inspection (DPI) and access control across various ports and protocols and has application and user awareness. URL filtering can block, allow, or restrict website access based on policies for users, user groups, or the entire organization, including specific destinations or URL categories (e.g., gambling, social media, etc.). Bandwidth control can enforce bandwidth policies and prioritize critical applications such as those related to entertainment traffic. DNS filtering can control and block DNS requests for known and malicious destinations.
[0040] Intrusion prevention and advanced threat protection can provide comprehensive threat protection against malicious content such as browser exploits, scripts, identified botnets, and malware callbacks. The sandbox can prevent zero-day vulnerabilities (just identified) by analyzing the malicious behavior of unknown files. Antivirus protection can include antivirus, anti-spyware, anti-malware, etc. protection for user 102 using signatures from sources that are continuously updated. DNS security can identify and route command and control connections to a threat detection engine for full content inspection. DLP can continuously monitor user 102 using standard and / or custom dictionaries, including compressed and / or Secure Sockets Layer (SSL) encrypted traffic.
[0041] In some embodiments, network configurations 100A, 100B, 100C can be multi-tenant and can serve a large number of users 102. Newly discovered threats can be published to all tenants almost immediately. Users 102 can be associated with a tenant, which can include enterprises, companies, organizations, etc. That is, a tenant is a group of users who share a common grouping with specific permissions, i.e., a unified group under a certain IT management. The present disclosure may use the terms tenant, enterprise, organization, unit, group, company, etc. interchangeably and refer to a certain group of users 102 managed by an IT group, department, administrator, etc., i.e., a certain group of users 102 managed together. One advantage of multi-tenancy is that globally, in many different organizations, network security threats to a large number of users 102 can be seen. This provides a large amount of data for analysis, using artificial intelligence techniques, comparison, etc.
[0042] Of course, the above network security technologies are only examples. Those skilled in the art will recognize that other technologies are also contemplated herein. That is, any network security method that can be implemented by any network configuration 100A, 100B, 100C. In addition, any one of network configurations 100A, 100B, 100C can be multi-tenant, with each tenant having its own users 102 and configurations, policies, rules, etc.
[0043] §1.1 Cloud Monitoring
[0044] The cloud 120 can extend network security monitoring and protection to the user 102 with near-zero latency. Additionally, the cloud 120 in network configuration 100C can be used with or without the application 110 in network configuration 100B and the server 200 in network configuration 100A. Logically, the cloud 102 can be regarded as an overlay network between the user 102 and the Internet 104 (as well as cloud services, SaaS, etc.). Previously, the IT deployment model included enterprise resources and applications stored within a data center (i.e., physical devices) behind a firewall (perimeter), which could be accessed on-site or remotely by employees, partners, contractors, etc. via a virtual private network (VPN), etc. The cloud 120 replaces the traditional deployment model. The cloud 120 can be used to implement these services in the cloud without the need for physical devices and their management by enterprise IT administrators. As an always-present overlay network, the cloud 120 can provide the same functionality as physical devices and / or appliances regardless of the geography or location of the user 102 and independently of the platform, operating system, network access technology, network access provider, etc.
[0045] There are various techniques for forwarding traffic between the user 102 and the cloud 120. A key aspect of the cloud 120 (as well as other network configurations 100A, 100B) is to monitor all traffic between the user 102 and the Internet 104. All kinds of monitoring methods can include log data 130 that can be accessed by a management system, management service, analysis platform, etc. For illustrative purposes, the log data 130 is shown as a data storage element, and those skilled in the art will recognize that the various computing platforms described herein can access the log data 130 to implement any of the techniques described herein for risk quantification. In an embodiment, the cloud 120 can be used with the log data 130 from any of the network configurations 100A, 100B, 100C and other data from external sources.
[0046] The cloud 120 can be a private cloud, a public cloud, a combination of private and public clouds (hybrid cloud), or the like. Cloud computing systems and methods abstract physical servers, storage, networks, etc., and provide them as on-demand and elastic resources. The National Institute of Standards and Technology (NIST) provides a concise and specific definition, which states that cloud computing is a model for enabling convenient, on-demand network access to a shared pool of configurable computing resources (such as networks, servers, storage, applications, and services), which can be rapidly configured and released with minimal management effort or service provider interaction. Cloud computing is different from the classical client-server model in that it provides applications from the server, which are executed and managed by the client's web browser or the like, without the need to install a client version of the application. Centralization enables the cloud service provider to have full control over the versions of browser-based applications and other applications provided to the client, thus eliminating the need for version upgrades or license management for individual client computing devices. The phrase "Software as a Service" (SaaS) is sometimes used to describe applications provided via cloud computing. A common shorthand for the cloud computing services provided (or even the aggregation of all existing cloud services) is "the cloud". The cloud 120 is contemplated to be implemented by any method known in the art.
[0047] The cloud 120 can be used to provide example cloud services, including Zscaler Internet Access (ZIA), Zscaler Private Access (ZPA), Zscaler Workload Segmentation (ZWS), and / or Zscaler Digital Experience (ZDX), all of which are from Zscaler, Inc. (the assignee and applicant of this application). Additionally, there can be multiple different clouds 120, including clouds with different architectures and multiple cloud services. The ZIA service can provide access control, threat defense, and data protection. ZPA can include access control, microservices segmentation, etc. The ZDX service can provide monitoring of the user experience, such as quality of experience (QoE), quality of service (QoS), etc., in a manner that can obtain insights based on continuous inline monitoring. For example, the ZIA service can provide Internet access to users, and the ZPA service can provide users with access to enterprise resources instead of a traditional virtual private network (VPN), i.e., ZPA provides zero trust network access (ZTNA). Those of ordinary skill in the art will recognize that various other types of cloud services can also be contemplated.
[0048] §1.2 Zero Trust
[0049] Figure 1BIt is a logical diagram of Cloud 120 operating as a zero-trust platform. Zero trust is a framework for protecting organizations in the cloud and mobile world, which claims that no user or application should be trusted by default. Following the principle of least privilege access, which is the key zero-trust principle, trust is established based on context (e.g., user identity and location, the security posture of the endpoint, the requested application or service), and at each step, a policy check is performed through Cloud 120. Zero trust is a cybersecurity policy where security policies are applied based on context established through least privilege access control and strict user authentication, rather than assumed trust. A well-tuned zero-trust architecture leads to a simpler network infrastructure, a better user experience, and improved network threat defense.
[0050] Building a zero-trust architecture requires visibility and control over the users and traffic in the environment, including encrypted traffic; monitoring and validating the traffic between parts of the environment; and strong multi-factor authentication (MFA) methods beyond passwords, such as biometrics or one-time codes. This is performed through Cloud 120. Crucially, in a zero-trust architecture, the network location of a resource is no longer the biggest factor in its security posture. Instead of strict network segmentation, your data, workflows, services, etc. are protected by software-defined micro-segmentation, enabling you to keep them secure anywhere (whether in a data center or in a distributed hybrid and multi-cloud environment).
[0051] The core concept of zero trust is simple: by default, assume that everything is hostile. This is very different from the cybersecurity models built on centralized data centers and secure network perimeters. These network architectures rely on approved IP addresses, ports, and protocols to establish access control and verify trusted content within the network, typically including anyone connected through a remote access VPN. In contrast, the zero-trust approach treats all traffic as hostile, even if it is already inside the perimeter. For example, a workload is blocked from communicating until it is authenticated through a set of attributes such as fingerprints or identity. Identity-based authentication policies bring stronger security regardless of where the workload communicates - in a public cloud, hybrid environment, container, or on-premises network architecture.
[0052] Since protection is environment-independent, zero trust can protect applications and services without the need for architectural changes or policy updates even when they communicate within a network environment. Zero trust enables secure digital transformation by securely connecting users, devices, and applications using business policies across any network. Zero trust is not just about user identity, segmentation, and secure access. It is a strategy for building a cybersecurity ecosystem.
[0053] At its core, there are three principles:
[0054] Terminate every connection: Technologies such as firewalls use a "pass-through" approach, checking files as they are delivered. If a malicious file is detected, the alert is often too late. Effective zero-trust solutions terminate every connection to allow an inline proxy architecture to inspect all traffic, including encrypted traffic, in real time before it reaches its destination to prevent ransomware, malware, etc.
[0055] Protect data with fine-grained context-based policies: Zero-trust policies validate access requests and permissions based on context, including user identity, device, location, content type, and the application being requested. The policies are adaptive, so as the context changes, user access rights are continuously re-evaluated.
[0056] Reduce risk by eliminating the attack surface: With a zero-trust approach, users can connect directly to the applications and resources they need without connecting to the network (see ZTNA). Direct user-to-application and application-to-application connections eliminate the risk of lateral movement and prevent infected devices from infecting other resources. Additionally, users and applications are not visible to the internet and thus cannot be discovered or attacked.
[0057] §2.0 Example Server Architecture
[0058] Figure 2 is a block diagram of server 200, which can be used as a destination on the internet for network configuration 100A, etc. Server 200 can be a digital computer, and in terms of its hardware architecture, it generally includes a processor 202, an input / output (I / O) interface 204, a network interface 206, a data memory 208, and a memory 210. Those of ordinary skill in the art should understand that Figure 2 Server 200 is depicted in an overly simplistic manner, and actual embodiments may include additional components and appropriately configured processing logic to support known or conventional operating features not described in detail herein. The components (202, 204, 206, 208, and 210) can be communicatively coupled via a local interface 212. The local interface 212 can be, for example but not limited to, one or more buses or other wired or wireless connections (as known in the art). The local interface 212 may have additional elements to enable communication, which are omitted for simplicity, such as controllers, buffers (caches), drivers, repeaters, and receivers, etc. Additionally, the local interface 212 can include address, control, and / or data connections to enable proper communication between the aforementioned components.
[0059] The processor 202 is a hardware device for executing software instructions. The processor 202 can be any custom or commercially available processor, a central processing unit (CPU), an auxiliary processor among multiple processors associated with the server 200, a semiconductor-based microprocessor (in the form of a microchip or chipset), or any device commonly used for executing software instructions. When the server 200 is in operation, the processor 202 is configured to execute software stored in the memory 210, communicate data to and from the memory 210, and generally control the operation of the server 200 according to the software instructions. The I / O interface 204 can be used to receive user input from and / or provide system output to one or more devices or components.
[0060] The network interface 206 can be used to enable the server 200 to communicate on a network such as the Internet 104. The network interface 206 can include, for example, an Ethernet card or adapter or a wireless local area network (WLAN) card or adapter. The network interface 206 can include address, control, and / or data connections to enable proper communication on the network. The data memory 208 can be used to store data. The data memory 208 can include any volatile memory elements (e.g., random access memory (RAM), such as DRAM, SRAM, SDRAM, and the like), non-volatile storage elements (e.g., ROM, hard disk drive, magnetic tape, CDROM, and the like), and combinations thereof. In addition, the data memory 208 can incorporate electronic, magnetic, optical, and / or other types of storage media. In one example, the data memory 208 can be located inside the server 200, such as an internal hard disk drive connected to the local interface 212 in the server 200. Additionally, in another embodiment, the data memory 208 can be located outside the server 200, such as an external hard disk drive (e.g., SCSI or USB connection) connected to the I / O interface 204. In another embodiment, the data memory 208 can be connected to the server 200 via a network, such as a network-connected file server.
[0061] The memory 210 may include any volatile memory elements (e.g., random access memory (RAM), such as DRAM, SRAM, SDRAM, etc.), non-volatile storage elements (e.g., ROM, hard disk drive, magnetic tape, CDROM, etc.), and combinations thereof. In addition, the memory 210 may incorporate electronic, magnetic, optical, and / or other types of storage media. Note that the memory 210 may have a distributed architecture, where various components are placed remotely from each other but can be accessed by the processor 202. The software in the memory 210 may include one or more software programs, each software program including an ordered list of executable instructions for implementing logical functions. The software in the memory 210 includes a suitable operating system (O / S) 214 and one or more programs 216. The operating system 214 substantially controls the execution of other computer programs (e.g., one or more programs 216) and provides scheduling, input / output control, file and data management, memory management, communication control, and related services. One or more programs 216 may be configured to implement the various processes, algorithms, methods, techniques, etc. described herein. Those skilled in the art will recognize that the cloud 120 ultimately runs on one or more physical servers 200, virtual machines, etc.
[0062] §3.0 Example User Equipment Architecture
[0063] Figure 3 is a block diagram of a user device 300, which can be used by a user 102. Specifically, the user device 300 may form a device used by one of the users 102, which may include common devices such as laptops, smartphones, tablets, netbooks, personal digital assistants, mobile phones, e-book readers, Internet of Things (IoT) devices, servers, desktops, printers, TVs, streaming devices, storage devices, and the like. The user device 300 may be a digital device, and in terms of its hardware architecture, it generally includes a processor 302, an I / O interface 304, a network interface 306, a data memory 308, and a memory 310. Those of ordinary skill in the art should understand that Figure 3 the user device 300 is depicted in an overly simplistic manner, and actual embodiments may include additional components and appropriately configured processing logic to support known or conventional operating features not detailed herein. The components (302, 304, 306, 308, and 302) may be communicatively coupled via a local interface 312. The local interface 312 may be, for example but not limited to, one or more buses or other wired or wireless connections (as known in the art). The local interface 312 may have additional elements to enable communication, which are omitted for simplicity, such as controllers, buffers (caches), drivers, repeaters, and receivers, etc. In addition, the local interface 312 may include address, control, and / or data connections to enable proper communication between the above components.
[0064] The processor 302 is a hardware device for executing software instructions. The processor 302 can be any custom or commercially available processor, a CPU, an auxiliary processor among multiple processors associated with the user device 300, a semiconductor-based microprocessor (in the form of a microchip or chipset), or any device commonly used for executing software instructions. When the user device 300 is in operation, the processor 302 is configured to execute software stored in the memory 310, communicate data to and from the memory 310, and generally control the operation of the user device 300 according to the software instructions. In an embodiment, the processor 302 may include a mobile-optimized processor, such as a processor optimized for power consumption and mobile applications. The I / O interface 304 can be used to receive user input and / or provide system output. User input can be provided through, for example, a keyboard, a touch screen, a scroll ball, a scroll bar, buttons, a barcode scanner, and the like. System output can be provided through display devices such as liquid crystal displays (LCDs), touch screens, and the like.
[0065] The network interface 306 enables wireless communication with an external access device or network. The network interface 306 can support any number of suitable wireless data communication protocols, technologies, or methods, including any protocol for wireless communication. The data memory 308 can be used to store data. The data memory 308 can include any volatile memory elements (e.g., random access memory (RAM), such as DRAM, SRAM, SDRAM, and the like), non-volatile storage elements (e.g., ROM, hard disk drive, magnetic tape, CDROM, and the like), and combinations thereof. In addition, the data memory 308 can incorporate electronic, magnetic, optical, and / or other types of storage media.
[0066] The memory 310 can include any volatile memory elements (e.g., random access memory (RAM), such as DRAM, SRAM, SDRAM, etc.), non-volatile storage elements (e.g., ROM, hard disk drive, etc.), and combinations thereof. In addition, the memory 310 can incorporate electronic, magnetic, optical, and / or other types of storage media. Note that the memory 310 can have a distributed architecture, where various components are placed remotely from each other but can be accessed by the processor 302. The software in the memory 310 can include one or more software programs, each software program including an ordered list of executable instructions for implementing logical functions. In Figure 3In the example, the software in the memory 310 includes a suitable operating system 314 and programs 316. The operating system 314 basically controls the execution of other computer programs and provides scheduling, input / output control, file and data management, memory management, communication control, and related services. The programs 316 can include various applications, add-ons, etc., which are configured to provide end-user functions to the user device 300. For example, the example programs 316 can include, but are not limited to, a web browser, a social network application, a streaming media application, games, a map and location application, an email application, a financial application, and the like. The application 110 can be one of the example programs.
[0067] §4.0 Data Loss
[0068] DLP involves monitoring an organization's sensitive data, including data on endpoint devices, data at rest (i.e., data stored somewhere), and data in motion (i.e., data being transmitted somewhere). DLP monitoring methods focus on various products, including software agents on endpoints, physical devices, virtual devices, etc. As applications migrate to the cloud and users can access them directly from anywhere they are connected, this inevitably leaves blind spots as users bypass the security controls in traditional DLP methods when they are off the network. Thus, U.S. Patent No. 11,829,347, titled "Cloud-based dataloss prevention" and issued on November 28, 2023, as previously cited, describes cloud-based technologies.
[0069] The present disclosure includes an AI-based DLP method that classifies data into one of multiple categories. Those skilled in the art will recognize that this method can be used in any system architecture, including network configurations 100A, 100B, 100C and their variants for network security monitoring and protection, as well as other methods known in the art. In addition, the AI-based method can be used in combination with existing DLP methods known in the art.
[0070] §4.1 Traditional DLP
[0071] Generally, all of these prior arts utilize DLP dictionaries, which include specific types of information and custom information in user traffic and information. For example, specific types of information can look for data types, such as personally identifiable information (PII), bank information, credit card information, etc. That is, certain content can be detected based on its format for specific information, a simple example being a social security number in the format XXX-XX-XXX. Custom information can be specific keywords of a company, such as customer names, product names, etc. In addition, custom information can be specific documents, i.e., the sensitive information itself. That is, DLP can detect keywords, specific types of information, actual documents, and parts of actual documents.
[0072] Using dictionaries, different techniques can be used to detect this information, including exact data matching (EDM), where specific keywords, data categories, etc. are marked. For example, DLP can detect social security numbers, credit card numbers, etc. based on the data format (such as structured documents, etc.). In unstructured documents, there can also be a method called index document matching (IDM) for identifying and protecting content that matches all or part of a document in a document repository. Additionally, any of these methods can be implemented through optical character recognition (OCR) to cover non-text data.
[0073] Similarly, these methods work well, but they also have some drawbacks. First, these methods require the use of a pre - defined dictionary. For specific types of information, DLP monitoring systems typically provide pre - defined dictionaries for specific types of messages. Thus, IT can pre - select these dictionaries. For a document repository, IT must provide this information. To address the desire to avoid sharing sensitive information, the method provides hashing techniques to allow the detection of sensitive information without sharing the actual sensitive information. However, a key point here is the need to provide information and / or select the dictionary in advance. Another drawback is that these methods tend to be either too strict (false positives) or miss key information (false negatives). In the case of being too strict, users 102 are prohibited from exchanging data that complies with the rules. For example, blocking and reporting emails that seemingly contain bank or PII information, but in fact the information belongs to user 102. Additionally, if a new document is not in the provided repository, it may be lost.
[0074] §4.2 Multimodal DLP Using Artificial Intelligence
[0075] Figure 4 FIG. is a schematic diagram of a multi - modal DLP system 400 that uses various tools 404 to analyze different input file formats 402. The multi - modal DLP system 400 is referred to as a system, and those skilled in the art will recognize that this can be implemented through a method with steps, through a non - transitory computer - readable medium having instructions that cause one or more processors to implement the steps, and through computing resources configured to implement the steps. For example, the computing resources can include a cloud 120, a server 200, a user device 300, etc.
[0076] The multimodal DLP system 400 is called multimodal, which means it can understand or generate information across multiple modalities or data types. In the context of artificial intelligence and machine learning, the multimodal DLP system 400 can process and integrate information from various modalities such as text, images, sound, video, etc. Traditional DLP solutions are limited to understanding and managing text and image-based data, and the world has transitioned to a wider range of visual and audio multimedia formats. The multimodal DLP system 400 enhances the way DLP operates by integrating generative AI and multimodal capabilities to protect customers' data from leakage in various media formats such as video and audio formats in addition to text and images.
[0077] Accordingly, the input file format 402 takes into account any type of content that can be used to convey information. The input file format 402 can be an image, text, audio, video, and their combinations. In particular, the input file format 402 can extend beyond anything that can be reduced to text. For example, traditional methods look for text in images or videos through OCR, etc., and look for text in audio by converting audio to text, etc. With artificial intelligence and machine learning, DLP detection is not limited to text and can extend to pure images and the like. That is, the output of the multimodal DLP system 400 is not just a judgment that some sensitive data is contained in the file, but can classify the type of content.
[0078] In various embodiments, the comprehensive input file format 402 can include but is not limited to image format, video format, text format, spreadsheet, comma-separated values (CSV) format, source code, presentation format, portable document format (PDF), and the like. The comprehensive input file format 402 can be a single input 406 to the tool 404 in the multimodal DLP system 400. The various tools 404 can include one or more large language models (LLMs) 410, OCR / computer vision (CV) system 412, speech detection system 414, and natural language processing (NLP) system 416. In some embodiments, a specific tool 404 can be used based on the file format 402. In other embodiments, multiple tools 404 can be used on the same file. For example, an audio file can be processed by the speech detection system 414 and then by the LLM 410 and / or NLP system 416. Similarly, in some embodiments, an image or video file can be processed by the OCR / CV system 412 and then by the LLM 410 and / or NLP system 416. In various embodiments, the LLM 410 can process all different file formats 402.
[0079] This disclosure contemplates using one or more tools 404 based on different file formats 402. In an embodiment, the following models are used individually and in combination with each other in the tool 404:
[0080] (1) BLIP (Bootstrapping Language-Image Pretraining), see, for example, Li, Junnan, et al.'s "Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation", International Conference on Machine Learning, PMLR, 2022; the entire content of which is incorporated herein by reference. The BLIP model is capable of processing images.
[0081] (2) Video LLaMA (Large Language Model Meta AI), see, for example, Zhang, Hang, Xin Li, and Lidong Bing's "Video-llama: An instruction-tuned audio-visual language model for video understanding", arXiv preprint arXiv:2306.02858 (2023), the entire content of which is incorporated herein by reference. The Video LLaMA model is capable of processing images and videos.
[0082] (3) LLaVa (Large Language and Vision Assistant), see, for example, Liu, Haotian, et al.'s "Visual instruction tuning", arXiv preprint arXiv:2304.08485 (2023). This is a novel end-to-end trained large multimodal that combines visual encoding and Vicuna for general vision and language understanding (Vicuna is a chatbot trained by fine-tuning LLaMA on user-shared conversations collected from ShareGTP, see, for example, Zheng, Lianmin, et al.'s "Judging LLM-as-a-judge with MT-Bench and Chatbot Arena", arXiv preprint arXiv:2306.05685 (2023), the entire content of which is incorporated herein by reference.
[0083] (4) BART (Bidirectional Auto-Regressive Transformer) zero-shot classifier. BART is pre-trained in two steps: (1) corrupting the text with an arbitrary noise function, and (2) learning the model to reconstruct the original text. Zero-shot classification is a machine learning method in which the model is able to classify data into multiple categories without any specific training examples for these categories.
[0084] (5) CLIP (Contrastive Language-Image Pretraining), see, e.g., Radford, Alec, et al., "Learning transferable visual models from natural language supervision", International Conference on Machine Learning, PMLR, 2021; the entire content of which is incorporated herein by reference. CLIP can predict the most relevant text snippet for a given image.
[0085] Figure 5 is a screenshot of an example output of the multimodal DLP system 400. Figure 5 is presented for illustrative purposes, and those skilled in the art will understand that this output can be used in the cloud 120, in any network configuration 100A, 100B, 100C, etc. for various purposes, including allowing / blocking content, providing notifications and alerts, crawling cloud services for detection, etc. Here, a single file (e.g., an image, video, document, CSV, source code, etc.) is input, and the tool 404 analyzes the file. For example, the file is an image in this case - a screenshot in the form of a Portable Network Graphics (PNG) file. The output includes the classification of the information as (1) sensitive and (2) in the category or supercategory of tax documents, as well as a confidence score (e.g., 80%), and other details, such as details derived from the LLM 410.
[0086] §4.3 Multimodal DLP Using Artificial Intelligence Process
[0087] Figure 6 is a flowchart of multimodal DLP utilizing the artificial intelligence process 450. The process 450 is envisioned to be implemented as a method with steps, via computing resources configured to implement these steps, and as a non-transitory computer-readable medium with instructions that, when executed, cause one or more processors to implement these steps. The process 450 can be implemented with the multimodal DLP system 400, and the actual implementation of the multimodal DLP system 400 and the process 450 can be achieved through network configurations 100A, 100B, 100C, and the like. That is, the multimodal DLP system 400 and the process 450 are contemplated for use with any network security monitoring platform, device, service, etc.
[0088] Process 450 is implemented by a two - stage classifier including a sensitive content recognizer step 452 and a sensitive data classifier step 454. Process 450 uses these two steps to improve the detection of sensitive data and enhance the user experience. Process 450 starts with an input (step 460). Similarly, the input can be some content in any file format or a combination of multiple formats. The sensitive content recognizer step 452 determines whether the input has sensitive data, emphasizing the precision and recall of sensitive categories to reduce false positives and false negatives. The sensitive content recognizer step 452 can include using LLM embeddings and machine learning classifiers to determine that the input is sensitive (step 464) or not sensitive (step 462). The LLM embeddings are used to detect and classify objects in the input, and the machine learning classifier can be used to classify text and objects. Of course, process 450 can terminate when it determines that the input is not sensitive (step 462), that is, there is no potential data loss. Thus, the sensitive content recognizer step 452 enables faster and more effective detection.
[0089] The sensitive data classifier step 454 needs to be executed only when the input is sensitive. The sensitive data classifier step 454 organizes sensitive information into predefined categories to enhance the user experience and for reporting. For example, the sensitive data classifier step 454 can determine the super - category (step 466) and the sub - category (step 468). For example, the super - categories can be finance, engineering, marketing, sales, human resources, tax, etc., that is, larger classifications. The sub - categories for each super - category may vary. For example, financial invoices, purchase orders, purchase agreements, financial statements, sales orders, loan agreements, etc. and the like.
[0090] The two steps 452, 454 can be used with various network security monitoring methods. The sensitive content recognizer step 452 can be the front - end and has shown more than 90% accuracy in tests. In the case of data transmission, the sensitive content recognizer step 452 can be used to block / allow files. In the case of data at rest, the sensitive content recognizer step 452 can be used to effectively identify and further process sensitive data, that is, there is no need for a comprehensive detection of non - sensitive data. The sensitive data classifier step 454 can be used by IT for policies. The two steps 452, 454 can use various combinations of tools 404, including the above - mentioned example machine learning models.
[0091] The following table provides some metrics associated with the implementation of process 450:
[0092]
[0093] §4.4 Image Data Denoising and Cleaning Algorithm for Image Classification
[0094] Both step 452 and step 454 can perform image classification using a machine learning model. In an embodiment, the present disclosure includes various techniques to enhance the cleanliness of image data and improve the quality of tasks related to image classification. These techniques can be used in conjunction with any image-based file format 402, tool 404 for processing images, and steps 452 and 454. It includes three aspects that can be used together or separately, including: OCR, file size filtering, and image hashing.
[0095] §4.4.1 OCR
[0096] OCR includes converting any text in an image into a computer-readable text format, that is, converting typed text, handwritten text, or printed text in an image or video screenshot into machine-encoded text. For example, a PDF document can be image-based, but it represents a text document. Typically, OCR is used to convert an image into text, and then the text is processed through a DLP dictionary. For example, see U.S. Patent No. 11,805,138, commonly assigned, titled "Data Loss Prevention on images," issued on October 31, 2023, the entire content of which is incorporated herein by reference. The present disclosure contemplates using OCR in the multimodal DLP system 400 for enhancing processing and / or improving efficiency, speed, etc.
[0097] In an embodiment, OCR can be used to train a machine learning model. As is known in the art, supervised learning involves a training process in which a machine learning model is trained using labeled samples. After training, the machine learning model can perform inference or classification in production. In the multimodal DLP system 400, one or more machine learning models can be used in the sensitive content recognizer step 452 to classify an image as sensitive or non-sensitive, and in the sensitive data classifier step 454 to classify the image into one of multiple supercategories and into one or more subcategories.
[0098] For both tasks, classification as sensitive / non-sensitive and classification by category require labeled training data. For example, a first set of documents is each labeled as sensitive or non-sensitive, and a second set of documents is labeled by category (of course, this second set of documents can be the same as, different from, or include some of the same images as the first set). A key aspect of this training is that it does not require individual companies to provide their sensitive information. Instead, the trained machine learning model is trained on a set of documents such as those from a public repository, model creator, etc.
[0099] For example, the model creator can use their own internal documents for training. The model creator can be a SaaS provider, a cloud service provider, etc., and the model creator has a large number of already labeled internal documents. The model creator can obtain documents from departments such as finance, HR, sales, engineering, etc., and attach labels to these categories. In other embodiments, other machine learning techniques such as clustering can be used to label the training documents.
[0100] In an embodiment, OCR can be used to identify mislabeled samples in one or both of the first set of training documents or the second set of training documents. Here, text can be extracted from the labeled sample images and then verified based on keywords related to sensitive data categories. For example, checking whether the extracted text contains words such as "property", "asset", "seller", "buyer", or "purchase agreement" for real estate documents. This can be extended to all supercategories and subcategories in the sensitive data classifier step 454. The output of this OCR can be a set of suspicious training documents, which can be provided to user input. The output can be used to further refine the keywords and relabel any images in the set of suspicious images.
[0101] OCR can also be used in the sensitive content recognizer step 452. In an embodiment, the sensitive / non-sensitive classification can be performed in combination with the category classification. For example, all HR documents are sensitive, etc. In another example, the sensitive / non-sensitive classification can be performed using a different model from the category classification. In either case, OCR can be used to identify mislabeled samples for the sensitive / non-sensitive classification. Here, there are a large number of sensitive and non-sensitive words.
[0102] The key to this technology is that since the input data is clearer in terms of labeling, it improves the training effect of the model used for classification.
[0103] §4.4.2 File Size Filtering
[0104] In various embodiments, the present disclosure can include techniques for filtering out images that are blurry, unreadable, etc. This filtering technique can be performed in both the first set of documents or the second set of documents used for training, and can also be performed in any input analyzed by the multimodal DLP system 400 in production. In particular, the filtering can be based on file size or image resolution. Smaller image files are usually blurry and difficult to visualize.
[0105] In training data, file size filtering can be used to exclude poor-quality images from the first set of documents or the second set of documents used for training. In production, file size filtering can be used to exclude images from the processing of the multimodal DLP system 400, where there is no risk due to the images not having any discernible content. In training, this method improves the quality of the training data, and in production, this approach improves the efficiency and resource cost of the multimodal DLP system 400.
[0106] §4.4.3 Image Hashing
[0107] In addition, the present disclosure can include image hashing, which can detect identical or very similar images in the training data (i.e., the first set of training data or the second set of training data) and eliminate duplicates. Figure 7 are screenshots of three sample images that are very similar to each other. In particular, these three sample images are government documents for assigning trademarks, collective trademarks, or service marks. That is, Figure 7 the highlighted part in shows the only difference among these three documents, that is, the words "Trademark", "Collective Mark", or "Service Mark" after "Assignment of". Through image hashing, these three documents are detected as being very similar to each other, and these three samples are grouped together - because they exhibit a high degree of similarity with only minor differences. To prevent data leakage problems, it is recommended to retain only one of them in the training data. In an embodiment, image hashing can utilize the ImageHash available at pypi.org / project / ImageHash / .
[0108] §4.4.4 Data Cleaning
[0109] For training data, the present disclosure can use various methods described herein for data cleaning. Some additional methods for data cleaning of image files include removing observed duplicates, removing logos, removing the cover page or the instruction page, etc., as described above.
[0110] §4.5 Combination of LLM and Zero-Shot Classifier
[0111] Figure 8It is a flowchart of process 480, which is an example implementation of sensitive data classifier step 454 using a combination of an LLM and a zero-shot classifier. Process 480 is considered to be implemented as a method with steps, via computing resources configured to implement these steps, and as a non-transitory computer-readable medium with instructions that, when executed, cause one or more processors to implement these steps. Process 480 can be implemented using multimodal DLP system 400, and the actual implementation of multimodal DLP system 400 and processes 450, 480 can be achieved through network configurations 100A, 100B, 100C, etc. That is, process 480 is considered to be used in conjunction with any network security monitoring platform, device, service, etc.
[0112] Process 480 includes receiving an input (step 482) and both of the following:
[0113] (1) Processing the input using an LLM to describe the image (step 484) and processing the description using a zero-shot classifier (step 486). For example, LlaVa can be used as the LLM to describe the image, and BART can be used as the zero-shot classifier.
[0114] (2) Processing the input using a zero-shot classifier such as CLIP-VIT (step 488).
[0115] Process 480 includes obtaining the outputs of the two zero-shot classifier steps 486, 488, assigning a weighted average (step 490), and providing an output classification (step 492).
[0116] §4.6 Combination of CLIP Embedding and Supervised Learning Xgboost
[0117] Figure 9 It is a flowchart of process 500, which is an example implementation of sensitive content recognizer step 452 using a combination of CLIP embeddings and supervised learning XGboost. Process 500 is considered to be implemented as a method with steps, via computing resources configured to implement these steps, and as a non-transitory computer-readable medium with instructions that, when executed, cause one or more processors to implement the steps. Process 500 can be implemented using multimodal DLP system 400, and the actual implementation of multimodal DLP system 400 and processes 450, 480, 500 can be achieved through network configurations 100A, 100B, 100C, etc. That is, process 500 is considered to be used in conjunction with any network security monitoring platform, device, service, etc.
[0118] Process 500 includes receiving an input (step 502), processing the input with a CLIP-VIT model (step 504) to obtain image and text embeddings, using an XGboost classifier trained with labeled data to process the image and text embeddings (step 506), and providing an output (step 508).
[0119] Figure 10 is an example table of the classification results using process 500. Here, process 500 is configured to classify the input into one of 13 categories such as finance, law, etc. Notably, both precision and recall are high. Figure 11 is an example table of sub-category results.
[0120] In various embodiments, the following models are used herein:
[0121]
[0122] §4.7 Reducing the Latency of LLM by Model Compression
[0123] The following table shows some example models, related performance, and costs for implementing sensitive data classifier step 454. The LLAVA+ zero-shot method uses a pre-trained model for classification. Embeddings are made using the Clip Base / Large method. Then, the data is split into training / test parts, and an ML classifier is built using the embedded feature vectors.
[0124]
[0125] §4.8 Multimodal DLP Process
[0126] Figure 12 is a flowchart of process 550 for multi-modal DLP. Process 550 is contemplated to be implemented as a method having steps, via computing resources configured to implement the steps, and as a non-transitory computer-readable medium having instructions that, when executed, cause one or more processors to implement the steps. Process 550 can be implemented using multi-modal DLP system 400, and the actual implementation of multi-modal DLP system 400 and processes 450, 480, 500, 550 can be achieved through network configurations 100A, 100B, 100C, etc. That is, process 550 is contemplated to be used with any network security monitoring platform, device, service, etc.
[0127] Process 550 includes receiving an input that includes data in any of a plurality of formats (step 552); processing the input to determine whether the data includes sensitive data (step 554); and in response to the input including sensitive data, performing the following steps: processing the input to classify the input into categories of a plurality of categories; and providing an indication of the category of the plurality of categories (step 556).
[0128] Process 550 can also include, in response to the input including non-sensitive data, providing an indication that the data is non-sensitive, thereby allowing the data in transit or the data at rest to be unmarked. Process 550 can also include, in response to the input including sensitive data, providing an indication of the category of the plurality of categories and sub-categories associated with the category.
[0129] The plurality of formats can include text format, image format, audio format, video format, source code, and combinations thereof. Processing the input to determine whether the data includes sensitive data can utilize: (1) a large language model (LLM) and embeddings, and (2) a machine learning model configured for classification. Processing the input to classify the input into categories can utilize: (1) a large language model (LLM) and (2) a zero-shot classifier.
[0130] Both processing the input to determine whether the data includes sensitive data and processing the input to classify the input into categories can utilize one or more machine learning models trained based on a set of training documents with labels. Process 550 can also include, before training one or more machine learning models with the set of training documents with labels, identifying any mislabeled documents by performing optical character recognition (OCR) and checking for the presence of relevant keywords. Process 550 can also include, before training one or more machine learning models with the set of training documents with labels, filtering out images in the set of training documents with labels based on file size. Process 550 can also include, before training one or more machine learning models with the set of training documents with labels, grouping images with minor differences in the training document set based on the hash of the images.
[0131] §4.9 Multimodal DLP and Traditional DLP
[0132] The present disclosure presents various methods for multi-modal DLP leveraging artificial intelligence. These techniques can be used in combination with existing DLP detection techniques (e.g., techniques with DLP dictionaries). In an embodiment, traditional DLP techniques can be used with multi-modal DLP leveraging machine learning to provide two answers that can be combined to give a final, more reliable answer (sensitive or non-sensitive). In another embodiment, multi-modal DLP leveraging machine learning and traditional DLP techniques can be used front-end to each other, i.e., in any way that reduces the computational workload. For example, the sensitive content recognizer step 452 can be used as a front-end classifier to determine whether further analysis of the given content (e.g., using a DLP dictionary) is needed to improve the efficiency and latency of DLP monitoring.
[0133] §4.10 Categories
[0134] Traditional DLP methods work by detecting specific predefined content, while the methods described herein provide classification and / or categorization of content. As described herein, the multi-modal DLP process 450 leveraging artificial intelligence can use the sensitive content recognizer step 452 to perform classification of whether the data is sensitive or non-sensitive, and the sensitive data classifier step 454 can perform categorization of the data into one of multiple categories. Advantages of this method include the ability to detect classes and categories of documents without the need to pre-provide sensitive data, as well as improved accuracy, reduced false negatives or false positives, etc.
[0135] Classification can be based on the training described herein. In an embodiment, the training can be performed by an entity hosting the model (referred to as the model creator). The model creator can use its own internal documents as well as publicly available documents to eliminate the need for a company to disclose its sensitive documents. Specifically, these models are trained for document types rather than specific content. In an embodiment, the categories can include immigration documents, corporate legal documents, court documents, legal documents, tax documents, insurance documents, invoice documents, resume documents, real estate documents, medical documents, technical documents, and financial documents. Of course, other categories are possible based on the training data and associated labels. That is, the number of labels determines the number of categories. Additionally, there can be an "other" category classification for files / content that do not belong to the various categories.
[0136] §4.11 DLP Rules and Analysis
[0137] Similarly, the methods described herein can be used in combination with traditional DLP techniques (i.e., using dictionaries). This enables an overall approach to DLP monitoring in network security. Specifically, in addition to being able to prevent the exposure of specific data, IT can gain in-depth insights into activities and identify disconnects that exist between data processing based on user functions. Using categories, DLP rules are no longer limited to a single document containing sensitive information. Instead, strategies can be developed to compare a person's function with their activities. For example, a person in the engineering department should not be handling a large number of HR documents, and conversely, a person in the HR department should not be handling a large number of engineering documents.
[0138] §4.12 Online Multimodal DLP
[0139] The present disclosure provides various methods for simplifying the multi-modal data loss protection (DLP) process described herein for better in-line utilization. These methods are aimed at improving the efficiency and effectiveness of DLP systems in handling multiple data types (e.g., text, images, audio, etc.). Key aspects of these methods include optimizing the data flow to ensure the effective management and protection of data from various sources without causing blockages or delays, and shortening the processing latency through advanced algorithms to speed up data processing and threat detection. Enhanced real-time inference enables the DLP system to classify faster and more accurately to ensure immediate protection against data breaches and leaks, which is particularly important for in-line applications where data needs to be protected during transmission or use. Improved preprocessing steps ensure that the data is correctly formatted and ready for analysis by the DLP system, employing techniques for handling different data formats and ensuring the integrity of the data before processing. Additionally, methods for training and tuning the DLP model are included to identify and defend against new and evolving threats, ensuring that the system remains effective over time and can handle various data types and scenarios. These methods are designed to seamlessly integrate with existing systems and workflows, enabling organizations to implement advanced DLP capabilities without significant disruption or overhaul to their current infrastructure. By focusing on these areas, the purpose of the present disclosure is to provide a comprehensive and effective multi-modal DLP approach that enables better protection of sensitive data in a wide range of real-world applications.
[0140] In addition to the above models (1)-(5), the use of the BERT model is also considered in various embodiments. BERT (Bidirectional Encoder Representations from Transformers) has created new benchmarks in multiple NLP tasks. This is due to the bidirectionality of context understanding. In addition to using text preprocessing, this model also attempts to capture the meaning from unseen words using WordPiece Tokenizer, and even makes it the preferred choice when dealing with noisy data. Similarly, as described herein, the various models used in this system can be trained with specific data or pre-trained. Then, various modifications described herein can be made to the models used in the system.
[0141] In various embodiments, the present inline multi-modal DLP technology includes modifying one or more models to maintain prediction accuracy while reducing latency. This is referred to herein as one or more modifications to the models used in the current multi-modal DLP system described herein. Various modifications are also described herein and test results are provided.
[0142] In various embodiments, one modification can include reducing the model vocabulary. For example, the BERT Tiny model vocabulary can be simplified to reduce latency. In addition, introducing a file size threshold can further reduce latency while maintaining prediction accuracy. For image and text classification, the BERT Tiny model can be fine-tuned for text classification, while the VisionTransformer (ViT) Tiny model can be fine-tuned for image classification. As further described herein, these models can also be used in a composite manner.
[0143] As described above, the various models described herein can be simplified through one or more modifications to reduce latency and maintain prediction accuracy. These methods are used to speed up text preprocessing without sacrificing prediction accuracy. These methods include removing non-English words from the model vocabulary, removing stop words, and lemmatization, i.e., reducing words to their root forms. The following table illustrates the impact of these methods on the model prediction accuracy. The example model is the BERT model.
[0144]
[0145] In addition, in various embodiments, to further reduce latency while maintaining prediction accuracy, these methods can include modifying the model by implementing text byte thresholds for lower and upper bounds, as well as text processing stop points. A text byte lower threshold can be employed to cause the model to skip files below this threshold and classify them as miscellaneous or "other". An upper text byte threshold can be implemented to cause the model to extract only a specified amount of text to input into the prediction model. Finally, a defined early stopping k value determines the stop point for text processing and vocabulary mapping iterations. That is, this innovative approach can be used to accelerate text vocabulary mapping by iterating only k tokens. The following table shows the iteration times associated with multiple file types and multiple specified early stopping k values.
[0146]
[0147] In addition, in various embodiments, for image processing models such as ViT, a maximum input file size can be implemented to control latency. These methods include using polynomial regression to predict trends and estimate thresholds. This is necessary because different image file types result in different loading times. Figure 13 is a graph showing the image size and loading time of multiple image file types. It can be seen that the loading time has no direct relationship with the image size of different image file types. Figure 14 Shows the trends of estimated loading times and image sizes for various image file types, with the estimated loading time and image size trends determined based on polynomial regression or other similar methods. Based on these trends, a threshold can be determined and implemented for each image file size based on the image file type. For example, an image file size threshold can be implemented, where image files with a size exceeding the determined threshold will not be processed. Similarly, due to the different loading time characteristics associated with each of the various image file types, the threshold can be based on the described image file type. That is, the threshold can be based on a set loading time, where the file size threshold can be determined according to the estimated trend for each image file type.
[0148] In addition, in various embodiments, composite text and image classification methods are envisioned. More specifically, composite BERT and ViT model architectures are used for more efficient text and image classification. Figure 15 is a flowchart of a composite text and image classification architecture 600. Figure 15The composite text and image classification architecture 600 shown in [Figure] operates based on the classification generated by the image model 602. If the image model generates a classification prediction of "other", the image will be passed to the OCR engine 604, then the text will be extracted and subsequently processed by the text model 606. For text-based documents, the text will be directly extracted and passed to the text model 606. This enables it to utilize the high accuracy of a single model, customize the post-processing steps of each model, and reduce OCR calls by first classifying the image data.
[0149] As described above, Figure 15 The image model 602 shown in [Figure] can include a ViT model, while the text model 606 can include the BERT model described herein, but the composite architecture can be used with any model for text and image processing described herein. Additionally, these models can include any modifications described in this disclosure to reduce latency while maintaining the accuracy of classification.
[0150] As previously mentioned, the inline multi-modal DLP system can be used for inline production data. That is, through network configurations 100A, 100B, 100C, etc. for processing data flowing through the cloud 120, i.e., as part of the inline monitoring described herein.
[0151] Figure 16 is a flowchart of a process 650 for inline multi-modal DLP. Process 650 contemplates being implemented as a method having steps, via computing resources configured to implement these steps, and as a non-transitory computer-readable medium having instructions that, when executed, cause one or more processors to implement the steps. Process 650 can be implemented using the multi-modal DLP systems 550, 400, and the actual implementation of the multi-modal DLP systems 400 and processes 450, 480, 500, 550 can be achieved through network configurations 100A, 100B, 100C, etc. That is, process 650 contemplates being used with any network security monitoring platform, device, service, etc.
[0152] Process 650 includes training one or more machine learning models for classifying input data into categories of multiple classes (step 652); performing one or more modifications to the one or more machine learning models, where the one or more modifications reduce the latency associated with the one or more machine learning models (step 654); receiving an input including data in any of multiple formats (step 656); processing the input to classify the input into categories of multiple classes (step 658); and providing an indication of the category of the multiple classes (step 660).
[0153] The process 650 can also include: wherein, one or more modifications include removing non-English words from the vocabulary of one or more machine learning models, removing stop words, and performing lemmatization. One or more modifications can include implementing any one of a text byte lower threshold, a text byte upper threshold, and an early stopping k value. One or more modifications can include implementing a maximum input file size. The maximum input file size can be based on the type of input file. The maximum input file size can be determined based on a trend graph of one or more estimated loading times versus image size. One or more machine learning models can include an image classification model and a text classification model, and wherein the steps further include: in response to the image model producing a classification prediction of "other", extracting text from the image via an optical character recognition (OCR) engine; and processing the extracted text via the text classification model. These steps can also include processing the input to determine whether the data includes sensitive data before processing the input for classification. The various formats can include text format, image format, audio format, video format, source code, and combinations thereof. These steps can also include, before training one or more machine learning models with a set of training documents with tags, identifying any mislabeled documents therein by performing optical character recognition (OCR) and checking for the presence of relevant keywords.
[0154] §5.0 Knowledge Distillation and LLM for Document Classification
[0155] As described above, current DLP processes include using various models to detect and classify sensitive data. These models can be optimized in various ways, such as the various mechanisms described herein. The present disclosure provides further optimization for these models to ensure efficiency, high performance, and small model size, thus making deployment easier and reliability higher. In various embodiments, the method also includes developing an LLM for document classification.
[0156] In various embodiments, knowledge distillation is used to create an efficient model that is suitable for running well with fewer computing resources. Additionally, high-quality relevant data points are selected and utilized to maximize the efficiency of the knowledge distillation process.
[0157] Knowledge distillation is a process that aims to transfer the broad knowledge of a large and complex model to a smaller and lighter model to generate a small distilled model. In the field of machine learning, larger models typically have a large number of learnable parameters, enabling them to achieve excellent performance in various tasks. In contrast, smaller models are characterized by a reduced number of parameters and usually sacrifice performance for efficiency. Even when training two models on the same dataset, this difference still exists, highlighting the challenge of achieving comparable performance in resource-constrained situations. Knowledge distillation serves as a bridge between these contrasting model sizes, enabling the smaller "student" model to not only learn from the dataset but also replicate the powerful performance achieved by its larger "teacher" model.
[0158] By distilling the essence of the teacher model's knowledge into a more compact form, the student model gains rich insight capabilities, enabling it to refine its prediction and decision-making processes. This deeper understanding ultimately leads to more accurate and reliable results, allowing the student model to reach a performance level comparable to that of the more powerful teacher. Additionally, knowledge distillation plays a crucial role in developing task-specific versions of large models for deployment in real-world environments.
[0159] Figure 17 is a flowchart of the knowledge distillation process 700. Process 700 includes obtaining predictions from the teacher model 702. Predictions are also obtained from the student model 704. These predictions can be document classification predictions as described herein or any other output generated by an LLM. Similarly, the teacher model 702 is a larger model compared to the smaller student model 704. Then, the predictions from the teacher model 702 and the student model 704 are compared to determine the differences between them, which are described as calculating the loss 706 between the two models. Calculating the loss can involve summing the differences between the teacher model and the student model. This information is then fed back into the student model via backpropagation 708, allowing the student model 704 to learn from its mistakes and attempt to replicate the teacher model 702. In various embodiments, the resulting model is referred to as a distilled model. That is, the optimized student model is called a distilled model.
[0160] Figure 18 is a flowchart for implementing knowledge distillation within the present system and method. The teacher model 702 is used to make predictions on a common dataset, e.g., a dataset including various inputs associated with various categories. These predictions are stored as a distilled dataset 710. Then, knowledge distillation 712 is performed using the student model 704 and the distilled dataset. This enables the system to create a new model 714 for classification. The distilled dataset 710 is envisioned as the output of the teacher model, where, as described herein, knowledge distillation is performed using the differences from the output of the student model.
[0161] Figure 19 Represents multiple experiments of models for optimizing content classification using different methods. In the first method 720-A, the teacher model 702 is used to make predictions 716 on the DLP dataset. In the second method 720-B, the student model 704 is used to make predictions 716 on the DLP dataset. In the third method 720-C, the student model 704 is optimized through knowledge distillation 712 and then used to make predictions 716 on the DLP dataset. In the fourth method 720-D, the student model 704 is optimized via knowledge distillation 712, then fine-tuned 718 on the DLP data, and then used to make predictions 716 on the DLP dataset. Finally, in the fifth method 720-E, the student model 704 is only fine-tuned 718 using the DLP data and then used to make predictions 716 on the DLP dataset.
[0162] Based on these different methods, the following accuracy measurements were obtained: [first method 720-A, 48.4%], [second method 720-B, 12.8%], [third method 720-C, 15%], [fourth method 720-D, 53%], [fifth method 720-E, 52.3%]. Similarly, these accuracy measurements are associated with the percentage of correct classifications that each model can provide after its respective optimization method. Based on this, it can be seen that using only the teacher model 702 results in an accuracy of 48.4%, while using only the unoptimized student model 704 results in an accuracy of 12.8%. Also, the purpose of the present disclosure is to reduce the size of the models used in production, so using the teacher model is not desirable. Additionally, it can be seen that simply using knowledge distillation 712 only slightly increases the accuracy of the student model from 12.8% to 15%, while using the combination of knowledge distillation 714 and fine-tuning 718 results in a more acceptable accuracy of 53%. Similarly, this accuracy is achieved by the optimized student model 704, which is much smaller in size than the teacher model. Furthermore, it can be seen that the difference in accuracy between using only fine-tuning 718 and using the combination of knowledge distillation 712 and fine-tuning 718 is small, although a closer look at the accuracy, recall, and scores for specific content categories provides more details.
[0163] Figure 20 Is a comparison of the category classification metrics between the fine-tuned model and the model that undergoes knowledge distillation and fine-tuning. It can be seen that when using the combination of knowledge distillation 712 and fine-tuning 718 instead of only fine-tuning, the accuracy metric increases significantly from 55% to 64%. Therefore, the accuracy metric for anything classified by the model optimized via the combination of knowledge distillation 712 and fine-tuning 718 increases by an average of 9%. Additionally, the F1 for 8 out of 12 categories in the DLP dataset also increases.
[0164] In addition, in various embodiments, the use of a custom dataset can be considered. Typically, only public / general datasets are used for training the model and / or for the knowledge distillation process. The present system and method include analyzing the drawbacks of the teacher model 702. For example, the teacher model may perform well in certain categories but poorly in other categories. Based on this, a synthetic dataset is created for the categories in which the teacher model performs well, i.e., a synthetic dataset is generated based on the strengths of the teacher model. This synthetic dataset can be created using an LLM. This can be achieved by querying the LLM to create specific data related to the categories in which the teacher model performs well. For example, the LLM can be asked to create a specific number of resume documents, tax documents, etc. The synthetic data is then passed to the teacher model to provide predictions, and the predictions are then stored in the distillation dataset 710. The student model for such specific category data can then be trained using this distilled data.
[0165] In various embodiments, multiple teacher models can be utilized to create a comprehensive distillation dataset, with each teacher model having an advantage in different categories. That is, each teacher model can be exposed to specific category data related to its strength, and the resulting prediction results can be combined into a single distillation dataset to optimize a single student model.
[0166] Figure 21 A performance graph of a standard student model and a standard teacher model is shown. Standard means the model is trained using general non-synthetic data (i.e., a general data loss protection (DLP) dataset). The output / predictions generated by these models can be regarded as general data predictions. It can be seen that the accuracy rate of the student model's execution is 35%, while that of the teacher model's execution is 67%. It can also be seen that the model performs relatively well except for 3 categories. For example, in the automotive category, the accuracy rate of the student model is 38%, while that of the teacher model is 85%. Based on this, in addition to all other well-performing categories, a synthetic dataset is created for the automotive category. Similarly, this synthetic data is further used to create a distillation dataset, which can be used to optimize the student model. In various embodiments, the "well-performing" categories can be determined based on an accuracy threshold.
[0167] Figure 22 It is a performance graph comparing student models optimized via various methods. Figure 22The different methods highlighted include the standard "unoptimized" student model 802, the student model 804 distilled with general data (i.e., data not including specific category synthetic data), the student model 806 distilled with specific category data, and the student model 808 distilled with specific category data and general data. Similarly, the standard student model has an accuracy rate of 35%. When using the student model distilled with general data, the accuracy rate increases to 37%. However, when only using the synthetic data of specific categories, the accuracy rate increases to 39%. Similarly, the student model distilled with specific category data and general data also shows an accuracy rate of 39%. It is worth noting that in this example, the general dataset includes 5000 data points, while the specific category synthetic data only includes 90 data points. Therefore, using the specific category synthetic data is more efficient and beneficial.
[0168] Based on the described process, various embodiments of the present disclosure utilize a combination of knowledge distillation and fine-tuning to optimize the student model. Fine-tuning can include creating synthetic data for specific categories, feeding the synthetic data into the teacher model to generate synthetic data predictions, and creating a distilled dataset including these synthetic data predictions. Then, the "fine-tuned" distilled dataset can be used to distill the student model and create a distilled model, thereby allowing existing DLP systems to use smaller models to perform content classification in production.
[0169] It should be understood that this model optimization process can be used to create optimized models for use in any DLP process described herein, such as processes 450, 480, 500, 550, and 650. In addition, the optimizations described in the present disclosure can be combined with additional optimizations such as those described in process 650. Furthermore, those skilled in the art will understand that the model optimization processes described herein can be used to optimize models outside of the DLP use cases described herein. That is, these processes can be used to optimize models performing various tasks and are not limited to content classification models.
[0170] By utilizing the present system and method, the performance of smaller models can be improved. By utilizing larger distilled datasets, more specific category datasets, and even larger and more accurate teacher models, the effectiveness of smaller models can be further enhanced. Similarly, the methods discussed are based on DLP datasets, although these techniques can be used in any area of LLM. Currently, most of the available high-precision LLMs are relatively large and not sufficient to meet the customer-centric user experience. However, by utilizing the present system and method, smaller, faster, and more accurate models can be built.
[0171] §5.1 LLM Knowledge Distillation Process
[0172] Figure 23It is a flowchart of process 850 for LLM knowledge distillation for data loss protection (DLP). Process 850 is envisioned as a method with steps, via computing resources configured to implement these steps, and implemented as a non-transitory computer-readable medium with instructions that, when executed, cause one or more processors to implement the steps. Process 850 can be implemented using multimodal DLP systems 550, 400, and the actual implementation of multimodal DLP systems 400 and processes 450, 480, 500, 550, and 650 can be achieved through network configurations 100A, 100B, 100C, etc. That is, process 850 is contemplated for use with any network security monitoring platform, device, service, etc.
[0173] Process 850 includes receiving multiple general data predictions from a teacher model (step 852); determining one or more advantages of the teacher model based on the received general data predictions (step 854); generating a synthetic dataset based on the one or more advantages of the teacher model (step 856); providing the synthetic dataset to the teacher model and receiving multiple synthetic data predictions from the teacher model based thereon (step 858); and performing knowledge distillation on a student model based on the synthetic data predictions received from the teacher model to produce a distilled model (step 860).
[0174] Process 850 can also include: wherein, the teacher model and the student model are large language models (LLMs). Before receiving multiple general data predictions from the teacher model, the steps can include providing a general data loss protection (DLP) dataset to the teacher model. The multiple general data predictions and the multiple synthetic data predictions can include content category classification predictions. Determining one or more advantages of the teacher model can include determining one or more categories for which the teacher model performs classification with an accuracy higher than a threshold. Generating the synthetic dataset can include using a large language model (LLM) to generate multiple inputs associated with the one or more advantages of the teacher model, wherein the synthetic dataset includes the multiple inputs. The steps can also include using the distilled model to classify inputs to a data loss protection (DLP) system in production. The steps can also include receiving an input including data in any one of multiple formats; processing the input via the distilled model to classify the input into a category of multiple categories; and providing an indication of the category of the multiple categories. The steps can also include, before processing the input for classification, processing the input to determine whether the data includes sensitive data. The multiple formats can include text format, image format, audio format, video format, source code, and combinations thereof.
[0175] §6.0 Conclusion
[0176] It should be understood that some embodiments described herein may include one or more general-purpose or special-purpose processors (“one or more processors”), such as: a microprocessor; a central processing unit (CPU); a digital signal processor (DSP); a custom processor (e.g., a network processor (NP) or a network processing unit (NPU)), a graphics processing unit (GPU), etc.; a field programmable gate array (FPGA); and unique stored program instructions for controlling their implementation (including software and / or firmware) to combine certain non-processor circuits to implement some, most, or all of the functions described in the methods and / or systems herein. Alternatively, some or all of the functions may be implemented by a state machine without stored program instructions, or in one or more application specific integrated circuits (ASICs), where each function or some combination of certain functions is implemented as custom logic or circuitry. Of course, combinations of these methods may be used. For some of the embodiments described herein, the hardware and optionally the corresponding devices in software, firmware, and their combinations can be referred to as “circuitry configured or adapted to”, “logic configured or adapted to”, “circuit configured to”, “one or more circuits configured to”, etc., for performing a set of operations, steps, methods, processes, algorithms, functions, techniques, etc. of the data described in the various embodiments herein.
[0177] In addition, some embodiments may include a non-transitory computer-readable storage medium having computer-readable code stored thereon for programming a computer, server, appliance, device, processor, circuitry, etc., where each of these may include a processor for performing the functions described and claimed herein. Examples of such computer-readable storage media include, but are not limited to, hard disks, optical storage devices, magnetic storage devices, read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash memory, etc. When stored in a non-transitory computer-readable medium, software may include instructions executable by a processor or device (e.g., any type of programmable circuitry or logic), and in response to such instructions being executed, cause the processor or device to perform a set of operations, steps, methods, processes, algorithms, functions, techniques, etc. described in the various embodiments herein.
[0178] Although the present disclosure has been illustrated and described with reference to embodiments and specific examples thereof, those of ordinary skill in the art will readily appreciate that other embodiments and examples can perform similar functions and / or achieve similar results. All such equivalent embodiments and examples are within the spirit and scope of the present disclosure and are thus contemplated to be covered by the appended claims. In addition, the various elements, operations, steps, methods, processes, algorithms, functions, techniques, etc. described herein are contemplated to be used in any and all combinations with each other, including use alone and partial use of various combinations of elements, operations, steps, methods, processes, algorithms, functions, techniques, etc.
Claims
1. A method (650) for multimodal data loss protection (DLP) includes the following steps: Receiving (652) an input including data in any one of a plurality of formats; Processing (654) the input to determine whether the data includes sensitive data, and In response to (656) the input including sensitive data, performing the following steps: Processing the input to classify the input into categories of a plurality of categories; and Providing an indication of the category of the plurality of categories.
2. The method (650) according to claim 1, wherein, The steps include: In response to the input including non-sensitive data, providing an indication that the data is non-sensitive, thereby allowing data in transit or not marking data at rest.
3. The method (650) according to any one of claims 1 to 2, wherein The steps include: In response to the input including sensitive data, providing an indication of the category of the plurality of categories and sub-categories associated with the category.
4. The method (650) according to any one of claims 1 to 3, wherein, The plurality of formats include text format, image format, audio format, video format, source code, and combinations thereof.
5. The method (650) according to any one of claims 1 to 4, wherein The processing of the input to determine whether the data includes sensitive data utilizes (1) a large language model (LLM) and embeddings, and (2) a machine learning model configured for classification.
6. The method (650) according to any one of claims 1 to 5, wherein, The processing of the input to classify the input into categories utilizes (1) a large language model (LLM) and (2) a zero-shot classifier.
7. The method (650) according to any one of claims 1 to 6, wherein, Both the processing of the input to determine whether the data includes sensitive data and the processing of the input to classify the input into categories utilize one or more machine learning models trained based on a set of training documents with labels.
8. The method (650) according to claim 7, wherein, The steps include: Before training one or more machine learning models with a set of training documents with labels, identifying any mislabeled documents in the training document set by performing optical character recognition (OCR) and checking for the presence of relevant keywords.
9. The method (650) according to any one of claims 7 to 8, wherein, The steps include: Before training one or more machine learning models with a set of training documents with labels, filtering out images in the training document set based on file size.
10. The method (650) according to any one of claims 7 to 9, wherein, The steps include: Before training one or more machine learning models with a set of training documents with labels, grouping images with minor differences in the training document set based on the hash of the images.
11. The method (650) according to any one of claims 1 to 10, wherein, The processing of the input to classify the input is performed by a machine learning model, and the steps further include: Training the machine learning model to classify input data into one of a plurality of categories; and Modifying the trained machine learning model to reduce latency therein.
12. The method (650) according to claim 11, wherein, The modification includes removing various words from the vocabulary of the trained machine learning model.
13. The method (650) according to claim 12, wherein, The words include any words among non-English words and stop words.
14. A server (200) includes one or more processors (202) and a memory (210) storing instructions that, when executed, cause the one or more processors (202) to implement the method (650) according to any one of claims 1 to 13.
15. Computer-readable instructions that, when executed, cause one or more processors (202) to implement the method (650) according to any one of claims 1 to 13.
Citation Information
Patent Citations
Data loss prevention on images
US11805138B2
Cloud-based data loss prevention
US11829347B2