Image-based identification of documents to protect confidential information
A deep learning-based CNN model enhances image classification in DLP systems by using online learning and storing extracted features, addressing accuracy and privacy issues in detecting sensitive information in images and screenshots.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- NETSKOPE INC
- Filing Date
- 2025-12-25
- Publication Date
- 2026-04-10
AI Technical Summary
Existing DLP technologies face challenges in accurately detecting sensitive information in images, particularly for non-ideal conditions, and require large numbers of labeled images, which raises privacy concerns and is computationally intensive.
A deep learning-based approach using a convolutional neural network (CNN) model is applied to detect sensitive information in images without requiring a large number of pre-labeled images, utilizing online learning and progressive refinement with specialized labeled images, and storing extracted features instead of raw images to maintain privacy.
Enhances detection accuracy and performance for sensitive information in images and screenshots, improving threat detection effectiveness by 20-25% while minimizing privacy concerns and computational overhead.
Smart Images

Figure 2026062834000001_ABST
Abstract
Description
Priority Claim
[0001] This application claims priority to U.S. Patent Application No. 16 / 891,647, filed on June 3, 2020, titled "Detection of Image-Derived Identification Documents for Protecting Confidential Information" (Attorney Docket No. NSKO1032-1) (currently U.S. Patent No. 10,990,856, issued on April 27, 2021), U.S. Patent Application No. 17 / 229,768, filed on April 13, 2021, titled "Deep Learning Stack Used in Production to Prevent Exfiltration of Image-Derived Identification Documents" (Attorney Docket No. NSKO1032-2), and
[0002] This application claims priority to U.S. Patent Application No. 16 / 891,678, filed on June 3, 2020, titled "Detection of Screenshot Images to Prevent Loss of Confidential Screenshot-Originated Data" (Attorney Docket No. NSKO1033-1) (currently U.S. Patent No. 10,949,961, issued on March 16, 2021), U.S. Patent Application No. 17 / 202,075, filed on March 15, 2021, titled "Training and Configuration of a DL Stack to Detect Attempted Exfiltration of Confidential Screenshot-Originated Data" (Attorney Docket No. NSKO1033-2), and
[0003] This application claims priority to U.S. Patent Application No. 16 / 891,698, filed on June 3, 2020, titled "Detection of Confidential Documents Derived from Organizational Images and Prevention of Loss of Confidential Documents" (Attorney Docket No. NSKO 1034-1) (currently U.S. Patent No. 10,867,073, issued on December 15, 2020), U.S. Patent Application No. 17 / 116,862, filed on December 9, 2020, titled "Deep Learning-Based Detection of Image-Derived Confidential Documents and Data Loss Prevention" (Attorney Docket No. NSKO 1034-2). These applications are hereby incorporated by reference in their entirety for all purposes. Combined materials
[0004] The following materials are hereby incorporated by reference into this application:
[0005] U.S. Patent Application No. 16 / 807,128, filed March 2, 2020, title of invention: "Load Balancing in Dynamically Scalable Service Meshes" (Agent Reference Number NSKO1025-3).
[0006] U.S. Patent Application No. 14 / 198,508, filed on March 5, 2014, with the title "Security for Network Distribution Services" (Agent Reference Number NSKO1000-3) (currently U.S. Patent No. 9,270,765, issued on February 23, 2016).
[0007] U.S. Patent Application No. 14 / 198,499, filed on March 5, 2014, with the title "Security for Network Distribution Services" (Agent Reference Number NSKO1000-2) (currently U.S. Patent No. 9,398,102, issued on July 19, 2016).
[0008] U.S. Patent Application No. 14 / 835,640, filed on August 25, 2015, with the title "System and Method for Monitoring and Controlling Enterprise Information Stored in Cloud Computing Services (CCS)" (Agent Reference Number NSKO1001-2) (currently U.S. Patent No. 9,928,377, issued on March 27, 2018).
[0009] U.S. Provisional Application No. 62 / 307,305, filed March 11, 2016, Title of Invention: "Multi-performance in Data Loss Transactions of Cloud Computing Services" U.S. Patent Application No. 15 / 368,246, filed 2 December 2016, title "Middleware Security Layer for Cloud Computing Services" (Agent Reference No. NSKO1003-3), claims the interest of a system and method for implementing a security policy (Agent Reference No. NSKO1003-1).
[0010] Chen, Itar, Narayanaswamy, and Malmskog, "Cloud Security for Dummy Data, Netscope Special Edition," John Wiley & Sons, 2015.
[0011] "Netskope Introspection," published by Netskope, Inc.
[0012] "Data Loss Prevention and Monitoring in the Cloud," published by Netskope, Inc.
[0013] "Cloud Data Loss Prevention Reference Architecture," published by Netskope, Inc.
[0014] "Five Steps to Cloud Confidence," published by NetScope, Inc.
[0015] "Netscope Active Platform" Netscope, Inc. Published by tskope, Inc.
[0016] "Netskope Advantage: Three Essential Requirements for Cloud Access Security Brokers" Netskope, Inc. issue.
[0017] "15 Important CASB Use Cases," published by Netskope, Inc.
[0018] "Netskope Active Cloud DLP," published by Netskope, Inc.
[0019] "Repairing the Collision Course of Cloud Data Breach," published by Netskope, Inc.
[0020] "Netskope Cloud Confidence Index (trademark)," published by Netskope, Inc.
[0021] The above materials are incorporated here by reference as if they were fully described. [Technical Field]
[0022] The disclosed technologies generally relate to security for network delivery services, and in particular to detecting identifiable documents within images, known as image-derived identifiable documents, and preventing the loss of such documents while applying security services. The disclosed technologies also relate to detecting screenshot images and preventing the loss of screenshot-derived data. Furthermore, separate organizations can use the disclosed technologies to detect image-derived identifiable documents and screenshot images from within their own organizations, so that images of the organization containing potentially sensitive data do not need to be shared with data loss prevention service providers. [Background technology]
[0023] The subject matter discussed in this section should not be assumed to be merely prior art as a result of its reference in this section. Similarly, the problems related to the subject matter presented in this section or the background art should not be assumed to be already recognized in the prior art. The subject matter in this section merely presents various approaches and may, on its own or spontaneously, correspond to the implementation of the claimed art.
[0024] Data loss prevention (DLP) technology is widely used in the security industry to prevent the leakage of sensitive information such as personally identifiable information (PII), protected health information (PHI), and intellectual property (IP). Both large and small businesses use DLP products. Such sensitive information resides in various sources, including documents and images. It is crucial that any DLP product can detect sensitive information within documents and images with high accuracy and computational efficiency.
[0025] For text documents, DLP products use string and regular expression-based pattern matching to identify sensitive information. For images, optical character recognition (OCR) technology was initially used to extract text characters. The extracted characters are then sent to the same pattern matching process to detect sensitive information. Historically, OCR has been computationally intensive and has not performed well, especially when images are not in ideal conditions, such as being blurry, dirty, rotated, or inverted, due to insufficient accuracy.
[0026] While training can be automated, the problem remains of assembling the training data in the correct format and sending it to a central computing node with sufficient memory capacity and computing power. In many fields, transmitting personally identifiable private data to any central authority raises data privacy concerns, including data security, data ownership, privacy protection, and the proper authorization and use of data.
[0027] Deep learning applies a multi-layer network to data. In recent years, deep learning technology has been increasingly used in image classification. Deep learning can detect images with confidential information without going through expensive OCR processing. An important issue with the deep learning approach is the need for a large number of high-quality labeled images that represent the real-world distribution. Unfortunately, in the case of DLP, high-quality labeled images typically utilize real images with confidential information such as genuine passport images and genuine driver's license images. These data sources are inherently difficult to obtain on a large scale. This limitation hinders the adoption of deep learning-based image classification in DLP products.
[0028] There is an opportunity to efficiently detect identification documents within an image, with an improvement in threat detection effectiveness of about 20 - 25%, and prevent the loss of confidential data of the image-derived identification document. Furthermore, there is an opportunity to detect screenshot images and prevent the loss of data from confidential screenshots, which could result in cost and time savings in the security systems utilized by customers using SaaS.
[0029] In the drawings, like reference numerals generally refer to like parts throughout different figures. Also, the drawings are not necessarily to scale; rather, emphasis is generally placed on illustrating the principles of the disclosed technology. In the following description, various embodiments of the disclosed technology are described below with reference to the following drawings.
Brief Description of the Drawings
[0030] [Figure 1A] A schematic diagram at the system architecture level for detecting identification documents within an image, called image-derived identification documents, while applying security services in the cloud, and preventing the loss of image-derived identification documents is shown. The disclosed system can also detect screenshot images and prevent the loss of data from confidential screenshots.
[0031] [Figure 1B] This document describes an architecture for detecting sensitive data originating from images, specifically identifying documents within images, known as image-derived identification documents, and preventing the loss of these documents while applying security services in the cloud. It also describes an architecture for detecting screenshot images and preventing the loss of sensitive screenshot-derived data.
[0032] [Figure 2] The diagram shows a configuration of a deep learning stack implemented using a convolutional neural network architecture model for image classification, which can be configured for use in a system for detecting identified documents within images and for detecting screenshot images, according to one embodiment of the disclosed technology.
[0033] [Figure 3] This shows the accuracy and recall results of the trained passport and driver's license classifiers.
[0034] [Figure 4] The execution time results for classifying images are shown as a graph of the image distribution.
[0035] [Figure 5] This shows benchmarking results for classifying sensitive images on U.S. driver's licenses.
[0036] [Figure 6] This example workflow demonstrates how to train a deep learning stack to detect identified documents within images, known as image-derived identified documents, and prevent the loss of these documents.
[0037] [Figure 7] An example screenshot showing an inventory list with listed costs is provided.
[0038] Figures 8A, 8B, 8C, and 8D show four false positive screenshot images.
[0039] [Figure 8A] This shows the Idaho map, which was misclassified as a screenshot due to the legend window and the dotted lines at the top and bottom.
[0040] [Figure 8B] The entire image is a window containing PII within a black background, and the UNITED STATES bar can be treated as a header bar, thus showing a driver's license image that has been misclassified as a screenshot.
[0041] [Figure 8C] The main window, including the PII, displays a passport image, and the shaded area at the bottom center could mislead the classifier into thinking it's an application bar.
[0042] [Figure 8D] This displays text information and characters within a measure window, including a uniform background.
[0043] [Figure 9] This is a simplified block diagram of a computer system that can be used to detect screenshot images and prevent the loss of image-derived identifiable documents, and which can be used to detect screenshot images and prevent the loss of image-derived screenshots, according to one embodiment of the disclosed technology.
[0044] [Figure 10] This example workflow demonstrates how to train a deep learning stack to detect identified documents within images, known as image-derived identified documents, and prevent the loss of these documents.
[0045] [Figure 11]This document describes a workflow for one or more computer systems that can be configured to perform identification document detection within images and prevent the loss of image-derived identification documents, and can be used to detect screenshot images and prevent the loss of image-derived screenshots. [Modes for carrying out the invention]
[0046] The following detailed description will be made with reference to the drawings. Exemplary embodiments are described not to limit the technical scope defined by the claims, but to illustrate the disclosed technology. Those skilled in the art will recognize various equivalent variations of the following description.
[0047] By using deep learning technology, the detection of sensitive information from documents and images can be enhanced, enabling the detection of images containing sensitive information without the need for existing, expensive OCR processing. Deep learning uses optimization to find the optimal parameter values for the model to make the best predictions. Image classification based on deep learning typically requires a large number of labeled images containing sensitive information, which are difficult to acquire on a large scale. This constraint hinders the adoption of image classification based on deep learning in DLP products.
[0048] The disclosed innovation applies deep learning-based image classification to data loss prevention (DLP) products without requiring a large number of pre-labeled images containing sensitive information. Many pre-trained general-purpose deep learning models available today use the public ImageNet dataset and other similar sources. These deep learning models are typically multi-layer convolutional neural networks (CNNs) capable of classifying common objects such as cats, dogs, and cars. The disclosed technique uses a small number of specialized labeled images, such as passport and driver's license images, to retrain the last few layers of the CNN model. In this way, the deep learning (DL) stack can detect these specific images with high accuracy without requiring a large number of labeled images containing sensitive data.
[0049] The DLP product deployed to the customer can handle the customer's production traffic and continuously generate new labels. To minimize privacy issues, new labels can be kept within the production environment using online learning, and whenever a sufficient batch of new labels has accumulated, a new, balanced incremental dataset can be created that can be used to progressively refine existing deep learning models using progressive learning by injecting a similar number of negative images.
[0050] Even with online learning and progressive learning, typical deep learning processes use original and newly added image inputs to create sophisticated models for predicting the presence of sensitive data within image documents or screenshots. This is necessary. This means the system needs to store newly labeled images generated in production for extended periods in production. While user private data is more secure in a production environment than when images and labels are stored offline, storing images presents privacy issues if sensitive data is stored in persistent storage.
[0051] The disclosed method stores the output of a deep learning stack, also known as a neural network, and stores the extracted features instead of the raw image. In a typical neural network, the raw image passes through many layers before the final set of features is extracted for the final classifier. These features cannot be reverse-transformed back into the original raw image. This feature of the disclosed technique allows for the protection of sensitive information in the production image, and the stored features of the model can be used to retrain the classifier in the future.
[0052] The disclosed technology provides accuracy and high performance when classifying images and screenshots containing sensitive information without requiring a large number of pre-labeled images. This technology also enables the use of production images to continuously improve accuracy and coverage without privacy concerns.
[0053] The disclosed innovations leverage machine learning classification to further enhance the ability to detect sensitive image content and enforce policies, applying advances in image classification and screenshot detection to network traffic proxied within the cloud, as described herein, in the context of Netscope Cloud Access Security Broker (N-CASB).
[0054] The following describes an example of a system that detects identification documents within images, known as image-derived identification documents, to prevent the loss of such documents in the cloud, and also detects screenshots to prevent the loss of confidential screenshot-derived data. [architecture]
[0055] Figure 1A shows a schematic, architecture-level diagram of System 100 for detecting identifiable documents within images, known as image-derived identifiable documents, and preventing the loss of image-derived identifiable documents in the cloud. System 100 can also detect screenshot images and prevent the loss of sensitive screenshot-derived data. As Figure 1A is an architecture diagram, certain details have been intentionally omitted to improve clarity of explanation. The explanation of Figure 1A is organized as follows: First, the elements of the diagram are described, then their interconnections are described, and then the use of the elements in the system is described in more detail. Figure 1B shows the system's method of detecting image-derived sensitive data, which will be described later.
[0056] System 100 includes an organizational network 102, a data center 152 having a Netscope Cloud Access Security Broker (N-CASB) 155, and cloud-based services 108. System 100 also includes multiple organizational networks 104 for multiple subscribers of a security service provider, also called a multi-tenant network, and multiple data centers 154, sometimes called branches. Organizational network 102 includes computers 112a-n, tablets 122a-n, mobile phones 132a-n, and smartwatches 142a-n. Additional devices may be available to organizational users in other organizational networks. Cloud services 108 include cloud-based hosting services 118, webmail services 128, video, messaging, and voice call services 138, streaming services 148, file transfer services 158, and cloud-based storage services 168. Data center 152 connects to organizational network 102 and cloud-based services 108 via a public network 145.
[0057] Continuing the explanation in Figure 1A, the disclosed Enhanced Netscope Cloud Access Security Broker (N-CASB) 155 manages access and activity in authorized and unauthorized cloud applications, protects sensitive data, prevents its loss, and protects against internal and external threats. In addition, it securely handles P2P traffic via BT, FTP, and UDP-based streaming protocols, as well as web traffic via SIP, Skype, voice, video, and messaging multimedia communication sessions, and other protocols. To prevent data loss, N-CASB 155 utilizes machine learning classification for identity detection and sensitive screenshot detection, and further extends its ability to detect sensitive image content and enforce policies. N-CASB 155 includes an Active Analyzer 165 and an Introspective Analyzer 175 that identify system users and set policies for applications. The Introspective Analyzer 175 directly interacts with cloud-based services 108 to inspect idle data. In polling mode, the introspective analyzer 175 uses API connectors to call cloud-based services, crawling data residing in those services and checking for changes. For example, the Box® storage application provides a management API called the Box Contents API®. This management API provides visibility into all users' organizational accounts, including audit logs for Box folders, and inspecting these can determine whether sensitive files have been downloaded since a specific date when credentials were compromised. The introspective analyzer 175 polls this API to discover any changes made to any of the accounts. When a change is detected, the Box Events API® is polled to discover more detailed data changes. In the callback model, the introspective analyzer 175 registers with cloud-based services via API connectors to be notified of critical events.For example, the Introspective Analyzer 175 can use the Microsoft Office 365 Webhooks API (trademark) to determine when a file was shared externally. The Introspective Analyzer 175 also has deep API inspection, deep packet inspection, and log inspection capabilities, and includes a DLP engine that applies various content inspection techniques to static files within cloud-based services to determine which documents and files are confidential based on policies and rules stored in storage 186. The inspection results from the Introspective Analyzer 175 generate per-user and per-file data.
[0058] Continuing the explanation in Figure 1A, the N-CASB155 further comprises a monitor 184 including an extraction engine 171, a classification engine 172, a security engine 173, a management plane 174, and a data plane 180. The N-CASB155 also further comprises storage 186 containing information for deep learning stack parameters 183, features and labels 185, content policies 187, content profiles 188, content inspection rules 189, corporate data 197, customers 198, and user identities 199. Corporate data 197 may include, but is not limited to, organizational data such as intellectual property, confidential financial information, strategic plans, customer lists, personally identifiable information (PII) belonging to customers or employees, patient health data, source code, trade secrets, reservation information, partnership agreements, corporate plans, merger and acquisition documents, and other confidential data. In particular, the term “corporate data” refers to documents, files, folders, web pages, collections of web pages, images, or other text-based documents. User identity refers to an indicator provided to a client device by a network security system, in the form of a token, a unique identity such as a UUID, or a public key certificate. In some cases, user identity can be linked to a specific user and a specific device. Thus, the same individual may have different user identities on a mobile phone and a computer. User identity can be linked to a corporate identity directory of entries or user IDs, but this is different. In one embodiment, a cryptographic certificate signed by the network security system is used as the user identity. In other embodiments, the user identity is unique to the user only and can be identical across devices.
[0059] The embodiments may also interoperate with single sign-on (SSO) solutions and / or corporate identity directories such as Microsoft Active Directory. Such embodiments may use custom attributes to allow policies to be defined within the directory, for example, at the group or user level. Host services configured in the system are also configured to require traffic to pass through the system. This can be done by setting IP range restrictions on the host service to the IP range of the system and / or the integration between the system and the SSO system. For example, integration with an SSO solution may enforce client presence requirements before allowing sign-on. Other embodiments may use a “proxy account” with the SaaS vendor, for example, a dedicated account held by a system that holds the sole credential for signing in to the service. In other embodiments, the client may encrypt the sign-on credential before passing the login to the host service, meaning that the network security system “owns” the password.
[0060] Storage 186 can store information from one or more tenants in a common database image table to form an on-demand database service (ODDS) that can be implemented in many ways, such as a multi-tenant database system (MTDS). The database image may contain one or more database objects. In other embodiments, the database may be a relational database management system (RDBMS), an object-oriented database management system (OODBMS), a distributed file system (DFS), a schema-less database, or any other data storage system or computing device. In some embodiments, the collected metadata is processed and / or normalized. In some cases, the metadata includes structured data and functional target-specific data structures provided by the cloud service 108. Unstructured data, such as free text, may also be provided by the cloud service 108 and can be returned to the cloud service 108. Both structured and unstructured data can be aggregated by the introspective analyzer 175. For example, assembled metadata can be stored in semi-structured data formats such as JSON (JavaScript Option Notation), BSON (Binary JSON), XML, Protobuf, Avro, or Thrift objects. These consist of string fields (or columns) and corresponding values of potentially various types, such as numbers, strings, objects, arrays, etc. JSON objects can be nested, and fields can be multi-valued into arrays, nested arrays, etc., in other embodiments.These JSON objects are stored in schema-less or NoSQL key-value metadata stores such as Apache Cassandra™ 158, Google's BigTable™, HBase™, Voldemort™, CouchDB™, MongoDB™, Redis™, Riak™, Neo4j™, etc., which store the parsed JSON objects using keyspaces equivalent to databases in SQL. Each keyspace is similar to a table and is divided into column families, each containing a set of rows and columns.
[0061] In one embodiment, the introspective analyzer 175 includes a metadata parser (not shown for clarity) that analyzes input metadata and identifies keywords, events, user IDs, locations, demographics, file types, timestamps, etc., in the received data. Because the metadata analyzed by the introspective analyzer 175 is not homogeneous (e.g., many different sources in many different formats), one embodiment uses at least one metadata parser per cloud service, and possibly multiple metadata parsers. In another embodiment, the introspective analyzer 175 uses monitor 184 to inspect the cloud service and assemble content metadata. In one use case, the identification of sensitive documents is based on pre-inspection of the documents. Users can manually tag documents as sensitive, and this manual tagging updates the document metadata in the cloud service. The document metadata can then be retrieved from the cloud service using a publicly available API and used as an indicator of sensitivity.
[0062] Continuing the explanation in Figure 1A, System 100 can include any number of cloud-based services 108, namely point-to-point streaming services, hosting services, cloud applications, cloud stores, cloud collaboration, and messaging platforms, as well as a cloud customer relationship management (CRM) platform. These services can include peer-to-peer file sharing (P2P) via portal traffic protocols such as BitTorrent (BT), User Data Protocol (UDP) streaming, and File Transfer Protocol (FTP), instant messaging via Internet Protocol (IP), and voice, video, and messaging multimedia communication sessions such as mobile phone calls via LTE (VoLTE) via Session Initiation Protocol (SIP) and Skype. These services can handle internet traffic, cloud application data, and General Purpose Routing Encapsulation (GRE) data. Network services or applications can be web-based (e.g., accessed via a Uniform Resource Locator (URL)) or native, such as a synchronization client. Examples include the provision of Software as a Service (SaaS), Platform as a Service (PaaS), and Infrastructure as a Service (IaaS), as well as internal enterprise applications exposed via URLs. Common examples of cloud-based services today include Salesforce.com®, Box®, Dropbox®, Google Apps®, Amazon AWS®, Microsoft Office 365®, Workday®, Oracle on Demand®, Taleo®, Yammer®, Jive®, and Concur®.
[0063] In the interconnection of the elements of system 100, network 145 connects computers 112a-n, tablets 122a-n, mobile phones 132a-n, smartwatches 142a-n, cloud-based hosting service 118, webmail service 128, video, messaging and voice call service 138, streaming service 148, file transfer service 158, cloud-based storage service 168, and N-CASB 155 into a communication state. The communication path can be point-to-point via public and / or private networks. Communication takes place over various networks such as private networks, VPNs, MPLS lines, or the internet, and can use appropriate application programming interfaces (APIs) and data exchange formats such as REST, JSON, XML, SOAP, and JMS. All communications can be encrypted. This communication generally takes place over networks such as LANs (Local Area Networks), WANs (Wide Area Networks), telephone networks (public switched telephone networks), Session Initiation Protocol (SIP), wireless networks, point-to-point networks, star networks, token ring networks, hub networks, and the internet, including mobile internet via protocols such as EDGE, 3G, 4G LTE, Wi-Fi, and WiMAX. Furthermore, communications can be protected using various authorization and authentication technologies such as usernames / passwords, open authentication (OAuth), Kerberos, SecureID, and digital certificates.
[0064] Furthermore, the description of the system architecture in Figure 1A continues. N-CASB155 includes monitor 184 and storage 186, which may include one or more computers and computer systems coupled to communicate with each other. They may also be one or more virtual computing and / or storage resources. For example, monitor 184 may be one or more Amazon EC2 instances, and storage 186 may be Amazon S3® storage. Rather than implementing N-CASB 155 directly on a physical computer or a conventional virtual machine, Other compute-as-a-service platforms such as Salesforce's Rackspace, Heroku, or Force.com can be used. Furthermore, one or more engines can be used, and one or more points of presence (POPs) can be established to implement security functions. The engine or system components in Figure 1A are implemented by software running on various types of computing devices. Examples of devices include workstations, servers, computing clusters, blade servers, server farms, or other data processing systems and computing devices. The engines can be communicatively coupled to the database via different network connections. For example, the extraction engine 171 can be coupled via network 145 (e.g., the internet), the classification engine 172 can be coupled via a direct network link, and the security engine 173 can be coupled via yet another different network connection. In the disclosed technology, the POPs of the data plane 180 are hosted on the client's premises or located within a virtual private network controlled by the client.
[0065] N-CASB155 provides various functions via a management plane 174 and a data plane 180. According to one embodiment, the data plane 180 includes an extraction engine 171, a classification engine 172, and a security engine 173. It can also provide other functions such as a control plane. Collectively, these functions provide a secure interface between the cloud service 108 and the organizational network 102. While the term "network security system" is used to describe N-CASB155, more generally, this system provides not only security but also application visibility and control functions. In one example, 35,000 cloud applications reside in a library that crosses servers used by computers 112a-n, tablets 122a-n, mobile phones 132a-n, and smartwatches 142a-n within the organizational network 102.
[0066] According to one embodiment, computers 112a-n, tablets 122a-n, mobile phones 132a-n, and smartwatches 142a-n within the organizational network 102 have a web browser with a secure web delivery interface provided by N-CASB 155 for defining and managing content policies 187. This includes the following. Because N-CASB155 is a multi-tenant system, users of the management client can, depending on several embodiments, only modify content policies associated with their organization. In some embodiments, an API can be provided for programmatically defining and updating policies. In such embodiments, the management client can include one or more servers, such as a corporate identity directory like Microsoft Active Directory, to push updates and / or respond to pull requests for updates to content policies. Both systems can coexist. For example, a corporate identity directory can be used to automate the identification of users within an organization while a web interface can be used to tailor policies to needs. Management clients are assigned roles, and access to N-CASB155 data is controlled based on the role, e.g., read-only versus read-write.
[0067] In addition to periodically generating user-specific and file-specific data and storing it in metadata store 178, the Active Analyzer and Introspective Analyzer (not shown) also enforce security policies on cloud traffic. For further information regarding the functionality of the Active Analyzer and Introspective Analyzer, refer to, for example, the following jointly owned documents: U.S. Patent No. 9,398,102 (Agent No. NSKO1000-2); U.S. Patent No. 9,270,765 (Agent No. NSKO1000-3); U.S. Patent No. 9,928,377 (Agent No. NSKO1001-2); and U.S. Patent Application No. 15 / 368,246 (Agent No. NSKO1003-3); Chen, Itar, Narayanaswamy, and Malmskog, "Cloud Security for Dummies, Netscope Special Edition," John Wiley & Sons, 2015; "Netscope "Introspection" Netskope, Inc. Published by Netskope, Inc.; "Data Loss Prevention and Monitoring in the Cloud"; "Cloud Data Loss Prevention Reference Architecture"; "Five Steps to Cloud Confidence"; "Netskope Active Platform"; "Netskope Advantage: Three Essential Requirements for Cloud Access Security Brokers"; "15 Key CASB Use Cases"; "Netskope Active Cloud DLP"; "Remediating the Collision Course of Cloud Data Breaches"; and "Netskope Cloud Confidence Index (trademark), published by Netskope, Inc. The above materials are incorporated herein by reference as if they were fully contained herein.
[0068] In the case of system 100, a control plane can be used together with, or instead of, the management plane 174 and the data plane 180. The specific division of functions between these groups is an option in the embodiment. Similarly, functionality can be highly distributed across several points of presence (POPs) to improve locality, performance, and / or security. In one embodiment, the data plane resides on a premises or virtual private network, and the management plane of the network security system is located on a cloud service or enterprise network, as described herein. In another secure network embodiment, POPs can be distributed in different ways.
[0069] In this specification, System 100 is described with reference to specific blocks, but it should be understood that these blocks are defined for illustrative purposes only and are not intended to require a specific physical arrangement of components. Furthermore, these blocks do not need to correspond to physically separate parts. As long as physically separate parts are used, the connections between components can be wired and / or wireless, as desired. Different elements or components can be combined into a single software module, and multiple software modules can run on the same hardware.
[0070] Furthermore, this technology can be implemented using two or more separate computer implementation systems that communicate with each other in cooperation. This technology can be implemented in numerous ways, including processes, methods, apparatus, systems, devices, computer-readable media such as computer-readable storage media that store computer-readable instructions or computer program code, or computer-usable media having computer-readable program code embodied therein, including computer program products. The disclosed technologies can be implemented in the context of any computer implementation system, including database systems, or relational database implementations such as Oracle®-compatible database implementations, IBM DB2 Enterprise Server®-compatible relational database implementations, MySQL® or PostgreSQL®-compatible relational database implementations, or Microsoft SQL Server®-compatible relational database implementations, or NoSQL non-relational database implementations such as Vampire®-compatible non-relational database implementations, Apache Cassandra®-compatible non-relational database implementations, BigTable®-compatible non-relational database implementations, or HBase® or Dynamo®-compatible non-relational database implementations. Furthermore, the disclosed technologies can be implemented using various programming models such as MapReduce®, bulk synchronization programs, and MPI primitives, or various scalable batch and stream management systems such as Amazon Web Services (AWS)® including Amazon Elasticsearch Service® and Amazon Kinesis®, Apache Storm®, Apache Spark®, Apache Kafka®, Apache Flink®, Truviso®, IBM Info-Sphere®, Borealis®, and Yahoo!S4®.
[0071] Early deep learning models can perform well on the datasets used for training. However, their performance is unpredictable for images that are not visible. There is a continuing need to increase dataset coverage of real-world scenarios.
[0072] Figure 1B shows the detection configuration of image-derived sensitive data in the system 100, previously described in relation to Figure 1A, which includes an organizational network 102, a data center 152, and a cloud-based service 108. Each individual organizational network 102 has a user interface 103 for interacting with data loss prevention functions and a deep learning stack trainer 162. The dedicated DL stack trainer can be configured to generate updated DL stacks for each organization under the organization's control. The deep learning stack trainer 162 enables customer organizations to perform update training for their image and screenshot classifiers without the organization transferring sensitive data in the images to a DLP provider that has performed pre-training of the master DL stack. This protects PII data and other sensitive data from being accessible by the data loss prevention provider, thus reducing the requirements for protecting stored sensitive data stored in the DLP center. DL stack training will be discussed further later.
[0073] Continuing the explanation of Figure 1B, the data center 152 includes a Netscope Cloud Access Security Broker (N-CASB) 155, which includes an image-derived sensitive data detection 156, an image generation robot 167, and a deep learning stack 157 with inference and backpropagation 166. The deep learning (DL) stack parameters 183 and features and labels 185 are stored in the storage 186 described in detail earlier. The deep learning stack 157 utilizes the stored features and labels 185, which are generated as output from the first set of layers of the stack and retained with their respective ground truth labels for progressive online deep learning, thereby eliminating the need to retain images of private image-derived identification documents. When a new image-derived identification document is received, the new document can be classified by the trained DL stack described later.
[0074] The image generation robot 167 generates examples of other image documents for use in training the deep learning stack 157, in addition to actual passport images and US driver's license images. For example, the image generation robot 167 crawls sample US driver's license images via a web-based search engine, inspects the images, and filters out low-fidelity images.
[0075] The image generation robot 167 also collects examples of screenshot and non-screenshot images, creates labeled ground truth data for the image examples, and applies re-rendering to at least some of the collected screenshot examples representing various variations of screenshots that may contain sensitive information, leveraging tools available for web UI automation to create synthetic data for training the deep learning stack 157. One example of such a tool is Selenium, an open-source tool that can open web browsers, access websites, open documents, and simulate clicks on pages. For example, this tool can simulate plain desktop Starting from the top, one or more web browsers of various sizes can be opened in various locations on the desktop to access live websites or open a given local document. These actions can then be repeated using randomized parameters such as the number of browser windows, the size and location of the browser windows, and the relative positioning of the browser windows. Next, the image generation robot 167 takes a screenshot of the desktop and re-renders the screenshot, including augmenting the generated sample images as training data to be fed to the DL stack 157. For example, this process can add noise to the image to increase the robustness of the DL stack 157. The augmentation applied to our training data includes cropping parts of the image and adjusting the hue, contrast, and saturation. Inversion or rotation has not been added to image augmentation to detect screenshot images that people use to secretly extract data. In different embodiments, inversion and rotation can be added to examples of other image documents.
[0076] Figure 2 shows a block diagram of a deep learning (DL) stack 157 implemented using a convolutional neural network (CNN) architecture model for image classification, configurable for use in a system to detect identified documents within images and for detecting screenshot images. The image of the CNN architecture model was downloaded on April 28, 2020 from https: / / towardsdatascience.com / covolutional-neural-network-cb0883dd6529. The input to the initial CNN layers is the image data itself, represented as a three-dimensional matrix with image dimensions and three color channels, namely red, green, and blue. The input image can be 224×224×3, as shown in Figure 2. In another embodiment, the input image can be 200×200×3. In the example embodiment where the results are shown later, the size of the image used is 160×160×3, and there are 88 layers in total.
[0077] Continuing the description of the DL stack 157, the feature extraction layer consists of a convolutional layer 245 and a pooling layer 255. The disclosed system stores the feature and label outputs of the feature extraction layer 185 as numerical values processed through many different iterations of the convolution operation, storing irreversible features instead of the raw image. The extracted features cannot be reverse-transformed back into the original image pixel data; that is, the stored features are irreversible features. By storing these extracted features instead of input image data, the DL stack does not store pixels of the original image that could carry sensitive and personal information such as personally identifiable information (PII), protected health information (PHI), and intellectual property (IP).
[0078] The DL stack 157 includes a first set of layers closer to the input layer and a second set of layers further away from the input layer. The first set of layers is pre-trained to perform image recognition before the second set of layers of the DL stack is fitted with labeled ground truth data for examples of image-derived identified documents and other image documents. The disclosed DL stack 157 freezes the first 50 layers as the first set of layers. The DL stack 157 is trained by forward inference and backpropagation 166 using labeled ground truth data for examples of image-derived identified documents and other image documents. For private image-derived identified documents and screenshot images, the CNN architecture model captures features generated as outputs from the first set of layers and retains the captured features along with their respective ground truth labels, thereby eliminating the need to retain images of the private image-derived identified documents. The 265 fully connected layers and 275 SoftMax layers comprise a second set of layers further away from the input layers of the CNN being trained. Together with the first set of layers, the model is used to detect identified documents in images and to detect screenshot images.
[0079] The DL stack 157 is trained using forward inference and backpropagation 166, utilizing labeled ground truth data for examples of image-derived identified documents and other image documents. The first set of layers is pre-trained to perform image recognition before the second set of layers of the DL stack are fed the labeled ground truth data for examples of image-derived identified documents and other image documents. The output of the image classifier can be used to train the second set of layers, and in one example, only images classified as the same type by both the OCR and the image classifier are fed to the deep learning stack as labeled images.
[0080] The disclosed technology stores the parameters of a DL stack 183 trained for inference from production images, and uses the production DL stack with the stored parameters to classify the production images by inference, in one use case, as sensitive image-derived identified documents, and in another use case, as containing screenshot images.
[0081] In one use case, the objective was to develop an image classification deep learning model for detecting passport images. Initial training data for building a deep learning-based binary image classifier for classifying passports was generated using approximately 550 passports from 55 countries as labeled ground truth data for detecting image-derived identification documents. Since the objective was to detect passports with a high detection rate, detecting other identity document types as passports was unacceptable. A negative dataset was used, consisting of images of other identity document types, including driver's licenses, ID cards, student identifiers, military ID cards, etc., as well as non-identity document images. These other identity document images were used in the negative dataset to satisfy the goal of minimizing the detection rate of other identity document types.
[0082] In the second use case, the objective was to develop an image classifier for detecting passport images and U.S. driver's license images. Training data for building a deep learning-based binary image classifier for classifying passports was generated using 550 passport images and 248 U.S. driver's license images. In addition to actual passport and U.S. driver's license images, sample U.S. driver's license images obtained by crawling the internet were included after inspection and filtering of low-fidelity images.
[0083] Cross-validation techniques were used to evaluate DL stack models by training several models on a subset of available input data and evaluating them on complementary subsets of that data. In k-fold cross-validation, the input data is divided into k subsets of data, also known as folds. 10-fold cross-validation was applied to check the performance of the resulting image classifiers. A cutoff value of 0.3 was chosen for US driver's licenses and 0.8 for passports, and the model accuracy and recall were checked.
[0084] Figure 3 shows the accuracy and recall results for a trained passport and driver's license classifier, graphed with an accuracy of 345 for driver's licenses and a recall of 355 for driver's licenses, an accuracy of 365 for passports and a recall of 375 for passport images, and an accuracy of 385 for non-identity documents (non-driver's licenses or passports), also called negative results, and a recall of 395 for negative results. As shown in the graph, recall decreases as accuracy increases. The designers used 10-fold cross-validation to check the performance of the passport image classifier. The false positive rate (FPR) was calculated for non-identity document images in the test, and the false negative rate (FNR) was calculated for passport and driver's license images in the test. The averaged FPR and FNR from the 10-fold cross-validation results are listed below. • Passport FPR (non-identity document image classified as passport): 0.7% • FPR (Frequently Recognized as a U.S. Driver's License) for U.S. Driver's Licenses: 0.3% • Passport FNR (Passport image not classified as a passport): 6% • FNR (Failure to classify US driver's license image as not being a driver's license): 6%
[0085] Figure 4 graphs the execution time results for classifying images using model inference on Google Cloud Platform (GCP) (n1-highcpu-64: 64 vCPU, 57.6GB memory) with over 1000 images of various file sizes, showing the distribution of execution times as images. For images with a file size of 2MB or less, the graph shows the execution time distribution as a function of file size. Execution time is counted from the time the image was read ("opencv") to the time the classifier finished its prediction on the image. The average execution time was 45ms, and the standard deviation was 56ms.
[0086] Figure 5 shows benchmark results comparing a commercially available classifier for classifying sensitive images of US driver's licenses with the significantly improved performance of the disclosed deep learning stack. The number of images to be classified is 334. Using a commercially available classifier that uses regular expression (Regex) OCR and pattern matching, 238 out of 334 images are detected, representing a 71.2% detection rate of 566. The majority of the sensitive images are detected, and the system performs only "reasonably" well. For some images, the classifier is unable to extract blurred or rotated text. In contrast, the disclosed technique utilizing the deep learning stack detects 329 out of 334 images, representing a 98.5% detection rate of 576 images, including identified documents derived from the sensitive images.
[0087] Figure 6 shows an example workflow 600 for detecting identified documents within an image, called image-derived identified documents, and preventing the loss of these documents. Step 605 selects a pre-trained network, such as a CNN, as previously discussed in relation to Figure 2. The DL stack includes at least a first set of layers closer to the input layer and a second set of layers further away from the input layer, with the first set of layers being pre-trained to perform image recognition. In the example described, a MobileNet CNN was selected for image detection. Different CNNs or different ML classifiers can be selected. Step 615 covers the collection of images containing sensitive information balanced against negative images, as described in the two use cases. In step 625, both the final layer of the pre-trained network and the classifier of the CNN model are retrained, and the CNN model is validated and tested. The DL stack is trained by forward inference and backpropagation using labeled ground truth data for examples of image-derived identified documents and other image documents collected in step 615, and the labeled ground truth data for examples of image-derived identified documents and other image documents is applied to a second set of layers of the DL stack. In step 635, the extracted features of the current CNN are saved for all images in the current dataset. In step 645, a new CNN model is deployed, which is a production DL stack with the stored parameters of the trained DL stack, for inference from production images. In step 655, a batch of new labels is collected from production OCR and negative images that do not contain image-derived information are added. In step 665, new images are added to the training dataset for the CNN model to form new inputs. In step 675, the classifier of the CNN model is retrained, the model is validated and tested, and then the production DL stack is used to classify by inference at least one production image as containing a sensitive image-derived identified document.
[0088] For use cases where screenshot images are detected and sensitive screenshot-derived data is prevented from being lost, the workflow is similar to workflow 600. To detect screenshot image scenarios, the image generation robot 167 is a screenshot robot that collects examples of screenshot and non-screenshot images and creates labeled ground truth data for the examples without requiring OCR for use when training the deep learning stack 157. The screenshot robot applies re-rendering to at least some of the collected screenshot examples to represent changes in screenshots that may contain sensitive information. The training data for training the DL stack using forward inference and backpropagation with the labeled ground truth data utilizes examples of screenshot and non-screenshot images. In one example, a full screenshot image contains a single application window, and the window size covers more than 50% of the full screen. In another example, a full screenshot image shows multiple application windows, and in yet another example, an application screenshot image shows a single application window.
[0089] Figure 7 shows an example screenshot image containing a customer inventory list with costs listed. By detecting screenshot images, confidential company data can be exfiltrated. It can prevent rations from being delivered.
[0090] The cross-validation of results obtained using the disclosed method for detecting screenshot images focuses on checking how well the DL stack model generalizes. The collection of screenshot and non-screenshot images was separated into training and test sets for screenshots with a MAC background. Images with a Windows background and images with a Linux® background were used exclusively for testing. Furthermore, application windows were divided into training and test sets based on their categories. The performance of five separate cross-validation cases is then described. The training data merge was a composite set of full screenshots mixed from MAC background training and App window training.
[0091] In cross-validation example 1, the test data was a set of composite full-screen screenshots mixed from tests with a Mac background and tests in an app window. The screenshot detection accuracy was measured at 93%. In cross-validation example 2, the test data was a set of composite full-screen screenshots mixed from tests with a Windows background and tests in an app window. The screenshot detection accuracy was measured at 92%. In cross-validation example 3, the test data was a set of composite full-screen screenshots mixed from tests with a Linux background and tests in an app window. The screenshot detection accuracy was measured at 86%. In cross-validation example 4, the test data was a set of composite full-screen screenshots mixed from tests with a Mac background and tests in multiple app windows. The screenshot detection accuracy using these training and test data sets was measured at 97%. In cross-validation example 5, the test data tested a different app than the training app window, and the accuracy was measured at 84%.
[0092] The performance of the deep learning stack model was tested for invisible background and app window types, and then the classifier was trained using 4,528 screenshots and 1,964 non-screenshot images with a synthetic full screenshot containing all background images and all app windows. The classifier was tested using 45,179 images. In the failure rate of detection (FNR) test using 45,179 screenshots, 90 images were classified as failures (FN) with an FNR of 0.2% at a threshold of 0.7. In the false positive rate (FPR) test of 1,336 non-screenshot images, 4 images were classified as false positives (FP) with an FPR of 0.374% at a threshold of 0.7. The 4 images in the test set were incorrectly classified as screenshots if they were non-screenshot images. Many layers in the disclosed deep learning stack model work to capture features to determine “screenshots” which include the following prominent features: (1) Screenshots tend to contain one or more main windows containing sensitive information. Such information may include personal information, code, text, pictures, etc. (2) Screenshots tend to include header / footer bars such as menus or application bars. (3) Screenshots tend to have a contrasting or uniform background compared to the content within the application window. For four FP images, the main reasons why the images were classified as screenshots are as follows: Figure 8A shows a map of Idaho that was misclassified as a screenshot image due to its legend window and dotted lines at the top and bottom. Figure 8B shows a driver's license image that was misclassified as a screenshot image because the entire image is a window containing PII on a black background and the UNITED STATES bar can be perceived as a header bar. Figure 8C shows a passport image as a main window containing PII, and the shaded area at the bottom center may mislead the classifier into thinking it is an application bar. Figure 8D shows text within a main window containing text information and a uniform background that was misclassified as a screenshot image.
[0093] In some use cases, separate organizations requiring DLP services can utilize a locally operating, dedicated DL stack trainer 162 configured to combine irreversible features from instances of organization-sensitive data in images with ground truth labels for those instances. The dedicated DL stack trainer transfers the irreversible features and ground truth labels to a deep learning stack that receives organization-sensitive training examples containing the irreversible features and ground truth labels from the dedicated DL stack trainer 162. The organization-sensitive training examples are used to further train a second set of layers in the trained master DL stack. The updated parameters of the second set of layers for inference from production images are stored and can be distributed to multiple separate organizations without compromising data security, as sensitive data is not accessible in the irreversible features.
[0094] Training the deep learning stack 157 can be started from scratch using training examples in a different order. Alternatively, in another example, training can be performed by further training a second set of layers of the trained master DL stack using an additional batch of labeled image examples.
[0095] In the added batch scenario, when samples are received back from the customer organization, the dedicated DL stack trainer may be configured to transfer updated coefficients from the second set of layers. The deep learning stack 157 can receive the respective updated coefficients from each of the second set of layers from multiple dedicated DL stack trainers and combine the updated coefficients from each of the second set of layers to train the second set of layers of the trained master DL stack. The deep learning stack 157 can then store the updated parameters of the second set of layers of the trained master DL stack for inference from production images and distribute the updated parameters of the second set of layers to separate customer organizations.
[0096] The dedicated DL stack trainer 162 can, in one example, handle training for detecting image-derived identification documents, and in another example, it can perform training for detecting screenshot images.
[0097] Next, we will describe an example of a computer system that can be used to detect identified documents in images, detect screenshots, and prevent the loss of sensitive image-derived documents in the cloud. [Receiving System]
[0098] Figure 9 is a simplified block diagram of a computer system 900 that can be used to detect identifiable documents within images, known as image-derived identifiable documents, and to prevent the loss of image-derived identifiable documents in the cloud. The computer system 900 can also be used to detect screenshot images and prevent the loss of sensitive screenshot-derived data. Furthermore, the computer system 900 can be used to detect organizational sensitive data within images and prevent the loss of image-derived organizational sensitive documents without requiring the transfer of potentially sensitive images to a centralized DLP service, by customizing the deep learning stack. The computer system 900 includes at least one central processing unit (CPU) 972 that communicates with several peripheral devices via a bus subsystem 955, and a Netscope Cloud Access Security Broker (N-CASB) 155 that provides the network security services described herein. These peripheral devices may include, for example, a storage subsystem 910 including a memory device and file storage subsystem 936, a user interface input device 938, a user interface output device 976, and a network interface subsystem 974. The input and output devices enable user interaction with the computer system 900. The network interface subsystem 974 provides an interface to an external network, including an interface to a corresponding interface device in another computer system.
[0099] In one embodiment, the Netscope Cloud Access Security Broker (N-CASB) 155 shown in Figures 1A and 1B is linked to the storage subsystem 910 and the user interface input device 938 in a communicative manner.
[0100] The user interface input device 938 may include pointing devices such as keyboards, mice, trackballs, touchpads, or graphics tablets, scanners, touchscreens integrated into displays, audio input devices such as voice recognition systems and microphones, and other types of input devices. In general, the use of the term “input device” is intended to include all possible types of devices and methods for inputting information into the computer system 900.
[0101] The user interface output device 976 may include a display subsystem, a printer, a fax machine, or a non-visual display such as an audio output device. The display subsystem may include a flat panel device such as an LED display, a cathode ray tube (CRT), or a liquid crystal display (LCD), a projection device, or any other mechanism for generating visible images. The display subsystem may also provide a non-visual display such as an audio output device. In general, the use of the term “output device” is intended to include all possible types of devices and methods for outputting information from the computer system 900 to a user or to another machine or computer system.
[0102] The storage subsystem 910 stores programming and data structures that provide some or all of the functionality of the modules and methods described herein. The subsystem 978 may be a graphics processing unit (GPU) or a programmable gate array (FPGA).
[0103] The memory subsystem 922 used in the storage subsystem 910 may include several memories, including main random access memory (RAM) 932 for storing instructions and data during program execution, and read-only memory (ROM) 934 for storing fixed instructions. The file storage subsystem 936 can provide persistent storage for program and data files and may include a hard disk drive, a floppy disk drive, a CD-ROM drive, an optical drive, or a removable media cartridge, along with associated removable media. Modules that perform the functions of a particular embodiment may be stored by the file storage subsystem 936 within the storage subsystem 910, or in other machines accessible by the processor.
[0104] The bus subsystem 955 provides a mechanism for various components and subsystems of the computer system 900 to communicate with each other as intended. Although the bus subsystem 955 is schematically shown as a single bus, other possible embodiments of the bus subsystem may utilize multiple buses.
[0105] Computer System 900 is a personal computer, a portable computer. This can include a variety of types, such as workstations, computer terminals, network computers, televisions, servers, mainframes, a widely distributed set of loosely coupled computers, or any other data processing systems or user devices. Due to the constantly changing nature of computers and networks, the description of the computer system 900 shown in Figure 9 is intended only as a specific example for the purpose of illustrating preferred embodiments of the invention. Many other configurations of the computer system 900 may have more or fewer components than the computer system shown in Figure 9.
[0106] Figure 10 shows a workflow 1000 for a system of one or more computers that can be configured to detect screenshot images and prevent the loss of screenshot data. Computers perform specific operations or actions by installing software, firmware, hardware, or a combination thereof on the system that causes the system to perform actions while in operation. One or more computer programs can be configured to perform specific operations or actions by containing instructions that cause the data processing device to perform actions when executed by the data processing device. In some embodiments, multiple actions can be combined. For convenience, this flowchart is illustrated with reference to a system that includes a Netscope Cloud Access Security Broker (N-CASB) and load balancing in a dynamic service chain while applying security services in the cloud. One common embodiment includes a method for detecting screenshot images and preventing the loss of sensitive screenshot-derived data, which includes collecting examples of screenshot and non-screenshot images and creating labeled ground truth data for the examples 1010. The method for detecting screenshot images also includes applying re-rendering 1020 to represent various variations of screenshots that may contain sensitive information. The method for detecting screenshot images also includes training a deep learning (DL) stack by forward inference and backpropagation using labeled ground truth data for instances of screenshot and non-screenshot images.1030 The DL stack pre-trains a first set of layers closer to the input layer to perform image recognition before a second set of layers further away from the input layer is fitted with labeled ground truth data for screenshot and non-screenshot images. The method for detecting screenshot images also includes storing the parameters of the trained DL stack for inference from production images.1040The method for detecting screenshot images also includes using a production DL stack along with stored parameters to classify, by inference, at least one production image as containing a screenshot image.1050 Other embodiments of this aspect include corresponding computer systems, devices, and computer programs recorded on one or more computer storage devices, each configured to perform the operation of the method.
[0107] Figure 11 shows a workflow 1100 for a system of one or more computers that can be configured to perform a specific operation or action by installing software, firmware, hardware, or a combination thereof on the system that causes the system to perform an action during operation. One or more computer programs can be configured to perform a specific operation or action by containing instructions that cause the data processing device to perform an action when executed by the data processing device. In some embodiments, multiple actions can be combined. For convenience, this flowchart is illustrated with reference to a system that includes a Netscope Cloud Access Security Broker (N-CASB) and load balancing in a dynamic service chain while applying security services in the cloud. One common embodiment includes a method for customizing a deep learning (DL) stack to detect organizational sensitive data in images, called image-derived organizational sensitive documents, and to prevent the loss of image-derived organizational sensitive documents, which includes pre-training a master DL stack by forward inference and backpropagation using labeled ground truth data for examples of image-derived confidential documents and other image documents (1110). The DL stack comprises at least a first set of layers closer to the input layer and a second set of layers further away from the input layer, and further comprises the first set of layers being pre-trained to perform image recognition before the second set of layers of the DL stack is fitted with labeled ground truth data for examples of image-derived sensitive documents and other image documents (1120). The method also comprises storing the parameters of the trained master DL stack for inference from production images (1130). The method also comprises distributing the trained master DL stack with the stored parameters to multiple organizations (1140).The method further includes enabling an organization to perform update training on a trained master DL stack using at least examples of organization-sensitive data in images and to save the parameters of the updated DL stack (1150). The organization uses each updated DL stack to classify, by inference, at least one production image as containing organization-sensitive documents (1160). The method may also optionally include, under the control of the organization, providing at least some organizations with a dedicated DL stack trainer and enabling the organization to perform update training using a dedicated DL stack trainer that can be configured to generate each updated DL stack without transferring examples of organization-sensitive data in images to a provider that performed pre-training of the master DL stack (1170). Other embodiments of this aspect include corresponding computer systems, devices, and computer programs recorded on one or more computer storage devices, each configured to perform the operations of the method. [Specific Embodiments]
[0108] Several specific embodiments and features for detecting identified documents within images and preventing the loss of image-derived identified documents are described in the following discussion.
[0109] In one disclosed embodiment, a method for detecting and preventing the loss of image-derived identified documents within an image, referred to as image-derived identified documents, includes training a deep learning (DL) stack by forward inference and backpropagation using labeled ground truth data for instances of image-derived identified documents and other image documents. The disclosed DL stack includes at least a first set of layers closer to the input layer and a second set of layers further away from the input layer, and further includes a first set of layers pre-trained to perform image recognition before the second set of layers of the DL stack is fitted with labeled ground truth data for instances of image-derived identified documents and other image documents. The disclosed method also includes storing parameters of the DL stack trained for inference from production images, and using the production DL stack with the stored parameters to classify by inference at least one production image as containing a sensitive image-derived identified document.
[0110] The methods described in this section and other sections of the disclosed technology may include one or more of the following features and / or features described in relation to additional methods disclosed. For brevity, the combinations of features disclosed in this application are not listed individually and are not repeated for each basic set of features. Readers can easily combine the features identified in this method with the sets of basic features identified as embodiments. They will understand whether they can do it.
[0111] Some disclosed embodiments of the method optionally include capturing features generated as output from a first set of layers for a private image-derived identification document, retaining the captured features along with their respective ground truth labels, thereby eliminating the need to retain the images of the private image-derived identification document.
[0112] Some embodiments of the disclosed method include limiting training by backpropagation using labeled ground truth data for examples of image-derived identification documents and other image documents to training parameters in a second set of layers.
[0113] In one disclosed embodiment of the present invention, optical character recognition (OCR) analysis of an image is applied to label the image as an identified or unidentified document. After the OCR analysis, a reliable classification can be selected for use in a training set. OCR and regular expression matching serve as an automated method for generating labeled data from a customer's production images. In one example, for a US passport, OCR first extracts text from the passport page. Then, a regular expression may match "PASSPORT", "UNITED STATES", "Department of State", "USA", "Authority", and other words on the page. In a second example, for a California driver's license, OCR first extracts text from the front of the driver's license. Then, a regular expression may match "California", "USA", "DRIVER LICENSE", "CLASS", "SEX", "HAIR", "EYES", and front It may match other words on the page. As a third example, in the case of a Canadian passport, OCR first extracts the text from the passport page. The regular expression can then match "PASSPORT", "PASSEPORT", "CANADA", and other words on that page.
[0114] In some disclosed embodiments of the present invention, when training the DL stack by backpropagation, the perspective of a first set of image-derived identified documents is distorted to generate a second set of image-derived identified documents, and the first and second sets are combined with labeled ground truth data.
[0115] In other embodiments of the disclosed method, when training the DL stack by backpropagation, a first set of image-derived identified documents is distorted by rotation to generate a third set of image-derived identified documents, and the first and third sets are combined with labeled ground truth data.
[0116] In one embodiment disclosed of the present invention, when training the DL stack by backpropagation, a first set of image-derived identified documents is distorted by noise to generate a fourth set of image-derived identified documents, and the first and fourth sets are combined with labeled ground truth data.
[0117] In some embodiments disclosed in the present invention, when training the DL stack by backpropagation, the focus of the first set of image-derived identified documents is distorted to generate a fifth set of image-derived identified documents, and the first and fifth sets are combined with labeled ground truth data.
[0118] In some embodiments, the disclosed method includes storing irreversible DL features of current training ground images rather than original ground images to avoid storing sensitive personal information, periodically adding irreversible DL features of new ground images to augment the training set, and augmenting the training dataset to make it more accurate. This includes regularly retraining the model. Irreversible DL features cannot be converted into images containing recognizable sensitive data.
[0119] Several specific embodiments and features for detecting screenshot images and preventing the loss of sensitive screenshot-derived data are described in the following discussion.
[0120] In one embodiment disclosed, a method for detecting screenshot images and preventing the loss of sensitive screenshot-derived data includes collecting examples of screenshot and non-screenshot images, and creating labeled ground truth data for the examples. The method also includes applying re-rendering to at least some of the collected screenshot examples to represent various variations of screenshots that may contain sensitive information, and training a DL stack by forward inference and backpropagation using the labeled ground truth data for the screenshot and non-screenshot examples. The method further includes storing parameters of a DL stack trained for inference from production images, and classifying at least one production image by inference as containing a sensitive image-derived screenshot using the production DL stack having the stored parameters.
[0121] Some embodiments of the disclosed method further include applying a screenshot robot to collect examples of screenshot images and non-screenshot images.
[0122] In one embodiment of the disclosed method, the DL stack includes at least a first set of layers closer to the input layer and a second set of layers further away from the input layer, and the first set of layers is pre-trained to perform image recognition before the second set of layers of the DL stack is fitted with labeled ground truth data for instances of screenshot and non-screenshot images.
[0123] Some embodiments of the disclosed method include applying an automatic re-rendering of at least a portion of the collected original screenshot image by cropping a portion of the image or by adjusting the hue, contrast, and saturation to represent changes in the screenshot. In some cases, the various changes in the screenshot include at least one of the window size, window position, number of open windows, and menu bar position.
[0124] In one embodiment of the disclosed method, when training a DL stack by backpropagation, a first set of screenshot images is surrounded by the boundaries of various photographic images of screenshots derived from two or more sensitive images to generate a third set of screenshot images, and the first and third sets are combined with labeled ground truth data. In another embodiment, when training a DL stack by backpropagation, a first set of screenshot images is surrounded by the boundaries of multiple overlaid program windows of screenshots derived from two or more sensitive images to generate a fourth set of screenshot images, and the first and fourth sets are combined with labeled ground truth data.
[0125] The following discussion describes several specific embodiments and features for detecting organizational confidential screenshot images and preventing the loss of image-derived organizational confidential screenshots.
[0126] In one disclosed embodiment, a deep learning stack is customized to detect organizational sensitive data within images, referred to as image-derived organizational sensitive documents. A method to prevent the loss of sensitive documents includes pre-training a master DL stack by forward inference and backpropagation using labeled ground truth data for instances of image-derived sensitive documents and other image documents. The DL stack includes at least a first set of layers closer to the input layer and a second set of layers further away from the input layer, and further includes pre-training the first set of layers to perform image recognition before the second set of layers of the DL stack is fitted with labeled ground truth data for instances of image-derived sensitive documents and other image documents. The disclosed method also includes storing the parameters of the master DL stack trained for inference from production images, distributing the trained master DL stack with the stored parameters to multiple organizations, and allowing organizations to perform update training on the master DL stack trained with at least instances of organization-sensitive data in images and store the parameters of the updated DL stack. Each organization uses its respective updated DL stack to classify, by inference, at least one production image as containing organization-sensitive documents.
[0127] Training of a deep learning stack can, in some cases, begin from scratch, and in other embodiments, training can further train a second set of layers of the trained master DL stack using an additional batch of labeled image examples utilized with previously determined coefficients. Some embodiments of the disclosed method further include providing at least a portion of an organization with a dedicated DL stack trainer under the organization's control, and enabling the organization to perform update training without transferring examples of organization-sensitive data in images to a provider that performed pre-training of the master DL stack. The dedicated DL stack trainer can be configured to generate each updated DL stack. In some cases, the dedicated DL stack trainer also includes a dedicated DL stack trainer configured to transfer irreversible features and ground truth labels from examples of organization-sensitive data in images, combining them with ground truth labels for said examples, and receiving organization-sensitive training examples containing irreversible features and ground truth labels from multiple dedicated DL stack trainers. In some embodiments, the disclosed method also includes further training a second set of layers of the trained master DL stack using organizational confidential training examples, storing updated parameters of the second set of layers for inference from production images, and distributing the updated parameters of the second set of layers to multiple organizations. Some embodiments further include performing update training to further train a second set of layers of the trained master DL stack. In other cases, the method includes further training a second set of layers of the trained master DL stack by performing training from the beginning using organizational confidential training examples in a different order. As one embodiment, the disclosed method further includes a dedicated DL stack trainer configured to transfer updated coefficients from a second set of layers, receiving the respective updated coefficients from each second set of layers from multiple dedicated DL stack trainers, and combining the updated coefficients from each second set of layers to train a second set of layers of the trained master DL stack.The disclosed method also includes storing updated parameters for a second set of layers of a trained master DL stack for inference from production images, and distributing the updated parameters for the second set of layers to multiple organizations.
[0128] Other embodiments of the disclosed technology described in this section may include a tangible, non-temporary, computer-readable storage medium containing program instructions loaded into memory, which, when executed on a processor, cause the processor to perform any of the methods described above. Yet another embodiment of the disclosed technology described in this section may include a system that includes memory and one or more processors operable to execute computer instructions stored in memory in order to perform any of the methods described above.
[0129] The foregoing description is provided to enable the use and implementation of the disclosed technology. Various modifications to the disclosed embodiments are evident, and the general principles expressed herein may be applied to other embodiments and applications without departing from the spirit and scope of the disclosed technology. Accordingly, the disclosed technology is not intended to be limited to the embodiments shown, but should be given the broadest scope consistent with the principles and features disclosed herein. The scope of the disclosed technology is defined by the appended claims. [Clause]
[0130] This document describes techniques for detecting identifiable documents within images and preventing the loss of image-derived identifiable documents.
[0131] The disclosed technology can be implemented as a system, method, device, or product. One or more features of an embodiment can be combined with the basic embodiment. Non-exclusive embodiments are taught to be combinable. One or more features of an embodiment can be combined with other embodiments. This disclosure periodically reminds the user of these options. The omission of repetitive descriptions of these options from some embodiments should not be construed as limiting the combinations taught in the preceding sections. These descriptions are incorporated by reference with consideration to each of the following embodiments.
[0132] One or more embodiments and provisions or elements of the disclosed technology may be implemented in the form of a computer product including a non-temporary computer-readable storage medium having computer-usable program code for performing the indicated method steps. Furthermore, one or more embodiments and provisions or elements of the disclosed technology may be implemented in the form of a device including memory and at least one processor coupled to the memory and operable to perform a typical method step. Furthermore, in another aspect, one or more embodiments and provisions or elements of the disclosed technology may be implemented in the form of means for performing one or more of the method steps described herein. The means may include (i) a hardware module, (ii) a software module running on one or more hardware processors, or (iii) a combination of a hardware module and a software module, any of (i) to (iii) implements the specific technology described herein, and the software module is stored in a computer-readable storage medium (or a number of such mediums).
[0133] The features described in this section can be combined as features. For brevity, feature combinations are not listed individually and are not repeated for each basic set of features. The reader will understand how the features identified in the features described in this section can be readily combined with sets of basic features identified as embodiments in other sections of this application. These features are not intended to be mutually exclusive, exhaustive, or restrictive, and the disclosed technology is not limited to these features, but rather encompasses all possible combinations, modifications, and variations within the scope of the claimed technology and its equivalents.
[0134] Other embodiments of the provisions described in this section may include a non-temporary computer-readable storage medium that stores instructions executable by a processor in order to perform any of the provisions described in this section. Yet another embodiment of the provisions described in this section may include a system that includes memory and one or more processors operable to execute instructions stored in that memory in order to perform any of the provisions described in this section.
[0135] The following clauses are disclosed: [Clause Set 1] 1. A method for detecting an identified document within an image, referred to as an image-derived identified document, and preventing the loss of said image-derived identified documents: Training a deep learning (DL) stack by forward inference and backpropagation using labeled ground truth data for the aforementioned image-derived identification documents and other image document examples; Here, the DL stack includes at least a first set of layers closer to the input layer and a second set of layers further away from the input layer, and further includes the first set of layers being pre-trained to perform image recognition before the second set of layers of the DL stack is fitted with the labeled ground truth data for the image-derived identification documents and other image documents; A method comprising: storing the parameters of the trained DL stack for inference from production images; and using the production DL stack having the stored parameters to classify by inference at least one production image as containing a sensitive image-derived identified document. 2. The method according to Clause 1, further comprising capturing features generated as outputs from the first set of layers for a private image-derived identification document, and retaining the captured features together with their respective ground truth labels, thereby eliminating the need to retain images of the private image-derived identification document. 3. The method according to any one of clauses 1 to 2, further comprising limiting backpropagation training using the labeled ground truth data for the aforementioned examples of the image-derived identification documents and other image documents to training of parameters in the second set of layers. 4. The method according to any one of the provisions of paragraphs 1 to 3, wherein the image is labeled as an identified document or a non-identified document by applying optical character recognition (OCR) analysis to the image. 5. The method according to any one of clauses 1 to 4, wherein when the DL stack is trained by backpropagation, the perspective of the first set of image-derived identification documents is distorted to generate a second set of image-derived identification documents, and the first set and the second set are combined with the labeled ground truth data. 6. The method according to any one of clauses 1 to 5, wherein when the DL stack is trained by backpropagation, the first set of image-derived identification documents is distorted by noise to generate a third set of image-derived identification documents, and the first set and the third set are combined with the labeled ground truth data. 7. The method according to any one of clauses 1 to 6, wherein when the DL stack is trained by backpropagation, the focus of the first set of image-derived identification documents is distorted to generate a fourth set of image-derived identification documents, and the first set and the fourth set are combined with the labeled ground truth data. 8. A tangible, non-temporary, computer-readable storage medium, which, when executed on a processor, includes program instructions loaded into memory, causing the processor to perform a method for detecting an identification document within an image, called an image-derived identification document, and for preventing the loss of the image-derived identification document: The method is Training a deep learning (DL) stack by forward inference and backpropagation using labeled ground truth data for the aforementioned image-derived identification documents and other image document examples; Here, the DL stack consists of a first set of layers closer to the input layer and layers further away from the input layer. The DL stack comprises at least a second set of layers, and further comprises the first set of layers being pre-trained to perform image recognition before the second set of layers of the DL stack is fitted with the labeled ground truth data of the image-derived identification documents and the other image documents; A tangible, non-temporary, computer-readable storage medium comprising storing the parameters of the trained DL stack for inference from production images, and using the production DL stack having the stored parameters to classify by inference at least one production image as containing a confidential image-derived identified document. 9. A tangible, non-temporary computer-readable storage medium as described in Clause 8, further comprising capturing features generated as output from the first set of layers for a private image-derived identification document, and retaining the captured features together with their respective ground truth labels, thereby eliminating the need to retain images of the private image-derived identification document. 10. A tangible, non-temporary, computer-readable storage medium as described in any one of clauses 8 to 9, further comprising limiting backpropagation training using the labeled ground truth data for the aforementioned examples of the image-derived identification documents and other image documents to training of parameters in the second set of layers. 11. A tangible, non-temporary computer-readable storage medium as described in any one of clauses 8 to 10, which applies optical character recognition (OCR) analysis of an image to label the image as an identified document or a non-identified document. 12. A tangible, non-temporary, computer-readable storage medium according to any one of clauses 8 to 11, wherein when the DL stack is trained by backpropagation, the perspective of the first set of image-derived identification documents is distorted to generate a second set of image-derived identification documents, and the first and second sets are combined with the labeled ground truth data. 13. A tangible, non-temporary, computer-readable storage medium according to any one of clauses 8 to 12, wherein when the DL stack is trained by backpropagation, the first set of image-derived identification documents is distorted by noise to generate a third set of image-derived identification documents, and the first set and the third set are combined with the labeled ground truth data. 14. A system for detecting an identifiable document within an image, referred to as an image-derived identifiable document, and for preventing the loss of such image-derived identifiable documents, the system comprising a processor, memory connected to the processor, and computer instructions loaded into the memory from a non-temporary computer-readable storage medium as described in Clause 8. 15. The system according to Clause 14, further comprising capturing features generated as outputs from the first set of layers for a private image-derived identification document, and retaining the captured features together with their respective ground truth labels, thereby eliminating the need to retain images of the private image-derived identification document. 16. The system according to any one of clauses 14-15, further comprising limiting backpropagation training using the labeled ground truth data for the aforementioned examples of image-derived identification documents and other image documents to training of parameters in the second set of layers. 17. A system described in any one of clauses 14 to 16, which applies optical character recognition (OCR) analysis of images to label them as identified or unidentified documents. 18. The system according to any one of clauses 14 to 17, wherein when the DL stack is trained by backpropagation, the perspective of the first set of image-derived identification documents is distorted to generate a second set of image-derived identification documents, and the first set and the second set are combined with the labeled ground truth data. 19. The system according to any one of clauses 14 to 18, wherein when the DL stack is trained by backpropagation, the first set of image-derived identification documents is distorted by noise to generate a third set of image-derived identification documents, and the first set and the third set are combined with the labeled ground truth data. 20. The system according to any one of clauses 14 to 19, wherein when the DL stack is trained by backpropagation, the focus of the first set of image-derived identification documents is distorted to generate a fourth set of image-derived identification documents, and the first set and the fourth set are combined with the labeled ground truth data. [Clause Set 2] 1. A method for detecting screenshot images and preventing the loss of sensitive data originating from screenshots: Collect examples of the aforementioned screenshot images and non-screenshot images, and create labeled ground truth data for the examples; To represent various variations of screenshots that may contain confidential information, re-rendering of at least some of the collected screenshot examples may be applied; Training a deep learning (DL) stack by forward inference and backpropagation using labeled ground truth data for the aforementioned screenshot and non-screenshot examples; A method comprising: storing the parameters of the trained DL stack for inference from production images; and using the production DL stack having the stored parameters to classify by inference that at least one production image contains a screenshot image. 2. The method according to Clause 1, further comprising applying a screenshot robot to collect the instances of the screenshot images and non-screenshot images. 3. The method according to any one of Clauses 1 to 2, wherein the DL stack comprises at least a first set of layers closer to the input layer and a second set of layers further away from the input layer, and further comprising the first set of layers being pre-trained to perform image recognition before the second set of layers of the DL stack is fitted with the labeled ground truth data for the instances of the screenshot images and non-screenshot images. 4. The method of any one of the clauses 1 to 3, further comprising applying an automatic re-rendering of at least a portion of the collected original screenshot image by cropping a portion of the image or by adjusting the hue, contrast, and saturation to represent the changes in the screenshot. 5. The various changes in the aforementioned screenshot include at least one of the following: window size, window position, number of open windows, and menu bar position, as described in any one of the methods in any one of the provisions of 1 to 4. 6. The method according to any one of clauses 1 to 5, wherein when the DL stack is trained by backpropagation, a first set of screenshot images is surrounded by the boundaries of various photographic images of multiple image-derived screenshots to generate a third set of screenshot images, and the first and third sets are combined with the labeled ground truth data. 7. The method according to any one of the clauses 1 to 6, wherein when the DL stack is trained by backpropagation, a first set of screenshot images is surrounded by the boundaries of multiple overlaid program windows of multiple image-derived screenshots to generate a fourth set of screenshot images, and the first and fourth sets are combined with the labeled ground truth data. 8. A tangible, non-temporary, computer-readable storage medium, which, when executed on a processor, includes program instructions loaded into memory, causing the processor to perform a method for detecting a screenshot image and preventing the loss of image-derived screenshots: The method is Collect examples of the aforementioned screenshot images and non-screenshot images, and create labeled ground truth data for the aforementioned examples; To represent various variations of screenshots that may contain confidential information, re-rendering of at least some of the collected screenshot examples may be applied; Training a deep learning (abbreviated as DL) stack by forward inference and backpropagation using labeled ground truth data for the aforementioned screenshot images and non-screenshot images; A tangible, non-temporary, computer-readable storage medium comprising storing the parameters of the trained DL stack for inference from production images, and using the production DL stack having the stored parameters to classify by inference that at least one production image contains an image-derived screenshot. 9. A tangible, non-temporary computer-readable storage medium as described in Clause 8, further comprising applying a screenshot robot to collect the aforementioned examples of screenshot images and non-screenshot images. 10. A tangible non-temporary computer-readable storage medium as described in any one of Clauses 8 to 9, wherein the DL stack comprises at least a first set of layers closer to the input layer and a second set of layers further away from the input layer, and further comprising the first set of layers being pre-trained to perform image recognition before the second set of layers of the DL stack is fitted with the labeled ground truth data for the examples of the screenshot images and the non-screenshot images. 11. A tangible, non-temporary computer-readable storage medium as described in any one of Clauses 8 to 10, further comprising applying automatic re-rendering of at least a portion of the collected original screenshot image by cropping a portion of the image or by adjusting the hue, contrast, and saturation to represent changes in the screenshot. 12. The various variations of the aforementioned screenshot include at least one of the following: window size, window position, number of open windows, and menu bar position, in a tangible, non-temporary computer-readable storage medium as described in any one of Clauses 8 to 11. 13. A tangible, non-temporary, computer-readable storage medium as described in any one of Clauses 8 to 12, wherein when the DL stack is trained by backpropagation, a first set of screenshot images is surrounded by the boundaries of various photographic images of multiple image-derived screenshots to generate a third set of screenshot images, and the first and third sets are combined with the labeled ground truth data. 14. A tangible, non-temporary, computer-readable storage medium as described in any one of Clauses 8 to 13, wherein when the DL stack is trained by backpropagation, a first set of screenshot images is surrounded by the boundaries of multiple overlaid program windows of multiple image-derived screenshots to generate a fourth set of screenshot images, and the first and fourth sets are combined with the labeled ground truth data. 15. A system for detecting screenshot images and preventing the loss of image-derived screenshots, comprising a processor, memory connected to the processor, and computer instructions loaded into the memory from a non-temporary computer-readable storage medium as described in Clause 8. 16. The system according to Clause 15, further comprising applying a screenshot robot to collect the aforementioned examples of screenshot images and non-screenshot images. 17. The DL stack includes at least a first set of layers closer to the input layer and a second set of layers further away from the input layer, wherein the first set of layers performs image recognition before the second set of layers of the DL stack applies the labeled ground truth data to the examples of the screenshot images and non-screenshot images. The system described in any one of clauses 15-16, further including being pre-trained. 18. The system described in any one of clauses 15 to 17, further comprising applying automatic re-rendering of at least a portion of the collected original screenshot images by cropping a portion of the image or by adjusting the hue, contrast, and saturation to represent changes in the screenshot. 19. The various changes in the aforementioned screenshot include at least one of the following: window size, window position, number of open windows, and menu bar position, as described in any one of clauses 15 to 18. 20. The system according to any one of clauses 15 to 19, wherein when the DL stack is trained by backpropagation, a first set of screenshot images is surrounded by the boundaries of various photographic images of multiple image-derived screenshots to generate a third set of screenshot images, and the first and third sets are combined with the labeled ground truth data. [Clause Set 3] 1. A method for customizing a deep learning (DL) stack to detect organizational confidential data in images, referred to as image-derived organizational confidential documents, and to prevent the loss of said image-derived organizational confidential documents: Pre-training a master DL stack using forward inference and backpropagation with labeled ground truth data on examples of image-derived sensitive documents and other image documents; Here, the DL stack comprises at least a first set of layers closer to the input layer and a second set of layers further away from the input layer, and further comprises the first set of layers being pre-trained to perform image recognition before the second set of layers of the DL stack is fitted with the labeled ground truth data for the image-derived sensitive documents and other image documents; To store the parameters of the trained master DL stack for inference from production images; Distributing the trained master DL stack, which has stored parameters, to multiple organizations; A method comprising enabling the organization to perform update training on the trained master DL stack using at least examples of the organization's confidential data in images, and to save the parameters of the updated DL stack, thereby enabling the organization to use each updated DL stack to classify, by inference, at least one production image as containing organization's confidential document. 2. The method according to Clause 1, comprising providing at least a portion of the organization with a dedicated DL stack trainer under the organization's control, enabling the organization to perform the update training without transferring examples of the organization's confidential data in images to a provider that performed pre-training of the master DL stack, wherein the dedicated DL stack trainer is configurable to generate the respective updated DL stacks. 3. The dedicated DL stack trainer is configured to combine irreversible features from the examples of the organizational confidential data in the image with the correct labels for the examples, and to transfer the irreversible features and the correct labels, The method of Clause 2, further comprising receiving organizational confidential training examples, including the irreversible features and correct labels, from a plurality of the dedicated DL stack trainers. 4. Use the organization-confidential training example to further train the second set of layers of the trained master DL stack; To store the updated parameters of the second set of layers for inference from the production images; and A clause including distributing the updated parameters of the second set of layers to multiple organizations. The method described in item 3. 5. The method according to any one of clauses 1 to 4, further comprising performing refresh training to further train the second set of layers of the trained master DL stack. 6. The method according to any one of the clauses 1 to 5, further comprising training the second set of layers of the trained master DL stack from the beginning using the organization confidential training example in a different order. 7. The dedicated DL stack trainer transfers the updated coefficients from the second set of layers. It is configured to send; and, Receiving the respective updated coefficients from each second set of layers from multiple dedicated DL stack trainers; To train the second set of layers of the trained master DL stack, combine the updated coefficients from each of the second set of layers; The method according to any one of clauses 2 to 4, comprising: storing updated parameters of the second set of layers of the trained master DL stack for inference from production images; and distributing the updated parameters of the second set of layers to the multiple organizations. 8. A tangible, non-temporary computer-readable storage medium containing program instructions loaded into memory, which, when executed on a processor, causes the processor to perform a method for customizing a deep learning (abbreviated as DL) stack to detect organizational confidential data within images, known as image-derived organizational confidential documents, and preventing the loss of said image-derived organizational confidential documents: the method Pre-training a master DL stack using forward inference and backpropagation with labeled ground truth data on examples of image-derived sensitive documents and other image documents; Here, the DL stack comprises at least a first set of layers closer to the input layer and a second set of layers further away from the input layer, and further comprises the first set of layers being pre-trained to perform image recognition before the second set of layers of the DL stack is fitted with the labeled ground truth data for the image-derived sensitive documents and other image documents; To store the parameters of the trained master DL stack for inference from production images; Distributing the trained master DL stack having the stored parameters to multiple organizations; The organization includes enabling the organization to perform update training on the trained master DL stack using at least examples of the organization's confidential data in images, and to save the parameters of the updated DL stack, thereby enabling the organization to use each updated DL stack to classify, by inference, at least one production image as containing organization's confidential document in a tangible, non-temporary, computer-readable storage medium. 9. A tangible, non-temporary computer-readable storage medium as described in Clause 8, comprising providing at least a portion of the organization with a dedicated DL stack trainer under the control of the organization, and enabling the organization to perform the update training without transferring examples of the organization's confidential data in images to a provider that performed pre-training of the master DL stack, wherein the dedicated DL stack trainer is configurable to generate the respective updated DL stacks. 10. A tangible, non-temporary, computer-readable storage medium according to Clause 9, further comprising: a dedicated DL stack trainer configured to combine irreversible features from the examples of the organization's confidential data in an image with the ground truth labels of the examples, and to transfer the irreversible features and ground truth labels; and receiving organization's confidential training examples containing the irreversible features and ground truth labels from a plurality of the dedicated DL stack trainers. 11. Use the organization-confidential training example to further train the second set of layers of the trained master DL stack; A tangible, non-temporary, computer-readable storage medium according to Clause 10, further comprising storing updated parameters of the second set of layers for inference from production images; and distributing the updated parameters of the second set of layers to multiple organizations. 12. A tangible, non-temporary, computer-readable storage medium as described in any one of clauses 8 to 11, further comprising performing update training to further train the second set of layers of the trained master DL stack. 13. A tangible, non-temporary, computer-readable storage medium as described in any one of Clauses 8 to 12, further comprising performing training from the beginning using the organization's confidential training example in a different order, in order to further train the second set of layers of the trained master DL stack. 14. The dedicated DL stack trainer is configured to transfer updated coefficients from the second set of layers; and to receive the respective updated coefficients from the respective second set of layers from a plurality of the dedicated DL stack trainers; To train the second set of layers of the trained master DL stack, combine the updated coefficients from each of the second set of layers; A tangible, non-temporary, computer-readable storage medium as described in any one of clauses 9 to 11, which includes storing updated parameters of the second set of layers of the trained master DL stack for inference from production images; and distributing the updated parameters of the second set of layers to the multiple organizations. 15. A system for customizing a deep learning (abbreviated as DL) stack to detect organizational confidential data within images, referred to as image-derived organizational confidential documents, and to prevent the loss of such image-derived organizational confidential documents, the system comprising a processor, memory connected to the processor, and computer instructions loaded into the memory from a non-temporary computer-readable storage medium as described in Clause 8. 16. The system according to Clause 15, further comprising providing at least a portion of the organization with a dedicated DL stack trainer under the organization's control, enabling the organization to perform the update training without transferring examples of the organization's confidential data in images to a provider that performed pre-training of the master DL stack, wherein the dedicated DL stack trainer is configurable to generate each updated DL stack. 17. The system according to Clause 16, further comprising: a dedicated DL stack trainer configured to combine irreversible features from the examples of the organization's confidential data in an image with the ground truth labels of the examples, and to transfer the irreversible features and ground truth labels; and receiving organization's confidential training examples containing the irreversible features and ground truth labels from a plurality of the dedicated DL stack trainers. 18. Use the organization-confidential training example to further train the second set of layers of the trained master DL stack; The system according to Clause 17, further comprising: storing updated parameters of the second set of layers for inference from production images; and distributing the updated parameters of the second set of layers to multiple organizations. 19. The system according to any one of clauses 15 to 18, further comprising performing refresh training to further train the second set of layers of the trained master DL stack. 20. The system according to any one of clauses 15 to 19, further comprising performing training from the beginning using the organization's confidential training example in a different order in order to further train the second set of layers of the trained master DL stack. 21. The dedicated DL stack trainer is configured to transfer updated coefficients from the second set of layers; and to receive the respective updated coefficients from the respective second set of layers from a plurality of the dedicated DL stack trainers; To train the second set of layers of the trained master DL stack, combine the updated coefficients from each of the second set of layers; A system according to any one of clauses 16 to 18, comprising: storing updated parameters of the second set of layers of the trained master DL stack for inference from production images; and distributing the updated parameters of the second set of layers to the multiple organizations.
Claims
1. A method for customizing a deep learning (DL) stack to detect organizational confidential data in images, referred to as image-derived organizational confidential documents, and to prevent the loss of said image-derived organizational confidential documents, wherein: Pre-training a master DL stack using forward inference and backpropagation with labeled ground truth data for examples of image-derived confidential documents and other image documents; Here, the DL stack includes at least a first set of layers closer to the input layer and a second set of layers further away from the input layer, and further includes the first set of layers being pre-trained to perform image recognition before the second set of layers of the DL stack is fitted with the labeled ground truth data for the image-derived sensitive documents and other image documents; To store the parameters of the trained master DL stack for inference from production images; Distributing the trained master DL stack, which has stored parameters, to multiple organizations; This includes enabling the organization to perform update training on the trained master DL stack using at least an example of the organization's confidential data in the image, and to save the parameters of the updated DL stack. This provides a method by which the organization uses its updated DL stack to classify, by inference, at least one production image as containing organizational confidential documents.
2. The method according to claim 1, comprising providing at least a portion of the organization with a dedicated DL stack trainer under the organization's control, enabling the organization to perform the update training without transferring examples of the organization's confidential data in images to a provider that performed pre-training of the master DL stack, wherein the dedicated DL stack trainer is configured to generate each updated DL stack.
3. The dedicated DL stack trainer is configured to combine irreversible features from the examples of the organizational confidential data in the image with the correct labels for the examples, and to transfer the irreversible features and the correct labels, and The method according to claim 2, further comprising receiving organizational confidential training examples, including the irreversible features and correct labels, from a plurality of the dedicated DL stack trainers.
4. To further train the second set of layers of the trained master DL stack, use the organization-confidential training example; Storing updated parameters of the second set of layers for inference from production images; and The method according to claim 3, comprising distributing the updated parameters of the second set of layers to a plurality of tissues.
5. The method according to any one of claims 1 to 4, further comprising performing update training to further train the second set of layers of the trained master DL stack.
6. The method according to any one of claims 1 to 5, further comprising training the second set of layers of the trained master DL stack from the beginning using the organization confidential training example in a different order, in order to further train the second set of layers of the trained master DL stack.
7. The dedicated DL stack trainer is configured to transfer updated coefficients from the second set of layers; and Receiving the updated coefficients from each second set of layers from multiple dedicated DL stack trainers; To train the second set of layers of the trained master DL stack, combine the updated coefficients from each of the second set of layers; To store updated parameters of the second set of layers of the trained master DL stack for inference from production images; and, The method according to any one of claims 2 to 4, comprising distributing the updated parameters of the second set of layers to the plurality of tissues.
8. A tangible, non-temporary, computer-readable storage medium comprising program instructions loaded into memory, which, when executed on a processor, cause the method according to any one of claims 1 to 7 to be carried out.
9. A system for customizing a deep learning (DL) stack to detect organizational confidential data in images, called image-derived organizational confidential documents, and for preventing the loss of said image-derived organizational confidential documents, the system comprising a processor, memory connected to the processor, and program instructions loaded into the memory, which, when executed on the processor, cause the processor to carry out the method according to any one of claims 1 to 7.