Image-based identification document detection for protecting confidential information

A deep learning-based system using a retrained CNN model with specialized images and online learning enhances image and screenshot detection in DLP, improving accuracy and reducing privacy concerns by storing extracted features, addressing computational inefficiencies and data privacy issues in existing DLP technologies.

JP7798488B2Active Publication Date: 2026-01-14NETSKOPE INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2021092862
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-04-13
Filing Date
2021-06-02
Publication Date
2026-01-14
Estimated Expiration
2041-06-02

AI Technical Summary

Technical Problem

Existing DLP technologies face challenges in accurately detecting sensitive information in images and screenshots due to the computational inefficiency of OCR and the difficulty in obtaining large-scale labeled datasets for deep learning, while also raising privacy concerns with centralized data storage.

Method used

A deep learning-based approach that reuses the last few layers of a pre-trained CNN model with a small number of specialized labeled images, combined with online learning and progressive refinement, stores extracted features instead of raw images to enhance detection accuracy without requiring extensive pre-labeled data, and uses a NetScope Cloud Access Security Broker (N-CASB) for network traffic analysis.

Benefits of technology

This method improves threat detection effectiveness by 20-25% and reduces privacy risks by enabling continuous model improvement without storing sensitive images, enhancing security in cloud-based environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007798488000001
    Figure 0007798488000001
  • Figure 0007798488000002
    Figure 0007798488000002
  • Figure 0007798488000003
    Figure 0007798488000003
Patent Text Reader

Abstract

To provide a method for detecting image-borne identification documents and prevent loss of the image-borne identification documents.SOLUTION: A method learns a DL stack by forward inference and back propagation using labelled ground truth data for image-borne identification documents and other image documents. The DL stack includes a first set of layers closer to an input layer and a second set of layers further from the input layer. The method performs pre-learning so that the first set of layers performs image recognition before exposing the second set of layers to the labelled ground truth data for the image-borne identification documents and examples of other image documents, stores parameters of the trained DL stack for inference from production images, and uses a production DL stack with the stored parameters to classify at least one production image by inference as containing a sensitive image-borne identification document.SELECTED DRAWING: Figure 11
Need to check novelty before this filing date? Find Prior Art

Description

Priority claim

[0001] This application claims priority to U.S. Patent Application No. 17 / 229,768, filed April 13, 2021, entitled "Deep Learning Stack for Production Use to Prevent Exfiltration of Image-Derived Identified Documents" (Attorney Docket No. NSKO1032-2), which is a continuation of U.S. Patent Application No. 16 / 891,647, filed June 3, 2020, entitled "Detection of Image-Derived Identified Documents for Protecting Sensitive Information" (Attorney Docket No. NSKO1032-1) (now U.S. Patent No. 10,990,856, issued April 27, 2021), and

[0002] Claiming priority to U.S. Patent Application No. 17 / 202,075, entitled "Training and Configuring a DL Stack to Detect Attempted Exfiltration of Sensitive Screenshot-Derived Data" (Attorney Docket No. NSKO1033-2), filed March 15, 2021, which is a continuation of U.S. Patent Application No. 16 / 891,678, filed June 3, 2020, entitled "Detecting Screenshot Images to Prevent Loss of Sensitive Screenshot-Derived Data" (Attorney Docket No. NSKO1033-1), now U.S. Patent No. 10,949,961, issued March 16, 2021; and

[0003] This application claims priority to U.S. Patent Application No. 17 / 116,862, entitled "Deep Learning-Based Image-Derived Sensitive Document Detection and Data Loss Prevention," filed December 9, 2020 (Attorney Docket No. NSKO 1034-2), which is a continuation of U.S. Patent Application No. 16 / 891,698, entitled "Detection of Sensitive Documents from Tissue Images and Prevention of Sensitive Document Loss," filed June 3, 2020 (Attorney Docket No. NSKO 1034-1), now U.S. Patent No. 10,867,073, issued December 15, 2020. These applications are incorporated by reference for all purposes. INCORPORATED MATTERS

[0004] The following documents are incorporated by reference into this application:

[0005] U.S. Patent Application No. 16 / 807,128, filed March 2, 2020, entitled "Load Balancing in a Dynamically Scalable Service Mesh" (Attorney Docket No. NSKO1025-3).

[0006] U.S. application Ser. No. 14 / 198,508, filed March 5, 2014, entitled "Security for Network Distribution Services" (Attorney Docket No. NSKO1000-3) (currently U.S. Patent No. 9,270,765, issued February 23, 2016).

[0007] U.S. application Ser. No. 14 / 198,499, filed March 5, 2014, entitled "Security for Network Distribution Services" (Attorney Docket No. NSKO1000-2) (currently U.S. Patent No. 9,398,102, issued July 19, 2016).

[0008] U.S. Application No. 14 / 835,640, filed August 25, 2015, entitled "System and Method for Monitoring and Controlling Business Information Stored in a Cloud Computing Service (CCS)" (Attorney Docket No. NSKO1001-2) (now U.S. Patent No. 9,928,377, issued March 27, 2018).

[0009] U.S. Provisional Application No. 62 / 307,305, filed March 11, 2016, entitled "System and Method for Enforcing Multipart Policies in Data-Loss Transactions in Cloud Computing Services" (Attorney Docket No. NSKO1003-1), claims the benefit of U.S. Provisional Application No. 15 / 368,246, filed December 2, 2016, entitled "Middleware Security Layer for Cloud Computing Services" (Attorney Docket No. NSKO1003-3).

[0010] Chen, Ital, Narayanaswamy, and Malmskog, "Cloud Security for Dummies, NetScope Special Edition," John Wiley & Sons, 2015.

[0011] "Netskope Introspection", published by Netskope, Inc.

[0012] "Data Loss Prevention and Monitoring in the Cloud," published by Netskope, Inc.

[0013] "Cloud Data Loss Prevention Reference Architecture," published by Netskope, Inc.

[0014] "Five Steps to Cloud Confidence," published by NetScope, Inc.

[0015] "Netskope Active Platform" published by Netskope, Inc.

[0016] "The Netskope Advantage: Three 'Must-Have' Requirements for Cloud Access Security Brokers" published by Netskope, Inc.

[0017] "15 Important CASB Use Cases," published by Netskope, Inc.

[0018] "Netskope Active Cloud DLP," published by Netskope, Inc.

[0019] "Remediating the Collision Course of Cloud Data Breaches," published by Netskope, Inc.

[0020] "Netskope Cloud Confidence Index(TM)", published by Netskope, Inc.

[0021] The above documents are incorporated by reference as if fully set forth herein. [Technical Field]

[0022] The disclosed technology generally relates to security for network-distributed services, and more particularly to detecting identifying documents within images, referred to as image-derived identifying documents, and applying security services to prevent the loss of the image-derived identifying documents. The disclosed technology also relates to detecting screenshot images and preventing the loss of screenshot-derived data. Furthermore, a separate organization can utilize the disclosed technology to detect image-derived identifying documents and detect screenshot images from within the organization, so that the organization's images, potentially containing sensitive data, do not need to be shared with a data loss prevention service provider. [Background technology]

[0023] It should not be assumed that the subject matter discussed in this section is prior art merely by virtue of its mention in this section. Likewise, it should not be assumed that any problem stated in this section or related to the subject matter provided as background art has already been recognized in the prior art. The subject matter in this section is merely illustrative of various approaches, and may, by itself or spontaneously, accommodate the implementation of the claimed technology.

[0024] Data Loss Prevention (DLP) technology is widely used in the security industry to prevent the leakage of sensitive information, such as personally identifiable information (PII), protected health information (PHI), and intellectual property (IP). DLP products are used by both large and small businesses. Such sensitive information resides in a variety of sources, including documents and images. It is important for any DLP product to be able to detect sensitive information in documents and images with high accuracy and computational efficiency.

[0025] For text documents, DLP products use string and regular expression-based pattern matching to identify sensitive information. For images, optical character recognition (OCR) technology has been used to first extract text characters. The extracted characters are then sent through the same pattern matching process to detect sensitive information. Historically, OCR has not performed well because it requires a lot of computational resources and has poor accuracy, especially when images are not in ideal conditions, such as when they are blurry, smudged, rotated, or inverted.

[0026] While training can be automated, there remains the problem of assembling the training data in the correct format and sending the data to a central computational node with sufficient storage and computing power. In many fields, transmitting personally identifiable private data to any central authority raises concerns about data privacy, including data security, data ownership, privacy protection, and appropriate permissions and use of the data.

[0027] Deep learning applies multi-layer networks to data. In recent years, deep learning techniques have been increasingly used in image classification. Deep learning can detect images containing sensitive information without expensive OCR processing. A key challenge of deep learning approaches is the need for a large number of high-quality labeled images that represent real-world distributions. Unfortunately, for DLP, high-quality labeled images typically utilize real images containing sensitive information, such as authentic passport images and authentic driver's license images. These data sources are inherently difficult to acquire on a large scale. This limitation hinders the adoption of deep learning-based image classification in DLP products.

[0028] An opportunity exists to efficiently detect identified documents in images with a 20-25% improvement in threat detection effectiveness and prevent the loss of sensitive data from image-derived identified documents. Additionally, an opportunity exists to detect screenshot images and prevent the loss of sensitive screenshot-derived data, potentially resulting in cost and time savings in security systems utilized by SaaS customers.

[0029] In the drawings, like reference numerals generally refer to like parts throughout the different views. Also, the drawings are not necessarily to scale, emphasis instead generally being placed upon illustrating the principles of the disclosed technology. In the following description, various embodiments of the disclosed technology are described with reference to the following drawings: [Brief explanation of the drawings]

[0030] [Figure 1A] 1 illustrates an architecture-level schematic diagram of a system for detecting identifying documents in images, called image-derived identifying documents, and preventing the loss of image-derived identifying documents while applying security services in the cloud. The disclosed system can also detect screenshot images and prevent the loss of sensitive screenshot-derived data.

[0031] [Figure 1B] This paper presents an image-derived sensitive data detection aspect of an architecture for detecting identifying documents in images, called image-derived identifying documents, and applying security services in the cloud to prevent the loss of image-derived identifying documents, detect screenshot images, and prevent the loss of sensitive screenshot-derived data.

[0032] [Figure 2]1 illustrates a block diagram of a deep learning stack implemented using a convolutional neural network architecture model for image classification that can be configured for use in a system for detecting identifying documents in images and detecting screenshot images, according to one embodiment of the disclosed technology.

[0033] [Figure 3] We present the precision and recall results of the trained passport and driver's license classifiers.

[0034] [Figure 4] We show the running time results for classifying images, graphed as a distribution of images.

[0035] [Figure 5] We present benchmarking results for classifying sensitive images on US driver's licenses.

[0036] [Figure 6] We present an example workflow for detecting discriminative documents in images, called image-derived discriminative documents, and training a deep learning stack to prevent the loss of image-derived discriminative documents.

[0037] [Figure 7] 1 shows an example screenshot with an inventory list with costs listed.

[0038] 8A, 8B, 8C, and 8D show four false positive screenshot images.

[0039] [Figure 8A] Shows a misclassified Idaho map as a screenshot for the legend window and dotted lines above and below.

[0040] [Figure 8B]This shows a misclassified driver's license image as a screenshot because the entire image is a window containing PII within a black background, and the UNITED STATES bar can be treated as a header bar.

[0041] [Figure 8C] It shows a passport image as the main window containing PII, and the shaded area at the bottom center could mislead the classifier into thinking it is an application bar.

[0042] [Figure 8D] Shows characters in a measure window containing text information and a uniform background.

[0043] [Figure 9] FIG. 1 is a simplified block diagram of a computer system that can be used to detect screenshot images and prevent image-derived screenshot loss, according to one embodiment of the disclosed technology, that can be used to perform detection of identification documents in images and prevent image-derived identification document loss.

[0044] [Figure 10] We present an example workflow for detecting discriminative documents in images, called image-derived discriminative documents, and training a deep learning stack to prevent the loss of image-derived discriminative documents.

[0045] [Figure 11] A workflow is shown for one or more computer systems that can be configured to perform detection of identifying documents in images and prevent loss of image-derived identifying documents and can be used to detect screenshot images and prevent loss of image-derived screenshots. DETAILED DESCRIPTION OF THE INVENTION

[0046] The following detailed description is provided with reference to the drawings. Exemplary embodiments are described to illustrate the disclosed technology, but not to limit the scope of the invention as defined by the claims. Those skilled in the art will recognize various equivalent variations of the following description.

[0047] Deep learning technology can be used to enhance the detection of sensitive information from documents and images, and to detect images containing sensitive information without the need for expensive OCR processes. Deep learning uses optimization to find optimal parameter values ​​for a model to make the best predictions. Deep learning-based image classification typically requires a large number of labeled images containing sensitive information, which are difficult to obtain on a large scale. This limitation hinders the adoption of deep learning-based image classification in DLP products.

[0048] The disclosed innovation applies deep learning-based image classification in data loss prevention (DLP) products without requiring a large number of pre-labeled images containing sensitive information. Many pre-trained, general-purpose deep learning models available today use the public ImageNet dataset and other similar sources. These deep learning models are typically multi-layer convolutional neural networks (CNNs) capable of classifying general objects such as cats, dogs, cars, etc. The disclosed technology uses a small number of specialized, labeled images, such as passport and driver's license images, to retrain the last few layers of the CNN model. In this way, the deep learning (DL) stack can detect these specific images with high accuracy without requiring a large number of pre-labeled images containing sensitive data.

[0049] DLP products in customer deployments can process customer production traffic and continuously generate new labels. To minimize privacy issues, new labels can be kept in production using online learning, and whenever a sufficient batch of new labels accumulates, a similar number of negative images can be injected to create a new, balanced, incremental dataset that can be used to incrementally refine existing deep learning models using progressive learning.

[0050] Even with online and progressive learning, a typical deep learning process requires input of the original image and newly added images to create a sophisticated model for predicting the presence of sensitive data in an image document or screenshot. This means that the system must store newly labeled images generated in production for long periods of time. In a production environment, users' private data is more secure than if the images and labels were stored offline, but storing images raises privacy concerns if sensitive data is stored in persistent storage.

[0051] The disclosed method saves the output of a deep learning stack, also known as a neural network, storing the extracted features instead of the raw image. In a typical neural network, the raw image passes through many layers before a final set of features is extracted for the final classifier. These features cannot be reverse-transformed back to the original raw image. This feature of the disclosed technique allows for the protection of sensitive information in production images, and the model's saved features can be used to retrain the classifier in the future.

[0052] The disclosed technique provides accuracy and high performance in classifying images containing sensitive information and screenshot images without requiring a large number of pre-labeled images. The technique also enables the use of production images to continuously improve accuracy and coverage without privacy concerns.

[0053] The disclosed innovation further expands the ability to utilize machine learning classification to detect sensitive image content and enforce policies, applying advances in image classification and screenshot detection to network traffic proxied in the cloud in the context of a NetScope Cloud Access Security Broker (N-CASB), as described herein.

[0054] An example system for detecting identifying documents in images, referred to as image-derived identifying documents, and preventing the loss of image-derived identifying documents in the cloud, as well as detecting screenshot images and preventing the loss of sensitive screenshot-derived data, is now described. [architecture]

[0055] FIG. 1A illustrates an architecture-level schematic diagram of a system 100 for detecting identification documents, referred to as image-derived identification documents, in images and preventing the loss of image-derived identification documents in the cloud. The system 100 can also detect screenshot images and prevent the loss of sensitive screenshot-derived data. Because FIG. 1A is an architecture diagram, certain details have been intentionally omitted to improve clarity of the description. The description of FIG. 1A is organized as follows: first, the elements of the diagram are described, and then their interconnections are described. Next, the use of the elements in the system is described in more detail. FIG. 1B illustrates the image-derived sensitive data detection aspect of the system, which is described later.

[0056] The system 100 includes an organizational network 102, a data center 152 with a NetScope Cloud Access Security Broker (N-CASB) 155, and cloud-based services 108. The system 100 includes multiple organizational networks 104, sometimes referred to as multi-tenant networks, for multiple subscribers of a security service provider, and multiple data centers 154, sometimes referred to as branches. The organizational network 102 includes computers 112a-n, tablets 122a-n, mobile phones 132a-n, and smartwatches 142a-n. In other organizational networks, organizational users may utilize additional devices. The cloud services 108 include a cloud-based hosting service 118, a webmail service 128, a video, messaging, and voice calling service 138, a streaming service 148, a file transfer service 158, and a cloud-based storage service 168. The data center 152 connects to the organizational network 102 and the cloud-based services 108 via a public network 145.

[0057] Continuing with FIG. 1A, the disclosed enhanced NetScope Cloud Access Security Broker (N-CASB) 155 manages access and activity in authorized and unauthorized cloud apps, secures sensitive data, prevents its loss, and protects against internal and external threats. It also securely handles Skype, voice, video, and messaging multimedia communication sessions over SIP, web traffic over other protocols, and P2P traffic over BT, FTP, and UDP-based streaming protocols. For data loss prevention, N-CASB 155 utilizes machine learning classification for identity detection and sensitive screenshot detection, further extending its ability to detect sensitive image content and enforce policies. N-CASB 155 includes an active analyzer 165 and an introspective analyzer 175 that identify system users and set policies for applications. Introspective analyzer 175 interacts directly with cloud-based services 108 to inspect data at rest. In polling mode, the introspective analyzer 175 calls cloud-based services using API connectors to crawl data residing in the cloud-based services and check for changes. For example, the Box™ storage application provides an administrative API called the Box Content API™. This administrative API provides visibility into all users' organization accounts, including audit logs for Box folders, and can be inspected to determine whether sensitive files have been downloaded since a specific date when credentials were compromised. The introspective analyzer 175 polls this API to discover any changes made to any of the accounts. If a change is discovered, the Box Event API™ is polled to discover detailed data changes. In the callback model, the introspective analyzer 175 registers with the cloud-based service via its API connector to be notified of significant events.For example, introspective analyzer 175 can use the Microsoft Office 365 Webhooks API™ to determine when files are shared externally. Introspective analyzer 175 also includes a DLP engine with deep API inspection, deep packet inspection, and log inspection capabilities that applies various content inspection techniques to files at rest within the cloud-based service to determine which documents and files are sensitive based on policies and rules stored in storage 186. Inspections by introspective analyzer 175 generate per-user and per-file data.

[0058] Continuing with FIG. 1A, N-CASB 155 further includes monitor 184, which includes extraction engine 171, classification engine 172, security engine 173, management plane 174, and data plane 180. N-CASB 155 also includes storage 186, which includes deep learning stack parameters 183, features and labels 185, content policies 187, content profiles 188, content inspection rules 189, enterprise data 197, customer 198, and user identities 199. Enterprise data 197 may include organizational data, including, but not limited to, intellectual property, non-public financial information, strategic plans, customer lists, personally identifiable information (PII) belonging to customers or employees, patient health data, source code, trade secrets, reservation information, collaboration agreements, corporate plans, merger and acquisition documents, and other confidential data. Specifically, the term "enterprise data" refers to documents, files, folders, web pages, collections of web pages, images, or other text-based documents. User identity refers to an indicator provided to a client device by a network security system in the form of a token, a unique identifier such as a UUID, a public key certificate, etc. In some cases, a user identity can be linked to a specific user and a specific device. Thus, the same individual can have different user identities on their mobile phone and computer. A user identity can be linked to, but is not limited to, a corporate identity directory of entries or user IDs. In one embodiment, a cryptographic certificate signed by the network security system is used as the user identity. In other embodiments, a user identity is unique only to the user and can be identical across devices.

[0059] Embodiments may also interoperate with single sign-on (SSO) solutions and / or enterprise identity directories, such as Microsoft Active Directory. Such embodiments may allow policies to be defined within the directory, for example, at either the group or user level, using custom attributes. Hosted services configured in the system are also configured to require traffic through the system. This can be done by setting IP range restrictions on the hosted service to the IP range of the system and / or integration between the system and the SSO system. For example, integration with an SSO solution can enforce a client presence requirement before allowing sign-on. Other embodiments may use a "proxy account" with a SaaS vendor, e.g., a dedicated account maintained by the system that holds the sole credentials for signing into the service. In other embodiments, the client may encrypt sign-on credentials before passing the login to the hosted service, meaning that the network security system "owns" the password.

[0060] Storage 186 can store information from one or more tenants in tables in a common database image to form an on-demand database service (ODDS), which can be implemented in many ways, such as a multi-tenant database system (MTDS). The database image can include one or more database objects. In other embodiments, the database can be a relational database management system (RDBMS), an object-oriented database management system (OODBMS), a distributed file system (DFS), a schema-less database, or any other data storage system or computing device. In some embodiments, the collected metadata is processed and / or normalized. In some cases, the metadata includes structured data and functional target-specific data structures provided by cloud service 108. Unstructured data, such as free text, can also be provided by and returned to cloud service 108. Both structured and unstructured data can be aggregated by introspective analyzer 175. For example, assembled metadata may be stored in a semi-structured data format such as JSON (JavaScript Option Notation), BSON (Binary JSON), XML, Protobuf, Avro, or Thrift objects, which consist of string fields (or columns) and corresponding values, potentially of various types, such as numbers, strings, objects, arrays, objects, etc. JSON objects can be nested, and fields can be multi-valued, in other implementations, into arrays, nested arrays, etc.These JSON objects are stored in a schemaless or NoSQL key-value metadata store 148, such as Apache Cassandra™ 158, Google's BigTable™, HBase™, Voldemort™, CouchDB™, MongoDB™, Redis™, Riak™, Neo4j™, etc., which uses keyspaces, the SQL equivalent of a database, to store the parsed JSON objects. Each keyspace is similar to a table and is divided into column families, each containing a set of rows and columns.

[0061] In one embodiment, introspective analyzer 175 includes a metadata parser (not shown for clarity) that analyzes input metadata and identifies keywords, events, user IDs, locations, demographics, file types, timestamps, etc. in the received data. Because the metadata analyzed by introspective analyzer 175 is not homogenous (e.g., from many different sources in many different formats), some embodiments use at least one metadata parser per cloud service, and possibly multiple metadata parsers. In other embodiments, introspective analyzer 175 uses monitors 184 to inspect cloud services and assemble content metadata. In one use case, identification of sensitive documents is based on a pre-inspection of the documents. A user can manually tag a document as sensitive, and this manual tagging updates the document metadata in the cloud service. The document metadata can then be retrieved from the cloud service using exposed APIs and used as an indicator of sensitivity.

[0062] Continuing with FIG. 1A, the system 100 can include any number of cloud-based services 108, such as point-to-point streaming services, hosted services, cloud applications, cloud stores, cloud collaboration and messaging platforms, and cloud customer relationship management (CRM) platforms. These services can include peer-to-peer file sharing (P2P) via portal traffic protocols such as BitTorrent (BT), User Data Protocol (UDP) streaming, and File Transfer Protocol (FTP); instant messaging over Internet Protocol (IP); and voice, video, and messaging multimedia communication sessions such as cellular calls over LTE (VoLTE) via Session Initiation Protocol (SIP) and Skype. These services can handle Internet traffic, cloud application data, and Generic Routing Encapsulation (GRE) data. Network services or applications can be web-based (e.g., accessed via Uniform Resource Locators (URLs)) or native, such as synchronization clients. Examples include Software as a Service (SaaS) offerings, Platform as a Service (PaaS) offerings, and Infrastructure as a Service (IaaS) offerings, as well as internal corporate applications exposed via a URL. Examples of cloud-based services common today include Salesforce.com™, Box™, Dropbox™, Google Apps™, Amazon AWS™, Microsoft Office 365™, Workday™, Oracle on Demand™, Taleo™, Yammer™, Jive™, and Concur™.

[0063] Interconnecting the elements of system 100, network 145 communicatively couples computers 112a-n, tablets 122a-n, mobile phones 132a-n, smartwatches 142a-n, cloud-based hosting service 118, web-based email service 128, video, messaging, and voice call service 138, streaming service 148, file transfer service 158, cloud-based storage service 168, and N-CASB 155. Communication paths can be point-to-point over public and / or private networks. Communications can occur over various networks, such as private networks, VPNs, MPLS lines, or the Internet, and can use appropriate application program interfaces (APIs) and data exchange formats, such as REST, JSON, XML, SOAP, and JMS. All communications can be encrypted. This communication typically occurs over networks such as LANs (Local Area Networks), WANs (Wide Area Networks), telephone networks (Public Switched Telephone Networks), Session Initiation Protocol (SIP), wireless networks, point-to-point networks, star networks, token ring networks, hubbed networks, and the Internet, including mobile Internet via protocols such as EDGE, 3G, 4G LTE, Wi-Fi, WiMAX, etc. Furthermore, communications can be secured using various authorization and authentication technologies such as username / password, open authentication (OAuth), Kerberos, SecureID, digital certificates, etc.

[0064] Continuing with the description of the system architecture of FIG. 1A, N-CASB 155 includes monitor 184 and storage 186, which may include one or more computers and computer systems communicatively coupled to each other. They may also be one or more virtual computing and / or storage resources. For example, monitor 184 may be one or more Amazon EC2 instances, and storage 186 may be Amazon S3™ storage. Rather than implementing N-CASB 155 directly on physical computers or traditional virtual machines, other compute-as-a-service platforms, such as Salesforce's Rackspace, Heroku, or Force.com, may be used. Furthermore, one or more engines may be used, and one or more points of presence (POPs) may be established to implement security functions. The engines or system components of FIG. 1A are implemented by software running on various types of computing devices. Examples of devices include workstations, servers, computing clusters, blade servers, server farms, or other data processing systems or computing devices. The engines may be communicatively coupled to a database via a different network connection. For example, extraction engine 171 may be coupled via network 145 (e.g., the Internet), classification engine 172 may be coupled via a direct network link, and security engine 173 may be coupled by a different network connection. In the disclosed technology, the POPs of data plane 180 are hosted on the client's premises or located within a virtual private network controlled by the client.

[0065] N-CASB 155 provides various functions via management plane 174 and data plane 180. According to one embodiment, data plane 180 includes extraction engine 171, classification engine 172, and security engine 173. Other functions, such as a control plane, may also be provided. Collectively, these functions provide a secure interface between cloud services 108 and the organizational network 102. While the term "network security system" is used to describe N-CASB 155, more generally, the system provides not only security but also application visibility and control. In one example, 35,000 cloud applications reside in a library across servers used by computers 112a-n, tablets 122a-n, mobile phones 132a-n, and smartwatches 142a-n within the organizational network 102.

[0066] According to one embodiment, the computers 112a-n, tablets 122a-n, mobile phones 132a-n, and smartwatches 142a-n in the organizational network 102 include administrative clients with web browsers that have a secure web delivery interface provided by the N-CASB 155 for defining and managing content policies 187. Because the N-CASB 155 is a multi-tenant system, users of the administrative clients can modify only the content policies 187 associated with their organization, depending on the implementation. In some implementations, an API can be provided for programmatically defining and updating policies. In such an implementation, the administrative clients can include one or more servers, e.g., an enterprise identity directory such as Microsoft Active Directory, that can push updates and / or respond to pull requests for updates to the content policies 187. Both systems can coexist. For example, the enterprise identity directory can be used to automate the identification of users within the organization, while a web interface can be used to tailor policies to their needs. Administrative clients are assigned roles and access to N-CASB 155 data is controlled based on the role, e.g., read-only vs. read-write.

[0067] In addition to periodically generating and maintaining per-user and per-file data in the metadata store 178, active and introspective analyzers (not shown) also enforce security policies on cloud traffic. For more information regarding the functionality of active and introspective analyzers, reference may be made to, for example, commonly owned U.S. Patent No. 9,398,102 (Attorney Docket No. NSKO1000-2); U.S. Patent No. 9,270,765 (Attorney Docket No. NSKO1000-3); U.S. Patent No. 9,928,377 (Attorney Docket No. NSKO1001-2); and U.S. Application No. 15 / 368,246 (Attorney Docket No. NSKO1003-3); Chen, Ital, Narayanaswamy, and Malmskog, "Cloud Security for Dummies, Netscope Special Edition," John Wiley & Sons, 2015; "Netscope Introspection." "Data Loss Prevention and Monitoring in the Cloud," published by Netskope, Inc.; "Cloud Data Loss Prevention Reference Architecture," published by Netskope, Inc.; "Five Steps to Cloud Confidence," published by Netskope, Inc.; "Netskope Active Platform," published by Netskope, Inc.; "The Netskope Advantage: Three 'Must-Have' Requirements for Cloud Access Security Brokers," published by Netskope, Inc.; "15 Critical CASB Use Cases," published by Netskope, Inc.; "Netskope Active Cloud DLP," published by Netskope, Inc.; "Remediating the Collision Course for Cloud Data Breaches," published by Netskope, Inc.; and "Netskope Cloud Confidence Index(TM)”, published by Netskope, Inc.The above documents are incorporated by reference as if fully set forth herein.

[0068] In the case of system 100, a control plane can be used in conjunction with or instead of management plane 174 and data plane 180. The specific division of functionality between these groups is an implementation choice. Similarly, functionality can be highly distributed across several points of presence (POPs) to improve locality, performance, and / or security. In one implementation, the data plane is on a premises or virtual private network, and the network security system's management plane is located in a cloud service or enterprise network, as described herein. In other secure network implementations, POPs can be distributed differently.

[0069] While system 100 is described herein with reference to particular blocks, it should be understood that the blocks are defined for convenience of description and are not intended to require a particular physical arrangement of components. Furthermore, the blocks need not correspond to physically separate components. To the extent physically separate components are used, connections between components can be wired and / or wireless, as desired. Different elements or components can be combined into a single software module, and multiple software modules can be executed on the same hardware.

[0070] Additionally, the techniques can be implemented using two or more separate and distinct computer-implemented systems that cooperate and communicate with each other. The techniques can be realized in numerous ways, including as a process, a method, an apparatus, a system, a device, a computer-readable medium such as a computer-readable storage medium that stores computer-readable instructions or computer program code, or a computer program product that includes a computer usable medium having computer-readable program code embodied therein. The disclosed technology may be implemented in the context of a database system or any computer-implemented system including a relational database implementation such as an Oracle™-compatible database implementation, an IBM DB2 Enterprise Server™-compatible relational database implementation, a MySQL™- or PostgreSQL™-compatible relational database implementation, or a Microsoft SQL Server™-compatible relational database implementation, or a NoSQL non-relational database implementation such as a Vampire™-compatible non-relational database implementation, an Apache Cassandra™-compatible non-relational database implementation, a BigTable™-compatible non-relational database implementation, or an HBase™- or Dynamo™-compatible non-relational database implementation. Additionally, the disclosed techniques may be implemented using a variety of programming models, such as MapReduce™, bulk synchronous programs, MPI primitives, etc., or a variety of scalable batch and stream management systems, such as Amazon Web Services (AWS™), including Amazon Elasticsearch Service™ and Amazon Kinesis™, Apache Storm™, Apache Spark™, ​​Apache Kafka™, Apache Flink™, Truviso™, IBM Info-Sphere™, Borealis™, and Yahoo! S4™.

[0071] Early deep learning models can perform well on the datasets used for training. For unseen images, performance is unpredictable. There is a continuing need to increase dataset coverage of real-world scenarios.

[0072] FIG. 1B illustrates the image-derived sensitive data detection aspect of the system 100 described above in connection with FIG. 1A, including an organization network 102, a data center 152, and a cloud-based service 108. Each individual organization network 102 has a user interface 103 for interacting with the data loss prevention functionality and includes a deep learning stack trainer 162. The dedicated DLP stack trainer can be configured to generate updated DLP stacks for each organization under the organization's control. The deep learning stack trainer 162 allows a client organization to perform updated training of its image and screenshot classifiers without the organization transferring sensitive data in images to the DLP provider that performed the pre-training of the master DLP stack. This reduces the requirements for protecting stored sensitive data stored at the DLP center, since PII and other sensitive data are protected from access by the data loss prevention provider. DLP stack training is discussed further below.

[0073] Continuing with FIG. 1B, data center 152 includes a NetScope Cloud Access Security Broker (N-CASB) 155, which includes image-derived sensitive data detection 156, a deep learning stack 157 with inference and backpropagation 166, and an image generation robot 167. Deep learning (DL) stack parameters 183 and features and labels 185 are stored in storage 186, as described in detail above. Deep learning stack 157 utilizes the stored features and labels 185, generated as output from the first set of layers of the stack and retained along with their respective ground truth labels for progressive online deep learning, thereby eliminating the need to retain images of private image-derived identified documents. As new image-derived identified documents are received, they can be classified by the trained DL stack, as described below.

[0074] Image generation robot 167 generates examples of actual passport images and U.S. driver's license images, as well as other image documents, for use in training deep learning stack 157. In one example, image generation robot 167 crawls sample U.S. driver's license images via a web-based search engine, inspects the images, and filters out low-fidelity images.

[0075] The image generation robot 167 also collects example screenshot and non-screenshot images, creates labeled ground truth data for the example images, and applies re-rendering of at least some of the collected example screenshot images representing various variations of screenshots that may contain sensitive information to create synthetic data for training the deep learning stack 157, leveraging available tools for web UI automation. One example tool is Selenium, an open-source tool that can simulate opening a web browser, visiting a website, opening a document, and clicking on a page. For example, this tool can start with a plain desktop, open one or more web browsers of various sizes in various locations on the desktop, and access a live website or open a predetermined local document. These actions can then be repeated with randomized parameters, such as the number of browser windows, the size and location of the browser windows, and the relative positioning of the browser windows. The image generation robot 167 then takes screenshots of the desktop and re-renders the screenshots, including augmenting the generated sample images as training data to feed to the DL stack 157. For example, this processing can add noise to the images to increase the robustness of the DL stack 157. The augmentations applied to our training data include cropping portions of the images and adjusting the hue, contrast, and saturation. To detect screenshot images that people use to secretly extract data, no flipping or rotation was added to the image augmentation. In a different example implementation, flipping and rotation could be added to other image document examples.

[0076] Figure 2 shows a block diagram of a deep learning (DL) stack 157 implemented using a convolutional neural network (CNN) architecture model for image classification, which can be configured for use in a system for detecting identifying documents in images and detecting screenshot images. The image of the CNN architecture model was downloaded from https: / / towardsdatascience.com / covolutional-neural-network-cb0883dd6529 on April 28, 2020. The input to the initial CNN layer is the image data itself, represented as a three-dimensional matrix with image dimensions and three color channels: red, green, and blue. The input image can be 224 x 224 x 3, as shown in Figure 2. In another implementation, the input image can be 200 x 200 x 3. In the example implementation whose results are presented later, the image dimensions used are 160 x 160 x 3, for a total of 88 layers. 1630995073802_0

[0077] Continuing with the description of the DL stack 157, the feature extraction layers are the convolution layer 245 and the pooling layer 255. The disclosed system stores the features and labels 185 output of the feature extraction layer as numerical values ​​processed through many different iterations of the convolution operation, preserving lossy features instead of the raw image. The extracted features cannot be reversed back to the original image pixel data; that is, the stored features are lossy features. By storing these extracted features instead of the input image data, the DL stack does not store pixels of the original image that may carry sensitive and private information, such as personally identifiable information (PII), protected health information (PHI), and intellectual property (IP).

[0078] The DL stack 157 includes a first set of layers closer to the input layer and a second set of layers further from the input layer. The first set of layers is pre-trained to perform image recognition before applying labeled ground truth data for image-derived identification documents and other image document examples to the second set of layers in the DL stack. The disclosed DL stack 157 freezes the first 50 layers as the first set of layers. The DL stack 157 is trained by forward inference and back propagation 166 using labeled ground truth data for image-derived identification documents and other image document examples. For private image-derived identification documents and screenshot images, the CNN architecture model captures the features generated as output from the first set of layers and retains the captured features along with their respective ground truth labels, thereby eliminating the need to retain images of private image-derived identification documents. The fully connected layer 265 and the SoftMax layer 275 comprise a second set of layers further from the input layer of the CNN being trained, and together with the first set of layers, the model is utilized to detect identifying documents in images and to detect screenshot images.

[0079] Training the DL stack 157 via forward inference and backpropagation 166 utilizes labeled ground truth data for image-derived identification documents and other image document examples. The first set of layers is pre-trained to perform image recognition before feeding the second set of layers in the DL stack with labeled ground truth data for image-derived identification documents and other image document examples. The output of the image classifier can be used to train the second set of layers; in one example, only images classified as the same type by both the OCR and the image classifier are fed as labeled images to the deep learning stack.

[0080] The disclosed technology stores parameters of a DL stack 183 trained for inference from production images, and uses the production DL stack with the stored parameters to infer and classify the production image as containing a sensitive image-derived identified document in one use case, or a screenshot image in another.

[0081] In one use case, the objective was to develop an image classification deep learning model for detecting passport images. Initial training data for building a deep learning-based binary image classifier for passport classification was generated using approximately 550 passports from 55 countries as labeled ground truth data for detecting image-based identification documents. Because the objective was to detect passports with a high detection rate, detecting other identification document types as passports was unacceptable. Images of other identification document types, including driver's licenses, identity cards, student identifiers, and military IDs, as well as images of non-identification documents, were used as negative datasets. To meet the goal of minimizing the detection rate of other identification document types, these other identification document images were used in the negative dataset.

[0082] In the second use case, the objective was to develop an image classifier for detecting passport and U.S. driver's license images. Training data for building a deep learning-based binary image classifier for passport classification was generated using 550 passport images and 248 U.S. driver's license images. In addition to the actual passport and U.S. driver's license images, sample U.S. driver's license images obtained by crawling the internet were included after inspection and filtering out low-fidelity images.

[0083] Cross-validation techniques were used to evaluate DL stack models by training several models on subsets of available input data and evaluating them on complementary subsets of the data. In k-fold cross-validation, the input data is divided into k subsets of data, also known as folds. To check the performance of the resulting image classifier, 10-fold cross-validation was applied. The cutoff values ​​for US driver's licenses and passports were selected as 0.3 and 0.8, respectively, to check the precision and recall of the models.

[0084] Figure 3 shows the precision and recall results for the trained passport and driver's license classifier, plotted with precision for driver's licenses 345 and recall for driver's licenses 355, precision for passports 365 and recall for passport images 375, and precision for non-ID documents (non-driver's licenses or passports), also called negative results 385 and recall for negative results 395. As shown in the graph, recall decreases as precision increases. The designers used 10-fold cross-validation to check the performance of the passport image classifier. The false positive rate (FPR) was calculated for the non-ID document images in the test, and the false negative rate (FNR) was calculated for the passport and driver's license images in the test. The results of the 10-fold cross-validation were averaged, and the averaged FPR and FNR are listed below. Passport FPR (non-ID photo classified as passport): 0.7% US Driver's License FPR (non-ID image classified as US Driver's License): 0.3% Passport FNR (passport image not classified as passport): 6% US Driver's License FNR (US Driver's License image not classified as a driver's license): 6%

[0085] Figure 4 shows the runtime results for classifying images tested using model inference on Google Cloud Platform (GCP) (n1-highcpu-64: 64 vCPUs, 57.6 GB memory) using over 1,000 images of various file sizes, plotted as a distribution of images. The graph shows the runtime distribution of images as a function of file size for images with file sizes of 2 MB or less. The runtime is counted from the time "opencv" reads the image to the time the classifier finishes its prediction on the image. The average runtime was 45 ms, with a standard deviation of 56 ms.

[0086] Figure 5 shows benchmarking results comparing a commercially available classifier with the significantly improved performance of the disclosed deep learning stack for classifying confidential images of U.S. driver's licenses. The number of images classified is 334. Using the commercial classifier, which uses regular expression (Regex) OCR and pattern matching, the number of images detected is 238 out of 334, representing a 71.2% detection rate. The majority of confidential images are detected, and the system performs only reasonably well. In some images, the classifier is unable to extract blurred or rotated text. In contrast, the disclosed technology utilizing the deep learning stack detects 329 out of 334 images, representing a 98.5% detection rate of images containing confidential image-derived identification documents.

[0087] FIG. 6 illustrates an example workflow 600 for detecting discriminative documents in images, referred to as image-derived discriminative documents, and preventing the loss of image-derived discriminative documents. Step 605 involves selecting a pre-trained network, such as the CNN described above in connection with FIG. 2. The DL stack includes at least a first set of layers closer to the input layer and a second set of layers further from the input layer, where the first set of layers are pre-trained to perform image recognition. In the example described, a MobileNet CNN was selected to detect images. A different CNN or a different ML classifier could be selected. Step 615 covers the collection of images containing sensitive information, balanced with negative images, as described for the two use cases. In step 625, both the final layer of the pre-trained network and the CNN model's classifier are retrained, and the CNN model is validated and tested. Using the labeled ground truth data for image-derived identification documents and other image document examples collected in step 615, a DL stack is trained by forward inference and backpropagation, and the labeled ground truth data for image-derived identification documents and other image document examples is applied to a second set of layers in the DL stack. In step 635, the extracted features of the current CNN are saved for all images in the current dataset. In step 645, a new CNN model, a production DL stack with the stored parameters of the trained DL stack, is deployed to infer from production images. In step 655, a batch of new labels is collected from the production OCR, and negative images that do not contain image-derived information are added. In step 665, new images are added to the training dataset for the CNN model to form new inputs. In step 675, after retraining the CNN model's classifier and validating and testing the model, the production DL stack is used to infer and classify at least one production image as containing sensitive image-derived identification documents.

[0088] For the use case of detecting screenshot images and preventing the loss of sensitive screenshot-derived data, the workflow is similar to workflow 600. To detect screenshot image scenarios, image generation robot 167 is a screenshot robot that collects examples of screenshot images and non-screenshot images and creates labeled ground truth data for use in training deep learning stack 157 without the need for OCR. The screenshot robot applies re-rendering to at least some of the collected screenshot image examples to represent screenshot changes that may contain sensitive information. Training data for training the DL stack through forward inference and backpropagation using the labeled ground truth data utilizes examples of screenshot images and non-screenshot images. In one example, a full screenshot image includes a single application window, with the window size covering more than 50% of the full screen. In another example, a full screenshot image shows multiple application windows, and in yet another example, an application screenshot image displays a single application window.

[0089] Figure 7 shows an example screenshot image containing a customer inventory list with costs listed. Screenshot image detection can prevent exfiltration of sensitive company data.

[0090] Cross-validation of the results obtained using the disclosed method for detecting screenshot images focuses on checking how well the DL stack model generalizes. The sample collection of screenshot and non-screenshot images was separated into a training set and a test set for screenshot images with MAC backgrounds. Images with Windows backgrounds and images with Linux backgrounds were used exclusively for testing. Furthermore, application windows were divided into training and test sets based on their categories. Next, we discuss the performance of five separate cross-validation cases. The combined training data was a synthetic set of full screenshots mixed with MAC background training and App window training.

[0091] For cross-validation case 1, the test data was a composite set of full screenshots mixed with tests on a MAC background and tests on an app window. Screenshot detection accuracy was measured at 93%. For cross-validation case 2, the test data was a composite set of full screenshots mixed with tests on a Windows background and tests on an app window. Screenshot detection accuracy was measured at 92%. For cross-validation case 3, the test data was a composite set of full screenshots mixed with tests on a Linux background and tests on an app window. Screenshot detection accuracy was measured at 86%. For cross-validation case 4, the test data was a composite set of full screenshots mixed with tests on a MAC background and tests on multiple app windows. Screenshot detection accuracy using these training and test data sets was measured at 97%. For cross-validation case 5, the test data tested a different app from the training app window, and accuracy was measured at 84%.

[0092] The performance of the deep learning stack model was tested for background and app window invisibility types. Next, a synthetic full screenshot with all background images and all app windows was used to train a classifier using 4,528 screenshots and 1,964 non-screenshot images. The classifier was tested using 45,179 images. The false negative rate (FNR) test using the 45,179 screenshots resulted in 90 images being classified as false negatives (FN) with an FNR of 0.2% at a threshold of 0.7. The false positive rate (FPR) test for 1,336 non-screenshot images resulted in four images being classified as false positives (FP) with an FPR of 0.374% at a threshold of 0.7. Four images in the test set were incorrectly classified as screenshots when they were non-screenshot images. The disclosed deep learning stack model's many layers work to capture features to determine "screenshots," including the following salient features: (1) Screenshots tend to include one or more main windows containing sensitive information. Such information can be personal information, code, text, pictures, etc. (2) Screenshots tend to include header / footer bars, such as menus or application bars. (3) Screenshots tend to have contrasting or uniform backgrounds compared to the content in the application window. For the four FP images, the main reasons for the classification of an image as a screenshot are as follows: Figure 8A shows a map of Idaho misclassified as a screenshot image due to the legend window and dotted lines above and below. Figure 8B shows a driver's license image misclassified as a screenshot image because the entire image is a window containing PII against a black background, and the UNITED STATES bar can be perceived as a header bar. Figure 8C shows a passport image as a primary window containing PII, with the shaded area at the bottom center potentially misleading the classifier into thinking it is an application bar. Figure 8D shows text in a primary window containing text information and a uniform background that was misclassified as a screenshot image.

[0093] In some use cases, separate organizations requiring DLP services can utilize a locally operating dedicated DL stack trainer 162 configured to combine irreversible features from examples of organization-sensitive data in images with ground-truth labels for the examples. The dedicated DL stack trainer forwards the irreversible features and ground-truth labels to a deep learning stack that receives organization-sensitive training examples, including the irreversible features and ground-truth labels, from the dedicated DL stack trainer 162. The organization-sensitive training examples are used to further train a second set of layers in the trained master DL stack. The updated parameters of the second set of layers for inference from production images can be stored and distributed to multiple separate organizations without compromising data security, since sensitive data is not accessible in the irreversible features.

[0094] Training of the deep learning stack 157 can start from scratch using training examples in a different order, or in another example, training can further train a second set of layers of the trained master DL stack using additional batches of labeled image examples.

[0095] In an added batch scenario, as samples are received back from a customer organization, the dedicated DL stack trainer can be configured to forward updated coefficients from the second set of layers. The deep learning stack 157 can receive updated coefficients from each of the second set of layers from multiple dedicated DL stack trainers and combine the updated coefficients from each of the second set of layers to train the second set of layers of the trained master DL stack. The deep learning stack 157 can then store the updated parameters of the second set of layers of the trained master DL stack for inferencing from production images and distribute the updated parameters of the second set of layers to separate customer organizations.

[0096] The dedicated DL stack trainer 162 may, in one example, handle training for detecting image-derived identified documents, and in another example, may handle training for detecting screenshot images.

[0097] Next, an embodiment of a computer system that can be used to detect identifying documents in images, detect screenshots, and prevent the loss of sensitive image-derived documents in the cloud is described. [Computer Systems]

[0098] FIG. 9 is a simplified block diagram of a computer system 900 that can be used to detect identified documents in images, referred to as image-derived identified documents, and prevent the loss of image-derived identified documents in the cloud. The computer system 900 can also be used to detect screenshot images and prevent the loss of sensitive screenshot-derived data. Furthermore, the computer system 900 can be used to customize a deep learning stack to detect organizationally sensitive data in images and prevent the loss of image-derived organizationally sensitive documents without requiring the transfer of potentially sensitive images to a centralized DLP service. The computer system 900 includes at least one central processing unit (CPU) 972 that communicates with several peripheral devices via a bus subsystem 955 and a Netscope Cloud Access Security Broker (N-CASB) 155 that provides the network security services described herein. These peripheral devices can include, for example, a storage subsystem 910, including a memory device and a file storage subsystem 936, a user interface input device 938, a user interface output device 976, and a network interface subsystem 974. The input and output devices allow user interaction with the computer system 900. The network interface subsystem 974 provides an interface to external networks, including interfaces to corresponding interface devices in other computer systems.

[0099] In one embodiment, the Netscope Cloud Access Security Broker (N-CASB) 155 of FIGS. 1A and 1B is communicatively linked to the storage subsystem 910 and the user interface input device 938.

[0100] User interface input devices 938 can include pointing devices such as keyboards, mice, trackballs, touchpads, or graphics tablets, scanners, touch screens integrated into displays, audio input devices such as voice recognition systems and microphones, and other types of input devices. In general, use of the term "input device" is intended to encompass all possible types of devices and methods for inputting information into computer system 900.

[0101] The user interface output devices 976 may include a display subsystem, a printer, a fax machine, or a non-visual display such as an audio output device. The display subsystem may include a flat panel device such as an LED display, a cathode ray tube (CRT), a liquid crystal display (LCD), a projection device, or some other mechanism for producing a visible image. The display subsystem may also provide a non-visual display such as an audio output device. In general, use of the term "output device" is intended to include all possible types of devices and methods for outputting information from the computer system 900 to a user or to another machine or computer system.

[0102] The storage subsystem 910 stores programming and data structures that provide the functionality of some or all of the modules and methods described herein. The subsystem 978 can be a graphics processing unit (GPU) or a programmable gate array (FPGA).

[0103] The memory subsystem 922 used in the storage subsystem 910 may include several memories, including a main random access memory (RAM) 932 for storing instructions and data during program execution and a read-only memory (ROM) 934 in which fixed instructions are stored. The file storage subsystem 936 may provide persistent storage for program and data files and may include a hard disk drive, a floppy disk drive with associated removable media, a CD-ROM drive, an optical drive, or a removable media cartridge. Modules that implement the functionality of particular embodiments may be stored by the file storage subsystem 936 within the storage subsystem 910 or in another machine accessible by the processor.

[0104] Bus subsystem 955 provides a mechanism for allowing the various components and subsystems of computer system 900 to communicate with each other as intended. Although bus subsystem 955 is shown schematically as a single bus, alternative implementations of the bus subsystem may use multiple buses.

[0105] Computer system 900 can be of various types, including a personal computer, a portable computer, a workstation, a computer terminal, a network computer, a television, a server, a mainframe, a widely distributed series of loosely coupled computers, or any other data processing system or user device. Due to the ever-changing nature of computers and networks, the description of computer system 900 shown in Figure 9 is intended only as a specific example for purposes of illustrating a preferred embodiment of the present invention. Many other configurations of computer system 900 are possible having more or fewer components than the computer system shown in Figure 9.

[0106] FIG. 10 illustrates a workflow 1000 for one or more computer systems that can be configured to detect screenshot images and prevent the loss of screenshot data. A computer performs a specific operation or action by installing software, firmware, hardware, or a combination thereof that causes the system to perform the action during operation. One or more computer programs can be configured to perform a specific operation or action by including instructions that, when executed by a data processing device, cause the device to perform the action. In some embodiments, multiple actions can be combined. For convenience, this flowchart is described with reference to a system including a Netskope Cloud Access Security Broker (N-CASB) and load balancing within a dynamic service chain while applying security services in the cloud. One general aspect includes a method for detecting screenshot images and preventing the loss of sensitive screenshot-derived data, including collecting 1010 instances of screenshot and non-screenshot images and creating labeled ground truth data for the instances. The method for detecting screenshot images also includes applying 1020 re-rendering of at least some of the collected screenshot image examples to represent various variations of the screenshot that may contain sensitive information. The method for detecting screenshot images also includes training 1030 a deep learning (DL) stack by forward inference and backpropagation using labeled ground truth data for examples of screenshot and non-screenshot images. The DL stack pre-trains a first set of layers closer to the input layer to perform image recognition before applying labeled ground truth data for screenshot and non-screenshot images to a second set of layers further from the input layer. The method for detecting screenshot images also includes storing 1040 parameters of the trained DL stack for inference from production images.The method for detecting screenshot images also includes using the production DL stack along with the stored parameters to classify, by inference, at least one production image as including a screenshot image 1050. Other embodiments of this aspect include corresponding computer systems, apparatus, and computer programs recorded on one or more computer storage devices, each configured to perform the operations of the method.

[0107] FIG. 11 illustrates a workflow 1100 for a system of one or more computers that can be configured to perform a particular operation or action by installing software, firmware, hardware, or a combination thereof that causes the system to perform an action during operation. One or more computer programs can be configured to perform a particular operation or action by including instructions that, when executed by a data processing device, cause the device to perform the action. In some implementations, multiple actions can be combined. For convenience, this flowchart is described with reference to a system including a NetScope Cloud Access Security Broker (N-CASB) and load balancing within a dynamic service chain while applying security services in the cloud. One general aspect includes a method for customizing a deep learning (DL) stack to detect organizational sensitive data in images, referred to as image-derived organizational sensitive documents, and prevent the loss of image-derived organizational sensitive documents, which includes pre-training a master DL stack using forward inference and backpropagation with labeled ground truth data for examples of image-derived sensitive documents and other image documents (1110). The DL stack includes at least a first set of layers closer to the input layer and a second set of layers further from the input layer, and further includes pre-training the second set of layers of the DL stack to perform image recognition before applying labeled ground truth data to examples of image-derived confidential documents and other image documents (1120). The method also includes storing parameters of the trained master DL stack for inference from production images (1130). The method also includes distributing the trained master DL stack with the stored parameters to multiple organizations (1140).The method further includes enabling the organization to perform update training of the trained master DL stack using at least examples of organizationally sensitive data in images and saving parameters of the updated DL stack (1150). The organization uses each updated DL stack to inferentially classify at least one production image as containing organizationally sensitive documents (1160). The method may also optionally include providing dedicated DL stack trainers, under the organization's control, to at least some of the organizations, and enabling the organization to perform update training using the dedicated DL stack trainers, which can be configured to generate each updated DL stack using examples of organizationally sensitive data in images without transferring the examples of organizationally sensitive data to the provider that performed the pre-training of the master DL stack (1170). Other embodiments of this aspect include corresponding computer systems, apparatuses, and computer programs recorded on one or more computer storage devices, each configured to perform the operations of the method. [Specific Implementation]

[0108] Some specific implementations and features for detecting identification documents in images and preventing loss of image-derived identification documents are described in the following discussion.

[0109] In one disclosed embodiment, a method for detecting identification documents in images, referred to as image-derived identification documents, and preventing loss of image-derived identification documents includes training a deep learning (DL) stack by forward inference and backpropagation using labeled ground truth data for the image-derived identification documents and other image document examples. The disclosed DL stack includes at least a first set of layers closer to an input layer and a second set of layers further from the input layer, and the second set of layers of the DL stack further includes the first set of layers pre-trained to perform image recognition before applying the labeled ground truth data for the image-derived identification documents and other image document examples. The disclosed method also includes storing parameters of the DL stack trained for inference from production images and classifying at least one production image as containing a sensitive image-derived identification document by inference using the production DL stack with the stored parameters.

[0110] The methods described in this and other sections of the disclosed technology may include one or more of the following features and / or features described in connection with the additional methods disclosed. For brevity, combinations of features disclosed in this application are not individually listed and are not repeated for each basic set of features. The reader will understand how features specified in the methods can be readily combined with the set of basic features specified as an embodiment.

[0111] Some disclosed embodiments of the method optionally include capturing features generated as output from the first set of layers for the private image-derived identification document and retaining the captured features along with their respective ground truth labels, thereby eliminating the need to retain images of the private image-derived identification document.

[0112] Some embodiments of the disclosed method include limiting training of parameters in the second set of layers to backward propagation using labeled ground truth data for image-derived identification documents and other image document examples.

[0113] In one disclosed embodiment of the present invention, optical character recognition (OCR) analysis of an image is applied to label the image as an identified or non-identified document. After the OCR analysis, a reliable classification can be selected for use in the training set. OCR and regular expression matching serve as an automated method for generating labeled data from customer production images. In one example, for a U.S. passport, OCR first extracts the text on the passport page. Then, a regular expression may match "PASSPORT," "UNITED STATES," "Department of State," "USA," "Authority," and other words on the page. As a second example, for a California driver's license, OCR first extracts the text from the front of the driver's license. Then, a regular expression may match "California," "USA," "DRIVER LICENSE," "CLASS," "SEX," "HAIR," "EYES," and other words on the front page. As a third example, for a Canadian passport, OCR first extracts the text on the passport page. A regular expression could then match "PASSPORT," "PASSEPORT," "CANADA," and other words on the page.

[0114] In some disclosed embodiments of the present invention, when training a DL stack by backpropagation, the perspective of a first set of image-derived identification documents is distorted to generate a second set of image-derived identification documents, and the first and second sets are combined with labeled ground truth data.

[0115] In another disclosed embodiment of the method, when training a DL stack by backpropagation, a first set of image-derived discriminative documents is distorted by rotation to generate a third set of image-derived discriminative documents, and the first and third sets are combined with labeled ground truth data.

[0116] In one disclosed embodiment of the present invention, when training a DL stack by backpropagation, a first set of image-derived discriminative documents is distorted with noise to generate a fourth set of image-derived discriminative documents, and the first and fourth sets are combined with labeled ground truth data.

[0117] In some disclosed embodiments of the present invention, when training a DL stack by backpropagation, the focus of the first set of image-derived discriminative documents is distorted to generate a fifth set of image-derived discriminative documents, and the first and fifth sets are combined with labeled ground truth data.

[0118] In some embodiments, the disclosed method includes storing the lossy DL features of the current training ground truth images rather than the original ground truth images to avoid storing sensitive personal information, periodically adding the lossy DL features of new ground truth images to augment the training set, and periodically retraining on the augmented training data set for greater accuracy. The lossy DL features cannot be transformed into images with recognizable sensitive data.

[0119] Some specific implementations and features for detecting screenshot images and preventing the loss of sensitive screenshot-derived data are described in the following discussion.

[0120] In one disclosed embodiment, a method for detecting screenshot images and preventing loss of sensitive screenshot-derived data includes collecting examples of screenshot images and non-screenshot images and creating labeled ground truth data for the examples. The method also includes re-rendering at least some of the collected example screenshot images to represent various variations of screenshots that may contain sensitive information, and training a DL stack by forward inference and backpropagation using the labeled ground truth data for the example screenshot images and non-screenshot images. The method further includes storing parameters of the DL stack trained for inference from production images and classifying at least one production image by inference as containing a sensitive image-derived screenshot using the production DL stack with the stored parameters.

[0121] Some embodiments of the disclosed method further include applying a screenshot robot to collect instances of screenshot images and non-screenshot images.

[0122] In one embodiment of the disclosed method, the DL stack includes at least a first set of layers closer to the input layer and a second set of layers further from the input layer, and further, the first set of layers is pre-trained to perform image recognition before applying labeled ground truth data for examples of screenshot images and non-screenshot images to the second set of layers of the DL stack.

[0123] Some implementations of the disclosed method include applying an automatic re-rendering of at least a portion of the original captured screenshot image by cropping a portion of the image or adjusting the hue, contrast, and saturation to represent changes in the screenshot, where in some cases the changes in the screenshot include at least one of window size, window position, number of open windows, and menu bar position.

[0124] In one embodiment of the disclosed method, when training the DL stack by backpropagation, the first set of screenshot images are surrounded by boundaries of various photographic images from the two or more sensitive image-derived screenshots to generate a third set of screenshot images, and the first and third sets are combined with labeled ground truth data. In another embodiment, when training the DL stack by backpropagation, the first set of screenshot images are surrounded by boundaries of a plurality of overlaid program windows from the two or more sensitive image-derived screenshots to generate a fourth set of screenshot images, and the first and fourth sets are combined with labeled ground truth data.

[0125] The following discussion describes several specific implementations and features for detecting tissue-sensitive screenshot images and preventing the loss of tissue-sensitive screenshot images.

[0126] In one disclosed embodiment, a method for customizing a deep learning stack to detect organizational sensitive data in images, referred to as image-derived organizational sensitive documents, and preventing the loss of image-derived organizational sensitive documents includes pre-training a master DL stack by forward inference and backpropagation using labeled ground truth data for examples of the image-derived sensitive documents and other image documents. The DL stack includes at least a first set of layers closer to an input layer and a second set of layers further from the input layer, and further includes pre-training the first set of layers to perform image recognition before applying the labeled ground truth data for examples of the image-derived sensitive documents and other image documents to the second set of layers of the DL stack. The disclosed method also includes storing parameters of the master DL stack trained for inference from production images, distributing the trained master DL stack with the stored parameters to multiple organizations, and allowing the organizations to perform updated training of the trained master DL stack using at least examples of the organizational sensitive data in images and saving the parameters of the updated DL stack. The organization uses each updated DL stack to inferentially classify at least one production image as containing organizationally sensitive documents.

[0127] Training of the deep learning stack can start from scratch in some cases, while in other embodiments, training can further train a second set of layers of the trained master DL stack using additional batches of labeled image examples utilized with the previously determined coefficients. Some embodiments of the disclosed method further include providing a dedicated DL stack trainer to at least a portion of the organization, under the organization's control, and enabling the organization to perform updated training without transferring examples of organization-sensitive data in images to the provider that performed the pre-training of the master DL stack. The dedicated DL stack trainer can be configured to generate each updated DL stack. In some cases, the method also includes a dedicated DL stack trainer configured to combine lossy features from examples of organization-sensitive data in images with ground truth labels for the examples and transfer the lossy features and ground truth labels, and receive organization-sensitive training examples including the lossy features and ground truth labels from the multiple dedicated DL stack trainers. In some embodiments, the disclosed method also includes further training a second set of layers of the trained master DL stack using tissue-confidential training examples, storing updated parameters for the second set of layers for inference from production images, and distributing the updated parameters for the second set of layers to multiple tissues. Some embodiments further include performing update training to further train the second set of layers of the trained master DL stack. In other cases, the method includes performing training from scratch using tissue-confidential training examples in a different order to further train the second set of layers of the trained master DL stack. In one embodiment, the disclosed method further includes a dedicated DL stack trainer configured to transfer updated coefficients from the second set of layers, receiving each updated coefficient from each of the second set of layers from the multiple dedicated DL stack trainers, and combining the updated coefficients from each of the second set of layers to train the second set of layers of the trained master DL stack.The disclosed method also includes storing updated parameters of the second set of layers of the trained master DL stack for inference from production images and distributing the updated parameters of the second set of layers to multiple organizations.

[0128] Other embodiments of the disclosed technology described in this section can include a tangible, non-transitory computer-readable storage medium containing program instructions loaded into a memory that, when executed on a processor, causes the processor to perform any of the methods described above. Yet another embodiment of the disclosed technology described in this section can include a system including a memory and one or more processors operable to execute computer instructions stored in the memory to perform any of the methods described above.

[0129] The foregoing description is presented to enable use and practice of the disclosed technology. Various modifications to the disclosed embodiments will be apparent, and the generic principles defined herein may be applied to other embodiments and applications without departing from the spirit and scope of the disclosed technology. Thus, the disclosed technology is not intended to be limited to the embodiments shown, but is to be accorded the widest scope consistent with the principles and features disclosed herein. The scope of the disclosed technology is defined by the appended claims. [Provisions]

[0130] Techniques are described for detecting identifying documents in images and preventing the loss of image-derived identifying documents.

[0131] The disclosed technology can be implemented as a system, method, device, or product. One or more features of an embodiment can be combined with the base embodiment. Non-mutually exclusive embodiments are taught as combinable. One or more features of an embodiment can be combined with other embodiments. The present disclosure periodically reminds the user of these options. The omission from some embodiments of statements repeating these options should not be construed as limiting the combinations taught in the preceding sections. These statements are incorporated by reference in consideration of each of the following embodiments.

[0132] One or more embodiments and clauses of the disclosed technology, or elements thereof, may be implemented in the form of a computer product including a non-transitory computer-readable storage medium having computer-usable program code for performing the method steps described herein. Furthermore, one or more embodiments and clauses of the disclosed technology, or elements thereof, may be implemented in the form of an apparatus including a memory and at least one processor coupled to the memory and operable to perform the exemplary method steps. In yet another aspect, one or more embodiments and clauses of the disclosed technology, or elements thereof, may be implemented in the form of a means for performing one or more of the method steps described herein. The means may include (i) hardware modules, (ii) software modules executing on one or more hardware processors, or (iii) a combination of hardware and software modules, any of which (i) through (iii) implement specific techniques described herein, and the software modules are stored on a computer-readable storage medium (or multiple such media).

[0133] The clauses described in this section can be combined as features. For brevity, combinations of features are not individually listed and are not repeated for each basic set of features. The reader will understand how features identified in clauses described in this section can be readily combined with sets of basic features identified as embodiments in other sections of this application. These clauses are not meant to be mutually exclusive, exhaustive, or limiting, and the disclosed technology is not limited to these clauses, but rather encompasses all possible combinations, modifications, and variations within the scope of the claimed technology and its equivalents.

[0134] Another implementation of the provisions described in this section may include a non-transitory computer-readable storage medium storing instructions executable by a processor to perform any of the provisions described in this section. Yet another implementation of the provisions described in this section may include a system including a memory and one or more processors operable to execute instructions stored in the memory to perform any of the provisions described in this section.

[0135] Disclose the following clauses: [Clause Set 1] 1. A method for detecting an identification document in an image, referred to as an image-derived identification document, and preventing loss of said image-derived identification document, comprising: training a deep learning (abbreviated as DL) stack by forward inference and backpropagation using labeled ground truth data for the image-derived identified documents and other image document examples; wherein the DL stack includes at least a first set of layers closer to an input layer and a second set of layers farther from the input layer, and the first set of layers of the second set of layers of the DL stack are pre-trained to perform image recognition before applying the labeled ground truth data to the image-derived identified document and other image document examples; storing parameters of the trained DL stack for inference from production images; and classifying at least one production image by inference as containing a sensitive image-derived identification document using a production DL stack having the stored parameters. 2. The method of clause 1, further comprising capturing features generated as output from the first set of layers for a private image-derived identification document, and retaining the captured features together with their respective ground truth labels, thereby eliminating the need to retain images of the private image-derived identification document. 3. The method of any one of clauses 1 to 2, further comprising limiting training by backpropagation using the labeled ground truth data for the examples of the image-derived identified documents and other image documents to training of parameters in the second set of layers. 4. The method according to any one of clauses 1 to 3, wherein an optical character recognition (abbreviated as OCR) analysis of the image is applied to label said image as an identified document or a non-identified document. 5. The method of any one of clauses 1 to 4, wherein when training the DL stack by backpropagation, the perspective of the first set of image-derived identified documents is distorted to generate a second set of image-derived identified documents, and the first set and the second set are combined with the labeled ground truth data. 6. The method of any one of clauses 1 to 5, wherein when training the DL stack by backpropagation, the first set of image-derived identification documents is distorted by noise to generate a third set of image-derived identification documents, and the first set and the third set are combined with the labeled ground truth data. 7. The method of any one of clauses 1 to 6, wherein when training the DL stack by backpropagation, the focus of the first set of image-derived identification documents is distorted to generate a fourth set of image-derived identification documents, and the first set and the fourth set are combined with the labeled ground truth data. 8. A tangible, non-transitory computer-readable storage medium containing program instructions loaded into memory that, when executed on a processor, causes said processor to perform a method for detecting an identification document in an image, referred to as an image-derived identification document, and preventing loss of said image-derived identification document, said method comprising: training a deep learning (abbreviated as DL) stack by forward inference and backpropagation using labeled ground truth data for the image-derived identified documents and other image document examples; wherein the DL stack includes at least a first set of layers closer to an input layer and a second set of layers farther from the input layer, and the first set of layers is pre-trained to perform image recognition before applying the labeled ground truth data of the image-derived identified document and the other image document examples to the second set of layers of the DL stack; Storing parameters of the trained DL stack for inference from production images, and classifying at least one production image by inference as containing a sensitive image-derived identification document using a production DL stack having the stored parameters. 9. The tangible, non-transitory computer-readable storage medium of clause 8, further comprising capturing features generated as output from the first set of layers for a private image-derived identification document, and retaining the captured features along with their respective ground truth labels, thereby eliminating the need to retain images of the private image-derived identification document. 10. A tangible, non-transitory computer-readable storage medium as described in any one of clauses 8 to 9, further comprising limiting training by backpropagation using the labeled ground truth data for the examples of the image-derived identification document and other image documents to training of parameters in the second set of layers. 11. A tangible, non-transitory computer-readable storage medium according to any one of clauses 8 to 10, which applies optical character recognition (abbreviated as OCR) analysis of an image to label said image as an identified document or a non-identified document. 12. A tangible, non-transitory computer-readable storage medium as described in any one of clauses 8 to 11, wherein when training the DL stack by backpropagation, the perspective of the first set of image-derived identified documents is distorted to generate a second set of image-derived identified documents, and the first set and the second set are combined with the labeled ground truth data. 13. A tangible, non-transitory computer-readable storage medium according to any one of clauses 8 to 12, wherein when training the DL stack by backpropagation, the first set of image-derived identification documents is distorted by noise to generate a third set of image-derived identification documents, and the first set and the third set are combined with the labeled ground truth data. 14. A system for detecting an identification document in an image, referred to as an image-derived identification document, and preventing loss of said image-derived identification document, the system comprising: a processor; a memory connected to said processor; and computer instructions loaded into said memory from a non-transitory computer-readable storage medium as described in clause 8. 15. The system described in clause 14, further comprising capturing features generated as output from the first set of layers for private image-derived identification documents and retaining the captured features along with their respective ground truth labels, thereby eliminating the need to retain images of the private image-derived identification documents. 16. The system of any one of clauses 14-15, further comprising limiting backpropagation training using the labeled ground truth data for the examples of the image-derived identification document and other image documents to training of parameters in the second set of layers. 17. A system according to any one of clauses 14 to 16, which applies an optical character recognition (abbreviated as OCR) analysis of the image to label the image as an identified or non-identified document. 18. The system of any one of clauses 14 to 17, wherein when training the DL stack by backpropagation, the perspective of the first set of image-derived identified documents is distorted to generate a second set of image-derived identified documents, and the first set and the second set are combined with the labeled ground truth data. 19. The system of any one of clauses 14 to 18, wherein when training the DL stack by backpropagation, the first set of image-derived identification documents is distorted by noise to generate a third set of image-derived identification documents, and the first set and the third set are combined with the labeled ground truth data. 20. The system of any one of clauses 14 to 19, wherein when training the DL stack by backpropagation, the focus of the first set of image-derived identification documents is distorted to generate a fourth set of image-derived identification documents, and the first set and the fourth set are combined with the labeled ground truth data. [Clause Set 2] 1. A method for detecting screenshot images and preventing loss of sensitive screenshot-derived data, comprising: collecting examples of the screenshot images and non-screenshot images and generating labeled ground truth data for the examples; applying re-rendering of at least some of said collected example screenshot images to represent various variations of the screenshots that may contain sensitive information; training a deep learning (abbreviated as DL) stack by forward inference and backpropagation using the labeled ground truth data for the screenshot and non-screenshot image examples; storing parameters of the trained DL stack for inference from production images; and classifying at least one production image by inference as comprising a screenshot image using a production DL stack having the stored parameters. 2. The method of clause 1, further comprising applying a screenshot robot to collect said instances of said screenshot images and non-screenshot images. 3. The method of any one of clauses 1 to 2, wherein the DL stack includes at least a first set of layers closer to an input layer and a second set of layers further from the input layer, and further comprising pre-training the first set of layers of the DL stack to perform image recognition before applying the labeled ground truth data to the examples of the screenshot images and the non-screenshot images. 4. The method of any one of clauses 1 to 3, further comprising applying automatic re-rendering of at least a portion of the collected original screenshot image by cropping a portion of the image or by adjusting hue, contrast, and saturation to represent changes in the screenshot. 5. The method of any one of clauses 1 to 4, wherein the various changes to the screenshot include at least one of window size, window position, number of open windows, and menu bar position. 6. The method of any one of clauses 1 to 5, wherein when training the DL stack by backpropagation, the first set of screenshot images are surrounded by boundaries of various photographic images of a plurality of image-derived screenshots to generate a third set of screenshot images, and the first and third sets are combined with the labeled ground truth data. 7. The method of any one of clauses 1 to 6, wherein when training the DL stack by backpropagation, the first set of screenshot images is surrounded by boundaries of a plurality of overlaid program windows of a plurality of image-derived screenshots to generate a fourth set of screenshot images, and the first and fourth sets are combined with the labeled ground truth data. 8. A tangible, non-transitory computer-readable storage medium comprising program instructions loaded into memory that, when executed on a processor, cause the processor to perform a method for detecting screenshot images and preventing loss of image-derived screenshots, the method comprising: collecting examples of said screenshot images and non-screenshot images and generating labeled ground truth data for said examples; applying re-rendering of at least some of said collected example screenshot images to represent various variations of the screenshots that may contain sensitive information; training a deep learning (abbreviated as DL) stack by forward inference and backpropagation using labeled ground truth data for the examples of the screenshot images and the non-screenshot images; storing parameters of the trained DL stack for inference from production images; and classifying, by inference, at least one production image as including an image-derived screenshot using a production DL stack having the stored parameters. 9. The tangible, non-transitory computer-readable storage medium of clause 8, further comprising applying a screenshot robot to collect said instances of said screenshot images and non-screenshot images. 10. The tangible, non-transitory computer-readable storage medium of any one of clauses 8 to 9, wherein the DL stack includes at least a first set of layers closer to an input layer and a second set of layers farther from the input layer, and further comprising pre-training the first set of layers to perform image recognition before applying the labeled ground truth data to the examples of the screenshot images and the non-screenshot images to the second set of layers of the DL stack. 11. The tangible, non-transitory computer-readable storage medium of any one of clauses 8 to 10, further comprising applying automatic re-rendering of at least a portion of the collected original screenshot image by cropping a portion of the image or by adjusting hue, contrast, and saturation to represent changes in the screenshot. 12. The tangible, non-transitory computer-readable storage medium of any one of clauses 8 to 11, wherein the various changes to the screenshot include at least one of window size, window position, number of open windows, and menu bar position. 13. The tangible, non-transitory computer-readable storage medium of any one of clauses 8 to 12, wherein when training the DL stack by backpropagation, the first set of screenshot images are surrounded by boundaries of various photographic images of a plurality of image-derived screenshots to generate a third set of screenshot images, and the first and third sets are combined with the labeled ground truth data. 14. The tangible, non-transitory computer-readable storage medium of any one of clauses 8 to 13, wherein when training the DL stack by backpropagation, the first set of screenshot images is surrounded by boundaries of a plurality of overlaid program windows of a plurality of image-derived screenshots to generate a fourth set of screenshot images, and the first and fourth sets are combined with the labeled ground truth data. 15. A system for detecting screenshot images and preventing loss of image-derived screenshots, the system including a processor, a memory coupled to the processor, and computer instructions loaded into the memory from a non-transitory computer-readable storage medium as described in clause 8. 16. The system of clause 15, further comprising applying a screenshot robot to collect the instances of the screenshot images and non-screenshot images. 17. The system of any one of clauses 15-16, wherein the DL stack includes at least a first set of layers closer to an input layer and a second set of layers farther from the input layer, and further including pre-training the first set of layers to perform image recognition before applying the labeled ground truth data to the examples of the screenshot images and the non-screenshot images to the second set of layers of the DL stack. 18. The system of any one of clauses 15 to 17, further comprising applying automatic re-rendering of at least a portion of the collected original screenshot image by cropping a portion of the image or by adjusting hue, contrast, and saturation to represent changes in the screenshot. 19. The system of any one of clauses 15 to 18, wherein the various changes to the screenshot include at least one of window size, window position, number of open windows, and menu bar position. 20. The system of any one of clauses 15 to 19, wherein when training the DL stack by backpropagation, the first set of screenshot images are surrounded by boundaries of various photographic images of a plurality of image-derived screenshots to generate a third set of screenshot images, and the first and third sets are combined with the labeled ground truth data. [Clause Set 3] 1. A method for detecting organizationally sensitive data in images, referred to as image-derived organizationally sensitive documents, and preventing loss of said image-derived organizationally sensitive documents, by customizing a deep learning (abbreviated as DL) stack, comprising: Pre-training a master DL stack using forward inference and backpropagation with labeled ground truth data for image-derived classified documents and other image document examples; wherein the DL stack includes at least a first set of layers closer to an input layer and a second set of layers farther from the input layer, and the method further includes pre-training the first set of layers in the second set of layers of the DL stack to perform image recognition before applying the labeled ground truth data to the image-derived confidential documents and other image document examples; storing parameters of the trained master DL stack for inferencing from production images; distributing the trained master DL stack with stored parameters to a plurality of organizations; and allowing the organization to perform update training of the trained master DL stack using at least examples of the organization sensitive data in images and save parameters of the updated DL stack, whereby the organization uses each updated DL stack to classify, by inference, at least one production image as containing organization sensitive documents. 2. The method of clause 1, comprising providing at least a portion of the organization, under organizational control, with a dedicated DL stack trainer to enable the organization to perform the updated training without transferring instances of the organization sensitive data in images to a provider that performed pre-training of the master DL stack, wherein the dedicated DL stack trainer is configurable to generate each updated DL stack. 3. the dedicated DL stack trainer is configured to combine lossy features from the examples of the tissue sensitive data in images with ground truth labels for the examples and forward the lossy features and ground truth labels; and 3. The method of claim 2, further comprising receiving tissue-sensitive training examples comprising the lossy features and ground truth labels from a plurality of the dedicated DL stack trainers. 4. using the tissue-confidential training examples to further train the second set of layers of the trained master DL stack; storing the updated parameters of the second set of layers for inference from production images; and 4. The method of clause 3, comprising distributing the updated parameters of the second set of layers to a plurality of organizations. 5. The method of any one of clauses 1 to 4, further comprising performing update training to further train the second set of layers of the trained master DL stack. 6. The method of any one of clauses 1 to 5, further comprising training from scratch using the organizationally sensitive training examples in a different order to further train the second set of layers of the trained master DL stack. 7. the dedicated DL stack trainer is configured to forward updated coefficients from the second set of layers; and receiving updated coefficients from each of the second set layers from the plurality of dedicated DL stack trainers; combining the updated coefficients from each second set of layers to train the second set of layers of the trained master DL stack; 5. The method of any one of clauses 2 to 4, comprising: storing updated parameters of the second set of layers of the trained master DL stack for inferencing from production images; and distributing the updated parameters of the second set of layers to the multiple organizations. 8. A tangible, non-transitory computer-readable storage medium containing program instructions loaded into a memory that, when executed on a processor, causes the processor to perform a method for customizing a deep learning (abbreviated as DL) stack to detect organizationally sensitive data in images, referred to as image-derived organizationally sensitive documents, and to prevent loss of said image-derived organizationally sensitive documents, said method comprising: Pre-training a master DL stack using forward inference and backpropagation with labeled ground truth data for image-derived classified documents and other image document examples; wherein the DL stack includes at least a first set of layers closer to an input layer and a second set of layers farther from the input layer, and the first set of layers of the second set of layers of the DL stack are pre-trained to perform image recognition before applying the labeled ground truth data to the image-derived confidential documents and other image document examples; storing parameters of the trained master DL stack for inferencing from production images; distributing the trained master DL stack with the stored parameters to a plurality of organizations; and enabling the organization to perform update training of the trained master DL stack using at least instances of the organization-sensitive data in images and save parameters of the updated DL stack, whereby the organization uses each updated DL stack to inferentially classify at least one production image as containing organization-sensitive documents. 9. The tangible, non-transitory computer-readable storage medium of clause 8, comprising providing at least a portion of the organization, under the control of the organization, a dedicated DL stack trainer, and enabling the organization to perform the updated training without transferring instances of the organization-sensitive data in images to a provider that performed pre-training of a master DL stack, wherein the dedicated DL stack trainer is configurable to generate the respective updated DL stack. 10. The tangible, non-transitory computer-readable storage medium of clause 9, further comprising: the dedicated DL stack trainer configured to combine lossy features from the examples of the tissue-sensitive data in images with ground-truth labels for the examples and forward the lossy features and ground-truth labels; and receiving tissue-sensitive training examples including the lossy features and ground-truth labels from a plurality of the dedicated DL stack trainers. 11. using the tissue-confidential training examples to further train the second set of layers of the trained master DL stack; 11. The tangible, non-transitory computer-readable storage medium of clause 10, further comprising: storing the updated parameters of the second set of layers for inference from production images; and distributing the updated parameters of the second set of layers to multiple organizations. 12. The tangible, non-transitory computer-readable storage medium of any one of clauses 8 to 11, further comprising performing update training to further train the second set of layers of the trained master DL stack. 13. The tangible, non-transitory computer-readable storage medium of any one of clauses 8 to 12, further comprising: performing training from scratch using the organizationally sensitive training examples in a different order to further train the second set of layers of the trained master DL stack. 14. The dedicated DL stack trainer is configured to forward updated coefficients from the second set of layers; and receiving updated coefficients from each of the second set of layers from a plurality of the dedicated DL stack trainers; combining the updated coefficients from each second set of layers to train the second set of layers of the trained master DL stack; 12. The tangible, non-transitory computer-readable storage medium of any one of clauses 9 to 11, comprising: storing updated parameters of the second set of layers of the trained master DL stack for inferencing from production images; and distributing the updated parameters of the second set of layers to the multiple organizations. 15. A system for customizing a deep learning (abbreviated as DL) stack to detect organizationally sensitive data in images, referred to as image-derived organizationally sensitive documents, and preventing loss of said image-derived organizationally sensitive documents, the system comprising: a processor; a memory connected to said processor; and computer instructions loaded into said memory from a non-transitory computer-readable storage medium as described in clause 8. 16. The system of clause 15, further comprising providing at least a portion of the organization, under its control, with a dedicated DL stack trainer to enable the organization to perform the updated training without transferring instances of the organization sensitive data in images to a provider that performed pre-training of the master DL stack, the dedicated DL stack trainer being configurable to generate each updated DL stack. 17. The system of clause 16, further comprising: the dedicated DL stack trainer configured to combine lossy features from the examples of the tissue-sensitive data in images with ground-truth labels for the examples and forward the lossy features and ground-truth labels; and receiving tissue-sensitive training examples including the lossy features and ground-truth labels from a plurality of the dedicated DL stack trainers. 18. using the tissue-confidential training examples to further train the second set of layers of the trained master DL stack; 18. The system of clause 17, further comprising: storing the updated parameters of the second set of layers for inference from production images; and distributing the updated parameters of the second set of layers to multiple tissues. 19. The system of any one of clauses 15 to 18, further comprising performing update training to further train the second set of layers of the trained master DL stack. 20. The system of any one of clauses 15-19, further comprising: performing training from scratch using the organizationally sensitive training examples in a different order to further train the second set of layers of the trained master DL stack. 21. The dedicated DL stack trainer is configured to forward updated coefficients from the second set of layers; and receiving updated coefficients from each of the second set of layers from a plurality of the dedicated DL stack trainers; combining the updated coefficients from each second set of layers to train the second set of layers of the trained master DL stack; 19. The system of any one of clauses 16 to 18, further comprising: storing updated parameters of the second set of layers of the trained master DL stack for inferencing from production images; and distributing the updated parameters of the second set of layers to the multiple organizations.

Claims

1. 1. A method for detecting an identification document in an image, referred to as an image-derived identification document, and preventing loss of said image-derived identification document, comprising: training a deep learning (abbreviated as DL) stack by forward inference and back propagation using labeled ground truth data for the image-derived identification document and other image document examples, without performing optical character recognition of words appearing on the image-derived identification document or other image document examples; wherein the DL stack includes at least a first set of layers closer to an input layer and a second set of layers farther from the input layer, and the first set of layers of the second set of layers of the DL stack are pre-trained to perform image recognition before applying the labeled ground truth data to the image-derived identified document and other image document examples; storing parameters of the trained DL stack for inference from production images; and classifying at least one production image by inference as containing a sensitive image-derived identification document using a production DL stack having the stored parameters.

2. 2. The method of claim 1, further comprising capturing features generated as output from the first set of layers for a private image-derived identified document, and retaining the captured features along with their respective ground truth labels, thereby eliminating the need to retain images of the private image-derived identified document.

3. The method of any one of claims 1 to 2, further comprising: limiting training by backpropagation using the labeled ground truth data for the examples of the image-derived identification document and other image documents to training of parameters in the second set of layers.

4. The method according to any one of claims 1 to 3, wherein an optical character recognition (abbreviated as OCR) analysis of the image is applied to label the image as an identified or non-identified document.

5. 5. The method of claim 1, wherein when training the DL stack by backpropagation, the perspective of the first set of image-derived identification documents is distorted to generate a second set of image-derived identification documents, and the first and second sets are combined with the labeled ground truth data.

6. 6. The method of claim 1, wherein when training the DL stack by backpropagation, the first set of image-derived identification documents is distorted by noise to generate a third set of image-derived identification documents, and the first and third sets are combined with the labeled ground truth data.

7. 7. The method of claim 1, wherein when training the DL stack by backpropagation, the focus of the first set of image-derived identification documents is distorted to generate a fourth set of image-derived identification documents, and the first and fourth sets are combined with the labeled ground truth data.

8. A tangible, non-transitory computer-readable storage medium comprising program instructions loaded into a memory which, when executed on a processor, causes the processor to perform the method of any one of claims 1 to 7, for detecting an identification document in an image, referred to as an image-derived identification document, and preventing loss of the image-derived identification document.

9. A system for detecting an identification document in an image, called an image-derived identification document, and preventing loss of said image-derived identification document, comprising a processor, a memory connected to said processor, and program instructions loaded into said memory that, when executed on said processor, cause said processor to perform a method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Method for training classifier designed for determining category of document

    CN109583463A

  • System, method, and computer program product for preventing image-related data loss

    US20120183174A1