Risk assessment techniques for controlling access to computing systems based on event analysis

The risk assessment system automates the analysis of data sources to generate accurate risk indicators for computing environment access, addressing inefficiencies and inaccuracies in manual methods.

WO2025159743A1PCT designated stage expired Publication Date: 2025-07-31EQUIFAX INC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/US2024/012600
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-01-23
Publication Date
2025-07-31

AI Technical Summary

Technical Problem

Existing systems for controlling access to computing environments rely on manual examination of data sources, leading to inefficiencies, inaccuracies, and limited scalability in risk assessment for entities involved in events.

Method used

A risk assessment system that analyzes data sources to identify entities, determine sentiment scores, and generate risk indicators based on event involvement, using automated sentiment analysis and machine learning to improve accuracy and scalability.

Benefits of technology

Enhances the accuracy and scalability of risk assessment by automating the analysis of large data volumes, providing flexible and robust control over access to computing environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2024012600_31072025_PF_FP_ABST
    Figure US2024012600_31072025_PF_FP_ABST
Patent Text Reader

Abstract

A system can generate a risk assessment associated with a target entity. For example, the system can receive a request for a risk indicator associated with a target entity. The system can determine that a data source contains a name associated with the target entity based on extracted text from the data source. The system can identify a sentence containing the name. The system can further determine a sentiment score for the sentence. The system can generate a classification associated with an event included in the sentence. The system can determine a confidence score that the name in the extracted text is associated with the target entity based on attributes associated with the name in the extracted text. The system can transmit, to a remote computing device, a message including the risk indicator based on the classification or sentiment score.
Need to check novelty before this filing date? Find Prior Art

Description

Attorney Docket No.096923-1414832 RISK ASSESSMENT TECHNIQUES FOR CONTROLLING ACCESS TO COMPUTING SYSTEMS BASED ON EVENT ANALYSIS TECHNICAL FIELD

[0001] The present disclosure relates generally to controlling interactions between computing systems. More specifically, but not by way of limitation, this disclosure relates to risk assessment based identification and analysis of events described in data sources for controlling interactions between computing systems. BACKGROUND

[0002] Various systems may use an entity’s involvement in certain events (e.g., crimes, research projects, joint ventures, etc.) to control access to restricted data or restricted computing environments. There are millions of data sources (e.g., news articles and publications) accessible via the Internet that contain information about entities and events in which the entities are involved. To parse these data sources for use in access control, systems rely on manual examination of the data sources or of summaries of the data sources, leading to inefficiencies, inaccuracies, high costs, and limitations in scalability. SUMMARY

[0003] Various aspects of the present disclosure provide systems and methods for risk assessment using a risk indicator. The system can receive a request for a risk indicator associated with a target entity. In some aspects, the system can determine that a data source contains a name associated with the target entity based on extracted text from the data source. The system can also divide the extracted text into sentences. The system can identify a sentence of the at least one sentence that includes the name associated with the target entity. The system can determine a sentiment score associated with the sentence. The system can also generate a classification associated with an event included in the sentence. In some aspects, the system can generate the risk indicator based on at least one of the classification or the sentiment score. In some aspects, the system can determine a confidence score that the name in the extracted text is associated with the target entity, where the confidence score is determined based on attributes associated with the US2008230018841Attorney Docket No.096923-1398661 name in the extracted text. Upon determining that the confidence score is above a confidence threshold, the system can transmit, to a remote computing device, a responsive message comprising at least the risk indicator used to control access of the target entity to one or more interactive computing environments.

[0004] This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to be used in isolation to determine the scope of the claimed subject matter. The subject matter should be understood by reference to appropriate portions of the entire specification, any or all drawings, and each claim.

[0005] The foregoing, together with other features and examples, will become more apparent upon referring to the following specification, claims, and accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0006] FIG. 1 is a block diagram depicting an example of an operating environment in which a risk assessment computing system can be used to provide a risk assessment associated with a target entity according to some aspects of the present disclosure.

[0007] FIG. 2 is a block diagram depicting an example of a risk assessment application for generating a risk assessment associated with a target entity according to some aspects of the present disclosure.

[0008] FIG.3 is a flow chart illustrating a method for generating a risk assessment associated with a target entity according to some aspects of the present disclosure.

[0009] FIG. 4 is a block diagram depicting an example of a computing device, which can be used to implement the embodiments described herein according to some aspects of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION

[0010] Certain aspects and examples of the present disclosure relate to risk assessment of a target entity based on an association of the target entity to one or more events as determined from one or more data sources. A target entity can be, for example, an individual, an organization, or a system. In some examples, the target entity can request access to a secured system or resource, US2008230018841Attorney Docket No.096923-1398661 such as a database. A risk assessment system can analyze one or more data sources (e.g., documents or publications) to determine an amount of risk associated with the target entity. The risk assessment can be used in determining whether to grant or deny the target entity access to the secure resource. For example, a data source, such as a news article, can provide details about an entity’s involvement in an event (e.g., an initial public offering (IPO), a government project, a crime, a public relations (PR) scandal, and the like). The type of event and the extent and nature of the entity’s involvement in the event can provide insights upon which to base a risk assessment of the entity.

[0011] Certain aspects described herein for performing risk assessments on target entities using event information extracted from one or more data sources can address one or more issues. For example, in certain aspects, disclosed systems and methods improve scalability and automation of data source review by analyzing large numbers of documents (i.e., data sources), which is both fast and less error-prone than manual examination of data sources. Further, disclosed systems and methods integrate various analytical functions, such as entity recognition, sentiment analysis, and event detection to provide a multifaceted risk assessment of an entity. Disclosed systems and methods are also flexible and have the ability to analyze large amounts of data sources in various formats.

[0012] In some examples, a risk assessment computing system can receive a request for a risk indicator associated with a target entity. The risk assessment computing system can receive a set of data sources (e.g., news articles or other publications). In some aspects, the data sources can be received from one or more databases storing the data sources. In another aspect, the data sources can be received as a result of a batched web scraping operation. The risk assessment computing system can extract the text from each data sources and employ a name recognition algorithm to identify any entities (e.g., individuals or organizations) mentioned in the data source.

[0013] In some aspects, the risk assessment computing system can analyze the extracted text and split the text into sentences. The risk assessment computing system can identify, categorize, and save various types of information associated with the identified target entity. For example, the risk assessment computing system can identify dates, locations, or events included in the extracted text. Additionally, the risk assessment computing system can apply each sentence that includes the target entity to a sentiment analysis engine to determine the nature of the content associated with US2008230018841Attorney Docket No.096923-1398661 the entity (e.g., to determine whether a sentence relates to an entity positively or negatively, or to determine an entity’s relation to or involvement with an event).

[0014] In some aspects, the risk assessment computing system can employ a function to identify types of events included in the extracted text. As an example, the risk assessment computing system could employ a function to parse the extracted text to identify events and to associate each event with a classification or category based on a mapping table. For example, the mapping table can include a set of events or event types that affect the target entity’s risk assessment. In a non-limiting example, a risk assessment can be based on an entity’s involvement in a financial crime. The risk assessment computing system can parse the text extracted form a data source to identify events (e.g., financial crimes) and categorize the events into event types (e.g., types of financial crimes such as embezzlement, insider trading, fraud, etc.).

[0015] The risk assessment computing system can also analyze the extracted text to identify attributes associated with the name of the target entity in the extracted text. Examples of attributes can include personally identifiable information (PII), headquarters locations, employee names, or other identifying information. Based on the attributes, the risk assessment system can determine a confidence score indicating a likelihood that the name in the extracted text is in fact associated with the target entity.

[0016] If the confidence score is above a predetermined confidence threshold, the risk assessment computing system can update a record associated with the target entity in a database to include the event and sentiment information identified in the data source. This information is then accessible to a user or program querying the database for information associated with the target entity. In some aspects, the risk assessment computing system can analyze the determined event type or sentiment score to determine a risk indicator associated with the target entity. For example, certain event types (e.g., crimes) may be indicative of a higher risk entity, while events, such as IPOs may be indicative of a lower risk entity.

[0017] The system can then transmit the risk indicator to a remote computing system. In some examples, this may be the system from which the risk indicator was requested. The risk indicator can be used to control access of the target entity to an interactive computing environment. For example, the risk indicator can be included in a responsive message to the request for evaluating the target entity such that the responsive message can be used to allow, challenge, or deny access US2008230018841Attorney Docket No.096923-1398661 to the target entity. For example, if the risk indicator is below a predefined threshold, a request by the target entity to access the interactive computing environment may be automatically denied or flagged for manual review.

[0018] Certain aspects described herein, which can include generating one or more risk indicators associated with target entities and providing a responsive message using the risk indicator, can improve at least the technical fields of controlling interactions between computing environments, access control for a computing environment, or a combination thereof. For instance, by generating and transmitting the responsive message, the risk assessment computing system can cause access to a computing system to be controlled more accurately. The risk indicator may be used to better predict whether the target entity requesting access is legitimate, and using the risk indicator may yield fewer malicious interactions than if the responsive message is not used. Further, the risk assessment computing system leverages distinctive components of the risk indicator to create a robust and easily implemented framework.

[0019] These illustrative examples are given to introduce the reader to the general subject matter discussed here and are not intended to limit the scope of the disclosed concepts. The following sections describe various additional features and examples with reference to the drawings in which like numerals indicate like elements, and directional descriptions are used to describe the illustrative examples but, like the illustrative examples, should not be used to limit the present disclosure. Operating Environment Example for Generating a Risk Indicator associated with a Target Entity

[0020] Referring now to the drawings, FIG. 1 is a block diagram depicting an example of an operating environment in which a risk assessment computing system can be used to provide a risk assessment associated with a target entity according to some aspects of the present disclosure. FIG. 1 depicts examples of hardware components of a risk assessment computing system 102, according to some aspects. The risk assessment computing system 102 can be a specialized computing system that may be used for processing large amounts of data using a large number of computer processing cycles. In other examples, the risk assessment computing system 102 may be or include a general- purpose computing system. The risk assessment computing system 102 can include a risk assessment server 104 for performing a risk assessment (e.g., predicting future risk associated with US2008230018841Attorney Docket No.096923-1398661 the target entity, predicting the legitimacy of the target entity, etc.) with respect to a target entity, such as a target individual or a user computing device.

[0021] The risk assessment server 104 can include one or more processing devices that can execute program code, such as a risk assessment application 106. The program code can be stored on a non-transitory computer-readable medium or other suitable medium. The risk assessment application 106 can include one or more modules or components executing software code to complete one or more steps for determining a risk indicator. For example, the risk assessment application 106 can include: a staging module 108; an analysis module 110; a data aggregation module 112, and a verification model 114. The staging module 108 receive one or more data sources and extract text from the data sources. In some aspects, the staging module 108 can split the extracted text into sentences or segments. The staging module 108 can, in some examples, clean and format the extracted text. The extracted text can be passed to the analysis module 110, which may analyze the extracted text to identify events included in the text and determine sentiment scores associated with the events and the target entity. The verification model 114 can be used to determine a confidence score indicating the likelihood that the entity included in the data source is the target entity. If the confidence score is above a predetermined confidence threshold, the data aggregation module 112 can update a database (e.g., data repository 118) record associated with the target entity to include the event and sentiment information. The updated information can be stored as entity data 132 in the data repository 118.

[0022] In some aspects, the risk assessment server 104 can perform risk assessment operations or access control operations for validating or otherwise authenticating the target entity, for example using other suitable modules, models, components, etc. of the risk assessment server 104. The risk assessment server 104 can receive data sources (e.g., documents, news articles, or other publications) from external databases 116, data repository 118, or any suitable combination thereof. In some aspects, the risk assessment application 106 can authenticate or deny a request for an interaction involving the target entity by generating a risk indicator using the target entity data retrieved from the external databases 116 and the data repository 118.

[0023] In some aspects, the target entity data can be determined or stored in one or more network-attached storage units on which various repositories, databases, or other structures are stored. An example of these data structures can include the data repository 118. Additionally or US2008230018841Attorney Docket No.096923-1398661 alternatively, training datasets 120 can be stored in the data repository 118. In some examples, the training datasets 120 can be used to train the verification model 114. The verification model 114 can be used to generate a confidence score that the target entity is included in the data source based on one or more attributes identified in the extracted text.

[0024] Network-attached storage units may store a variety of different types of data organized in a variety of different ways and from a variety of different sources. For example, the network- attached storage unit may include storage other than primary storage located within the risk assessment server 104 that is directly accessible by processors located therein. In some aspects, the network-attached storage unit may include secondary, tertiary, or auxiliary storage, such as large hard drives, servers, and virtual memory, among other types of suitable storage. Storage devices may include portable or non-portable storage devices, optical storage devices, and various other mediums capable of storing and containing data. A machine-readable storage medium or computer-readable storage medium may include a non-transitory medium in which data can be stored and that does not include carrier waves or transitory electronic signals. Examples of a non- transitory medium may include, for example, a magnetic disk or tape, optical storage media such as a compact disk or digital versatile disk, flash memory, memory devices, or other suitable media.

[0025] Furthermore, the risk assessment computing system 102 can communicate with various other computing systems. The other computing systems can include user computing systems 122, such as smartphones, personal computers, etc., client computing systems 124, and other suitable computing systems. For example, user computing systems 122 may transmit, such as in response to receiving input from the target entity, requests for accessing the interactive computing environment 126 to the client computing systems 124. In response, the client computing systems 124 can send authentication queries to the risk assessment server 104, and the risk assessment server 104 can receive data associated with the target entity used in the request and generate a risk indicator associated with the target entity. While FIG. 1 illustrates that the risk assessment computing system 102 and the client computing systems 124 are separate systems, the risk assessment computing system 102 and the client computing systems 124 can be one system. For example, the risk assessment computing system 102 can be a part of the client computing systems 124, or vice versa. US2008230018841Attorney Docket No.096923-1398661

[0026] As illustrated in FIG. 1, the risk assessment computing system 102 may interact with the client computing systems 124, the user computing systems 122, or a combination thereof via one or more public data networks 128 to facilitate interactions between users of the user computing systems 122 and the interactive computing environment 126. For example, the risk assessment computing system 102 can facilitate the client computing systems 124 providing a user interface to the user computing system 122 for receiving various data from the user. The risk assessment computing system 102 can transmit validated risk assessment data, for example similarity- preserving hashes, comparisons or scores determined therefrom, etc., to the client computing systems 124 for providing, challenging, or rejecting, etc. access of the target entity to the interactive computing environment 126. In some examples, the risk assessment computing system 102 can additionally communicate with third-party systems to receive risk assessment data, entity data, and the like, through the public data network 128. In some examples, the third-party systems can provide real-time (e.g., streamed) data about the target entity, historical data about the target entity, etc. to the risk assessment computing system 102.

[0027] Each client computing system 124 may include one or more devices such as individual servers or groups of servers operating in a distributed manner. A client computing system 124 can include any computing device or group of computing devices operated by a seller, lender, or other suitable entity that can provide products or services. The client computing system 124 can include one or more server devices. The one or more server devices can include or can otherwise access one or more non-transitory computer-readable media.

[0028] The client computing system 124 can further include one or more processing devices that can be capable of providing an interactive computing environment 126, such as a user interface, etc., that can perform various operations. The interactive computing environment 126 can include executable instructions stored in one or more non-transitory computer-readable media. The instructions providing the interactive computing environment 126 can configure one or more processing devices to perform the various operations. In some aspects, the executable instructions for the interactive computing environment 126 can include instructions that provide one or more graphical interfaces. The graphical interfaces can be used by a user computing system 122 to access various functions of the interactive computing environment 126. For instance, the interactive computing environment 126 may transmit data to and receive data, such as via the graphical interface, from a user computing system 122 to shift between different states of the interactive US2008230018841Attorney Docket No.096923-1398661 computing environment 126, where the different states allow one or more electronic interactions between the user computing system 122 and the client computing system 124 to be performed.

[0029] In some examples, the client computing system 124 may include other computing resources associated therewith (e.g., not shown in FIG. 1), such as server computers hosting and managing virtual machine instances for providing cloud computing services, server computers hosting and managing online storage resources for users, server computers for providing database services, and others. The interaction between the user computing system 122, the client computing system 124, and the risk assessment computing system 102, or any suitable sub-combination thereof may be performed through graphical user interfaces, such as the user interface, presented by the risk assessment computing system 102, the client computing system 124, other suitable computing systems of the computing environment 100, or any suitable combination thereof. The graphical user interfaces can be presented to the user computing system 122. Application programming interface (API) calls, web service calls, or other suitable techniques can be used to facilitate interaction between any suitable combination or sub-combination of the client computing system 124, the user computing system 122, and the risk assessment computing system 102.

[0030] A user computing system 122 can include any computing device or other communication device that can be operated by a user or entity, such as the user entity, which may include a consumer or a customer. The user computing system 122 can include one or more computing devices such as laptops, smartphones, and other personal computing devices. A user computing system 122 can include executable instructions stored in one or more non-transitory computer-readable media. The user computing system 122 can additionally include one or more processing devices configured to execute program code to perform various operations. In various examples, the user computing system 122 can allow a user to access certain online services or other suitable products, services, or computing resources from a target entity, such as the client computing system 124, to engage in mobile commerce with the client computing system 124, to obtain controlled access to electronic content, such as the interactive computing environment 126, hosted by the client computing system 124, etc.

[0031] In some examples, the user or a target entity can use the user computing system 122 to engage in an electronic interaction with the client computing system 124 via the interactive computing environment 126. The risk assessment computing system 102 can receive a request, for US2008230018841Attorney Docket No.096923-1398661 example from the user computing system 122, to access the interactive computing environment 126 and can use target entity data or any other suitable data or signals determined therefrom, to determine whether to provide access, to challenge the request, to deny the request, etc. An electronic interaction between the user computing system 122 and the client computing system 124 can include, for example, the user computing system 122 being used to request a financial loan or other suitable services or products from the client computing system 124, and so on. An electronic interaction between the user computing system 122 and the client computing system 124 can also include, for example, one or more queries for a set of sensitive or otherwise controlled data, accessing online financial services provided via the interactive computing environment 126, submitting an online credit card application or other digital application to the client computing system 124 via the interactive computing environment 126, operating an electronic tool within the interactive computing environment 126 (e.g., a content-modification feature, an application- processing feature, etc.), etc.

[0032] In some aspects, an interactive computing environment 126 implemented through the client computing system 124 can be used to provide access to various online functions. As a simplified example, a user interface or other interactive computing environment 126 provided by the client computing system 124 can include electronic functions for requesting computing resources, online storage resources, network resources, database resources, or other types of resources. In another example, a website or other interactive computing environment 126 provided by the client computing system 124 can include electronic functions for obtaining one or more financial services, such as an asset report, management tools, credit card application and transaction management workflows, electronic fund transfers, etc.

[0033] A user computing system 122 can be used to request access to the interactive computing environment 126 provided by the client computing system 124. The client computing system 124 can submit a request, such as in response to a request made by the user computing system 122 to access the interactive computing environment 126, for risk assessment to the risk assessment computing system 102 and can selectively grant or deny access to various electronic functions based on risk assessment performed by the risk assessment computing system 102. Based on the request, or continuously or substantially contemporaneously, the risk assessment computing system 102 can determine one or more risk signals or risk indicators for data associated with the target entity, which may submit or may have submitted the request via the user computing system US2008230018841Attorney Docket No.096923-1398661 122. Based on a risk indicator determined from the sentiment score and event information in a data sources, the risk assessment computing system 102, the client computing system 124, or a combination thereof can determine whether to grant the access request of the user computing system 122 to certain features of the interactive computing environment 126. The risk assessment computing system 102, the client computing system 124, or a combination thereof can use the risk indicator for other suitable purposes such as identifying a manipulated identity, controlling a real- world interaction, and the like.

[0034] In a simplified example, the system illustrated in FIG. 1 can configure the risk assessment server 104 to be used for controlling access to the interactive computing environment 126. The risk assessment server 104 can retrieve data sources associated with the target entity in response to a request to access the interactive computing environment 126. The data sources may, for example, be retrieved from databases 116 or received via other suitable computing systems. The databases 116 can store, for example, data sources such as articles or other publications periodically scraped from the Internet. The risk assessment server 104 can determine a risk indicator associated with the target entity by extracting and analyzing text from a data source to determine the target entity’s involvement in certain events and sentiment scores associated with the events. The risk assessment server 104 can transmit the risk indicator, or any inference derived therefrom, to the client computing system 124 for use in controlling access to the interactive computing environment 126.

[0035] The risk indicator associated with the target entity, or any suitable score or comparison determined therefrom, can be used, for example by the risk assessment computing system 102, the client computing system 124, etc., to determine whether the risk associated with the target entity accessing a good or a service provided by the client computing system 124 using exceeds a threshold, thereby granting, challenging, or denying access by the target entity to the interactive computing environment 126. For example, if the risk assessment computing system 102 determines that the risk indicator indicates that risk associated with the identity element is lower than a threshold value, then the client computing system 124 associated with the service provider can generate or otherwise provide access permission to the user computing system 122 that requested the access. The access permission can include, for example, cryptographic keys used to generate valid access credentials or decryption keys used to decrypt access credentials. The client computing system 124 can also allocate resources to the target entity and provide a dedicated web US2008230018841Attorney Docket No.096923-1398661 address for the allocated resources to the user computing system 122, for example, by adding the user computing system 122 in the access permission. With the obtained access credentials or the dedicated web address, the user computing system 122 can establish a secure network connection to the interactive computing environment 126 hosted by the client computing system 124 and access the resources via invoking API calls, web service calls, HTTP requests, other suitable mechanisms or techniques, etc.

[0036] In some examples, the risk assessment computing system 102 may determine whether to grant, challenge, or deny the access request made by the user computing system 122 for accessing the interactive computing environment 126. For example, based on the risk indicator associated with the target entity, the risk assessment computing system 102 can determine that the target entity is a legitimate entity that made the access request and may authenticate the request. In other examples, the risk assessment computing system 102 can challenge or deny the access attempt if the risk assessment computing system 102 determines that the target entity may not be a legitimate entity.

[0037] In some examples, the risk indicator used to determine access to the interactive computing environment 126 may be determined at least in part based on output from one or more machine learning models. For example, the risk assessment application 106 can analyze text from one or more data sources to determine the target entity’s involvement in a particular event. Based on this involvement and a sentiment associated with the involvement, the risk assessment application 106 can determine a risk indicator. To prevent a target entity from being incorrectly associated with an event, the risk assessment application can analyze the text extracted from the data source to identify one or more attributes associated with the entity named in the data source. The one or more attributes can be applied to a machine learning model to generate a confidence score that the entity named in the data source is the target entity. If the confidence score is greater than a confidence threshold, the risk indicator can be transmitted to the requesting system.

[0038] Each communication within the computing environment 100 may occur over one or more data networks, such as a public data network 128, a network 130 such as a private data network, or some combination thereof. A data network may include one or more of a variety of different types of networks, including a wireless network, a wired network, or a combination of a wired and wireless network. Examples of suitable networks include the Internet, a personal area US2008230018841Attorney Docket No.096923-1398661 network, a local area network (“LAN”), a wide area network (“WAN”), or a wireless local area network (“WLAN”). A wireless network may include a wireless interface or a combination of wireless interfaces. A wired network may include a wired interface. The wired or wireless networks may be implemented using routers, access points, bridges, gateways, or the like, to connect devices in the data network.

[0039] The number of devices illustrated in FIG. 1 is provided for illustrative purposes. Different numbers of devices may be used. For example, while certain devices or systems are shown as single devices in FIG.1, multiple devices may instead be used to implement these devices or systems. Similarly, devices or systems that are shown as separate may be instead implemented in a signal device or system. Exemplary Application for Generating a Risk Indicator associated with a Target Entity

[0040] FIG. 2 is a block diagram depicting an example risk assessment application 106 for generating a risk assessment associated with a target entity according to some aspects of the present disclosure. As discussed with reference to FIG.1, the risk assessment application 106 can be stored on a risk assessment server 104 and can be used by the risk assessment system 102 to generate risk indicators for target entities in response to requests from the client computing systems 124 or the user computing systems 122. Other implementations or architectures, however, are possible.

[0041] As discussed with reference to FIG.1, the risk assessment application 106 can include a staging module 108, an analysis module 110, a data aggregation module 112, and a verification model 114. In response to receiving a request for a risk assessment of a target entity, the risk assessment application 106 can receive a set of one or more data sources 212 from databases 116. In some aspects, the data sources 212 can be pulled from databases 116 in response to a request for a risk indicator. In other aspects, the data sources 212 can be automatically received by the risk assessment application at predetermined time intervals.

[0042] The staging module 108 can include sub-modules such as: a file reading and text extraction sub-module 202; a named entity recognition sub-module 204; and a sentence splitting sub-module 206. The file reading and text extraction sub-module 202 can extract text from each data source 212. For example, the sub-module 202 can extract text from various data source types such as WORD documents or PDF documents. In some aspects, the databases 116 can include US2008230018841Attorney Docket No.096923-1398661 links to data sources on the Internet such that the sub-module 202 can extract text from the linked data sources using a web scraping or text recognition tool.

[0043] The named entity recognition sub-module 204 can receive the extracted text from the file reading and text extraction sub-module 202. The named entity recognition sub-module 204 can identify entities (e.g., individuals or organizations) included or mentioned in the extracted text, for example, by employing a pipeline component for named entity recognition such as spaCy. In some aspects, the sub-module 204 can categorize the identified entities into people and organizations. In certain aspects in which a request identifies a target entity, the named entity recognition sub-module 204 can parse the extracted text for names or variations of names associated with the target entity. The sub-module 204 can also identify all entities in a data source and generate a ranked list of the entities by category and by number of mentions in the data source. For example, where both individuals and organizations are identified in a data source, the sub- module 204 can generate two ranked lists – one for individuals and one for organizations. The top individual and organization can be categorized as a “main person” and “main organization” of the data source. The other entities mentioned in the data source can be categorized as “other entities.”

[0044] In some aspects, the sub-module 204 can also identify one or more attributes associated with one or more “main” entities of the data source. For example, attributes can include PII (e.g., name, date of birth, address, phone number, etc.) for an individual. For an organization attributes can include an organization name, stock symbol, headquarters location, and the like. This attribute information can later be used by the verification model 114 to generate a confidence score that either the main person or main organization is the target entity. In some aspects, the sub-module 204 can further determine relationships between entities, such as an employer-employee relationship.

[0045] In some aspects, the sub-module 204 can further include functionality to recognize names associated with performers or titles (e.g., movie or book titles). In cases in which a performer or title are recognized in the data source text, the text can be applied to a machine learning model to determine a context in which the performer or title is discussed. For example, the machine learning model can be a large language model and can determine the context in which an entity is mentioned. Thus, the sub-module 204 can analyze the data source and determine to US2008230018841Attorney Docket No.096923-1398661 exclude the data source if the text discusses fictional characters, performances, or titles of fictional works.

[0046] The staging module 108 can also include the sentence splitting sub-module 206. The sentence splitting sub-module 206 can parse the extracted text and divide the text into sentences. For example, the sentence splitting sub-module 206 can analyze the text to recognize punctuation indicating the end of a sentence. Sentences containing the target entity name, or, in some aspects, the main person or main organization, can be provided to the analysis module 110. The sentence splitting module 206 can divide the extracted text into one or more sentences, for example, using software such as TextBlob.

[0047] In some examples, the sentence splitting module 206 can include further clean and deduplicate the extracted text. For example, once one or more sentences associated with the target entity are generated, the sentence splitting module 206 can remove repetitive sentences from the one or more sentences. Repetitive sentences can be, for example, sentences including repeated instances of the entity name, attribute information, or other details (e.g., events, dates, locations) identified by the sub-module 204.

[0048] The analysis module 110 can include: a sentiment analysis sub-module 208; and an event classification sub-module 210. The sentiment analysis sub-module 208 can include, in some aspects, a sentiment analysis engine. For example, a sentiment analysis engine can be an engine, such as Valence Aware Dictionary and sEntiment Reasoner (VADER), that applies syntactical and grammatical rules to extracted text to determine sentiment and generate a sentiment score, although other sentiment analysis tools can be used.

[0049] The sentiment analysis sub-module 208 can be applied to each sentence of the data source that contains the target entity name (or the main person or main organization). A score can be determined for each sentence such that each sentence can be categorized. For example, if the sentiment score is greater than a predetermined threshold, the sentence can be categorized as “positive,” and if the sentiment score is below the predetermined threshold, the sentence can be categorized as “negative.” In some aspects, a sentence having a sentiment score within a predetermined range can be categorized as “neutral.”

[0050] The analysis module 110 can also include the event classification sub-module 210. The event classification sub-module 210 can parse each sentence of the data source (or each sentence US2008230018841Attorney Docket No.096923-1398661 containing the target entity name) to determine if the sentence includes an event. If the sentence includes an event, the event classification sub-module 210 can categorize the event based on a mapping table which may be stored, for example, in the data repository 118. As an example, the event classification sub-module 210 can use a text recognition function to identify a crime included, or mentioned, in the text. Based on the text associated with the mention of the crime, the event classification sub-module can map that crime to a particular type of crime (e.g., financial or non-financial).

[0051] Once the entity, attribute, and event data are identified in the extracted text, the attribute data can be applied to the verification model 114 to generate a confidence score. The verification model 114 can be a machine learning model trained on training data (e.g., training datasets 120) associated with entity identities. For example, the verification model 114 can be trained to generate a confidence score that the entity named in the data source is the target entity based on the attributes associated with the entity in the data source. In some examples, certain attributes can be weighted such that attributes having a higher likelihood of being associated with a single entity have a higher weight.

[0052] In some examples, the verification model 114 can be used to determine if an entity named in the data source is actually associated with the target entity (or main person or main organization). For example, if the confidence score is greater than a confidence threshold, the entity named in the data source is associated with the target entity (or main person or organization). Based on this determination, the sentiment scores and event information associated with the entity can be transmitted to the data aggregation module 112.

[0053] The data aggregation module 112 can either generate a new record for the entity in the data repository 118 or can update an existing record in the data repository 118. Thus, once a target entity if verified, a record can be created, or an existing record associated with the target entity can be updated to include the sentiment scores and event information (e.g., event type, event location, event date, and the like). The data aggregation module 112 can also deduplicate any information already stored in the data repository 118 that is associated with the entity to prevent the creation of multiple records.

[0054] In some aspects, the sentiment score and event information can be used to generate a risk indicator. The risk assessment application 106 can retrieve the sentiment score and event US2008230018841Attorney Docket No.096923-1398661 information for the target entity from the data repository 118 in response to receiving a request for a risk indicator associated with the target entity. The risk indicator can be determined, for example, in part from the type of event with which the target entity is associated. A target entity associated with a financial crime can have a higher risk indicator for a request to access a financial system. In some examples, an entity’s relationship or involvement in a crime can also affect the risk indicator. For example, an organization doing business with an organization involved in insider trading may have a lower risk indicator than the organization involved in insider trading.

[0055] Accordingly, the risk assessment application 106 can be used to generate risk indicators, sentiment scores, and event information associated with entities. This information can be beneficial in determining whether to allow a target entity to access an interactive computing environment by providing automatic and up to date information about events in which an entity is involved and the nature of their involvement. This information can be used in more accurately predicting risk associated with a target entity. Techniques for Generating a Risk Indicator associated with a Target Entity

[0056] FIG. 3 is a flow chart illustrating an example of a process 300 for generating a risk assessment associated with a target entity according to some aspects of the present disclosure. In some examples, the operations of the process 300, or any subset thereof, may be performed by the risk assessment computing system 102 via the risk assessment server 104, but other suitable systems, devices, or subsets or combinations thereof may perform one or more operations described with respect to the process 300. For illustrative purposes, the process 300 is described with reference to certain examples depicted in the figures. Other implementations, however, are possible.

[0057] At block 302, the process 300 involves receiving a request for a risk indicator associated with a target entity. The request may be generated as part of an authentication process initiated when the target entity attempts to access an interactive computing environment 126. In some aspects, the process 300 can begin at step 304, and the request for a risk indicator can be received after the entity data is aggregated, such that the risk indicator can be determined from existing sentiment scores and event information.

[0058] At block 304, the process 300 involves determining that a data source contains a name associated with the target entity based on extracted text from the data source. As discussed above, US2008230018841Attorney Docket No.096923-1398661 the staging module 108 can extract text from a data source and analyze the text to identify an entity mentioned in the text. In some examples, the main entity for each data source can be determined by preprocessing, such that data sources containing a name associated with the target entity are retrieved from data bases 116.

[0059] At block 306, the process 300 involves, dividing the extracted text into at least one sentence. As discussed above, the risk assessment application 106 can parse the extracted text to identify punctuation and divide the text into sentences. The risk assessment application 106 can, for example, employ one or more software packages configured for natural language processing (NLP) to divide the extracted text into sentences.

[0060] At block 308, the process 300 involves identifying a sentence containing the name associated with the target entity. For example, the risk assessment application 106 can apply a name recognition algorithm to the sentence to identify the target entity within the sentence. In some examples, multiple sentences may contain the name associated with the target entity. In some aspects, the risk assessment application 106 can deduplicate the sentences containing the name associated with the target entity by determining a number of sentences that contain the name. The risk assessment application 106 can then remove any sentences of the number of sentences that exceed a deduplication threshold, thereby improving efficiency by reducing replicated analysis.

[0061] At block 310, the process 300 involves determining a sentiment score associated with the sentence. As discussed above, the sentiment score can be determined by applying the sentence to a sentiment analysis engine configured to recognize one or more sentiments expressed in the sentence. The sentiment analysis engine can be used to generate a score based on grammar, syntax, characters, punctuation, etc. contained in the sentence. In some aspects, the sentiment of the sentence can be categorized based a comparison of the sentiment score to various sentiment category thresholds.

[0062] At block 312, the process 300 involves generating classification associated with an event included in the sentence. For example, the risk assessment application 106 can use a mapping table stored in the data repository 118 to map an event identified in the extracted text to a classification (e.g., an event type or event category). The event type can further provide a basis for determining the risk indicator depending on the mapping table, which can include an indicator of how the event affects the entity’s risk indicator (e.g., positively or negatively). US2008230018841Attorney Docket No.096923-1398661

[0063] At block 314, the process 300 involves generating the risk indicator based on at least one of the classification or the sentiment score. As discussed above the risk indicator can be determined based on at least one of the sentiment score or the event information. As an example, the risk assessment system 102 can store a mapping of event types and sentiment scores that result in a risk indicator that is below a threshold for access to an interactive computing environment. In other examples, the mapping table can include a listing of event types that do not affect a risk indicator. Thus, the risk assessment system 102 can generate a risk indicator based on several factors, one of which being the target entity’s involvement in an event and the sentiment associated with that event.

[0064] At block 316, the process 300 involves determining a confidence score that the name in the extracted text is associated with the target entity, where the confidence score is determined based on attributes associated with the name in the extracted text. As discussed above, the risk assessment application 106 can identify one or more attributes in the extracted text that are associated with the target entity. The identified attributes can be used to correlate the entity mentioned in the data source to the target entity. The attributes can be applied to a machine learning model (e.g., verification model 114) to generate the confidence score. The confidence score can indicate a likelihood that the entity named in the data source is referring to the target entity.

[0065] At block 318, upon determining that the confidence score is above a confidence threshold, the process 300 involves transmitting, to a remote computing device, a responsive message comprising at least the risk indicator for use in controlling access of the target entity to one or more interactive computing environments. For example, the risk indicator can be used in controlling an interaction involving a target entity or access of the target entity to a restricted system.

[0066] Systems and methods described herein provide advantages over traditional risk assessment systems. In some examples, the risk assessment system 102 can provide flexible and scalable means for determining risk based on an entity’s involvement with certain events. Additionally, by leveraging robust sentiment analysis engines, the system can make more accurate determinations regarding the nature of an entity’s involvement in an event. Further, the risk assessment system 102 can be automated such that the data associated with each entity is continually updated and is retrievable for generating risk assessments. US2008230018841Attorney Docket No.096923-1398661 Example of Computing System

[0067] Any suitable computing system or group of computing systems can be used to perform the operations for the techniques described herein. For example, FIG. 4 is a block diagram depicting an example of a computing device 400, which can be used to implement the risk assessment server 104. The computing device 400 can include various devices for communicating with other devices in the computing environment 100, as described with respect to FIG. 1. The computing device 400 can include various devices for performing one or more operations, such as risk assessment operations, described above with respect to FIGs.1-3.

[0068] The computing device 400 can include a processor 402 that can be communicatively coupled to a memory 404. The processor 402 can execute computer-executable program code stored in the memory 404, can access information stored in the memory 404, or both. Program code may include machine-executable instructions that may represent a procedure, a function, a subprogram, a program, a routine, a subroutine, a module, a software package, a class, or any combination of instructions, data structures, or program statements. A code segment may be coupled to another code segment or a hardware circuit by passing or receiving information, data, arguments, parameters, or memory contents. Information, arguments, parameters, data, etc., may be passed, forwarded, or transmitted via any suitable means including memory sharing, message passing, token passing, network transmission, among others.

[0069] Examples of a processor 402 can include a microprocessor, an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or any other suitable processing device. The processor 402 can include any suitable number of processing devices, including one. The processor 402 can include or communicate with a memory 404. The memory 404 can store program code that, when executed by the processor 402, causes the processor 402 to perform the operations described herein.

[0070] The memory 404 can include any suitable non-transitory computer-readable medium. The computer-readable medium can include any electronic, optical, magnetic, or other storage device capable of providing a processor with computer-readable program code or other program code. Non-limiting examples of a computer-readable medium can include a magnetic disk, memory chip, optical storage, flash memory, storage class memory, ROM, RAM, an ASIC, magnetic storage, or any other medium from which a computer processor can read and execute US2008230018841Attorney Docket No.096923-1398661 program code. The program code may include processor-specific program code generated by a compiler or an interpreter from code written in any suitable computer-programming language. Examples of suitable programming language can include Hadoop, C, C++, C#, Visual Basic, Java, Python, Perl, JavaScript, ActionScript, etc.

[0071] The computing device 400 may also include a number of external or internal devices such as input or output devices. For example, the computing device 400 is illustrated with an input / output interface 408 that can receive input from input devices or provide output to output devices. A bus 406 can also be included in the computing device 400. The bus 406 can communicatively couple one or more components of the computing device 400.

[0072] The computing device 400 can execute program code 414 that can include risk assessment application 106. The program code 414 for the risk assessment application 106 may be resident in any suitable computer-readable medium and may be executed on any suitable processing device. For example, and as illustrated in FIG. 4, the program code 414 for the risk assessment application 106 can reside in the memory 404 at the computing device 400 along with the program data 416 associated with the program code 414. Executing the risk assessment application 106 can configure the processor 402 to perform at least a portion of the operations described herein.

[0073] In some aspects, the computing device 400 can include one or more output devices. One example of an output device can be or include the network interface device 410 illustrated in FIG. 4. A network interface device 410 can include any device or group of devices suitable for establishing a wired or wireless data connection to one or more data networks described herein. Non-limiting examples of the network interface device 410 can include an Ethernet network adapter, a modem, etc.

[0074] Another example of an output device can include the presentation device 412 depicted in FIG. 4. A presentation device 412 can include any device or group of devices suitable for providing visual, auditory, or other suitable sensory output. Non-limiting examples of the presentation device 412 can include a touchscreen, a monitor, a speaker, a separate mobile computing device, etc. In some aspects, the presentation device 412 can include a remote client- computing device that communicates with the computing device 400 using one or more data networks described herein. In other aspects, the presentation device 412 can be omitted. US2008230018841Attorney Docket No.096923-1398661

[0075] The foregoing description of some examples has been presented only for the purpose of illustration and description and is not intended to be exhaustive or to limit the disclosure to the precise forms disclosed. Numerous modifications and adaptations thereof will be apparent to those skilled in the art without departing from the spirit and scope of the disclosure. US2008230018841

Claims

Attorney Docket No.096923-1411973 Claims What is claimed is:

1. A system comprising: a processor; and a non-transitory computer-readable medium comprising instructions that are executable by the processor for causing the processor to perform operations comprising: receiving a request for a risk indicator associated with a target entity; determining that a data source contains a name associated with the target entity based on extracted text from the data source; dividing the extracted text into at least one sentence; identifying a sentence of the at least one sentence that includes the name associated with the target entity; determining a sentiment score associated with the sentence; generating a classification associated with an event included in the sentence; generating the risk indicator based on at least one of the classification or the sentiment score; determining a confidence score that the name in the extracted text is associated with the target entity, the confidence score determined based on attributes associated with the name in the extracted text; and upon determining that the confidence score is above a confidence threshold, transmitting, to a remote computing device, a responsive message comprising at least the risk indicator used to control access of the target entity to one or more interactive computing environments.

2. The system of claim 1, wherein the operation of determining that the data source contains the name associated with the target entity comprises: analyzing the text from the data source to determine a set of entities referenced in the text; ranking each entity of the set of entities based on a number of times each entity is referenced in the text; and US2008230018841Attorney Docket No.096923-1411973 determining that the data source is associated with the target entity based on the number of times the target entity is referenced in the text being above a predetermined threshold.

3. The system of claim 1, wherein the target entity comprises at least one of an individual or an organization.

4. The system of claim 1, wherein the operations further comprise: deduplicating the at least one sentence of the extracted text by determining a number of sentences of the at least one sentence that include the name associated with the target entity; and removing, from the number of the at least one sentence, sentences causing the number to exceed a deduplication threshold.

5. The system of claim 1, wherein the sentiment score indicates that a sentence relates positively to the target entity based on the sentiment score being above a sentiment threshold.

6. The system of claim 1, wherein the operation of generating the classification associated with the event comprises: applying the sentence to a language model to determine an event type associated with the event; and determining that the event type is included in a list of events associated with a risk.

7. The system of claim 1, wherein the operation of generating the risk indicator further comprises: identifying, in an entity database, a previous risk indicator associated with the target entity; and updating the previous risk indicator based on the risk indicator.

8. A method comprising: US2008230018841Attorney Docket No.096923-1411973 receiving, by a processor, a request for a risk indicator associated with a target entity; determining, by the processor, that a data source contains a name associated with the target entity based on extracted text from the data source; dividing, by the processor, the extracted text into at least one sentence; identifying, by the processor, a sentence of the at least one sentence that includes the name associated with the target entity; determining, by the processor, a sentiment score associated with the sentence; generating, by the processor and based on the sentence, a classification associated with an event included in the sentence; generating, by the processor, the risk indicator based on at least one of the classification or the sentiment score; determining, by the processor, a confidence score that the name in the extracted text is associated with the target entity, the confidence score determined based on attributes associated with the name in the extracted text; and upon determining that the confidence score is above a confidence threshold, transmitting, by the processor to a remote computing device, a responsive message comprising at least the risk indicator used to control access of the target entity to one or more interactive computing environments.

9. The method of claim 8, wherein determining that the data source contains the name associated with the target entity comprises: analyzing, by the processor, the text from the data source to determine a set of entities referenced in the text; ranking, by the processor, each entity of the set of entities based on a number of times each entity is referenced in the text; and determining, by the processor, that the data source is associated with the target entity based on the number of times the target entity is referenced in the text being above a predetermined threshold.

10. The method of claim 8, wherein the target entity comprises at least one of an individual or an organization. US2008230018841Attorney Docket No.096923-1411973 11. The method of claim 8, wherein the method further comprises: deduplicating, by the processor, the at least one sentence of the extracted text by determining a number of the at least one sentence that include the name associated with the target entity; and removing, from the number of the at least one sentence by the processor, sentences causing the number to exceed a deduplication threshold.

12. The method of claim 8, wherein the sentiment score indicates that a sentence relates positively to the target entity based on the sentiment score being above a sentiment threshold.

13. The method of claim 8, wherein generating the classification associated with the event comprises: applying, by the processor, the sentence to a language model to determine an event type associated with the event; and determining, by the processor, that the event type is included in a list of events associated with a risk.

14. The method of claim 8, wherein generating the risk indicator further comprises: identifying, in an entity database by the processor, a previous risk indicator associated with the target entity; and updating, by the processor, the previous risk indicator based on the risk indicator.

15. A non-transitory computer-readable medium comprising instructions that are executable by a processor for causing the processor to perform operations comprising: receiving a request for a risk indicator associated with a target entity; determining that a data source contains a name associated with the target entity based on extracted text from the data source; dividing the extracted text into at least one sentence; US2008230018841Attorney Docket No.096923-1411973 identifying a sentence of the at least one sentence that includes the name associated with the target entity; determining a sentiment score associated with the sentence; generating a classification associated with an event included in the sentence; generating the risk indicator based on at least one of the classification or the sentiment score; determining a confidence score that the name in the extracted text is associated with the target entity, the confidence score determined based on attributes associated with the name in the extracted text; and upon determining that the confidence score is above a confidence threshold, transmitting, to a remote computing device, a responsive message comprising at least the risk indicator used to control access of the target entity to one or more interactive computing environments.

16. The non-transitory computer-readable medium of claim 15, wherein the operation of determining that the data source contains the name associated with the target entity comprises: analyzing the text from the data source to determine a set of entities referenced in the text; ranking each entity of the set of entities based on a number of times each entity is referenced in the text; and determining that the data source is associated with the target entity based on the number of times the target entity is referenced in the text being above a predetermined threshold.

17. The non-transitory computer-readable medium of claim 15, wherein the target entity comprises at least one of an individual or an organization.

18. The non-transitory computer-readable medium of claim 15, wherein the operations further comprise: deduplicating the at least one sentence of the extracted text by determining a number of the at least one sentence that include the name associated with the target entity; and US2008230018841Attorney Docket No.096923-1411973 removing, from the number of the at least one sentence, sentences causing the number to exceed a deduplication threshold.

19. The non-transitory computer-readable medium of claim 15, wherein generating the classification associated with the event comprises: applying the sentence to a language model to determine an event type associated with the event; and determining that the event type is included in a list of events associated with a risk.

20. The non-transitory computer-readable medium of claim 15, wherein the operation of generating the risk indicator further comprises: identifying, in an entity database, a previous risk indicator associated with the target entity; and updating the previous risk indicator based on the risk indicator. US2008230018841

Citation Information

Patent Citations

  • Methods and systems for generating composite index using social media sourced data and sentiment analysis

    US20120296845A1

  • System and semi-supervised methodology for performing machine driven analysis and determination of integrity due diligence risk associated with third party entities and associated individuals and stakeholders

    US20210026835A1

  • Systems and methods for identifying, quantifying, and mitigating risk

    US20230186213A1