Methods, systems, and computer programs for correlating regulatory data using a processor in a computing environment (identification of regulatory data corresponding to executable rules).
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- INTERNATIONAL BUSINESS MACHINE CORPORATION
- Filing Date
- 2023-01-20
- Publication Date
- 2026-08-06
AI Technical Summary
【0006】 本発明の利点を容易に理解できるようにするため、上で簡単に記載した本発明を、添付図面に示す特定の実施形態を参照しながらさらに詳細に説明する。これらの図面は、本発明の典型的な実施形態のみを示しており、したがってその範囲を限定するものとはみなされないことを理解した上で、本発明は、添付図面の使用により追加的な特定事項および詳細事項を伴って記載および説明される。
Smart Images

Figure 0007901418000002 
Figure 0007901418000003 
Figure 0007901418000004
Abstract
Description
[Technical Field]
[0001] The present invention generally relates to a computing system, and more particularly to various embodiments for using a computing processor to identify regulatory data and correlate that regulatory data with executable rules. [Background technology]
[0002] Computing systems may be found in the workplace, home, or school. Computer systems may include data storage systems or disk storage systems for processing and storing data. Large amounts of data must be processed daily, and current trends suggest that these amounts will continue to increase to unprecedented levels in the near future. Due to recent advances in information technology and the growing popularity of the internet, vast amounts of information are now available in digital format. Such availability of information presents numerous opportunities. Digital and online information is a valuable source of business intelligence, crucial for entities to survive and adapt in a highly competitive environment. Furthermore, many businesses and organizations involved in the use of computing systems and online data must ensure that their operations, procedures, or combinations thereof comply with common business protocols, corporate compliance, or legal regulations, policies, or requirements, or combinations thereof. [Overview of the project] [Problems that the invention aims to solve]
[0003] Many businesses and organizations that use computing systems and online data must ensure that their operations, business practices, or procedures, or any combination thereof, comply with common business protocols, corporate compliance, or legal regulations, policies, or requirements, or any combination thereof. [Means for solving the problem]
[0004] Various embodiments are provided for a processor to identify regulatory data and correlate it with executable rules in a computing environment. In one embodiment, as merely an example, a method is provided for a processor to identify regulatory data and correlate it with executable rules. The rules may be associated with one or more text paragraphs extracted from a policy document that contains at least a portion of the rules.
[0005] In other embodiments, semantic data extracted from a law, policy, regulation, or combination thereof may be associated with text data from one or more data sources that describe at least a portion of that law, policy, regulation, or combination thereof. [Brief explanation of the drawing]
[0006] To facilitate understanding of the advantages of the present invention, the invention, briefly described above, will be further described with reference to specific embodiments shown in the accompanying drawings. While it should be understood that these drawings only illustrate typical embodiments of the invention and are therefore not intended to limit its scope, the invention will be described and explained with additional specifics and details using the accompanying drawings.
[0007] [Figure 1] This is a block diagram showing an exemplary cloud computing node according to an embodiment of the present invention.
[0008] [Figure 2] This is an additional block diagram illustrating an exemplary cloud computing environment according to an embodiment of the present invention.
[0009] [Figure 3] This is an additional block diagram showing an abstraction model layer according to an embodiment of the present invention.
[0010] [Figure 4] An additional block diagram showing exemplary functional relationships between various aspects of the present invention.
[0011] [Figure 5] A block diagram showing an exemplary operation for identifying regulatory data and correlating it with executable rules to implement aspects of the present invention.
[0012] [Figure 6] A block diagram showing an exemplary operation and workflow for identifying regulatory data and correlating it with executable rules to implement aspects of the present invention.
[0013] [Figure 7] A diagram showing an exemplary knowledge graph for identifying regulatory data and correlating it with executable rules to implement aspects of the present invention.
[0014] [Figure 8] A diagram showing an exemplary operation using forward and inverse transformations for identifying regulatory data and correlating it with executable rules to implement aspects of the present invention.
[0015] [Figure 9] A flowchart diagram showing an exemplary method for a processor to identify regulatory data and correlate it with executable rules to implement aspects of the present invention.
Best Mode for Carrying Out the Invention
[0016] As the amount of electronic information continues to increase, the demand for advanced information access systems is also growing. Digital or "online" data is becoming increasingly accessible via real-time global computer networks. Data can reflect various aspects of topics spanning science, law, education, finance, travel, shopping and entertainment activities, healthcare, and more. Many data-intensive applications require the extraction of information from data sources. Information extraction may be obtained through a knowledge generation process, which may include initial data collection, data normalization and aggregation, and final data extraction across different sources.
[0017] Furthermore, an entity (e.g., a business, government, organization, academic research institution, etc.) may be subject to specific processes, policies, guidelines, rules, laws, or regulations, or combinations thereof, related to that entity. Complying with these processes, policies, guidelines, rules, laws, or regulations, or combinations thereof, is also important and essential to ensure the company's compliance and avoid violations, fines, or legal penalties. For example, since new regulations occur continuously, regulatory compliance management is the most important and critical matter for an organization. In one aspect, regulatory compliance means that an entity adheres to laws, regulations, guidelines, and specifications related to its purpose or business. These companies / entities often require human interaction with people having various skills and expertise (e.g., subject matter experts (SMEs)) to support compliance among enterprises.
[0018] Furthermore, to support regulatory compliance, governments and businesses are automating policies in the form of coded rules (for example, to verify the eligibility of residents for certain benefits they may be entitled to). For example, "Rules as Code" (RaC) is an initiative that envisions formal versions of rules (e.g., laws and regulations) in a machine-consumable form, enabling them to be understood and addressed by computer systems in a consistent manner. This forms part of a larger movement toward digital government and has attracted the attention of a wide range of public sectors. Recent OECD reports on Rules as Code[3] identify ways to address this, such as bringing together lawmakers, policy analysts, and software developers to co-create versions of policies and machine-consumable rules, or using AI and automation to shorten the route from policy to code.
[0019] Organizations automating policies are responsible for addressing errors in the "translation of intent" from laws to policies, business requirements, and executable rules and code, but often the link between the implementation (executable rules / code) and the original policy text from which it originated is missing. Thus, given the vast amounts of text data and the pace at which regulatory documents change, various embodiments for identifying regulatory data and correlating them with executable rules are provided herein. Semantic data extracted from laws, policies, regulations, or combinations thereof may be associated with text data from one or more data sources that describe at least a portion of those laws, policies, regulations, or combinations thereof.
[0020] In several implementations, the present invention enables the alignment and correlation of existing policy rules (code) with specific sections / paragraphs of policy and legal text. and text BetweenThis enables mapping, addresses the issue of multi-transformation in legal domains, and ensures the following:
[0021] 1) The present invention enables rules and their dependencies on other rules to be learned and understood, and corrected by machine learning operations. Note that a section / paragraph may be translated into multiple rules, and a rule may depend on other rules originating from other sections within a policy, or even on other rules spanning a policy manual, legal regulations, or a combination thereof. 2) The present invention enables support and detection of legal errors, such as missing policy sections in the code. 3) The present invention enables support, maintenance, and updating of rules when policy changes are introduced. In particular, suitable The rules of personality Latest 4) The invention enables a more efficient and faster process for making these updates to accurately reflect the policy intent. 5) The invention enables support and identification of rule hierarchies, duplicate rules, and attributes or definitions. 6) The invention enables support and reusability by identifying which rules already exist when a new policy is introduced. That is, existing rules that are similar to parts of the new policy text, i.e., have the same definition or subtle differences, for example, the enforcement of residency rules across two states or the enforcement between different policies on cash benefits and child allowances may be very similar. 7) The invention also enables the generation of training data for machine learning operations that support entity annotation or rule extraction by correlating and aligning rules with text.
[0022] Thus, the present invention enables the identification of regulatory data that correlates with executable rules in a computing environment. In one embodiment, laws, statutes, policies, regulations, or combinations thereof may be extracted from one or more segments of text data from one or more data sources that can be identified as obligations that entities are required to fulfill. Semantic data extracted from laws, policies, regulations, or combinations thereof may be associated with text data from one or more data sources that describe at least a portion of the laws, policies, regulations, or combinations thereof. In one embodiment, “obligation” may represent a legal requirement (including laws, policies, regulations, orders, responsibilities, or combinations thereof). In other words, the obligation target / content extraction component may execute a perceptron algorithm for extracting entity classes. Note that “obligation” represents a legal requirement (including laws, policies, regulations, orders, responsibilities, or combinations thereof).
[0023] In one embodiment, one or more entities perform one or more natural language processing (NLP) operations, or solid Yes Named entities may be extracted using a Name Recognition (NER) operation or a combination thereof. The NER operation recognizes named entities within the text. Special This may be a subtask of information extraction that can be determined and classified into predefined categories such as people, entities, organizations, and locations. A set of sentences having content such as obligations may be determined (e.g., calculated) using an extraction operation from sentences and one or more filtering operations applied to content of semantic roles. A machine learning (ML) classifier may be used to determine whether selected segments / sections of the ingested text data are obligations / requirements.
[0024] In one embodiment, as used herein, the term "regulation" may be a document written in natural language that encompasses a set of laws, statutes, policies, regulations, regulatory targets, entities, and requirements that specify obligations, obligation targets, constraints, and priorities relating to an entity's desired structure and behavior. A regulation may specify the domain elements to which it applies. For example, a regulation may be a law (e.g., a medical law, an environmental protection law, an aviation law), a standardization document, a contract, etc. Also, as used herein, the type of entity in question may include, for example, a “definition,” which is a section / text segment that defines / represents a stakeholder or specific equipment. That is, a “definition” may be a section in legal text that defines a specific party, entity, or stakeholder, or a combination thereof, that is regulated by text (e.g., legal text such as rules / laws / policies of a particular jurisdiction). A definition target may be a definition entity, and definition content may be a section / text segment that lists the target. Also, an “obligation” may be a section that represents a legal requirement. An obligation target is the legal target of a particular section. Obligation content may be the requirement that applies to the target.
[0025] Furthermore, the term “domain” is intended to have its usual meaning. In addition, the term “domain” may include areas of expertise concerning a set of systems, materials, information, content, or other resources, or combinations thereof, related to a particular subject. For example, a domain may refer to information specific to regulation, law, policy, government, finance, healthcare, advertising, commerce, science, industry, education, medicine, biomedicine, or other areas or information defined by subject experts. A domain may refer to information concerning any particular subject or a selected combination of subjects.
[0026] The term "ontology" is also intended to have its usual meaning. For example, an ontology may include information or content related to the content of a target domain or a particular class or concept. The content can be any searchable information, such as information distributed on a computer-accessible network, such as the Internet. Concepts or topics may generally be classified into one of many content concepts or topics, and these content concepts or topics may also include one or more subconcepts, one or more subtopics, or a combination thereof. Examples of concepts or topics may include, but are not limited to, regulatory compliance information, policy information, legal information, government information, business information, educational information, or any other group of information. An ontology can be continuously updated with information synchronized with the source, and information from the source can be used as a model, Mo Dell attributes, or associations between models within an ontology as ontology Add to it.
[0027] Where used herein, the term “intelligent” (or “cognitive”) may relate to, be, or include conscious intellectual activities that may be performed using machine learning, such as thinking, reasoning, or memory. In additional embodiments, intelligent or “intelligence” may be the mental process of knowing, including aspects such as recognition, perception, reasoning, and judgment. A machine learning system may use artificial reasoning to interpret data from one or more data sources (e.g., sensor-based devices or other computing systems) and may learn topics, concepts, judgment-reasoning knowledge, or processes, or combinations thereof, which may be determined or derived, or both, by machine learning.
[0028] Generally as used herein, “optimization” may mean, or may be defined as, the “maximization,” “minimization,” “most likely,” “best,” or achievement of one or more specific targets, objectives, goals, or intentions. Optimization may also mean maximizing the benefit to the user (e.g., maximizing the benefit of a trained machine learning pipeline / model). Optimization may also mean using a situation, opportunity, or resource in the most efficient or functional way.
[0029] Furthermore, optimization does not necessarily refer to the best solution or outcome, but may refer to a "sufficiently good" or "most likely" solution or outcome for a particular use case, for example. In some implementations, the objective is to suggest the best combination of preprocessing operations ("preprocessors") or machine learning models / pipelines or both, but there may be various factors that lead to alternative suggestions for preprocessing operations ("preprocessors") or machine learning models or combinations of both that produce better results. In this specification, the term "optimized" may refer to such an outcome based on a minimum (or maximum, depending on the parameters considered in the optimization problem). In additional embodiments, the terms "optimize" or "optimize" or both may refer to an operation performed to achieve an improved outcome, such as a reduction in execution cost or an increase in resource utilization, regardless of whether the optimal outcome is actually achieved. Similarly, the term "optimize" may refer to a component for performing such an improvement operation, and the term "optimized" may be used to describe the outcome of such an improvement operation.
[0030] While this disclosure includes a detailed description of cloud computing, it should be understood that the implementations of the teachings described herein are not limited to cloud computing environments. Rather, embodiments of the present invention can be implemented in conjunction with any other type of computing environment that is currently known or may be developed in the future.
[0031] Cloud computing is a service delivery model that enables convenient, on-demand network access to a shared pool of configurable computing resources (e.g., networks, network bandwidth, servers, processing, memory, storage, applications, virtual machines, and services) that can be rapidly provisioned and deployed with minimal management effort or interaction with service providers. This cloud model may include at least five characteristics, at least three service models, and at least four deployment models.
[0032] The characteristics are as follows:
[0033] On-demand self-service: Cloud consumers can unilaterally provision computing power, such as server time and network storage, automatically as needed, without requiring human interaction with service providers.
[0034] Broad network access: Capabilities are available over the network and access to them is provided through standard mechanisms that facilitate use by heterogeneous thin client platforms or thick client platforms (e.g., mobile phones, laptops, and PDAs).
[0035] Resource pooling: A provider's computing resources are pooled to serve multiple consumers using a multi-tenant model, with different physical and virtual resources dynamically allocated and reallocated according to demand. Consumers generally have no control over or awareness of the exact location of the resources provided, but they have a sense of location independence in that they may be able to specify the location at a higher level of abstraction (e.g., country, state, or data center).
[0036] Rapid Flexibility: Capabilities can be provisioned quickly and flexibly, sometimes automatically, and scale out instantly, or be rapidly released and scale in instantly. To consumers, the available capacity for provisioning often appears unlimited, and any amount can be purchased at any time.
[0037] Measurement Services: Cloud systems automatically control and optimize resource usage by leveraging measurement capabilities at a level of abstraction appropriate to the type of service (e.g., storage, processing, bandwidth, and active user accounts). Resource usage may be monitored, controlled, and reported, providing transparency to both service providers and consumers.
[0038] The service model is as follows:
[0039] Software as a Service (SaaS): The capability offered to consumers is the use of a provider's applications running on cloud infrastructure. These applications are accessible from various client devices via thin client interfaces, such as web browsers (e.g., web-based email). Consumers do not manage or control the underlying cloud infrastructure, including networks, servers, operating systems, storage, or even individual application capabilities, with the exception of limited, user-specific application configuration settings.
[0040] Platform as a Service (PaaS): The capability offered to consumers is the ability to deploy applications they have created or acquired, written using programming languages and tools supported by the provider, onto a cloud infrastructure. Consumers do not manage or control the underlying cloud infrastructure, including networks, servers, operating systems, or storage, but they do control the deployed applications and, in some cases, the configuration of the application hosting environment.
[0041] Infrastructure as a Service (IaaS): The ability provided to consumers is to provision processing, storage, networking, and other fundamental computing resources, allowing consumers to deploy and run any software that may include operating systems and applications. Consumers do not manage or control the underlying cloud infrastructure, but they do control the operating system, storage, and deployed applications, and in some cases have limited control over selected networking components (e.g., host firewalls).
[0042] The deployment model is as follows:
[0043] Private Cloud: A cloud infrastructure operated for a single organization only. It may be managed by the organization or a third party, and may reside on-premises or off-premises.
[0044] Community Cloud: A cloud infrastructure shared by several organizations to support a specific community with common interests (e.g., mission, security requirements, policies, and compliance considerations). It may be managed by the organization or a third party and may reside on-premises or off-premises.
[0045] Public cloud: Cloud infrastructure is made available to the general public or large industry groups and is owned by an organization that sells cloud services.
[0046] Hybrid Cloud: A cloud infrastructure consisting of two or more clouds (private, community, or public) that remain distinct entities but are coupled together by standardized or proprietary technologies that enable data and application portability (e.g., cloud bursting for load balancing across clouds).
[0047] Cloud computing environments are service-oriented, emphasizing statelessness, low coupling, modularity, and semantic interoperability. At the heart of cloud computing is an infrastructure with a network of interconnected nodes.
[0048] Referring here to Figure 1, a schematic diagram of an example cloud computing node is shown. Cloud computing node 10 is merely one example of a suitable cloud computing node and is not intended to imply any limitation on the scope of use or functionality of the embodiments of the present invention described herein. Nevertheless, cloud computing node 10 can be implemented to perform any of the functions described above, or both.
[0049] The cloud computing node 10 has a computer system / server 12 that operates with a number of other general-purpose or dedicated computing system environments or configurations. Examples of well-known computing systems, environments, or configurations, or combinations thereof, that may be suitable for use with the computer system / server 12 include, but are not limited to, personal computer systems, server computer systems, thin clients, thick clients, handheld or laptop devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputer systems, mainframe computer systems, and distributed cloud computing environments that include any of the above systems or devices.
[0050] The computer system / server 12 may be described in the general context of computer system executable instructions, such as program modules, that are executed by the computer system. Generally, a program module may include routines, programs, objects, components, logic, data structures, etc., that perform a specific task or implement a specific abstract data type. The computer system / server 12 may be implemented in a distributed cloud computing environment, where tasks are executed by remote processing devices linked over a communication network. In a distributed cloud computing environment, program modules may reside on both local computer system storage media, including memory storage devices, and remote computer system storage media.
[0051] As shown in Figure 1, the computer system / server 12 of the cloud computing node 10 is shown in the form of a general-purpose computing device. The components of the computer system / server 12 may include, but are not limited to, one or more processors or processing units 16, system memory 28, and a bus 18 that connects various system components, including the system memory 28, to the processor 16.
[0052] Bus 18 represents one or more of any several types of bus structures, including memory buses or memory controllers, peripheral buses, accelerated graphics ports, and processor buses or local buses using any various bus architectures. Examples, but not limited to, such architectures include Industry Standard Architecture (ISA) buses, Microchannel Architecture (MCA) buses, Extended ISA (EISA) buses, Video Electronics Standards Association (VESA) local buses, and Peripheral Component Interconnect (PCI) buses.
[0053] The computer system / server 12 typically includes various computer system-readable media. Such media may be any available media accessible by the computer system / server 12, and may include both volatile and non-volatile media, and both removable and non-removable media.
[0054] The system memory 28 may include computer system-readable media in the form of volatile memory, such as random access memory (RAM) 30 or cache memory 32 or both. The computer system / server 12 may further include other removable / non-removable, volatile / non-volatile computer system storage media. For example, a storage system 34 may be provided for reading from and writing to a non-removable, non-volatile magnetic medium (not shown, but typically called a “hard drive”). Not shown, a magnetic disk drive may be provided for reading from and writing to a removable, non-volatile magnetic disk (e.g., a “floppy disk”), and an optical disk drive may be provided for reading from or writing to a removable, non-volatile optical disk, such as a CD-ROM, DVD-ROM, or other optical medium. In such cases, each may be connected to the bus 18 by one or more data medium interfaces. As further illustrated and described below, the system memory 28 may include at least one program product having a set of program modules (e.g., at least one of which) configured to perform the functions of embodiments of the present invention.
[0055] A program / utility 40 having a set of program modules 42 (at least one of them) may be stored in system memory 28, as well as an operating system, one or more application programs, other program modules, and program data. Each of these operating systems, one or more application programs, other program modules, and program data, or any combination thereof, may include an implementation of a networking environment. The program modules 42 generally perform the functions or methodologies, or both, of the embodiments of the present invention described herein.
[0056] The computer system / server 12 may also communicate with one or more external devices 14, such as a keyboard, pointing device, or display 24; one or more devices that allow a user to interact with the computer system / server 12; or any device that allows the computer system / server 12 to communicate with one or more other computing devices (e.g., a network card, modem, etc.); or a combination thereof. Such communication may occur via an input / output (I / O) interface 22. Furthermore, the computer system / server 12 may communicate with one or more networks, such as a local area network (LAN), a general wide area network (WAN), or a public network (e.g., the Internet), or a combination thereof, via a network adapter 20. As shown, the network adapter 20 communicates with other components of the computer system / server 12 via a bus 18. It should be understood that other hardware and / or software components, or both, that are not shown may be used in conjunction with the computer system / server 12. Examples include, but are not limited to, microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data archive storage systems.
[0057] Referring here to Figure 2, an exemplary cloud computing environment 50 is shown. As shown, the cloud computing environment 50 includes one or more cloud computing nodes 10 that can communicate with local computing devices used by cloud consumers, such as personal digital assistants (PDAs), or cellular phones 54A, desktop computers 54B, laptop computers 54C, or automotive computer systems 54N, or a combination thereof. The nodes 10 may communicate with each other. They may be grouped physically or virtually in one or more networks, such as the private cloud, community cloud, public cloud, or hybrid cloud, or a combination thereof, as described above (not shown). This enables the cloud computing environment 50 to provide infrastructure, platforms, or software, or a combination thereof, as a service, without requiring cloud consumers to maintain resources on their local computing devices. The types of computing devices 54A-N shown in Figure 2 are intended to be illustrative only, and it should be understood that the computing node 10 and the cloud computing environment 50 can communicate with any type of computerized device via any type of network or network addressable connection or both (for example, using a web browser).
[0058] Referring now to Figure 3, the set of function abstraction layers provided by the cloud computing environment 50 (Figure 2) is shown. The components, layers, and functions shown in Figure 3 are intended to be illustrative only, and it should be understood in advance that embodiments of the present invention are not limited thereto. As shown, the following layers and corresponding functions are provided:
[0059] The device layer 55 includes physical or virtual devices, or both, embedded or standalone electronic devices, sensors, actuators, and other objects for performing various tasks in the cloud computing environment 50. Each device in the device layer 55 incorporates networking capabilities to other functional abstraction layers, thereby providing them with information acquired from the device, or information from other abstraction layers, or both. In one embodiment, the various devices including the device layer 55 may incorporate a network of entities collectively referred to as the “Internet of Things” (IoT). As those skilled in the art will understand, such a network of entities enables the communication, collection, and dissemination of data to achieve a wide variety of purposes.
[0060] The device layer 55 shown includes, as shown, a sensor 52, an actuator 53, a “learning” thermostat 56 integrating processing, sensor, and network-forming electronics, a camera 57, a controllable household outlet / receptacle 58, and a controllable electrical switch 59. Other possible devices may include, but are not limited to, various additional sensor devices, network-forming devices, electronic devices (e.g., remote control devices), additional actuator devices, so-called “smart” appliances, such as refrigerators or washer / dryers, and a variety of other possible interconnected objects.
[0061] The hardware and software layer 60 includes hardware components and software components. Examples of hardware components include a mainframe 61, a RISC (Reduced Instruction Set Computer) architecture-based server 62, a server 63, a blade server 64, a storage device 65, and network and network-forming components 66. In some embodiments, the software components include network application server software 67 and database software 68.
[0062] The virtualization layer 70 provides an abstraction layer, from which examples of the following virtual entities can be provided: virtual servers 71, virtual storage 72, virtual networks 73 including virtual private networks, virtual applications and operating systems 74, and virtual clients 75.
[0063] In one example, the management layer 80 may provide the following functions: Resource provisioning 81 provides dynamic procurement of computing resources and other resources used to perform tasks within the cloud computing environment. Measurement and pricing 82 provides cost tracking as resources are used within the cloud computing environment and billing or invoicing for the consumption of these resources. In one example, these resources may include licenses for application software. Security provides identity verification of cloud consumers and tasks, as well as protection of data and other resources. User portal 83 provides consumers and system administrators with access to the cloud computing environment. Service level management 84 provides allocation and management of cloud computing resources to ensure that the required service levels are met. Service level agreement (SLA) planning and execution 85 provides proactive arrangement and procurement of cloud computing resources for which future requirements are anticipated, in accordance with the SLA.
[0064] The workload layer 90 provides examples of functions that can take advantage of the cloud computing environment. Examples of workloads and functions that may be provided from this layer include mapping and navigation 91, software development and lifecycle management 92, provision of virtual classroom education 93, data analysis processing 94, transaction processing 95, and, in the context of the shown embodiments of the invention, various workloads and functions 96 for identifying regulatory data and correlating it with actionable rules. Furthermore, the workloads and functions 96 for identifying regulatory data and correlating it with actionable rules may include operations such as analysis, entity and obligation analysis, and, as further described, user and device management functions. Those skilled in the art will understand that the workloads and functions 96 for identifying regulatory data and correlating it with actionable rules may work in conjunction with other parts of various abstraction layers, e.g., hardware and software 60, virtualization 70, other parts of management 80, and other workloads 90 (e.g., data analysis processing 94) to achieve various objectives of the shown embodiments of the invention.
[0065] Figure 4 shows a block diagram of an exemplary function 400 relating to identifying regulatory data and correlating it with executable rules. As shown, the various blocks of the function are indicated with arrows to show the relationships between the blocks 400 and illustrate the process flow. Furthermore, descriptive information associated with each of the function blocks 400 is also seen. As can be seen, many of the function blocks may also be considered "modules" of the function in the same descriptive sense as previously described in Figures 1 to 3. With the above in mind, module blocks 400 may also be incorporated into various hardware and software components of the determination method and system for feature extraction and summarization according to the present invention, such as those described in Figures 1 to 3. Many of the function blocks 400 may run as background processes on various components, either within a distributed computing component, on a user device, or elsewhere.
[0066] Multiple data sources 401–403 may be provided by one or more content contributors. Data sources 401–403 may be provided as a corpus or group of data sources that are defined, identified, or both. Data sources 401–403 may include, but are not limited to, data sources relating to one or more documents, emails, books, scientific papers, online journals, journals, articles, manuscripts, audio data, video data, or various other documents or data sources that can be published, displayed, interpreted, transcribed, or converted into text data, or combinations thereof. Data sources 401–403 may all be of the same type, for example, wiki pages or articles, or blog pages. Alternatively, data sources 401–403 may be of different types, such as Word documents, wikis, web pages, PowerPoint presentations, printable document formats, or any documents that can be parsed by a natural language processing system.
[0067] In addition to text-based documents, other data sources such as audio, video, and image sources may also be used, and these audio, video, or image sources may undergo pre-analysis, such as audio-to-text conversion, image analysis, or both, in order to extract or transcribe their content for natural language processing. For example, voice commands uttered by content contributors may be detected by a voice-activated detection device 404, and each voice command or communication may be recorded. The recorded voice commands / communications may then be transcribed into text data for natural language processing. As an additional example, one or more of the data sources 401-403 may be audio, video, or both capture devices (e.g., cameras with microphones) that record audio, video, or a combination thereof in an online seminar or meeting with a camera set up in the room, for example, to broadcast the meeting to a remote location where contributors of various intellectual property content may collaborate remotely. Video data captured by the video capture device may be analyzed and transcribed into image or text data for natural language processing.
[0068] The data sources 401-403 are consumed by regulatory data correlation systems, such as regulatory data correlation system 430, using natural language processing (NLP) and artificial intelligence (AI) to provide processed content.
[0069] To display the information in a more user-friendly format, or to provide the information in a more searchable format, or both, data sources 401-403 may be parsed by NLP component 410 (and optionally transcript component 439) to data-mined or transcribe relevant information from the content of data sources 401-403 (e.g., documents, emails, reports, notes, audio recordings, video recordings, live streaming communications, etc.). NLP component 410 may be provided as a cloud service or a local service.
[0070] The regulatory data correlation system 430 may include an NLP component 410, a content consumption component 411, a characteristic association component 412, and a post-processing component. The NLP component 410 may be associated with the consumption component 411. The content consumption component 411 is used to input data sources 401-403 and to operate NLP and AI tools on them, and may learn the content, for example, by using a machine learning component 438. Note that the other components in Figure 4 may also use one or more NLP systems, and the NLP component 410 is shown merely as an example of how an NLP system can be used. As the NLP component 410 (including the machine learning component 438) learns different sets of data, the characteristic association component 412 (or "intelligent characteristic association component") may create associations or links between data sources 401-403 by using artificial intelligence to determine common concepts, methods, features, similar characteristics, or underlying common topics, or combinations thereof.
[0071] "Intelligence" or "cognition" is the mental process of knowing, including aspects such as recognition, perception, reasoning, and judgment. The AI system uses artificial reasoning to interpret data sources 401–403 and extract their topics, ideas, or concepts. Learned decisions, decision factors, alternative decisions, alternative options / choices, decision criteria, concepts, implications, topics, and subtopics of the subject domain, obligations, regulations, laws, policies, legal texts, or other content do not need to be specifically named or mentioned in data sources 401–403, and are derived or inferred by the AI interpretation.
[0072] The learned content of data sources consumed by the NLP system, along with the learned concepts, methods, or features of data sources 401-403, or a combination thereof, is integrated into database 420 (or knowledge store or both), or other data storage methods for the consumed content, to provide associations with the content that referenced the original data sources 401-403.
[0073] Database 420 may operate in conjunction with the transcript component 439 included in the regulatory data correlation system 430 to maintain a timestamped record of all dialogues and submissions for each content contributor, decision, alternative, standard, subject, topic, or idea. Database 420 may record and maintain the progress of decisions, obligations, regulations, laws, policies, legal texts, alternatives, standards, subjects, topics, ideas, or content discussed in data sources 401-403. For example, the transcript component 439 may be used to transcribe various types of data from data sources 401-403, such as audio data or image / video data. For example, voice commands / communications captured by a voice-activated detection device may be transcribed into text data for natural language processing by the transcript component 439. As an additional example, video data captured by a video capture device may be analyzed by the transcript component 439 and transcribed into text data for natural language processing.
[0074] Database 420 may track, identify, and associate all communication threads, messages, transcripts, etc., of all data generated during all stages of the development or “lifecycle” of decisions, obligations, regulations, laws, policies, legal texts, alternatives, standards, subjects, topics, or ideas. By integrating the data into a single database 420 (which may include domain knowledge), the regulatory data correlation system 430 can function like a search engine, but the regulatory data correlation system 430 uses an AI method that creates cognitive associations between data sources using inferred concepts rather than keyword lookup.
[0075] The regulatory data correlation system 430 may include a user interface ("UI") component 434 (e.g., an interactive graphical user interface "GUI") that provides user interaction with indexed content for mining and navigation, or receives one or more inputs / queries from the user, or both. More specifically, the user interface component 434 may communicate with a wireless communication device 455 (see also PDA or cellular phone 54A, desktop computer 54B, laptop computer 54C, or automotive computer system 54N, or combination thereof, in Figure 2) to provide user input for inputting data such as data sources 401-403, or to provide user interaction with summaries of decision factors, alternatives, or criteria, or combinations thereof. The wireless communication device 455 may use the UI component 434 (e.g., GUI) to provide data input, or to provide querying functionality such as interactive GUI functionality that allows the user to input queries for GUI 422 regarding a subject domain, topic, decision, alternative, criteria, decision summary, or associated purpose, or combination thereof, or both. For example, GUI422 may display data related to identifying regulatory data and correlating it with actionable rules.
[0076] The regulatory data correlation system 430 may also include an identification component 432. The identification component 432 may use data extracted directly from one or more data sources, or it may use data stored in the database 420 (or multiple immutable ledgers). The identification component 432 may identify segments, sentences, phrases, paragraphs, and topics relating to one or more decisions, identify each decision element and each criterion of one or more decisions, or identify and extract criteria and one or more alternative implications relating to one or more decisions, or a combination thereof.
[0077] The regulatory data correlation system 430 may also include a matching component 435, a mapping component 436, and a filtering component 437.
[0078] The regulatory data correlation system 430 may use the alignment component 435 to align, group, cluster, or organize, or combine, decisions, obligations, regulations, laws, policies, legal texts, alternatives, standards, subjects, topics, ideas, or content according to similar decisions, obligations, regulations, laws, policies, legal texts, alternatives, standards, subjects, topics, ideas, or content. Alignment component 435 may align, group, cluster, organize, or combine obligations, regulations, laws, policies, legal texts, alternatives, standards, subjects, topics, ideas, or content based on context, similar intentions, similar concepts, similar obligations, similar regulations, similar laws, similar policies, similar legal texts, similar alternatives, similar subjects, similar topics, similar ideas, similar content, or communication timestamps (e.g., audio / video data, text data or both, with timestamps indicating that communications took place at the same time, such as video, audio, memos, or text data of a meeting that took place at a selected time, or a combination thereof), or a combination thereof. Alignment component 435 may track the progress of ideas, topics / subtopics, decisions, decision factors, alternatives, standards, obligations, regulations, laws, policies, legal texts, or content, or a combination thereof, that may be discussed in documents or records in database 420 (e.g., from the start to the end of the Legislative Assembly).
[0079] The regulatory data correlation system 430 may include a mapping component 436. The mapping component may map topics / subtopics, decisions, decision factors, alternatives, standards, obligations, regulations, laws, policies, legal texts, or content, or combinations thereof, with decisions, implications, obligations, regulations, laws, policies, legal texts, and choices that are similar, alternative, or both.
[0080] The alignment component 435 associates semantic data extracted from laws, policies, regulations, or combinations thereof with text data from one or more data sources that describe at least a portion of those laws, policies, regulations, or combinations thereof.
[0081] In one embodiment, once the NLP component 410 performs data linking, the identification component 432 may mine associated concepts, topics, obligations, regulations, laws, policies, legal texts, or similar characteristics from the consumed content database 420 to identify regulatory data and correlate them with actionable rules.
[0082] The regulatory data correlation system 430 may also include a filtering component 437 for filtering decisions, decision factors, alternative decisions, alternative implications, alternative choices, criteria, obligations, regulations, laws, policies, legal texts, or summaries of multiple decision factors, or combinations thereof, into domain knowledge, which may be contained in, associated with, or both of, the database 420.
[0083] In some embodiments, the regulatory data correlation system 430 may use a matching component 435, a filtering component 437, and a machine learning component 438 to extract one or more entities from text data and, based on one or more entities, identify a portion of the semantic data as candidate semantic data.
[0084] In some embodiments, the regulatory data correlation system 430 may use a matching component 435, a filtering component 437, and a machine learning component 438 to extract one or more logical structures from text data and identify a portion of the semantic data as candidate semantic data by comparing one or more logical structures with semantic data.
[0085] In some embodiments, the regulatory data correlation system 430 may use a matching component 435, a filtering component 437, and a machine learning component 438 to generate clusters of entities identified from one or more segments of text data.
[0086] In some embodiments, the regulatory data correlation system 430 may use a matching component 435, a filtering component 437, and a machine learning component 438 to assign a match score between semantic data and text data, which indicates the degree of correspondence between the semantic data and text data.
[0087] In some embodiments, the regulatory data correlation system 430 may use a matching component 435, a filtering component 437, and a machine learning component 438 to transform semantic data into a set of named entities, each of which is assigned a ranking score; to identify selected portions of text data based on the ranking score assigned to each named entity; and to assign confidence scores to the selected portions of text data indicating the degree of confidence that the selected portions of text data match the text data.
[0088] In some embodiments, the regulatory data correlation system 430 may initialize a machine learning component 438 to associate semantic data of laws, policies, regulations, or combinations thereof with text data using machine learning operations.
[0089] The machine learning component 438 may assign a score indicating the degree of relevance based entirely on unsupervised learning. A given score consists of two parts: P(r|t) + P(t|r), where the first part P(r|t) is the likelihood of rule r given text t, and the second part P(t|r) is the likelihood of text t given rule r. The machine learning model used to estimate these likelihood values is a pre-trained language model, such as a bidirectional encoded representation by transformers ("BERT"), which is then fine-tuned to maximize a given likelihood objective.
[0090] The machine learning component 438 may apply heuristic and / or machine learning-based models using a combination of various methods, such as supervised learning, unsupervised learning, delayed learning, and reinforcement learning. Some non-exclusive examples of supervised learning that may be used in this technique include AODE (average one-dependence estimators), artificial neural networks, Bayesian statistics, naive Bayesian classifiers, Bayesian networks, case-based inference, decision trees, inductive logic programming, Gaussian process regression, gene expression programming, and group method of data handling. This includes handling: GMDH), learning automata, learning vector quantization, minimum message length (decision trees, decision graphs, etc.), lazy learning, example-based learning, nearest neighbors, analogical modeling, probabilistic and approximately correct (PAC) learning, ripple-down rules, knowledge acquisition methodologies, symbolic machine learning algorithms, subsymbolic machine learning algorithms, support vector machines, random forests, classifier ensembles, bootstrap aggregation (bagging), boosting (meta-algorithm), ordinal classification, regression analysis, information fuzzy networks (IFN), statistical classification, linear classifiers, Fisher's linear discriminant, logical regression, perceptrons, quadratic classifiers, K nearest neighbors, hidden Markov models, and boosting. Some non-exclusive examples of unsupervised learning that may be used in this technology include artificial neural networks, data clustering, expectation maximization, self-organizing maps, radial basis function networks, vector quantization, generative topographic maps, information bottleneck methods, IBSEAD (distributed autonomous entity systems based interaction), correlation rule learning, a priori algorithms, éclat algorithms, FP growth algorithms, hierarchical clustering, simply connected clustering, conceptual clustering, partitional clustering, K-means algorithms, fuzzy clustering, and reinforcement learning.Some non-exclusive examples of delayed learning may include Q-learning and learning automata. Specific details relating to any supervised, unsupervised, delayed learning, or other machine learning examples described in this paragraph are known and are considered to be within the scope of this disclosure.
[0091] In one embodiment, domain knowledge may be an ontology of concepts representing a domain of knowledge. A thesaurus or ontology may be used as domain knowledge and may also be used to identify semantic relationships between observed variables, between unobserved variables, or both. In one embodiment, the term “domain” is a term intended to have its ordinary meaning. Furthermore, the term “domain” may include an area of expertise concerning a set of systems, or materials, information, content, or other resources, or combinations thereof, relating to a particular subject. A domain may refer to information concerning any particular subject or a selected combination of subjects.
[0092] The term "ontology" is also intended to have its usual meaning. In one aspect, the term ontology, in its broadest sense, may include anything that can be modeled as an ontology, including but not limited to classifications, thesauruses, and vocabulary. For example, an ontology may include information or content related to the content of a domain or a particular class or concept. An ontology can be continuously updated with information synchronized with a source, and information from the source can be added to the ontology as models, attributes of models, or associations between models within the ontology.
[0093] Furthermore, domain knowledge may include links to one or more external resources, such as one or more internet domains, web pages, etc. For example, text data may be hyperlinked to web pages that may contain, explain, or provide additional information about that text data. Thus, the summary may be extended through links to external resources that further explain, indicate, illustrate, or provide contextual or additional information, or both, to support decisions, alternative implications, alternative choices, standards, obligations, regulations, laws, policies, legal texts, alternatives, criteria, subjects, topics, ideas, or content, or any combination thereof.
[0094] In one embodiment, the regulatory data correlation system 430 may perform one or more different types of calculations or calculations. The calculation or calculation operations may be performed using a function that may include a variety of mathematical operations, or one or more mathematical operations (for example, solving differential or partial differential equations by analysis or computation, by finding minimum, maximum, or similar thresholds of a combinational variable using addition, subtraction, division, multiplication, standard deviation, mean, average, percentage, statistical modeling using statistical distributions). Note that each component of the regulatory data correlation system 430 may be an individual component of the regulatory data correlation system 430, a separate component, or both.
[0095] For further explanation, Figure 5 is a block diagram illustrating exemplary operations for identifying regulatory data and correlating them with actionable rules, enabling embodiments of the present invention. In one embodiment, one or more of the components, modules, services, applications, or functions, or combinations thereof, described in Figures 1 to 4 may be used in Figure 5. For example, the computer system / server 12 of Figure 1, incorporating the processing unit 16, may be used to perform the various computational data processing and other functions described in Figure 5.
[0096] For example, policy documents 510 (e.g., policy documents from one or more data sources) and rule documents 520 (e.g., a set of rules, laws, or regulations from one or more data sources) may be parsed, and text data (e.g., T1, T2, and T n Entities such as "T" may be included. One or more segments of text data may be extracted from policy document 510 and legal document 520. Each entity within the regulatory text (e.g., legal document 402) may be identified and extracted.
[0097] In other words, policy document 510 and Expressed in a formal / executable format (including metadata about variables within the rule expression) Extracted from rule document 520, Te A list of short paragraphs may be extracted and compiled.
[0098] Each rule may be aligned, correlated, and associated with one or more relevant text paragraphs from policy document 510 that best represent part or all of each rule from rule document 520.
[0099] Furthermore, each named entity from the text data may be identified and extracted from the policy document 510 and the rule document 520, and these named entities may be used to: From rule document 520, It has entities that have significant overlap with the text from policy document 510. weather Supplemental rules are retrieved.
[0100] In some embodiments, each logical structure may be extracted from the text of the policy document 510 and the rule document 520, and this logical structure is used to compare with the rule logical structure from the policy document 510 to obtain candidate rules with similar structures from the rule document 520. Using the extracted entities and logical structures, rules, paragraphs, and candidate rules / paragraphs for each paragraph / rule may be used as input, and all rules extracted from the rule document 520 (e.g., r1, r2, r n ) and paragraphs extracted from policy document 510, for example, T1, T2, and T n Matching score table 530 is output between these and others.
[0101] To determine a match score (e.g., paragraph-rule match score) for correlating semantic data (e.g., paragraphs of policy text) with rule data, the determination may be performed as follows: In one implementation, given a set of rules (e.g., legal rules, policies, regulations, etc.), R = r1, r2, ..., r n and a set of texts T=T1, T2, ..., T n For example, for a text paragraph of legal policy text, identify the most optimal, best, or closest match between the rule and the text and Special The rule R determines the best matching text T, as shown in Match Score Table 530.
[0102] Among the paragraph / segment data (e.g., "Best Start Tax Credit"), the data with the closest match or "optimal" or "best" match score, e.g., T1, T2, and T nThese may be learned using machine learning operations. For example, a machine learning operation learns and identifies that paragraph T2 from policy document 510 has the most "optimal" or best match score correlated with r2 extracted from rule document 520, but has the lowest match score correlated with r1 extracted from rule document 520. Also, as just an example, a machine learning operation learns and identifies that paragraph Tn from policy document 510 has the most "optimal" or best match score correlated with r1 extracted from rule document 520, but has the lowest match score correlated with r2 extracted from rule document 520.
[0103] In this way, machine learning operations are used to align the formal rules from rule document 520 with the text referenced from policy document 510, and to understand the degree to which each rule depends on other rules, evidence, or input data, code and rate tables, time and eligibility constraints, or decisions, or combinations thereof.
[0104] In other words, machine learning operations are used to align formal rules from rule document 520 with text referenced from policy document 510, and to understand the degree to which each rule depends on other decisions, obligations, regulations, laws, policies, legal texts, alternatives, standards, subjects, topics, ideas, or content (including alternatives or similar ones).
[0105] Referring to FIG. 6 here, there is shown a block / flow diagram 600 for identifying regulatory data and correlating it with executable rules to implement aspects of the present invention. In one aspect, one or more of the components, modules, services, applications, or functions described in FIGS. 1-5, or combinations thereof, may be used in FIG. 6. For example, the computer system / server 12 of FIG. 1 incorporating the processing unit 16 may be used to perform various computational data processing and other functions described in FIG. 5.
[0106] In one implementation, in block 610, entity-based candidates may be filtered. For example, to filter entity-based candidates, in step 1) rule R may be transformed into a set S = {E1, E2, ··· E n} of unique expressions. Each unique expression S may be associated with a score (e.g., a real number such as in the range [0, 1]) indicating its level or importance (ranking). The set S of entities may be used to re-identify one or more paragraphs of text from a policy document that may correspond to rule R. Each paragraph P of text is associated with a score (e.g., a real number in the range [0, 1]) indicating the confidence that P corresponds to rule R.
[0107] In some implementations, as in block 620, the logical structure of rule R may be derived from the set S of entities. As in block 620, paragraph candidates (C) (e.g., C(r1), C(r2), and C(r n )) and rule candidates (C) (e.g., C(T1), C(T2), and C(T n )) may be derived from the set S of entities. As shown in block 630, an unsupervised forward-inverse transformation operation is performed to obtain a paragraph-rule match score (e.g., M(t j ,r1)=0.5*P(T j |r j )+0.5*P(r j |Tj )) may be generated.
[0108] Thus, candidate filtering operations 610 and 620 are used to create finer matching candidates for each rule or paragraph, and the unsupervised forward-inverse transform operation is used to obtain the matching score M(t) via the machine learning operation. j An estimation of r1) is performed.
[0109] In stage 1 of block 610, Rules Set of named entities to To perform the conversion, the following operations may be performed. In some implementations, rule R may be converted into a set of named entities. Rule R may be, for example, It may have a text description, One or more operations In the text description, Extract named entities do The rules may have conditions, and the variables may have "expressive" names, such as "isCitizenOf" (citizens of), "hasAge" (age is), "person.name" (person's name), "personAddress" (person's address), "income" (income), etc., and the names of the conditions / variables may be transformed into entities by mapping them to one or more standard vocabulary or knowledge graphs or both. The rule R may have variables bound to data / values, and named entities may be inferred using the data / values.
[0110] Furthermore, the data may be in tabular format, and one or more operations may be used to retrieve the meaning of the table, such as tagging columns, based on their values. In other implementations, the data may be in XML / JSON / YAML form or other “descriptive” format. Thus, data formats such as XML tags or JSON / YAML may be converted to attribute names to named entities (similar to the names of “expressive” conditions / variables in Rule R). Also, the meaning (e.g., tags) may be retrieved using the values of XML tags or JSON / YAML values.
[0111] In other implementations, named entity extensions (by habitual / standard vocabulary or knowledge graphs or both) may be used to extend the set of entities. For example, wikiification or other semantic matching or both, and query extension operations may be used. To associate scores with named entities, one or more types of operations may be used, such as frequency scores, named entity extractor confidence scores, or operations used to extract meaning from data / values, or a combination of these scores.
[0112] In some implementations, the following operations may be performed to map the set of named entities in stage 2 to one or more paragraphs of text. In some implementations, the set of named entities may be mapped to one or more paragraphs of text. For example, the set of entities S may be the input set of named entities (derived from rule R). Assuming paragraphs of text, a set of paragraphs of the text P of named entities may be extracted from the paragraphs.
[0113] A similarity metric between S and P is often determined (for example, using naive metrics, including Jackard similarity between S and P, or the F1 score between S (ground truth) and P (prediction)), where S is the set of entities extracted from rule R and P is the set of entities extracted from text T. Both F1 and Jackard similarity are two metrics that indicate how similar the two sets are.
[0114] The variable S may be converted to a vector VS using a pre-trained embedding model. The variable P of a text paragraph may be converted to a vector VP in the same embedding space. A similarity metric may be determined between VS and VP (e.g., cosine similarity). The pre-trained embedding model may be used, for example, to convert the variable S to a vector VS. An intermediate step may be required, in which the variable S is converted to a sentence (text), and then this sentence is converted to a vector VS using the embedding model. This may be done using named entities of the set of entities S.
[0115] In some implementations, the logical structure of rule R may be derived from a set of entities S, as shown in block 630, because it is known which entities correspond to the conditions or variables of rule R, and the logical connectors of rule R are known. Once pairs of S and P match and candidate paragraphs are identified, a “collective” match / scoring (in other words, matching a set of named entities to all candidate paragraphs) may be applied to identify a subset (one or more) of paragraphs that cumulatively define the rule. This can be implemented by a hierarchical approach, for example, by using a dependency tree or semantic knowledge graph to independently link all S named entities to all P entities, and then examining the “coverage / match” of each paragraph to confirm that the paragraphs add to / contribute to the rule.
[0116] For further explanation, Figure 7 is a diagram 700 showing an exemplary knowledge graph for identifying regulatory data and correlating it with actionable rules (e.g., the logical structure-based candidate filter 620 in Figure 6) that enable embodiments of the present invention. In one embodiment, one or more of the components, modules, services, applications, or functions, or combinations thereof, described in Figures 1 to 6 may be used in Figure 7.
[0117] As shown, in order to filter logical structure-based candidates (block 620 in Figure 6), rule R may be transformed into a knowledge graph tree Tr710 representing a logical expression, where nodes may be arithmetic or logical operators such as AND, OR, greater than, less than, addition, and subtraction, and nodes may be variables or constants. For example, the expression age > 19 AND SSN = true is represented as a tree in the knowledge graph tree Tr710.
[0118] Next, a can be transformed into a tree Ts, where each node is either a detected logical operator or a fragment of text. For example, the sentence "To be eligible, participants must be older than 19 years old and have a valid SSN" can be represented as a tree in the knowledge graph tree Ts720.
[0119] Knowledge graph tree Tr710 may be compared to knowledge graph tree Ts720 using a tree similarity metric.
[0120] For further explanation, Figure 8 shows an exemplary operation using forward and inverse transforms to identify regulatory data and correlate it with actionable rules, which enables embodiments of the present invention (e.g., the unsupervised forward-inverse operation 630 in Figure 6). In one embodiment, one or more of the components, modules, services, applications, or functions, or combinations thereof, described in Figures 1 to 6 may be used in Figure 7.
[0121] In some implementations, for each rule R, a set of candidate texts C(r) is identified or determined such that the likelihood (e.g., the probability) of the candidate text C(r) containing the optimal or best match (e.g., the maximized match) is high. For each text T, a set of candidate rule C(t) is identified or determined such that the likelihood (e.g., the probability) of the candidate rule C(t) containing the optimal or best match (e.g., the maximized match) is high.
[0122] In some implementations, language models may be used to estimate probabilities P(T|r) and P(r|T), and the paragraph-rule match score is maximized based on the following formula.
number
[0123] Figure 9 shows a method 900 that uses a processor to identify regulatory data and correlate it with executable rules, which can implement various aspects of the embodiment shown. Function 900 may be implemented as an instruction executed on a machine, the instruction contained in at least one computer-readable medium or at least one non-temporary machine-readable storage medium. Function 900 begins in block 902.
[0124] As shown in block 904, semantic data extracted from laws, policies, regulations, or combinations thereof may be associated with text data from one or more data sources that describe at least a portion of those laws, policies, regulations, or combinations thereof. The text data from one or more data sources may be ingested when processing the text data using lexical analysis, syntactic analysis, concept extraction, semantic analysis, machine learning operations, or a combination thereof. Function 900 may end in block 906.
[0125] In one embodiment, in conjunction with or as part of at least one block in Figure 9, or both, Operation 900 may include each of the following: Operation 900 may take text data from one or more data sources when processing text data using lexical analysis, syntactic analysis, concept extraction, semantic analysis, machine learning operations, or a combination thereof; Operation 900 may identify the scope of compliance required by laws, policies, regulations, or a combination thereof (e.g., legal text).
[0126] Furthermore, the 900 operations may use natural language processing (NLP) to determine semantic and text data that have obligation content. The semantic and text data may be sentences, and the obligation content may be a requirement that the obligation must comply with an obligation, law, policy, regulation, or a combination thereof, or initialize a machine learning mechanism to learn, determine, or identify an obligation from one or more segments of text data, or generate a compliance corpus from training one or more machine learning models to manage regulatory compliance, or a combination thereof. The 900 operations may define an obligation as an action required to comply with a law, policy, regulation, or a combination thereof, a prohibition on an entity's conduct, behavior, or activity, an entity's legal rights, restrictions, or a combination thereof.
[0127] The 900 operation may extract one or more entities from text data and, based on one or more entities, identify a portion of semantic data as candidate semantic data. The 900 operation may extract one or more logical structures from text data and, by comparing one or more logical structures with semantic data, identify a portion of semantic data as candidate semantic data. In other words, one or more segments of text data may be extracted from one or more data sources representing one or more objects that describe compliance corpora or compliance named entities of an organization expected to comply with obligations, laws, policies, regulations, or a combination thereof.
[0128] Operation 900 may generate clusters of entities identified from one or more segments of text data. Operation 900 may assign a matching score between semantic data and text data, which indicates the degree of correspondence between the semantic data and text data.
[0129] Operation 900 may involve transforming semantic data into a set of named entities, each of which is assigned a ranking score; identifying selected portions of text data based on the ranking score assigned to each named entity; and assigning confidence scores to the selected portions of text data, indicating the degree of confidence that they match the text data.
[0130] The 900 operations may initialize machine learning mechanisms to associate semantic data of laws, policies, regulations, or combinations thereof with text data using machine learning operations.
[0131] The present invention may be a system, a method, a computer program product, or a combination thereof. The computer program product may include a computer-readable storage medium having computer-readable program instructions for causing a processor to implement an aspect of the present invention.
[0132] A computer-readable storage medium can be a tangible device capable of holding and storing instructions for use by an instruction execution device. A computer-readable storage medium may be, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination thereof. A non-exhaustive list of more specific examples of computer-readable storage media includes portable computer diskettes, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disk read-only memory (CD-ROM), digital multipurpose disks (DVDs), memory sticks, floppy disks, mechanically encoded devices such as punch cards or grooved raised structures on which instructions are recorded, and any suitable combination thereof. Computer-readable storage media, when used herein, should not be interpreted as transient signals themselves, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses passing through optical fiber cables), or electrical signals transmitted through wires.
[0133] The computer-readable program instructions described herein may be downloaded from a computer-readable storage medium to each computing device / processing device, or they may be downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, or a wireless network, or a combination thereof. The network may include copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers, or edge servers, or a combination thereof. The network adapter card or network interface of each computing device / processing device receives computer-readable program instructions from the network and transfers those computer-readable program instructions for storage in a computer-readable storage medium within each computing device / processing device.
[0134] The computer-readable program instructions for performing the operations of the present invention may be assembler instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, or source code or object code, written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk®, C++, and conventional procedural programming languages such as the C programming language or similar programming languages. The computer-readable program instructions may run entirely on the user's computer as a standalone software package, partially on the user's computer, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter case, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or wide area network (WAN), or a connection to an external computer may be made (for example, via the Internet using an Internet service provider). In some embodiments, an electronic circuit including, for example, a programmable logic circuit, a field-programmable gate array (FPGA), or a programmable logic array (PLA) may execute computer-readable program instructions by personalizing the electronic circuit using state information of computer-readable program instructions in order to perform aspects of the present invention.
[0135] Aspects of the present invention are described herein with reference to flowcharts or block diagrams, or both, of methods, apparatus (systems), and computer program products according to embodiments of the present invention. It will be understood that each block in a flowchart or block diagram, or both, and combinations of blocks in a flowchart or block diagram, or both, can be implemented by computer-readable program instructions.
[0136] These computer-readable program instructions may be provided to a processor of a general-purpose computer, a dedicated computer, or other programmable data processing device to create a machine, thereby creating means for implementing functions / operations explicitly shown in a flowchart or block diagram or both, through which the instructions are executed via the processor of the computer or other programmable data processing device. Furthermore, these computer-readable program instructions may be stored in a computer-readable storage medium capable of directing a computer, a programmable data processing device, or other device, or a combination thereof, to function in a particular manner, thereby comprising a manufactured article containing instructions for implementing functions / operations explicitly shown in a flowchart or block diagram or both.
[0137] Furthermore, computer-readable program instructions may be loaded into a computer, other programmable data processing device, or other device to execute a series of operational steps on that computer, other programmable device, or other device, thereby creating a computer implementation process in which the instructions executed on that computer, other programmable device, or other device implement the functions / operations explicitly shown in the flowchart or block diagram or both.
[0138] The flowcharts and block diagrams in the figures illustrate the architecture, functions, and operations of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or part of an instruction containing one or more executable instructions for implementing an explicit logical function. In some alternative implementations, the functions appended to a block may be performed in an order different from the order appended to the figure. For example, two blocks shown consecutively may actually be executed substantially simultaneously, or multiple blocks may be executed in reverse order depending on the functions they contain. It should also be noted that each block in a block diagram or flowchart, or both, and any combination of blocks in a block diagram or flowchart, or both, can be implemented by a dedicated hardware-based system that performs an explicit function or operation, or implements a combination of dedicated hardware and computer instructions.
Claims
1. A method for correlating regulatory data with rules in a computer-executable code format using a processor in a computing environment, The steps include associating the aforementioned rule with one or more text paragraphs extracted from a policy document containing at least a portion of the aforementioned rule, A step of extracting one or more logical structures from one or more text paragraphs, wherein the one or more logical structures extracted from the one or more text paragraphs include logical operators or text fragments. A step of extracting one or more logical structures from the aforementioned rules, wherein the one or more logical structures extracted from the rules include at least one of arithmetic operators, logical operators, variables, and constants. A step in which a portion of one or more text paragraphs is identified as a candidate text paragraph by comparing one or more logical structures extracted from one or more text paragraphs with one or more logical structures extracted from the rule. A method for providing this.
2. The method according to claim 1, further comprising the step of assigning a match score between the identified candidate text paragraph and the rule based on the likelihood of the rule given the candidate text paragraph and the likelihood of the candidate text paragraph given the rule, wherein the match score indicates the degree of correspondence between the candidate text paragraph and the rule.
3. The method according to claim 2, further comprising the step of identifying the candidate text paragraph having the highest match score for the rule as the text paragraph correlated with the rule.
4. The steps include extracting one or more entities from the one or more text paragraphs, A step of identifying a portion of the one or more text paragraphs as candidate text paragraphs based on the one or more entities. The method according to claim 1, further comprising:
5. The step of mapping the aforementioned rule to the candidate text paragraph using a forward-to-back transformation operation. The method according to claim 1, further comprising:
6. The step of converting the aforementioned rule into a set of entities, wherein the set of entities is a candidate set of entities, The steps include: identifying and matching one or more candidate text paragraphs extracted from the policy document using the candidate set of entities of the rule; The method according to claim 1, further comprising:
7. A step of assigning a confidence score to the candidate text paragraph, indicating the degree of confidence that the candidate text paragraph matches the rule. The method according to any one of claims 1 to 6, further comprising:
8. A system for correlating regulatory data with rules in the form of code that can be executed on a computer in a computing environment, One or more processors having executable instructions, When the aforementioned executable instruction is executed, the system, Associating the aforementioned rule with one or more text paragraphs extracted from a policy document containing at least a portion of the aforementioned rule, Extracting one or more logical structures from one or more text paragraphs, wherein the one or more logical structures extracted from the one or more text paragraphs include logical operators or text fragments. Extracting one or more logical structures from the aforementioned rules, wherein the one or more logical structures extracted from the aforementioned rules include at least one of arithmetic operators, logical operators, variables, and constants. By comparing the one or more logical structures extracted from the one or more text paragraphs with the one or more logical structures extracted from the rules, a portion of the one or more text paragraphs is identified as a candidate text paragraph. One or more processors that perform this task A system that includes these features.
9. When the aforementioned executable instruction is executed, the system will, Extracting one or more entities from the aforementioned one or more text paragraphs, Based on the one or more entities, a portion of the one or more text paragraphs is identified as the candidate text paragraph. The system according to claim 8, which causes the following to be performed.
10. When the aforementioned executable instruction is executed, the system will, Mapping the aforementioned rules to the candidate text paragraphs using a forward-to-back transformation operation. The system according to claim 8, which causes the following to be performed.
11. When the aforementioned executable instruction is executed, the system will, Assigning a matching score between the candidate text paragraph and the rule, wherein the matching score indicates the degree of correspondence between the candidate text paragraph and the rule. The system according to claim 8, which causes the following to be performed.
12. When the aforementioned executable instruction is executed, the system will, The process involves converting the aforementioned rule into a set of entities, wherein the set of entities is a candidate set of entities. Using the candidate set of entities in the rule, identify and match one or more candidate text paragraphs extracted from the policy document. The system according to claim 8, which causes the following to be performed.
13. When the aforementioned executable instruction is executed, the system will, Assigning a confidence score to the candidate text paragraph that indicates the degree of confidence that the candidate text paragraph matches the aforementioned rule. A system according to any one of claims 8 to 12, which causes the following to be performed.
14. A computer program for correlating regulatory data with rules in a computer-executable code format in a computing environment, wherein the computer A procedure for associating the aforementioned rule with one or more text paragraphs extracted from a policy document containing at least a portion of the aforementioned rule, A procedure for extracting one or more logical structures from one or more text paragraphs, wherein the one or more logical structures extracted from the one or more text paragraphs include logical operators or text fragments. A procedure for extracting one or more logical structures from the aforementioned rules, wherein the one or more logical structures extracted from the rules include at least one of arithmetic operators, logical operators, variables, and constants. A procedure for identifying a portion of one or more text paragraphs as a candidate text paragraph by comparing one or more logical structures extracted from one or more text paragraphs with one or more logical structures extracted from the rule. A computer program that executes something.
15. To the aforementioned computer, A procedure for extracting one or more entities from the aforementioned one or more text paragraphs, A procedure for identifying a portion of one or more text paragraphs as candidate text paragraphs based on the one or more entities, The computer program according to claim 14, which further performs the following.
16. To the aforementioned computer, A procedure for mapping the aforementioned rule to the candidate text paragraph using a forward-to-back transformation operation. The computer program according to claim 14, which further performs the following.
17. To the aforementioned computer, A procedure for assigning a matching score between a candidate text paragraph and a rule, wherein the matching score indicates the degree of correspondence between the candidate text paragraph and the rule. The computer program according to claim 14, which further performs the following.
18. To the aforementioned computer, A procedure for converting the aforementioned rule into a set of entities, wherein the set of entities is a candidate set of entities, A procedure for identifying and matching one or more candidate text paragraphs extracted from the policy document using the candidate set of entities of the rule, and The computer program according to claim 14, which further performs the following.
19. To the aforementioned computer, A procedure for assigning a confidence score to a candidate text paragraph, indicating the degree of confidence that the candidate text paragraph matches the aforementioned rule. A computer program according to any one of claims 14 to 18, which further performs the following:
Citation Information
Patent Citations
Label printer, label diagnostic system and label diagnostic program
JP2021039743A
System for managing community provided in information processing system, and method thereof
US20090013376A1
Facilitating mapping of control policies to regulatory documents
US20180137107A1
Technologies for dynamically creating representations for regulations
US20210117621A1
Determining syntax parse trees for extracting nested hierarchical structures from text data
US20210319173A1