Method and device for generating a knowledge graph of cross-level data processing activities
Patent Information
- Application Number
- CN202610248270.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-03-02
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2046-03-02
AI Technical Summary
[0003]传统的方法主要依赖领域专家通过人工阅读海量文档,像拼图一样手动建立这些联系,不仅耗时费力、成本高昂,而且难以保证一致性和完整性
[0009]本申请实施例通过分别定义每一层的实体及关系,以及跨层之间的关系,为后续生成实体关系模型提供依据。
Smart Images

Figure CN122088647B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing technology, and more specifically, to a method and apparatus for generating knowledge graphs for cross-level data processing activities. Background Technology
[0002] In current enterprise data management and security compliance practices, information at the three levels—business operations, data manipulation, and technical implementation—often exists in a fragmented manner. Business personnel use process documents and operation manuals to describe "who, where, and what business is being handled"; system logs record "which account, from which IP, and what instructions were executed." There is a lack of direct, structured connection between these two.
[0003] Traditional methods primarily rely on domain experts manually reading massive amounts of documents, like piecing together... Figure 1 Manually establishing these connections is not only time-consuming, labor-intensive, and costly, but also makes it difficult to guarantee consistency and completeness. On the other hand, methods based entirely on machine learning require a large amount of labeled data as learning material, which is often difficult to obtain in the early stages of many business scenarios or in fields involving sensitive information.
[0004] This leads to a core dilemma: we struggle to quickly, accurately, and comprehensively answer questions such as, "Which specific business process does this access to sensitive data correspond to? Which position is responsible?" or "What key instructions from which systems are actually involved in this business process?" The gaps in information make comprehensive data tracing, risk analysis, and compliance auditing exceptionally difficult and inefficient. Summary of the Invention
[0005] The purpose of this application is to provide a method and apparatus for generating knowledge graphs for cross-level data processing activities, so as to establish the relationship between multi-level data and thereby improve the efficiency of data analysis.
[0006] In a first aspect, embodiments of this application provide a method for generating a knowledge graph for cross-level data processing activities, including: Obtain the entities and relationships corresponding to the business logic layer, data operation layer, and technical implementation layer, as well as the cross-layer relationships between entities in adjacent layers; whereby, the business logic layer is used to describe the data processing activities involved in the business logic layer; the data operation layer is used to describe the specific data operations in the business activities; and the technical implementation layer is used to describe the underlying technical execution details of the data operation behavior; Generate entity-relationship models based on entities, relationships, and cross-level relationships; Entity data is extracted from the data processing activity documents corresponding to the business logic layer, data operation layer, and technical implementation layer to obtain entity relationship data. A knowledge graph is generated based on the entity relationship model and entity relationship data.
[0007] This application embodiment achieves full-link correlation analysis through three-layer cross-level entity modeling, forming a structured, traceable, high-quality knowledge graph, providing accurate reference and data support for risk analysis of data processing behavior, and helping to quickly identify abnormal operations and potential risks.
[0008] In one possible implementation of the first aspect, the technical implementation layer is the lowest layer, and the business logic layer is the highest layer; the entities and relationships corresponding to the business logic layer, data operation layer, and technical implementation layer, as well as the cross-layer relationships between entities in adjacent layers, are obtained, including: Obtain the first-level relationships between the first-level entities in the business logic layer, the second-level relationships between the second-level entities in the data operation layer, and the third-level relationships between the third-level entities in the technology implementation layer; and obtain the first cross-layer relationships between the first-level entities and the second-level entities, and the second cross-layer relationships between the second-level entities and the third-level entities.
[0009] This application provides a basis for generating an entity-relationship model by defining the entities and relationships at each layer, as well as the relationships between layers.
[0010] In one possible implementation of the first aspect, an entity-relationship model is generated based on entities, relationships, and cross-layer relationships, including: The entity relationship model is generated by using the first-level entity, the second-level entity, and the third-level entity as vertices, and the first-level relationship, the second-level relationship, the third-level relationship, the first cross-level relationship, and the second cross-level relationship as edges.
[0011] This application embodiment generates an entity relationship model by using the first-layer entity, the second-layer entity, and the third-layer entity as vertices, and the first-layer relationship, the second-layer relationship, the third-layer relationship, the first cross-layer relationship, and the second cross-layer relationship as edges, providing a basis for subsequent mapping of entities in specific documents.
[0012] In one possible implementation of the first aspect, the method further includes: Generate modeling scripts based on entity relationship models; Generate a knowledge graph based on the entity relationship model and entity relationship data, including: Knowledge graphs are generated based on modeling scripts and entity relationship data.
[0013] This application embodiment generates a modeling script from the actual relationship model, providing a foundation for subsequent entity relationship import.
[0014] In one possible implementation of the first aspect, entity data is extracted from the data processing activity documents corresponding to the business logic layer, data operation layer, and technical implementation layer respectively to obtain entity relationship data, including: The entity data is extracted from the business process description class document corresponding to the business logic layer using a large language model. The business entity data includes the specific data corresponding to the first-level entities contained in the business process description class document. Entity data is extracted from the system ledger documents by keyword matching to obtain supplementary entity data. The supplementary entity data includes the specific data corresponding to the second-level entities, the specific data corresponding to the third-level entities, and the relationships between entities. The entity relationship data includes business entity data and supplementary entity data.
[0015] This application embodiment extracts entity data from relevant documents in the business logic layer using a large language model, enabling zero-sample automated extraction and significantly reducing the manpower and time costs of knowledge graph construction.
[0016] In one possible implementation of the first aspect, entity data is extracted from the business process description class document corresponding to the business logic layer using a large language model, including: Input prompt words into the large language model. The prompt words include: extraction task, information extraction rules, data processing activity recognition knowledge base, output specifications and constraints. Entity data is extracted from business process description documents using a large language model based on prompt words.
[0017] In the information extraction stage, this application employs a large language model as the core inference engine. Through a series of creative prompting and rule designs—including a pre-built domain knowledge base, structured output prompts with header specifications, splitting rules, and source tracing constraints, as well as defined field mapping trigger rules—the model is guided to accurately and structurally extract entities and relationships conforming to the pre-defined model from unstructured documents without the need for labeled samples. This method overcomes the strong dependence of traditional algorithms on large-scale labeled data and achieves high-quality transformation from free text to standardized knowledge.
[0018] In one possible implementation of the first aspect, entity data is extracted from the system ledger documents through keyword matching to obtain supplementary entity data, including: By matching preset keywords with system ledger documents, the target fields contained in the system ledger documents can be obtained; Map the target field to the target entities in the first-level, second-level, and third-level entities; Extract the target relationships between target entities; Read the specific data corresponding to the target field from the system ledger document; Supplementary entity data is generated based on the target entity, target relationship, and the specific data corresponding to the target relationship.
[0019] This application embodiment extracts entity relationships from system ledger documents through keyword matching. As a supplement to the business logic layer, the extracted entities and relationships are associated with the entities in the business logic layer to form complete three-layer entity relationship data.
[0020] In one possible implementation of the first aspect, the first-layer entities include environment entities, role entities, information entities, and business activity entities; the second-layer entities include terminal entities, user entities, data entities, and information system entities; and the third-layer entities include terminal identifier entities, user identifier entities, data identifier entities, and system identifier entities.
[0021] This application embodiment predefines entities at each layer based on actual business scenarios, which facilitates the generation of subsequent entity relationship models.
[0022] In one possible implementation of the first aspect, after generating the knowledge graph, the method further includes: Receive query requests; The query request is processed based on the knowledge graph to generate query results.
[0023] In this embodiment of the application, after generating a knowledge graph, all entities related to the query conditions in the business logic layer, data operation layer, and technical implementation layer can be queried based on the knowledge graph.
[0024] Secondly, embodiments of this application provide a knowledge graph generation apparatus for cross-level data processing activities, including: The entity relationship acquisition module is used to acquire the entities and relationships corresponding to the business logic layer, data operation layer, and technical implementation layer, as well as the cross-layer relationships between entities in adjacent layers. Among them, the business logic layer is used to describe the data processing activities involved in the business logic layer; the data operation layer is used to describe the specific data operations in the business activities; and the technical implementation layer is used to describe the underlying technical execution details of the data operation behavior. The model generation module is used to generate entity-relationship models based on entities, relationships, and cross-level relationships. The entity relationship data acquisition module is used to extract entity data from the data processing activity documents corresponding to the business logic layer, data operation layer, and technical implementation layer to obtain entity relationship data. The knowledge graph generation module is used to generate knowledge graphs based on entity relationship models and entity relationship data.
[0025] Thirdly, embodiments of this application provide an electronic device, including: a processor, a memory, and a bus, wherein: The processor and memory communicate with each other via a bus; The memory stores program instructions that can be executed by the processor, and the processor can execute the method of the first aspect by calling the program instructions.
[0026] Fourthly, embodiments of this application provide a non-transitory computer-readable storage medium, comprising: A non-transitory computer-readable storage medium stores computer instructions that cause the computer to perform the methods in the various possible implementations of the first aspect.
[0027] Fifthly, embodiments of this application provide a computer program product, including computer program instructions, which, when read and executed by a processor, perform the methods in various possible implementations of the first aspect.
[0028] Other features and advantages of this application will be set forth in the following description and will be apparent in part from the description or may be learned by practicing embodiments of this application. Attached Figure Description
[0029] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments of this application will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0030] Figure 1 A flowchart for generating a knowledge graph for a cross-level data processing activity is provided in this application embodiment; Figure 2 A schematic diagram of a knowledge graph generation method for cross-level data processing activities provided in this application embodiment; Figure 3 A schematic diagram of a cross-level data processing activity entity relationship model provided in this application embodiment; Figure 4 This is a schematic diagram of another entity relationship model provided in an embodiment of this application; Figure 5 A schematic diagram of a knowledge graph generation device for cross-level data processing activities provided in this application embodiment; Figure 6 This is a schematic diagram of the physical structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0031] The embodiments of the technical solution of this application will now be described in detail with reference to the accompanying drawings. These embodiments are only used to more clearly illustrate the technical solution of this application and are therefore merely examples, and should not be used to limit the scope of protection of this application.
[0032] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains; the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the application; the terms “comprising” and “having”, and any variations thereof, in the specification, claims, and foregoing description of the drawings are intended to cover non-exclusive inclusion.
[0033] In the description of the embodiments of this application, technical terms such as "first" and "second" are used only to distinguish different objects and should not be construed as indicating or implying relative importance or implicitly specifying the number, specific order, or primary and secondary relationship of the indicated technical features. In the description of the embodiments of this application, "multiple" means two or more, unless otherwise explicitly defined.
[0034] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0035] In the description of the embodiments in this application, the term "and / or" is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Additionally, the character " / " in this document generally indicates that the preceding and following related objects have an "or" relationship.
[0036] In the description of the embodiments of this application, the term "multiple" refers to two or more (including two), similarly, "multiple sets" refers to two or more (including two sets), and "multiple pieces" refers to two or more (including two pieces).
[0037] In the description of the embodiments of this application, unless otherwise expressly specified and limited, technical terms such as "installation," "connection," "joining," and "fixing" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral part; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; they can refer to the internal communication of two components or the interaction between two components. For those skilled in the art, the specific meaning of the above terms in the embodiments of this application can be understood according to the specific circumstances.
[0038] To address the problem of fragmented data processing activities across business logic, data operations, and technical implementation, making it difficult to correlate, analyze, and trace, this application constructs a unified entity relationship model with pre-defined cross-layer mapping relationships. Based on field mapping rules, it implements entity expansion and association to solve the core technical bottleneck of semantic fragmentation across layers of information, inability to effectively correlate and trace globally.
[0039] Figure 1 A knowledge graph generation flowchart for cross-level data processing activities provided in this application embodiment may specifically include the following steps: Step (1) Define a cross-layered data processing activity entity relationship model and generate a knowledge graph entity relationship model. In a knowledge graph, vertices are the basic units that carry knowledge, which can represent both concrete entities and abstract concepts. Edges (also known as relationships) are the links that connect different vertices and are used to define and express the semantic relationships between vertices. This method first designs the vertices and edges involved in the three layers of business logic, data operation, and technical implementation.
[0040] Step (2) Utilize the large language model and mapping rules to extract entities and relationships under zero-sample conditions. Use the large language model to extract entity and relationship information from business activity-related documents, define prompt words, and refine them into structured data required for knowledge graph modeling.
[0041] Step (3) imports the relational model and entity relation data into the graph database to generate a knowledge graph. The structured entity relation data is imported into the graph database to construct the knowledge graph. This step, as the implementation link of knowledge graph construction in this technical solution, has the core objective of completing the standardized import and structured organization of data into the graph database based on the cross-level data processing activity entity relation model defined in step (1) and the entity relation data obtained through zero-sample extraction from the large language model in step (2), ultimately generating a knowledge graph that can support the traceability and analysis of data processing activities. This step ensures the integrity, hierarchical relevance, and storage effectiveness of entity relations through a standardized import process and data mapping mechanism.
[0042] The following is a detailed explanation of the three steps mentioned above: Figure 2 This is a schematic flowchart illustrating a method for generating a knowledge graph for cross-level data processing activities, provided in an embodiment of this application. It is understood that the method provided in this embodiment can be applied to a data processing system, which includes a terminal and a server; the terminal can specifically be a smartphone, tablet computer, computer, personal digital assistant (PDA), etc.; the server can specifically be an application server or a web server.
[0043] The knowledge graph generation methods specifically include: Step 201: Obtain the entities and relationships corresponding to the business logic layer, data operation layer, and technical implementation layer, as well as the cross-layer relationships between entities in adjacent layers.
[0044] The business logic layer describes the data processing activities involved in the business logic layer, and its knowledge can serve as an analytical basis for macro-level business management. The data operation layer describes the specific data operations in business activities, focusing on the tracking and management of actual data operation behaviors. Its knowledge can serve as an analytical basis for data practice, such as operation behavior auditing and security incident investigation. The technical implementation layer, located below the data operation layer, describes the underlying technical execution details of data operation behaviors, focusing on the accurate recording of core technical dimensions such as account identity, terminal IP, operation instructions, and target field names. Its knowledge can serve as an analytical basis for technical practice scenarios such as source tracing, operation behavior authenticity verification, and underlying security vulnerability investigation.
[0045] Define the first-level entities and relationships within the business logic layer. First-level entities include vertices such as roles, environments, information, and business activities. First-level relationships include edges such as "located in," "processed," and "used in." Entities in the "role" vertex category are job titles; entities in the "environment" vertex category are office location names; entities in the "information" vertex category are business information names; and entities in the "business activity" vertex category are business names. The "processed" relationship between roles and information can be defined based on data processing concepts in data security, defining data content as collection, storage, use, processing, transmission, provision, disclosure, and deletion. The data content of the "located in" and "used in" relationships is their category names themselves.
[0046] Define the second-level entities and second-level relationships of the data operation layer. The second-level entities include entities such as users, terminals, data, and information systems. The second-level relationships include relationships such as usage, operation, and storage. The entity in the user vertex category is the user name; the entity in the terminal vertex category is the terminal name; the entity in the data vertex category is the data field name; and the entity in the information system vertex category is the information system name. The "operation" relationship between users and data includes common data operation names such as "access, query, filter, fill or enter, import or upload, archive, restore, preview, export or download, print, submit, approve, annotate or note, modify or edit, statistics or calculation, analysis, visualization, modeling, share, transfer or distribute, and delete." The data content of the "use" and "store" relationships is the category name itself.
[0047] This section defines the third-layer entities and relationships within the technical implementation layer. The third-layer entities include user identifiers, terminal identifiers, data identifiers, and system identifiers. The third-layer relationships include usage, operation instructions, and storage relationships. The entity in the user identifier vertex category is the user account; the entity in the terminal identifier vertex category is the terminal IP address; the entity in the data vertex category is the data field name; and the entity in the system identifier vertex category is the information system IP address. The "operation instructions" relationship between user identifiers and data identifiers consists of keyword names in the information system logs for operations such as "access," "query," "filter," "fill," or "enter." These can be mapped and organized according to the information system log description file, for example, "access," "query," "filter," "add," etc. The data content of the "usage" and "storage" relationships is the category name itself.
[0048] Define the cross-layer relationships between entities in the business logic layer and the data operation layer. For example, users "belong" to roles, terminals "are located" in environments, information "contains" data, and information systems "serve" business activities.
[0049] Define the cross-layer relationships between entities in the data operation layer and the technical implementation layer. For example: a user "owns" a user identifier, a terminal "is configured with" a terminal identifier, data "contains" a data identifier, and an information system "is configured with" a system identifier.
[0050] The entity-relationship structure at each level, as well as the entity-relationship structure across levels, are shown in the table below:
[0051] Step 202: Generate an entity-relationship model based on entities, relationships, and cross-layer relationships.
[0052] The entity relationship model is generated by using the first-layer entities, second-layer entities, and third-layer entities as vertices, and the first-layer relationships, second-layer relationships, third-layer relationships, first cross-layer relationships, and second cross-layer relationships as edges. Figure 3 This is a schematic diagram of a cross-level data processing activity entity relationship model provided in an embodiment of this application. It should be noted that the entities and relationships listed in the embodiments of this application are examples. In practical applications, specific entities and relationships can be set according to the actual scenario, and this embodiment of the application does not impose specific limitations on them.
[0053] After obtaining the entity relationship model, a modeling script can be generated based on the entity relationship model. The modeling script is as follows: {"schema":[{"label":"Character","primary":"Character ID", "type":"VERTEX", "properties": [{"name":"role ID" , "type":"INT32","optional":false,"index":true, "unique":true}, {"name":"Job Title","type":"STRING","optional":false,"index":false, "unique":false}]}, {"label":"Information","primary":"Information ID","type":"VERTEX","properties": [{"name":"Information ID","type":"INT32","optional":false,"index":true, "unique":true}, {"name":"Information name","type":"STRING","optional":false,"index":false, "unique":false}]}, {"label":"Business Activity","primary":"Business Activity ID","type":"VERTEX", "properties": [{"name":"Business Activity ID","type":"INT32","optional":false,"index":true, "unique":true}, {"name":"Business Activity Name","type":"STRING","optional":false,"index":false,"unique":false}]}, {"label":"Environment","primary":"EnvironmentID","type":"VERTEX","properties": [{"name":"Environment ID","type":"INT32","optional":false,"index":true, "unique":true}, {"name":"Office Location Name", "type": "STRING", "optional": false, "index":false,"unique":false}]}, {"label":"User","primary":"UserID","type":"VERTEX","properties": [{"name":"User ID", "type": "INT32", "optional":false,"index":true, "unique":true}, {"name":"User Name","type":"STRING","optional":false,"index":false, "unique":false}]}, {"label":"data","primary":"dataID","type":"VERTEX","properties": [{"name":"data ID","type":"INT32","optional":false,"index":true, "unique":true}, {"name":"Data Name","type":"STRING","optional":false,"index":false, "unique":false}]}, {"label":"Information System","primary":"Information System ID","type":"VERTEX", "properties": [{"name":"Information System ID","type":"INT32","optional":false,"index":true, "unique":true}, {"name":"Information System Name","type":"STRING","optional":false,"index":false, "unique":false}]}, {"label":"Terminal","primary":"Terminal ID","type":"VERTEX","properties": [{"name":"Terminal ID","type":"INT32","optional":false,"index":true, "unique":true}, {"name":"Terminal Name","type":"STRING","optional":false,"index":false, "unique":false}]}, {"label":"Terminal Identifier","primary":"Terminal Identifier ID","type":"VERTEX", "properties": [{"name":"Terminal Identifier ID","type":"INT32","optional":false,"index":true, "unique":true}, {"name":"Terminal IP address","type":"STRING","optional":false,"index":false,"unique":false}]}, {"label":"User ID","primary":"User ID","type":"VERTEX", "properties": [{"name":"User ID","type":"INT32","optional":false,"index":true, "unique":true}, {"name":"Employee ID","type":"STRING","optional":false,"index":false, "unique":false}]}, {"label":"Data Identifier","primary":"Data Identifier ID","type":"VERTEX", "properties": [{"name":"Data Identifier ID","type":"INT32","optional":false,"index":true, "unique":true}, {"name":"Field Name","type":"STRING","optional":false,"index":false, "unique":false}]}, {"label":"System Identifier","primary":"System Identifier ID","type":"VERTEX", "properties": [{"name":"System Identifier ID","type":"INT32","optional":false,"index":true, "unique":true}, {"name":"System IP address","type":"STRING","optional":false,"index":false,"unique":false}]}, {"label":"Located at","type":"EDGE","properties":[{"name":"Located at","type":"STRING","optional":true,"index":false,"unique":false}],"constraints":[["role","environment"],["terminal","environment"]]}, {"label":"Processing","type":"EDGE","properties":[{"name":"Data Processing","type":"STRING","optional":false,"index":false,"unique":false}],"constraints":[["role","information"]]}, {"label":"Used for","type":"EDGE","properties":[{"name":"Used for","type":"STRING","optional":false,"index":false,"unique":false}],"constraints":[["Information","Business Activities"]]}, {"label":"belongs to","type":"EDGE","properties":[{"name":"belongs to","type":"STRING","optional":false,"index":false,"unique":false}],"constraints":[["user","role"]]}, {"label":"Usage","type":"EDGE","properties":[{"name":"Usage","type":"STRING","optional":false,"index":false,"unique":false}],"constraints":[["User","Terminal"],["User Identifier","Terminal Identifier"]]}, {"label":"Operation","type":"EDGE","properties":[{"name":"Data Operation","type":"STRING","optional":false,"index":false,"unique":false}],"constraints":[["User","Data"]]}, {"label":"Configured","type":"EDGE","properties":[{"name":"Configured","type":"STRING","optional":false,"index":false,"unique":false}],"constraints":[["Terminal","Terminal Identifier"],["Information System","System Identifier"]]}, {"label":"Serves","type":"EDGE","properties":[{"name":"Serves","type":"STRING","optional":false,"index":false,"unique":false}],"constraints":[["Information System","Business Activities"]]}, {"label":"Contains","type":"EDGE","properties":[{"name":"Contains","type":"STRING","optional":false,"index":false,"unique":false}],"constraints":[["Information","Data"],["Data","Data Identifier"]]}, {"label":"Stored in","type":"EDGE","properties":[{"name":"Stored in","type":"STRING","optional":false,"index":false,"unique":false}],"constraints":[["Data","Information System"],["Data Identifier","System Identifier"]]}, {"label":"Own","type":"EDGE","properties":[{"name":"Own","type":"STRING","optional":false,"index":false,"unique":false}],"constraints":[["User","User Identifier"]]}, {"label":"Operation Command","type":"EDGE","properties":[{"name":"Operation Command","type":"STRING","optional":false,"index":false,"unique":false}],"constraints":[["User Identifier","Data Identifier"]]}]} Step 203: Extract entity data from the data processing activity documents corresponding to the business logic layer, data operation layer, and technical implementation layer to obtain entity relationship data.
[0054] The data processing documents required in this embodiment are mainly internal enterprise documents and tables related to business operations; these documents are all basic documents required for enterprise information technology construction and management. In subsequent embodiments, the key information corresponding to the document data sources required for each stage will be explained. If the document to be identified has missing data sources, a survey and interview should be conducted beforehand to supplement the relevant data sources. Identification can only begin after the data sources have been supplemented.
[0055] Extracting entity relationships from the business logic layer is the main step in this application's embodiments. Entity relationships in the data operation layer and technical implementation layer are supplements and extensions to these extractions. The core data sources are business process description documents and system operation description documents. Business process description documents contain detailed step descriptions of the business process, including but not limited to: job role names, business office names, business information names, business names, business operation descriptions, and other relevant information. System operation description documents need to include the operation signatures of each functional module and data in the system, used to supplement and revise the business operations in the business process.
[0056] The data operation layer entity relationships extend the business logic layer. They are supplemented by job description and staffing documents, data resource catalog documents, and information system asset list documents, which connect the entities in the business logic layer to their corresponding entities in the data operation layer. Job description and staffing documents provide the names of the business personnel associated with each job role, as well as the names of their assigned office terminals. Data resource catalog documents provide the names of the data tables or data fields contained in the business information. Information system asset list documents provide the names of the information systems associated with the business activities. The operational relationships between personnel names and their corresponding data can be inherited from the business information corresponding to the job roles. These data operation relationships are identified along with the business process description documents in the business logic layer.
[0057] The technical implementation layer entity relationships are the concrete identifiers of the data operation layer entity relationships within an information system. They can be identified and confirmed using documents such as job descriptions and personnel allocation documents, information system asset lists, data asset lists, and information system log descriptions. Specifically, the core function of job descriptions and personnel allocation documents is to identify the terminal identification information corresponding to the terminal name. This identification information must be unique to the terminal, using features such as IP address, network segment, and terminal ID to uniquely identify the terminal. The information system asset list document is used to identify the system identification information corresponding to the information system name. Similarly, this identification information must meet the uniqueness requirement of the information system, specifically including features such as IP address, network segment, and information system ID to uniquely define the information system. The data asset list document focuses on data-level identification, accurately matching the English names of data fields or data tables, providing a basis for the technical identification of data entities. The purpose of information system log documentation is to extract log keywords associated with various data operations. These keywords can be directly used as data operation identifiers at the technical implementation level, providing support for the dynamic association of entity relationships.
[0058] It should be noted that the document names mentioned above are just examples. In practical applications, appropriate documents can be selected for entity data extraction based on the actual situation.
[0059] In this embodiment of the application, when extracting entity relationships in the business logic layer, a large language model is used as the core processing component. Domain expert knowledge is pre-configured as dedicated prompt words and input into the large language model. The semantic understanding and reasoning capabilities of the large language model are used to identify core entity relationships and obtain entity relationship data output by the large language model.
[0060] Specifically, a large language model is used to extract entities such as "roles," "environments," "information," and "business activities," as well as their interrelationships. This step is a core prerequisite for the business data processing activity parsing technical solution. It primarily involves connecting to the large language model and designing standardized prompts for the model. This provides clear logical guidance and constraints for the model to perform targeted parsing of business process documents and accurately extract key information related to data operations, ensuring that the model's output meets the technical requirements for traceability and verifiability of business data processing activities. The core logic of the prompts designed in this step revolves around task extraction, information extraction rules, knowledge base matching mechanisms, output specifications, and constraints.
[0061] Regarding task extraction, the core functional positioning of the large model is clearly defined in the prompts: after receiving the input business process document, it focuses on the "detailed process description" section of the document to accurately extract and infer nine types of key information, specifically including the business activity name, role name, data operation type name, role's office location name, business information name, operation content overview, data processing activity type name, index, and explanation; and generates a standardized data operation list based on the extracted information to achieve full-process traceability and standardized verification of business data processing activities.
[0062] Regarding information extraction rules, to ensure the accuracy and consistency of information extraction by the large model, the following standardized extraction rules are preset in the prompt words to guide the large model in locating and refining key information.
[0063] Business Activity Name Extraction Rules: The main model uses the title of the section containing the business activity process description as the sole extraction basis, refining it into a standardized name in the format of "XXXX Business". Index Information Extraction Rules: The main model clearly records the section number and page number of the business process document corresponding to each extracted piece of information, ensuring that all extracted information can be accurately traced back to the original document. Operation Content Overview Extraction Rules: The main model uses the textual description of the corresponding business operation in the original document as a basis to extract and concisely summarize the core information, retaining the three key elements: the operation subject, the operation object, and the core action. Explanation Generation Rules: The main model uses the keywords describing data operations in the original document to explain the reasoning process of the data operation type, clearly marking the relationship between the keywords and the matched operation type.
[0064] The design of the prompt word embedded knowledge base matching mechanism addresses the problem of inconsistent identification of data operation types and data processing activity types in existing technologies. This step designs a prompt word embedded with a dual-standardized knowledge base matching mechanism to guide the large model to achieve accurate matching based on the preset knowledge base. The prompt model calls the preset "data processing activity identification knowledge base" to identify the specific data processing activities in the original document according to Chinese laws and regulations. The matching scope is strictly limited to: data collection, data storage, data use, data processing, data transmission, data provision, data disclosure, and data deletion.
[0065] To ensure the standardization and usability of the large model's output, standardized output requirements are explicitly designed in the prompts. Specifically, the large model is guided to generate a data operation list in CSV table format. The table must contain ten headers: "Serial Number, Business Activity Name, Role Name, Data Operation Type Name, Role's Office Location Name, Business Information Name, Operation Content Overview, Data Processing Activity Type Name, Index, and Explanation." Furthermore, to refine and standardize the output content, the following two core specifications are defined: First, the single-entry output specification requires that different roles, data operation types, and business information must be split and generated as independent row records, and multiple values are strictly prohibited in the same field. Second, the business information splitting output specification requires that when multiple sets of business information are identified in the original document, they must be split into multiple independent table records for each piece of information.
[0066] To ensure the consistency and accuracy of the large-scale model's output, the following core constraints are embedded in the prompts to strictly limit its parsing and output behavior: First, information traceability constraints are strengthened, explicitly stipulating that the large-scale model must ensure that the index field of each record can be accurately traced to the specific section and page number in the business process document; information entries that cannot be traced must not be output. Second, format uniformity constraints are enforced, requiring the large-scale model to strictly follow the preset table headers, field types, and record splitting rules for output, and prohibiting it from adjusting the format itself. Through the above constraint design, the large-scale model can be guided to achieve standardized and accurate parsing of data operation information in business process documents, thereby effectively solving the technical problems of inconsistent operation types, chaotic information classification, and poor traceability in current business data extraction, providing high-quality and highly reliable basic data support for the standardized analysis of subsequent business data processing activities.
[0067] In addition to the data sources mentioned above, the "Data Processing Activity Definition Specification" can be used to standardize the definition of data processing activities in the business logic layer, and the "Data Operation Specification" can be used to standardize the definition of data operation names in the data operation layer. The "Data Processing Activity Definition Specification" document describes the definitions and reasoning logic of each stage of a data processing activity, enabling the large language model to infer the names of the relevant data processing activities based on the business activities. The rules for defining data processing activities are shown in the table below.
[0068]
[0069] The "Data Operation Instructions" document describes the names and reasoning logic of various data operations, enabling the large language model to infer the names of the relevant data operations based on business activities. The rules for defining data operations are shown in the table below:
[0070] After preparing the relevant data source documents, input the documents into the document reading module of the data processing system. The document reading module, as the basic data input module of the data processing system, is configured to read various types of documents related to the enterprise's internal business, including but not limited to Word documents, Excel spreadsheets, and PDF documents. This module receives the storage path or data stream of the document to be parsed, extracts the valid content from the document, and outputs it as a standardized data source in a unified format, which is then transmitted to the subsequent data processing modules.
[0071] After extracting entity data from the business logic layer documents, entity mapping and expansion can be performed using standardized enterprise management ledgers and explanatory documents, enabling automated extraction of entity relationships under zero-sample conditions.
[0072] The specific extraction steps are as follows: The system ledger documents are matched with preset keywords to obtain the target fields contained within them. Different system ledger documents have different preset keywords. For example, for documents related to job postings and personnel allocation, the preset keywords could be name, employee ID, job title, department information, etc. For documents related to IT resource allocation, the preset keywords could be office computer name, user, office location, etc. Therefore, keywords can be set according to the specific document. The target fields are the fields in the system ledger documents that successfully match the preset keywords. For example, if a document related to job postings and personnel allocation contains name and employee ID, then name and employee ID are the target fields.
[0073] Map the target fields to target entities in the first-level entity, second-level entity, and third-level entity; continuing with the above example, if the target fields are name and employee number, the name can be mapped to the "user" entity and the employee number can be mapped to the "user identifier" entity.
[0074] Extract the target relationships between target entities, for example: the relationship between "user" and "role" is "belongs to"; the relationship between "user" and "user ID" is "owns".
[0075] Retrieves the specific data corresponding to the target field from the system ledger document. Specific data refers to the specific values of the target field; for example, the specific data for a name could include Zhang San, Li Si, etc., and the specific data for an employee ID could include 001, 002, etc.
[0076] Supplementary entity data is generated based on the target entity, target relationship, and the specific data corresponding to the target relationship.
[0077] To clearly define the process of generating structured entity relationships based on predefined configurations and field mapping rules, the output entity relationship triples can be represented by the following formula, combining the aforementioned source data document set, entity and relationship type set, target field set, and field mapping relationship.
[0078] 1. Import document d, and associate field f in document d with a target field k in the target field set K; ; 2. Once the ƒ field in document d is mapped, the extraction of the corresponding entity e∈E and relation r∈R is triggered;
[0079] 3. Read the specific value of field f from document d;
[0080] 4. Output structured entity relation triples;
[0081] in D Represents the collection of source data documents. E Represents a collection of entity types. R Represents a set of relation types. d Represents the collection of source data documents D A single document instance in f Document d The original data fields included. K This represents a predefined set of target fields. k Represents the target field set K A single target field in μ This represents a field mapping function. r Represents a set of relation types R A single relation type instance in v Represents the original field f In the document dThe specific values that can be taken in the value.
[0082] The extracted entity relationship data is generated in CSV format to serve as the data input for the graph database.
[0083] The entity relationship extraction rules are as follows:
[0084] Step 204: Generate a knowledge graph based on the entity relationship model and entity relationship data.
[0085] The core objective of this step is to standardize and structure the import of the cross-level data processing activity entity relationship model generated in step 202 and the entity relationship data obtained through zero-sample extraction from the large language model in step 203 into a graph database, ultimately generating a knowledge graph that can support the traceability and analysis of data processing activities. This step ensures the integrity, hierarchical relevance, and storage effectiveness of entity relationships through a standardized import process and data mapping mechanism.
[0086] The specific method is as follows: Preprocessed entities are mapped to node types with different labels, and the mapping relationships of their attribute fields are defined. The corresponding CSV format entity data source files are imported in batches using the import module, generating all nodes and storing their attributes.
[0087] Based on the pre-defined cross-level model, different types of relationship edges are defined, and the node IDs and edge attributes that their start point (Src) and end point (Dst) depend on are specified. The import module will batch create edges from the relationship data source file with the node ID as the association key, ensuring accurate connection of relationships between entities and complete recording of attributes.
[0088] Once all entity and relation data has been imported, the graph database, based on its graph storage engine and indexing mechanism, automatically completes the structured organization of entities and relations, generating a knowledge graph that meets the requirements of the cross-level model in step 202. This knowledge graph visually presents various cross-level entities in the form of nodes, and presents the hierarchical and business relationships between entities in the form of edges with attributes.
[0089] The knowledge graph constructed in this application clearly presents the relationship between roles, environment, data, and business, providing a clear business logic basis for data security access control, achieving precise matching between access allocation and business needs, and avoiding redundancy or missing permissions. Furthermore, this cross-layered modeling relationship can establish semantic connections between the business logic layer and the data operation layer with traditional technical layer log auditing and flow auditing mechanisms, providing core support for building a more comprehensive baseline for data processing activities.
[0090] This application's embodiments achieve zero-sample automated extraction by introducing a large language model and combining it with a graph database batch construction process, significantly reducing the manpower and time costs of knowledge graph construction while significantly improving the overall efficiency of data processing activities, modeling, and analysis. Specifically, automated extraction replaces the high-cost, long-cycle process of traditional manual annotation and modeling relying on domain experts; standardized batch import and verification mechanisms avoid errors and repeated corrections from manual input; and the resulting structured knowledge graph enables cross-level relational queries, risk insights, and compliance verification to be completed in seconds, providing immediate and reliable data support for business decisions and security operations, saving operating costs and improving response efficiency from the source.
[0091] Based on the above embodiments, after the knowledge graph is constructed, query requests can be received, processed based on the knowledge graph, and query results can be generated. For example, if the query request is to find out what data the sales manager operated on using his terminal, the knowledge graph can determine the sales manager's specific name, employee number, and the terminal he used, as well as the specific data he operated on that terminal.
[0092] This application, based on a recruitment management business activity modeling scenario, illustrates the use of a cross-level data processing activity modeling method based on a large language model, as proposed in this invention. In recruitment management business modeling, this method effectively addresses pain points such as fragmented business activity data, weak cross-stage correlations, and high reliance on manual modeling. This results in a comprehensive improvement in the accuracy, process traceability, and management efficiency of recruitment business modeling.
[0093] According to the method described in detail in this invention, the specific process is as follows: 1. Define the vertex-edge model and generate the vertex-edge modeling script, as follows: {"schema":[{"label":"Character","primary":"Character ID","type":"VERTEX", "properties": [{"name":"role ID","type":"INT32","optional":false,"index":true, "unique":true}, {"name":"Job Title","type":"STRING","optional":false,"index":false, "unique":false}]}, {"label":"Information","primary":"Information ID","type":"VERTEX","properties": [{"name":"Information ID","type":"INT32","optional":false,"index":true, "unique":true}, {"name":"Information name","type":"STRING","optional":false,"index":false, "unique":false}]}, {"label":"Business Activity","primary":"Business Activity ID","type":"VERTEX", "properties": [{"name":"Business Activity ID","type":"INT32","optional":false,"index":true, "unique":true}, {"name":"Business Activity Name","type":"STRING","optional":false,"index":false, "unique":false}]}, {"label":"Environment","primary":"EnvironmentID","type":"VERTEX","properties": [{"name":"Environment ID","type":"INT32","optional":false,"index":true, "unique":true}, {"name":"Office Location Name","type":"STRING","optional":false,"index":false, "unique":false}]}, {"label":"User","primary":"UserID","type":"VERTEX","properties": [{"name":"User ID","type":"INT32","optional":false,"index":true, "unique":true}, {"name":"User Name","type":"STRING","optional":false,"index":false, "unique":false}]}, {"label":"data","primary":"dataID","type":"VERTEX","properties": [{"name":"data ID","type":"INT32","optional":false,"index":true, "unique":true}, {"name":"Data Name","type":"STRING","optional":false,"index":false, "unique":false}]}, {"label":"Information System","primary":"Information System ID","type":"VERTEX", "properties": [{"name":"Information System ID","type":"INT32","optional":false,"index":true, "unique":true}, {"name":"Information System Name","type":"STRING","optional":false,"index":false, "unique":false}]}, {"label":"Terminal","primary":"Terminal ID","type":"VERTEX","properties": [{"name":"Terminal ID","type":"INT32","optional":false,"index":true, "unique":true}, {"name":"Terminal Name","type":"STRING","optional":false,"index":false, "unique":false}]}, {"label":"Terminal Identifier","primary":"Terminal Identifier ID","type":"VERTEX", "properties": [{"name":"Terminal Identifier ID","type":"INT32","optional":false,"index":true, "unique":true}, {"name":"Terminal IP address","type":"STRING","optional":false,"index":false,"unique":false}]}, {"label":"User ID","primary":"User ID","type":"VERTEX", "properties": [{"name":"User ID","type":"INT32","optional":false,"index":true, "unique":true}, {"name":"Employee ID","type":"STRING","optional":false,"index":false, "unique":false}]}, {"label":"Data Identifier","primary":"Data Identifier ID","type":"VERTEX", "properties": [{"name":"Data Identifier ID","type":"INT32","optional":false,"index":true, "unique":true}, {"name":"Field Name","type":"STRING","optional":false,"index":false, "unique":false}]}, {"label":"System Identifier","primary":"System Identifier ID","type":"VERTEX", "properties": [{"name":"System Identifier ID","type":"INT32","optional":false,"index":true, "unique":true}, {"name":"System IP address","type":"STRING","optional":false,"index":false,"unique":false}]}, {"label":"Located at","type":"EDGE","properties":[{"name":"Located at","type":"STRING","optional":true,"index":false,"unique":false}],"constraints":[["role","environment"],["terminal","environment"]]}, {"label":"Processing","type":"EDGE","properties":[{"name":"Data Processing","type":"STRING","optional":false,"index":false,"unique":false}],"constraints":[["role","information"]]}, {"label":"Used for","type":"EDGE","properties":[{"name":"Used for","type":"STRING","optional":false,"index":false,"unique":false}],"constraints":[["Information","Business Activities"]]}, {"label":"belongs to","type":"EDGE","properties":[{"name":"belongs to","type":"STRING","optional":false,"index":false,"unique":false}],"constraints":[["user","role"]]}, {"label":"Usage","type":"EDGE","properties":[{"name":"Usage","type":"STRING","optional":false,"index":false,"unique":false}],"constraints":[["User","Terminal"],["User Identifier","Terminal Identifier"]]}, {"label":"Operation","type":"EDGE","properties":[{"name":"Data Operation","type":"STRING","optional":false,"index":false,"unique":false}],"constraints":[["User","Data"]]}, {"label":"Configured","type":"EDGE","properties":[{"name":"Configured","type":"STRING","optional":false,"index":false,"unique":false}],"constraints":[["Terminal","Terminal Identifier"],["Information System","System Identifier"]]}, {"label":"Serves","type":"EDGE","properties":[{"name":"Serves","type":"STRING","optional":false,"index":false,"unique":false}],"constraints":[["Information System","Business Activities"]]}, {"label":"Contains","type":"EDGE","properties":[{"name":"Contains","type":"STRING","optional":false,"index":false,"unique":false}],"constraints":[["Information","Data"],["Data","Data Identifier"]]}, {"label":"Stored in","type":"EDGE","properties":[{"name":"Stored in","type":"STRING","optional":false,"index":false,"unique":false}],"constraints":[["Data","Information System"],["Data Identifier","System Identifier"]]}, {"label":"Own","type":"EDGE","properties":[{"name":"Own","type":"STRING","optional":false,"index":false,"unique":false}],"constraints":[["User","User Identifier"]]}, {"label":"Operation Instruction","type":"EDGE","properties":[{"name":"Operation Instruction","type":"STRING","optional":false,"index":false,"unique":false}],"constraints":[["User Identifier","Data Identifier"]]}]}.
[0094] 2. Generate a basic model view, which presents the entities and relationship framework to be extracted, but does not contain data, such as... Figure 4 As shown.
[0095] 3. Configure prompt words for the large language model. The prompt words are as follows: # Role You are a business data processing activity analysis expert who can analyze information from provided documents, including data operations, business information, and data processing activities, across various business steps.
[0096] # Skill The model needs to extract the following information items one by one from the "Detailed Process Description" section of the business process document. The definition and extraction rules for each information item are as follows: Serial number: A continuous and unique identifier arranged in the order in which data records are generated; Business Activity Name: Extracted from the section title corresponding to the business activity process description, it should reflect the core attributes of the business activity (such as "XXXX business" or "XXXX process operation"). Role Name: The name of the main role that performs the corresponding operation in the business process. The original text description must be extracted accurately. Data operation type name: Strictly refer to the preset names in the "Data Operation Recognition Knowledge Base" for identification and matching. Only the following names can be selected: access, query, filter, fill in or enter, import or upload, archive, restore, preview, export or download, print, submit, review, comment or note, modify or edit, statistics or calculation, analysis, visualization, modeling, share, provide or distribute, delete (except for adding or using non-preset data operation types); Name of the office location of the role: Information on the office location of the role performing the operation, which must be fully extracted from the original text; Business Information Name: The specific business information content involved in the business process must be extracted independently as a single piece of information (multiple pieces of business information are prohibited from being separated by commas or other symbols and presented together). Operation Summary: A concise summary of the original text of the corresponding business operation, retaining the core elements of the operation; Data processing activity type name: Strictly refer to the preset names in the "Data Processing Activity Identification Knowledge Base" for identification and matching. Only the following names can be selected: data collection, data storage, data use, data processing, data transmission, data provision, data disclosure, and data deletion (adding or using non-preset data processing activity types is prohibited). Index: Records the section number and page number of the business process document corresponding to this data entry, clarifying the source of the information; Explanation: This section explains the keywords in the original text upon which the reasoning corresponds to the "data operation type name" is based, establishing the association between the operation type and the original text description.
[0097] # Output format specifications Overall format: Output in CSV table format, with the first row being the table header. The header fields are as follows: serial number, business activity name, role name, data operation type name, name of the office where the role is located, business information name, operation content overview, data processing activity type name, index, and explanation. Record generation rules: (1) Different roles, different data operation types, and different business information should generate independent data record rows respectively, and it is prohibited to include multiple similar information in a single field; (2) The business information name field must follow the principle of "single information per line of record". For example, "company strategy, industry layout, human resource planning, staffing and establishment, etc." in the original text must be split into four independent records: "company strategy", "industry layout", "human resource planning" and "staffing and establishment"; "annual employee recruitment plan" and "annual employee recruitment scheme" in the original text must be split into two independent records: "annual employee recruitment plan" and "annual employee recruitment scheme". Content constraints: The output content contains only CSV table data, without any additional symbols, comments or redundant descriptions, to ensure the purity and standardization of the data.
[0098] 4. Import the business process specification document "Human Resources Management Business Specification - Recruitment Management Volume" as the business process specification document for this example. Start the large language model recognition and reasoning. This case focuses on sorting out and analyzing the "Campus Recruitment Written Test Process" in the business process specification document to obtain the roles, environment, information, business entities, the relationship between roles and environment, the relationship between roles and information, and the relationship between information and business activities.
[0099] 5. Business personnel confirm the accuracy of data processing activities and make adjustments accordingly.
[0100] Prepare the associated identification file, configure field mapping, and gradually expand to obtain the relationships between various entities. The specific steps are as follows: 1) Import the "Human Resources Department Position Establishment and Personnel Allocation Table", map the "Name" field to the "User" entity, map the "Employee ID" field to the "User Identifier" entity, associate the "Position Name" field and the "Role" entity, obtain the relationship between roles and users, and obtain the relationship between users and user identifiers.
[0101] 2) Import the "HR Department IT Resource Allocation Table", map the "Asset Alias" field to the terminal entity, map the "Fixed IP" field to the terminal identifier entity, associate the "Terminal Location" field to the environment entity, associate the "User" field to the user entity, obtain the relationship between the terminal and the environment, obtain the relationship between the terminal and the terminal identifier, obtain the relationship between the terminal and the user, and the relationship between the terminal identifier and the user identifier inherits the relationship between the terminal and the user.
[0102] 3) Import the "Recruitment Management Business Data Resource Catalog" document, associate the "Business Object" field with the information entity, map the "Main Field" field with the data entity, and obtain the relationship between data and information.
[0103] 4) Import the "Human Resources Department Information System Asset List" document, map the "Asset Name" field to the information system entity, map the "Server IP / Address" field to the system identifier entity, associate the "Involved Business" field to the business activity entity, obtain the relationship between the information system and the business activity, and obtain the relationship between the information system and the system identifier.
[0104] 5) Import the "Recruitment Management Subsystem Data Asset List" document, map "Field English Name" to the data identifier entity, associate "Field Chinese Name" to the "Data" entity, associate the "System Name" field to the information system entity, obtain the relationship between data and data identifier, obtain the relationship between data and information system, and obtain the relationship between data identifier and information system identifier.
[0105] 6) Import the "Log Description Document" to obtain the relationship between data operations and data operation instruction keywords, and the relationship between operation instructions from user identifier to data identifier inherits the operation relationship from user to data.
[0106] 5. The system parses the mapping relationship, generates a graph database import file, and generates an entity relationship rule table in graph database format.
[0107] 6. Import the generated entity-relationship file into the graph database.
[0108] 7. This system supports the Cypher syntax, a dedicated query language for graph databases. To conduct a query test, input the full query command "MATCH (n) RETURN n LIMIT 99999" to obtain a complete view of the knowledge graph of the data processing activities of "Campus Recruitment Written Test Business".
[0109] 8. Query the entities related to the business logic layer and the relationships between them. The query syntax is as follows.
[0110] MATCH (src)-[rel]-(dst) / / Key optimization: Add parentheses to the tag conditions of src and dst to clarify the logical relationship.
[0111] WHERE (src:role OR src:information OR src:environment OR src:business activity) AND (dst: role OR dst: information OR dst: environment OR dst: business activity) / / Returns complete information about the source node, relationship, and target node.
[0112] RETURN / / Source node information (tag + attributes).
[0113] src AS source entity, labels(src) AS source entity tags, / / Relationship information (type + attributes).
[0114] rel AS relationship, `type(rel)` is the relation type. / / Target node information (label + attributes).
[0115] dst AS target entity, labels(dst) AS target entity labels / / Optional: Limit the number of returned records to avoid excessive data volume (adjust according to actual needs).
[0116] LIMIT 1000; 9. To query related entities in the data operation layer and the relationships between them, the query syntax is as follows: MATCH (src)-[rel]-(dst) WHERE (src:user OR src:data OR src:informationsystem OR src:terminal) AND (dst: user OR dst: data OR dst: information system OR dst: terminal) / / Returns complete information about the source node, relationship, and target node.
[0117] RETURN / / Source node information (tag + attributes).
[0118] src AS source entity, labels(src) AS source entity tags, / / Relationship information (type + attributes).
[0119] rel AS relationship, `type(rel)` is the relation type. / / Target node information (label + attributes).
[0120] dst AS target entity, labels(dst) AS target entity labels / / Optional: Limit the number of returned records to avoid excessive data volume (adjust according to actual needs).
[0121] LIMIT 2000; 10. To query the relationship between entity keys in the technology implementation layer, the query syntax is as follows: MATCH (src)-[rel]-(dst) / / Key optimization: Add parentheses to the tag conditions of src and dst to clarify the logical relationship.
[0122] WHERE (src:user identifier OR src:data identifier OR src:system identifier OR src:terminal identifier) AND (dst: User ID OR dst: Data ID OR dst: System ID OR dst: Terminal ID) / / Returns complete information about the source node, relationship, and target node.
[0123] RETURN / / Source node information (tag + attributes) src AS source entity, labels(src) AS source entity tags, / / Relationship information (type + attributes).
[0124] rel AS relationship, `type(rel)` is the relation type. / / Target node information (label + attributes).
[0125] dst AS target entity, labels(dst) AS target entity labels / / Optional: Limit the number of returned records to avoid excessive data volume (adjust according to actual needs).
[0126] LIMIT 2000; 11. Specified Condition Query Test. Query what data Zhang San from the Human Resources Department manipulated using his terminal in the recruitment management process, and what those operations were.
[0127] MATCH (user:user{username: 'Zhang San'})-[r]->(data:data) RETURN user AS Zhang San User Entity TYPE(r) AS Relationship Type, PROPERTIES(r) AS relational property, LABELS(data) AS associated entity tags, data AS Related Entity Details UNION ALL MATCH (user:user{username: 'Zhang San'})-[r]->(terminal:terminal) RETURN user AS Zhang San User Entity TYPE(r) AS Relationship Type, PROPERTIES(r) AS relational property, LABELS (terminal) AS associated entity tags, Terminal AS associated entity details, ORDER BY associated entity tags.
[0128] Figure 5 This is a schematic diagram of a knowledge graph generation device for cross-level data processing activities provided in an embodiment of this application. The device can be a module, program segment, or code on an electronic device. It should be understood that this device is similar to the one described above. Figure 2 The method implementation corresponds to this and can be executed. Figure 2 The specific functions of the device involved in each step of the method embodiment can be found in the description above; to avoid repetition, detailed descriptions are omitted here. The device includes: an entity relationship acquisition module 501, a model generation module 502, an entity relationship data acquisition module 503, and a knowledge graph generation module 504, wherein: The entity relationship acquisition module 501 is used to acquire the entities and relationships corresponding to the business logic layer, data operation layer, and technical implementation layer, as well as the cross-layer relationships between entities in adjacent layers. Among them, the business logic layer is used to describe the data processing activities involved in the business logic layer; the data operation layer is used to describe the specific data operations in the business activities; and the technical implementation layer is used to describe the underlying technical execution details of the data operation behavior. Model generation module 502 is used to generate entity-relationship models based on entities, relationships, and cross-layer relationships; The entity relationship data acquisition module 503 is used to extract entity data from the data processing activity documents corresponding to the business logic layer, data operation layer and technical implementation layer to obtain entity relationship data. The knowledge graph generation module 504 is used to generate a knowledge graph based on the entity relationship model and entity relationship data.
[0129] Based on the above embodiments, the technical implementation layer is the lowest layer, and the business logic layer is the highest layer; the entity relationship acquisition module 501 is specifically used for: Obtain the first-level relationships between the first-level entities in the business logic layer, the second-level relationships between the second-level entities in the data operation layer, and the third-level relationships between the third-level entities in the technology implementation layer; and obtain the first cross-layer relationships between the first-level entities and the second-level entities, and the second cross-layer relationships between the second-level entities and the third-level entities.
[0130] Based on the above embodiments, the model generation module 502 is specifically used for: The entity relationship model is generated by using the first-level entity, the second-level entity, and the third-level entity as vertices, and the first-level relationship, the second-level relationship, the third-level relationship, the first cross-level relationship, and the second cross-level relationship as edges.
[0131] Based on the above embodiments, the device is also used for: Generate modeling scripts based on entity relationship models; Knowledge graphs are generated based on modeling scripts and entity relationship data.
[0132] Based on the above embodiments, the entity relationship data acquisition module 503 is specifically used for: The entity data is extracted from the business process description class document corresponding to the business logic layer using a large language model. The business entity data includes the specific data corresponding to the first-level entities contained in the business process description class document. Entity data is extracted from the system ledger documents by keyword matching to obtain supplementary entity data. The supplementary entity data includes the specific data corresponding to the second-level entities, the specific data corresponding to the third-level entities, and the relationships between entities. The entity relationship data includes business entity data and supplementary entity data.
[0133] Based on the above embodiments, the entity relationship data acquisition module 503 is specifically used for: Input prompt words into the large language model. The prompt words include: extraction task, information extraction rules, data processing activity recognition knowledge base, output specifications and constraints. Entity data is extracted from business process description documents using a large language model based on prompt words.
[0134] Based on the above embodiments, the entity relationship data acquisition module 503 is specifically used for: By matching preset keywords with system ledger documents, the target fields contained in the system ledger documents can be obtained; Map the target field to the target entities in the first-level, second-level, and third-level entities; Extract the target relationships between target entities; Read the specific data corresponding to the target field from the system ledger document; Supplementary entity data is generated based on the target entity, target relationship, and the specific data corresponding to the target relationship.
[0135] Based on the above embodiments, the first layer of entities includes environment entities, role entities, information entities, and business activity entities; the second layer of entities includes terminal entities, user entities, data entities, and information system entities; and the third layer of entities includes terminal identifier entities, user identifier entities, data identifier entities, and system identifier entities.
[0136] Figure 6 This is a schematic diagram of the physical structure of the electronic device provided in the embodiments of this application, such as... Figure 6 As shown, the electronic device includes: a processor 601, a memory 602, and a bus 603; wherein: The processor 601 and the memory 602 communicate with each other through the bus 603; The processor 601 is used to call program instructions in the memory 602 to execute the methods provided in the above-described method embodiments, including, for example, obtaining entities and relationships corresponding to the business logic layer, data operation layer, and technical implementation layer, as well as cross-layer relationships between entities in adjacent layers; wherein, the business logic layer is used to describe the data processing activities involved in the business logic layer; the data operation layer is used to describe the specific data operations in the business activities; the technical implementation layer is used to describe the underlying technical execution details of the data operation behavior; generating an entity relationship model based on entities, relationships, and cross-layer relationships; extracting entity data from the data processing activity documents corresponding to the business logic layer, data operation layer, and technical implementation layer to obtain entity relationship data; and generating a knowledge graph based on the entity relationship model and entity relationship data.
[0137] Processor 601 can be an integrated circuit chip with signal processing capabilities. The processor 601 can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor can be a microprocessor or any conventional processor.
[0138] The memory 602 may include, but is not limited to, random access memory (RAM), read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), etc.
[0139] This embodiment discloses a computer program product, which includes a computer program stored on a non-transitory computer-readable storage medium. The computer program includes program instructions, and when the program instructions are executed by a computer, the computer can execute the methods provided in the above-described method embodiments, such as: obtaining entities and relationships corresponding to the business logic layer, data operation layer, and technical implementation layer, as well as cross-layer relationships between entities in adjacent layers; wherein, the business logic layer is used to describe the data processing activities involved in the business logic layer; the data operation layer is used to describe the specific data operations in the business activities; the technical implementation layer is used to describe the underlying technical execution details of the data operation behavior; generating an entity relationship model based on entities, relationships, and cross-layer relationships; extracting entity data from the data processing activity documents corresponding to the business logic layer, data operation layer, and technical implementation layer to obtain entity relationship data; and generating a knowledge graph based on the entity relationship model and entity relationship data.
[0140] This embodiment provides a non-transitory computer-readable storage medium storing computer instructions that cause the computer to execute the methods provided in the above-described method embodiments. These instructions include, for example, obtaining entities and relationships corresponding to the business logic layer, data operation layer, and technical implementation layer, as well as cross-layer relationships between entities in adjacent layers. The business logic layer describes data processing activities involved in the business logic layer; the data operation layer describes specific data operations within the business activities; and the technical implementation layer describes the underlying technical execution details of the data operation behavior. An entity-relationship model is generated based on entities, relationships, and cross-layer relationships. Entity data is extracted from the data processing activity documents corresponding to the business logic layer, data operation layer, and technical implementation layer to obtain entity-relationship data. A knowledge graph is generated based on the entity-relationship model and the entity-relationship data.
[0141] In the embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. The apparatus embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. Furthermore, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Additionally, the displayed or discussed mutual couplings, direct couplings, or communication connections may be through some communication interfaces; indirect couplings or communication connections between devices or units may be electrical, mechanical, or other forms.
[0142] Furthermore, the units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0143] Furthermore, the functional modules in the various embodiments of this application can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.
[0144] In this document, relational terms such as first and second are used only to distinguish one entity or operation from another entity or operation, without necessarily requiring or implying any such actual relationship or order between these entities or operations.
[0145] The above description is merely an embodiment of this application and is not intended to limit the scope of protection of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application.
Claims
1. A method for generating knowledge graphs for cross-level data processing activities, characterized in that, The method is applied in the field of data security, and the method includes: Obtain the entities and relationships corresponding to the business logic layer, data operation layer, and technical implementation layer, as well as the cross-layer relationships between entities in adjacent layers; wherein, the business logic layer is used to describe the data processing activities involved in the business logic layer; the data operation layer is used to describe the specific data operations in the business activities; and the technical implementation layer is used to describe the underlying technical execution details of the data operation behavior; Generate an entity-relationship model based on the entity, the relationship, and the cross-layer relationship; Entity data is extracted from the data processing activity documents corresponding to the business logic layer, the data operation layer, and the technical implementation layer to obtain entity relationship data. A knowledge graph is generated based on the entity relationship model and the entity relationship data; The technical implementation layer is the lowest layer, and the business logic layer is the highest layer; the acquisition of entities and relationships corresponding to the business logic layer, data operation layer, and technical implementation layer, as well as cross-layer relationships between entities in adjacent layers, includes: Obtain the first layer entity of the business logic layer and the first layer relationship between the first layer entity, the second layer entity of the data operation layer and the second layer entity, the third layer entity of the technology implementation layer and the third layer relationship between the third layer entity, and obtain the first cross-layer relationship between the first layer entity and the second layer entity, and the second cross-layer relationship between the second layer entity and the third layer entity. The generation of the entity-relationship model based on the entity, the relationship, and the cross-layer relationship includes: The entity relationship model is generated by using the first-layer entities, the second-layer entities, and the third-layer entities as vertices, and the first-layer relationships, the second-layer relationships, the third-layer relationships, the first cross-layer relationships, and the second cross-layer relationships as edges. The first-layer entities include environment entities, role entities, information entities, and business activity entities. The second-layer entities include terminal entities, user entities, data entities, and information system entities. The third-layer entities include terminal identifier entities, user identifier entities, data identifier entities, and system identifier entities. Entity data is extracted from the data processing activity documents corresponding to the business logic layer, the data operation layer, and the technical implementation layer to obtain entity relationship data, including: The entity data of the business process description class document corresponding to the business logic layer is extracted using a large language model to obtain business entity data; the business entity data includes the specific data corresponding to the first-layer entity contained in the business process description class document. Supplementary entity data is obtained by extracting entity data from the system ledger documents through keyword matching. The supplementary entity data includes the specific data corresponding to the second-level entities, the specific data corresponding to the third-level entities, and the relationships between the entities. The entity relationship data includes the business entity data and the supplementary entity data.
2. The method according to claim 1, characterized in that, The method further includes: Generate a modeling script based on the entity relationship model; The step of generating a knowledge graph based on the entity relationship model and the entity relationship data includes: The knowledge graph is generated based on the modeling script and the entity relationship data.
3. The method according to claim 1, characterized in that, The step of extracting entity data from the business process description document corresponding to the business logic layer using a large language model includes: Input prompt words into the large language model, the prompt words including: extraction task, information extraction rules, data processing activity identification knowledge base, output specifications and constraints; The entity data of the business process description document is extracted using the large language model based on the prompt words.
4. The method according to claim 1, characterized in that, The step of extracting entity data from system ledger documents through keyword matching to obtain supplementary entity data includes: By matching preset keywords with the system ledger document, the target fields contained in the system ledger document can be obtained; Map the target field to the target entities in the first-level entity, the second-level entity, and the third-level entity; Extract the target relationships between the target entities; Read the specific data corresponding to the target field from the system ledger document; The supplementary entity data is generated based on the target entity, the target relationship, and the specific data corresponding to the target relationship.
5. The method according to any one of claims 1-4, characterized in that, After generating the knowledge graph, the method further includes: Receive query requests; The query request is processed based on the knowledge graph to generate query results.
6. An electronic device, characterized in that, include: Processor, memory, and bus, among which: The processor and the memory communicate with each other via the bus; The memory stores program instructions that can be executed by the processor, and the processor can execute the method as described in any one of claims 1-5 by calling the program instructions.
7. A non-transitory computer-readable storage medium, characterized in that, The non-transitory computer-readable storage medium stores computer instructions, which, when executed by a computer, cause the computer to perform the method as described in any one of claims 1-5.
8. A computer program product, characterized in that, It includes computer program instructions, which, when read and executed by a processor, perform the method as described in any one of claims 1-5.
Citation Information
Patent Citations
Knowledge graph construction method and system, information processing system, terminal and medium
CN112966057A