Enterprise data risk processing method and system based on dynamic knowledge graph

By constructing dynamic knowledge graphs and adaptive propagation rules, the problems of incomplete coverage of static graphs and one-sided risk quantification are solved, real-time and accurate identification and early warning of enterprise data risk processing are achieved, and the initiative of risk prevention and control and the accuracy of decision-making are improved.

CN120494538BActive Publication Date: 2025-09-12GUIZHOU UNIVERSITY OF FINANCE AND ECONOMICS
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510991556.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-18
Publication Date
2025-09-12
Estimated Expiration
2045-07-18

AI Technical Summary

Technical Problem

In existing technologies, enterprise data risk processing methods have incomplete coverage of risk propagation paths due to the difficulty of static graphs in reflecting the dependencies between entities in real time. Traditional risk quantification methods fail to comprehensively consider the path structure density and the superposition effect of risk attributes, resulting in significant deviations between risk assessment results and the actual impact range, making it difficult to support accurate risk warning and disposal decisions.

Method used

By collecting real-time data streams from enterprise operations, building a dynamic knowledge graph, combining the risk impact range description to generate adaptive propagation rules, calling the pre-trained model to traverse the dynamic graph, integrating the path structure density and the risk attribute superposition effect, calculating the comprehensive risk value and triggering a risk warning.

Benefits of technology

It achieves real-time capture of dynamic dependencies between entities in a complex and ever-changing enterprise environment, accurately identifies cross-regional and cross-level risk transmission chains, improves the timeliness and scenario adaptability of risk warnings, enhances the ability to explore hidden risk transmission paths, and ensures that risk disposal strategies are deeply aligned with the actual operating status of the enterprise.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120494538B_ABST
    Figure CN120494538B_ABST
Patent Text Reader

Abstract

The present invention provides an enterprise data risk processing method and system based on a dynamic knowledge graph. By collecting the real-time data stream generated during the operation of the enterprise, the entity objects and association relationships in the real-time data stream are extracted to form initial nodes and initial edges. A risk processing request is received, the initial node corresponding to the target entity is located from the dynamic knowledge graph according to the risk attribute identifier, and the association propagation path generation rule of the target entity is determined based on the risk impact range description. The pre-trained risk propagation model is called, and the adjacent nodes in the dynamic knowledge graph are traversed according to the association propagation path generation rule with the initial node as the starting point to generate a set of risk propagation paths corresponding to the target entity. According to the node connection density and the risk attribute superposition strength of each risk propagation path, the comprehensive risk value corresponding to the target entity is calculated, and a risk warning signal is triggered when it exceeds the preset threshold. The present invention can improve the initiative and decision-making accuracy of enterprise risk prevention and control.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of text processing, and specifically to a method and system for enterprise data risk processing based on a dynamic knowledge graph. Background Art

[0002] Enterprise data risk management involves identifying and evaluating the relationships between entities and risk transmission paths during enterprise operations. Existing technologies typically generate risk transmission paths by traversing fixed node connections using predefined rules, and rely on simple indicators to calculate risk values. However, due to the dynamic nature of internal and external data flows within an enterprise, static graphs struggle to promptly reflect the real-time dependencies between entities, resulting in incomplete coverage of risk transmission paths and a lag behind actual business status. Furthermore, the fixed-rule-driven path generation mechanism lacks the ability to dynamically adapt to geographic distribution, business levels, and timeliness constraints, making it prone to generating redundant paths or missing key nodes. Furthermore, traditional risk quantification methods focus only on one-sided features such as node attributes or path length, without comprehensively considering the combined effects of path structure density and risk attributes. This results in significant deviations between risk assessment results and the actual scope of impact, making it difficult to support accurate risk warning and disposal decisions. Summary of the Invention

[0003] This application provides an enterprise data risk management method and system based on a dynamic knowledge graph.

[0004] According to one aspect of the present application, a method for enterprise data risk processing based on a dynamic knowledge graph is provided, the method comprising: collecting real-time data streams generated during the operation of the enterprise, extracting entity objects and association relationships in the real-time data streams, and forming initial nodes and initial edges of the dynamic knowledge graph; receiving a risk processing request, the risk processing request including a risk attribute identifier and a risk impact range description of the target entity; locating the initial node corresponding to the target entity from the dynamic knowledge graph according to the risk attribute identifier, and determining an association propagation path generation rule for the target entity based on the risk impact range description; calling a pre-trained risk propagation model, taking the initial node as the starting point, traversing adjacent nodes in the dynamic knowledge graph according to the association propagation path generation rule, and generating a risk propagation path set corresponding to the target entity; calculating a comprehensive risk value corresponding to the target entity based on the node connection density and risk attribute superposition strength of each risk propagation path in the risk propagation path set, and triggering a risk warning signal when the comprehensive risk value exceeds a preset threshold.

[0005] According to another aspect of the present application, a computer system is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method described above.

[0006] This application has at least the following beneficial effects:

[0007] The present invention dynamically constructs a network of knowledge graph nodes and edges by collecting multi-source heterogeneous data streams in real time during enterprise operation, generates adaptive propagation rules based on the geographical, business level and timeliness constraints in the description of the risk impact range, calls the pre-trained model to traverse the node connection path in the dynamic graph, and integrates the path structure density and the risk attribute superposition effect to comprehensively quantify the risk value. It can capture the dynamic dependency relationship between entities in a complex and changeable enterprise environment in real time, accurately identify cross-regional and cross-level risk transmission chains, and effectively solve the problems of incomplete path coverage and one-sided risk quantification caused by static data modeling in traditional methods; through the continuous evolution characteristics of the dynamic knowledge graph and the adaptive matching mechanism of risk propagation rules, the timeliness and scenario adaptability of risk warning are significantly improved; at the same time, based on the computational logic coupled with path topology characteristics and business attributes, the ability to mine hidden risk transmission paths is enhanced, ensuring that the risk disposal strategy is deeply aligned with the actual operating status of the enterprise, forming a full-process governance from data collection, graph update to risk warning, and comprehensively improving the initiative and decision-making accuracy of enterprise risk prevention and control. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] Figure 1 A schematic diagram of an application scenario of an enterprise data risk processing method based on a dynamic knowledge graph according to an embodiment of the present application is shown.

[0009] Figure 2 A flowchart of a method for handling enterprise data risks based on a dynamic knowledge graph according to an embodiment of the present application is shown.

[0010] Figure 3 A schematic diagram of the composition of a computer system according to an embodiment of the present application is shown. DETAILED DESCRIPTION

[0011] Figure 1 A schematic diagram of an application scenario provided according to an embodiment of the present application is shown. The application scenario includes one or more data acquisition systems 101, a computer system 120, and one or more networks 110 coupling the one or more data acquisition systems 101 to the computer system 120. The data acquisition system 101 can be configured to execute one or more application programs.

[0012] In an embodiment of the present application, the computer system 120 may run one or more services or software applications that enable execution of an enterprise data risk processing method based on a dynamic knowledge graph.

[0013] exist Figure 1 In the configuration shown, the computer system 120 may include one or more components that implement the functions performed by the computer system 120. These components may include software components, hardware components, or a combination thereof that can be executed by one or more processors. A user operating the data acquisition system 101 may, in turn, utilize one or more applications to interact with the computer system 120 to utilize the services provided by these components. It should be understood that a variety of different system configurations are possible, which may vary depending on the application scenario. Therefore, Figure 1 is one example of a system for implementing the various methods described herein and is not intended to be limiting.

[0014] Computer system 120 may include one or more general-purpose computers, dedicated server computers (e.g., PC (personal computer) servers, UNIX servers, mid-range servers), blade servers, mainframe computers, server clusters, or any other suitable arrangement and / or combination. Computer system 120 may include one or more virtual machines running virtual operating systems, or other computing architectures involving virtualization (e.g., one or more flexible pools of logical storage devices that may be virtualized to maintain a server's virtual storage device). In various embodiments, computer system 120 may run one or more services or software applications that provide the functionality described below.

[0015] In some embodiments, computer system 120 can be a server in a distributed system, or a server integrated with blockchain. Computer system 120 can also be a cloud server, or an intelligent cloud computing server or intelligent cloud host equipped with artificial intelligence technology. A cloud server is a host product within the cloud computing service system that addresses the management difficulties and poor scalability of traditional physical hosts and virtual private servers (VPS) services.

[0016] The application scenario may also include one or more databases 130. In certain embodiments, these databases may be used to store data and other information. For example, one or more of databases 130 may be used to store real-time enterprise data streams. Databases 130 may reside in a variety of locations. For example, the database used by computer system 120 may be local to computer system 120, or it may be remote from computer system 120 and communicate with computer system 120 via a network-based or dedicated connection.

[0017] Please refer to Figure 2The enterprise data risk processing method based on the dynamic knowledge graph provided by the embodiment of the present invention includes the following steps:

[0018] Step S100: Collect the real-time data stream generated during the operation of the enterprise, extract the entity objects and association relationships in the real-time data stream, and form the initial nodes and initial edges of the dynamic knowledge graph.

[0019] Real-time data streams refer to the continuous data generated by an enterprise during its operations. They include information from various aspects, such as enterprise business system logs, equipment status monitoring data, and external market environment data. Enterprise business system logs record detailed information on various operations and events within the enterprise business system, such as user logins and transaction operations; equipment status monitoring data reflects the operating status of enterprise equipment, such as equipment temperature, pressure, and operating hours; and external market environment data covers various external market information related to the enterprise, such as market price fluctuations and competitor dynamics. Entity objects refer to individuals or concepts with practical significance identified from real-time data streams, such as users in the enterprise business system, equipment in equipment status monitoring, and market entities in the external market environment. Associations describe the connections between entity objects, such as the operational dependency between users and equipment, and the market influence relationship between market entities and users.

[0020] In this step, real-time data streams from the enterprise's operations are collected and analyzed to extract entity objects and relationships. These entity objects are then mapped as initial nodes of the dynamic knowledge graph, and the relationships are mapped as initial edges. For example, in an e-commerce company, real-time data streams may include user shopping records (business system logs), server operating status (device status monitoring data), and price changes for similar products on the market (external market environment data). By analyzing this data, entity objects such as users, servers, and products can be identified, as well as relationships such as the usage relationship between users and servers and the influence between product prices and user purchasing behavior. Using these entity objects as initial nodes and relationships as initial edges, the initial structure of the dynamic knowledge graph can be constructed.

[0021] As an implementation method, step S100 collects real-time data streams generated during the operation of the enterprise, extracts entity objects and association relationships in the real-time data streams, and forms initial nodes and initial edges of the dynamic knowledge graph, which may specifically include the following steps:

[0022] Step S110: Receive real-time data streams through a distributed data acquisition interface. The real-time data streams include enterprise business system logs, equipment status monitoring data, and external market environment data.

[0023] A distributed data collection interface is an interface that can simultaneously collect data from multiple data sources, improving data collection efficiency and reliability. Enterprise business system logs are information about various operations and events automatically recorded during the operation of an enterprise's business system. This information can help enterprises understand business operations and user behavior patterns. Equipment status monitoring data is obtained through real-time monitoring of the operating status of enterprise equipment and can reflect the health and performance of the equipment. External market environment data refers to various external market information related to the enterprise, such as market prices, supply and demand, policies and regulations, etc. In this step, the distributed data collection interface is used to receive real-time data streams generated by the enterprise's operations from different data sources. For example, for a manufacturing enterprise, the distributed data collection interface can simultaneously collect data from multiple sources, including the enterprise's production management system (for business system logs), equipment monitoring sensors (for equipment status monitoring data), and the market information platform (for external market environment data). This ensures comprehensive and timely acquisition of various data and information from the enterprise's operations, providing a rich data foundation for subsequent entity object and relationship extraction.

[0024] Step S120: Perform multimodal feature analysis on the real-time data stream to identify the operation subject features in the enterprise business system log, the equipment identification features in the equipment status monitoring data, and the market entity features in the external market environment data, and map the operation subject features, equipment identification features, and market entity features into entity objects in the dynamic knowledge graph.

[0025] Multimodal feature parsing involves comprehensively analyzing different types of data to extract key features. Operation subject features refer to information in enterprise business system logs that uniquely identifies the operation subject, such as the user's account number, name, and role. Device identification features refer to information in device status monitoring data that uniquely identifies a device, such as the device's serial number, model, and name. Market entity features refer to information in external market environment data that uniquely identifies a market entity, such as the company's name, product brand, or market category.

[0026] In this step, the collected real-time data stream is subjected to multimodal feature analysis to identify operational subject features, device identification features, and market entity features from different types of data. These features are then mapped to entity objects in the dynamic knowledge graph. For example, in a financial enterprise, business system logs are analyzed to identify operational subject features such as user accounts and names; device status monitoring data is analyzed to identify device identification features such as server numbers and models; and external market environment data is analyzed to identify market entity features such as stock names and codes. These features are mapped to entity objects such as user nodes, server nodes, and stock nodes in the dynamic knowledge graph, thereby constructing the node system of the dynamic knowledge graph.

[0027] Step S130: Generate an operational dependency relationship between the operating subject characteristics and the device identification characteristics based on the behavior sequence of the operating subject characteristics within a preset time window, and generate a market impact relationship between the market entity characteristics and the operating subject characteristics based on the fluctuation trend of the market entity characteristics in historical data.

[0028] The preset time window refers to a pre-set time range used to analyze the behavior sequence of the operator. A behavior sequence refers to the sequence and record of a series of operations performed by the operator within the preset time window. Operational dependencies describe the degree of dependency between the operator and the device. For example, if a user frequently uses a device, it indicates a strong operational dependency between the user and the device. The fluctuation trend of market entity characteristics refers to the changes in market entity indicators such as price and sales volume in historical data. Market impact relationships describe the degree to which changes in market entities affect the operator. For example, a price increase for a certain commodity in the market may affect user purchasing behavior. In this step, the behavioral sequence of the operator's characteristics within the preset time window is first analyzed. By statistically analyzing information such as the operator's frequency of device usage and duration, an operational dependency relationship is generated between the operator's characteristics and the device's identification features. For example, in an internet company, the preset time window is one month. By analyzing user operation records on a server within a month, it is found that a user frequently accesses a particular server, thereby determining a strong operational dependency between the user and the server. Then, we analyze the fluctuation trends of market entity characteristics in historical data. By comparing changes in market entities with changes in the behavior of the operating entities, we can generate a market impact relationship between the characteristics of the market entities and the operating entities. For example, in a retail enterprise, by analyzing the price fluctuations of a certain category of goods and user purchase records over the past year, we found that when the price of the goods increased, the purchase volume of users decreased significantly, thus confirming the existence of a market impact relationship between the goods and users.

[0029] As an embodiment, step S130 generates an operation dependency relationship between the operation subject characteristics and the device identification characteristics based on the behavior sequence of the operation subject characteristics within a preset time window, and generates a market impact relationship between the market entity characteristics and the operation subject characteristics based on the fluctuation trend of the market entity characteristics in historical data. The steps may specifically include the following:

[0030] Step S131: extracting the operation behavior log of the operation subject characteristics within a preset time window, where the operation behavior log includes the operation timestamp, operation type identification, and associated device identification characteristics.

[0031] An operation log is a log file that records detailed information about various operations performed by an operator within a preset time window. The operation timestamp records the specific time the operation occurred. The operation type identifier distinguishes different types of operations, such as login, query, and modification. The associated device identifier identifies the device involved in the operation.

[0032] In this step, we extract operational behavior logs from the enterprise's business system logs, capturing the characteristics of the operator within a preset time window. For example, in a logistics company, the preset time window is one week. We then extract the operational behavior logs of a particular operator within that week from the logistics management system logs. These logs contain the timestamp of each operation, the operation type (e.g., goods inbound, goods outbound), and the associated equipment identifiers (e.g., warehouse number, handling equipment number, etc.). These operational behavior logs provide detailed data support for subsequent analysis of operational dependencies.

[0033] Step S132: sorting the operation behavior logs in time sequence according to the operation timestamps, identifying the device access patterns that appear repeatedly in the continuous operation type identifiers, and counting the occurrence frequency of the device access patterns within a preset time window.

[0034] Chronological sorting arranges operation logs by the order of their timestamps to better analyze the temporal order and continuity of operations. Device access patterns refer to the patterns of device access by the operator during continuous operations, such as multiple consecutive accesses to the same device or accessing multiple devices in a preset sequence. Occurrence frequency refers to the number of times a device access pattern occurs within a preset time window.

[0035] In this step, the extracted operation logs are sorted by operation timestamps, and then analyzed for recurring device access patterns within the consecutive operation type identifiers. For example, in a manufacturing company, after sorting a worker's operation logs over a month, it was found that the worker always accessed device A first, then device B, in multiple consecutive operations. This access sequence constitutes a device access pattern. The frequency of this device access pattern over the course of the month is counted to provide a basis for subsequent assessment of the strength of the operation dependency.

[0036] Step S133: When the occurrence frequency exceeds the preset pattern threshold, the device access pattern between the operation subject feature and the device identification feature is marked as a high-frequency dependency, and an initial strength value of the operation dependency relationship is generated based on the time distribution uniformity of the high-frequency dependency.

[0037] The preset pattern threshold refers to a pre-set critical value of the frequency of occurrence, which is used to determine whether the device access pattern is a high-frequency dependency. A high-frequency dependency indicates that there is a relatively frequent and stable usage relationship between the operating subject and the device. The time distribution uniformity refers to whether the occurrence time of the high-frequency dependency within the preset time window is evenly distributed. The initial strength value is used to indicate the strength of the operation dependency. In this step, the occurrence frequency of the device access pattern is compared with the preset pattern threshold. If the occurrence frequency exceeds the preset pattern threshold, the device access pattern is marked as a high-frequency dependency. Then, the time distribution uniformity of the high-frequency dependency is analyzed. For example, in an e-commerce company, a user accesses a server multiple times within a fixed time period within a month, indicating that the time distribution of the high-frequency dependency is relatively uniform. The initial strength value of the operation dependency is generated based on the time distribution uniformity. The more uniform the time distribution, the higher the initial strength value, indicating that the operation dependency between the operating subject and the device is stronger.

[0038] Step S134: Synchronously obtain the price fluctuation records and trading volume change records of the market entity characteristics in the historical data, align the trends of the price fluctuation records and the trading volume change records, and identify the overlapping periods of the peak and trough intervals in the price fluctuation records and the sudden increase and decrease intervals in the trading volume change records.

[0039] Price fluctuation records refer to price changes of market entities in historical data, typically recorded in a time series format. Trading volume change records refer to changes in trading volume of market entities in historical data, also recorded in a time series format. Trend alignment involves matching price fluctuation records and trading volume change records over time to better analyze the relationship between them. Peak-trough intervals refer to the time periods in price fluctuation records when prices reach their highest and lowest points. Sudden increase / decrease intervals refer to the time periods in trading volume change records when trading volume suddenly increases or decreases.

[0040] In this step, we simultaneously obtain historical data on price fluctuations and trading volume changes for market entities and align their trends over time. For example, in a stock market, we obtain the price fluctuations and trading volume change records for a particular stock over the past year and arrange them chronologically. We then identify peaks and troughs in the price fluctuation records and periods of sudden increases and decreases in the trading volume change records, and identify any overlaps between them. These overlaps may reveal the inherent connection between the price and trading volume changes of market entities, providing important clues for subsequent analysis of market influence relationships.

[0041] Step S135: extract the operation behavior log of the operation subject characteristics corresponding to the overlapping period, count the access frequency change rate of the operation subject characteristics to the target device identification characteristics during the overlapping period, and match the access frequency change rate with the fluctuation range of the market entity characteristics for positive correlation.

[0042] The access frequency change rate refers to the ratio of the change in the frequency of accesses to the target device by the operator during the overlapping time period relative to other time periods. Fluctuation amplitude refers to the degree of change in the price or trading volume of the market entity during the overlapping time period. Positive correlation matching involves analyzing whether there is a positive correlation between the access frequency change rate and the fluctuation amplitude. Specifically, does the access frequency change rate of the operator to the target device increase as the fluctuation amplitude of the market entity increases? In this step, the operation behavior logs corresponding to the operator's characteristics for the overlapping time period are extracted. The frequency of accesses to the target device's identification characteristics during the overlapping time period by the operator are counted, and the access frequency change rate is calculated. For example, within a financial enterprise, for a trader, the operation behavior logs for the trader's trading device during the overlapping time period of stock price fluctuations and trading volume changes are extracted, and the frequency change rate of access to the trading device is counted. The access frequency change rate is then positively correlated with the fluctuation amplitude of the market entity characteristics. If the access frequency change rate of the trader to the trading device increases as the stock price fluctuation amplitude increases, then a positive correlation exists between the two.

[0043] Step S136: If the positive correlation matching degree exceeds the preset correlation threshold, the fluctuation response relationship between the market entity characteristics and the operation subject characteristics is marked as a market impact relationship, and an initial impact factor of the market impact relationship is generated based on the ratio of the fluctuation amplitude to the access frequency change rate.

[0044] The preset correlation threshold is a pre-set critical value for the degree of positive correlation, used to determine whether a market influence relationship exists between the characteristics of a market entity and the characteristics of an operator. The volatility response relationship refers to the impact of the fluctuations of the market entity on the behavior of the operator. The initial impact factor is used to indicate the strength of the market influence relationship.

[0045] In this step, the positive correlation match is compared with a preset correlation threshold. If the positive correlation match exceeds the preset correlation threshold, it is considered that a market influence relationship exists between the market entity characteristics and the operating subject characteristics, and this fluctuation response relationship is marked as a market influence relationship. Then, an initial impact factor of the market influence relationship is generated based on the ratio of the fluctuation amplitude to the access frequency change rate. For example, in a real estate market, when the positive correlation match between the housing price fluctuation amplitude and the change rate of the homebuyer's access frequency to the property information query device exceeds the preset correlation threshold, the relationship between the housing price and the homebuyer is marked as a market influence relationship, and an initial impact factor is generated based on the ratio of the housing price fluctuation amplitude to the access frequency change rate. The larger the factor, the stronger the market influence relationship.

[0046] Step S137: Normalize the initial strength value of the operation dependency relationship and the initial impact factor of the market impact relationship, and map them to the confidence weight and time decay coefficient of the initial edge in the dynamic knowledge graph respectively.

[0047] Normalization involves converting data of varying scales and magnitudes to a uniform scale for comparison and processing. Confidence weights represent the credibility of the relationships represented by initial edges in a dynamic knowledge graph. Higher weights indicate more reliable relationships. The time decay coefficient indicates how much a relationship changes over time. A larger coefficient indicates slower decay.

[0048] In this step, the initial strength values ​​of the operational dependency relationships and the initial impact factors of the market influence relationships are normalized and converted to a uniform numerical range. For example, both the initial strength values ​​and the initial impact factors are normalized to the range of [0, 1]. The normalized initial strength values ​​are then mapped to the confidence weights of the initial edges corresponding to the operational dependency relationships in the dynamic knowledge graph, and the normalized initial impact factors are mapped to the time decay coefficients of the initial edges corresponding to the market influence relationships. In this way, each initial edge in the dynamic knowledge graph has a corresponding confidence weight and time decay coefficient, which can more accurately represent the associations between entity objects and their changes over time.

[0049] Step S140: Encode the operation dependency relationship and the market impact relationship as initial edges connecting the corresponding entity objects in the dynamic knowledge graph, and record the generation timestamp and confidence weight of each initial edge.

[0050] Encoding involves converting operational dependencies and market impact relationships into a form that can be represented in the dynamic knowledge graph. The generation timestamp records the specific time when the initial edge was generated, and the confidence weight indicates the credibility of the relationship represented by the initial edge. In this step, the previously generated operational dependencies and market impact relationships are encoded as initial edges connecting the corresponding entity objects in the dynamic knowledge graph. For example, the operational dependency relationship between a user and a device is encoded as an edge connecting the user node and the device node, and the market impact relationship between a market entity and a user is encoded as an edge connecting the market entity node and the user node. The generation timestamp and confidence weight of each initial edge are also recorded. This information is crucial for subsequent risk propagation analysis and graph structure updates within the dynamic knowledge graph. Recording the generation timestamp allows us to understand the timeliness of the relationship, while recording the confidence weight allows us to assess the reliability of the relationship.

[0051] Step S200: receiving a risk handling request, which includes a risk attribute identifier of a target entity and a description of the risk impact scope.

[0052] A risk handling request is a request from a user or system to address enterprise data risks. A target entity is the specific object requiring risk assessment and handling, such as a company department, a piece of equipment, or a market entity. A risk attribute identifier uniquely identifies the risk attributes of the target entity, such as the risk type and level. A risk impact scope description details the potential impact of the risk on the target entity, including constraints such as geographic region, business level, and timeliness.

[0053] In this step, the system receives a risk handling request and extracts the target entity's risk attribute identifier and risk impact description from it. For example, within a large enterprise, a department may discover a potential data leak risk and issue a risk handling request. The request includes the department's (target entity's) risk attribute identifier (e.g., a high data leak risk level) and a risk impact description (e.g., the affected geographic area is the company's headquarters city, the business level is the department and its sub-departments, and the timeliness is within the next week). By receiving and parsing this request, the system provides a basis for subsequent risk identification and risk transmission path analysis.

[0054] Step S300: Locate the initial node corresponding to the target entity from the dynamic knowledge graph according to the risk attribute identifier, and determine the associated propagation path generation rule of the target entity based on the risk impact range description.

[0055] The initial node is the node in the dynamic knowledge graph that corresponds to the target entity and serves as the starting point for risk propagation path analysis. The associated propagation path generation rules guide the generation of the target entity's risk propagation path in the dynamic knowledge graph, including directional constraints, hop count constraints, and node weight calculation functions.

[0056] In this step, the initial nodes that match the target entity are first screened out from the dynamic knowledge graph based on the risk attribute identifier. For example, based on the entity type code in the risk attribute identifier, a set of candidate nodes whose node attributes match it is screened out from the dynamic knowledge graph. Then, based on the risk impact range description, the associated propagation path generation rule of the target entity is determined. For example, based on the geographical area constraints, business level constraints, and timeliness constraints in the risk impact range description, the candidate node set is screened and weighted to obtain the initial node corresponding to the target entity; at the same time, based on the propagation direction indicator and path depth limit parameter in the risk impact range description, the direction constraint condition and hop count constraint condition in the associated propagation path generation rule are generated; and the node weight calculation function in the associated propagation path generation rule is configured to calculate the path selection priority between adjacent nodes.

[0057] As an implementation method, step S300, locating the initial node corresponding to the target entity from the dynamic knowledge graph according to the risk attribute identifier, and determining the associated propagation path generation rule of the target entity based on the risk impact range description, may specifically include the following steps:

[0058] Step S310: Parse the entity type code in the risk attribute identifier, and filter a set of candidate nodes whose node attributes match the entity type code from the dynamic knowledge graph.

[0059] The entity type code is the encoding information used to represent the target entity type in the risk attribute identifier. For example, different codes can be used to represent different types of entities such as enterprise departments, equipment, and market entities. The candidate node set is the set of nodes selected from the dynamic knowledge graph whose node attributes match the entity type code.

[0060] In this step, the entity type code in the risk attribute identifier is parsed, and then nodes whose node attributes match this code are searched in the dynamic knowledge graph. For example, if the entity type code in the risk attribute identifier is "01," representing a corporate department, the system will filter all nodes in the dynamic knowledge graph whose node attributes are of the corporate department type to form a candidate node set. This narrows the scope of the initial node for subsequent target entity location, improving the accuracy and efficiency of location.

[0061] Step S320: Based on the geographical area constraints, business level constraints and timeliness constraints in the risk impact range description, the candidate node set is spatially filtered and time-attenuated weighted to obtain the initial node corresponding to the target entity.

[0062] Geographical constraints refer to restrictions on the affected geographical area in the risk impact range description, for example, they can be expressed by administrative region codes and physical location radius. Business-level constraints refer to restrictions on the affected business level, such as department affiliation identifiers and business line classification identifiers. Timeliness constraints refer to restrictions on the time range of risk impact, such as the effective time interval. Spatial range filtering refers to filtering out nodes that meet the geographical area requirements from the candidate node set based on geographic area constraints. Time decay weighting refers to adjusting the priority of a node based on the difference between the node's update timestamp and the current time.

[0063] In this step, the candidate node set is first spatially filtered according to the geographic area constraint. For example, the administrative area code and physical location radius in the geographic area constraint are parsed, and the nodes whose node location attributes match the administrative area code and whose distance from the center coordinate of the target entity is less than the physical location radius are filtered out from the candidate node set to generate the first candidate subset after spatial filtering. Then, the first candidate subset is hierarchically filtered according to the business hierarchy constraint. For example, the department affiliation identifier and the business line classification identifier in the business hierarchy constraint are extracted, and the nodes whose node attributes include the department affiliation identifier and whose business labels are consistent with the business line classification identifier are retained from the first candidate subset to generate the second candidate subset after hierarchical filtering. Then, the second candidate subset is time-filtered according to the timeliness constraint. For example, the valid time interval in the timeliness constraint is parsed, and the nodes whose node update timestamps fall within the valid time interval are extracted from the second candidate subset to generate the third candidate subset after timeliness filtering. Finally, the nodes in the third candidate subset are time-attenuated weighted. For example, the difference between the update timestamp and the current time of each node in the third candidate subset is obtained. The time decay factor is calculated based on the difference, and the time decay factor is multiplied by the initial activity weight of the node to generate the time decay weighted node priority weight. The nodes in the third candidate subset are sorted in descending order according to the node priority weight, and a preset number of nodes with the highest weights are selected as the initial nodes corresponding to the target entity. The attributes of the initial nodes are marked as the starting points for the risk propagation path calculation.

[0064] Step S321: parse the administrative area code and physical location radius in the geographic area constraint, and select nodes from the candidate node set whose node location attributes match the administrative area code and whose distance from the center coordinate of the target entity is less than the physical location radius, to generate a first candidate subset after spatial filtering.

[0065] Administrative region codes are used to uniquely identify different administrative regions, such as country, province, city, and county codes. A physical location radius is a distance range defined by the target entity's center coordinates. Node location attributes are attributes of nodes in a dynamic knowledge graph that represent their geographic location.

[0066] In this step, the administrative area code and physical location radius in the geographic area constraint are parsed, and then the nodes whose node location attributes match the administrative area code and whose distance from the center coordinates of the target entity is less than the physical location radius are screened out from the candidate node set. For example, assuming that the administrative area code in the geographic area constraint is "110101", which represents Dongcheng District, Beijing, the physical location radius is 10 kilometers, and the center coordinates of the target entity are the coordinates of a certain building. The system screens out the nodes in Dongcheng District, Beijing and less than 10 kilometers away from the building from the candidate node set, forming the first candidate subset after spatial filtering. In this way, the range of candidate nodes can be further narrowed to nodes that meet the geographic area requirements, thereby improving the accuracy of the subsequent positioning of the initial node of the target entity.

[0067] Step S322: Extract the department affiliation identifier and business line classification identifier in the business hierarchy constraint, retain the nodes whose node attributes include the department affiliation identifier and whose business labels are consistent with the business line classification identifier from the first candidate subset, and generate a second candidate subset after hierarchical filtering.

[0068] The department affiliation identifier is used to identify the department to which a node belongs, such as the department's number or name. The business line classification identifier is used to identify the business line to which a node belongs, such as its type or number. A business label is a label that indicates the business type of a node in the dynamic knowledge graph.

[0069] In this step, the department affiliation identifier and the business line classification identifier in the business hierarchy constraint are extracted, and then the nodes whose node attributes include the department affiliation identifier and whose business label is consistent with the business line classification identifier are filtered out in the first candidate subset. For example, suppose the department affiliation identifier in the business hierarchy constraint is "001", which represents the company's sales department, and the business line classification identifier is "02", which represents the electronic product sales business line. The system filters out the nodes that belong to the sales department and have the business label of electronic product sales in the first candidate subset to form the second candidate subset after hierarchical filtering. In this way, the range of candidate nodes can be further narrowed down to nodes that meet the business hierarchy requirements, thereby improving the accuracy of the initial node positioning of the target entity.

[0070] Step S323: parsing the valid time interval in the timeliness constraint, extracting nodes whose node update timestamps fall within the valid time interval from the second candidate subset, and generating a third candidate subset after timeliness filtering.

[0071] The effective time interval refers to the time range of the risk impact specified in the risk impact range description, such as from a certain time point to another time point. The node update timestamp refers to the specific time when the node information is updated in the dynamic knowledge graph. In this step, the effective time interval in the timeliness constraint is parsed, and then the nodes whose node update timestamps fall within the effective time interval are filtered out in the second candidate subset. For example, assuming that the effective time interval in the timeliness constraint is "2024-01-01 00:00:00" to "2024-01-31 23:59:59", the system filters out the nodes whose node update timestamps are within this time interval in the second candidate subset, forming the third candidate subset after time filtering. This ensures that the filtered nodes have been updated within the effective time range of the risk impact, thereby improving the timeliness and accuracy of the initial node positioning of the target entity.

[0072] Step S324: Obtain the difference between the update timestamp and the current time of each node in the third candidate subset, calculate the time decay factor based on the difference, and multiply the time decay factor by the initial activity weight of the node to generate the time decay weighted node priority weight.

[0073] The difference between the update timestamp and the current time represents the time interval between node information updates and the present. The time decay factor is calculated based on this time interval and represents the degree to which the timeliness of the node information affects its priority. The longer the time interval, the smaller the time decay factor. The initial activity weight is the weight that a node originally possesses in the dynamic knowledge graph, representing its activity level. In this step, the difference between the update timestamp and the current time of each node in the third candidate subset is first obtained. For example, if the current time is "2024-02-01 00:00:00" and a node's update timestamp is "2024-01-15 00:00:00," the difference is 16 days. The time decay factor is then calculated based on this difference. For example, an exponential decay function can be used to calculate the time decay factor; the longer the time interval, the smaller the decay factor. Finally, the time decay factor is multiplied by the node's initial activity weight to obtain the time-decay-weighted node priority weight. This method comprehensively considers the node's activity and information timeliness, allowing for a more accurate assessment of the node's priority in risk propagation path analysis.

[0074] Step S325: Arrange the nodes in the third candidate subset in descending order according to the node priority weights, select a preset number of nodes with the highest weights as the initial nodes corresponding to the target entity, and mark the attributes of the initial nodes as the starting points for risk propagation path calculation.

[0075] The preset number is the number of initial nodes to be selected that is set in advance. By arranging the nodes in the third candidate subset in descending order according to the node priority weights, the nodes with higher priorities can be placed in front. In this step, the nodes in the third candidate subset are arranged in descending order according to the node priority weights. For example, the nodes are sorted using a sorting algorithm. Then, a preset number of nodes with the highest weights are selected as the initial nodes corresponding to the target entity. Assuming that the preset number is 5, the top 5 nodes are selected as the initial nodes. Finally, the attributes of these initial nodes are marked as the starting points for the risk propagation path calculation. In this way, in the subsequent risk propagation path analysis, traversal and calculation can be started from these initial nodes to improve the accuracy and efficiency of the risk propagation path analysis.

[0076] Step S330: extracting the propagation direction indicator and the path depth limit parameter in the risk impact range description, and generating the direction constraint condition and the hop count constraint condition in the associated propagation path generation rule.

[0077] A propagation direction indicator is information used within the risk impact scope description to indicate the direction of risk propagation, such as "forward," "reverse," or "bidirectional." A path depth constraint specifies the maximum depth of a risk propagation path, such as the maximum number of hops. Direction constraints restrict the direction of risk propagation within the associated propagation path generation rules, while hop constraints restrict the number of hops in the risk propagation path.

[0078] In this step, the propagation direction indicator and path depth limit parameter are extracted from the risk impact scope description. For example, assume the propagation direction indicator in the risk impact scope description is "forward" and the path depth limit parameter is 3. Based on the propagation direction indicator, a direction constraint is generated, specifying that the risk can only propagate in the forward direction. Based on the path depth limit parameter, a hop constraint is generated, specifying that the maximum number of hops in the risk propagation path is 3. These constraints will play an important role in guiding the subsequent risk propagation path generation process, ensuring that the generated path meets the requirements of the risk impact scope.

[0079] Step S340: configuring a node weight calculation function in the associated propagation path generation rule, the node weight calculation function calculates the path selection priority between adjacent nodes based on the confidence weight of the initial edge and the dynamic attenuation coefficient of the generated timestamp.

[0080] The node weight calculation function is used to calculate the path selection priority between adjacent nodes. The confidence weight of the initial edge indicates the credibility of the relationship represented by the initial edge, and the dynamic decay coefficient of the generated timestamp indicates the degree to which the relationship changes over time. Path selection priority is used to determine which path is preferred for risk propagation.

[0081] In this step, you configure the node weight calculation function for the associated propagation path generation rule. For example, you can define a node weight calculation function that calculates the node weight based on the confidence weight of the initial edge and the dynamic decay coefficient of the generation timestamp. Specifically, the confidence weight of the initial edge can be multiplied by a factor related to the dynamic decay coefficient of the generation timestamp to obtain the path selection priority between adjacent nodes. This allows the system to prioritize paths with high path selection priorities during risk propagation, thereby more accurately simulating the propagation of risks in a dynamic knowledge graph.

[0082] Step S400: Call the pre-trained risk propagation model, take the initial node as the starting point, traverse the adjacent nodes in the dynamic knowledge graph according to the associated propagation path generation rules, and generate a risk propagation path set corresponding to the target entity.

[0083] A pretrained risk propagation model is a pretrained model used to simulate risk propagation in a dynamic knowledge graph. Adjacent nodes are nodes directly connected to the initial node in the dynamic knowledge graph. The risk propagation path set is the set of all possible risk propagation paths obtained by traversing the initial node according to the associated propagation path generation rules.

[0084] In this step, the pre-trained risk propagation model is called, and the initial node corresponding to the target entity is used as the starting point to traverse the adjacent nodes in the dynamic knowledge graph according to the associated propagation path generation rules. For example, the model determines the list of accessible adjacent nodes of the current node based on the directional constraint, and then prioritizes each node in the accessible adjacent node list according to the node weight calculation function, and generates a node access sequence according to the scoring result. Then, according to the hop count constraint, the node access operation is recursively performed in the path generation layer until the preset path depth is reached or the accessible adjacent node list is empty, and the path fragments generated by each recursion are recorded. Finally, the continuous node sequences that meet the directional constraints in all path fragments are merged to generate a risk propagation path set, and each path in the set is sorted according to the comprehensive priority. In this way, all possible risk propagation paths for the target entity can be found comprehensively and accurately.

[0085] As an implementation method, step S400 calls a pre-trained risk propagation model, takes the initial node as the starting point, traverses adjacent nodes in the dynamic knowledge graph according to the associated propagation path generation rules, and generates a risk propagation path set corresponding to the target entity. The steps may specifically include the following:

[0086] Step S410: inputting the feature vector of the initial node into the path generation layer of the risk propagation model, and the path generation layer determines a list of accessible adjacent nodes of the current node according to the direction constraint condition.

[0087] The feature vector of the initial node is a vector representation obtained by encoding the various attributes of the initial node. It contains important information about the initial node. The path generation layer is the part of the risk propagation model used to generate risk propagation paths. The accessible neighboring node list is a list of accessible nodes selected from all neighboring nodes of the current node based on directional constraints.

[0088] In this step, the feature vector of the initial node is input into the path generation layer of the risk propagation model. For example, information such as the node attributes of the initial node and the confidence weights of the associated edges are encoded into the feature vector. The path generation layer selects accessible adjacent nodes from all adjacent nodes of the current node based on the directional constraints in the associated propagation path generation rules to form a list of accessible adjacent nodes. For example, if the directional constraint is "forward," the path generation layer only selects the positive adjacent nodes of the current node as accessible adjacent nodes. This ensures that risk propagation proceeds in the specified direction and improves the accuracy of risk propagation path generation.

[0089] Step S420: Through the weight allocation module of the risk propagation model, a priority score is performed on each node in the accessible adjacent node list according to the node weight calculation function, and a node access sequence is generated according to the scoring results.

[0090] The weight assignment module is used to calculate node priorities in the risk propagation model. The node weight calculation function is the previously configured function used to calculate the path selection priority between adjacent nodes. The priority score is a score calculated for each node in the accessible adjacent node list based on the node weight calculation function, representing the node's priority in risk propagation. The node access sequence is the order in which accessible adjacent nodes are accessed, sorted according to the priority score.

[0091] In this step, the risk propagation model's weight assignment module assigns a priority score to each node in the accessible adjacent node list based on the node weight calculation function. For example, the confidence weight of the initial edge corresponding to each node and the dynamic decay coefficient of the generated timestamp are substituted into the node weight calculation function to obtain a priority score for each node. The accessible adjacent nodes are then sorted in descending order according to the priority score results to generate a node access sequence. This allows the system to access adjacent nodes sequentially according to the node access sequence during the subsequent risk propagation process, prioritizing nodes with higher priorities for propagation, thereby improving the efficiency and accuracy of risk propagation.

[0092] Step S430: According to the hop count constraint, the node access operation is recursively performed in the path generation layer until the preset path depth is reached or the accessible adjacent node list is empty, and the path segment generated each time is recorded.

[0093] The hop count constraint limits the number of hops in the risk propagation path generation rule. The preset path depth refers to the maximum depth of the risk propagation path, i.e., the maximum number of hops. A path segment refers to the portion of the risk propagation path generated during each recursive node visit. In this step, the path generation layer recursively performs node visits based on the hop count constraint. After each node visit, it determines whether the preset path depth has been reached or the list of accessible adjacent nodes is empty. If the preset path depth has not been reached and the list of accessible adjacent nodes is not empty, the next node in the node visit sequence is selected for visit and the current path segment is recorded. For example, assuming the preset path depth is 3, starting from the initial node, after the first visit to an adjacent node, the path segment consisting of these two nodes is recorded. The next adjacent nodes of this adjacent node are then visited, and new path segments are recorded. If the preset path depth has been reached or the list of accessible adjacent nodes is empty, the recursive operation stops. This constrains the risk propagation path to extend indefinitely while recording all possible path segments, providing a foundation for the subsequent generation of a complete set of risk propagation paths.

[0094] Step S440: Merge the continuous node sequences that meet the direction constraint conditions in all path segments to generate a risk propagation path set, and sort the paths in the set according to the comprehensive priority.

[0095] A continuous node sequence refers to a sequence consisting of a series of consecutive nodes. The comprehensive priority refers to the priority obtained by evaluating the risk propagation path after considering multiple factors such as node weight and path length.

[0096] In this step, all recorded path segments are merged, and continuous node sequences that satisfy the directional constraints are selected to form a risk propagation path set. For example, all path segments containing forward propagation node sequences are merged. Each path in the risk propagation path set is then sorted according to its overall priority. This overall priority can be calculated by taking into account factors such as the weights of the nodes in the path and the length of the path. For example, the higher the node weights and the shorter the path, the higher the overall priority. This sorting allows the system to more clearly understand the importance of each risk propagation path, providing a reference for subsequent risk assessment and resolution.

[0097] Based on step S400, it can be understood that the risk propagation model is a pre-trained model designed to traverse adjacent nodes in the dynamic knowledge graph, starting from an initial node, according to the associated propagation path generation rules, thereby generating a set of risk propagation paths corresponding to the target entity. This model primarily consists of a path generation layer and a weight assignment module.

[0098] The path generation layer's main function is to determine the list of accessible neighboring nodes for the current node based on directional constraints. It can adopt a deep neural network (DNN) architecture, specifically comprising an input layer, hidden layers, and an output layer. The input layer receives the feature vector of the initial node. These feature vectors encode the various attributes of the initial node and contain important node information, such as the node's attributes and the confidence weights of associated edges. The hidden layer can be composed of multiple fully connected layers, each containing multiple neurons. The input is transformed using nonlinear activation functions (such as the ReLU function) to learn more complex feature representations. The output layer outputs the list of accessible neighboring nodes for the current node. This list is selected from all of the current node's neighboring nodes based on the directional constraints in the associated propagation path generation rules.

[0099] The weight assignment module prioritizes each node in the accessible neighboring node list based on a node weight calculation function and generates a node access sequence based on the scoring results. The weight assignment module can employ an attention-based architecture to dynamically adjust node priorities by learning the importance of each node. Specifically, the weight assignment module can include an attention layer and a scoring layer.

[0100] The attention layer calculates an attention weight for each accessible neighboring node, representing the node's importance in risk propagation. The attention layer employs a multi-head attention mechanism, capturing diverse information through parallel computation across multiple attention heads. Each attention head calculates the correlations between a node and other nodes, then weights these correlations to obtain the node's attention weight. The scoring layer prioritizes each accessible neighboring node based on the node weight calculation function and the attention weights. The node weight calculation function calculates the path selection priority between adjacent nodes based on the confidence weights of the initial edges and the dynamic decay coefficient of the generated timestamps. The scoring layer combines the attention weights with the path selection priority to derive the final node priority score.

[0101] Step S500: Calculate the comprehensive risk value corresponding to the target entity based on the node connection density and risk attribute superposition strength of each risk propagation path in the risk propagation path set, and trigger a risk warning signal when the comprehensive risk value exceeds a preset threshold.

[0102] Node connection density refers to the closeness of connections between adjacent nodes in a risk propagation path, typically expressed as the ratio of the number of initial edges between adjacent nodes to the total number of hops along the path. Risk attribute stacking strength refers to the strength value obtained by accumulating the risk coefficients corresponding to the risk attribute labels carried by each node in the risk propagation path along the path. The comprehensive risk value is a numerical value representing the risk level of the target entity, calculated by comprehensively considering the node connection density and the risk attribute stacking strength. The preset threshold is a pre-set critical value used to determine whether a risk warning signal should be triggered.

[0103] In this step, the ratio of the number of initial edges between adjacent nodes in each risk propagation path to the total number of hops along the path is calculated to obtain the node connection density. For example, if a risk propagation path contains five nodes, has four initial edges between adjacent nodes, and has a total path hop count of four, the node connection density is 4 / 4 = 1. Next, the risk attribute labels carried by each node in the risk propagation path are extracted. The label type is matched to a preset risk coefficient table, and the risk coefficients of the same type are accumulated along the path to obtain the risk attribute stacking strength. For example, if three nodes in the path carry the same type of risk attribute label, and the risk coefficient for that type is 0.2, the risk attribute stacking strength is 0.2 × 3 = 0.6. Next, based on the preset density weight and stacking weight, the node connection density and risk attribute stacking strength are weighted and summed to obtain the local risk contribution value for each risk propagation path. For example, if the preset density weight is 0.4 and the stacking weight is 0.6, the local risk contribution value for this path is 1 × 0.4 + 0.6 × 0.6 = 0.76. The local risk contribution values ​​of all risk propagation paths in the risk propagation path set are normalized and summed using a decaying weight based on the path ranking results to generate a comprehensive risk value for the target entity. Finally, this comprehensive risk value is compared with a preset threshold. If the comprehensive risk value exceeds the threshold, a risk warning signal is triggered. This allows for a comprehensive and accurate assessment of the target entity's risk level and timely issuance of warnings.

[0104] As an embodiment, step S500, calculating the comprehensive risk value corresponding to the target entity based on the node connection density and the risk attribute superposition strength of each risk propagation path in the risk propagation path set, may specifically include the following steps:

[0105] Step S510: Count the ratio of the number of initial edges between adjacent nodes in each risk propagation path to the total number of hops in the path as the node connection density.

[0106] The number of initial edges refers to the number of edges connecting adjacent nodes in the risk propagation path. The total number of path hops refers to the number of nodes traversed from the starting node to the ending node in the risk propagation path minus 1. Node connection density reflects the closeness of the connections between nodes in the risk propagation path. Higher connection density indicates closer connections between nodes and a greater likelihood of risk propagation.

[0107] In this step, for each risk propagation path in the risk propagation path set, the number of initial edges between adjacent nodes and the total number of hops along the path are counted. For example, if a risk propagation path includes nodes A, B, C, and D, with three initial edges (AB, BC, CD) between adjacent nodes and a total hop count of three, then the node connection density is 3 / 3 = 1. By calculating the node connection density, we can quantify the closeness of connections between nodes in the risk propagation path, providing an important basis for the subsequent calculation of the comprehensive risk value.

[0108] Step S520: Extract the risk attribute label carried by each node in the risk propagation path, match the preset risk coefficient table according to the label type, accumulate the same type of risk coefficients along the path, and obtain the risk attribute superposition strength.

[0109] Risk attribute labels are tags carried by nodes that indicate they possess certain risk attributes, such as "data leakage risk" and "system failure risk." The preset risk factor table is a pre-defined table containing risk factors corresponding to various risk attribute labels. The risk attribute stacking intensity is the intensity value obtained by adding together risk factors of the same type along the risk propagation path. It reflects the degree of risk accumulation along the propagation path.

[0110] In this step, the risk attribute labels carried by each node in the risk propagation path are extracted. For example, node A in a risk propagation path carries the label "data leakage risk", node B also carries the label "data leakage risk", and node C carries the label "system failure risk". Then, the preset risk coefficient table is matched according to the label type to find the risk coefficient corresponding to each risk attribute label. Assume that the risk coefficient of "data leakage risk" is 0.3 and the risk coefficient of "system failure risk" is 0.2. The same type of risk coefficients are accumulated along the path. For "data leakage risk", the accumulated result is 0.3+0.3=0.6; for "system failure risk", the accumulated result is 0.2. The accumulated results of all risk coefficients of the same type are added together to obtain the risk attribute superposition strength, that is, 0.6+0.2=0.8. By calculating the risk attribute superposition strength, the degree of risk accumulation on the propagation path can be evaluated, providing important information for the calculation of the comprehensive risk value.

[0111] Step S530: According to the preset density weight and overlay weight, the node connection density and the risk attribute overlay intensity are weighted and summed to obtain the local risk contribution value of each risk propagation path.

[0112] The preset density weights and overlay weights are pre-set values ​​used to adjust the importance of node connection density and risk attribute overlay strength when calculating the local risk contribution value. The local risk contribution value is a numerical value that represents the contribution of each risk propagation path to the target entity's risk, calculated by comprehensively considering node connection density and risk attribute overlay strength.

[0113] In this step, a weighted summation of the node connection density and the risk attribute overlay intensity is performed based on the preset density weight and overlay weight. For example, if the preset density weight is 0.4, the overlay weight is 0.6, the node connection density of a risk propagation path is 0.8, and the risk attribute overlay intensity is 0.9, then the local risk contribution value of this path is 0.8 × 0.4 + 0.9 × 0.6 = 0.86. This weighted summation method comprehensively considers the impact of node connection density and risk attribute overlay intensity on risk propagation, and more accurately assesses the contribution of each risk propagation path to the target entity's risk.

[0114] Step S540: normalize the local risk contribution values ​​of all risk propagation paths in the risk propagation path set, and perform attenuated weighted summation based on the path sorting results to generate a comprehensive risk value corresponding to the target entity.

[0115] Normalization involves converting local risk contribution values ​​of varying scopes and magnitudes into a unified scale for easier comparison and processing. Attenuated weighted summation involves weighting the local risk contribution values ​​of risk propagation paths of varying priority levels based on the path ranking results. Paths with higher priorities receive greater weights, and the weights gradually decay as the path priority decreases.

[0116] In this step, the local risk contribution values ​​of all risk propagation paths in the risk propagation path set are first normalized. For example, all local risk contribution values ​​are converted to the range of [0, 1]. Then, an attenuation-weighted summation is performed based on the path sorting results. Assuming that the risk propagation paths are sorted according to comprehensive priority, the weight of the path with the highest priority is 1, and as the priority decreases, the weight gradually decreases according to the preset attenuation coefficient. For example, the weight of the path with the second priority is 0.8, the weight of the path with the third priority is 0.6, and so on. Multiply the normalized local risk contribution value of each path by the corresponding weight, and then add all the results to obtain the comprehensive risk value corresponding to the target entity. Through normalization and attenuation-weighted summation, the impact of all risk propagation paths on the target entity can be comprehensively considered, and the comprehensive risk level of the target entity can be more accurately assessed.

[0117] As an implementation method, the risk propagation model training process may include the following steps:

[0118] Step S401: Obtain a historical risk event dataset, which includes event triggering entity identifiers, risk propagation path records, and actual impact range labels.

[0119] Historical risk event datasets contain information related to past risk events. Event trigger entity identifiers uniquely identify the entity that triggered the risk event, such as the company department number or device name. Risk propagation path records detail the path along which a risk propagates within the dynamic knowledge graph, including the nodes and edges it passes through. Actual impact scope labels denote the actual impact of a risk event, such as the affected geographic region or business level.

[0120] In this step, we collect and organize historical risk event datasets, ensuring they contain key information such as the entity that triggered the event, the risk propagation path, and labels for the actual impact scope. For example, we can extract information about all risk events that occurred over the past year from a company's risk log to form a historical risk event dataset. This data will serve as the foundation for training the risk propagation model, helping it learn the patterns and laws of risk propagation.

[0121] Step S402: Extract the training nodes corresponding to the event-triggered entity identifiers from the historical version of the dynamic knowledge graph, and generate a training path set based on the risk propagation path records.

[0122] The historical version of the dynamic knowledge graph refers to the state of the dynamic knowledge graph at different points in time. Training nodes are nodes corresponding to event-triggering entity identifiers extracted from the historical version of the dynamic knowledge graph. They serve as input for training the risk propagation model. The training path set is a set of paths generated from the risk propagation path records and used for model training.

[0123] In this step, first, the corresponding training nodes are extracted from the historical version of the dynamic knowledge graph according to the event-triggered entity identifier. For example, according to the entity type code and name in the event-triggered entity identifier, matching nodes are searched in the historical version of the dynamic knowledge graph. Then, a training path set is generated based on the risk propagation path record. The specific process includes parsing the entity jump sequence in the risk propagation path record, identifying the node identifier and jump direction corresponding to each entity in the sequence, and if there is no node matching the entity in the historical version of the dynamic knowledge graph, then the attribute filling rule of the virtual node is generated based on the contextual relationship of the adjacent entities; according to the node identifier, the connection edge list of the corresponding node is extracted from the historical version of the dynamic knowledge graph, and each jump action in the entity jump sequence is matched with the connection edge list. If there is an unrecorded jump edge, a temporary edge is inserted according to the attribute filling rule of the virtual node; the matched entity jump sequence is traversed to detect the loop structure of repeated access to the same node in continuous jumps. When the number of occurrences of a loop structure exceeds a preset tolerance threshold, the jump sequence after the loop structure is truncated and recorded as an invalid path segment. All jump sequences that do not contain invalid path segments are extracted as candidate training paths. The proportion of temporary edges in each candidate training path is counted. If the proportion exceeds a preset simulation threshold, a low-confidence label is added to the candidate training path. Stratified sampling is performed on the candidate training paths based on the low-confidence labels to ensure that the distribution ratio of paths at different confidence levels in the training and validation sets is consistent. Node attributes are perturbed on the sampled paths to generate adversarial example paths. The adversarial example paths are then merged with the original candidate training paths to form a training path set. This generates a rich and accurate training path set, improving the training effect of the risk propagation model.

[0124] In step S402, extracting the training node corresponding to the event-triggered entity identifier from the historical version of the dynamic knowledge graph may specifically include the following steps:

[0125] Step S4021: Based on the occurrence timestamp of the historical risk event, load the graph structure data of the corresponding time interval from the version snapshot library of the dynamic knowledge graph.

[0126] The timestamp of a historical risk event records the specific time when the risk event occurred. The version snapshot library of the dynamic knowledge graph is a database used to store version snapshots of the dynamic knowledge graph at different points in time. Graph structure data refers to information such as nodes, edges, and their attributes within a dynamic knowledge graph over a specific time period.

[0127] In this step, based on the timestamp of the historical risk event, the graph structure data for the corresponding time interval is loaded from the version snapshot library of the dynamic knowledge graph. For example, if the timestamp of a historical risk event is "2023-06-15 12:00:00," the graph structure data for the time interval containing this time point (e.g., "2023-06-15 00:00:00" to "2023-06-15 23:59:59") is loaded from the version snapshot library. This ensures that the extracted training nodes and training paths are consistent with the state of the dynamic knowledge graph at the time of the historical risk event, improving the accuracy of model training.

[0128] Step S4022: parse the node attributes of the event triggering entity identifier in the graph structure data. If there are multiple nodes with the same name, filter the multiple nodes with the same name based on the proximity between the node activity index and the event occurrence time to obtain the filtered nodes.

[0129] Node attributes refer to the various attribute information possessed by nodes in a dynamic knowledge graph, such as name, type, and activity. Node activity metrics are used to measure node activity, such as node update frequency and the number of connected edges. Event proximity refers to the time interval between a node's update time and the occurrence of a historical risk event.

[0130] In this step, the node attributes of the event triggering entity identifier in the graph structure data are parsed. If there are multiple nodes with the same name, these nodes with the same name are filtered out based on the proximity of the node activity index to the time of the event. For example, assuming that the event triggering entity identifier is "Department A", there are three nodes named "Department A" in the graph structure data. The activity index of these three nodes and the proximity to the time of the event are calculated respectively. The nodes with higher activity index and higher proximity to the time of the event are selected as the filtered nodes. This ensures that the filtered nodes are the nodes most relevant to the historical risk events, thereby improving the quality of the training nodes.

[0131] Step S4023: Perform feature enhancement processing on the selected nodes, and encode the degree centrality index, betweenness centrality index and timeliness attenuation factor of the selected nodes in the graph structure data into feature vectors.

[0132] The degree centrality index refers to the number of connections a node has with other nodes in the dynamic knowledge graph, reflecting the node's importance. The betweenness centrality index refers to the number of times a node acts as an intermediary for the shortest path between other nodes in the dynamic knowledge graph, reflecting the node's role as a bridge in information dissemination. The timeliness decay factor is a factor calculated based on the difference between the node's update time and the current time, reflecting the timeliness of the node's information. The feature vector is a vector representation obtained by encoding the various features of the node. In this step, feature enhancement processing is performed on the selected nodes. The degree centrality index, betweenness centrality index, and timeliness decay factor of the selected nodes in the graph structure data are calculated and encoded into feature vectors. For example, the degree centrality index, betweenness centrality index, and timeliness decay factor are used as different dimensions of the feature vector. This enriches the feature information of the node and improves the risk propagation model's ability to understand and process the node.

[0133] Step S4024: concatenate the feature vector with the original attribute vector of the node to generate an input feature representation of the training node.

[0134] The original attribute vector refers to the vector representation obtained by encoding the attribute information originally possessed by the node. By splicing the feature vector with the original attribute vector, the enhanced feature information can be combined with the original attribute information of the node to obtain a more comprehensive and accurate input feature representation of the training node. In this step, the feature vector generated previously is spliced ​​with the original attribute vector of the node. For example, assuming that the original attribute vector contains information such as the name and type of the node, and the feature vector contains information such as the degree centrality index, the betweenness centrality index, and the timeliness attenuation factor, they are spliced ​​together to form a longer vector as the input feature representation of the training node. This can provide richer information for the risk propagation model, helping the model to better learn and predict the risk propagation path.

[0135] In step S402, generating a training path set based on the risk propagation path record may specifically include the following steps:

[0136] Step S4025: parse the entity jump sequence in the risk propagation path record, identify the node identifier and jump direction corresponding to each entity in the sequence, and if there is no node matching the entity in the historical version of the dynamic knowledge graph, generate attribute filling rules for the virtual node based on the contextual relationship between adjacent entities.

[0137] The entity jump sequence is the order of jumps between entities recorded in the risk propagation path record. A node identifier uniquely identifies a node in the dynamic knowledge graph. A jump direction indicates the direction of the entity jump, such as forward or reverse. A virtual node is a node that does not exist in previous versions of the dynamic knowledge graph but is created based on contextual needs. Attribute population rules are used to determine the attributes of virtual nodes. In this step, the entity jump sequence in the risk propagation path record is parsed to identify the node identifier and jump direction corresponding to each entity. For example, if the entity jump sequence is "A->B->C," the node identifiers corresponding to entities A, B, and C are identified, and the jump direction is forward. If a node matching an entity, such as entity D, does not exist in previous versions of the dynamic knowledge graph, attribute population rules for the virtual node are generated based on the contextual relationships between adjacent entities (e.g., entities C and E). For example, if both entities C and E are related to business line X, it can be inferred that virtual node D is also related to business line X, and the attribute population rules for virtual node D are determined accordingly. This can handle the problem of missing nodes that may exist in the risk propagation path records and ensure the integrity of the training path set.

[0138] Step S4026: Extract the connection edge list of the corresponding node from the historical version of the dynamic knowledge graph according to the node identifier, match each jump action in the entity jump sequence with the connection edge list, and if there are unrecorded jump edges, insert temporary edges according to the attribute filling rules of the virtual node.

[0139] A connection edge list is a list of all edges connected to a node in the dynamic knowledge graph. A jump action is a jump operation between entities in an entity jump sequence. A temporary edge is an edge that is not recorded in the historical version of the dynamic knowledge graph but needs to be inserted based on context.

[0140] In this step, the connection edge list of the corresponding node is extracted from the historical version of the dynamic knowledge graph based on the node identifier. For example, for node A, its connection edge list is extracted. Then, each jump action in the entity jump sequence is matched with the connection edge list. For example, for the jump action "A->B", check whether there is an edge from node A to node B in the connection edge list. If there is an unrecorded jump edge, such as an edge from node A to node D, a temporary edge is inserted according to the attribute filling rules of the virtual node D. This ensures that the paths in the training path set can be fully represented in the dynamic knowledge graph, improving the accuracy of model training.

[0141] Step S4027: traverse the matched entity jump sequence and detect the loop structure that repeatedly accesses the same node in continuous jumps. When the number of occurrences of the loop structure exceeds a preset tolerance threshold, truncate the jump sequence after the loop structure and record it as an invalid path segment.

[0142] A loop structure refers to a situation where the same node is repeatedly visited during consecutive jumps in a physical jump sequence, for example, "A->B->A." The preset tolerance threshold is a pre-set threshold for the number of occurrences of a loop structure. An invalid path segment is a path segment containing a loop structure that exceeds the preset tolerance threshold.

[0143] In this step, the matched entity jump sequence is traversed to detect whether there is a loop structure that repeatedly visits the same node in consecutive jumps. For example, the jump sequence "A->B->A->C->A" is detected, and a loop structure that repeatedly visits node A is found. The number of occurrences of the loop structure is counted and compared with the preset tolerance threshold. If the number of occurrences of the loop structure exceeds the preset tolerance threshold, for example, the preset tolerance threshold is 1, and the loop structure appears 2 times, the jump sequence after the loop structure is truncated, the part after "A->B->A" is truncated, and recorded as an invalid path fragment. This can prevent the model from learning invalid loop paths and improve the quality of the training path set.

[0144] Step S4028: extract all jump sequences that do not contain invalid path segments as candidate training paths, count the proportion of temporary edges in each candidate training path, and add a low confidence label to the candidate training path if the proportion exceeds a preset simulation threshold.

[0145] A candidate training path is a jump sequence extracted from a matched entity jump sequence that does not contain invalid path segments. The temporary edge ratio is the ratio of the number of temporary edges to the total number of edges in a candidate training path. The preset simulation threshold is a pre-set critical value for the temporary edge ratio. A low confidence label is used to indicate that a candidate training path has low reliability.

[0146] In this step, all jump sequences that do not contain invalid path segments are extracted as candidate training paths. For example, sequences that do not contain invalid path segments are filtered out from the matched entity jump sequences. Then, the proportion of temporary edges in each candidate training path is counted. For example, a candidate training path has a total of 10 edges, of which 3 are temporary edges, and the proportion of temporary edges is 3 / 10=0.3. The proportion of temporary edges is compared with the preset simulation threshold. If the proportion exceeds the preset simulation threshold, for example, the preset simulation threshold is 0.2, a low confidence label is added to the candidate training path. In this way, candidate training paths of different reliabilities can be distinguished, providing a reference for subsequent stratified sampling and model training.

[0147] Step S4029: Perform stratified sampling on the candidate training paths according to the low-confidence labels to ensure that the distribution ratios of paths at different confidence levels in the training set and the validation set are consistent, and perform node attribute perturbations on the sampled paths to generate adversarial sample paths.

[0148] Stratified sampling is a method for sampling at different levels. Here, candidate training paths are divided into different confidence levels based on low-confidence labels. Samples are then drawn from each level according to a preset ratio. This ensures that paths of different confidence levels are distributed equally across the training and validation sets, preventing the model from learning paths of certain confidence levels poorly due to uneven data distribution. Node attribute perturbation, which slightly modifies the attributes of nodes in a path, generates adversarial example paths that are similar but different from the original path, enhancing the model's generalization and robustness.

[0149] In this step, the candidate training paths are stratified according to low-confidence labels. For example, paths with low-confidence labels are grouped into one layer, and paths without low-confidence labels are grouped into another layer. Next, samples are extracted from each layer according to a pre-set ratio to form a training set and a validation set, ensuring that the distribution ratio of paths at different confidence levels in the two sets is consistent. Afterwards, the node attributes of the sampled paths are perturbed. For example, for a node in the path, its attributes include node activity, the number of associated edges, etc. These attribute values ​​are slightly randomly changed to generate new node attributes, thereby obtaining an adversarial sample path. Taking a simple e-commerce scenario as an example, assuming that a candidate training path contains user nodes and product nodes, and the attributes of user nodes include purchase frequency, consumption amount, etc. These attributes are slightly randomly increased or decreased to generate new user node attributes, thereby obtaining an adversarial sample path.

[0150] Step S40210: Merge the adversarial sample path and the original candidate training path into a training path set.

[0151] The training path set is the set of all paths used to train the risk propagation model. Merging the adversarial sample paths with the original candidate training paths can enrich the training data, enable the model to learn more different types of path features, and improve the model's generalization ability and ability to handle complex situations.

[0152] In this step, the generated adversarial paths are combined with the original candidate training paths to form the final training path set. For example, if there are 100 original candidate training paths, and 20 adversarial paths are generated by perturbing node attributes, these 20 adversarial paths are added to the 100 original candidate training paths, resulting in a training path set containing 120 paths. This set will serve as an important data foundation for subsequent risk propagation model training.

[0153] Step S403: input the training nodes into the initial risk propagation model to generate a set of predicted risk propagation paths, and calculate the first training loss based on the path overlap between the predicted risk propagation path set and the training path set.

[0154] The initial risk propagation model is an incompletely trained model, and its parameters have not yet reached their optimal state. The predicted risk propagation path set is the set of risk propagation paths generated by the initial risk propagation model based on its own rules and parameters after the training nodes are input into the model. Path overlap refers to the proportion of identical paths in the predicted risk propagation path set and the training path set, and is used to measure the degree of similarity between the model's predictions and the actual situation. The first training loss is calculated based on the path overlap and is used to evaluate the model's accuracy in predicting risk propagation paths.

[0155] In this step, the training nodes generated previously are input into the initial risk propagation model. For example, the input feature representations of the training nodes are fed into the model's input layer. The model performs calculations based on its own structure and parameters to generate a set of predicted risk propagation paths. The path overlap between the predicted risk propagation path set and the training path set is then calculated. The number of identical paths can be determined by comparing the node sequences and edge connectivity of each path in the two sets, thereby calculating the path overlap. For example, if there are 50 paths in the predicted risk propagation path set and 60 paths in the training path set, and 30 of them are identical, the path overlap is 30 / min(50, 60) = 0.6. The first training loss is calculated based on the path overlap. A loss function, such as the mean squared error loss function, is typically used to quantify the difference between the path overlap and the ideal overlap (e.g., 1). This loss value is used to subsequently adjust model parameters to improve the accuracy of the model's risk propagation path predictions.

[0156] Step S404: Extract the node connection density and risk attribute superposition strength of each risk propagation path in the predicted risk propagation path set, calculate the predicted comprehensive risk value, and calculate the second training loss based on the difference between the predicted comprehensive risk value and the actual risk value recorded in the actual impact range label.

[0157] The calculation method of node connection density and risk attribute superposition strength is the same as the method described in the previous steps S510 and S520, that is, the node connection density is obtained by counting the ratio of the number of initial edges between adjacent nodes to the total number of hops in the path, the risk attribute label carried by each node in the path is extracted, the preset risk coefficient table is matched according to the label type, and the risk coefficients of the same type are accumulated to obtain the risk attribute superposition strength. The predicted comprehensive risk value is a comprehensive risk value calculated according to the method described in steps S530 and S540 based on the node connection density and risk attribute superposition strength of each path in the predicted risk propagation path set. The actual risk value recorded in the actual impact range label is a quantitative value of the risk level actually caused by the historical risk event. The second training loss is a loss value calculated based on the difference between the predicted comprehensive risk value and the actual risk value, which is used to evaluate the accuracy of the model in predicting the comprehensive risk value. In this step, the node connection density and risk attribute superposition strength of each risk propagation path are first extracted from the predicted risk propagation path set. For example, for a predicted risk propagation path, the number of initial edges between adjacent nodes and the total number of hops along the path are counted to calculate the node connection density. The risk attribute label for each node in the path is extracted, and the risk coefficients of the same type are accumulated to obtain the risk attribute stacking strength. Then, based on the preset density weight and stacking weight, the node connection density and the risk attribute stacking strength are weighted and summed to obtain the local risk contribution value for each path. The local risk contribution values ​​of all paths are then normalized and summed with an attenuated weight based on the path ranking results to obtain the predicted comprehensive risk value. Next, the true risk value is obtained from the actual impact range label. Finally, the difference between the predicted comprehensive risk value and the true risk value is calculated and quantified as a second training loss using an appropriate loss function, such as the absolute error loss function. This loss value is used together with the first training loss to adjust the model parameters to enable the model to more accurately predict the comprehensive risk value.

[0158] Step S405: The first training loss and the second training loss are merged into a total training loss according to a preset ratio, and the parameters of the initial risk propagation model are adjusted through back propagation until the total training loss converges to obtain a risk propagation model.

[0159] The preset ratio is a pre-determined ratio used to fuse the first and second training losses. By properly setting this ratio, the model's learning can be balanced in predicting risk propagation paths and overall risk values. The total training loss is the sum of the first and second training losses according to the preset ratio, reflecting the model's errors in both areas. Backpropagation is an algorithm used to adjust model parameters. It calculates the gradient of the parameters based on the total training loss and then updates the parameters in the opposite direction of the gradient, gradually reducing the total training loss. When the total training loss converges, it means that the model parameters have been adjusted to a relatively stable state, and the resulting model is now a risk propagation model.

[0160] In this step, the first and second training losses are first combined according to a preset ratio. For example, if the ratio of the first training loss is preset to 0.6 and the ratio of the second training loss is preset to 0.4, the total training loss = 0.6 × first training loss + 0.4 × second training loss. Then, the parameters of the initial risk propagation model are adjusted using the backpropagation algorithm. Specifically, the gradient of each parameter is calculated based on the total training loss, and the partial derivative of the loss function with respect to each parameter is calculated using the chain rule. Next, the parameters are updated in the opposite direction of the gradient, for example using the stochastic gradient descent algorithm. The parameter update formula is: parameter = parameter - learning rate × gradient, where the learning rate is a hyperparameter that controls the parameter update step size. This process is repeated until the total training loss converges. A convergence threshold can be set; when the change in the total training loss is less than this threshold, the total training loss is considered to have converged. At this point, the resulting model is a trained risk propagation model that can more accurately predict risk propagation paths and overall risk values.

[0161] As an implementation method, in step S405, adjusting the parameters of the initial risk propagation model by backpropagation until the total training loss converges may specifically include the following steps:

[0162] Step S4051: Obtain the path overlap difference between the predicted path set generated by the initial risk propagation model in the current training cycle and the training path set, and dynamically adjust the node traversal step threshold of the next training cycle based on the difference.

[0163] A training cycle is a complete parameter update process for the model. The path overlap difference is the difference between the path overlap between the predicted and training paths in the current training cycle and the path overlap in the previous training cycle. The node traversal step threshold is used to control the node traversal speed in the risk propagation model. It determines the number of nodes that can be visited during each traversal.

[0164] In this step, the path overlap between the predicted path set generated by the initial risk propagation model in the current training cycle and the training path set is first calculated. This path overlap is then subtracted from the path overlap in the previous training cycle to obtain the path overlap difference. This difference is used to dynamically adjust the node traversal step threshold for the next training cycle. For example, if the path overlap difference is positive, it indicates that the model's prediction performance is improving. In this case, the node traversal step threshold can be appropriately increased to speed up traversal and improve training efficiency. If the path overlap difference is negative, it indicates that the model's prediction performance is declining. In this case, the node traversal step threshold needs to be reduced to fine-tune model parameters. For a simple numerical example, assuming the path overlap in the previous training cycle was 0.5 and the path overlap in the current training cycle was 0.6, the path overlap difference is 0.1. If the current node traversal step threshold is 5, the node traversal step threshold for the next training cycle can be increased to 6 according to the preset adjustment rules.

[0165] Step S4052: monitor the decreasing slope of the total training loss within a preset number of consecutive training cycles. If the decreasing slope is lower than a preset stagnation threshold, trigger an early stopping signal and extract a snapshot of the current model parameters as intermediate candidate parameters.

[0166] The decline slope is the rate of change of the total training loss over a preset number of training cycles, reflecting the rate at which the total training loss is decreasing. The preset stagnation threshold is a pre-set critical value for the decline slope, used to determine whether the total training loss has stopped decreasing. The early stopping signal is used to stop model training. When the decline slope of the total training loss falls below the preset stagnation threshold, it indicates that the model's training performance is unlikely to improve further. Triggering the early stopping signal at this time can prevent overfitting. The intermediate candidate parameters are a snapshot of the current model's parameters extracted when the early stopping signal is triggered. They may represent a relatively optimal parameter combination.

[0167] In this step, the downward slope of the total training loss over a preset number of training cycles is continuously monitored. For example, the preset number is 10, and the rate of change of the total training loss over the last 10 training cycles is calculated. This downward slope is compared with the preset stagnation threshold. If the downward slope is lower than the preset stagnation threshold, for example, the preset stagnation threshold is 0.01, and the current downward slope is 0.005, an early stopping signal is triggered. At the same time, a parameter snapshot of the current model is extracted as an intermediate candidate parameter. In this way, training can be stopped in time when the model training effect is no longer significantly improved, avoiding over-training and overfitting of the model, while retaining a relatively optimal parameter combination to provide a basis for subsequent model optimization.

[0168] Step S4053: Load the intermediate candidate parameters into the path generation layer of the initial risk propagation model, recalculate the path overlap between the predicted path set and the training path set, and if the overlap exceeds the preset rollback threshold, freeze the parameters of the path generation layer and lock the generation logic of the node access sequence.

[0169] The path generation layer is the component of the risk propagation model used to generate risk propagation paths. After loading the intermediate candidate parameters into the path generation layer, the model regenerates a set of predicted paths based on these parameters. Path overlap improvement refers to the difference between the recalculated path overlap after loading the intermediate candidate parameters and the previous path overlap. The preset rollback threshold is a pre-set critical value for path overlap improvement, used to determine whether the path generation layer parameters need to be frozen. Freezing the path generation layer parameters means that these parameters will no longer be updated. Locking the node access sequence generation logic means fixing the order and rules for node access.

[0170] In this step, the intermediate candidate parameters are loaded into the path generation layer of the initial risk propagation model. Then, the path overlap between the predicted path set and the training path set is recalculated. This new path overlap is compared with the previous path overlap to calculate the improvement in path overlap. If the improvement in path overlap exceeds the preset rollback threshold, for example, the preset rollback threshold is 0.05, and the path overlap is improved to 0.06, the parameters of the path generation layer are frozen, that is, the parameters of the path generation layer are no longer updated. At the same time, the generation logic of the node access sequence is locked to ensure that the node access order and rules remain unchanged during subsequent training. This can fix the model's performance in path generation and focus on optimizing other parts.

[0171] Step S4054: In the parameter freezing state, only the node weight allocation module of the initial risk propagation model is opened for incremental training, and the priority rules of the node weight calculation function are updated by injecting new node connection relationships into the real-time data stream.

[0172] The parameter freeze state means that the parameters of the path generation layer are locked and no longer updated. The node weight assignment module is used to calculate node priorities in the risk propagation model. Incremental training refers to training parts of the model using new data based on the existing model. New node connections in the real-time data stream refer to new node connection information collected in real time over time during model training. The priority rule of the node weight calculation function is used to determine node priorities. By updating this rule, the model can better adapt to new node connections.

[0173] In this step, when the parameters of the path generation layer are frozen, only the node weight distribution module of the initial risk propagation model is opened for incremental training. The newly added node connection relationships in the real-time data stream are injected into the training process. For example, some new connection relationships between users and devices are collected in real time, and these relationships are input into the node weight distribution module as new data. The priority rules of the node weight calculation function are updated according to these newly added node connection relationships. For example, the original node weight calculation function only considered the confidence weight of the initial edge and the dynamic attenuation coefficient of the generated timestamp. Now, according to the newly added connection relationships, some new factors such as the stability of the connection, the frequency of the connection, etc. can be added to adjust the priority rules. This enables the node weight distribution module of the model to adapt to the new node connection situation in a timely manner and improve the performance of the model.

[0174] Step S4055: When the incremental training loss of the node weight distribution module continuously reaches a stable state, the frozen parameters of the path generation layer are combined with the updated node weight calculation function for verification to generate the final model parameters and export the path traversal configuration template that is compatible with the dynamic knowledge graph version.

[0175] Incremental training loss refers to the loss value calculated during the incremental training process of the node weight assignment module. A stable state refers to the change in incremental training loss over multiple consecutive training cycles being less than a preset stability threshold, indicating that the model's training effect at this stage has stabilized. Combined validation combines the frozen parameters of the path generation layer with the updated node weight calculation function, validates the model using a validation set, and evaluates the model's overall performance. The final model parameters are the optimal parameter combination determined after combined validation. The path traversal configuration template is a template compatible with the dynamic knowledge graph version for guiding risk propagation path traversal. It contains information such as node priority rules and hop count constraints.

[0176] In this step, the incremental training loss of the node weight distribution module is continuously monitored. When the loss continuously reaches a stable state, the frozen parameters of the path generation layer are combined with the updated node weight calculation function. The combined model is verified using the validation set, and the various performance indicators of the model on the validation set are calculated, such as path overlap, comprehensive risk value prediction accuracy, etc. Based on the verification results, the parameters are adjusted until the optimal parameter combination, i.e., the final model parameters, is obtained. Then, based on the final model parameters, a path traversal configuration template that is compatible with the dynamic knowledge graph version is derived. For example, information such as node priority rules and hop count constraints is organized into a configuration file as a path traversal configuration template. This template will be used for subsequent online risk propagation path generation to ensure the accuracy and stability of the model in practical applications.

[0177] Step S4056: Synchronize the node priority rules and hop count constraints in the path traversal configuration template to the risk warning signal trigger module to keep the online risk propagation path generation consistent with the rules of offline model training.

[0178] The risk warning signal trigger module is used to trigger risk warning signals based on the risk propagation path and the overall risk value. Synchronizing the node priority rules and hop count constraints in the path traversal configuration template to this module ensures that the rules used during online risk propagation path generation are the same as those used during offline model training, thereby ensuring the accuracy and consistency of the model in actual applications. In this step, the node priority rules and hop count constraints in the path traversal configuration template are extracted and synchronized to the risk warning signal trigger module. For example, the node priority rules can be embedded in the program of the risk warning signal trigger module in the form of code, and the hop count constraints can be passed to the module as parameters. In this way, during the actual online risk propagation path generation process, the risk warning signal trigger module will perform path traversal and risk assessment according to the same rules as offline model training. When the overall risk value exceeds the preset threshold, the risk warning signal is accurately triggered. In this way, the stability and reliability of the entire enterprise data risk processing system can be guaranteed, and the model training results can be effectively applied in real-world scenarios.

[0179] Step S500: After the risk warning signal is triggered, the method provided by the embodiment of the present invention may further include:

[0180] Step S600: Based on the preceding path segment with the highest node connection density in the risk propagation path set, extract the upstream node set that has a direct dependency relationship with the target entity, and identify the resource scheduling authority identifier corresponding to each node in the upstream node set.

[0181] The preceding path segment with the highest node connection density refers to the preceding path segment within the risk propagation path set, where the node connection density is high. A direct dependency relationship refers to a situation where the state or behavior of one node directly affects the state or behavior of another node. An upstream node refers to a node that precedes the target entity in the risk propagation path and has a direct dependency relationship with the target entity. A resource scheduling permission identifier uniquely identifies a node's permissions for enterprise resource scheduling, such as a permission code or name.

[0182] In this step, the preceding path segments with the highest node connection density are screened out from the risk propagation path set. For example, by calculating the node connection density of each risk propagation path, the preceding segments of the paths with the highest connection density are selected as the preceding path segments with the highest node connection density. Then, based on these path segments, the upstream node set that has a direct dependency relationship with the target entity is extracted. For example, in an enterprise's supply chain risk propagation path, the target entity is a production workshop. By analyzing the preceding path segments with the highest node connection density, it is found that the raw material supplier nodes have a direct dependency relationship with the production workshop, and these raw material supplier nodes are used as the upstream node set. Next, the resource scheduling authority identifiers corresponding to each node in the upstream node set are identified. For example, for each raw material supplier node, the enterprise's authority management system is queried to obtain its authority identifiers for resource scheduling such as raw material procurement and transportation. These resource scheduling authority identifiers will provide an important basis for subsequent risk isolation and resource scheduling.

[0183] Step S700: Match the preset risk isolation rules from the enterprise policy library according to the resource scheduling authority identifier, and generate a resource access flow limiting threshold and an operation prohibition list for each node in the upstream node set.

[0184] The enterprise policy library is a database of pre-defined policies and rules, including risk isolation rules for different risk scenarios and resource scheduling permissions. Resource access limit thresholds limit the amount of traffic or quantity a node can access when accessing enterprise resources, such as limiting the amount of raw materials a node can purchase within a set timeframe. An operation ban list is a list of prohibited operations for a node, such as prohibiting a node from performing equipment upgrades during a risk period.

[0185] In this step, based on the resource scheduling authority identifier corresponding to each node in the upstream node set, the matching preset risk isolation rules are searched in the enterprise policy library. For example, for a raw material supplier node with resource scheduling authority, the risk isolation rules for this authority level are found in the enterprise policy library. Then, based on these rules, the resource access flow limiting threshold and operation ban list for each node are generated. For example, according to the rules, for this raw material supplier node, its raw material procurement volume will be limited to 50% of the normal situation during the risk period, as the resource access flow limiting threshold; at the same time, it is prohibited to conduct new supplier cooperation negotiations during the risk period, and this operation is included in the operation ban list. In this way, the resource access and operations of upstream nodes can be effectively restricted, reducing the spread of risks.

[0186] Step S800: Collect the latest connection edge update records of each node in the upstream node set in the dynamic knowledge graph in real time. If it is detected that the newly added connection edge involves a restricted operation type in the operation ban list, a dynamic downward adjustment instruction of the resource access current limiting threshold is triggered.

[0187] The latest edge update record refers to the latest changes in the edges of each node in the upstream node set in the dynamic knowledge graph, including newly added and deleted edges. Restricted operation types refer to the types of operations prohibited by the operation ban list. The dynamic lowering instruction for resource access throttling threshold refers to the instruction to lower the resource access throttling threshold when a newly added edge involving a restricted operation type is detected, thereby further restricting the node's resource access and reducing risk.

[0188] In this step, the latest connection edge update records of each node in the upstream node set in the dynamic knowledge graph are collected in real time. For example, by monitoring the update log of the dynamic knowledge graph, the changes in the connection edges of each node are obtained. Then, check whether the newly added connection edges involve restricted operation types in the operation ban list. For example, the operation ban list prohibits nodes from performing equipment leasing operations. If it is detected that a certain upstream node has newly added a connection edge related to equipment leasing, it means that the node involves a restricted operation type. At this time, the dynamic downward adjustment instruction of the resource access flow limit threshold is triggered. For example, the raw material procurement flow limit threshold of the node is further reduced from 50% under normal circumstances to 30% to enhance the risk isolation effect.

[0189] Step S900: Compare the dynamic downgrade instruction with the historical access frequency of the current node. When the historical access frequency exceeds the adjusted resource access current limit threshold, send a node service downgrade request to the associated business system and record the downgrade timestamp.

[0190] Historical access frequency refers to the number of times a node has accessed enterprise resources over a period of time. Associated business systems are enterprise business systems related to upstream nodes, such as procurement and production systems. A node service downgrade request is a request sent to the associated business system to downgrade the node's service level when the node's historical access frequency exceeds the adjusted resource access limit threshold. This request may involve restricting the node's procurement permissions or reducing its production tasks. The downgrade timestamp records the time when the node's service downgrade occurred.

[0191] In this step, the resource access limit threshold adjusted in the dynamic downgrade instruction is compared with the historical access frequency of the current node. For example, the adjusted raw material procurement limit threshold is 30%, and the raw material procurement volume of this node in the past month has reached 40% of the normal level, indicating that the historical access frequency has exceeded the adjusted resource access limit threshold. At this time, a node service downgrade request is sent to the associated business system. For example, a request is sent to the procurement system to restrict the node to only carry out emergency raw material procurement for a period of time in the future. At the same time, the downgrade timestamp is recorded to facilitate subsequent tracing and analysis of the risk disposal process.

[0192] Step S1000: backtracking the affected path segments in the risk propagation path set according to the downgrade timestamp, removing the path branches containing the downgraded nodes and recalculating the comprehensive risk value of the target entity.

[0193] An affected path segment is a path segment within a risk propagation path set that is affected by a node service degradation. A path branch is a portion of a risk propagation path that contains a degraded node. Recalculating the comprehensive risk value involves recalculating the target entity's comprehensive risk value according to the method described in step S500 after removing the path branch containing the degraded node to assess the effectiveness of the risk isolation measures.

[0194] In this step, the affected path segments are traced back in the risk propagation path set according to the downgrade timestamp. For example, by finding the update timestamps of the nodes in the risk propagation path, determine which path segments are affected by the node service degradation. Then, remove the path branches containing the downgraded nodes. For example, in a risk propagation path, a raw material supplier node is downgraded, and the node and its subsequent path branches are removed from the risk propagation path set. Finally, recalculate the comprehensive risk value of the target entity. According to the previous method, the node connection density and risk attribute superposition intensity of the remaining paths are counted, the local risk contribution value is calculated, and normalization and attenuation weighted summation are performed to obtain the updated comprehensive risk value. By recalculating the comprehensive risk value, the effect of the node service degradation measures on reducing the risk of the target entity can be evaluated.

[0195] Step S1100: The updated comprehensive risk value is associated with the downgrade operation record and stored in the risk disposal case library, and a priority improvement mark is generated for the nodes in the upstream node set that have not triggered downgrade to optimize the subsequent path traversal order.

[0196] The Risk Disposal Case Library is a database used to store information related to enterprise risk disposal, including descriptions of risk events, disposal measures, and disposal results. Degradation operation records contain information related to node service degradation, such as degradation timestamp, degradation reason, and degradation measures. Priority Boosting tags are used to mark nodes in the upstream node set that have not triggered degradation, increasing their priority in subsequent risk propagation path traversals to more rapidly identify potential risks.

[0197] In this step, the updated comprehensive risk value is associated with the downgrade operation record and stored in the risk disposal case library. For example, the updated comprehensive risk value, downgrade timestamp, downgrade measures and other information are organized into a record and stored in the corresponding table of the risk disposal case library. This can provide a reference for the company's subsequent risk disposal and accumulate risk disposal experience. At the same time, a priority increase mark is generated for the nodes in the upstream node set that have not triggered downgrades. For example, a priority increase attribute mark is added to these nodes in the dynamic knowledge graph. In the subsequent risk propagation path traversal process, these marked nodes are visited first, the path traversal order is optimized, and the efficiency of risk discovery is improved. In this way, the company's risk handling mechanism can be continuously improved and the company's ability to deal with data risks can be enhanced.

[0198] Please refer to Figure 3 , is a block diagram of the structure of the computer system 120 of the present application. The computer system 120 includes a computing unit 1001, which can perform various appropriate actions and processes according to a computer program stored in a ROM (read-only memory) 1002 or a computer program loaded from a storage unit 1008 into a RAM (random access memory) 1003. The RAM 1003 can also store various programs and data required for the operation of the computer system 120. The computing unit 1001, ROM 1002, and RAM 1003 are connected to each other via a bus 1004. An input / output (I / O) interface 1005 is also connected to the bus 1004.

[0199] Multiple components in the computer system 120 are connected to the I / O interface 1005, including an input unit 1006, an output unit 1007, a storage unit 1008, and a communication unit 1009. The input unit 1006 can be any type of device capable of inputting information into the computer system 120. The input unit 1006 can receive input digital or character information and generate key signal input related to user settings and / or function control of the server. The output unit 1007 can be any type of device capable of presenting information. The storage unit 1008 can include, but is not limited to, a magnetic disk or an optical disk. The communication unit 1009 allows the computer system 120 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks. The computing unit 1001 can be various general-purpose and / or specialized processing components with processing and computing capabilities, such as a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1001 performs the various methods and processes described above, such as the enterprise data risk processing method based on the dynamic knowledge graph. For example, in some embodiments, the enterprise data risk processing method based on the dynamic knowledge graph can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit 1008. In some embodiments, part or all of the computer program can be loaded and / or installed on the computer system 120 via the ROM 1002 and / or the communication unit 1009. When the computer program is loaded into the RAM 1003 and executed by the computing unit 1001, one or more steps of the enterprise data risk processing method based on the dynamic knowledge graph described above can be performed. Alternatively, in other embodiments, the computing unit 1001 can be configured to execute the enterprise data risk processing method based on the dynamic knowledge graph by any other appropriate means (for example, by means of firmware).

[0200] The program code for implementing the method of the present application can be written in any combination of one or more programming languages. Such program code can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented.

[0201] That is, the computer system provided by an embodiment of the present invention includes at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the above method.

Claims

1. A method for handling enterprise data risks based on dynamic knowledge graph, characterized in that: The method comprises: The method collects real-time data streams generated during the operation of an enterprise, extracts entity objects and association relationships in the real-time data streams, and forms initial nodes and initial edges of a dynamic knowledge graph. The method specifically comprises: receiving the real-time data streams through a distributed data acquisition interface, the real-time data streams comprising enterprise business system logs, equipment status monitoring data, and external market environment data; performing multimodal feature analysis on the real-time data streams, identifying operation subject features in the enterprise business system logs, device identification features in the equipment status monitoring data, and market entity features in the external market environment data, and mapping the operation subject features, the device identification features, and the market entity features into entity objects in the dynamic knowledge graph; generating an operation dependency relationship between the operation subject features and the device identification features based on the behavior sequence of the operation subject features within a preset time window, and generating a market impact relationship between the market entity features and the operation subject features based on the fluctuation trend of the market entity features in historical data; encoding the operation dependency relationship and the market impact relationship into initial edges connecting corresponding entity objects in the dynamic knowledge graph, and recording the generation timestamp and confidence weight of each initial edge; Receive a risk handling request, wherein the risk handling request includes a risk attribute identifier of a target entity and a description of a risk impact scope; According to the risk attribute identifier, the initial node corresponding to the target entity is located from the dynamic knowledge graph, and the associated propagation path generation rule of the target entity is determined based on the risk impact range description; specifically comprising: parsing the entity type code in the risk attribute identifier, and screening a set of candidate nodes whose node attributes match the entity type code from the dynamic knowledge graph; performing spatial range filtering and time attenuation weighting on the candidate node set according to the geographical area constraints, business level constraints and timeliness constraints in the risk impact range description to obtain the initial node corresponding to the target entity; extracting the propagation direction indicator and path depth limit parameter in the risk impact range description to generate the direction constraint condition and the hop count constraint condition in the associated propagation path generation rule; configuring the node weight calculation function in the associated propagation path generation rule, and the node weight calculation function calculates the path selection priority between adjacent nodes based on the confidence weight of the initial edge and the dynamic attenuation coefficient of the generated timestamp; Call the pre-trained risk propagation model, take the initial node as the starting point, traverse the adjacent nodes in the dynamic knowledge graph according to the associated propagation path generation rules, and generate a risk propagation path set corresponding to the target entity; specifically including: inputting the feature vector of the initial node into the path generation layer of the risk propagation model, and the path generation layer determines the accessible adjacent node list of the current node according to the direction constraint; through the weight allocation module of the risk propagation model, each node in the accessible adjacent node list is given a priority score according to the node weight calculation function, and a node access sequence is generated according to the score result; according to the hop count constraint, the node access operation is recursively performed in the path generation layer until the preset path depth is reached or the accessible adjacent node list is empty, and the path fragment generated each time recursively is recorded; the continuous node sequence that meets the direction constraint in all path fragments is merged to generate the risk propagation path set, and each path in the set is sorted according to the comprehensive priority; According to the node connection density and risk attribute superposition strength of each risk propagation path in the risk propagation path set, the comprehensive risk value corresponding to the target entity is calculated, and a risk warning signal is triggered when the comprehensive risk value exceeds a preset threshold; specifically, the method includes: counting the ratio of the number of initial edges between adjacent nodes in each risk propagation path to the total number of path hops as the node connection density; extracting the risk attribute label carried by each node in the risk propagation path, matching the preset risk coefficient table according to the label type, and accumulating the same type of risk coefficients along the path to obtain the risk attribute superposition strength; according to the preset density weight and superposition weight, performing weighted summation on the node connection density and the risk attribute superposition strength to obtain the local risk contribution value of each risk propagation path; normalizing the local risk contribution values ​​of all risk propagation paths in the risk propagation path set, and performing attenuated weighted summation according to the path sorting result to generate the comprehensive risk value corresponding to the target entity.

2. The method according to claim 1, characterized in that The training process of the risk propagation model includes: Obtain a historical risk event dataset, which includes event triggering entity identifiers, risk propagation path records, and actual impact range labels; Extracting the training node corresponding to the event triggering entity identifier from the historical version of the dynamic knowledge graph, and generating a training path set based on the risk propagation path record; Inputting the training node into an initial risk propagation model to generate a set of predicted risk propagation paths, and calculating a first training loss based on a degree of path overlap between the set of predicted risk propagation paths and the set of training paths; Extracting the node connection density and risk attribute superposition strength of each risk propagation path in the predicted risk propagation path set, calculating a predicted comprehensive risk value, and calculating a second training loss based on the difference between the predicted comprehensive risk value and the actual risk value recorded in the actual impact range label; The first training loss and the second training loss are fused into a total training loss according to a preset ratio, and the parameters of the initial risk propagation model are adjusted through back propagation until the total training loss converges, thereby obtaining the risk propagation model.

3. The method according to claim 2, characterized in that The step of extracting the training node corresponding to the event-triggered entity identifier from the historical version of the dynamic knowledge graph includes: According to the occurrence timestamp of the historical risk event, the graph structure data of the corresponding time interval is loaded from the version snapshot library of the dynamic knowledge graph; Parsing the node attributes of the event triggering entity identifier in the graph structure data, if there are multiple nodes with the same name, filtering the multiple nodes with the same name according to the proximity between the node activity index and the event occurrence time to obtain the filtered nodes; Performing feature enhancement processing on the selected nodes, encoding the degree centrality index, betweenness centrality index and timeliness attenuation factor of the selected nodes in the graph structure data into feature vectors; The feature vector is concatenated with the original attribute vector of the node to generate an input feature representation of the training node.

4. The method according to claim 2, characterized in that The generating of a training path set based on the risk propagation path record includes: Parse the entity jump sequence in the risk propagation path record, identify the node identifier and jump direction corresponding to each entity in the sequence, and if there is no node matching the entity in the historical version of the dynamic knowledge graph, generate attribute filling rules for virtual nodes based on the contextual relationship between adjacent entities; Extracting a connection edge list of the corresponding node from the historical version of the dynamic knowledge graph according to the node identifier, matching each jump action in the entity jump sequence with the connection edge list, and inserting a temporary edge according to the attribute filling rule of the virtual node if there is an unrecorded jump edge; Traverse the matched entity jump sequence and detect the loop structure that repeatedly visits the same node in continuous jumps. When the number of occurrences of the loop structure exceeds a preset tolerance threshold, truncate the jump sequence after the loop structure and record it as an invalid path segment. Extract all jump sequences that do not contain invalid path segments as candidate training paths, count the proportion of temporary edges in each candidate training path, and add a low confidence label to the candidate training path if the proportion exceeds a preset simulation threshold; Performing stratified sampling on candidate training paths based on the low-confidence labels to ensure that the distribution ratios of paths at different confidence levels in the training set and the validation set are consistent, and performing node attribute perturbations on the sampled paths to generate adversarial sample paths; The adversarial sample path and the original candidate training path are merged into the training path set.

5. The method according to claim 2, characterized in that The adjusting the parameters of the initial risk propagation model by backpropagation until the total training loss converges includes: Obtaining the path overlap difference between the predicted path set generated by the initial risk propagation model in the current training cycle and the training path set, and dynamically adjusting the node traversal step threshold for the next training cycle based on the difference; monitoring a decreasing slope of the total training loss over a preset number of consecutive training cycles, and if the decreasing slope is lower than a preset stagnation threshold, triggering an early stopping signal and extracting a snapshot of the current model parameters as intermediate candidate parameters; Loading the intermediate candidate parameters into the path generation layer of the initial risk propagation model, recalculating the path overlap between the predicted path set and the training path set, and if the overlap exceeds a preset rollback threshold, freezing the parameters of the path generation layer and locking the generation logic of the node access sequence; In the parameter freezing state, only the node weight allocation module of the initial risk propagation model is opened for incremental training, and the priority rules of the node weight calculation function are updated by injecting new node connection relationships into the real-time data stream; When the incremental training loss of the node weight distribution module continuously reaches a stable state, the frozen parameters of the path generation layer are combined with the updated node weight calculation function for verification to generate the final model parameters and export a path traversal configuration template compatible with the dynamic knowledge graph version; The node priority rules and hop count constraints in the path traversal configuration template are synchronized to the risk warning signal trigger module to keep the rules for online risk propagation path generation consistent with offline model training.

6. A computer system, characterized in that: include: at least one processor; and a memory communicatively coupled to the at least one processor; The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Automatic construction method of end-to-end agent based on graph structure semantic fusion

    CN120235181A

  • Knowledge graph generation method and system for science and technology project risk control

    CN120296180A