Privacy protection of cross-institutional data sharing-oriented bad asset risk prediction system
Patent Information
- Application Number
- CN202610797476.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-04
- Publication Date
- 2026-09-25
AI Technical Summary
[0005]本申请的目的在于提供一种面向跨机构数据共享的隐私保护不良资产风险预测系统,以至少解决现有技术在可信执行环境中处理大规模不良资产担保关系图时,因安全内存容量受限而导致页面换入换出频繁、风险预测计算效率较低,以及页面访问模式可能暴露图拓扑结构、担保链路或风险传播路径的问题
1.本申请通过建立不良资产担保关系图数据与内存页之间的映射关系,并结合图神经网络当前计算阶段确定候选访问页面集合,使内存调度不再仅依赖页面热度或最近访问时间,而是能够结合图拓扑关系和风险预测计算过程进行页面调度,有利于降低无效页面换入换出,提高可信执行环境中图神经网络风险预测计算的效率。
Smart Images

Figure CN122818399A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the interdisciplinary fields of data security and artificial intelligence, and in particular to a privacy-protected non-performing asset risk prediction system for cross-institutional data sharing. Background Technology
[0002] In cross-institutional joint risk identification scenarios such as financial risk control and non-performing asset management, a Trusted Execution Environment (TEE) is typically used to perform graph neural network calculations on the guarantee relationship graph constructed from shared data, in order to uncover potential default risks within a relatively safe space of hardware isolation.
[0003] However, the available secure memory capacity of a TEE is typically limited. Graph neural networks exhibit strong topological dependence and require frequent cross-page data access when performing node sampling and message passing. This not only leads to numerous page faults and page swaps, severely reducing the efficiency of risk prediction computation, but also allows untrusted external environments to infer key guarantor nodes or risk propagation paths by observing external characteristics such as page fault frequency and access concentration, triggering side-channel leakage risks.
[0004] Existing memory scheduling methods typically rely solely on page access frequency, failing to fully consider the topological relationships of the bad asset graph and struggling to balance data anonymization with secure memory budget constraints when introducing intelligent prefetching. Furthermore, the random noise masquerading methods used in existing side-channel defenses deviate from the true graph traversal patterns, resulting in limited defense effectiveness and excessive overhead. Therefore, there is an urgent need for a privacy protection system within a TEE that balances graph computation efficiency with side-channel defense. Summary of the Invention
[0005] The purpose of this application is to provide a privacy-preserving non-performing asset risk prediction system for cross-institutional data sharing, so as to at least solve the problems of existing technologies in processing large-scale non-performing asset guarantee relationship graphs in trusted execution environments, such as frequent page switching due to limited secure memory capacity, low efficiency of risk prediction calculation, and the possibility that page access patterns may expose graph topology, guarantee links, or risk propagation paths.
[0006] To achieve the above objectives, this application provides a privacy-preserving non-performing asset risk prediction system for cross-institutional data sharing. The system is deployed in a computing architecture that includes an untrusted environment and a Trusted Execution Environment (TEE), and is used to process encrypted non-performing asset guarantee relationship graph data constructed based on cross-institutional shared data. The system includes: A graph neural network computing engine and memory scheduler deployed within the Trusted Execution Environment (TEE); and a large language model that communicates with the TEE via a secure channel. The memory scheduler is configured as follows: Based on the mapping relationship between the non-performing asset guarantee relationship graph data and memory pages, the set of candidate access pages for the graph neural network computing engine in the risk prediction calculation process is determined. Calculate the performance gain metrics and side-channel risk metrics for each memory page in the candidate access page set; Based on the performance gain indicators and the side channel risk indicators, real-time page scheduling is performed within the Trusted Execution Environment (TEE) through a lightweight prefetch decision mechanism. The system sends a state observation vector, which includes the performance gain metric, the side channel risk metric, and the candidate access page set, to the large language model at an asynchronous period, after page-level anonymization processing, and receives prefetch bias parameters generated by the large language model. The prefetch bias parameters are subjected to security verification, and the boundary conditions of the lightweight prefetch decision mechanism are dynamically updated based on the verified prefetch bias parameters in order to perform scheduling on memory pages. When the side-channel risk index exceeds a preset risk threshold, based on the guarantee association node, guarantee link or graph community of the graph node corresponding to the accessed memory page, a set of fake crawled pages that matches the message passing rules of the graph neural network computing engine is generated, and fake crawling action is performed to confuse the memory access patterns observable in the untrusted environment through the simulated graph traversal sequence, thereby protecting the topological privacy of the bad asset guarantee relationship graph data.
[0007] Optionally, the system further includes a page mapping table, which is used to record the correspondence between graph nodes, edge data, node feature data and memory pages in the non-performing asset guarantee relationship graph data.
[0008] Optionally, the page mapping table includes node identifiers, page identifiers, organizational domain labels, node feature storage locations, adjacency list storage locations, and adjacent page sets; the memory scheduler determines the candidate access page set for the next computation stage from the adjacent page set based on the current graph neural network computation layer number, the current batch node set, and the sampling fan-out number.
[0009] Optionally, the performance benefit metric includes risk topology-page dwell density, which is calculated based at least on the node degree of the graph nodes contained in the memory page and the neighbor expansion probability in the current graph neural network computation stage, and in combination with at least one of the following: feature vector centrality, guarantee amount weight, guarantee link depth, and graph distance with historical bad nodes.
[0010] Optionally, the side-channel risk indicator includes a side-channel exposure index, which is calculated based at least on the page fault frequency of memory pages within a preset time window, and in combination with at least one of the following: page fault burst rate, page access interval variance, and correlation between page access sequence and graph topology path.
[0011] Optionally, the memory scheduler is configured to calculate a comprehensive scheduling utility value and perform scheduling on memory pages based on the comprehensive scheduling utility value, which is calculated based on the performance benefit metric, the confidence level of the prefetch bias parameter, the side-channel risk metric, and the cost-effectiveness ratio of page swapping in and out.
[0012] Optionally, the state observation vector is a de-identified page-level state observation vector, which includes anonymized page identifier, page dwell state, performance benefit index, side-channel risk index, memory budget, current graph neural network computation stage, and candidate access page set, but does not include the original financial plaintext data and the complete guarantee relationship graph.
[0013] Optionally, the prefetch bias parameters include target page identifier, prefetch priority, confidence level, effective time window, and reason code; the security verification includes candidate set verification, organization domain permission verification, memory budget verification, and effective time window verification.
[0014] Optionally, when the prefetch bias parameter fails the security check, or when the large language model does not return the prefetch bias parameter within the asynchronous update cycle, the memory scheduler maintains the boundary conditions of the current lightweight prefetch decision mechanism unchanged and continues to perform page scheduling based on the mapping relationship and the performance benefit metric.
[0015] Optionally, the spoofing action includes tiered spoofing: when the side-channel risk indicator is in the first risk range, the memory scheduler selects cold pages from the set of low-access-hot pages to perform spoofing; when the side-channel risk indicator is in the second risk range, the memory scheduler selects the page containing the first-order or second-order guarantee association node based on the first-order or second-order guarantee association node of the graph node corresponding to the accessed memory page to perform spoofing; when the side-channel risk indicator is in the third risk range, the memory scheduler generates a spoofing access path based on the guarantee link or graph community in the non-performing asset guarantee relationship graph data, and initiates spoofing on multiple memory pages corresponding to the spoofing access path according to the access rhythm matching the message passing stage of the graph neural network computing engine.
[0016] Compared with the prior art, the technical solution provided in this application has the following beneficial effects: 1. This application establishes a mapping relationship between non-performing asset guarantee relationship graph data and memory pages, and determines the candidate access page set by combining the current computing stage of the graph neural network. This enables memory scheduling to no longer rely solely on page popularity or recent access time, but to perform page scheduling by combining graph topology relationships and risk prediction calculation process. This helps reduce invalid page swapping in and out, and improves the efficiency of graph neural network risk prediction calculation in trusted execution environments.
[0017] 2. This application calculates performance gain indicators and side-channel risk indicators, and uses them together with security-verified prefetch bias parameters for memory page scheduling. This enables the system to balance page access performance and side-channel risk control under limited secure memory conditions, reducing the possibility of page access pattern leakage graph topology, key guarantee nodes, or risk propagation paths.
[0018] 3. This application limits the state observation vector sent to the large language model to the de-identified page-level state observation vector, and performs security checks on the prefetch bias parameters generated by the large language model, such as candidate set, institutional domain permissions, memory budget, and effective time window. This can reduce the risk of improper exposure or bypassing of original financial data, complete guarantee relationship graph, and cross-institutional data permissions while using the large language model to assist prefetching.
[0019] 4. When the side-channel risk increases, this application generates a set of fake crawled pages based on the graph topology relationship of the bad asset guarantee relationship graph. In particular, it can generate a simulated graph traversal fake access based on first-order or second-order guarantee association, guarantee link or graph community. This is closer to the real access pattern in graph neural network calculation than random noise page access, which is conducive to improving the access pattern confusion effect and reducing the additional overhead caused by excessive random noise access. Attached Figure Description
[0020] Figure 1 This is a schematic diagram of the system operation process provided in the embodiments of this application; Figure 2 This is a schematic diagram of the side-channel exposure index calculation and defense triggering logic provided in the embodiments of this application. Detailed Implementation
[0021] To better understand the technical solutions of this application, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. It is understood that the specific embodiments described herein are only for explaining this application and not for limiting this application.
[0022] Example 1 This embodiment provides a privacy-preserving non-performing asset risk prediction system for cross-institutional data sharing. This embodiment aims to fully demonstrate the technical solution of the system under its basic workflow, addressing the dual challenges of performance and security when performing large-scale non-performing asset collateral relationship graph calculations within a Trusted Execution Environment (TEE).
[0023] The system in this embodiment is deployed within a hybrid computing architecture, which logically includes an untrusted environment and a Trusted Execution Environment (TEE). The untrusted environment can be a host operating system, external memory, a regular database environment, or other execution environments not protected by hardware isolation mechanisms. The TEE is used to provide isolation protection for code and data in sensitive computing processes, thereby reducing the risk of the untrusted environment directly obtaining plaintext financial data.
[0024] In untrusted environments, a graph database is primarily deployed. This graph database stores encrypted non-performing asset guarantee relationship graph data. In scenarios involving financial risk control, non-performing asset management, or cross-institutional joint risk identification, the non-performing asset guarantee relationship graph data can be constructed from data provided by multiple participating institutions. This data may include enterprise operating data, guarantee relationship data, counter-guarantee relationship data, asset disposal data, repayment records, risk tags, and other data related to non-performing asset risk prediction. Before entering the graph database, the data from each participating institution can undergo cleaning, alignment, de-identification, and encryption to ensure the confidentiality of the data in its static storage state.
[0025] Within the Trusted Execution Environment (TEE), the system's core processing unit is deployed. This core processing unit includes a graph neural network computing engine, a memory scheduler, and an encryption / decryption engine. The graph neural network computing engine performs computational tasks such as node sampling, neighbor aggregation, message passing, graph representation learning, and risk probability prediction on the non-performing asset collateral relationship graph data. The memory scheduler schedules the limited secure memory resources within the TEE based on the graph data's topological characteristics, page-level performance gains, side-channel risk status, and prefetch bias parameters returned by the large language model. The encryption / decryption engine decrypts ciphertext data pages loaded into the TEE from the untrusted environment and encrypts relevant data when it needs to be written back to the untrusted environment.
[0026] Furthermore, the Trusted Execution Environment (TEE) also maintains a page mapping table. This page mapping table can be maintained by the memory scheduler or implemented as an internal data structure of the memory scheduler. The page mapping table records the correspondence between graph nodes, edge data, node feature data, and memory pages in the non-performing asset guarantee relationship graph data. Specifically, the page mapping table may include information such as node identifiers, page identifiers, institutional domain labels, node feature storage locations, adjacency list storage locations, and adjacent page sets.
[0027] For example, for a given enterprise node, its node features can be stored in one or more node feature pages, its guarantee relationship edges can be stored in edge data pages, and its first- or second-order guarantee association nodes can be obtained through adjacency list page indexes. By establishing the above page mapping table, the memory scheduler can determine the set of candidate access pages that may be accessed in the next calculation stage based on the current batch of node sets, the current number of graph neural network calculation layers, and the sampling fan-out number when the graph neural network computing engine performs risk prediction calculations.
[0028] In addition, the system includes a large language model. This large language model communicates with a Trusted Execution Environment (TEE) via a secure channel and is used to generate prefetch bias parameters, policy suggestions, or semantic attribution explanations based on anonymized page-level state observation vectors. In this embodiment, the information sent to the large language model is preferably page-level, summarized, or anonymized state information, rather than raw financial plaintext data or complete guarantee relationship diagrams. This reduces the risk of new data leakage introduced when the large language model participates in collaborative reasoning.
[0029] The following is combined with Figure 1 The schematic diagram shown illustrates the system operation flow of this embodiment.
[0030] First, the system executes step S101 to construct a encrypted graph. After obtaining raw data from various data providers, the system cleans, aligns, de-identifies, and encrypts the raw data, and integrates and stores it in a graph database in an untrusted environment, forming an encrypted graph of non-performing asset guarantee relationships. Nodes in the non-performing asset guarantee relationship graph can represent enterprises, asset entities, guarantor entities, or other risk-related objects, or they can represent guarantee relationships, counter-guarantee relationships, related-party transaction relationships, financial relationships, or other relationships that can reflect risk transmission.
[0031] Subsequently, the system executes step S102, transmitting topology metadata and establishing a page mapping table. The graph's topology metadata is extracted from the graph database and transmitted to the memory scheduler within the Trusted Execution Environment (TEE). The topology metadata may include node identifiers, edge identifiers, node degrees, adjacency relationships, guaranteed link levels, graph community partitioning results, graph distances to historically problematic nodes, and statistical information related to risk propagation.
[0032] In step S102, the memory scheduler further establishes or updates a page mapping table based on the storage location of the graph data in memory. The page mapping table associates graph nodes, edge data, node feature data, and adjacency lists with corresponding memory pages. Through this page mapping table, the memory scheduler can determine the memory pages corresponding to a target node, its first-order guaranteed associated nodes, second-order guaranteed associated nodes, and related node features, thereby providing a basis for subsequent candidate page set determination, page prefetching, page swapping, and spoofing crawling.
[0033] Subsequently, the system executes step S103, performing graph neural network computation. The graph neural network computation engine begins risk prediction for target enterprise nodes, asset nodes, or other risky objects. When the memory page that the graph neural network computation engine needs to access is not located in the secure memory of the Trusted Execution Environment (TEE), a page fault is triggered. At this time, the memory scheduler, in conjunction with the encryption / decryption engine, loads the corresponding encrypted data page from the graph database in the untrusted environment or from an external memory region. The encryption / decryption engine then decrypts the data within the TEE for use by the graph neural network computation engine.
[0034] During graph neural network computation, node sampling, neighbor aggregation, and multi-layer message passing may lead to access to a large number of scattered memory pages. For example, in the first-layer message passing phase, the graph neural network computation engine may need to access first-order neighbor pages of the current batch of nodes; in the next layer message passing phase, this may continue to extend to second-order neighbor pages. The memory scheduler determines the candidate page set for access based on the page map, the current batch of nodes, the current graph neural network computation layer number, and the sampling fan-out number, in order to assess the performance benefits and security risks of these pages in advance.
[0035] During the calculation process, the system continuously executes step S104, which involves calculating page metrics and performing dynamic memory scheduling and defense. The memory scheduler monitors and calculates several core metrics to guide decisions regarding page prefetching, retention, swapping out, lazy loading, and spoofed crawling.
[0036] The first core metric is the performance benefit metric. As a preferred implementation, this performance benefit metric includes risk topology-page dwell density, used to quantify the scheduling value of a specific memory page in the current risk prediction task. This metric considers not only the general topological importance of graph nodes but also the risk semantic information in the non-performing asset collateral relationship graph.
[0037] In one exemplary implementation, for memory pages Its risk topology - page dwell density This is defined as the Markov steady-state residence probability of risk semantics in the graph topology. Specifically, it can be iteratively calculated using the random walk formula with biased restart:
[0038] in, Represents memory page The set of graph nodes contained therein Indicates the number of nodes; Represents a node The set of neighboring nodes; For nodes The degree of departure; The edge weight represents the propagation attenuation rate of the guaranteed link; For nodes in the previous iteration round The probability of risky residence; This is a constant for the restart probability (typically between 0.15 and 0.3). For nodes The inherent risk bias prior value is calculated by normalizing the node's own guarantee amount exposure and historical default records.
[0039] This formula simulates the transition matrix operation of a Markov chain. The actual physical process of node features propagating along the guarantee link during computation. The larger the value, the higher the memory page's position. The higher the theoretical probability of being sampled frequently in multi-hop message passing, the more adaptively the topological value of a page can be quantified without relying on manual experience for weighting.
[0040] For example, in risk prediction tasks that focus on preventing large-scale funding chain disruptions, the inherent risk bias prior value can be adjusted. This guides scheduling preferences. For core nodes with extremely high collateral exposure, the system will assign them higher priority. The initial values significantly increase the probability of dwelling on the pages containing these high-risk nodes during the random walk's iterative convergence.
[0041] In standards that focus on computational graph topological connectivity During the task, the restart probability constant can be increased. The value of (e.g., set to 0.3) is used to limit the excessive spread of risk along the long-tail link, making the prefetching behavior more focused on the local first- or second-order neighborhood of the visited node. Through the above physical adjustment of the Markov walk parameters, the system can adaptively quantify the topological value of the page under different business scenarios.
[0042] In this way, the memory scheduler can prioritize retaining or prefetching memory pages that have a greater impact on risk prediction results. For example, if a page contains multiple highly central nodes, nodes located at the core of the guarantee link, nodes with high guarantee amounts, or nodes close to historically problematic nodes, then the risk topology of that page—the page residency density—is high, making it more suitable to retain it in safe memory or prefetch it in advance.
[0043] The second core metric is the side-channel risk metric. As a preferred implementation, this side-channel risk metric includes a side-channel exposure index, used to quantify the risk of inferring graph topology, key guarantor nodes, or risk propagation paths in an untrusted environment by observing page access patterns.
[0044] In one exemplary implementation, the system quantifies the amount of information leakage in access patterns based on relative entropy (KL divergence) from information theory. For memory pages... Within the preset time window Side-channel exposure index The following formula can be used for calculation:
[0045] in, Indicates a memory page The access event sequence state space; This indicates an untrusted environment within a time window. Empirical distribution of actual page access frequency observed within the database; This represents the random white noise access distribution under the security baseline. The Kolb-Leibler divergence measures the difference between the actual access patterns and random white noise. The larger the difference, the easier it is for the topological patterns to be externally reconstructed. The aim is to apply a non-linear multiplicative penalty to the exposure index when a high-risk page experiences a high-frequency page fault, thereby rigorously integrating the security risks of the two orthogonal dimensions of "access frequency" and "pattern regularity" within a single indicator.
[0046] Scaling factor in the formula This is used to measure the system's tolerance limit to high-frequency page fault bursts. In this embodiment, since abnormally frequent page switching is a core prerequisite for triggering side-channel monitoring, when the page fault frequency... When it is at a low level, the exponential term Approximately The exposure index is mainly determined by the distribution difference (KL divergence); while when When the exponential term rises sharply, it rapidly amplifies previously minor topological distribution anomalies. An example of a preferred set of system configuration parameters is: random white noise access distribution at a secure baseline. Set to uniform distribution, scaling factor Set to 0.05 The preset security baseline threshold The method for determining this is as follows: when the Trusted Execution Environment (TEE) is idle or running a non-sensitive graph algorithm (such as a random walk), a preset time window (e.g., ...) is statistically analyzed. The average exposure index within the range is used as the baseline. The value is set to 1.2 to 1.5 times the average baseline value (e.g., 0.5). The defense mechanism is only triggered when the calculated composite exposure index exceeds this dynamic baseline.
[0047] In this way, the side-channel exposure index can not only reflect whether a page is frequently accessed, but also whether the page access sequence exhibits a regularity that may expose guarantor links, graph communities, or risk propagation paths. When the side-channel exposure index of a page or set of pages is high, it indicates that the untrusted environment may be able to infer sensitive graph topology information by observing page faults, page access order, or page access concentration.
[0048] like Figure 2 As shown, after calculating the side-channel exposure index, the system compares the side-channel exposure index with a preset risk threshold. If the side-channel exposure index does not exceed the preset risk threshold, the system maintains normal page scheduling; if the side-channel exposure index exceeds the preset risk threshold, the system triggers a spoofing and capture defense to obfuscate observable memory access patterns in an untrusted environment.
[0049] The third core indicator is the comprehensive scheduling utility value. This is used to guide decisions regarding page prefetching, retention, swapping out, or lazy loading. To avoid dimensional imbalance caused by linear weighting, the system constructs a scheduling utility function based on a non-linear cost-benefit ratio, which can be calculated using the following formula:
[0050] in, Represents memory page Risk topology - page dwell density; This represents the prefetch bias confidence of the output of a large language model. To adjust the hyperparameters for the large model policy gain. The numerator represents the overall graph computation throughput gain brought by the safe memory where the page resides; in the denominator, For the physical I / O and encryption / decryption overhead of memory pages, This is a penalty for side-channel risk. It applies to the page's side-channel exposure index. Exceeding the safety threshold At that time, the denominator grows exponentially, causing the overall scheduling utility value of the page to decrease. Rapid decay forces direct prefetching of high-risk pages to be blocked at the underlying level, achieving a mathematical fusion of performance and security hard constraints.
[0051] The memory scheduler bases its decisions on the overall scheduling utility value. Page scheduling can be implemented. For example, pages with high overall scheduling utility and low side-channel risk can be prefetched or retained; pages with low overall scheduling utility or high page swapping costs can be delayed or swapped out; and pages with side-channel exposure exceeding a preset risk threshold require defense in conjunction with spoofing and capture strategies.
[0052] In step S105, the system and the large language model perform asynchronous collaborative updates and verify the prefetch bias parameters. To avoid the large language model's inference latency blocking the underlying real-time memory scheduling flow, the system deconstructs the page scheduling logic into a two-layer architecture of "real-time local decision-making" and "asynchronous policy optimization." Specifically, the memory scheduler has a lightweight prefetch decision mechanism pre-built in the Trusted Execution Environment (TEE) to handle real-time page faults and page prefetching actions at the nanosecond to microsecond level; while the large language model acts as an asynchronous policy optimizer, providing macro-level scheduling guidance through a secure channel. The memory scheduler periodically organizes the underlying state information into state observation vectors in the background and sends them to the large language model through a secure channel. Unlike directly sending raw financial data or complete guarantee relationship diagrams, the state observation vectors in this embodiment are preferably anonymized page-level observation information.
[0053] Specifically, the state observation vector may include anonymized page identifiers, page dwell status, performance benefit indicators, side-channel risk indicators, memory budget, current graph neural network computation stage, and candidate page set, but does not include the original company name, original guarantee amount in plaintext, complete adjacency list, or complete guarantee relationship graph. In this way, the large language model can analyze topology access trends and optimize scheduling boundaries based on page-level state, while reducing the risk of cross-institutional financial data and complete graph topology leakage.
[0054] After receiving the state observation vector, the large language model generates prefetch bias parameters through semantic reasoning and returns them to the memory scheduler. In this embodiment, the prefetch bias parameters may include the target page identifier, prefetch priority, confidence level, effective time window, and reason code. The reason code can be used to indicate the reason for the prefetch suggestion, such as neighbor expansion, guarantee link extension, graph community relevance, increased industry risk, or the upcoming access in the current computation phase.
[0055] It should be noted that, in this embodiment, the prefetch bias parameters generated by the large language model do not directly control the underlying physical execution; instead, they are first subjected to security checks by the memory scheduler. These security checks may include candidate set verification, institutional domain permission verification, memory budget verification, side-channel risk verification, and effective time window verification.
[0056] Among them, candidate set verification is used to determine whether the target page identifier in the prefetch bias parameter belongs to the current candidate access page set, or whether it belongs to the adjacent page set that can be deduced from the page mapping table and the current graph neural network calculation stage; organization domain permission verification is used to determine whether the organization domain label corresponding to the target page meets the access permissions of the current cross-organization shared task; memory budget verification is used to determine whether adopting the bias parameter to perform prefetching will exceed the available safe memory budget of the Trusted Execution Environment (TEE); side channel risk verification is used to determine whether the target page is in a high-risk state that is not suitable for direct prefetching; and effective time window verification is used to determine whether the bias parameter is still within the effective time range for execution.
[0057] Only after the prefetch bias parameters pass the aforementioned security checks will the memory scheduler incorporate them into the comprehensive scheduling utility value calculation. This is used to dynamically update the boundary conditions of the lightweight prefetch decision mechanism and, based on this, preload relevant memory pages, thus achieving a shift from passive response to proactive planning. If the prefetch bias parameters returned by the large language model fail the security checks, or if the large language model does not return a result within the preset response time, the memory scheduler maintains the parameter configuration of the current lightweight prefetch decision mechanism or switches to a local heuristic prefetch strategy based on the page mapping table and performance benefit metrics. The local heuristic prefetch strategy can determine the prefetch pages based on the page mapping table, risk topology-page dwell density, the current number of computational layers in the graph neural network, the current batch node set, and the sampling fan-out number.
[0058] Furthermore, when the side-channel exposure index of any memory page or set of pages exceeds a preset risk threshold, the system executes a disguised crawling defense action. Unlike simply randomly reading irrelevant noisy pages, the disguised crawling action in this embodiment generates a set of disguised crawling pages based on the graph topology relationship of the non-performing asset guarantee relationship graph data, so that the real accessed pages and the disguised crawling pages together form a simulated graph traversal access sequence.
[0059] Specifically, the memory scheduler can determine the graph node corresponding to the accessed memory page based on the page mapping table, find the first-order or second-order guaranteed association nodes of that graph node, and then select the pages containing these guaranteed association nodes as the spoofed crawling pages. If further obfuscation is required, the memory scheduler can also generate spoofed access paths based on guaranteed links or graph communities, and initiate spoofed crawling of multiple memory pages corresponding to the spoofed access paths according to the access rhythm that matches the message passing phase of the graph neural network computing engine.
[0060] In one basic implementation, when the side-channel risk is low but requires perturbation, the memory scheduler can select several cold pages from a set of low-access-hot pages to perform low-frequency masquerading crawls, increasing background noise with low overhead. As the side-channel risk increases further, the memory scheduler prioritizes pages with first- or second-order guarantee associations in the graph topology to perform neighborhood masquerading crawls, making the access sequence observed in the untrusted environment closer to a normal graph traversal. When the side-channel risk is even higher, the memory scheduler can generate one or more false guarantee propagation paths and initiate continuous or quasi-continuous masquerading crawls on multiple pages corresponding to these false guarantee propagation paths to confuse the real risk propagation paths.
[0061] Through the aforementioned disguised scraping methods, even if an untrusted environment can observe the page swapping sequence, page fault frequency, or page access order, it will be difficult to distinguish which accesses belong to the real graph neural network computing needs and which accesses belong to disguised scraping behavior. This increases the difficulty for attackers to reconstruct the real guarantee relationship graph topology, key guarantee nodes, or risk propagation paths.
[0062] Finally, in step S106, the system completes risk prediction. The graph neural network computing engine outputs the predicted risk probability of the target node or target asset. Optionally, the system can also combine semantic attribution explanations generated by a large language model as the final output. These semantic attribution explanations can be generated based on anonymized graph topology summaries, risk propagation path summaries, page-level state information, or business rule templates to enhance the transparency and credibility of the risk prediction results. In practical applications, the system can output risk probability, risk level, main risk sources, possible risk transmission paths, and related explanatory information to the user.
[0063] Through steps S101 to S106 above, the system of this embodiment can complete the following steps in the graph neural network risk prediction process: ciphertext graph construction, page mapping maintenance, candidate access page set determination, performance benefit index calculation, side-channel risk index calculation, LLM collaborative prefetching, prefetch bias parameter security verification, dynamic memory scheduling, spoofing crawling, and prediction result output. Thus, this embodiment establishes a correspondence between non-performing asset guarantee relationship graph data and memory pages through a page mapping table, identifies high-value pages through risk topology-page residency density, identifies high-risk access patterns through the side-channel exposure index, constrains the prefetch bias parameters generated by the large language model through a security verification mechanism, and obfuscates observable page access patterns in an untrusted environment through spoofing crawling actions based on graph topology relationships. This achieves a balance between computational performance and privacy security under the TEE secure memory conditions of a limited trusted execution environment.
[0064] Example 2 This embodiment is a variant of Embodiment 1, the main difference being the deepening and refinement of the "disguised scraping" defense action performed by the system. The system is internally configured to execute a dynamic, layered defense strategy to more intelligently and efficiently respond to security threats of varying intensities.
[0065] In Example 1, the triggering of the system's spoofing and capture action can be based on whether the side-channel exposure index exceeds a preset risk threshold. As an optional system configuration method, in this example, the memory scheduler implements more complex response logic. The system presets at least two risk thresholds: a lower first risk threshold. and a higher second risk threshold ,in, According to memory pages In the time window Side-channel exposure index Based on the relationship with the aforementioned risk thresholds, the system divides the side-channel risk status of memory pages or page sets into a first risk interval, a second risk interval, and a third risk interval with progressively increasing risk levels, and triggers different levels of spoofing and capture strategies.
[0066] Specifically, when a memory page Side-channel exposure index satisfy:
[0067] If the system determines that the current page access pattern is slightly abnormal, or that it may be under low-intensity attack probing, then the system will trigger the low-intensity camouflage mode under the first risk zone.
[0068] In this low-intensity camouflage mode, the memory scheduler is configured to send camouflage fetch requests to the untrusted environment at a low frequency, such as 1 to 5 times per second. The targets of the fetch can be 1 to 2 low-access-hot pages, which can be selected from a set of cold pages in the current risk prediction calculation task that will not be actually accessed in the short term. These cold pages can be isolated nodes, low-degree nodes, nodes that have not been accessed for a long time, or pages that are not directly related to the current batch of nodes. The system performs this action to introduce a small amount of low-cost and irregular background noise from the perspective of an external observer, interfering with the initial attack detection with low performance overhead.
[0069] To avoid excessive overhead caused by low-intensity spoofing, this embodiment can further set spoofing budget parameters. These parameters may include the maximum number of spoofing crawls allowed per unit time, the upper limit on the number of spoofed pages, and the upper limit on the proportion of secure memory used by spoofing crawls. When the current secure memory usage is high or the graph neural network computing engine is under high load, the memory scheduler can reduce the frequency of low-intensity spoofing without falling below the minimum defense requirements, thus avoiding impacting normal risk prediction calculations.
[0070] And when a certain memory page Side-channel exposure index satisfy:
[0071] At this point, the system determines that the current page access pattern has obvious observable regularity, and the untrusted environment may be analyzing the current graph neural network computation process through page access order, page fault frequency, or page access concentration. At this time, the system enters the neighborhood masquerading mode under the second risk zone.
[0072] In this neighborhood masquerading mode, the system's defensive actions no longer rely solely on random cold pages. The memory scheduler analyzes memory pages that are currently being accessed or attacked. And using the page mapping table maintained in Example 1, the memory page is determined. The graph contains one or more graph nodes. Subsequently, the memory scheduler, based on the set of adjacent pages recorded in the page mapping table, searches for first-order or second-order guarantee association nodes of these graph nodes in the non-performing asset guarantee relationship graph, and selects the memory page containing the first-order or second-order guarantee association node as the masquerading crawl page.
[0073] For example, if the currently accessed page Includes target enterprise nodes And the enterprise node If a first-order guarantee relationship exists with several guarantor nodes, counter-guarantor nodes, or affiliated enterprise nodes, the memory scheduler can select the pages containing these first-order guarantee-related nodes as disguised crawling pages. If the current computation stage has entered the second-layer message passing stage of the graph neural network, or if an attacker may infer the guarantee link based on continuous page visits, the memory scheduler can further select the pages containing second-order guarantee-related nodes as disguised crawling pages.
[0074] In the second risk zone, the frequency of spoofed crawls can be higher than in the first risk zone, for example, 5 to 10 times per second. The selection of spoofed pages can balance page topological similarity and scheduling cost. Page topological similarity can be determined by factors such as the degree of nodes within the page, the range of guaranteed amounts, the depth of guaranteed links, the centrality of eigenvectors, the graph community ID, and the distance to historically problematic nodes. By selecting pages that have a certain similarity to the real accessed pages in terms of graph topological features as spoofed pages, the access patterns of the real accessed pages observed in the untrusted environment can be made closer to those of the spoofed accessed pages, thereby increasing the difficulty of differentiation.
[0075] From the perspective of an external attacker, the series of page faults and page reads generated by the system includes both genuine graph neural network computation requirements and topology-related spoofed crawling requests. Because spoofed pages also have collateral associations or topological similarities with the real accessed pages, attackers cannot simply exclude spoofed pages based on whether they are related to the current graph traversal path. Therefore, this neighborhood spoofing pattern is more suitable for obfuscating topology restoration attacks targeting bad asset collateral relationship graphs than simple random cold page spoofing.
[0076] Furthermore, when a certain memory page Side-channel exposure index of page set satisfy:
[0077] When the system determines that it may be under a high-intensity, persistent, or adaptive side-channel attack, it immediately upgrades to the high-intensity simulated graph traversal camouflage mode under the third risk zone.
[0078] In this high-intensity simulated graph traversal masquerading mode, the system's defensive actions are enhanced in multiple dimensions. First, the frequency of the system's masquerading crawls is further increased, for example, to 10 to 20 times per second, to generate more concentrated masquerading access traffic. Second, and more importantly, the target of the system's masquerading crawls is no longer just a single cold page or a single neighboring page, but rather one or more masquerading access paths are generated based on the guarantee links or graph communities in the bad asset guarantee relationship graph.
[0079] Specifically, the memory scheduler can adjust based on the currently accessed page. For each corresponding graph node, a set of candidate nodes similar to its risk topology characteristics is determined. These risk topology characteristics may include node degree, guarantee amount weight, guarantee link depth, graph community number, neighbor expansion probability, and graph distance to historically problematic nodes. Subsequently, the memory scheduler selects one or more starting nodes from the candidate node set and constructs a pseudo-guarantee propagation path based on the first-order guarantee association, second-order guarantee association, or associated nodes within the same graph community of that starting node.
[0080] For example, the actual access path might correspond to: Target enterprise node → Direct guarantor node → Counter-guarantor node.
[0081] To obfuscate the real path, the memory scheduler can choose another set of nodes with similar topological characteristics to construct a fake access path, for example: False target node → False guarantor node → False counter-guarantor node.
[0082] Subsequently, the memory scheduler initiates spoofing fetches on multiple memory pages corresponding to the spoofed access paths, following an access rhythm that matches the message passing phase of the graph neural network computing engine. In other words, while the graph neural network computing engine aggregates messages in the order of first-order neighbors and second-order neighbors during actual computation, the spoofing fetches can also proceed in a similar first-order page and second-order page access rhythm, making the externally observed page access sequence closer to a normal depth graph traversal operation.
[0083] In another alternative implementation, the memory scheduler can generate community-level spoofed access based on graph communities. Specifically, when real access is concentrated in a certain risky community, the memory scheduler can select another graph community with similar node size, average degree, guaranteed link level, or risk topology-page dwell density, and initiate spoofed crawling of several pages in that graph community. Therefore, even if the untrusted environment can observe that page access is concentrated in a certain type of graph community, it will be difficult to determine the location of the graph community targeted by the actual risk prediction calculation.
[0084] Through the aforementioned path-level or community-level spoofing, the page fault sequences, page access co-occurrence relationships, and access path continuity generated by the system are all subject to topology-related spoofing interference. From the perspective of an external attacker, both the real and spoofed access paths exhibit access patterns similar to message passing in a graph neural network. Therefore, it is more difficult for attackers to reconstruct the real guarantee link, key guarantee nodes, or risk propagation path based on the access sequences. This spoofing configuration method, which "simulates real business graph traversal," makes it difficult for attackers to distinguish between genuine and fake memory access sequences, greatly increasing the difficulty of analyzing and reconstructing the real graph structure, thus achieving a stronger obfuscation and defense effect.
[0085] Furthermore, the first risk threshold in the system of this embodiment Second risk threshold It is not fixed, but configured to be dynamically adjusted. For example, when analyzing external security information, system operation logs, or anonymized page-level state information, if the large language model discovers that a new side-channel attack method targeting trusted execution environments has been disclosed, or detects that the current page access sequence is highly similar to a known attack pattern, it can generate policy recommendations to lower the first risk threshold. Second risk threshold This puts the system into a more sensitive and vigilant defensive state.
[0086] It should be noted that threshold adjustment suggestions generated by large language models can also be adopted after security verification and policy constraints by the memory scheduler. For example, the memory scheduler can determine whether the threshold adjustment suggestion exceeds the preset allowable adjustment range, whether it will cause the spoofing capture frequency to exceed the spoofing budget, and whether it will affect the minimum throughput requirements of the current graph neural network computing engine. Only when the threshold adjustment suggestion meets the above constraints will the memory scheduler execute the corresponding adjustment. This avoids the large language model outputting abnormal policies that could lead to excessive system defense or severe performance degradation.
[0087] Conversely, when the system load is extremely high, the security environment is stable, the side-channel exposure index is consistently below the security baseline, or the current graph neural network computation task has high real-time requirements, the memory scheduler can appropriately increase the first risk threshold. Second risk threshold Alternatively, reduce the frequency of spoofing and capturing in the first and second risk zones to reduce defense overhead and prioritize business performance.
[0088] In a further implementation, the memory scheduler can dynamically select tiered masquerading strategies based on historical defense effectiveness. For example, when the system finds that random cold page masquerading is insufficient to interfere with the current attack pattern, it can switch from the first risk zone to the neighborhood masquerading mode in advance. When the system finds that neighborhood masquerading still cannot effectively improve the uncertainty of page access sequences, it can further introduce path-level masquerading or community-level masquerading. Correspondingly, when the system finds that the additional I / O overhead generated by high-intensity masquerading is too high, it can shorten the length of the masquerading access path, reduce the number of masquerading paths, or only perform high-intensity masquerading on risk topologies—pages with high page dwell density.
[0089] For ease of implementation, the hierarchical camouflage strategy in this embodiment can be executed through the following logic.
[0090] First, the memory scheduler periodically obtains the side-channel exposure index of each memory page or set of pages. And determine its risk range.
[0091] Secondly, the memory scheduler selects the corresponding camouflage strategy based on the risk range. The camouflage strategy includes random camouflage of cold pages, camouflage of neighborhoods, camouflage of paths, or camouflage of communities.
[0092] Next, the memory scheduler determines the set of pages to be spoofed based on the page mapping table, and determines the spoofing frequency by combining the current safe memory budget, spoofing budget, and graph neural network computation stage.
[0093] Finally, the memory scheduler initiates a fake crawling request to the untrusted environment, so that the real accessed page and the fake accessed page form a mixed access sequence at the external observable level.
[0094] By configuring this hierarchical and dynamically adjustable defense strategy, the system in this embodiment can finely allocate defense resources according to the threat level, avoiding the performance waste that may result from a "one-size-fits-all" defense. Compared to simply randomly reading irrelevant noisy pages, this embodiment introduces neighborhood masquerading based on first-order or second-order guaranteed associations, path-level masquerading based on guaranteed links, and community-level masquerading based on graph communities. This makes the masquerading access more similar to the real graph neural network message passing process, thereby improving the system's resistance and resilience against persistent and adaptive side-channel attacks.
[0095] Example 3 This embodiment is also a variation or preferred implementation based on embodiment 1. Its core lies in further deepening the role of the large language model in the whole system. It is not only used for page prefetching and result interpretation, but also integrated into risk attribution analysis and high-level strategy generation, thereby improving the intelligence level of the system and the commercial value of the output results.
[0096] Compared to Example 1, this example does not change the basic deployment architecture of the system. That is, the system is still deployed in a hybrid computing architecture containing an untrusted environment and a Trusted Execution Environment (TEE). The graph database is deployed in the untrusted environment, while the memory scheduler, graph neural network computing engine, and encryption / decryption engine are deployed in the TEE. The large language model communicates with the TEE through a secure channel. The main improvement of this example is that, when the system interacts with the large language model, the construction method of the page-level state observation vector, the privacy protection method, the prefetch bias parameter output format, the security verification mechanism, and the method by which attribution results participate in scheduling control are further defined.
[0097] In this embodiment, the system enhances the construction process of the page-level state observation vector sent to the large language model. It should be noted that the page-level state observation vector can also be called a structured state observation vector or a multimodal cue word, depending on the actual implementation; it does not necessarily mean that the original plaintext financial data or the complete guarantee relationship diagram must be exposed to the large language model. Instead, in cross-institutional data sharing scenarios, preferably, the system sends anonymized page-level state observation vectors to the large language model. These page-level state observation vectors can be generated by the memory scheduler within the Trusted Execution Environment (TEE) based on the page mapping table, the candidate access page set, the page residency status, performance benefit indicators, and side-channel risk indicators.
[0098] In this embodiment, the page-level state observation vector sent to the large language model may include the following categories.
[0099] First, page-level graph topology state information. This information may include anonymized page identifiers, anonymized node identifiers, a set of candidate pages to be accessed, page dwell status, a set of adjacent pages, the current stage of the graph neural network computation, and the neighbor expansion status corresponding to the current batch of nodes. Using this information, the large language model can understand the access relationships between candidate pages and potential subsequent access trends without accessing the complete guarantee relationship graph.
[0100] Second, page-level performance and security status information. This information may include risky topology-page dwell density, side-channel exposure index, security memory budget, page swapping in and out costs, prefetch queue status, and masquerading fetch budget usage. Using this information, the large language model can combine page value, side-channel risk, and memory resource constraints to generate prefetch suggestions or strategy recommendations that better meet system scheduling requirements.
[0101] Third, the textual attribute information of nodes. For key nodes in the graph, such as the enterprise node being analyzed, the system can extract its unstructured textual description information, such as enterprise name, industry classification (e.g., "real estate" or "information technology"), registered address, description of main business scope, and publicly available legal records, and add this textual information to the page-level state observation vector.
[0102] In an implementation more suitable for cross-institutional privacy protection scenarios, the aforementioned node text attribute information is not sent directly to the large language model in its original plaintext form. Instead, it undergoes anonymization, binning, or summarization processing within a Trusted Execution Environment (TEE). For example, company names can be replaced with anonymized node identifiers, industry classifications can be retained as industry tags, registered locations can be mapped to regional tags, guarantee amounts can be mapped to amount ranges, and legal litigation records can be extracted as "litigation risk level" or "litigation event summary." Through this approach, the large language model can understand the semantic context required for risk analysis while reducing the risk of exposing raw financial data, company identity information, or complete business details.
[0103] Fourth, external macro-level textual information. The system can access a real-time information stream, such as financial news, industry research reports, and government policy documents. When the system needs to analyze a certain type of node or page set in conjunction with the external macro environment, it can dynamically obtain the latest macro-level textual summaries related to the industry, region, or asset class to which the node belongs, and concatenate them as contextual information into the page-level state observation vector.
[0104] In this embodiment, the external macro-level text information is preferably a summary of information related to the industry, region, or risk type of the anonymized node, rather than containing complete external documents or sensitive plaintext directly linked to a specific company. For example, the system can process macro-level information into summary descriptions such as "recent tightening of financing policies in the target node's industry," "increased default rates of similar companies within the target region," and "extended disposal cycles for related asset classes." In this way, the large language model can combine changes in the external environment to assist in judging risk transmission trends, page prefetching scope, and defense strategies.
[0105] Fifth, natural language mapping of state instructions. To enable the large language model to better understand the semantics of the underlying system state, the system can discretize continuous numerical indicators and map them to natural language phrases. For example, a side-channel exposure index value between 0 and 0.5 can be mapped to "access pressure: low", between 0.5 and 2.0 to "access pressure: medium", and greater than 2.0 to "access pressure: high". These natural language state instructions make page-level state observation vectors more suitable for logical reasoning by the large language model.
[0106] Furthermore, in this embodiment, the natural language mapping of state instructions can be extended to page scheduling state and security state. For example, a risky topology—a page with high page dwell density—can be mapped to "page value: high," a Trusted Execution Environment (TEE) with low remaining security memory can be mapped to "memory budget: tight," and a candidate page with a high side-channel exposure index can be mapped to "direct prefetch risk: high." Through this approach, the large language model can generate prefetch suggestions and policy suggestions that better conform to system constraints based on page-level state and natural language state labels without directly acquiring underlying sensitive data.
[0107] In a preferred implementation, the page-level state observation vector sent to the large language model may include: task identifier, current graph neural network computation stage, anonymized node identifier, candidate access page set, page dwell state, risk topology-page dwell density, side-channel exposure index, memory budget, institution domain label, state instruction natural language mapping, and external macro-text summary. The page-level state observation vector does not include raw financial plaintext data, complete guarantee relationship graphs, complete adjacency lists, or unauthorized cross-institutional data.
[0108] An enhanced page-level state observation vector can be illustrated as follows: Task Identifier: Target node: Anonymous node Industry: Real Estate; Region: East China; Graph Topology Features: Anonymized high-dimensional vector summary; Candidate Access Page Set: Page dwell time status: Already stationed, No stay [Not resident]; Recent side-channel exposure index sequence: [0.1, 0.2, 1.5, 3.4, 3.2]; Mapping status: 'Access pressure: High'; Memory budget: 'Tight'; Related macro news summary: 'Real estate enterprise development loan approvals are becoming stricter'. In authorized and risk-permissible on-premises deployment scenarios, the system can also employ richer business-related prompts. For example, an enhanced page-level state observation vector could be illustrated as follows: Target node: XX Real Estate Group; Industry: Real Estate; Graph topological features: [a high-dimensional vector]; Candidate visit page set: Recent side-channel exposure index sequence: [0.1, 0.2, 1.5, 3.4, 3.2]; Mapping status: 'Access pressure: High'; Related macro news: 'Stricter approval process for real estate development loans'. It should be noted that this example is only used to illustrate that large language models can combine node business semantics and external macro semantics for analysis. In cross-institutional privacy protection scenarios, the actual sent content can be processed using the aforementioned anonymization, bucketing, or summarization methods to avoid exposing the original plaintext financial data or the complete guarantee relationship graph.
[0109] Once the large language model receives the aforementioned page-level state observation vector, it is configured to perform more complex reasoning tasks than those in Example 1.
[0110] First, it generates structured attribution reports. At this point, the large language model no longer simply outputs a single explanation, but can generate a structured report containing multiple fields according to a preset template. For example, the report could include sections such as "Risk Profile," "Main Risk Sources," "Risk Transmission Path Analysis," and "Supporting Materials." In the "Supporting Materials" section, the large language model can cite news summaries, policy summaries, industry status, page-level status, or graph topology summaries obtained from the page-level state observation vectors as the basis for its judgments.
[0111] In this embodiment, the structured attribution report can be further combined with the risk prediction probability output by the graph neural network computing engine. For example, when the graph neural network computing engine outputs a high risk probability for a certain target node, the large language model can generate explanations such as "the main sources of risk are concentrated guarantee links, increased risk of associated nodes, and tightening of industry financing environment" based on the anonymized graph topology summary and external macro-text summary. This explanation can not only be used to display the final risk prediction results, but also serve as auxiliary input for the memory scheduler to adjust the prefetch range, defense level, or candidate page priority.
[0112] Secondly, it generates higher-level policy recommendations. Besides generating prefetch bias parameters to guide memory scheduling, the large language model can also provide business-level policy recommendations from a more macro-level perspective. For example, in an authorized and risk-tolerant on-premises deployment scenario, based on the aforementioned page-level state observation vectors, the large language model might generate the following recommendations: "The overall risk in the building materials industry has been detected to be rising due to the tightening of real estate credit, and the access pressure on YY Cement Plant, the direct upstream supplier of the target node XX Real Estate Group, has also begun to increase. Recommendations: 1. Raise the risk monitoring level of all upstream supplier nodes in the building materials industry that have a guarantee relationship with XX Real Estate Group to the highest level; 2. Prefetch the memory pages of all these supplier nodes and their primary guarantor nodes into the Trusted Execution Environment (TEE) to prepare for the upcoming in-depth correlation investigation." In cross-institutional privacy protection implementations, the above strategy recommendations can also be translated into anonymized forms. For example, a large language model can output: "Increased page access pressure was detected in the target industry's corresponding risky communities. It is recommended to increase interaction with anonymous nodes." Prioritize the prefetching of supply chain-related pages with first-order guarantee associations, and prioritize adding pages with high risk topology (high page dwell density and low side-channel risk) from the candidate page set to the prefetching queue. It should be noted that the aforementioned business-layer strategy suggestions do not directly alter the system's business decisions or page scheduling status. Instead, they need to be transformed into parameter adjustment suggestions within page prefetching scope, defense level, candidate page priority, or comprehensive scheduling utility value before they can participate in system scheduling control. In this way, the system retains the high-level strategy generation capabilities of the large language model while avoiding the exposure of unnecessary raw financial plaintext data to the large language model and preventing it from bypassing the security controls within the Trusted Execution Environment (TEE).
[0113] In this embodiment, the prefetch bias parameters generated by the large language model are preferably structured prefetch bias parameters. These structured prefetch bias parameters may include a target page identifier, prefetch priority, confidence level, effective time window, reason code, and optional defense level recommendations. Specifically, the target page identifier indicates the page to be prefetched, the prefetch priority indicates the prefetch order, the confidence level indicates the large language model's degree of certainty regarding the prefetch recommendation, the effective time window limits the time range within which the instruction can be executed, and the reason code indicates the reason for the recommendation, such as neighbor expansion, extended guarantor links, graph community relevance, increased industry risk, or the upcoming access during the current computation phase.
[0114] As an example, the structured prefetch bias parameters returned by a large language model may include the following: Target page identifier: Prefetch priority: 0.86; Confidence level: 0.79; Effective time window: 300ms; Reason code: . In another example, the structured prefetch bias parameters returned by the large language model may include the following: Target page identifier: Prefetch priority: 0.74; Confidence level: 0.72; Effective time window: 500ms; Reason code: Defense level recommendation: Neighborhood camouflage. It should be noted that the prefetch bias parameters or policy suggestions generated by the large language model are not directly executed by the system. To avoid issues such as unauthorized prefetching, invalid prefetching, excessive prefetching, or execution of expired instructions, the memory scheduler needs to perform security checks on them before execution. These security checks may include candidate set verification, institutional domain permission verification, memory budget verification, side-channel risk verification, and effective time window verification.
[0115] Among them, candidate set verification is used to determine whether the target page belongs to the current candidate access page set, or whether it belongs to the adjacent page set that can be deduced from the page mapping table and the current graph neural network calculation stage; organization domain permission verification is used to determine whether the organization domain label corresponding to the target page meets the access permissions of the current cross-organization shared task; memory budget verification is used to determine whether the execution of prefetch will exceed the secure memory budget of the Trusted Execution Environment (TEE); side channel risk verification is used to determine whether the target page is suitable for direct prefetching, or whether it needs to switch to masquerading prefetching, delayed prefetching, or defense mode; and valid time window verification is used to determine whether the prefetch bias parameter has expired.
[0116] The memory scheduler only includes the prefetch bias parameters in the overall scheduling utility value calculation if the prefetch bias parameters pass the aforementioned security checks, and determines whether to perform prefetching, holding, swapping out, or lazy loading based on the overall scheduling utility value. If the prefetch bias parameters fail the security checks, or the large language model does not return the prefetch bias parameters within a preset response time, the memory scheduler switches to a local heuristic prefetch strategy. The local heuristic prefetch strategy can determine the prefetch pages based on the page mapping table, risk topology-page dwell density, the current graph neural network computation layer number, the current batch node set, and the sampling fan-out number.
[0117] Furthermore, in this embodiment, the large language model can also participate in the generation of defense strategies, but its generated defense suggestions are also subject to the constraints of the memory scheduler. For example, when the large language model determines, based on the page-level state observation vector, that the current access pressure is high, the page access path has strong continuity, and the memory budget is sufficient, it can suggest increasing the masquerading capture level or expanding the neighborhood masquerading range. After receiving this suggestion, the memory scheduler can combine the side-channel exposure index, masquerading budget, and system throughput requirements to decide whether to execute cold page random masquerading, neighborhood masquerading, path-level masquerading, or community-level masquerading.
[0118] In a specific scenario, when a large language model identifies an increase in industry risk for a target node, and pages in the candidate page set with first- or second-order guaranteed associations to that target node are likely to be accessed consecutively, the large language model can generate two types of outputs: one is a prefetching suggestion, which increases the prefetching priority of relevant guaranteed pages; the other is a defense suggestion, which simultaneously mixes in topologically similar fake pages when prefetching these pages. After verification, the memory scheduler can mix real prefetched pages and fake crawled pages according to a rhythm matching the message passing stage of the graph neural network, making the access sequence observed in the untrusted environment closer to the normal graph traversal process.
[0119] Through the solution in this embodiment, the role of the large language model in the system transforms from a passive "query-response" tool to an intelligent analysis component that actively participates in system decision-making. The final prediction output of the system is no longer a single number, but an intelligent report with in-depth analysis, supporting evidence, and actionable suggestions. This improves the interpretability and operability of the entire system, enabling it to better serve practical risk analysis scenarios.
[0120] Meanwhile, in this embodiment, the active participation of the large language model does not mean that it can bypass the security controls within the Trusted Execution Environment (TEE). On the contrary, the large language model receives state observation vectors that have been anonymized, summarized, or expressed at the page level. Its output prefetch bias parameters, defense suggestions, or policy suggestions all need to undergo security verification and scheduling constraints by the memory scheduler before they can affect the system execution process. Therefore, this embodiment can reduce the risks of cross-organizational data leakage, unauthorized prefetching, and abnormal instruction execution while retaining the semantic understanding and policy generation capabilities of the large language model.
[0121] Furthermore, attribution reports and strategy recommendations can also influence the system's page scheduling strategy. For example, when a structured attribution report indicates a rapid risk spread trend in a certain risk community, the memory scheduler can appropriately increase the risk topology-page dwell density weight of related pages in subsequent scheduling cycles; when strategy recommendations indicate that a certain type of industry node may have a concentrated risk screening need in the near future, the system can increase the candidate priority of first-order or second-order guaranteed associated pages of that type of node; when attribution results show that page access pressure is too high and the side-channel exposure index continues to rise, the system can increase the camouflage capture level or reduce the direct prefetch ratio.
[0122] In this embodiment, the semantic reasoning ability, external macro-information understanding ability, structured attribution ability, and policy generation ability of the large language model are combined with the page-level security scheduling mechanism within the Trusted Execution Environment (TEE). This enables the large language model to not only improve the interpretability of risk prediction results but also to assist in optimizing page prefetching, defense level, and candidate page selection within a controlled range. This further enhances the system's intelligence level, computational efficiency, and privacy protection capabilities in complex cross-organizational risk prediction scenarios.
[0123] Example 4 This embodiment proposes an adaptive optimization mechanism for the system described in Embodiment 1. Its core principle lies in introducing reinforcement learning configuration to automatically adjust the weight coefficients in the comprehensive scheduling utility value calculation formula. This enables the system to learn and evolve based on the actual operating business load and security environment, dynamically finding the current optimal performance-security balance point without frequent manual intervention.
[0124] In Example 1, the overall scheduling utility value can be used to guide decisions regarding the prefetching, retention, swapping out, or lazy loading of memory pages. For example, the overall scheduling utility value can be obtained by weighting risk topology-page residency density, prefetch bias parameter confidence, side-channel exposure index, and page swapping-in / swap-out costs. The optimal values for these weighting coefficients may differ for different service scenarios, different graph sizes, different Trusted Execution Environment (TEE) secure memory capacities, and different side-channel threat intensities.
[0125] For example, when the system is subjected to a strong side-channel attack, it may be necessary to increase the penalty weight of the side-channel risk term to make the memory scheduler prioritize reducing the risk of page access pattern exposure. When the system performs batch offline risk assessment tasks, it may be necessary to increase the weight of the performance gain term and the prefetch confidence term to improve the computational throughput of the graph neural network. When the remaining safe memory space is small or the I / O overhead is high, it may be necessary to increase the weight of the page swapping in and out cost term to reduce unnecessary page swapping and spoofing crawling overhead. Manually adjusting these parameters is not only inefficient, but also difficult to dynamically optimize in real time according to changes in the system state.
[0126] Therefore, this embodiment models the dynamic configuration process of scheduling weights at the system's underlying level as an internal reinforcement learning mechanism. Through this mechanism, the memory scheduler within the Trusted Execution Environment (TEE) can adaptively adjust the weight coefficients in the comprehensive scheduling utility value during continuous system operation, based on page fault conditions, side-channel risk levels, secure memory occupancy, graph neural network computational throughput, prefetch bias parameter security verification results, and policy boundary suggestions output by the large language model.
[0127] It should be noted that the above formula is only one exemplary implementation. In other implementations, the overall scheduling utility value can further incorporate factors such as the amount of safe memory remaining, page dwell time, institutional domain permission level, spoofing budget, effective time window of prefetch bias parameters, or the current computation layer of the graph neural network. Any factor that can be used to measure the retention, prefetching, swapping out, or spoofing value of memory pages in the current scheduling cycle can be included as a component of the overall scheduling utility value.
[0128] To achieve the aforementioned adaptive adjustment, this embodiment configures the memory scheduler within the Trusted Execution Environment (TEE) as a reinforcement learning agent, serving as the primary agent for learning and decision-making. This agent observes the system state in each scheduling cycle, selects appropriate weight adjustment actions, and receives rewards based on the impact of these actions on system performance, security, and cost in subsequent cycles, thereby continuously updating its strategy.
[0129] Specifically, the reinforcement learning mechanism in this embodiment may include the following elements.
[0130] First, the agent. The system configures the memory scheduler within the Trusted Execution Environment (TEE) as an agent for reinforcement learning. This agent is responsible for performing tasks such as page mapping maintenance, candidate page set determination, performance benefit calculation, side-channel risk calculation, prefetch bias parameter security verification, memory page scheduling, and masquerade capture control as described in Example 1. It is also responsible for dynamically adjusting the weight coefficients in the comprehensive scheduling utility value according to the reinforcement learning strategy.
[0131] Second, state. The environmental information upon which the agent bases its decisions constitutes the state. A state can be composed of a set of key performance indicators and safety metrics that comprehensively describe the current operating status of the system. For example, a state can be represented as a vector:
[0132] in, Indicates the first The state of each scheduling cycle This indicates the current secure memory usage of the Trusted Execution Environment (TEE). This represents the average page fault rate within the most recent time window. This represents the average side-channel exposure index of all resident memory pages or candidate access pages. This represents the computational throughput of the graph neural network computing engine. Indicates the current I / O overhead. This indicates the current usage of the fake data scraping budget. This represents the average risk topology-page dwell density of the candidate page set. It indicates the response status of the large language model in the current or most recent period, such as normal, timeout, low confidence, or abnormal output.
[0133] In other implementations, the state can also include information such as the current graph neural network computation stage, the current batch node set size, the candidate page set size, page access sequence entropy, page access path continuity, the number of institutional domain access permission conflicts, and the proportion of prefetch bias parameters that pass security checks. By incorporating these metrics into the state representation, the reinforcement learning agent can simultaneously perceive computational performance, side-channel risks, page scheduling costs, large language model-assisted prefetch reliability, and masquerading crawling overhead.
[0134] Third, actions. The operations that the scheduler agent can perform are called actions. Actions are configured as the core hyperparameter vector for adjusting the cost-effectiveness scheduling formula and the defense system:
[0135] in, The corresponding large language model policy gain coefficient (controls the degree of confidence in the prefetch suggestions of the large model). Corresponding side-channel risk penalty attenuation rate (controlling the blocking strength of high-risk page scheduling). Corresponding dynamic security baseline threshold (controlling the sensitivity of triggering defensive actions).
[0136] To simplify the adaptive learning process of the system, the continuous hyperparameter adjustment space can be discretized. For example, a set of discrete actions can be defined, including: keeping the current parameters unchanged; slightly increasing... To accelerate computation (suitable for stable, safe environments); significantly increase And lower (Suitable for forcibly blocking direct reading of high-risk pages with an extremely high attenuation rate when encountering strong side-channel attack probes); or for use when a large language model experiences continuous timeouts or output anomalies. The decay to zero degrades the system into a conservative scheduling mode that is purely local and topological. Each action corresponds to a set of fine-tuning of the parameter vector A.
[0137] In a further implementation, actions may also include limited adjustments to the masquerade capture strength or prefetch depth. For example, when side-channel risk is high but the secure memory budget is sufficient, actions may include increasing the masquerade capture budget or expanding the neighborhood masquerade range; when secure memory is tight and the page fault rate is high, actions may include reducing the path-level masquerade length or reducing the prefetching of low-value pages. These actions still focus on adjusting the overall scheduling utility value and related scheduling boundaries, without changing the overall system architecture.
[0138] Fourth, rewards. Rewards evaluate the compensation an agent receives for performing a certain action in a given state and are central to guiding the system optimization process. The design of the reward function should reflect the comprehensive objectives of maximizing computational performance, minimizing side-channel risk, and controlling I / O overhead.
[0139] To ensure the policy convergence of the reinforcement learning agent in complex memory environments, the system is configured with a composite reward function based on a logarithmic potential function and a quadratic barrier penalty. This reward function can be expressed as:
[0140] in, Indicates the first The reward value for each scheduling cycle; the first item To calculate the core performance gain, throughput is calculated using a logarithmic function compression plot. Page fault rate The difference in the ratio between them makes the gradient update smoother; the second term is the security constraint barrier function, which is the system's average side-channel exposure exponent. Exceeding the threshold At that time, a severe punishment of quadratic magnitude is imposed. This forces the agent to quickly learn and retreat to a safe zone; the third item Used to suppress the extreme I / O overhead caused by high-frequency spoofing and grabbing; This design employs a hard out-of-bounds penalty term to execute prefetched actions without security checks. It eliminates the cumbersome manual intervention of multiple linear weights, ensuring the system can autonomously explore performance extrema while meeting stringent topological security boundaries.
[0141] During the training and execution of reinforcement learning agents, due to the huge differences in the dimensions and numerical distribution of various indicators (such as the reciprocal of page fault rate, throughput, side channel exposure index, etc.), the above hyperparameters not only play a weighting role, but also play a role in dimension normalization.
[0142] As an example in this application, when throughput The unit is Node / s, total I / O overhead When the unit is MB / s, the initial settings for each hyperparameter in the logarithmic composite reward function can be: quadratic security penalty factor. Log-cost repressor Abnormal out-of-bounds penalty factor Minimal positive number Pick This configuration increases the reward value within a single cycle. Stable convergence at Within the interval, ensure that the gradient updates of the policy network (such as the DQN network) are protected by log smoothing and strongly corrected by the quadratic function when the safety threshold is reached.
[0143] When an action by the scheduler results in a lower page fault rate, higher computational throughput, lower side-channel exposure index, and lower I / O overhead, the system awards a positive reward. Conversely, if an action leads to an increased page fault rate, more easily observable page access patterns, excessive overhead for spoofing, or over-reliance on anomalous outputs from large language models, the system awards a negative reward or penalty.
[0144] Furthermore, this embodiment can also incorporate the prefetch bias parameter security verification results into the reward function. For example, when the prefetch bias parameters generated by the large language model pass candidate set verification, institutional domain permission verification, memory budget verification, side-channel risk verification, and effective time window verification, and actually reduce the subsequent page fault rate, the system can improve its performance. (Large Model Policy Gain Coefficient) Rewards for related actions; when the prefetch bias parameters output by the large language model frequently fail security checks or time out, the system can reduce the reward for accepting the prefetch confidence of the large language model and prompt the scheduler agent to reduce... The value of is degraded to a pure local heuristic prefetching strategy.
[0145] The system's adaptive scheduling process can be configured with the following execution logic.
[0146] The system runs at a fixed scheduling cycle, for example, every 5 minutes. At the beginning of each scheduling cycle, the memory scheduler, acting as a reinforcement learning agent, first observes the current environment state. This state can be composed of runtime monitoring data within the Trusted Execution Environment (TEE), page mapping tables, candidate access page sets, risk topology-page dwell density statistics, side-channel exposure index statistics, large language model response status, and spoofing capture budget usage.
[0147] Then, the memory scheduler selects an action based on its internally deployed policy network. The policy network can be a Q-table, a linear function approximation model, a deep neural network, or other policy model capable of outputting actions based on the state. The action... This corresponds to an adjustment to the hyperparameter vector A, such as increasing the penalty decay rate. Lower the baseline threshold Or adjust the strategy gain coefficient Throughout the subsequent scheduling cycle, the system uses the adjusted hyperparameter vector. The overall scheduling utility value is calculated, and memory pages are prefetched, held, swapped out, lazy-loaded, or faked-fetched based on this value. During this process, the large language model can still generate structured prefetch bias parameters and policy suggestions as described in Examples 1 and 3, but these prefetch bias parameters and policy suggestions must undergo security verification by the memory scheduler before influencing scheduling decisions.
[0148] After the scheduling cycle ends, the system calculates the overall reward based on the page fault rate, graph neural network throughput, side channel exposure index, page swapping cost, masquerading crawling overhead, large language model prefetch bias parameter pass rate, and abnormal scheduling penalty item during that cycle. Subsequently, the system obtains the experience tuple:
[0149] in, This represents the new state after the cycle ends. This empirical tuple is used to update the policy network of the scheduler agent. Understandably, reinforcement learning algorithms such as Q-learning, SARSA, or deep Q-networks, deep deterministic policy gradients, and proximal policy optimization for more complex state spaces can all be used as built-in modules of the system for policy update configuration here.
[0150] Through continuous system operation, exploration, and adaptive adjustment, the scheduler agent gradually learns which scheduling actions to take under different system states. For example, under the conditions of "high safe memory occupancy, high page fault rate, and low side-channel risk," the system may tend to increase the weights of risky topology-page dwell density and prefetch confidence to improve computational performance. Under the conditions of "low page fault rate, high side-channel exposure index, and strong page access path continuity," the system may tend to increase the side-channel exposure index penalty weight and increase the priority of masquerading and crawling strategies. Under the conditions of "high proportion of large language model response timeouts or prefetch bias parameter verification failures," the system may reduce the weight of prefetch confidence in large language models and rely more on page mapping tables and local heuristic prefetch strategies.
[0151] In this embodiment, the large language model can also form a hierarchical collaborative relationship with the reinforcement learning mechanism. Specifically, the system can form an adaptive scheduling structure that combines short-cycle, medium-cycle, and long-cycle scheduling.
[0152] At the short-cycle level, the memory scheduler performs page-level real-time scheduling based on the current page state, the set of candidate accessed pages, the risky topology-page dwell density, the side-channel exposure index, and the security-verified large language model prefetch bias parameters. The goal at this level is to quickly respond to the current graph neural network computation needs, reduce page faults, and control the side-channel risk of the current page access pattern.
[0153] At the mid-cycle level, the reinforcement learning agent dynamically adjusts the weighting coefficients in the overall scheduling utility value based on factors such as page fault rate, side-channel risk, I / O overhead, masquerading and fetching budget usage, and graph neural network throughput over several scheduling cycles. The goal at this level is to learn optimal scheduling strategies under different service loads and attack intensities during continuous operation.
[0154] At the long-term level, large language models can generate policy boundary suggestions based on anonymized page-level state summaries, historical operation log summaries, external security posture summaries, or business load types. For example, large language models can suggest lowering the risk threshold, increasing the maximum camouflage budget, or shortening the effective time window of prefetch bias parameters when the known side-channel attack situation is escalating; and when there are batch offline computing tasks and the security environment is stable, they can suggest relaxing the prefetch depth or increasing the priority of performance gain items. These policy boundary suggestions also need to be subject to security constraints and range verification by the memory scheduler before they can be adopted.
[0155] To prevent the large language model from causing instability in the reinforcement learning process, this embodiment can also set policy boundary constraints. For example, the large language model can only suggest adjustments to the risk threshold, masquerade budget, prefetch depth, or effective time window within a preset range; it cannot directly modify page access permissions, bypass candidate set verification, or directly change cross-institutional data access policies. After the memory scheduler verifies the policy suggestions of the large language model, it can use them as part of the reinforcement learning state or as constraints when selecting actions, rather than executing them unconditionally.
[0156] In one alternative implementation, the reinforcement learning mechanism can also be linked with the hierarchical masquerading and crawling mechanism in Example 2. For example, when the reinforcement learning agent discovers that path-level masquerading can significantly reduce the distinguishability between page access sequences and real risk propagation paths, and the additional I / O overhead is within an acceptable range, it can increase the selection probability of path-level masquerading in the third risk interval; when the system load increases and the masquerading and crawling overhead is too high, the path-level masquerading length or the number of pages masqueraded at the community level can be reduced. Thus, the system can dynamically balance the defense effect and operating cost between different threat intensities and different computational loads.
[0157] In another optional implementation, the reinforcement learning mechanism can also be linked with the structured attribution report in Example 3. For example, when the attribution report generated by the large language model indicates that a certain graph community has a rapid risk diffusion trend, the system can use the average risk topology-page dwell density or candidate priority of the relevant pages in that graph community as part of the reinforcement learning state; if subsequent facts prove that such pages are frequently accessed and that prefetching significantly reduces the page miss rate, the reinforcement learning mechanism can increase the performance gain weight associated with such pages. Conversely, if a certain strategy suggestion fails to improve performance or reduce risk in the long term, the system can reduce its impact on weight adjustment.
[0158] This embodiment can also include a security degradation mechanism. Specifically, when the large language model continuously outputs prefetch bias parameters that fail security checks, or fails to return valid results within the preset response time across multiple scheduling cycles, the memory scheduler can temporarily reduce the security level. The value of will switch the system to an operating mode that primarily relies on the page mapping table, risk topology-page dwell density, and local heuristic prefetching strategies. Once the large language model resumes stable output, and its prefetch bias parameter pass rate and actual hit rate reach the preset requirements again, the system can gradually recover. The value.
[0159] Similarly, when the system detects a persistently high side-channel exposure index, or when external page access observation sequences exhibit strong path continuity and reproducibility, the memory scheduler can trigger a safety fallback constraint outside of the reinforcement learning strategy. For example, it can set... The minimum value of the attenuation rate ensures that the side-channel risk penalty term is not excessively weakened; alternatively, a maximum direct prefetch ratio can be set to prevent high-risk pages from being continuously prefetched, thus exposing the true graph traversal path. Through these methods, the reinforcement learning mechanism can adaptively optimize without breaching the system's safety limits.
[0160] Through continuous system operation, exploration, and policy updates in this embodiment, the memory scheduler can gradually learn which weight adjustment actions should be taken under different system states to obtain the greatest long-term returns. For example, under conditions of high memory utilization and high side-channel risk, the system can automatically increase the penalty weight of the side-channel exposure index and appropriately reduce the direct prefetch ratio; under conditions of low side-channel risk and high page fault rate, the system can automatically increase the weight of risky topology-page dwell density and large language model prefetch confidence to improve the computational performance of graph neural networks; under conditions of high I / O cost and masquerading budget approaching the upper limit, the system can automatically increase the weight of cost items to reduce unnecessary page swapping in and out and masquerading crawling.
[0161] This embodiment endows the system with adaptive and self-optimizing capabilities. It can automatically adjust its core memory scheduling strategy under different business loads, security threat environments, large language model response states, and Trusted Execution Environment (TEE) security memory pressures, maintaining a consistently optimal balance between performance and security without frequent manual intervention. Compared to fixed-weight scheduling methods, this embodiment can dynamically balance the value of risky topology pages, large language model prefetch confidence, side-channel exposure risk, and page scheduling costs based on operational feedback, thereby significantly improving the system's robustness, intelligence, and long-term operational stability.
[0162] The above description is merely a preferred embodiment of this application and is not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
Claims
1. A privacy-preserving non-performing asset risk prediction system for cross-institutional data sharing, characterized in that, The system is deployed in a computing architecture that includes an untrusted environment and a trusted execution environment (TEE), and is used to process encrypted non-performing asset guarantee relationship graph data constructed based on cross-institutional shared data. The system includes: A graph neural network computing engine and memory scheduler deployed within the Trusted Execution Environment (TEE); and a large language model that communicates with the TEE via a secure channel. The memory scheduler is configured as follows: Based on the mapping relationship between the non-performing asset guarantee relationship graph data and memory pages, the set of candidate access pages for the graph neural network computing engine in the risk prediction calculation process is determined. Calculate the performance gain metrics and side-channel risk metrics for each memory page in the candidate access page set; Based on the performance gain indicators and the side channel risk indicators, real-time page scheduling is performed within the Trusted Execution Environment (TEE) through a lightweight prefetch decision mechanism. The system sends a state observation vector, which includes the performance gain metric, the side channel risk metric, and the candidate access page set, to the large language model at an asynchronous period, after page-level anonymization processing, and receives prefetch bias parameters generated by the large language model. The prefetch bias parameters are subjected to security verification, and the boundary conditions of the lightweight prefetch decision mechanism are dynamically updated based on the verified prefetch bias parameters in order to perform scheduling on memory pages. When the side-channel risk index exceeds a preset risk threshold, based on the guarantee association node, guarantee link or graph community of the graph node corresponding to the accessed memory page, a set of fake crawled pages that matches the message passing rules of the graph neural network computing engine is generated, and fake crawling action is performed to confuse the memory access patterns observable in the untrusted environment through the simulated graph traversal sequence, thereby protecting the topological privacy of the bad asset guarantee relationship graph data.
2. The system according to claim 1, characterized in that, The system also includes a page mapping table, which is used to record the correspondence between graph nodes, edge data, node feature data and memory pages in the non-performing asset guarantee relationship graph data.
3. The system according to claim 2, characterized in that, The page mapping table includes node identifier, page identifier, organization domain label, node feature storage location, adjacency list storage location, and adjacent page set; the memory scheduler determines the candidate access page set for the next calculation stage from the adjacent page set based on the current graph neural network calculation layer number, the current batch node set, and the sampling fan-out number.
4. The system according to claim 1, characterized in that, The performance benefit metric includes risk topology-page dwell density, which is calculated based at least on the node degree of the graph nodes contained in the memory page and the neighbor expansion probability in the current graph neural network computation stage, and combined with at least one of the following: feature vector centrality, guarantee amount weight, guarantee link depth, and graph distance with historical bad nodes.
5. The system according to claim 1, characterized in that, The side-channel risk index includes the side-channel exposure index, which is calculated based on at least one of the following: page fault frequency of memory pages within a preset time window, page fault burstiness, page access interval variance, and correlation between page access sequence and graph topology path.
6. The system according to claim 1, characterized in that, The memory scheduler is configured to calculate a comprehensive scheduling utility value and perform scheduling on memory pages based on the comprehensive scheduling utility value. The comprehensive scheduling utility value is calculated based on the performance benefit index, the confidence level of the prefetch bias parameter, the side channel risk index, and the cost-effectiveness ratio of page swapping in and out.
7. The system according to claim 1, characterized in that, The state observation vector is a de-identified page-level state observation vector, which includes anonymized page identifier, page dwell state, performance benefit index, side-channel risk index, memory budget, current graph neural network calculation stage, and candidate access page set, but does not include the original financial plaintext data and the complete guarantee relationship graph.
8. The system according to claim 1, characterized in that, The prefetch bias parameters include target page identifier, prefetch priority, confidence level, effective time window, and reason code; the security verification includes candidate set verification, organization domain permission verification, memory budget verification, and effective time window verification.
9. The system according to claim 8, characterized in that, When the prefetch bias parameter fails the security check, or the large language model does not return the prefetch bias parameter within the asynchronous update cycle, the memory scheduler maintains the boundary conditions of the current lightweight prefetch decision mechanism unchanged and continues to perform page scheduling based on the mapping relationship and the performance benefit metric.
10. The system according to claim 1, characterized in that, The spoofing and crawling action includes tiered spoofing and crawling: when the side-channel risk indicator is in the first risk range, the memory scheduler selects cold pages from the set of low-access-hot pages to perform spoofing and crawling; when the side-channel risk indicator is in the second risk range, the memory scheduler selects the page containing the first-order or second-order guarantee association node based on the first-order or second-order guarantee association node of the graph node corresponding to the accessed memory page to perform spoofing and crawling; when the side-channel risk indicator is in the third risk range, the memory scheduler generates a spoofing access path based on the guarantee link or graph community in the non-performing asset guarantee relationship graph data, and initiates spoofing and crawling on multiple memory pages corresponding to the spoofing access path according to the access rhythm matching the message passing stage of the graph neural network computing engine.