A data security intelligent protection method based on deep learning

CN122698342APending Publication Date: 2026-09-04CHONGQING HEHUOREN TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610960655.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-30
Publication Date
2026-09-04

AI Technical Summary

Technical Problem

[0003]然而,现有技术大多侧重于对单次异常数据暴露行为或静态风险指标进行检测,缺乏对数据暴露状态连续演化过程的分析,难以识别数据暴露过程偏离正常业务路径后的风险演化趋势;同时,对历史数据暴露事件之间的关联关系以及数据分类分级约束利用不足,难以实现未来数据暴露风险的动态推理与持续预警,导致企业数据安全防护的准确性和主动性仍有待提高

Benefits of technology

本发明通过采集数据访问日志、数据下载日志、数据共享日志、数据外发日志、权限变更日志、审批信息以及数据分类分级信息,构建数据暴露状态链,并结合历史正常业务路径识别数据暴露偏离结构,在此基础上构建改进的数据暴露风险演化神经霍克斯过程模型,实现对数据暴露风险的动态演化推理,相比于现有技术仅针对单次异常行为进行检测的方式,本发明能够从数据暴露状态连续演化的角度识别数据暴露过程中的偏离位置、偏离方向及偏离跨度,充分利用历史数据暴露事件之间的演化关系以及数据分类分级约束信息,对未来数据暴露事件的发展趋势进行预测,从而提高数据暴露风险识别的连续性和完整性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122698342A_ABST
    Figure CN122698342A_ABST
Patent Text Reader

Abstract

The application discloses a kind of data security intelligent protection methods based on deep learning, including the following steps: data log is collected, and data security behavior dataset is generated;Exposure statistics processing is executed, and data exposure state unit set is generated;State connection relationship is established, state evolution sequence is formed, and data exposure state chain set is generated;Deviation node, deviation position, deviation direction and deviation span are identified, and data exposure deviation structure is generated;Data exposure risk evolution NHP model is constructed;Risk evolution analysis processing is executed, and risk evolution result is generated;Data security protection processing is executed, and data security protection result is generated.The application utilizes data exposure state chain, data exposure deviation structure and data exposure risk evolution NHP model, realizes data exposure risk dynamic analysis and intelligent security protection, with the advantages of early warning accuracy, proactive protection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent data security protection, and in particular to a data security intelligent protection method based on deep learning. Background Technology

[0002] As enterprises continue to advance their digital transformation, data sharing and business collaboration among business platforms such as financial systems, auditing systems, legal systems, and collaborative office systems are becoming increasingly frequent. A large amount of business data is constantly flowing during access, download, sharing, external transmission, and permission transfer. To ensure enterprise data security, existing technologies typically monitor and analyze data exposure behavior by collecting data access logs, data download logs, data sharing logs, data external transmission logs, and permission change logs, and combine this with a data classification and grading mechanism to achieve data security risk warnings.

[0003] However, most existing technologies focus on detecting single abnormal data exposure behaviors or static risk indicators, lacking analysis of the continuous evolution of data exposure status, making it difficult to identify the risk evolution trend after the data exposure process deviates from the normal business path; at the same time, the correlation between historical data exposure events and the data classification and grading constraints are not fully utilized, making it difficult to achieve dynamic reasoning and continuous early warning of future data exposure risks, resulting in the need to improve the accuracy and initiative of enterprise data security protection. Summary of the Invention

[0004] One objective of this invention is to propose a data security intelligent protection method based on deep learning. This invention utilizes the data exposure state chain, data exposure deviation structure, and data exposure risk evolution NHP model to achieve dynamic analysis of data exposure risks and intelligent security protection, which has the advantages of accurate early warning and proactive protection.

[0005] A data security intelligent protection method based on deep learning according to an embodiment of the present invention includes the following steps: Collect data access logs, data download logs, data sharing logs, data outbound logs, permission change logs, approval and data classification and grading information, and preprocess them to generate a data security behavior dataset; Perform exposure statistics processing on each data object in the data security behavior dataset to generate a set of data exposure status units; For multiple data exposure state units corresponding to the same data object within a continuous time window, establish state connection relationships, form a state evolution sequence, and generate a set of data exposure state chains; By comparing and analyzing the data exposure state chain with the historical normal business path, the deviation nodes, deviation positions, deviation directions and deviation spans are identified, and a data exposure deviation structure is generated. A data exposure risk evolution NHP model is constructed, which includes a data exposure deviation coding layer, a risk induction intensity calculation layer, a data sensitivity constraint layer, and a risk evolution inference layer. The data exposure state chain set is used as the event evolution input, and the data exposure deviation structure is used as the deviation constraint input. The risk evolution analysis is performed using the data exposure risk evolution NHP model to generate risk evolution results. Based on the risk evolution results, the corresponding risk level is determined, and the protection strategy corresponding to the risk level is invoked to perform data security protection processing and generate data security protection results.

[0006] Optionally, the data access log includes a data object identifier, accessor identifier, department, and access time; the data download log includes a data object identifier, downloader identifier, and download time; the data sharing log includes a data object identifier, sharing initiator identifier, sharing recipient identifier, and sharing time; the data outbound log includes a data object identifier, outbound personnel identifier, outbound target identifier, and outbound time; the permission change log includes a data object identifier, changing personnel identifier, change time, and change result; the approval information includes a data object identifier, approval node name, approval processing personnel identifier, approval processing time, and approval result; the data classification and grading information includes a data object identifier, data category, and data level; the preprocessing specifically includes time alignment, missing data completion, abnormal record removal, and format standardization.

[0007] Optionally, the generation of the data exposure state unit set specifically includes: Read the data access logs, data download logs, data sharing logs, data outgoing logs, and permission change logs from the data security behavior dataset, extract the data object identifiers corresponding to each behavior record, and perform merging processing on each behavior record according to the data object identifiers to generate a data object behavior record set; For each data object behavior record in the data object behavior record set, count the number of corresponding visitor identifiers and the number of departments to which they belong, and generate contact statistics results. Count the number of download records, sharing records, outgoing records, and permission change records for each data object, and generate data exposure behavior statistics results. The contact statistics and data exposure behavior statistics are correlated according to the data object identifier to generate a data exposure status unit corresponding to each data object; Perform summary processing on multiple data exposure status units to generate a set of data exposure status units.

[0008] Optionally, the generation of the data exposure state chain set specifically includes: Read multiple data exposure state units from the data exposure state unit set, extract the data object identifier corresponding to each data exposure state unit, and perform classification and merging processing on multiple data exposure state units according to the data object identifier to generate a data object state unit set; For each data object in the data object state unit set, multiple data exposure state units corresponding to each data object are sorted according to the time sequence of the corresponding behavior records to generate a state time series. Traverse adjacent data exposure state units in the state time series, determine the connection relationship between the preceding and subsequent data exposure state units as state connection relationships, and perform aggregation processing on multiple state connection relationships to generate a set of state connection relationships; Based on multiple data-exposed state units and state connection relationships in the state time series, a continuous state connection is established to form a state evolution sequence corresponding to each data object; The state evolution sequence corresponding to each data object is determined as the corresponding data exposure state chain, and multiple data exposure state chains are aggregated to generate a set of data exposure state chains.

[0009] Optionally, the generation of the data exposure deviation structure specifically includes: Read the data object identifier, approval node name, approval personnel identifier, approval processing time and approval result from the approval information, filter the approval information with the approval result as passed, perform merging processing on the approval information according to the data object identifier, perform sequential processing on multiple approval node names according to the approval processing time, and generate the historical normal business path corresponding to each data object. Obtain multiple data exposure state chains from the data exposure state chain set, extract the data object identifier corresponding to each data exposure state chain, and match the data exposure state chain corresponding to the same data object with the historical normal business path; Traverse the data exposure state chain and the corresponding historical normal business path, identify data exposure state units in the data exposure state chain that are not in the business flow process corresponding to the historical normal business path, and generate deviation nodes. The deviation position is determined according to the order of the deviation nodes in the corresponding data exposure state chain, and the preceding data exposure state unit corresponding to the deviation node is extracted. The process of the preceding data exposure state unit changing to the data exposure state of the deviation node is determined as the deviation direction. Perform a count of consecutively occurring deviation nodes in the same data exposure state chain to generate the deviation span; The deviation nodes, deviation positions, deviation directions, and deviation spans are associated with corresponding data object identifiers to generate a data exposure deviation structure.

[0010] Optionally, the construction of the neural Hawkes process model for the evolution of data exposure risk specifically includes: The NHP infrastructure is invoked, which includes an event input layer, a historical event memory layer, an intensity function calculation layer, and an event prediction output layer, and multiple data exposure state units in the data exposure state chain set are set as event input content. Add a data exposure deviation coding layer before the event input layer. Write the deviation node, deviation position, deviation direction and deviation span in the data exposure deviation structure into the data exposure deviation coding layer to form the deviation constraint input path. The memory processing path of the historical event memory layer for the exposed state unit of the preceding data is preserved, and the exposed state unit of the subsequent data is used as the input intensity function calculation layer for the content of the event to be reasoned. The intensity function calculation layer is transformed into a risk-induced intensity calculation layer, which reads the data exposure state change process between adjacent data exposure state units and performs calculation processing on the risk-induced intensity between data access, data download, data sharing, data outgoing, and permission changes. After the risk-induced intensity calculation layer, a data sensitivity constraint layer is added to read the data object identifier, data category, and data level from the data classification and grading information, and perform data level constraint processing on the risk-induced intensity. The event prediction output layer is transformed into a risk evolution reasoning layer, and the deviation constraint input path, the output results of the historical event memory layer, the risk induction intensity, and the data level constraint processing results are input into the risk evolution reasoning layer. The data exposure risk evolution NHP model is generated by establishing interlayer connections in the order of data exposure deviation coding layer, event input layer, historical event memory layer, risk induction intensity calculation layer, data sensitivity constraint layer, and risk evolution inference layer.

[0011] Optionally, the generation of the risk evolution result specifically includes: Read multiple data exposure state chains from the data exposure state chain set, extract multiple data exposure state units from each data exposure state chain, and input the multiple data exposure state units into the event input layer according to the data object identifier; Read the deviation nodes, deviation positions, deviation directions, and deviation spans in the data exposure deviation structure, and input the deviation nodes, deviation positions, deviation directions, and deviation spans into the data exposure deviation coding layer to generate deviation constraint input results; The historical event memory layer performs historical memory processing on the preceding data exposure state unit in each data exposure state chain, and inputs the data exposure state change process between the subsequent data exposure state unit and the preceding data exposure state unit into the risk induction intensity calculation layer. At the same time, the risk induction intensity result is input into the data sensitivity constraint layer, and the risk constraint result is generated based on the data category and data level in the data classification and grading information. The deviation constraint input results, the output results of the historical event memory layer, and the risk constraint results are input into the risk evolution reasoning layer to generate a candidate set of future data exposure events corresponding to each data object. Based on the candidate set of future data exposure events, determine the occurrence trend of future data exposure events for data objects and the corresponding degree of risk evolution, and generate risk evolution results corresponding to each data object.

[0012] Optionally, the generation of the data security protection result specifically includes: Based on the risk evolution results corresponding to each data object, the trend of future data exposure events and the degree of risk evolution are extracted to determine the risk level corresponding to each data object; The corresponding protection strategy is invoked based on the risk level of each data object, and the execution content of the corresponding protection strategy is determined in combination with the future trend of data exposure events. In accordance with the corresponding protection policy, data security protection measures are implemented for each data object, including data access restrictions, data download restrictions, data sharing restrictions, data outbound restrictions, or permission change restrictions. The risk level, corresponding protection strategy, and execution results of data security protection processing are associated with the data object identifier to generate intelligent data security protection results corresponding to each data object.

[0013] The beneficial effects of this invention are: This invention constructs a data exposure state chain by collecting data access logs, data download logs, data sharing logs, data outbound logs, permission change logs, approval information, and data classification and grading information. It then identifies data exposure deviation structures by combining these with historical normal business paths. Based on this, an improved neural Hawkes process model for data exposure risk evolution is built, enabling dynamic evolutionary reasoning of data exposure risks. Compared to existing technologies that only detect single abnormal behaviors, this invention can identify the deviation position, direction, and span in the data exposure process from the perspective of continuous evolution of the data exposure state. It fully utilizes the evolutionary relationships between historical data exposure events and data classification and grading constraints to predict the development trend of future data exposure events, thereby improving the continuity and completeness of data exposure risk identification.

[0014] Furthermore, this invention combines the data exposure state chain, data exposure deviation structure, and improved neural Hawkes process model to uniformly model various data exposure behaviors in enterprise business processes, such as data access, data download, data sharing, data outreach, and permission changes. It also generates a candidate set of future data exposure events by incorporating historical event memory and risk constraint mechanisms, further determining the degree of risk evolution and outputting the risk evolution results. This achieves a shift from post-event detection to pre-event prediction of data security risks, enabling more accurate identification of potential risks in the continuous flow of enterprise business data, improving the accuracy and timeliness of data security risk warnings, and providing reliable technical support for enterprise data security supervision and proactive protection. Attached Figure Description

[0015] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings: Figure 1 This is a flowchart of a data security intelligent protection method based on deep learning proposed in this invention; Figure 2 This is a schematic diagram illustrating the generation of data exposure deviation structure in a deep learning-based intelligent data security protection method proposed in this invention. Figure 3 This is a schematic diagram of the data exposure risk evolution (NHP) model of a deep learning-based intelligent data security protection method proposed in this invention. Detailed Implementation

[0016] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.

[0017] refer to Figures 1-3 A data security intelligent protection method based on deep learning includes the following steps: Collect data access logs, data download logs, data sharing logs, data outbound logs, permission change logs, approval and data classification and grading information, and preprocess them to generate a data security behavior dataset; Perform exposure statistics processing on each data object in the data security behavior dataset to generate a set of data exposure status units; For multiple data exposure state units corresponding to the same data object within a continuous time window, establish state connection relationships, form a state evolution sequence, and generate a set of data exposure state chains; By comparing and analyzing the data exposure state chain with the historical normal business path, the deviation nodes, deviation positions, deviation directions and deviation spans are identified, and a data exposure deviation structure is generated. A data exposure risk evolution NHP model is constructed, which includes a data exposure deviation coding layer, a risk induction intensity calculation layer, a data sensitivity constraint layer, and a risk evolution inference layer. The data exposure state chain set is used as the event evolution input, and the data exposure deviation structure is used as the deviation constraint input. The risk evolution analysis is performed using the data exposure risk evolution NHP model to generate risk evolution results. Based on the risk evolution results, the corresponding risk level is determined, and the protection strategy corresponding to the risk level is invoked to perform data security protection processing and generate data security protection results.

[0018] In this embodiment, the data access log includes the data object identifier, the accessor identifier, the department, and the access time; the data download log includes the data object identifier, the downloader identifier, and the download time; the data sharing log includes the data object identifier, the sharing initiator identifier, the sharing recipient identifier, and the sharing time; the data outbound log includes the data object identifier, the outbound personnel identifier, the outbound target identifier, and the outbound time; the permission change log includes the data object identifier, the changing personnel identifier, the change time, and the change result; the approval information includes the data object identifier, the approval node name, the approval processing personnel identifier, the approval processing time, and the approval result; the data classification and grading information includes the data object identifier, the data category, and the data level; the preprocessing specifically includes time alignment, missing data completion, abnormal record removal, and format standardization. Data objects refer to electronic data entities generated, stored, transmitted, shared, or used in the course of an enterprise's business activities. These include contract documents, financial vouchers, audit reports, business documents, and other business data files. They are used to carry business information and serve as the objects of data access, data download, data sharing, data outreach, permission changes, and approval processes.

[0019] In this embodiment, the generation of the data exposure status unit set specifically includes: Read the data access logs, data download logs, data sharing logs, data outgoing logs, and permission change logs from the data security behavior dataset, extract the data object identifiers corresponding to each behavior record, and perform merging processing on each behavior record according to the data object identifiers to generate a data object behavior record set; For each data object behavior record in the data object behavior record set, count the number of corresponding visitor identifiers and the number of departments to which they belong, and generate contact statistics results. Count the number of download records, sharing records, outgoing records, and permission change records for each data object, and generate data exposure behavior statistics results. The contact statistics and data exposure behavior statistics are correlated according to the data object identifier to generate a data exposure status unit corresponding to each data object; The contact statistics and data exposure behavior statistics reflect the contact range and spread of the data object, respectively. By performing corresponding associations according to the data object identifier, multiple statistical results of the same data object can be collectively reflected to reflect the current data exposure of the data object. A data exposure status unit is a centralized description of the data exposure status of the same data object after a statistical processing of exposure conditions. Data access logs, data download logs, data sharing logs, data outbound logs, and permission change logs generated during enterprise business activities are usually independent behavioral records, which are difficult to directly reflect the current data exposure level of the data object. By associating the statistical results of contact conditions corresponding to the same data object with the statistical results of data exposure behavior, a data exposure status unit is formed. This allows the scattered data security behavior records to be summarized into a unified status description that can reflect the current data exposure status of the data object. The same data object can form multiple data exposure status units in different exposure condition statistical processing, and provides a basic status for the subsequent establishment of a data exposure status chain. Perform summary processing on multiple data exposure status units to generate a set of data exposure status units.

[0020] In this embodiment, the generation of the data exposure state chain set specifically includes: Read multiple data exposure state units from the data exposure state unit set, extract the data object identifier corresponding to each data exposure state unit, and perform classification and merging processing on multiple data exposure state units according to the data object identifier to generate a data object state unit set; For each data object in the data object state unit set, multiple data exposure state units corresponding to each data object are sorted according to the time sequence of the corresponding behavior records to generate a state time series. Traverse adjacent data exposure state units in the state time series, determine the connection relationship between the preceding and subsequent data exposure state units as state connection relationships, and perform aggregation processing on multiple state connection relationships to generate a set of state connection relationships; The generation of the set of state connection relationships specifically includes: Traverse adjacent data exposure state units in the state time series, determine the data exposure state unit that appears earlier in time as the preceding data exposure state unit, and determine the data exposure state unit that appears later in time as the following data exposure state unit, and establish a state connection relationship between the preceding data exposure state unit and the following data exposure state unit; then perform aggregation processing on multiple state connection relationships corresponding to the same data object to form a set of state connection relationships. State connection relationships are used to characterize the evolution of two data exposure state units of the same data object at adjacent times. Since multiple data exposure state units can only reflect the data exposure situation at each time and cannot describe the data exposure change process between adjacent times, it is necessary to establish state connection relationships to characterize the continuous evolution process of the data exposure state of the same data object. Based on multiple data-exposed state units and state connection relationships in the state time series, a continuous state connection is established to form a state evolution sequence corresponding to each data object; The generation of the state evolution sequence specifically includes: Starting with the first data-exposed state unit in the state time series, and following multiple state connection relationships in the state connection relationship set, subsequent data-exposed state units are sequentially connected until the last data-exposed state unit in the state time series, thus establishing continuous state connections and forming a state evolution sequence. Even after multiple data exposure state units in a state time series are arranged in chronological order, they still represent discrete state descriptions and do not yet form a continuous data exposure process. The state evolution sequence records the continuous change process of the data exposure state of the same data object. Compared to a single data exposure state unit, which only reflects the data exposure situation at a certain moment, the continuous change process can reflect the expansion trend of the data exposure range and the development process of data exposure behavior. Based on the continuous change process, subsequent steps can identify the deviation between the data exposure process and the historical normal business path, and further analyze the inducing effect of previous data exposure behaviors on subsequent data exposure behaviors, providing an event evolution basis for risk evolution analysis. The state evolution sequence corresponding to each data object is determined as the corresponding data exposure state chain, and multiple data exposure state chains are aggregated to generate a set of data exposure state chains. The state evolution sequence formed by state connections already possesses a complete time sequence and a continuous change process of data exposure state, which can directly describe the data exposure process of the same data object. Therefore, the state evolution sequence corresponding to each data object is determined as the corresponding data exposure state chain. In the process of enterprise business activities, there are usually multiple data objects at the same time, and the data exposure process corresponding to each data object is independent of each other. After performing summary processing on multiple data exposure state chains, a set of data exposure state chains is formed. Subsequently, data exposure deviation analysis and risk evolution analysis are performed one by one according to the data object.

[0021] In this embodiment, the generation of the data exposure deviation structure specifically includes: Read the data object identifier, approval node name, approval personnel identifier, approval processing time and approval result from the approval information, filter the approval information with the approval result as passed, perform merging processing on the approval information according to the data object identifier, perform sequential processing on multiple approval node names according to the approval processing time, and generate the historical normal business path corresponding to each data object. Approval information with a passing result reflects the normal business flow process that has actually occurred and been recognized by the enterprise. The historical normal business path is formed by arranging the names of multiple approval nodes in the order of execution according to the approval processing time. Obtain multiple data exposure state chains from the data exposure state chain set, extract the data object identifier corresponding to each data exposure state chain, and match the data exposure state chain corresponding to the same data object with the historical normal business path; After matching the data exposure state chain corresponding to the same data object with the historical normal business path, the actual data exposure process of the data object is made to correspond with the normal business flow process, providing a basis for comparison for subsequent identification of deviations. Traverse the data exposure state chain and the corresponding historical normal business path, identify data exposure state units in the data exposure state chain that are not in the business flow process corresponding to the historical normal business path, and generate deviation nodes. The data exposure state chain can describe the data exposure process of a data object, but it is difficult to directly determine whether the data exposure behavior has exceeded the normal business scope based solely on the data exposure process. By traversing the data exposure state chain and the corresponding historical normal business path, the data exposure state unit that is not in the corresponding business flow process of the historical normal business path is identified as a deviation node. This distinguishes the data exposure behavior that has already occurred from the normal business flow process and determines the position where the risk begins to accumulate from the data exposure process, providing a basis for subsequently determining the deviation position, deviation direction, and deviation span. Deviation nodes can provide a clear starting point for subsequent risk evolution analysis, enabling the Data Exposure Risk Evolution NHP model to continue to infer the future trend of data exposure events based on the data exposure state that has deviated from the normal business path, thereby triggering corresponding data security protection actions before the data exposure expands further. The deviation position is determined according to the order of the deviation nodes in the corresponding data exposure state chain, and the preceding data exposure state unit corresponding to the deviation node is extracted. The process of the preceding data exposure state unit changing to the data exposure state of the deviation node is determined as the deviation direction. The deviation position reflects the order of the deviation node in the corresponding data exposure state chain. When the same data exposure behavior occurs in different orders, the subsequent data exposure process will differ, so it is necessary to determine the deviation position. The deviation direction reflects the change in data exposure state from the preceding data exposure state unit to the data exposure state unit corresponding to the deviation node. It is difficult to reflect the changes in data exposure state before the deviation occurs based solely on the deviation node. By determining the deviation direction, the changes in data exposure state before the deviation can be preserved. By simultaneously determining the deviation position and deviation direction, not only can the order in which the deviation node appears in the data exposure process be determined, but the changes in data exposure state before the deviation can also be described. This allows subsequent risk evolution analysis to combine the data exposure process that has already occurred to determine the development trend of future data exposure events, and further determine the data access, data download, data sharing, data outsourcing, or permission change behaviors that need to be restricted. Perform a count of consecutively occurring deviation nodes in the same data exposure state chain to generate the deviation span; Deviation span reflects the number of consecutive deviation nodes in the same data exposure state chain. A single deviation node only indicates that the data exposure process has deviated at a certain point, while multiple consecutive deviation nodes indicate that the data exposure process has continuously deviated from the business flow process corresponding to the historical normal business path. By statistically analyzing the number of consecutive deviation nodes, we can describe the situation where the data exposure process continuously deviates from the normal business flow process, and provide a basis for subsequent judgment on whether the data exposure process is still in a state of continuous deviation. In combination with the continuous deviation situation reflected by the deviation span, we can adjust the execution scope of data access, data download, data sharing, data outbound transmission, and permission change behaviors in the data exposure process. The deviation nodes, deviation positions, deviation directions, and deviation spans are associated with corresponding data object identifiers to generate a data exposure deviation structure; The data exposure deviation structure is a unified summary of the deviation nodes, deviation locations, deviation directions, and deviation spans corresponding to the same data object. By associating them according to the data object identifier, the results of different dimensions generated during the deviation analysis are aggregated under the same data object, thereby forming a complete data exposure deviation description result. This result is used to uniformly organize the deviation information identified in the data exposure state chain, serving as the input basis for the subsequent data exposure risk evolution NHP model.

[0022] In this embodiment, constructing the neural Hawkes process model of data exposure risk evolution specifically includes: The NHP infrastructure is invoked. The NHP infrastructure includes an event input layer, a historical event memory layer, an intensity function calculation layer, and an event prediction output layer. Multiple data exposure state units in the data exposure state chain set are set as event input content. In the process of enterprise data security management, data exposure risks are usually formed gradually by multiple data exposure behaviors such as data access, data download, data sharing, data outreach, and permission changes, and continue to evolve along the data exposure process. It is difficult to accurately analyze the future development trend of data exposure risks based solely on the current data exposure situation. The NHP model can use the temporal dependencies between historical events to describe the continuous evolution process of events, which is consistent with the business characteristics of continuous evolution of data exposure risks. Therefore, the NHP model is selected as the risk evolution analysis model. The original NHP model primarily relies on the temporal relationships between events to perform event evolution analysis. It does not incorporate historical normal business paths, data exposure deviations, and data classification and grading information, making it difficult to directly meet the needs of enterprises for data security risk evolution analysis. To address these issues, while retaining the historical event memory mechanism, a data exposure deviation encoding layer and a data sensitivity constraint layer are added. The intensity function calculation layer is transformed into a risk-induced intensity calculation layer, and the event prediction output layer is transformed into a risk evolution inference layer. This enables the model to comprehensively consider data exposure deviations, data exposure processes, and data classification and grading information to perform risk evolution analysis on the future development of data exposure risks, generating data security risk evolution results that are more consistent with the actual business operations of enterprises. Add a data exposure deviation coding layer before the event input layer. Write the deviation node, deviation position, deviation direction and deviation span in the data exposure deviation structure into the data exposure deviation coding layer to form the deviation constraint input path. Since the original NHP model mainly performs event evolution analysis based on the order of events, its input can only reflect multiple data exposure state units in the data exposure state chain, and cannot distinguish which data exposure states have deviated from the historical normal business path during the data exposure process. Therefore, when performing risk evolution analysis, both the data exposure process corresponding to normal business and the data exposure process that has deviated from the historical normal business path participate in the subsequent analysis, making it difficult to highlight the data exposure process that actually has risks and reducing the pertinence of risk evolution analysis on data exposure deviation. To address this, a data exposure deviation coding layer is added to uniformly encode the deviation nodes, deviation positions, deviation directions, and deviation spans in the data exposure deviation structure. Since the deviation nodes, deviation positions, deviation directions, and deviation spans are business analysis results obtained from data exposure state chain analysis, and are not data exposure state information that the model can directly use, the data exposure deviation structure is converted into deviation constraint information that the model can use through encoding processing, which serves as the constraint basis for subsequent risk evolution analysis. Furthermore, since the historical event memory layer, the risk induction intensity calculation layer, and the risk evolution reasoning layer are all based on event input, the data exposure deviation coding layer is set before the event input layer. This allows the data exposure risk evolution NHP model to obtain data exposure deviation information before receiving the data exposure state chain, and to form a deviation constraint input path that runs through the model analysis process. This ensures that the deviation constraint information continuously participates in subsequent historical event memory, risk induction intensity calculation, and risk evolution reasoning, rather than just participating in one event input processing. By improving the model, the Data Exposure Risk Evolution (NHP) model is transformed from performing a unified risk evolution analysis on all data exposure states to performing risk evolution analysis on data exposure processes that have deviated from the historical normal business path. This improves the model's relevance to the real data security risk evolution process and provides a foundation for generating subsequent risk evolution results. The memory processing path of the historical event memory layer for the exposed state unit of the preceding data is preserved, and the exposed state unit of the subsequent data is used as the input intensity function calculation layer for the content of the event to be reasoned. Data exposure risks typically exhibit continuous evolution. Subsequent data exposure state units are gradually formed based on the continuous changes of preceding data exposure state units. Relying solely on the current data exposure state unit is insufficient to fully reflect the data exposure process. Therefore, the historical event memory layer in the original NHP model is retained, memory processing is performed on preceding data exposure state units, and subsequent data exposure state units are input as the content to be inferred into the risk induction intensity calculation layer. The historical event memory layer is used to retain the historical evolutionary relationships between multiple data exposure state units, enabling the risk induction intensity calculation to combine the historical information corresponding to preceding data exposure state units, analyze the impact of preceding data exposure state units on subsequent data exposure state units, and provide a continuous data exposure process basis for subsequent risk evolution inference. The intensity function calculation layer is transformed into a risk-induced intensity calculation layer, which reads the data exposure state change process between adjacent data exposure state units and performs calculation processing on the risk-induced intensity between data access, data download, data sharing, data outgoing, and permission changes. The intensity function calculation layer in the original NHP model primarily calculates the intensity of the next data exposure state unit based on the order of occurrence between data exposure state units. Its results mainly reflect the evolutionary relationship between data exposure state units, but cannot reflect the risk formation process between data exposure behaviors. In actual business operations, data security risks are usually not directly formed by a single data exposure behavior, but rather gradually accumulate and continuously evolve during the continuous occurrence of multiple data exposure behaviors such as data access, data download, data sharing, data outreach, and permission changes. Therefore, analyzing only the order of occurrence between data exposure state units is insufficient to reflect the formation and evolution process of data exposure risks. To address this, the intensity function calculation layer in the original NHP model is transformed into a risk-induced intensity calculation layer. This layer uses the data exposure state change process between adjacent data exposure state units as the calculation object, analyzes the data exposure state change process from the preceding data exposure state to the subsequent data exposure state, and calculates the risk-induced intensity between data access, data download, data sharing, data outreach, and permission changes, thereby reflecting the degree of risk inducement between different data exposure behaviors. Since data exposure risk gradually forms and evolves during the continuous change of data exposure status, taking the process of data exposure status change between adjacent data exposure status units as the calculation object can completely preserve the process of gradual evolution of data exposure risk and avoid the loss of intermediate data exposure risk evolution process due to direct analysis of non-adjacent data exposure status units. The risk-induced intensity calculation layer is set before the data sensitivity constraint layer, so that the data exposure risk evolution NHP model first completes the data exposure risk-induced relationship analysis, and then performs data level constraint processing on the risk-induced intensity in combination with data classification and grading information. This allows the subsequent risk evolution reasoning to not only reflect the evolution process of data exposure risk, but also to reflect the differences in data security risks corresponding to different data levels. By transforming the intensity function calculation layer into a risk-induced intensity calculation layer, the original NHP model, which calculates the intensity of data exposure state units, is transformed into a calculation of the risk-induced relationship between data exposure behaviors. This changes the focus of the data exposure risk evolution NHP model from whether data exposure state units occur to how data exposure risks are formed and how they continue to evolve, providing a risk evolution basis for the subsequent risk evolution inference layer to generate risk evolution results. After the risk-induced intensity calculation layer, a data sensitivity constraint layer is added to read the data object identifier, data category, and data level from the data classification and grading information, and perform data level constraint processing on the risk-induced intensity. In enterprise data security management, different data objects have different data categories and data levels. Even if the data exposure process is the same, the data security risks corresponding to different data objects are still significantly different. For example, when the same data access, data download, data sharing, or data leakage occurs on data objects of different levels, the impact on enterprise business is not the same. It is difficult to reflect the differences in data security risks between different data objects by only performing subsequent risk evolution analysis based on the risk induction intensity. Based on this, a data sensitivity constraint layer is added after the risk induction intensity calculation layer. The data object identifier, data category, and data level in the data classification and grading information are read, and data level constraint processing is performed on the risk induction intensity. Among them, the data object identifier is used to establish the association between the risk induction intensity and the corresponding data object, the data category is used to reflect the data type of the data object, and the data level is used to reflect the importance of the data object. By combining the data classification and grading information, the risk induction intensity can be consistent with the data security requirements of the corresponding data object. The data sensitivity constraint layer does not change the calculation result of the risk induction intensity. Instead, it constrains the subsequent risk evolution analysis by combining data classification and grading information while maintaining the calculation result of the risk induction intensity. This enables the subsequent risk evolution reasoning to simultaneously consider the development process of data exposure risk and the importance of the data object. The data sensitivity constraint layer is placed after the risk induction intensity calculation layer because risk induction intensity reflects the objective risk induction relationship between data exposure behaviors, while data classification and grading information reflects the data security management requirements corresponding to different data objects. These two layers correspond to two different levels of information: risk evolution analysis and business management. Calculating risk induction intensity first ensures that the risk induction relationship is calculated solely based on the data exposure state change process, truly reflecting the data exposure risk evolution pattern between data access, data download, data sharing, data outreach, and permission changes. Then, data level constraints are applied to the risk induction intensity in conjunction with data classification and grading information, enabling risk evolution analysis to... This approach further highlights the differences in data security risks corresponding to different data categories and levels. If data classification and grading information is introduced before calculating the risk induction intensity, the data level will directly participate in the calculation, affecting the objective risk induction relationship between data exposure behaviors and making it difficult to accurately reflect the development process of data exposure risks themselves. By implementing risk induction intensity calculation and data level constraints in stages, while maintaining the integrity of the data exposure risk evolution law, risk constraint processing is combined with the enterprise's data classification and grading requirements. This ensures that the generated risk evolution results can both truly reflect the development process of data exposure risks and meet the actual data security management requirements of the enterprise. After improvement, the Data Exposure Risk Evolution (NHP) model can not only reflect the development process of data exposure risk, but also reflect the differences in data security risks corresponding to different data categories and data levels. This makes the generated risk evolution results more in line with the enterprise's data classification and grading management requirements, and provides a basis for subsequent risk level determination and data security protection processing. The event prediction output layer is transformed into a risk evolution reasoning layer, and the deviation constraint input path, the output results of the historical event memory layer, the risk induction intensity, and the data level constraint processing results are input into the risk evolution reasoning layer. The original NHP model's event prediction output layer primarily predicts the outcome of the next event based on information from preceding events. Its output mainly reflects the development direction of events, making it difficult to comprehensively analyze subsequent data exposure risks by integrating data exposure deviation, the evolution process of data exposure risks, and enterprise data classification and grading requirements. Therefore, it cannot meet the actual needs of enterprise data security risk evolution analysis. Based on this, the event prediction output layer is transformed into a risk evolution inference layer. The deviation constraint input path, the output results of the historical event memory layer, the risk induction intensity, and the data level constraint processing results are all input into the risk evolution inference layer. Among them, the deviation constraint input path is used to provide information that the data exposure process has deviated from the historical normal business path; the output results of the historical event memory layer are used to provide historical evolution information of the data exposure process; the risk induction intensity is used to provide the evolutionary relationship of data exposure risks; and the data level constraint processing results are used to provide the data security management requirements corresponding to data classification and grading. Based on the above information, the risk evolution reasoning layer performs comprehensive reasoning on the development process of data exposure risk, rather than predicting the subsequent data exposure state based solely on a single data exposure state. The risk evolution reasoning layer takes the deviation of data exposure as the starting point of reasoning, the historical data exposure process as the basis of reasoning, the risk induction relationship as the direction of evolution, and the data classification and grading requirements as business constraints. It generates risk evolution results according to the reasoning process of "deviation positioning, historical matching, risk evolution and business constraints", rather than directly predicting the next data exposure state. The risk evolution reasoning layer is placed at the last layer of the data exposure risk evolution NHP model because deviation constraint input, historical event memory, risk induction intensity calculation, and data level constraint processing are all fundamental to risk evolution reasoning. Only after the analysis is completed can a comprehensive reasoning be made about the future development process of data exposure risk. If risk evolution reasoning is executed in advance, the model cannot obtain complete data exposure risk analysis basis, affecting the completeness of the risk evolution analysis results. Through improvements, the Data Exposure Risk Evolution (NHP) model has transformed from predicting the development of the next data exposure state to comprehensively analyzing data exposure deviations, the data exposure risk evolution process, and data classification and grading requirements. It performs risk evolution reasoning on the future development process of data exposure risks, making the generated risk evolution results more consistent with the actual business process of the gradual formation and continuous evolution of enterprise data security risks, and providing a basis for subsequent risk level determination and data security protection measures. Establish interlayer connections according to the order of data exposure deviation coding layer, event input layer, historical event memory layer, risk induction intensity calculation layer, data sensitivity constraint layer, and risk evolution reasoning layer to generate the data exposure risk evolution NHP model; The improvements to the Data Exposure Risk Evolution NHP model include: 1. Regarding model input, the original NHP model only performs event evolution analysis based on event input, failing to distinguish data exposure states that have deviated from historical normal business paths during the data exposure process. The NHP model adds a data exposure deviation encoding layer, encoding the data exposure deviation structure to form deviation constraint input paths, enabling the model to perform risk evolution analysis around data exposure processes that have deviated from historical normal business paths, improving the targeting of data security risk identification. 2. Regarding risk calculation, the intensity function calculation layer in the original NHP model mainly calculates the intensity of event occurrence, making it difficult to reflect the formation and continuous evolution process of data exposure risks. The NHP model transforms the intensity function calculation layer into a risk-induced intensity calculation layer, based on the data exposure state change process between adjacent data exposure state units, calculating the risk-induced intensity between data access, data download, data sharing, data outreach, and permission changes, enabling the model to reflect the evolutionary pattern of data exposure risks and improving risk evolution analysis capabilities. 3. Regarding business constraints, the original NHP model did not consider the data classification and grading requirements for different data objects, making it difficult to reflect the differences in data security risks corresponding to different data objects in the same data exposure process. The Data Exposure Risk Evolution NHP model adds a data sensitivity constraint layer. While maintaining the calculation results of risk induction intensity, it combines data classification and grading information to perform data level constraint processing on the risk induction intensity, making the risk evolution results more in line with the enterprise's data security management requirements. Fourth, in terms of inference output, the event prediction output layer of the original NHP model mainly predicts the occurrence of the next event, making it difficult to comprehensively analyze the risk by considering data exposure deviation, the data exposure risk evolution process, and data classification and grading requirements. The Data Exposure Risk Evolution NHP model transforms the event prediction output layer into a risk evolution inference layer. It integrates the deviation constraint input path, the output results of the historical event memory layer, the risk induction intensity, and the data level constraint processing results to perform comprehensive inference on the future development process of data exposure risk, transforming the model from event prediction to risk evolution analysis and generating data security risk evolution results that are more in line with the actual business of enterprises. The Data Exposure Risk Evolution (NHP) model, while retaining the historical event memory mechanism of the NHP model, has been improved around four key aspects: model input, risk calculation, business constraints, and inference output. This makes the NHP model more in line with the actual business needs of enterprises in terms of the gradual formation, continuous evolution, and hierarchical management of data security risks. The Data Exposure Risk Evolution (NHP) model comprises a data exposure deviation coding layer, an event input layer, a historical event memory layer, a risk induction intensity calculation layer, a data sensitivity constraint layer, and a risk evolution inference layer. Multiple data exposure state units from the data exposure state chain set are input to the event input layer, and the data exposure deviation structure is input to the data exposure deviation coding layer. The data exposure deviation coding layer outputs the deviation constraint input path. The event input layer connects to the historical event memory layer, which in turn connects to the risk induction intensity calculation layer, which connects to the data sensitivity constraint layer, and finally, the data sensitivity constraint layer connects to the risk evolution inference layer. Simultaneously, the deviation constraint input path, the output results of the historical event memory layer, the risk induction intensity, and the data level constraint processing results are all input to the risk evolution inference layer. The risk evolution inference layer performs risk evolution analysis and generates risk evolution results. The entire model completes the data flow transmission in the order of data exposure state unit input, historical information retention, risk induction intensity calculation, data level constraints, and risk evolution inference. The output results of each layer serve as the input basis for subsequent layers, enabling the data exposure risk to gradually complete the risk evolution analysis along the data exposure state change process between multiple data exposure state units. The training data for the Data Exposure Risk Evolution (NHP) model comes from the enterprise's historical data security behavior data. First, data processing is performed on historical data access logs, data download logs, data sharing logs, data outbound logs, permission change logs, approval information, and data classification and grading information to generate a set of historical data exposure state chains, historical data exposure deviation structures, and historical data classification and grading information. Combined with the risk level of each data object in the subsequent actual occurrence, a model training sample is formed. There are a total of 102,368 training samples, including 81,894 training samples, 10,237 validation samples, and 10,237 test samples. During training, multiple data exposure state units from the historical data exposure state chain set are input into the event input layer, the historical data exposure deviation structure is input into the data exposure deviation encoding layer, and the data classification and grading information is input into the data sensitivity constraint layer. The historical event memory layer continuously updates the historical event memory information according to the order of the data exposure state units in the data exposure state chain. The risk induction intensity calculation layer calculates the risk induction intensity based on the data exposure state change process between the preceding and subsequent data exposure state units. The data sensitivity constraint layer performs data level constraint processing on the risk induction intensity in combination with the data category and data level. The risk evolution reasoning layer integrates the deviation constraint input path, the output result of the historical event memory layer, the risk induction intensity, and the data level constraint processing result to perform risk evolution reasoning on the future development process of the data exposure risk corresponding to each data object and generate the risk evolution result. The model's predicted risk-induced intensity is compared with historical risk-induced intensity. The model loss is calculated based on the difference between the two. Backpropagation training is used to synchronously update the model parameters of the data exposure deviation encoding layer, historical event memory layer, risk-induced intensity calculation layer, data sensitivity constraint layer, and risk evolution inference layer. The Adam optimizer is used for training, with an initial learning rate of 0.0005, a batch size of 32, a weight decay coefficient of 0.0001, and a maximum training epoch of 100 epochs. When the model loss decreases by less than 0.001 for 8 consecutive epochs on the validation set, model training is stopped, and the corresponding model parameters are saved, resulting in the trained data exposure risk evolution NHP model.

[0023] In this embodiment, the generation of risk evolution results specifically includes: Read multiple data exposure state chains from the data exposure state chain set, extract multiple data exposure state units from each data exposure state chain, and input the multiple data exposure state units into the event input layer according to the data object identifier; Read the deviation nodes, deviation positions, deviation directions, and deviation spans in the data exposure deviation structure, and input the deviation nodes, deviation positions, deviation directions, and deviation spans into the data exposure deviation coding layer to generate deviation constraint input results; The generation of deviation constraint input results specifically includes: Read the deviation node, deviation position, deviation direction, and deviation span respectively, and create corresponding fields for each; write the deviation node, deviation position, deviation direction, and deviation span into the corresponding fields respectively; combine the contents of each field in the order of the fields to generate the deviation constraint input results; The historical event memory layer performs historical memory processing on the preceding data exposure state unit in each data exposure state chain, and inputs the data exposure state change process between the subsequent data exposure state unit and the preceding data exposure state unit into the risk induction intensity calculation layer. At the same time, the risk induction intensity result is input into the data sensitivity constraint layer, and the risk constraint result is generated based on the data category and data level in the data classification and grading information. The generation of risk-induced intensity results specifically includes: Extract the data access, data download, data sharing, data outgoing, and permission change behaviors that have changed in the preceding and subsequent data exposure state units respectively; count the number of consecutive occurrences of each behavior change in the historical data exposure state chain; determine the risk induction intensity between corresponding behaviors based on the number of consecutive occurrences of each behavior change; summarize the risk induction intensity between each behavior and generate the risk induction intensity result; The induction of subsequent data exposure state changes by changes in preceding data exposure state is usually manifested in the continuous occurrence of corresponding behavioral changes. When the same behavioral change occurs a large number of times in the historical data exposure state chain, it indicates that the corresponding subsequent behavioral change continues to occur after the preceding behavioral change, and the behavioral change has a strong and continuous inducing effect. Conversely, when the number of consecutive occurrences is small, it indicates that the inducing relationship between corresponding behavioral changes is weak. Therefore, the number of consecutive occurrences of behavioral changes is used as the basis for calculating the risk induction intensity. The risk induction intensity of the corresponding behavioral change relationship is determined based on the number of consecutive occurrences, so that the risk induction intensity can reflect the actual induction degree between each data exposure behavior. The generation of risk constraint results specifically includes: Read the data object identifier, data category, and data level from the data classification and grading information, and establish a correspondence between the data object identifier and the risk induction intensity result; determine the corresponding data security management requirements based on the data category, which include data access behavior, data download behavior, data sharing behavior, data external release behavior, and permission change behavior, and determine the corresponding risk constraint level based on the data level; map the data security management requirements and risk constraint levels to the risk induction intensity result to generate the risk constraint result; The deviation constraint input results, the output results of the historical event memory layer, and the risk constraint results are input into the risk evolution reasoning layer to generate a candidate set of future data exposure events corresponding to each data object. The generation of the candidate set of future data exposure events specifically includes: Read the deviation constraint input results, the historical event memory layer output results, and the risk constraint results, and establish a corresponding relationship according to the data object identifier; based on the deviation constraint input results, filter the data exposure state chains that meet the deviation constraints from the historical event memory layer output results, and filter the data exposure state chains that meet the risk constraints according to the risk constraint results; extract the corresponding subsequent data exposure events in each data exposure state chain, and perform summary processing according to the data object identifier to generate a candidate set of future data exposure events; Based on the candidate set of future data exposure events, determine the occurrence trend of future data exposure events for data objects and the corresponding degree of risk evolution, and generate risk evolution results corresponding to each data object; The generation of risk evolution results specifically includes: Read multiple future data exposure events from the candidate set of future data exposure events and perform classification and merging processing according to the data object identifier; perform sorting processing according to the order of occurrence of multiple future data exposure events corresponding to each data object to determine the occurrence trend of future data exposure events; determine the corresponding risk evolution degree based on the occurrence trend of future data exposure events; perform corresponding association between the occurrence trend of future data exposure events and the risk evolution degree according to the data object identifier to generate the risk evolution result.

[0024] In this embodiment, the generation of data security protection results specifically includes: Based on the risk evolution results corresponding to each data object, the trend of future data exposure events and the degree of risk evolution are extracted to determine the risk level corresponding to each data object; The corresponding protection strategy is invoked based on the risk level of each data object, and the execution content of the corresponding protection strategy is determined in combination with the future trend of data exposure events. In accordance with the corresponding protection policy, data security protection measures are implemented for each data object, including data access restrictions, data download restrictions, data sharing restrictions, data outbound restrictions, or permission change restrictions. The risk level, corresponding protection strategy, and execution results of data security protection processing are associated with the data object identifier to generate intelligent data security protection results corresponding to each data object.

[0025] Example 1: To verify the feasibility of this invention in practice, it was applied to the data security management of a large state-owned enterprise. This enterprise routinely operates finance, legal, audit, contract management, collaborative office work, and electronic archives. Data flows between these businesses through a unified business platform. Important data objects such as contract documents, financial statements, audit materials, budget documents, procurement materials, and business analysis reports are continuously accessed, downloaded, shared, distributed, and have their permissions adjusted during business processes. As cross-departmental collaboration increases, the same data object often flows continuously between multiple business stages. For example, after the finance department generates a budget document, it is downloaded and shared by the business departments to the procurement department. After the procurement department completes its business processing, the document continues to flow to the legal department for review, then to the management department for approval, and finally archived. Existing data security management methods mainly rely on triggering risk alerts based on single abnormal behaviors. When a data object experiences multiple data exposure behaviors consecutively, only the current abnormal behavior can be identified. It is difficult to analyze the continuous evolution of the data exposure state or determine whether the current data flow has gradually deviated from the normal business path. Therefore, it is prone to problems such as delayed risk detection and insufficient early warning capabilities.

[0026] This embodiment continuously collects data security behavior data generated during the company's business operations over 30 days, collecting a total of 4.627 million data access logs, 386,000 data download logs, 274,000 data sharing logs, 43,000 data outbound logs, 21,000 permission change logs, and 168,000 approval records, involving 102,361 data objects, including 4,862 first-level data objects, 21,837 second-level data objects, and 75,662 third-level data objects. First, all logs were processed for time alignment, missing data completion, abnormal record removal, and format standardization to form a standardized data security behavior dataset. Then, the number of accessors, the number of departments involved, the number of downloads, the number of shares, the number of outbound transmissions, and the number of permission changes for each data object were counted to generate multiple data exposure state units, and a data exposure state chain corresponding to each data object was established in chronological order.

[0027] In actual operation, taking a certain level contract document as an example, the contract document should, according to the normal business process, go through the data flow path of "business department access, legal approval, business department download, procurement department sharing, and contract archiving" in sequence. However, in the actual business operation, the data exposure state chain corresponding to this contract document is "business department access, business department download, procurement department sharing, download again and permission adjustment → data outgoing". This invention first performs comparative analysis on the current data exposure state chain based on the historical normal business path, identifies the business department download as a deviation node, determines that the deviation position is located at the 4th state position of the data exposure state chain, identifies that the deviation direction changes from the normal archiving direction to the continued outgoing direction, and determines that the deviation span is two data exposure state units, thus forming the corresponding data exposure deviation structure. Subsequently, the data exposure state chain is input into the historical event memory layer, and the data exposure deviation structure is input into the data exposure deviation coding layer to generate deviation constraint input results. Then, combined with the fact that the contract document belongs to the first-level contract data and the corresponding first-level data level, risk constraint results are generated, and these are jointly input into the improved neural Hawkes process model to perform risk evolution reasoning.

[0028] During the risk evolution reasoning process, the historical event memory layer first reads the evolution records of data exposure states that are the same as the current contract document in the historical business process. The risk induction intensity calculation layer calculates the corresponding risk induction intensity based on the changing relationships of data access behavior, data download behavior, data sharing behavior, data outbound behavior, and permission change behavior between continuous data exposure states. The data sensitivity constraint layer combines the data category and first-level data level corresponding to the first-level contract documents to form risk constraint results. The risk evolution reasoning layer integrates the deviation constraint input results, historical event memory results, and risk constraint results to screen historical evolution records that meet the evolution characteristics of the current data exposure state, generate a candidate set of future data exposure events, and further determine the trend of future data exposure events and the corresponding degree of risk evolution. When the model identifies that the contract document has continuously exhibited data exposure behaviors such as downloading, sharing, re-downloading, and permission adjustments, and predicts that there is a trend of continued data outbound behavior in the next stage, a risk warning is completed before the actual outbound behavior occurs. Data security management personnel promptly suspend the outbound permission of the contract document and conduct manual review of the relevant operators. Finally, it is confirmed that the outbound behavior is an abnormal data flow behavior, preventing the continued spread of important contract documents.

[0029] Table 1. Comprehensive Comparison of Data Exposure Risk Identification Performance

[0030] As shown in Table 1, all three methods can identify enterprise data exposure risks, but their overall performance differs significantly. Traditional rule-based detection methods primarily rely on fixed rules to judge behaviors such as data access, data download, data sharing, data outreach, and permission changes. They can only identify abnormal behaviors that have already occurred and lack the ability to continuously analyze the evolution of data exposure states resulting from multiple consecutive data exposure behaviors. Therefore, the risk identification accuracy is only 84.68%, and the risk trend identification rate is only 71.36%. Furthermore, because the rules are independent of each other, it is impossible to combine historical business processes to determine whether data exposure behaviors have gradually deviated from the normal business path. Therefore, the average warning lead time is only 13.2 minutes, and risk alarms are typically triggered only after multiple abnormal behaviors occur consecutively, resulting in a certain lag in risk detection. In addition, the false alarm rate of traditional rule-based detection methods reaches 11.74%, requiring data security managers to manually review an average of 19.4 times per day, which not only increases manual analysis costs but also easily interferes with normal business operations.

[0031] Traditional deep learning methods utilize historical data to train anomaly detection models, enabling them to learn the characteristic information of different data exposure behaviors. This improves risk identification accuracy to 86.42%, a 1.74 percentage point increase compared to traditional rule-based detection methods; risk trend identification accuracy to 79.52%, an 8.16 percentage point increase; average warning lead time to 24.1 minutes; false alarm rate to 8.63%; and manual review workload to 15.9 times / day. This demonstrates that traditional deep learning methods can improve anomaly behavior identification capabilities and have certain advantages in identifying single data exposure behaviors. However, this method primarily relies on current behavioral characteristics for judgment, lacking modeling of the continuous evolution of data exposure states and failing to fully utilize the sustained influence relationships between historical events and data classification and grading information for joint analysis. Therefore, its ability to predict the development trend of future data exposure events remains limited.

[0032] This invention constructs a data exposure state chain to continuously model the data exposure status of enterprise data objects during continuous business processes. It also identifies trends in data flow processes that deviate from normal business paths by combining data exposure deviation structure identification. Furthermore, it utilizes the data exposure deviation coding layer, risk induction intensity calculation layer, data sensitivity constraint layer, and risk evolution reasoning layer in the Data Exposure Risk Evolution (NHP) model to perform collaborative reasoning on historical event memories, data exposure deviation information, and data classification and grading constraint information. This enables prediction of future data exposure events and risk evolution analysis. Therefore, compared with traditional rule-based detection methods, this invention achieves a risk identification accuracy of 89.58%, an improvement of 4.90 percentage points; a risk trend identification rate of 88.41%, an improvement of 17.05 percentage points; an average early warning time of 35.4 minutes, completing risk warnings 22.2 minutes ahead of schedule; a false alarm rate reduced to 5.58%, a decrease of 6.16 percentage points; and a reduction in manual review workload to 12.5 times / day, a decrease of 6.9 times / day. Compared with traditional deep learning methods, this invention improves the risk identification accuracy by 3.16 percentage points, the risk trend identification rate by 8.89 percentage points, the average early warning time by 11.3 minutes, the false alarm rate by 3.05 percentage points, and reduces the workload of manual review by 3.4 times per day.

[0033] The experimental results above demonstrate that the advantages of this invention are not only reflected in the improved accuracy of risk identification, but also in its ability to describe the continuous evolution of data exposure behavior using a data exposure state chain. By accurately identifying the development characteristics of data flow deviating from normal business paths through data exposure deviation structures, and combining the data exposure risk evolution NHP model with collaborative reasoning based on historical event memory, risk induction intensity, and data classification and grading constraints, this invention enables prediction of future data exposure event trends and risk evolution analysis. Therefore, this invention can identify potential risks before they escalate further, significantly improving risk trend identification capabilities and timely warnings while ensuring the accuracy of risk identification. It also effectively reduces false alarm rates and manual review workload, better meeting the actual business needs of enterprises for data security risk warning and proactive protection.

[0034] The above are merely preferred embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. A data security intelligent protection method based on deep learning, characterized in that, Includes the following steps: Collect data access logs, data download logs, data sharing logs, data outbound logs, permission change logs, approval and data classification and grading information, and preprocess them to generate a data security behavior dataset; Perform exposure statistics processing on each data object in the data security behavior dataset to generate a set of data exposure status units; For multiple data exposure state units corresponding to the same data object within a continuous time window, establish state connection relationships, form a state evolution sequence, and generate a set of data exposure state chains; By comparing and analyzing the data exposure state chain with the historical normal business path, the deviation nodes, deviation positions, deviation directions and deviation spans are identified, and a data exposure deviation structure is generated. A data exposure risk evolution NHP model is constructed, which includes a data exposure deviation coding layer, a risk induction intensity calculation layer, a data sensitivity constraint layer, and a risk evolution inference layer. The data exposure state chain set is used as the event evolution input, and the data exposure deviation structure is used as the deviation constraint input. The risk evolution analysis is performed using the data exposure risk evolution NHP model to generate risk evolution results. Based on the risk evolution results, the corresponding risk level is determined, and the protection strategy corresponding to the risk level is invoked to perform data security protection processing and generate data security protection results.

2. The data security intelligent protection method based on deep learning according to claim 1, characterized in that, The data access log includes the data object identifier, the accessor identifier, the department, and the access time; the data download log includes the data object identifier, the downloader identifier, and the download time; the data sharing log includes the data object identifier, the sharing initiator identifier, the sharing recipient identifier, and the sharing time; the data outbound log includes the data object identifier, the outbound sender identifier, the outbound target identifier, and the outbound time; the permission change log includes the data object identifier, the changing sender identifier, the change time, and the change result; the approval information includes the data object identifier, the approval node name, the approval processing personnel identifier, the approval processing time, and the approval result; the data classification and grading information includes the data object identifier, the data category, and the data level; the preprocessing specifically includes time alignment, missing data completion, abnormal record removal, and format standardization.

3. The data security intelligent protection method based on deep learning according to claim 1, characterized in that, The generation of the data exposure state unit set specifically includes: Read the data access logs, data download logs, data sharing logs, data outgoing logs, and permission change logs from the data security behavior dataset, extract the data object identifiers corresponding to each behavior record, and perform merging processing on each behavior record according to the data object identifiers to generate a data object behavior record set; For each data object behavior record in the data object behavior record set, count the number of corresponding visitor identifiers and the number of departments to which they belong, and generate contact statistics results. Count the number of download records, sharing records, outgoing records, and permission change records for each data object, and generate data exposure behavior statistics results. The contact statistics and data exposure behavior statistics are correlated according to the data object identifier to generate a data exposure status unit corresponding to each data object; Perform summary processing on multiple data exposure status units to generate a set of data exposure status units.

4. The data security intelligent protection method based on deep learning according to claim 1, characterized in that, The generation of the data exposure state chain set specifically includes: Read multiple data exposure state units from the data exposure state unit set, extract the data object identifier corresponding to each data exposure state unit, and perform classification and merging processing on multiple data exposure state units according to the data object identifier to generate a data object state unit set; For each data object in the data object state unit set, multiple data exposure state units corresponding to each data object are sorted according to the time sequence of the corresponding behavior records to generate a state time series. Traverse adjacent data exposure state units in the state time series, determine the connection relationship between the preceding and subsequent data exposure state units as state connection relationships, and perform aggregation processing on multiple state connection relationships to generate a set of state connection relationships; Based on multiple data-exposed state units and state connection relationships in the state time series, a continuous state connection is established to form a state evolution sequence corresponding to each data object; The state evolution sequence corresponding to each data object is determined as the corresponding data exposure state chain, and multiple data exposure state chains are aggregated to generate a set of data exposure state chains.

5. The data security intelligent protection method based on deep learning according to claim 1, characterized in that, The generation of the data exposure deviation structure specifically includes: Read the data object identifier, approval node name, approval personnel identifier, approval processing time and approval result from the approval information, filter the approval information with the approval result as passed, perform merging processing on the approval information according to the data object identifier, perform sequential processing on multiple approval node names according to the approval processing time, and generate the historical normal business path corresponding to each data object. Obtain multiple data exposure state chains from the data exposure state chain set, extract the data object identifier corresponding to each data exposure state chain, and match the data exposure state chain corresponding to the same data object with the historical normal business path; Traverse the data exposure state chain and the corresponding historical normal business path, identify data exposure state units in the data exposure state chain that are not in the business flow process corresponding to the historical normal business path, and generate deviation nodes. The deviation position is determined according to the order of the deviation nodes in the corresponding data exposure state chain, and the preceding data exposure state unit corresponding to the deviation node is extracted. The process of the preceding data exposure state unit changing to the data exposure state of the deviation node is determined as the deviation direction. Perform a count of consecutively occurring deviation nodes in the same data exposure state chain to generate the deviation span; The deviation nodes, deviation positions, deviation directions, and deviation spans are associated with corresponding data object identifiers to generate a data exposure deviation structure.

6. The data security intelligent protection method based on deep learning according to claim 1, characterized in that, The construction of the neural Hawkes process model for the evolution of data exposure risk specifically includes: The NHP infrastructure is invoked, which includes an event input layer, a historical event memory layer, an intensity function calculation layer, and an event prediction output layer, and multiple data exposure state units in the data exposure state chain set are set as event input content. Add a data exposure deviation coding layer before the event input layer. Write the deviation node, deviation position, deviation direction and deviation span in the data exposure deviation structure into the data exposure deviation coding layer to form the deviation constraint input path. The memory processing path of the historical event memory layer for the exposed state unit of the preceding data is preserved, and the exposed state unit of the subsequent data is used as the input intensity function calculation layer for the content of the event to be reasoned. The intensity function calculation layer is transformed into a risk-induced intensity calculation layer, which reads the data exposure state change process between adjacent data exposure state units and performs calculation processing on the risk-induced intensity between data access, data download, data sharing, data outgoing, and permission changes. After the risk-induced intensity calculation layer, a data sensitivity constraint layer is added to read the data object identifier, data category, and data level from the data classification and grading information, and perform data level constraint processing on the risk-induced intensity. The event prediction output layer is transformed into a risk evolution reasoning layer, and the deviation constraint input path, the output results of the historical event memory layer, the risk induction intensity, and the data level constraint processing results are input into the risk evolution reasoning layer. The data exposure risk evolution NHP model is generated by establishing interlayer connections in the order of data exposure deviation coding layer, event input layer, historical event memory layer, risk induction intensity calculation layer, data sensitivity constraint layer, and risk evolution inference layer.

7. The data security intelligent protection method based on deep learning according to claim 1, characterized in that, The generation of the risk evolution result specifically includes: Read multiple data exposure state chains from the data exposure state chain set, extract multiple data exposure state units from each data exposure state chain, and input the multiple data exposure state units into the event input layer according to the data object identifier; Read the deviation nodes, deviation positions, deviation directions, and deviation spans in the data exposure deviation structure, and input the deviation nodes, deviation positions, deviation directions, and deviation spans into the data exposure deviation coding layer to generate deviation constraint input results; The historical event memory layer performs historical memory processing on the preceding data exposure state unit in each data exposure state chain, and inputs the data exposure state change process between the subsequent data exposure state unit and the preceding data exposure state unit into the risk induction intensity calculation layer. At the same time, the risk induction intensity result is input into the data sensitivity constraint layer, and the risk constraint result is generated based on the data category and data level in the data classification and grading information. The deviation constraint input results, the output results of the historical event memory layer, and the risk constraint results are input into the risk evolution reasoning layer to generate a candidate set of future data exposure events corresponding to each data object. Based on the candidate set of future data exposure events, determine the occurrence trend of future data exposure events for data objects and the corresponding degree of risk evolution, and generate risk evolution results corresponding to each data object.

8. The data security intelligent protection method based on deep learning according to claim 1, characterized in that, The generation of the data security protection results specifically includes: Based on the risk evolution results corresponding to each data object, the trend of future data exposure events and the degree of risk evolution are extracted to determine the risk level corresponding to each data object; The corresponding protection strategy is invoked based on the risk level of each data object, and the execution content of the corresponding protection strategy is determined in combination with the future trend of data exposure events. In accordance with the corresponding protection policy, data security protection measures are implemented for each data object, including data access restrictions, data download restrictions, data sharing restrictions, data outbound restrictions, or permission change restrictions. The risk level, corresponding protection strategy, and execution results of data security protection processing are associated with the data object identifier to generate intelligent data security protection results corresponding to each data object.