Data leakage behavior checking and controlling method and device, computer equipment and storage medium
By using intelligent behavior analysis and machine learning, the problem of insufficient monitoring of internal threats in traditional security protection systems has been solved, enabling precise prevention and efficient response to data breaches and improving the security defense capabilities of enterprises.
Patent Information
- Application Number
- CN202511352622.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-22
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2045-09-22
AI Technical Summary
Traditional security systems lack effective monitoring and analysis of internal threats, making it difficult to warn and prevent data breaches. They also suffer from problems such as data silos, alarm fatigue, and a lack of intelligent analysis.
By using intelligent behavior analysis, machine learning, and association rules, raw user behavior data is obtained, cleaned, parsed, and normalized. Machine learning algorithms are then used for real-time risk rule matching to generate a visual risk profile and output alarm information.
It enables precise prevention and control of internal data leakage risks, improves threat detection efficiency, breaks down data silos, provides early warning and in-process intervention capabilities, and protects the core value of enterprises.
Smart Images

Figure CN120880778A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data security prevention technology in the logistics industry, and in particular to a method, device, computer equipment and storage medium for investigating and controlling data leakage behavior. Background Technology
[0002] Traditional security systems primarily target external attacks, while internal threats, due to their legitimate privileges and difficulty in prevention, pose a greater threat. Malicious operations by internal personnel, account theft, data theft, and unintentional violations are among the core causes of data breaches. Currently, there is a lack of effective technical means to continuously monitor and analyze internal personnel behavior, making it impossible to provide pre-emptive warnings or real-time intervention. Furthermore, most enterprises currently rely on logs from traditional security devices (such as firewalls and IDS) and system logs for post-incident auditing, which presents the following pain points: Data silo problem: Logs are scattered across various systems with different formats, making it difficult to correlate and analyze them, resulting in extremely low investigation efficiency.
[0003] Alarm fatigue and false alarms: Traditional rule-based alarms are mostly single-point, low-level events that lack contextual correlation, generating a lot of noise and masking the real threat.
[0004] Lack of intelligent analysis: It relies mainly on fixed rules and cannot effectively identify unknown threats and slow, lurking attacks, among other problems. Summary of the Invention
[0005] This application provides a method, apparatus, computer device, and storage medium for investigating and controlling data leakage behavior. It aims to accurately prevent and control data leakage risks, protect core corporate value, and maintain brand reputation and customer trust through intelligent behavior analysis, machine learning, association rules, and risk profiling.
[0006] The technical solution is as follows: In a first aspect, embodiments of this application provide a method for investigating and controlling data leakage behavior, including: This includes obtaining raw user behavior data from VPNs, bastion hosts, databases, office automation systems, cloud platforms, and terminals. The raw user data obtained includes abnormal behaviors of users obtaining data by operating documents, taking screenshots, or recording videos. These abnormal behaviors include access to sensitive data, unauthorized operations, and login at abnormal times. Preprocessing of acquired raw user behavior data includes cleaning, parsing, enriching, and normalizing storage; Based on preset risk rules, the preprocessed data is matched in real time using machine learning algorithms. A risk score is obtained based on the matching results. The risk score is used to determine whether there is a risk of violation. A visual risk profile is generated based on the risk of violation. Output alarm information according to the alarm rules.
[0007] Furthermore, the preprocessing of the acquired raw user behavior data includes cleaning, parsing, enriching and normalizing the storage. It also includes cleaning, parsing, enriching and normalizing the acquired raw user behavior data by means of log parsing, field standardization, invalid data filtering, reverse IP address resolution and information completion by associating user and asset metadata.
[0008] Furthermore, the step of performing real-time risk rule matching on the preprocessed data using a machine learning algorithm based on preset risk rules, obtaining a risk score based on the matching results, determining whether there is a risk of violation based on the risk score, and generating a visual risk profile based on the risk of violation, includes: Furthermore, the step of performing real-time risk rule matching on the preprocessed data using a machine learning algorithm based on preset risk rules, obtaining a risk score based on the matching results, determining whether there is a risk of violation based on the risk score, and generating a visual risk profile based on the risk of violation, includes: Obtain the risk scenario triggered by each user, and obtain the first risk score configured according to that risk scenario; Based on the first risk score, obtain the second risk score for the user after it has decayed over time according to the decay rule for this risk scenario; Based on the second risk score, a risk ranking list is generated and sorted according to the second risk score, and displayed in the form of a visual risk profile.
[0009] Further, the step of acquiring each user-triggered risk scenario and acquiring a first risk score configured according to that risk scenario includes: Obtain the types of events triggered in this risk scenario and the corresponding weight scores assigned to each type; Based on the weighted scores, a comprehensive score is obtained for the risk scenario, based on the chronological order of various triggering events and the number of triggers in that chronological order. The number of triggers includes the number of times the events are triggered repeatedly in the chronological order of various triggering events or the number of times a single triggering event is triggered repeatedly. This comprehensive score is the first risk score.
[0010] Furthermore, the step of generating a risk ranking list sorted by the second risk score and displaying it in the form of a visual risk profile includes: when generating the visual risk profile of the risk ranking list sorted by the second risk score, also generating a first-level cascaded personal risk details visual profile for each user, including the user's organizational structure information, risk summary, risk composition analysis, self-data leakage risk, list of key risk events, name of the triggered risk scenario, data source, and risk score contributed by each event; wherein the organizational structure information includes displaying the user's name, department, position, and job title; the risk summary includes displaying the current total risk score, risk ranking within the enterprise, and a recent risk change trend chart; the risk composition analysis includes the source composition of the user's risk score displayed in the form of a pie chart or bar chart; the self-data leakage risk includes privilege abuse risk and violation operation risk; the list of key events includes all high-risk events triggered by the user in reverse chronological order, including event time, name of the triggered risk scenario, data source, and risk score contributed by each event.
[0011] Furthermore, the cascading personal risk details visualization profile for each user also includes a deep investigation button window set in the first cascading personal risk details visualization profile page for each user. Through this window, a second cascading personal risk details visualization profile can be cascaded. The content of the second cascading personal risk details visualization profile includes displaying behavioral trajectory and alarm details. The behavioral trajectory includes all operation logs of the user at a specific time. The alarm details include alarm records generated by the user's risky behavior and their processing status.
[0012] Furthermore, the step of outputting alarm information according to alarm rules includes outputting alarm notifications based on alarm triggering conditions, including notification method, alarm assignment, alarm release time, and time limit, and displaying processing results including processing opinions, evidence upload status, and status tracking status.
[0013] Secondly, embodiments of the present invention also provide a data leakage behavior detection and control device, comprising: The raw data acquisition unit is used to acquire raw user behavior data from VPN, bastion host, database, office OA, cloud platform, and terminal. The acquired raw user data includes abnormal behaviors of users acquiring data by operating documents, taking screenshots, or shooting videos. The abnormal behaviors include access to sensitive data, unauthorized operations, and login at abnormal times. The data preprocessing unit is used to preprocess the acquired raw user behavior data, including cleaning, parsing, enriching and normalizing the data before storage. The violation risk assessment unit is used to perform real-time risk rule matching on the preprocessed data obtained according to preset risk rules using machine learning algorithms, obtain a risk score based on the matching results, determine whether there is a violation risk based on the risk score, and generate a visual risk profile based on the violation risk. The alarm unit determines the alarm information to output based on the alarm rules.
[0014] Thirdly, this embodiment also provides a computer storage medium storing a plurality of instructions adapted for loading and execution by a processor as described above.
[0015] Fourthly, this embodiment also provides a computer device, including: a processor and a memory; wherein the memory stores a computer program adapted to be loaded by the processor and executed as described above.
[0016] The beneficial effects of the technical solutions provided in some embodiments of this application include at least the following: By acquiring raw user behavior data from VPNs, bastion hosts, databases, office automation systems, cloud platforms, and terminals, including abnormal user behaviors such as accessing sensitive data, unauthorized operations, and logins at unusual times, the system achieves precise data leakage risk prevention and control. This protects the company's core value, maintains its brand reputation, and gains high customer trust. The system also acquires raw user behavior data through preprocessing, including cleaning, parsing, enriching, and normalizing storage. Based on preset risk rules, machine learning algorithms are used to perform real-time risk rule matching on the preprocessed data. Risk scores are obtained based on the matching results, and the presence of violations is determined based on these scores. A visual risk profile is generated based on the violation risks. Alarm information is output according to alarm rules. Attached Figure Description
[0017] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0018] Figure 1 This is a flowchart of the method provided in the embodiments of this application; Figure 2 This application provides a first-level visualized profile of a user's personal risk details in an embodiment of the application. Figure 3 This is a schematic diagram of the data leakage behavior detection and control device provided in the embodiments of this application; Figure 4 This is a schematic diagram of the data leakage behavior detection and control device provided in the embodiments of this application; Figure 5 This is a schematic diagram of a computer storage medium provided in an embodiment of this application. Detailed Implementation
[0019] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0020] The flowchart shown in the attached diagram is for illustrative purposes only and does not necessarily include all content and operations / steps, nor does it necessarily have to be performed in the order described. For example, some operations / steps can be broken down, combined, or partially merged, so the actual execution order may change depending on the actual situation.
[0021] It should be understood that, in order to clearly describe the technical solutions of the embodiments of the present invention, the terms "first" and "second" are used in the embodiments of the present invention to distinguish identical or similar items with essentially the same function and effect. Those skilled in the art will understand that the terms "first" and "second" do not limit the quantity or execution order, and the terms "first" and "second" are not necessarily different.
[0022] It should be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the scope of the application. As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise.
[0023] It should also be understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0024] In the digital economy era, enterprises' core assets are increasingly digitized, and their internal operating environments are becoming increasingly complex. Security threats from internal personnel, such as unauthorized operations, data breaches, and malicious damage, are among the weakest links in enterprise security systems due to their high privilege levels, high trust levels, and difficulty in prevention. Traditional security protection systems primarily focus on responding to external attacks, lacking effective perception, analysis, and response capabilities to internal threats, resulting in pain points of being "invisible, unclear, and uncontrollable." This application aims to construct an enterprise-level intelligent internal security hub, integrating real-time stream computing, machine learning modeling, correlation graph analysis, and multimodal content recognition into a comprehensive solution. By integrating, standardizing, and intelligently analyzing user behavior logs scattered across various isolated systems, it constructs a user-centric, multi-dimensional behavioral trajectory and risk profile, enabling real-time early warning, proactive discovery, in-depth tracing, and efficient auditing of potential internal threats. This elevates the enterprise's security defenses from "passive response" to "proactive defense," and from post-incident investigation to in-process intervention and even pre-incident prediction. This application integrates real-time data collection, intelligent behavioral analysis, precise risk warning, and visualized traceability auditing, providing powerful technical tools for internal control and compliance departments, Security Operations Centers (SOCs), and IT administrators. It can accurately depict the digital behavior patterns of employees, quantify the security posture through a multi-dimensional risk weighting model, and intelligently capture and alert on abnormal behaviors deviating from the normal baseline. Ultimately, it breaks down data silos, achieves cross-system behavioral data correlation and fusion, improves threat discovery efficiency, transforming "finding a needle in a haystack" into "precision fishing"; and solidifies the audit process, providing a complete chain of evidence and decision support throughout the pre-event, during-event, and post-event stages.
[0025] See Figure 1 , Figure 1 This is a schematic flowchart illustrating a data leakage behavior detection and control method according to an embodiment of this application. This data leakage behavior detection and control method can be implemented using computer equipment, which can be deployed on a single server or a server cluster. It can also be deployed on handheld terminals, laptops, wearable devices, or robots, etc.
[0026] It should be noted that the acquisition of any information involved in the provided methods is in compliance with relevant regulations and is carried out with the user's consent, and will not infringe on the user's privacy or violate relevant laws and regulations.
[0027] This application adopts a microservice architecture, which ensures high cohesion and low coupling of the system. Each core component, including data acquisition, stream processing, rules engine, and API service, can be developed, deployed, extended, and upgraded independently.
[0028] Specifically, such as Figure 1 As shown, the data leakage behavior detection and control method provided in this embodiment may include the following steps: S101. Obtain raw user behavior data from VPN, bastion host, database, office OA, cloud platform, and terminal. The raw user data obtained includes abnormal behaviors of users obtaining data by operating documents, taking screenshots, or shooting videos. The abnormal behaviors include access to sensitive data, unauthorized operations, and login at abnormal times.
[0029] The acquisition of user's original behavior is based on Apache Kafka. Kafka was originally developed by LinkedIn and is a distributed, partitioned, multi-replica, multi-subscriber distributed log system (which can also be used as a message queue system) coordinated by ZooKeeper. It is commonly used for web / nginx logs, access logs, message services, etc. Its main application scenarios are log collection systems and message systems. Kafka is characterized by providing message persistence capabilities with a time complexity of O(1), ensuring constant-time access performance even for data of TB or more. High throughput. Even on very inexpensive commercial machines, it can support the transmission of 100K messages per second on a single machine, supports message partitioning between Kafka Servers and distributed consumption, and ensures the sequential transmission of messages within each partition. It also supports offline data processing and real-time data processing, and supports online horizontal scaling.
[0030] Kafka is used as a unified message queue, serving as the "log data bus." Various data sources (such as network devices, application systems, security products, and endpoints), specifically VPNs, bastion hosts, databases, office automation systems, cloud platforms, and mobile phones, push log data to the corresponding Kafka topics in real time via agents, Syslog, and APIs. This decouples data collection and processing. Data producers do not need to know who the downstream consumers are, and changes or expansions in the data processing workflow will not affect the normal reporting of data sources, greatly improving the system's flexibility and robustness. Acquiring various types of raw data from multiple platforms and terminals broadens the coverage of data leakage investigation and control, and makes analysis and control more accurate.
[0031] In addition, by providing standardized APIs (such as RESTful), supporting common data protocols (such as Syslog, JDBC, Kafka), and developing lightweight data collection agents, it can interface with the vast majority of existing business systems, network devices, and security products.
[0032] By transforming data source integration from a "hard-coded" development model to a "configuration-based" management operation, data processing gains powerful adaptability and flexibility. Simultaneously, by configuring data source connection information, data structure definitions, and parsing rules through the management interface, most of the code for data acquisition and processing tasks can be automatically generated, significantly reducing the cost and development cycle of integrating new data sources.
[0033] The acquired raw user data includes abnormal user behaviors such as accessing sensitive data, unauthorized operations, and logging in at unusual times by manipulating documents, taking screenshots, and recording videos. The implementation involves deeply integrating user behavior log analysis with multimodal content insights. Unstructured data objects such as images and documents involved in the user's actions are automatically processed in real time. OCR (Optical Character Recognition) technology is integrated to extract text information from images; Natural Language Processing (NLP) technology is combined to identify sensitive information and perform semantic analysis on the text content. This embodiment solves the problem of detecting sensitive information leaks through screenshots and screen captures by acquiring raw user behavior data, achieving a qualitative leap from "behavior monitoring" to "behavior + content monitoring," and filling a key blind spot in content security auditing.
[0034] S102, Preprocess the acquired raw user behavior data, including cleaning, parsing, enriching and normalizing the data before storing it; This embodiment utilizes a hybrid architecture of "Flink for Real-Time + Spark for Batch" for data processing. Apache Flink is a framework and distributed processing engine for stateful computation on unbounded and bounded data streams. Apache Flink is designed to run in all common cluster environments, performing computations at memory speeds and any scale. Compared to other stream processing frameworks (such as Spark Streaming), Apache Flink offers significant advantages in low latency, high throughput, exactly-once semantics, and robust state management. Its EventTime and ProcessingTime models perfectly handle scenarios where logs arrive out of order, ensuring the accuracy of risky computations. Its active community and rich ecosystem (such as Flink CDC) provide strong support for integrating complex data sources with minimal technical risk. Apache Spark is a fast and general-purpose computing engine designed specifically for large-scale data processing. Apache Spark has three main characteristics: First, its high-level API decouples developers from the cluster itself, allowing them to focus on the computations their applications perform. Second, Apache Spark is fast and supports interactive computation and complex algorithms. Finally, Apache Spark is a general-purpose engine that can be used to perform a wide variety of operations, including SQL queries, text processing, machine learning, and more.
[0035] This embodiment preprocesses the acquired raw user behavior data, including cleaning, parsing, enriching, and normalizing its storage. This involves consuming raw logs from Kafka using Apache Flink and performing a series of streaming ETL (Extract, Transform, Load) operations, including log parsing, field normalization, invalid data filtering, reverse IP address lookup, and information completion by associating with user / asset metadata (from CMDB or HR system). The cleaned and normalized data is then written back to a new Kafka topic for subsequent consumption.
[0036] S103: Based on preset risk rules, the preprocessed data is matched in real time using machine learning algorithms. A risk score is obtained based on the matching results. The presence of a violation risk is determined based on the risk score. A visual risk profile is generated based on the violation risk.
[0037] This embodiment utilizes Apache Flink to process real-time streaming data for real-time rule matching and second-level alerts; Apache Spark performs periodic batch processing of all historical data for complex computational tasks such as machine learning model training, deep correlation analysis, and periodic report generation. The combination of these two technologies balances low-latency real-time response with complex deep analysis capabilities, achieving an optimal solution for different computational needs, resulting in higher resource utilization efficiency and a more robust system architecture.
[0038] S104 outputs alarm information based on the risk of violation.
[0039] This embodiment establishes a standardized workflow from alarm triggering to handling completion. Based on the alarm triggering conditions, it outputs an alarm notification including notification method, alarm assignment, alarm release time, and time limit, and displays the handling results including handling opinions, evidence upload status, and status tracking, forming an effective closed-loop risk alarm. The handling result record can be used for subsequent auditing and analysis. Notification methods include various means such as email, SMS, and instant messaging tools to notify relevant personnel.
[0040] In some feasible embodiments, the step of performing real-time risk rule matching on the preprocessed data obtained according to preset risk rules using machine learning algorithms, obtaining a risk score based on the matching results, determining whether there is a violation risk based on the risk score, and generating a visual risk profile based on the violation risk includes: Obtain the risk scenario triggered by each user, and obtain the first risk score configured according to that risk scenario; Based on the first risk score, obtain the second risk score for the user after it has decayed over time according to the decay rule for this risk scenario; Based on the second risk score, a risk ranking list is generated and sorted according to the second risk score, and displayed in the form of a visual risk profile.
[0041] By analyzing preprocessed data from each user, and considering various factors such as assigning different weights to events and risk values that decay over time, a visualized risk profile is created for each user. By incorporating time decay functions and contextual weights, this approach abandons the traditional static, cumulative risk scoring mechanism and constructs a dynamic model that more closely resembles human risk perception. This significantly reduces "false alarms" caused by historical issues or one-time misoperations, making risk scoring more accurately reflect the current and immediate threat level. This greatly improves the scientific rigor and accuracy of risk assessment, which is then presented in the form of a visualized risk profile.
[0042] The step of obtaining each user's triggered risk scenario and obtaining the first risk score configured according to that risk scenario includes: Obtain the types of events triggered in this risk scenario and the corresponding weight scores assigned to each type; Based on the weighted scores, a comprehensive score is obtained for each risk scenario, calculated according to the chronological order of various triggering events and the number of triggers in that chronological order. The number of triggers includes the number of times each triggering event is triggered repeatedly in the chronological order or the number of times a single triggering event is triggered repeatedly. This comprehensive score is the first risk score. Different risk scenarios first consider the type of triggering event. Different types of risk scenarios have different severity levels, and the weight of their triggering events also differs. For example, the weight of "batch export of core database" is much higher than that of "logging into the OA system outside of working hours". In addition, a risk scenario may not be triggered by a single event, but may be a combination of multiple events in a specific order, chronological order, and number of triggers. The corresponding first risk score is the comprehensive score for that scenario. For example: "Within 5 minutes, the user first logged into the VPN, then immediately accessed the core financial system, and attempted to query a sensitive data table." This risk scenario mainly includes two risk points: logging into a VPN and accessing the core financial system and querying sensitive data. The severity levels of logging into the VPN and accessing the core financial system are different (different types of triggering events), and the risk scores are also different. For example, accessing the VPN is assigned 10 points, accessing the core financial system is assigned 20 points, and accessing both sequentially is assigned 30 points. Based on this, if the VPN and the core financial system are accessed once within one hour, the scenario will have a base risk score, such as 30 points + 10 points. The second trigger score is 30 + 25 points, and the third trigger score is 30 + 50 points. This score represents the first risk score, which is the total score in this scenario. Therefore, the first risk score is a comprehensive risk assessment score based on multiple factors such as scenario category or type (i.e., accessing one or more), number of accesses, etc., and according to the different weights assigned to each factor.
[0043] Different events in the scenario can come from completely different data sources. For example, one event might come from VPN logs (successful login), while the next event might come from database audit logs (query execution), enabling true correlation analysis.
[0044] In this embodiment, a comprehensive risk score is maintained for each user and updated in real time. This score is not a simple summation, but a comprehensive score derived according to the aforementioned assignment rules. All users are sorted according to their risk scores to form a "risk ranking list".
[0045] Based on the initial risk score, a second risk score is obtained for the user, after decaying over time according to the decay rule for that risk scenario. An "decay factor" is introduced, where the risk score of a historical event automatically decays exponentially over time. This avoids a user being permanently "listed" due to a single historical misoperation, making risk assessment more scientific and fair.
[0046] Based on the second risk score, a risk ranking list is generated and displayed in the form of a visualized risk profile. Unlike traditional log-list-based auditing methods, this embodiment aggregates multi-source data and applies risk models to construct a dynamic, quantitative, and visualized risk profile for each user (including entities such as servers and application accounts). This transforms abstract, massive behavioral data into a clear and concise risk landscape, enabling security analysts to quickly answer the three key questions: "Who is currently at the highest risk? Why? What did they do?" This achieves precise positioning and efficient response.
[0047] For example, the enterprise's internal risk ranking is presented in columns and ranked from top to bottom according to the second risk score. It also supports filtering by time, scenario, organization, and other defined scopes. Through multi-condition and multi-event complex logic, false alarms caused by single-point event alarms are greatly reduced, making each alarm more valuable for investigation.
[0048] In some embodiments, generating a risk ranking list sorted by the second risk score and displaying it in the form of a visual risk profile further includes: When generating a visual risk profile of a risk ranking sorted by the second risk score, a first-level personal risk details visualization profile is also generated, displaying each user's organizational structure information, risk summary, risk composition analysis, self-data breach risk, list of key risk events, names of triggered risk scenarios, data sources, and risk scores contributed by each event. The organizational structure information includes the user's name, department, position, and job title. The risk summary displays the current total risk score, risk ranking within the company, and a recent risk trend chart. The risk composition analysis shows the source composition of the user's risk score in pie chart or bar chart format. Self-data breach risk includes risks such as privilege abuse and unauthorized operations; for example, privilege abuse risk accounts for 30% of all risks, unauthorized operations risk accounts for 20%, and so on. This allows management and analysis departments to immediately and clearly identify the user's main risk types. The list of key events includes all high-risk events triggered by the user, arranged in reverse chronological order, including event time, names of triggered risk scenarios, data sources, risk scores contributed by each event, etc. For example,... Figure 2As shown, this is a visualization of a user's first-level personal risk details. The content displayed includes personal information such as employment status, start date, and job type, as well as information such as the time of risk occurrence, risk behavior, risk level, risk scenario, risk score, and handling status.
[0049] In some embodiments, the cascading personal risk details visualization profile for each user also includes a deep investigation button window on the first cascading personal risk details visualization profile page for each user. This window allows for cascading to a second cascading personal risk details visualization profile. The second cascading personal risk details visualization profile displays behavioral trajectories and alarm details. The behavioral trajectories include all operation logs of the user at a specific time, and the alarm details include alarm records generated by the user's risky behavior and their processing status. For example, alarm records include previously set alarm rules, such as outputting alarm notifications based on alarm triggering conditions, including notification method, alarm assignment, alarm release time, and time limit. The processing status includes processing opinions, evidence upload status, and processing results of status tracking. For example, filtering conditions such as time and operation type can be added or changed through the alarm editing box on the second cascading personal risk details visualization profile page for convenient advanced queries. This embodiment implements a standardized workflow from alarm triggering to handling completion, ensuring that each risk alarm is effectively closed-loop and forms a complete handling record for subsequent auditing and analysis.
[0050] The risk assessment method in this embodiment is a dynamic and quantifiable user risk assessment approach. It not only records individual events but also focuses on the sequence, context, and correlation of events. Through configured scoring rules, such as assigning different weights to different events, score decay over time, and adding points for multiple events, a comprehensive risk value is calculated for each user, forming a visualized risk profile. Security personnel can easily see the high-risk personnel ranking list and drill down to view their detailed risk composition, behavioral patterns, and related events.
[0051] In summary, this application's method, through multimodal data fusion analysis, pioneered the deep integration of log behavior analysis and image content recognition (OCR), constructing a dual detection dimension of "behavior + content," thus solving the data leakage blind spots that pure behavior analysis cannot reach. Simultaneously, through a dynamic scoring algorithm with a decay factor, the risk profile not only reflects historical behavior but also accurately depicts the current risk situation, significantly improving the accuracy of risk assessment. It transforms data access from a "hard-coded" development model to a "configurable" management operation, endowing this method with powerful evolutionary capabilities and flexibility. Data leakage investigation and control achieves fully online, closed-loop management, truly transforming analysis results into security actions.
[0052] The following are embodiments of the apparatus described in this application, which can be used to execute the embodiments of the method described in this application. For details not disclosed in the apparatus embodiments of this application, please refer to the embodiments of the method described in this application.
[0053] Please see Figure 3 This application illustrates a data leakage behavior detection and control device provided in an exemplary embodiment, comprising: The raw data acquisition unit 301 is used to acquire raw user behavior data from VPN, bastion host, database, office OA, cloud platform, and terminal. The acquired raw user data includes abnormal behaviors of users acquiring data by operating documents, taking screenshots, or shooting videos. The abnormal behaviors include access to sensitive data, unauthorized operations, and login at abnormal times. Data preprocessing unit 302 is used to preprocess the acquired raw user behavior data, including cleaning, parsing, enriching and normalizing the storage. The violation risk judgment unit 303 is used to perform real-time risk rule matching on the preprocessed data obtained according to preset risk rules through machine learning algorithms, obtain a risk score based on the matching results, determine whether there is a violation risk based on the risk score, and generate a visual risk profile based on the violation risk. Alarm unit 304 outputs alarm information according to the alarm rules.
[0054] Based on the above methods for investigating and controlling data breaches, such as Figure 4 As shown in the diagram, this embodiment of the invention also provides a structural schematic of a data leakage behavior detection and control device 4, which includes a processor 41 and a memory 42 coupled to the processor 41. The memory 42 stores a computer program, which, when executed by the processor 41, causes the processor 41 to execute the data leakage behavior detection and control method described in the above embodiment.
[0055] For other details regarding the implementation of the above technical solution by the processor 41 in the above data leakage behavior detection and control device, please refer to the description in the data leakage behavior detection and control method provided in the above invention embodiments, which will not be repeated here.
[0056] The processor 41 can also be called a CPU (Central Processing Unit). The processor 41 may be an integrated circuit chip with signal processing capabilities. The processor 41 may also be a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The general-purpose processor may be a microprocessor, or the processor 41 may be any conventional processor.
[0057] like Figure 5 As shown in the diagram, this embodiment of the invention also provides a schematic diagram of a computer-readable storage medium, on which a readable computer program 51 is stored. The computer program 51 can be stored in the storage medium in the form of a software product, including several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute all or part of the steps of the methods described in various embodiments of the invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, magnetic disks or optical disks, ROM (Read-Only Memory), RAM (Random Access Memory), or terminal devices such as computers, servers, mobile phones, and tablets.
[0058] In the several embodiments provided in this application, it should be understood that the disclosed apparatus, devices, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple modules or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces, indirect coupling or communication connection between devices or modules, and may be electrical, mechanical, or other forms.
[0059] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical modules; that is, they may be located in one place or distributed across multiple network modules. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0060] Furthermore, the functional modules in the various embodiments of this application can be integrated into one processing module, or each module can exist physically separately, or two or more modules can be integrated into one module. The integrated module can be implemented in hardware or as a software functional module. If the integrated module is implemented as a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage medium.
[0061] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product.
[0062] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a server or data center that integrates one or more available media. The available medium may be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state disk (SSD)).
[0063] The technical solutions provided in this application have been described in detail above. Specific examples have been used in this application to illustrate the principles and implementation methods of this application. The description of the above embodiments is only for the purpose of helping to understand the method and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.
[0064] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0065] This application is described with reference to flowchart illustrations and / or block diagrams of the methods, apparatus, and computer program products according to this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0066] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0067] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0068] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.
Claims
1. A method for investigating and controlling data leakage behavior, characterized in that: This includes obtaining raw user behavior data from VPNs, bastion hosts, databases, office automation systems, cloud platforms, and terminals. The raw user data obtained includes abnormal behaviors of users obtaining data by operating documents, taking screenshots, or recording videos. These abnormal behaviors include access to sensitive data, unauthorized operations, and login at abnormal times. Preprocessing of acquired raw user behavior data includes cleaning, parsing, enriching, and normalizing storage; Based on preset risk rules, the preprocessed data is matched in real time using machine learning algorithms. A risk score is obtained based on the matching results. The risk score is used to determine whether there is a risk of violation. A visual risk profile is generated based on the risk of violation. Output alarm information according to the alarm rules.
2. The data leakage behavior investigation and control method according to claim 1, characterized in that: The preprocessing of the acquired raw user behavior data includes cleaning, parsing, enriching and normalizing the storage. It also includes cleaning, parsing, enriching and normalizing the acquired raw user behavior data by means of log parsing, field standardization, invalid data filtering, reverse IP address resolution and information completion by associating user and asset metadata.
3. The data leakage behavior investigation and control method according to claim 1, characterized in that: The process involves real-time risk rule matching of the preprocessed data obtained according to preset risk rules using machine learning algorithms, obtaining a risk score based on the matching results, determining whether there is a risk of violation based on the risk score, and generating a visual risk profile based on the risk of violation. This includes: Obtain the risk scenario triggered by each user, and obtain the first risk score configured according to that risk scenario; Based on the first risk score, obtain the second risk score for the user after it has decayed over time according to the decay rule for this risk scenario; Based on the second risk score, a risk ranking list is generated and sorted according to the second risk score, and displayed in the form of a visual risk profile.
4. The data leakage behavior investigation and control method according to claim 3, characterized in that: The step of obtaining each user's triggered risk scenario and obtaining the first risk score configured according to that risk scenario includes: The step of obtaining each user's triggered risk scenario and obtaining the first risk score configured according to that risk scenario includes: Obtain the types of events triggered in this risk scenario and the corresponding weight scores assigned to each type; Based on the weighted scores, a comprehensive score is obtained for the risk scenario, based on the chronological order of various triggering events and the number of triggers in that chronological order. The number of triggers includes the number of times the events are triggered repeatedly in the chronological order of various triggering events or the number of times a single triggering event is triggered repeatedly. This comprehensive score is the first risk score.
5. The data leakage behavior detection and control method according to claim 4, characterized in that: The process of generating a risk ranking list based on the second risk score and displaying it in the form of a visual risk profile includes: When generating a visual risk profile of a risk ranking sorted by the second risk score, a first-level personal risk details visual profile is also generated, displaying each user's organizational structure information, risk summary, risk composition analysis, self-data breach risk, list of key risk events, names of triggered risk scenarios, data sources, and risk scores contributed by each event. The organizational structure information includes the user's name, department, position, and job title. The risk summary displays the current total risk score, risk ranking within the company, and a recent risk trend chart. The risk composition analysis shows the source composition of the user's risk score in pie chart or bar chart format. The self-data breach risk includes risks of privilege abuse and unauthorized operations. The list of key risk events includes all high-risk events triggered by the user, listed in reverse chronological order, including event time, names of triggered risk scenarios, data sources, and risk scores contributed by each event.
6. The data leakage behavior investigation and control method according to claim 5, characterized in that: The first-level personal risk details visualization profile for each user also includes a deep investigation button window on the first-level personal risk details visualization profile page for each user. Through this window, a second-level personal risk details visualization profile can be cascaded. The content of the second-level personal risk details visualization profile includes displaying behavioral trajectory and alarm details. The behavioral trajectory includes all operation logs of the user at a specific time. The alarm details include alarm records generated by the user's risky behavior and their processing status.
7. The data leakage behavior investigation and control method according to claim 1, characterized in that: The alarm information is output according to the alarm rules, including outputting alarm notifications based on alarm triggering conditions, including notification method, alarm assignment, alarm release time, and time limit, and displaying processing results including processing opinions, evidence upload status, and status tracking status.
8. A data leakage behavior detection and control device, characterized in that, include: The raw data acquisition unit is used to acquire raw user behavior data from VPN, bastion host, database, office OA, cloud platform, and terminal. The acquired raw user data includes abnormal behaviors of users acquiring data by operating documents, taking screenshots, and shooting videos. The abnormal behaviors include access to sensitive data, unauthorized operations, and login at abnormal times. The data preprocessing unit is used to preprocess the acquired raw user behavior data, including cleaning, parsing, enriching and normalizing the data before storage. The violation risk assessment unit is used to perform real-time risk rule matching on the preprocessed data obtained according to preset risk rules using machine learning algorithms, obtain a risk score based on the matching results, determine whether there is a violation risk based on the risk score, and generate a visual risk profile based on the violation risk. The alarm unit determines the alarm information to output based on the alarm rules.
9. A computer device, characterized in that, include: A processor and a memory; wherein the memory stores a computer program adapted to be loaded by the processor and to execute the steps of the method as described in any one of claims 1 to 7.
10. A computer storage medium, characterized in that, The computer storage medium stores a plurality of instructions adapted for loading by a processor and executing the steps of the method as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Data theft risk analyzing method and analyzing system
CN106469274A
Method for implementing scalper risk control of internet medical care
CN107147621A
Data security risk monitoring method and device, electronic equipment and storage medium
CN118157880A
Effective network security event monitoring method and system based on big data model
CN119167358A
Network data leakage monitoring system and method
CN119561794A