Data processing method and device
Through data processing methods and dynamic score cards, high-risk data are automatically identified, solving the problem of huge DLP alarm data and low manual processing efficiency, and achieving efficient and accurate risk identification and processing.
Patent Information
- Application Number
- CN202210265154.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-17
- Publication Date
- 2025-08-19
- Estimated Expiration
- 2042-03-17
AI Technical Summary
In the prior art, the alarm data generated by the DLP system is of huge magnitude, the manual processing efficiency is low and the cost is high, and it is unable to respond to real risks in a timely manner.
Using data processing methods, by obtaining the data scenario of initial sensitive data, the corresponding data processing algorithm and target scoring units are determined. The dynamic scoring card covers risk identification factors, and automatically scores to identify high-risk data and reduce labor costs.
It realizes rapid and accurate processing of DLP alarm data, reduces response time, improves risk and hemostasis efficiency, and saves labor costs.
Smart Images

Figure CN114611149B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of this specification relate to the field of data processing technology, and in particular to a data processing method. Background Art
[0002] With the implementation of data protection laws, individuals and organizations are increasingly aware of the importance of protecting sensitive data. Organizations typically invest heavily in DLP (Data Leak Prevention) systems to monitor sensitive data. However, the volume of daily DLP alerts generated is enormous. Manually processing this data is labor-intensive and unable to promptly respond to real risks, resulting in significant overall cost savings. Summary of the Invention
[0003] In view of this, embodiments of this specification provide a data processing method. One or more embodiments of this specification also relate to a data processing apparatus, a computing device, a computer-readable storage medium, and a computer program to address technical deficiencies in the prior art.
[0004] According to a first aspect of an embodiment of this specification, there is provided a data processing method, including:
[0005] Obtaining a sensitive data set including initial sensitive data, and determining a data scenario corresponding to the initial sensitive data;
[0006] Determining a data processing algorithm corresponding to the initial sensitive data and at least two target scoring units according to the data scenario;
[0007] When it is determined according to the data processing algorithm that the initial sensitive data meets the preset processing conditions, the initial sensitive data is subjected to sensitivity scoring according to the at least two target scoring units to obtain a target scoring result.
[0008] According to a second aspect of the embodiments of this specification, there is provided a data processing device, including:
[0009] A data acquisition module is configured to acquire a sensitive data set including initial sensitive data and determine a data scenario corresponding to the initial sensitive data;
[0010] an algorithm determination module, configured to determine a data processing algorithm corresponding to the initial sensitive data and at least two target scoring units according to the data scenario;
[0011] The scoring module is configured to, when it is determined according to the data processing algorithm that the initial sensitive data meets the preset processing conditions, perform sensitivity scoring on the initial sensitive data according to the at least two target scoring units to obtain a target scoring result.
[0012] According to a third aspect of an embodiment of this specification, a computing device is provided, including:
[0013] memory and processor;
[0014] The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the above-mentioned data processing method are implemented.
[0015] According to a fourth aspect of the embodiments of this specification, a computer-readable storage medium is provided, which stores computer-executable instructions, and when the instructions are executed by a processor, the steps of the above-mentioned data processing method are implemented.
[0016] According to a fifth aspect of the embodiments of this specification, a computer program is provided, wherein when the computer program is executed in a computer, the computer is caused to execute the steps of the above-mentioned data processing method.
[0017] One embodiment of the present specification implements a data processing method and device, wherein the method includes obtaining a sensitive data set including initial sensitive data, and determining a data scenario corresponding to the initial sensitive data; determining a data processing algorithm corresponding to the initial sensitive data and at least two target scoring units based on the data scenario; and when it is determined that the initial sensitive data meets preset processing conditions based on the data processing algorithm, performing a sensitivity score on the initial sensitive data based on the at least two target scoring units to obtain a target scoring result.
[0018] Specifically, the data processing method adopts a data processing algorithm and a target scoring unit to implement a sensitivity score for each initial sensitive data. Subsequently, the target sensitive data can be quickly and accurately determined based on the target scoring result determined by the sensitivity score, so that the target sensitive data can be adjusted subsequently. This automated method can greatly save labor costs on the basis of improving efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 This is a schematic diagram of a specific application scenario of a data processing method provided by an embodiment of this specification;
[0020] Figure 2 is a flow chart of a data processing method provided by one embodiment of this specification;
[0021] Figure 3 This is a flowchart of a data processing method provided by one embodiment of this specification;
[0022] Figure 4This is a schematic diagram of the structure of a data processing device provided by one embodiment of this specification;
[0023] Figure 5 This is a structural block diagram of a computing device provided by one embodiment of this specification. DETAILED DESCRIPTION
[0024] The following description sets forth many specific details to facilitate a thorough understanding of this specification. However, this specification can be implemented in many other ways than those described herein, and those skilled in the art can make similar generalizations without violating the scope of this specification. Therefore, this specification is not limited to the specific implementations disclosed below.
[0025] The terms used in one or more embodiments of this specification are for the purpose of describing specific embodiments only and are not intended to limit one or more embodiments of this specification. The singular forms "a," "the," and "the" used in one or more embodiments of this specification and the appended claims are also intended to include plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more associated listed items.
[0026] It should be understood that although the terms first, second, etc. may be used to describe various information in one or more embodiments of this specification, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of one or more embodiments of this specification, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".
[0027] First, the terms involved in one or more embodiments of this specification are explained.
[0028] Sensitive data: refers to personal privacy data, such as ID number, mobile phone number, bank card number, etc.
[0029] DLP: Data loss prevention system.
[0030] SFTP: Secure File Transfer Protocol.
[0031] Each organization will assess whether sensitive data is public, internal, confidential, secret, or top secret based on the potential harm it would cause if leaked, and the value it brings to the organization. Through a data classification and grading program, sensitive data is classified and categorized across hosts, terminals, applications, and other environments. The categorized data will be organized into data scenarios, covering a wide range of data leakage strategies. The embodiments of this specification address the technical issues surrounding the massive volume of alert data generated by these data leakage strategies, the low efficiency, and high cost of manual processing.
[0032] Based on this, a data processing method is provided in this specification. This specification also involves a data processing device, a computing device, and a computer-readable storage medium, which are described in detail one by one in the following embodiments.
[0033] See also Figure 1 , Figure 1 A schematic diagram of a specific application scenario of a data processing method provided according to an embodiment of this specification is shown.
[0034] The data processing method provided in the embodiment of this specification is applied to Figure 1 In the data processing system, the data processing system includes a data storage unit 102, an algorithm unit 104, a scoring card unit 106, a data analysis unit 108, and a data processing unit 110.
[0035] The data storage unit 102 stores multiple DLP warning data generated by DLP, such as sensitive chat records, file transfers, and system calls;
[0036] The algorithm unit 104 stores multiple algorithms, such as semantic rule algorithms, model algorithms, and business logic algorithms. During specific calculations, the data processing system will pre-select an algorithm based on the data scenario corresponding to each DLP alert data. The selected algorithm will be used to analyze each DLP alert data, further screening each DLP alert data to determine whether it is high-risk data. If so, it will be sent to the scoring card unit 106. If not, it will not be processed further.
[0037] The scoring card unit 106 stores multiple scoring cards, such as personnel scoring cards, time scoring cards, file scoring cards, application scoring cards, data scoring cards, address scoring cards, equipment scoring cards, and stakeholder scoring cards. Each scoring card has a corresponding scoring range (e.g., 1-10 points, 0-5 points, etc.), scoring dimensions (e.g., job responsibilities, resignation intention, resignation status), and scoring strategies (e.g., calculation factors are addition or multiplication, etc.). Subsequently, the sensitivity score of each DLP alert data can be calculated based on the scoring card corresponding to it to obtain the target scoring result for each DLP alert data.
[0038] The data analysis unit 108 can dynamically update the scoring result of each DLP warning data according to the target scoring result, and judge whether each DLP warning data is target sensitive data in combination with the preset processing rules, thereby determining which DLP warning data requires subsequent processing. The preset processing rules can be set according to actual applications and are not limited in this embodiment of the specification.
[0039] The data processing unit 110 will perform post-response processing on the initial sensitive data (ie, target sensitive data) that needs to be processed as determined by the data analysis unit 108, such as issuing a warning message.
[0040] The data processing method provided in the embodiments of this specification utilizes the risk identification factors covered by dynamic scorecards to comprehensively and accurately process each DLP alert data. Furthermore, the calculation factors of each scorecard can be dynamically adjusted based on the data scenario. For example, in scenarios involving device sharing and misappropriation, the calculation factor weight of the device scorecard can be automatically increased. Furthermore, the risk identification factors within each scorecard can also be dynamically adjusted. The following uses three scorecards as examples for illustration:
[0041] Personnel Scorecard: Employees are the root cause of data breaches. Risk factors such as employee job responsibilities, willingness to resign, and resignation status are weighted. The final score is dynamically adjusted based on the employee's real-time resignation application and job transfer information.
[0042] File Scorecard: Files are important carriers of sensitive data. We assign weighted scores based on risk factors such as file size, content, and encrypted and compressed records. The weights and scores are analyzed based on the file's historical operations.
[0043] Stakeholder Scorecard: Stakeholders can be employees' relatives, colleagues, supervisors, friends, etc. To prevent collusion in data theft, the behavior of stakeholders is an important factor in assessing this risk.
[0044] In the examples in this manual, the sensitivity score assigned by the scoring card for each DLP alert is incorporated into the risk scoring formula to calculate the final risk score. The risk score also calculates risk acceleration and deceleration, thereby predicting employee risk fluctuations. The final score influences the final disposition decision, such as terminating the disposition process, contacting a supervisor for investigation, or automatically issuing a blocking policy. This significantly reduces labor costs, shortens response times, and improves risk mitigation efficiency.
[0045] See also Figure 2 , Figure 2 A flow chart of a data processing method provided according to an embodiment of the present specification is shown, which specifically includes the following steps.
[0046] Step 202: Acquire a sensitive data set including initial sensitive data, and determine a data scenario corresponding to the initial sensitive data.
[0047] The data processing method provided in the embodiments of this specification is applied to data leakage scenarios, and the data leakage scenarios within each organization can occur in multiple locations and environments such as office terminal computers, application systems, data warehouses, ecological organizations, host servers, etc. The data leakage scenarios in the above environments have been sorted out and clarified in the initial stage. The organization will purchase basic monitoring software such as DLP data leakage protection software, network DLP, application system operation logs, host HIDS, etc. Based on the original log data produced by these basic monitoring software, these original log data can be understood as the initial sensitive data of the embodiments of this specification. The specific implementation method is as follows:
[0048] In a specific implementation, obtaining a sensitive data set including initial sensitive data includes:
[0049] Obtain a sensitive data set containing initial sensitive data determined by a preset risk identification terminal.
[0050] Among them, the preset risk identification terminals include but are not limited to DLP data leakage protection software, network DLP, application system operation logs, host HIDS and other basic monitoring software.
[0051] In actual applications, the data scenarios corresponding to each initial sensitive data have been determined in the initial stage. As mentioned above, they include data leakage scenarios of office terminal computers, data leakage scenarios of application systems, data leakage scenarios of data warehouses, data leakage scenarios of ecological organizations, data leakage scenarios of host servers, etc.
[0052] Step 204: Determine a data processing algorithm corresponding to the initial sensitive data and at least two target scoring units according to the data scenario.
[0053] In practical applications, data processing algorithms include but are not limited to semantic rule algorithms, machine model algorithms, and business logic algorithms. Semantic rule methods include keyword, key word, and regular rule recognition capabilities. Machine model algorithms include algorithms such as relationship graph analysis, frequency increase analysis, and employee resignation model analysis. These algorithms can be used to implement strategies involving relationship, frequency, and deviation. Business logic algorithms are primarily based on internal organizational business attributes. For example, if a query on a personal account by a department that issues corporate loans is considered abnormal due to financial attributes, this is a strategy formulated using business logic. Target scoring units include but are not limited to personnel scoring units, time scoring units, document scoring units, application scoring units, data scoring units, address scoring units, device scoring units, and person-to-person scoring units.
[0054] Specifically, according to the different data scenarios corresponding to each initial sensitive data, the data processing algorithm and target scoring unit used in the subsequent processing of the initial sensitive data are also different.
[0055] Because the attributes of the data scenarios corresponding to each initial sensitive data are different, the data processing algorithms and target scoring units that each initial sensitive data will be analyzed through are determined from the initial stage; for example, if the data scenario corresponding to the initial sensitive data is a data leakage scenario of an office terminal computer (such as the risk of data leakage due to device misappropriation), this data scenario is strongly related to the device, and the focus is on identifying keywords and keywords based on semantic rule algorithms and calling the device scoring unit for analysis; if the data scenario corresponding to the initial sensitive data is a data leakage scenario of the application system (such as the risk of data leakage due to application violation queries), this data scenario is strongly related to the relationship network of the queried customer, and the focus is on analyzing the relationship network of the relationship person based on the machine algorithm model, and calling the application scoring unit for analysis.
[0056] Among them, each scoring unit corresponds to a scoring interval, for example, the scoring interval corresponding to the personnel scoring unit is 1-10 points, the scoring interval corresponding to the time scoring unit is 0-5 points, the scoring interval corresponding to the file scoring unit is 0-100 points, the scoring interval corresponding to the application scoring unit is 0-100 points, the scoring interval corresponding to the data scoring unit is 0-100 points, the scoring interval corresponding to the address scoring unit is 0-10 points, the scoring interval corresponding to the device scoring unit is 0-10 points, the scoring interval corresponding to the related person scoring unit is 0-10 points, etc.
[0057] The weight of the calculation factor of each scoring unit can be dynamically adjusted according to different data scenarios. For example, in the scenario of device sharing and misappropriation, the weight of the calculation factor of the device scoring unit will be automatically increased. The specific implementation method is as follows:
[0058] The determining, according to the data scenario, a data processing algorithm corresponding to the initial sensitive data and at least two target scoring units includes:
[0059] Determining a data processing algorithm corresponding to the initial sensitive data and at least two initial scoring units according to the data scenario;
[0060] Calculating the relevance of each initial scoring unit with the data scenario;
[0061] The scoring strategy of each initial scoring unit is adjusted according to the correlation degree, and each adjusted initial scoring unit is determined as at least two target scoring units corresponding to the initial sensitive data.
[0062] The scoring strategy can be understood as the final scoring result of the initial scoring unit being addition, or multiple growth, etc.
[0063] Specifically, adjusting the scoring strategy of each initial scoring unit according to the correlation degree can be understood as adjusting the weight value of the calculation factor of each initial scoring unit according to the correlation degree, and determining the scoring strategy of each initial scoring unit according to the weight value. For example, after the weight value is adjusted, the final scoring result of the initial scoring unit can be multiplied, etc., so that the score calculated by the initial scoring unit is more prominent.
[0064] In actual applications, each data scenario corresponds to two or more initial scoring units. In the case of multiple initial scoring units, the correlation between each initial scoring unit and the data scenario can be calculated. The weight value of the calculation factor of each initial scoring unit can be adjusted according to the correlation, and its scoring strategy can be adjusted according to the weight value to obtain the target scoring unit corresponding to the initial sensitive data. The specific calculation method of the correlation between the initial scoring unit and the data scenario can adopt any method, such as calculation by keyword similarity, etc., and the embodiments of this specification do not impose any restrictions.
[0065] For example, when the data scenario is a device sharing and misappropriation scenario, the target scoring unit corresponding to the data scenario includes a device scoring unit and a related person scoring unit, etc. The similarity of the keywords can determine that the device scoring unit has a high correlation with the data scenario, and the weight value of the calculation factor of the device scoring unit can be dynamically increased.
[0066] Step 206: When it is determined according to the data processing algorithm that the initial sensitive data meets the preset processing conditions, the initial sensitive data is subjected to a sensitivity score according to the at least two target scoring units to obtain a target scoring result.
[0067] The role of the data processing algorithm in the embodiments of this specification can be understood as further screening of initial sensitive data. That is, before assigning a sensitivity score to a piece of initial sensitive data, further sensitivity processing is performed on that data using the corresponding data processing algorithm to determine whether to proceed to the next sensitivity score. Specifically, only when the corresponding data processing algorithm determines that the initial sensitive data is at high risk of data leakage will it be assigned a sensitivity score using the corresponding target scoring algorithm.
[0068] Specifically, determining that the initial sensitive data meets the preset processing conditions according to the data processing algorithm can be understood as determining that the initial sensitive data meets the preset processing conditions when the initial sensitive data is determined to be high-risk leakage data according to the data processing algorithm.
[0069] In actual applications, data processing algorithms include but are not limited to semantic rule algorithms, machine model algorithms, business logic algorithms, etc.; among them, semantic rule algorithms include keyword, key word, and regular rule recognition capabilities, and can determine whether the initial sensitive data is high-risk leakage data based on its above-mentioned recognition capabilities; machine model algorithms include relationship graph analysis, frequency increase analysis, resignation model analysis and other algorithms. When the initial sensitive data involves relationship class, frequency class, and deviation class strategies, this method can be used for analysis to determine whether the initial sensitive data is high-risk leakage data; business logic algorithms are based on the business attributes within the organization. For example, financial-related attributes make it abnormal for a department that issues loans to an enterprise to query personal accounts. This is a strategy formulated by business logic, and this algorithm is used to determine whether the initial sensitive data is high-risk leakage data.
[0070] In a specific implementation, when it is determined according to the data processing algorithm that the initial sensitive data meets the preset processing conditions, the initial sensitive data is subjected to sensitivity scoring according to the at least two target scoring units to obtain a target scoring result. The specific implementation method is as follows:
[0071] The performing sensitivity scoring on the initial sensitive data according to the at least two target scoring units includes:
[0072] Perform a sensitivity score on the initial sensitive data according to the scoring dimension and scoring strategy corresponding to each target scoring unit in the at least two target scoring units.
[0073] Among them, the scoring dimensions corresponding to each target scoring unit are different. Taking the target scoring unit as a personnel scoring unit as an example, the scoring dimensions corresponding to the target scoring unit include but are not limited to the personnel's job responsibilities, willingness to resign, resignation status, etc.; taking the target scoring unit as a file scoring unit as an example, the scoring dimensions corresponding to the target scoring unit include but are not limited to file size, file content, encrypted compressed records, etc.; taking the target scoring unit as a relationship scoring unit as an example, the scoring dimensions corresponding to the relationship scoring unit include but are not limited to employees' relatives, colleagues, supervisors, friends, etc.
[0074] In practical applications, each target scoring unit corresponds to a scoring interval, and the scoring dimensions corresponding to each target scoring unit are also matched to corresponding scores based on this scoring interval. In specific applications, each target scoring unit corresponding to each initial sensitive data will assign a sensitivity score to each initial sensitive data based on its corresponding scoring dimension and scoring strategy. This allows for subsequent sensitivity scoring of each initial sensitive data based on its corresponding target scoring unit, resulting in an accurate target scoring result.
[0075] In specific implementation, since each initial sensitive data will correspond to two or more target scoring units, after obtaining the sensitivity score of each target scoring unit, the target scoring result of each initial sensitive data must be accurately calculated through addition, multiplication or other calculation methods. The specific implementation method is as follows:
[0076] The performing sensitivity scoring on the initial sensitive data according to the at least two target scoring units to obtain a target scoring result includes:
[0077] Performing a sensitivity score on the initial sensitive data according to each of the at least two target scoring units, to obtain an initial scoring result of each target scoring unit for the initial sensitive data;
[0078] The initial scoring result of the initial sensitive data is calculated according to a preset calculation rule to obtain a target scoring result of the initial sensitive data.
[0079] The preset calculation rules include but are not limited to addition, multiplication or other calculation rules.
[0080] Specifically, taking the preset calculation rule of addition as an example, when there are two or more target scoring units corresponding to a certain initial sensitive data, the initial sensitive data is subjected to a sensitivity score according to each adjusted target scoring unit to obtain the initial scoring result of each target scoring unit for the initial sensitive data; and then the initial scoring results of each target scoring unit for the initial sensitive data are added together to obtain the target scoring result of the initial sensitive data.
[0081] In actual applications, after obtaining the target scoring result for each initial sensitive data, it is possible to determine whether the initial sensitive data is the target sensitive data based on the target scoring result. Subsequently, specific business processing can be performed on the target sensitive data to ensure the security of the overall business. The specific implementation method is as follows:
[0082] After obtaining the target scoring result, the method further includes:
[0083] Determine whether the initial sensitive data is target sensitive data based on the target scoring result.
[0084] In specific implementation, there are at least three ways to determine whether the initial sensitive data is the target sensitive data based on the target scoring result. The following takes these three as examples to specifically introduce each way of determining whether the initial sensitive data is the target sensitive data based on the target scoring result.
[0085] The first method is to determine the average score result based on the target score result, and use the average score result to determine whether the initial sensitive data is the target sensitive data. The specific implementation method is as follows:
[0086] The determining whether the initial sensitive data is target sensitive data according to the target scoring result includes:
[0087] Determining target scoring results for all initial sensitive data in the sensitive data set;
[0088] Determining an average scoring result based on target scoring results of all initial sensitive data in the sensitive data set;
[0089] Determine whether the initial sensitive data is target sensitive data based on the average scoring result.
[0090] During specific implementation, the target scoring results of all initial sensitive data in the sensitive data set are obtained, and then the average scoring result is calculated based on the target scoring results of all initial sensitive data in the sensitive data set; the target scoring result of each initial sensitive data is then compared with the average scoring result to determine whether the initial sensitive data is the target sensitive data.
[0091] For example, if the target score is greater than the average score, the initial sensitive data can be considered the target sensitive data; if the target score is less than or equal to the average score, the initial sensitive data can be considered not the target sensitive data. This method can quickly determine whether each initial sensitive data is the target sensitive data.
[0092] The second method is to obtain the target scoring results of all the initial sensitive data in the sensitive data set, and determine whether it is the target sensitive data based on the ranking results of the target scoring results of all the initial sensitive data. The specific implementation method is as follows:
[0093] The determining whether the initial sensitive data is target sensitive data according to the target scoring result includes:
[0094] Determining target scoring results for all initial sensitive data in the sensitive data set;
[0095] sorting all the initial sensitive data in the sensitive data set according to the target scoring results of all the initial sensitive data in the sensitive data set;
[0096] Determine whether the initial sensitive data is target sensitive data based on the sorting result.
[0097] For example, the target scoring results of all initial sensitive data in the sensitive data set are obtained, and then all initial sensitive data in the sensitive data set are sorted in descending order based on the target scoring results. Then, the top 30 or top 50 initial sensitive data are selected from the sorting results as the target sensitive data. This method can also quickly obtain the target sensitive data.
[0098] The third method is to determine whether the initial sensitive data is the target sensitive data based on the historical scoring results. The specific implementation method is as follows:
[0099] The determining whether the initial sensitive data is target sensitive data according to the target scoring result includes:
[0100] In the case where it is determined that there is a historical scoring result for the initial sensitive data, whether the initial sensitive data is the target sensitive data is determined according to the historical scoring result of the initial sensitive data and the target scoring result.
[0101] Specifically, after obtaining the target scoring result of the initial sensitive data, if it is determined that the initial sensitive data has a historical scoring result, it can be determined whether each initial sensitive data is the target sensitive data based on the historical scoring result and the target scoring result of each initial sensitive data.
[0102] One possible implementation method is that determining whether the initial sensitive data is target sensitive data based on the historical scoring result of the initial sensitive data and the target scoring result includes:
[0103] Calculating a difference between the historical scoring result and the target scoring result based on the historical scoring result of the initial sensitive data and the target scoring result;
[0104] If the difference is greater than a preset difference threshold, determining that the initial sensitive data is target sensitive data; and
[0105] When the difference is less than or equal to the preset difference threshold, it is determined that the initial sensitive data is not the target sensitive data.
[0106] The preset difference threshold value can be set according to actual application and is not limited in this specification, for example, 60 or 70.
[0107] During specific implementation, the historical scoring results of each initial sensitive data (such as the previous scoring result of the current scoring result) and the target scoring result are obtained; the difference between the historical scoring results of each initial sensitive data and the target scoring result is obtained, and whether each initial sensitive data is the target sensitive data is determined based on the difference.
[0108] For example, if the preset difference threshold is 70 points, and the difference between the historical scoring result of the initial sensitive data and the target scoring result is 90 points, it can be determined that the difference is greater than the preset difference threshold, and the initial sensitive data can be determined to be the target sensitive data; if the difference between the historical scoring result of the initial sensitive data and the target scoring result is 50 points, it can be determined that the difference is less than the preset difference threshold, and the initial sensitive data can be determined not to be the target sensitive data.
[0109] Another possible implementation method is to obtain the previous historical scoring result of each initial sensitive data, and the previous historical scoring result of the previous historical scoring result; then calculate the difference between it and the previous historical scoring result, as well as the difference between its previous historical scoring result and the previous historical scoring result of the previous historical scoring result, and compare the two differences. If the value after subtracting the two differences is greater than a preset threshold, the initial sensitive data can also be determined to be the target sensitive data.
[0110] The data processing method provided in the embodiments of this specification adopts a data processing algorithm and a target scoring unit to achieve sensitivity scoring for each sensitive dimension of each initial sensitive data. Finally, the target sensitive data can be quickly and accurately determined based on the target scoring result determined by the sensitivity score, so that the target sensitive data can be adjusted subsequently. This automated method saves a lot of labor costs on the basis of improving efficiency.
[0111] The following combined Figure 3 , taking the data processing method provided in this specification to process DLP alarm data as an example, the data processing method is further explained. Figure 3 A flowchart of a data processing method provided in one embodiment of this specification is shown, which specifically includes the following steps.
[0112] Step 302: Obtain DLP alarm data.
[0113] Step 304: Determine the data scenario of each piece of DLP alarm data, and determine the data processing algorithm and at least two score cards corresponding to each piece of DLP alarm data based on the data scenario of each piece of DLP alarm data.
[0114] Among them, the scoring card can be understood as the above-mentioned scoring unit.
[0115] Step 306: When it is determined according to the data processing algorithm corresponding to each piece of DLP warning data that the piece of DLP warning data is high-risk leakage data, the piece of DLP warning data is sent to the corresponding at least two target scoring units.
[0116] Step 308: Perform a sensitivity score on each piece of DLP alert data based on at least two score cards corresponding to each piece of DLP alert data to obtain a target score result for each piece of DLP alert data.
[0117] Step 310: Determine whether each piece of DLP warning data is target sensitive data based on the target scoring result of each piece of DLP warning data.
[0118] The data processing method provided in the embodiments of this specification uses a dynamic scoring card mechanism to comprehensively and accurately obtain the target scoring results for each DLP alarm data. Subsequently, in actual application scenarios, the DLP alarm data that needs to be post-processed can be obtained based on the target scoring results to ensure the security of the system.
[0119] Corresponding to the above method embodiment, this specification also provides a data processing device embodiment, Figure 4 FIG1 shows a schematic diagram of the structure of a data processing device provided by an embodiment of this specification. Figure 4 As shown, the device includes:
[0120] The data acquisition module 402 is configured to acquire a sensitive data set including initial sensitive data and determine a data scenario corresponding to the initial sensitive data;
[0121] An algorithm determination module 404 is configured to determine a data processing algorithm corresponding to the initial sensitive data and at least two target scoring units according to the data scenario;
[0122] The scoring module 406 is configured to, when it is determined according to the data processing algorithm that the initial sensitive data meets the preset processing conditions, perform sensitivity scoring on the initial sensitive data according to the at least two target scoring units to obtain a target scoring result.
[0123] Optionally, the data acquisition module 402 is further configured to:
[0124] Obtain a sensitive data set containing initial sensitive data determined by a preset risk identification terminal.
[0125] Optionally, the algorithm determination module 404 is further configured to:
[0126] Determining a data processing algorithm corresponding to the initial sensitive data and at least two initial scoring units according to the data scenario;
[0127] Calculating the relevance of each initial scoring unit with the data scenario;
[0128] The scoring strategy of each initial scoring unit is adjusted according to the correlation degree, and each adjusted initial scoring unit is determined as at least two target scoring units corresponding to the initial sensitive data.
[0129] Optionally, the scoring module 406 is further configured to:
[0130] Perform a sensitivity score on the initial sensitive data according to the scoring dimension and scoring strategy corresponding to each target scoring unit in the at least two target scoring units.
[0131] Optionally, the scoring module 406 is further configured to:
[0132] Performing a sensitivity score on the initial sensitive data according to each of the at least two target scoring units, to obtain an initial scoring result of each target scoring unit for the initial sensitive data;
[0133] The initial scoring result of the initial sensitive data is calculated according to a preset calculation rule to obtain a target scoring result of the initial sensitive data.
[0134] Optionally, the device further includes:
[0135] The target data determination module is configured to:
[0136] Determine whether the initial sensitive data is target sensitive data based on the target scoring result.
[0137] Optionally, the target data determination module is further configured to:
[0138] Determining target scoring results for all initial sensitive data in the sensitive data set;
[0139] Determining an average scoring result based on target scoring results of all initial sensitive data in the sensitive data set;
[0140] Determine whether the initial sensitive data is target sensitive data based on the average scoring result.
[0141] Optionally, the target data determination module is further configured to:
[0142] Determining target scoring results for all initial sensitive data in the sensitive data set;
[0143] sorting all the initial sensitive data in the sensitive data set according to the target scoring results of all the initial sensitive data in the sensitive data set;
[0144] Determine whether the initial sensitive data is target sensitive data based on the sorting result.
[0145] Optionally, the target data determination module is further configured to:
[0146] In the case where it is determined that there is a historical scoring result for the initial sensitive data, whether the initial sensitive data is the target sensitive data is determined according to the historical scoring result of the initial sensitive data and the target scoring result.
[0147] Optionally, the target data determination module is further configured to:
[0148] Calculating a difference between the historical scoring result and the target scoring result based on the historical scoring result of the initial sensitive data and the target scoring result;
[0149] If the difference is greater than a preset difference threshold, determining that the initial sensitive data is target sensitive data; and
[0150] When the difference is less than or equal to the preset difference threshold, it is determined that the initial sensitive data is not the target sensitive data.
[0151] The data processing device provided in the embodiments of this specification adopts a data processing algorithm and a target scoring unit to implement sensitivity scoring for each initial sensitive data. Subsequently, the target sensitive data can be quickly and accurately determined based on the target scoring result determined by the sensitivity score, so that the target sensitive data can be adjusted subsequently. This automated method can greatly save labor costs on the basis of improving efficiency.
[0152] The above is a schematic diagram of a data processing device according to this embodiment. It should be noted that the technical solution of the data processing device and the technical solution of the above-mentioned data processing method are based on the same concept. For details not described in detail in the technical solution of the data processing device, please refer to the description of the technical solution of the above-mentioned data processing method.
[0153] Figure 5 The block diagram of a computing device 500 according to one embodiment of the present disclosure is shown. Components of the computing device 500 include, but are not limited to, a memory 510 and a processor 520. The processor 520 is connected to the memory 510 via a bus 530, and a database 550 is used to store data.
[0154] The computing device 500 also includes an access device 540 that enables the computing device 500 to communicate via one or more networks 560. Examples of such networks include a public switched telephone network (PSTN), a local area network (LAN), a wide area network (WAN), a personal area network (PAN), or a combination of communication networks such as the Internet. The access device 540 may include one or more of any type of network interface (e.g., a network interface card (NIC)), whether wired or wireless, such as an IEEE 802.11 wireless local area network (WLAN) wireless interface, a Worldwide Interoperability for Microwave Access (Wi-MAX) interface, an Ethernet interface, a universal serial bus (USB) interface, a cellular network interface, a Bluetooth interface, a near field communication (NFC) interface, and the like.
[0155] In one embodiment of the present specification, the above components of the computing device 500 and Figure 5 Other components not shown in the figure may also be connected to each other, for example, via a bus. Figure 5 The computing device structure block diagram shown is for illustrative purposes only and is not intended to limit the scope of this specification. Those skilled in the art may add or replace other components as needed.
[0156] Computing device 500 can be any type of stationary or mobile computing device, including a mobile computer or mobile computing device (e.g., a tablet computer, a personal digital assistant, a laptop computer, a notebook computer, a netbook computer, etc.), a mobile phone (e.g., a smartphone), a wearable computing device (e.g., a smartwatch, smart glasses, etc.), or other types of mobile devices, or a stationary computing device such as a desktop computer or PC. Computing device 500 can also be a mobile or stationary server.
[0157] The processor 520 is configured to execute the following computer-executable instructions, which implement the steps of the above-mentioned data processing method when executed by the processor.
[0158] The above is a schematic scheme of a computing device of this embodiment. It should be noted that the technical scheme of the computing device and the technical scheme of the above-mentioned data processing method are based on the same concept. For details not described in detail in the technical scheme of the computing device, please refer to the description of the technical scheme of the above-mentioned data processing method.
[0159] An embodiment of the present specification further provides a computer-readable storage medium storing computer-executable instructions, which implement the steps of the above-mentioned data processing method when executed by a processor.
[0160] The above is a schematic scheme of a computer-readable storage medium of this embodiment. It should be noted that the technical scheme of the storage medium and the technical scheme of the above-mentioned data processing method are based on the same concept. For details not described in detail in the technical scheme of the storage medium, please refer to the description of the technical scheme of the above-mentioned data processing method.
[0161] An embodiment of the present specification further provides a computer program, wherein when the computer program is executed in a computer, the computer is caused to execute the steps of the above-mentioned data processing method.
[0162] The above is an illustrative solution of a computer program of this embodiment. It should be noted that the technical solution of the computer program and the technical solution of the above-mentioned data processing method are of the same concept. For details not described in detail in the technical solution of the computer program, please refer to the description of the technical solution of the above-mentioned data processing method.
[0163] The foregoing description of this specification describes specific embodiments. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0164] The computer instructions include computer program code, which may be in source code form, object code form, executable file, or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content contained in the computer-readable medium may be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media do not include electric carrier signals and telecommunication signals.
[0165] It should be noted that for the aforementioned method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the embodiments of this specification are not limited by the order of the actions described, because according to the embodiments of this specification, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in this specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the embodiments of this specification.
[0166] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0167] The preferred embodiments disclosed above are intended only to help illustrate this specification. The optional embodiments do not exhaustively describe all details, nor do they limit the invention to the specific embodiments described. Obviously, many modifications and variations can be made based on the content of the embodiments of this specification. This specification selects and specifically describes these embodiments in order to better explain the principles and practical applications of the embodiments of this specification, so that those skilled in the art can better understand and utilize this specification. This specification is limited only by the claims and their full scope and equivalents.
Claims
1. A data processing method, comprising: Obtaining a sensitive data set including initial sensitive data, and determining a data scenario corresponding to the initial sensitive data; Determining a data processing algorithm corresponding to the initial sensitive data and at least two initial scoring units according to the data scenario, and calculating a degree of association between each initial scoring unit and the data scenario; adjusting a scoring strategy of each initial scoring unit according to the correlation degree, and determining each adjusted initial scoring unit as at least two target scoring units corresponding to the initial sensitive data; When it is determined according to the data processing algorithm that the initial sensitive data meets the preset processing conditions, the initial sensitive data is subjected to sensitivity scoring according to the at least two target scoring units to obtain a target scoring result.
2. The data processing method according to claim 1, wherein obtaining a sensitive data set including initial sensitive data comprises: Obtain a sensitive data set containing initial sensitive data determined by a preset risk identification terminal.
3. The data processing method according to claim 1, wherein the step of performing sensitivity scoring on the initial sensitive data according to the at least two target scoring units comprises: Perform a sensitivity score on the initial sensitive data according to the scoring dimension and scoring strategy corresponding to each target scoring unit in the at least two target scoring units.
4. The data processing method according to claim 1, wherein the step of performing sensitivity scoring on the initial sensitive data according to the at least two target scoring units to obtain a target scoring result comprises: Performing a sensitivity score on the initial sensitive data according to each of the at least two target scoring units, to obtain an initial scoring result of each target scoring unit for the initial sensitive data; The initial scoring result of the initial sensitive data is calculated according to a preset calculation rule to obtain a target scoring result of the initial sensitive data.
5. The data processing method according to claim 1, further comprising: Determine whether the initial sensitive data is target sensitive data based on the target scoring result.
6. The data processing method according to claim 5, wherein determining whether the initial sensitive data is target sensitive data according to the target scoring result comprises: Determining target scoring results for all initial sensitive data in the sensitive data set; Determining an average scoring result based on target scoring results of all initial sensitive data in the sensitive data set; Determine whether the initial sensitive data is target sensitive data based on the average scoring result.
7. The data processing method according to claim 5, wherein determining whether the initial sensitive data is target sensitive data according to the target scoring result comprises: Determining target scoring results for all initial sensitive data in the sensitive data set; sorting all the initial sensitive data in the sensitive data set according to the target scoring results of all the initial sensitive data in the sensitive data set; Determine whether the initial sensitive data is target sensitive data based on the sorting result.
8. The data processing method according to claim 5, wherein determining whether the initial sensitive data is target sensitive data according to the target scoring result comprises: In the case where it is determined that there is a historical scoring result for the initial sensitive data, whether the initial sensitive data is the target sensitive data is determined according to the historical scoring result of the initial sensitive data and the target scoring result.
9. The data processing method according to claim 8, wherein determining whether the initial sensitive data is target sensitive data based on the historical scoring results of the initial sensitive data and the target scoring results comprises: Calculating a difference between the historical scoring result and the target scoring result based on the historical scoring result of the initial sensitive data and the target scoring result; When the difference is greater than a preset difference threshold, determining that the initial sensitive data is target sensitive data; as well as When the difference is less than or equal to the preset difference threshold, it is determined that the initial sensitive data is not the target sensitive data.
10. A data processing device comprising: A data acquisition module is configured to acquire a sensitive data set including initial sensitive data and determine a data scenario corresponding to the initial sensitive data; an algorithm determination module configured to determine, based on the data scenario, a data processing algorithm corresponding to the initial sensitive data and at least two initial scoring units, and calculate a degree of association between each initial scoring unit and the data scenario; adjust a scoring strategy for each initial scoring unit based on the degree of association, and determine each adjusted initial scoring unit as the at least two target scoring units corresponding to the initial sensitive data; The scoring module is configured to, when it is determined according to the data processing algorithm that the initial sensitive data meets the preset processing conditions, perform sensitivity scoring on the initial sensitive data according to the at least two target scoring units to obtain a target scoring result.
11. A computing device comprising: memory and processor; The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the data processing method according to any one of claims 1 to 9 are implemented.
12. A computer-readable storage medium storing computer-executable instructions, wherein the computer-executable instructions, when executed by a processor, implement the steps of the data processing method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Data leakage identification method, device and equipment
CN110717189A
Data grading method and device, and related equipment
CN110941956A