Data processing method and device
By acquiring dynamic scoring cards for data scenarios and scoring units, DLP alarm data is processed automatically, solving the problems of large alarm data volume and low manual processing efficiency in DLP systems, and achieving efficient risk identification and response.
Patent Information
- Application Number
- CN202511064115.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-17
- Publication Date
- 2025-11-07
AI Technical Summary
In existing technologies, DLP systems generate massive amounts of alarm data, which are inefficient and costly to process manually, and cannot respond to real risks in a timely manner.
By employing data processing methods, the data scenario of acquiring initial sensitive data is used to determine the data processing algorithm and target scoring unit. The dynamic scoring card covers risk identification factors, and automated scoring is used to identify high-risk data, thereby reducing labor costs.
It enables fast and accurate processing of DLP alarm data, reduces response time, improves risk control efficiency, and saves manpower costs.
Smart Images

Figure CN120910909A_ABST
Abstract
Description
[0001] This application is a divisional application of application No. 202210265154.8, filed on March 17, 2022, entitled "Data processing method and device". TECHNICAL FIELD
[0002] Embodiments of the present specification relate to the technical field of data processing, in particular to a data processing method. BACKGROUND
[0003] With the coming into effect of data protection laws, individuals or organizations are increasingly aware of the protection of sensitive data. Organizations generally spend a large budget to purchase DLP (Data Loss Prevention System) to monitor sensitive data, but the amount of alert data output by DLP is often huge every day. If manual methods are used to process these alert data, a large amount of manpower will be consumed, and the real risks cannot be responded to in a timely manner, and the overall cost is seriously wasted. SUMMARY
[0004] Therefore, the embodiments of the present specification provide a data processing method. One or more embodiments of the present specification also relate to a data processing device, a computing device, a computer-readable storage medium, and a computer program to solve the technical defects in the prior art.
[0005] According to a first aspect of the embodiments of the present specification, a data processing method is provided, comprising: obtaining a sensitive data set containing initial sensitive data, and determining a data scenario corresponding to the initial sensitive data; determining a data processing algorithm corresponding to the initial sensitive data and at least two target scoring units according to the data scenario; in a case where it is determined according to the data processing algorithm that the initial sensitive data meets a preset processing condition, performing sensitivity scoring on the initial sensitive data according to the at least two target scoring units to obtain a target scoring result.
[0006] According to a second aspect of the embodiments of the present specification, a data processing device is provided, comprising: a data acquisition module configured to obtain a sensitive data set containing initial sensitive data, and determine a data scenario corresponding to the initial sensitive data; an algorithm determination module configured to determine a data processing algorithm corresponding to the initial sensitive data and at least two target scoring units according to the data scenario; a scoring module configured to, in a case where it is determined according to the data processing algorithm that the initial sensitive data meets a preset processing condition, perform sensitivity scoring on the initial sensitive data according to the at least two target scoring units to obtain a target scoring result.
[0007] According to a third aspect of the embodiments of the present specification, a computing device is provided, comprising: a memory and a processor; the memory is configured to store computer-executable instructions, and the processor is configured to execute the computer-executable instructions, which, when executed by the processor, implement the steps of the above data processing method.
[0008] According to a fourth aspect of the embodiments of the present specification, a computer-readable storage medium is provided, which stores computer-executable instructions, which, when executed by a processor, implement the steps of the above data processing method.
[0009] According to a fifth aspect of the embodiments of the present specification, a computer program is provided, which, when executed in a computer, causes the computer to perform the steps of the above data processing method.
[0010] One embodiment of the present specification implements a data processing method and device, wherein the method comprises obtaining a sensitive data set containing initial sensitive data, and determining a data scenario corresponding to the initial sensitive data; determining a data processing algorithm corresponding to the initial sensitive data and at least two target scoring units according to the data scenario; in the case where it is determined according to the data processing algorithm that the initial sensitive data satisfies a preset processing condition, performing sensitivity scoring on the initial sensitive data according to the at least two target scoring units to obtain a target scoring result.
[0011] Specifically, the data processing method uses a data processing algorithm and a target scoring unit to implement sensitivity scoring on each initial sensitive data, and subsequently, a target sensitive data can be quickly and accurately determined according to the target scoring result determined by the sensitivity scoring, so as to subsequently adjust the target sensitive data. Through this kind of automatic way, the efficiency is improved, and a large amount of human cost is saved. BRIEF DESCRIPTION OF DRAWINGS
[0012] Figure 1 is a specific application scenario diagram of a data processing method provided by one embodiment of the present specification; Figure 2 is a flowchart of a data processing method provided by one embodiment of the present specification; Figure 3 is a process flowchart of a data processing method provided by one embodiment of the present specification; Figure 4 is a structure diagram of a data processing device provided by one embodiment of the present specification; Figure 5is a structural block diagram of a computing device provided by one embodiment of the present specification. DETAILED DESCRIPTION
[0013] In the following description, numerous specific details are set forth in order to provide a thorough understanding of the present specification. However, the present specification can be practiced without the specific details, other than those described herein, and it is understood that the scope of the present specification is not limited to the details of the description.
[0014] The terminology used in one or more embodiments of the present specification is for the purpose of describing particular embodiments only and is not intended to be limiting of one or more embodiments of the present specification. As used in one or more embodiments of the present specification and the accompanying claims, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms "comprises" and / or "comprising," when used in one or more embodiments of the present specification, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0015] It will be understood that, although the terms first, second, etc. can be employed in one or more embodiments of the present specification, these terms are used to distinguish one information from another and are not intended to denote a particular order or sequence unless otherwise indicated by the context. For example, without departing from the scope of one or more embodiments of the present specification, first can be termed second and, similarly, second can be termed first. The word "if' as used herein means "when" or "upon" or "in response to a determination" depending on the context.
[0016] First, the noun terms related to one or more embodiments of the present specification are explained.
[0017] Sensitive data: refers to personal privacy data, such as: ID number, mobile phone number, bank card number, etc.
[0018] DLP: Data Loss Prevention System.
[0019] SFTP: Secure File Transfer Protocol.
[0020] Each organization will evaluate the level of data as public, internal, confidential, secret or top secret according to the harm caused by sensitive data leakage to the organization and the value of data to the organization. Through the data classification and grading marking program, sensitive data is classified and graded and marked in the host, terminal, application and other environments. The marked data will sort out the data scene, covering many data leakage strategies. The technical problem to be solved by the embodiments of the present specification is that the alarm data output by these data leakage strategies is huge in magnitude, and the manual processing efficiency is low and the cost is high.
[0021] Based on this, in the present specification, a data processing method is provided, the present specification also relates to a data processing device, a computing device, and a computer readable storage medium, which are described in detail one by one in the following embodiments.
[0022] Referring to Figure 1 , Figure 1 A specific application scenario diagram of a data processing method according to an embodiment of the present specification is shown.
[0023] The data processing method provided by the embodiment of the present specification is applied to a data processing system in Figure 1 , which includes a data storage unit 102, an algorithm unit 104, a score card unit 106, a data analysis unit 108, and a data processing unit 110.
[0024] The data storage unit 102 stores a plurality of DLP alert data output by DLP, such as sensitive chat records, file transmission, and system calls. The algorithm unit 104 stores a plurality of algorithms, such as semantic rule algorithm, model algorithm, and business logic algorithm. When calculating, the data processing system will select the corresponding algorithm according to the data scene corresponding to each DLP alert data, analyze each DLP alert data through the selected algorithm, further filter whether each DLP alert data is high-risk data, and if so, send it to the score card unit 106, and if not, do not process it subsequently. The score card unit 106 stores a plurality of score cards, such as personnel score card, time score card, file score card, application score card, data score card, address score card, device score card, and relationship person score card. Each score card corresponds to a corresponding score interval (such as 1-10 points, 0-5 points, etc.), a score dimension (such as post responsibility, willingness to resign, and resignation status), and a score strategy (such as calculation factor is addition or multiplication, etc.). The sensitivity score of each DLP alert data can be obtained according to the score card corresponding to each DLP alert data, and the target score result of each DLP alert data can be obtained. The data analysis unit 108 can update the dynamic score result of each DLP alert data according to the target score result, and judge whether each DLP alert data is target sensitive data combined with the preset processing rule, so as to determine which DLP alert data needs subsequent processing. The preset processing rule can be set according to actual application, and the embodiment of the present specification is not limited thereto. The data processing unit 110 will process the post-response of the initial sensitive data (i.e. target sensitive data) determined by the data analysis unit 108, such as issuing warning information.
[0025] The data processing method provided by the embodiments of the present specification can comprehensively and accurately process each DLP alarm data by using the risk identification factors covered by the dynamic scorecard, and the calculation factors of each scorecard can be dynamically adjusted according to different data scenarios, such as the data scenario of device sharing and moving. The weight value of the calculation factor of the device scorecard can be automatically increased. In addition, the risk identification factors in each scorecard can also be dynamically adjusted. The following will take three scorecards as examples for introduction: Personnel scorecard: the staff is the root cause of the data leakage event. The risk factors such as the post responsibility of the staff, the willingness to leave, and the state of leaving will be attached with weights. The final score is dynamically adjusted according to the real-time submission of the staff's resignation application and the change information of job transfer; File scorecard: the file is an important carrier of loading sensitive data. The risk factors such as file size, file content, encryption compression record will be attached with weights. The weight and score are analyzed according to the historical operation of the file; Relationship scorecard: the relationship person can be a relative, colleague, supervisor, friend, etc. of the staff. In order to prevent collusion and data theft behavior, the behavior of the relationship person is an important factor to evaluate this risk.
[0026] In the embodiments of the present specification, the sensitivity score of each DLP alarm data corresponding to the scorecard will be brought into the risk score formula to calculate the final risk score. The risk score can also calculate the risk acceleration and deceleration, so as to estimate the risk abnormality of the staff. The final score will affect the final disposal decision, such as ending the disposal process, reaching the supervisor for investigation, automatically publishing the interception strategy, etc. This will greatly reduce the labor cost, shorten the response time, and improve the efficiency of risk hemostasis, etc.
[0027] Referring to Figure 2 , Figure 2 A flowchart of a data processing method according to an embodiment of the present specification is shown, which specifically includes the following steps.
[0028] Step 202: obtaining a sensitive data set containing initial sensitive data, and determining a data scenario corresponding to the initial sensitive data.
[0029] The data processing method provided by the embodiments of the present specification is applied to a data leakage scene, and the data leakage scene in each organization can occur in multiple site environments such as office terminal computers, application systems, data warehouses, ecological institutions, host servers, and the like. The data leakage data scene in the above environments has been sorted out and clarified in the initial stage. The organization will purchase basic monitoring software such as DLP data leakage protection software, network DLP, application system operation logs, host HIDS, and the like. The original log data output based on the basic monitoring software can be understood as the initial sensitive data of the embodiments of the present specification. The specific implementation manner is as follows: In a specific implementation, the sensitive data set containing the initial sensitive data comprises: The sensitive data set containing the initial sensitive data determined by the preset risk identification terminal is acquired.
[0030] The preset risk identification terminal includes but is not limited to DLP data leakage protection software, network DLP, application system operation logs, host HIDS, and the like.
[0031] In actual application, the data scene corresponding to each initial sensitive data is determined in the initial stage. As described above, the data leakage scene of the office terminal computer, the data leakage scene of the application system, the data leakage scene of the data warehouse, the data leakage scene of the ecological institution, and the data leakage scene of the host server are included.
[0032] Step 204: determining a data processing algorithm corresponding to the initial sensitive data and at least two target scoring units according to the data scene.
[0033] In actual application, the data processing algorithm includes but is not limited to a semantic rule algorithm, a machine model algorithm, and a business logic algorithm. The semantic rule method includes keyword, keyword, and regular rule recognition capability. The machine model algorithm includes relationship graph analysis, frequency increase analysis, and resignation model analysis algorithms, and involves relationship, frequency, and deviation strategies. The business logic algorithm is mainly based on the business attributes of the organization, such as the financial related attributes of the department of the enterprise issuing loans to query personal accounts as abnormal, which is a strategy formulated by the business logic. The target scoring unit includes but is not limited to a personnel scoring unit, a time scoring unit, a file scoring unit, an application scoring unit, a data scoring unit, an address scoring unit, a device scoring unit, and a relationship person scoring unit.
[0034] Specifically, according to the difference of the data scene corresponding to each initial sensitive data, the data processing algorithm and the target scoring unit used for subsequent processing of the initial sensitive data are also different.
[0035] Because the attributes of the data scene corresponding to each initial sensitive data are different, it is determined from the initial stage which data processing algorithm and target scoring unit each initial sensitive data is analyzed by; for example, if the data scene corresponding to the initial sensitive data is a data leakage scene of an office terminal computer (such as a data leakage risk of device moving), the data scene is strongly related to the device, and the identification of the keyword and the calling of the device scoring unit are analyzed by giving priority to the semantic rule algorithm; if the data scene corresponding to the initial sensitive data is a data leakage scene of an application system (such as a data leakage risk of application violation query), the data scene is strongly related to the query customer's relationship network, and the relationship network is analyzed by giving priority to the machine algorithm model and the calling of the application scoring unit.
[0036] Each scoring unit corresponds to a scoring interval, for example, the scoring interval corresponding to the personnel scoring unit is 1-10 points, the scoring interval corresponding to the time scoring unit is 0-5 points, the scoring interval corresponding to the file scoring unit is 0-100 points, the scoring interval corresponding to the application scoring unit is 0-100 points, the scoring interval corresponding to the data scoring unit is 0-100 points, the scoring interval corresponding to the address scoring unit is 0-10 points, the scoring interval corresponding to the device scoring unit is 0-10 points, and the scoring interval corresponding to the relationship person scoring unit is 0-10 points.
[0037] The weight value of the calculation factor of each scoring unit can be dynamically adjusted according to different data scenes, such as the device sharing moving scene, and the weight value of the calculation factor of the device scoring unit is also automatically adjusted. The specific implementation manner is as follows: The data processing algorithm corresponding to the initial sensitive data and the at least two target scoring units are determined according to the data scene, and the data processing algorithm corresponding to the initial sensitive data and the at least two target scoring units are determined according to the data scene. The data processing algorithm corresponding to the initial sensitive data and the at least two target scoring units are determined according to the data scene, and the data processing algorithm corresponding to the initial sensitive data and the at least two target scoring units are determined according to the data scene. The correlation degree of each initial scoring unit and the data scene is calculated. The scoring strategy of each initial scoring unit is adjusted according to the correlation degree, and each initial scoring unit after adjustment is determined as the at least two target scoring units corresponding to the initial sensitive data.
[0038] The scoring strategy can be understood as the final scoring result of the initial scoring unit being addition or multiple growth.
[0039] Specifically, the score strategy of each initial scoring unit is adjusted according to the correlation degree, which can be understood as that the weight value of the calculation factor of each initial scoring unit is adjusted according to the correlation degree, and the score strategy of each initial scoring unit is determined according to the weight value, for example, the final score result of the initial scoring unit can be multiplied after the weight value adjustment, and the like, so that the score calculated by the initial scoring unit is more prominent.
[0040] In actual application, two or more initial scoring units corresponding to each data scene, in the case of corresponding to multiple initial scoring units, the correlation degree of each initial scoring unit and the data scene can be calculated, and the weight value of the calculation factor of each initial scoring unit can be adjusted according to the correlation degree, and the score strategy is adjusted according to the weight value, so as to obtain the target scoring unit corresponding to the initial sensitive data, wherein the specific calculation method of the correlation degree of the initial scoring unit and the data scene can adopt any method, for example, the similarity of the keyword is calculated, and the like, which is not limited in the embodiments of the present application.
[0041] For example, when the data scene is a device sharing misappropriation scene, the target scoring unit corresponding to the data scene includes a device scoring unit and a relationship person scoring unit, and the correlation degree of the device scoring unit and the data scene can be determined by the similarity of the keyword, and then the weight value of the calculation factor of the device scoring unit can be dynamically increased.
[0042] Step 206: In the case that the initial sensitive data meets the preset processing condition according to the data processing algorithm, the sensitivity of the initial sensitive data is scored according to the at least two target scoring units, and a target scoring result is obtained.
[0043] The role of the data processing algorithm in the embodiments of the present application can be understood as further screening of the initial sensitive data. That is, before the sensitivity of a certain initial sensitive data is scored, the sensitivity of the initial sensitive data needs to be further processed according to the data processing algorithm corresponding to the initial sensitive data, so as to determine whether to perform the next step of sensitivity scoring. That is, in the case that the initial sensitive data is determined to be high-risk data leakage by the data processing algorithm corresponding to the initial sensitive data, the sensitivity of the initial sensitive data is scored according to the target scoring algorithm corresponding to the initial sensitive data.
[0044] Specifically, the initial sensitive data meets the preset processing condition according to the data processing algorithm, which can be understood as that the initial sensitive data meets the preset processing condition in the case that the initial sensitive data is determined to be high-risk leakage data according to the data processing algorithm.
[0045] In practical applications, the data processing algorithm includes but is not limited to semantic rule algorithm, machine model algorithm, business logic algorithm, etc.; wherein, the semantic rule algorithm contains keyword, keyword, regular rule recognition ability, which can determine whether the initial sensitive data is high-risk leakage data according to the above-mentioned recognition ability; the machine model algorithm contains relationship graph analysis, frequency amplitude analysis, and off-job model analysis algorithm. When the initial sensitive data involves relationship, frequency, and deviation, the strategy can be analyzed using this method to determine whether the initial sensitive data is high-risk leakage data; the business logic algorithm is based on the business attributes of the organization, such as the financial related attributes. The department of the enterprise that issues loans queries personal accounts as an exception, which is a strategy formulated by business logic. Through this algorithm, it is determined whether the initial sensitive data is high-risk leakage data.
[0046] In specific implementation, in the case where it is determined according to the data processing algorithm that the initial sensitive data meets the preset processing condition, the sensitivity of the initial sensitive data is scored according to the at least two target scoring units, and a target scoring result is obtained. The specific implementation is as follows: The sensitivity of the initial sensitive data is scored according to the at least two target scoring units, including: The sensitivity of the initial sensitive data is scored according to the scoring dimension and the scoring strategy corresponding to each target scoring unit in the at least two target scoring units.
[0047] Wherein, the scoring dimension corresponding to each target scoring unit is different. Taking a target scoring unit as a personnel scoring unit as an example, the scoring dimension corresponding to the target scoring unit includes but is not limited to the post responsibility, the willingness to leave, and the off-job state of the personnel. Taking a target scoring unit as a file scoring unit as an example, the scoring dimension corresponding to the target scoring unit includes but is not limited to the file size, the file content, and the encryption compression record. Taking a target scoring unit as a relationship person scoring unit as an example, the scoring dimension corresponding to the relationship person scoring unit includes but is not limited to the employee's relatives, colleagues, supervisors, and friends.
[0048] In practical applications, each target scoring unit corresponds to a scoring interval, and the scoring dimension corresponding to each target scoring unit is also matched with a corresponding score value based on the scoring interval. In specific applications, each initial sensitive data corresponds to a target scoring unit, which scores the sensitivity of each initial sensitive data according to the corresponding scoring dimension and scoring strategy. So that the target scoring result can be accurately obtained according to the sensitivity score of each initial sensitive data by the target scoring unit.
[0049] In a specific implementation, since each initial sensitive data corresponds to two or more target scoring units, after obtaining the sensitivity score of each target scoring unit, the target scoring result of each initial sensitive data is accurately obtained by summation, multiplication or other calculation methods. The specific implementation is as follows: The sensitivity of the initial sensitive data is scored according to the at least two target scoring units, and a target scoring result is obtained, including: The sensitivity of the initial sensitive data is scored according to each target scoring unit of the at least two target scoring units, and an initial scoring result of the initial sensitive data for each target scoring unit is obtained. The initial scoring result of the initial sensitive data is calculated according to a pre-designed calculation rule, and a target scoring result of the initial sensitive data is obtained.
[0050] The pre-designed calculation rule includes but is not limited to summation, multiplication or other calculation rules.
[0051] Specifically, taking summation as the pre-designed calculation rule, in the case where the target scoring unit corresponding to a certain initial sensitive data is two or more, the sensitivity of the initial sensitive data is scored according to each adjusted target scoring unit, and an initial scoring result of the initial sensitive data for each target scoring unit is obtained. The initial scoring result of the initial sensitive data for each target scoring unit is added to obtain the target scoring result of the initial sensitive data.
[0052] In actual application, after obtaining the target scoring result of each initial sensitive data, it can be determined whether the initial sensitive data is target sensitive data according to the target scoring result, and subsequent specific business processing can be performed on the target sensitive data to ensure the security of the overall business. The specific implementation is as follows: After obtaining the target scoring result, it further includes: According to the target scoring result, it is determined whether the initial sensitive data is target sensitive data.
[0053] In a specific implementation, there are at least three ways to determine whether the initial sensitive data is target sensitive data according to the target scoring result, and the following three are taken as examples to introduce each way of determining whether the initial sensitive data is target sensitive data according to the target scoring result.
[0054] The first way is to determine the average score result according to the target scoring result, and to determine whether the initial sensitive data is target sensitive data by the average score result. The specific implementation is as follows: According to the target scoring result, it is determined whether the initial sensitive data is target sensitive data, including: determining a target score result of all initial sensitive data in the sensitive data set; determining an average score result according to the target score result of all initial sensitive data in the sensitive data set; determining whether the initial sensitive data is target sensitive data according to the average score result.
[0055] In specific implementation, the target score result of all initial sensitive data in the sensitive data set is obtained, and then the average score result is calculated according to the target score result of all initial sensitive data in the sensitive data set; and then the target score result of each initial sensitive data is compared with the average score result to determine whether the initial sensitive data is target sensitive data.
[0056] For example, if the target score result is greater than the average score result, it can be considered that the initial sensitive data is target sensitive data; if the target score result is less than or equal to the average score result, it can be considered that the initial sensitive data is not target sensitive data. In this way, it can be quickly determined whether each initial sensitive data is target sensitive data.
[0057] Secondly, the target score result of all initial sensitive data in the sensitive data set is obtained, and whether it is target sensitive data is determined according to the sorting result of the target score result of all initial sensitive data. The specific implementation is as follows: The determining whether the initial sensitive data is target sensitive data according to the target score result comprises: determining a target score result of all initial sensitive data in the sensitive data set; sorting all initial sensitive data in the sensitive data set according to the target score result of all initial sensitive data in the sensitive data set; determining whether the initial sensitive data is target sensitive data according to the sorting result.
[0058] For example, the target score result of all initial sensitive data in the sensitive data set is obtained, all initial sensitive data in the sensitive data set is arranged in descending order according to the target score result of all initial sensitive data in the sensitive data set, and then the top 30 or top 50 initial sensitive data is selected as target sensitive data according to the sorting result. In this way, the target sensitive data can also be quickly obtained.
[0059] Thirdly, whether the initial sensitive data is target sensitive data is determined according to a historical score result. The specific implementation is as follows: The determining whether the initial sensitive data is target sensitive data according to the target score result comprises: In a case where it is determined that the initial sensitive data has a historical scoring result, whether the initial sensitive data is target sensitive data is determined according to the historical scoring result of the initial sensitive data and the target scoring result.
[0060] Specifically, after obtaining the target scoring result of the initial sensitive data, if it is determined that the initial sensitive data has a historical scoring result, whether each initial sensitive data is target sensitive data can be determined according to the historical scoring result and the target scoring result of each initial sensitive data.
[0061] One implementation manner is that the determining whether the initial sensitive data is target sensitive data according to the historical scoring result of the initial sensitive data and the target scoring result comprises: calculating a difference between the historical scoring result and the target scoring result according to the historical scoring result of the initial sensitive data and the target scoring result; in a case where the difference is greater than a preset difference threshold, determining that the initial sensitive data is target sensitive data; and in a case where the difference is less than or equal to the preset difference threshold, determining that the initial sensitive data is not target sensitive data.
[0062] The preset difference threshold can be set according to actual application, and the present specification does not make any limitation on this. For example, 60 or 70, etc.
[0063] In a specific implementation, after obtaining the historical scoring result (such as the last scoring result of the current scoring result) and the target scoring result of each initial sensitive data, the difference between the historical scoring result and the target scoring result of each initial sensitive data is obtained, and whether each initial sensitive data is target sensitive data is determined according to the difference.
[0064] For example, the preset difference threshold is 70 points, and the difference between the historical scoring result and the target scoring result of the initial sensitive data is 90 points, it can be determined that the difference is greater than the preset difference threshold, and it can be determined that the initial sensitive data is target sensitive data; and if the difference between the historical scoring result and the target scoring result of the initial sensitive data is 50 points, it can be determined that the difference is less than the preset difference threshold, and it can be determined that the initial sensitive data is not target sensitive data.
[0065] Another implementation manner is that a previous historical scoring result of each initial sensitive data can be obtained, and a previous historical scoring result of the previous historical scoring result is obtained; then, a difference between the previous historical scoring result and the previous historical scoring result of the previous historical scoring result is calculated, and the difference between the previous historical scoring result and the difference between the previous historical scoring result and the previous historical scoring result of the previous historical scoring result is compared, and if a value obtained by subtracting the two differences is greater than a preset threshold, it is determined that the initial sensitive data is the target sensitive data.
[0066] The data processing method provided by the embodiments of the present specification adopts a data processing algorithm and a target scoring unit to score the sensitivity of each sensitive dimension of each initial sensitive data, and finally determines the target sensitive data quickly and accurately according to the target scoring result determined by the sensitivity score, so as to adjust the target sensitive data subsequently. Through this kind of automatic way, the efficiency is improved, and a large amount of manpower cost is saved.
[0067] The following describes the embodiments of the present specification in conjunction with the accompanying Figure 3 The data processing method provided by the embodiments of the present specification is further described by taking the processing of the DLP alarm data as an example. Among them, Figure 3 FIG. 3 shows a process flow diagram of a data processing method provided by an embodiment of the present specification, which specifically includes the following steps.
[0068] Step 302: Obtain DLP alarm data.
[0069] Step 304: Determine the data scene of each DLP alarm data, and determine the data processing algorithm corresponding to each DLP alarm data and at least two scoring cards according to the data scene of each DLP alarm data.
[0070] Among them, the scoring card can be understood as the scoring unit described above.
[0071] Step 306: If it is determined that the DLP alarm data is high-risk leakage data according to the data processing algorithm corresponding to each DLP alarm data, the DLP alarm data is sent to the corresponding at least two target scoring units.
[0072] Step 308: According to the at least two scoring cards corresponding to each DLP alarm data, the sensitivity of each DLP alarm data is scored to obtain a target scoring result of each DLP alarm data.
[0073] Step 310: According to the target scoring result of each DLP alarm data, it is determined whether each DLP alarm data is target sensitive data.
[0074] The data processing method provided by the embodiment of the present specification can comprehensively and accurately obtain the target scoring result of each DLP alarm data by using the dynamic scoring card mechanism, and subsequently, the DLP alarm data that needs to be post-processed can be obtained according to the target scoring result in the actual application scene, thereby ensuring the security of the system.
[0075] Corresponding to the method embodiments described above, the present specification also provides data processing device embodiments, Figure 4 The structure diagram of a data processing device provided by one embodiment of the present specification is shown. As shown in the figure, Figure 4 The device comprises: The data acquisition module 402 is configured to acquire a sensitive data set containing initial sensitive data and determine a data scene corresponding to the initial sensitive data. The algorithm determination module 404 is configured to determine a data processing algorithm corresponding to the initial sensitive data and at least two target scoring units according to the data scene. The scoring module 406 is configured to, in the case where it is determined that the initial sensitive data meets a preset processing condition according to the data processing algorithm, perform sensitivity scoring on the initial sensitive data according to the at least two target scoring units to obtain a target scoring result.
[0076] Optionally, the data acquisition module 402 is further configured to: Acquire a sensitive data set containing initial sensitive data determined by a preset risk identification terminal.
[0077] Optionally, the algorithm determination module 404 is further configured to: Determine a data processing algorithm corresponding to the initial sensitive data and at least two initial scoring units according to the data scene. Calculate the correlation degree of each initial scoring unit and the data scene. Adjust the scoring strategy of each initial scoring unit according to the correlation degree, and determine the adjusted each initial scoring unit as the at least two target scoring units corresponding to the initial sensitive data.
[0078] Optionally, the scoring module 406 is further configured to: Perform sensitivity scoring on the initial sensitive data according to the scoring dimension and the scoring strategy corresponding to each target scoring unit in the at least two target scoring units.
[0079] Optionally, the scoring module 406 is further configured to: score the initial sensitive data according to each target scoring unit to obtain an initial scoring result of the initial sensitive data by the each target scoring unit; calculate the initial scoring result of the initial sensitive data according to a preset calculation rule to obtain a target scoring result of the initial sensitive data.
[0080] Optionally, the apparatus further comprises: a target data determination module configured to: determine whether the initial sensitive data is target sensitive data according to the target scoring result.
[0081] Optionally, the target data determination module is further configured to: determine target scoring results of all initial sensitive data in the sensitive data set; determine an average scoring result according to the target scoring results of all initial sensitive data in the sensitive data set; determine whether the initial sensitive data is target sensitive data according to the average scoring result.
[0082] Optionally, the target data determination module is further configured to: determine target scoring results of all initial sensitive data in the sensitive data set; sort all initial sensitive data in the sensitive data set according to the target scoring results of all initial sensitive data in the sensitive data set; determine whether the initial sensitive data is target sensitive data according to the sorting result.
[0083] Optionally, the target data determination module is further configured to: in a case where it is determined that the initial sensitive data has a historical scoring result, determine whether the initial sensitive data is target sensitive data according to the historical scoring result of the initial sensitive data and the target scoring result.
[0084] Optionally, the target data determination module is further configured to: calculate a difference between the historical scoring result of the initial sensitive data and the target scoring result according to the historical scoring result of the initial sensitive data and the target scoring result; in a case where the difference is greater than a preset difference threshold, determine that the initial sensitive data is target sensitive data; and in a case where the difference is less than or equal to the preset difference threshold, determine that the initial sensitive data is not target sensitive data.
[0085] The data processing apparatus provided by the embodiment of the present specification adopts a data processing algorithm and a target scoring unit to score the sensitivity of each initial sensitive data, and subsequently, the target sensitive data can be quickly and accurately determined according to the target scoring result determined by the sensitivity score, so as to subsequently adjust the target sensitive data. Through this kind of automatic way, the efficiency is improved, and a large amount of manpower cost is saved.
[0086] The above is a schematic scheme of the data processing apparatus of the embodiment. It should be noted that the technical scheme of the data processing apparatus and the technical scheme of the data processing method described above belong to the same concept, and the details of the technical scheme of the data processing apparatus which are not described in detail can be referred to the description of the technical scheme of the data processing method.
[0087] Figure 5 A structural block diagram of a computing device 500 according to an embodiment of the present specification is shown. The components of the computing device 500 include, but are not limited to, a memory 510 and a processor 520. The processor 520 is connected to the memory 510 through a bus 530, and a database 550 is used to save data.
[0088] The computing device 500 also includes an access device 540, which enables the computing device 500 to communicate via one or more networks 560. Examples of these networks include a public switched telephone network (PSTN), a local area network (LAN), a wide area network (WAN), a personal area network (PAN), or a combination of communication networks such as the Internet. The access device 540 can include one or more of any type of network interface (e.g., network interface card (NIC)) such as an IEEE 802.11 wireless local area network (WLAN) wireless interface, a worldwide interoperability for microwave access (Wi-MAX) interface, an Ethernet interface, a universal serial bus (USB) interface, a cellular network interface, a Bluetooth interface, a near field communication (NFC) interface, etc.
[0089] In an embodiment of the present specification, the above-mentioned components of the computing device 500 and other components not shown in the Figure 5 may be connected to each other, for example, through a bus. It should be understood that Figure 5 The structural block diagram of the computing device shown is only for the purpose of example, and is not a limitation on the scope of the present specification. Other components can be added or replaced as needed by those skilled in the art.
[0090] The computing device 500 can be any type of stationary or mobile computing device, including a mobile computer or mobile computing device (e.g., a tablet computer, a personal digital assistant, a laptop computer, a notebook computer, a netbook, etc.), a mobile phone (e.g., a smartphone), a wearable computing device (e.g., a smart watch, smart glasses, etc.), or other type of mobile device, or a stationary computing device such as a desktop computer or PC. The computing device 500 can also be a mobile or stationary server.
[0091] The processor 520 is configured to execute computer-executable instructions to perform the steps of the data processing method described above.
[0092] The above is a schematic solution of the computing device of the embodiment. It should be noted that the technical solution of the computing device and the technical solution of the data processing method described above belong to the same concept, and the details of the technical solution of the computing device that are not described in detail can be referred to the description of the technical solution of the data processing method.
[0093] An embodiment of the present specification further provides a computer-readable storage medium, which stores computer-executable instructions, and the computer-executable instructions are executed by a processor to implement the steps of the data processing method described above.
[0094] The above is a schematic solution of the computer-readable storage medium of the embodiment. It should be noted that the technical solution of the storage medium and the technical solution of the data processing method described above belong to the same concept, and the details of the technical solution of the storage medium that are not described in detail can be referred to the description of the technical solution of the data processing method.
[0095] An embodiment of the present specification further provides a computer program, which causes a computer to perform the steps of the data processing method described above when the computer program is executed in the computer.
[0096] The above is a schematic solution of the computer program of the embodiment. It should be noted that the technical solution of the computer program and the technical solution of the data processing method described above belong to the same concept, and the details of the technical solution of the computer program that are not described in detail can be referred to the description of the technical solution of the data processing method.
[0097] The above-described embodiments of the application have several aspects, no single one of which is solely responsible for the application's desirable attributes. Without limiting the scope of the application as expressed by the claims which follow, some further embodiments make these aspects even more useful. Other embodiments can result in less desirable attributes.
[0098] The computer readable medium can include any entity or apparatus capable of carrying the computer program code, recording medium, U disk, mobile hard disk, magnetic disk, optical disk, computer memory, Read-Only Memory (ROM), Random Access Memory (RAM), electrical carrier signal, telecommunication signal, software distribution medium, etc. It should be noted that the computer readable medium can include appropriate contents according to the requirements of legislation and patent practice in the jurisdiction, for example, in some jurisdictions, according to legislation and patent practice, the computer readable medium does not include electrical carrier signals and telecommunication signals.
[0099] It should be noted that for the foregoing method embodiments, the acts described therein can be performed in a different order from the order described, and that some acts can be performed in parallel or concurrently. In addition, some of the acts described above can not be performed in all embodiments. Furthermore, the acts described above can be performed by different parties in some embodiments. Furthermore, each of the acts described above can be performed by specialized hardware components or modules or can be embodied in a software-specific or a generalized computing system or module. Similarly, general-purpose computing systems and modules can be configured to constitute one or more specialized components or modules described above.
[0100] In the above embodiments, the description of each embodiment is focused on different aspects, and the parts not described in detail in a certain embodiment can be referred to the relevant description of other embodiments.
[0101] The above disclosed preferred embodiments of the present application are only used to help explain the present application. Alternative embodiments do not describe all the details, nor limit the application to only the specific embodiments described. Obviously, according to the content of the embodiments of the present application, many modifications and changes can be made. The present application selects and specifically describes these embodiments in order to better explain the principles and practical applications of the embodiments of the present application, so that those skilled in the art can well understand and use the present application. The present application is limited by the claims and their entire scope and equivalents.
Claims
1. A data processing method, comprising: obtaining a sensitive data set containing initial sensitive data, and determining a data scenario corresponding to the initial sensitive data; determining a data processing algorithm corresponding to the initial sensitive data and at least two target scoring units according to the data scenario; in a case where it is determined according to the data processing algorithm that the initial sensitive data satisfies a preset processing condition, performing sensitivity scoring on the initial sensitive data according to the at least two target scoring units to obtain a target scoring result; determining target scoring results of all initial sensitive data in the sensitive data set, and determining whether the initial sensitive data is target sensitive data according to the target scoring results of all initial sensitive data in the sensitive data set.
2. The data processing method of claim 1, wherein the obtaining a sensitive data set containing initial sensitive data comprises: obtaining a sensitive data set containing initial sensitive data determined by a preset risk identification terminal.
3. The data processing method of claim 1, wherein the determining a data processing algorithm corresponding to the initial sensitive data and at least two target scoring units according to the data scenario comprises: determining a data processing algorithm corresponding to the initial sensitive data and at least two initial scoring units according to the data scenario; calculating a correlation degree of each initial scoring unit and the data scenario; adjusting a scoring strategy of each initial scoring unit according to the correlation degree, and determining each adjusted initial scoring unit as at least two target scoring units corresponding to the initial sensitive data.
4. The data processing method of claim 1, wherein the performing sensitivity scoring on the initial sensitive data according to the at least two target scoring units comprises: performing sensitivity scoring on the initial sensitive data according to a scoring dimension and a scoring strategy corresponding to each target scoring unit in the at least two target scoring units.
5. The data processing method of claim 1, wherein the performing sensitivity scoring on the initial sensitive data according to the at least two target scoring units to obtain a target scoring result comprises: performing sensitivity scoring on the initial sensitive data according to each target scoring unit in the at least two target scoring units to obtain an initial scoring result of the initial sensitive data by each target scoring unit; performing calculation on the initial scoring result of the initial sensitive data according to a preset calculation rule to obtain a target scoring result of the initial sensitive data.
6. The data processing method of claim 1, wherein the determining whether the initial sensitive data is target sensitive data according to the target scoring results of all initial sensitive data in the sensitive data set comprises: determining an average scoring result according to the target scoring results of all initial sensitive data in the sensitive data set; determining whether the initial sensitive data is target sensitive data according to the average scoring result.
7. The data processing method of claim 1, wherein determining whether the initial sensitive data is a target sensitive data based on the target score result of all initial sensitive data in the sensitive data set comprises: ranking all initial sensitive data in the sensitive data set according to the target score result of all initial sensitive data in the sensitive data set; and determining whether the initial sensitive data is a target sensitive data based on the ranking result.
8. The data processing method of any one of claims 1-7, further comprising: in a case where it is determined that the initial sensitive data has a historical score result, determining whether the initial sensitive data is a target sensitive data based on the historical score result of the initial sensitive data and the target score result.
9. The data processing method of claim 8, wherein determining whether the initial sensitive data is a target sensitive data based on the historical score result of the initial sensitive data and the target score result comprises: calculating a difference between the historical score result and the target score result based on the historical score result of the initial sensitive data and the target score result; determining that the initial sensitive data is a target sensitive data in a case where the difference is greater than a preset difference threshold; and determining that the initial sensitive data is not a target sensitive data in a case where the difference is less than or equal to the preset difference threshold.
10. A data processing apparatus, comprising: a data acquisition module configured to acquire a sensitive data set containing initial sensitive data and determine a data scenario corresponding to the initial sensitive data; an algorithm determination module configured to determine a data processing algorithm corresponding to the initial sensitive data and at least two target scoring units based on the data scenario; a scoring module configured to perform sensitivity scoring on the initial sensitive data based on the at least two target scoring units to obtain a target score result in a case where it is determined that the initial sensitive data satisfies a preset processing condition based on the data processing algorithm; and a target data determination module configured to determine a target score result of all initial sensitive data in the sensitive data set and determine whether the initial sensitive data is a target sensitive data based on the target score result of all initial sensitive data in the sensitive data set.
11. A computing device, comprising: a memory and a processor; the memory is configured to store computer executable instructions, and the processor is configured to execute the computer executable instructions, which, when executed by the processor, implement the steps of the data processing method of any one of claims 1-9.
12. A computer readable storage medium storing computer executable instructions, which, when executed by a processor, implement the steps of the data processing method of any one of claims 1-9.