A sensitive data protection business processing method, device and electronic equipment
By adopting a decision-making approach that integrates data sensitivity and the public disclosure requirements of business scenarios, the problem of a disconnect between business needs and the rigidity of enterprise sensitive data desensitization strategies has been solved, thereby improving user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- CHINA MOBILE GRP HENAN CO LTD
- Filing Date
- 2021-09-14
- Publication Date
- 2026-08-04
AI Technical Summary
In existing technologies, enterprises' rigid strategies for desensitizing sensitive data lead to a disconnect from actual business needs, affecting the user's business experience.
By comprehensively considering the sensitivity of the data itself and the disclosure requirements of the current business scenario, a decision is made on whether to perform anonymization processing for users, including the calculation of information sensitivity index and information transparency index, to determine whether it is necessary to anonymize the target content.
This avoids the disconnect between traditional desensitization strategies and business needs, and improves the business experience on the user side.
Smart Images

Figure CN115809475B_ABST
Abstract
Description
Technical Field
[0001] This document relates to the field of information security technology, and in particular to a business processing method, apparatus and electronic equipment for sensitive data protection. Background Technology
[0002] With the advent of the big data era, information security has received increasing attention. Currently, the most common information security measure used by enterprises is data anonymization, and whether or not enterprise decision-making data needs anonymization depends entirely on whether the data itself contains sensitive information. However, in practice, it has been found that users' needs for the same sensitive data vary across different business scenarios. Some business scenarios require enterprises to anonymize sensitive data before providing it to users, while in other scenarios, enterprises need to fully disclose the sensitive data to users in order to assist them with related business.
[0003] Therefore, the rigidity of existing sensitive data anonymization strategies by enterprises can lead to a disconnect from actual business needs, and in many cases, negatively impact the user experience. Thus, how to more intelligently decide whether data needs to be anonymized for users during business processes is the technical problem this application aims to solve. Summary of the Invention
[0004] The purpose of this invention is to provide a business processing method, apparatus, and electronic device for sensitive data protection, which can comprehensively consider the sensitivity of the data itself and the current business scenario's requirements for data disclosure to decide whether the data needs to be anonymized for users.
[0005] To achieve the above objectives, the embodiments of the present invention are implemented as follows:
[0006] Firstly, a business processing method for sensitive data protection is provided, including:
[0007] Receive business requests initiated by the business requester for the target business scenario;
[0008] Determine the information sensitivity index of the target content requested by the business request and the information transparency index of the target content in relation to the target business scenario, wherein the information transparency index represents the degree to which the target content is disclosed to the business requester under the business logic requirements of the target business scenario;
[0009] Based on the information sensitivity index and the information transparency index, it is determined whether the target content needs to be de-identified for the business requester in the target business scenario;
[0010] If de-identification is required, the target content will be de-identified and then provided to the requesting party; otherwise, the target content will be provided directly to the requesting party.
[0011] Secondly, a business processing device for sensitive data protection is provided, comprising:
[0012] The business request module receives business requests initiated by the business requester for the target business scenario;
[0013] The data anonymization analysis module determines the information sensitivity index of the target content requested by the business request and the information transparency index of the target content in relation to the target business scenario. The information transparency index represents the degree to which the target content is disclosed to the business requester under the business logic requirements of the target business scenario.
[0014] The data desensitization decision module determines, based on the information sensitivity index and the information transparency index, whether the target content needs to be desensitized for the business requester in the target business scenario.
[0015] If the business request feedback module requires de-identification processing, it will provide the target content to the business requester after de-identification processing; otherwise, it will provide the target content directly to the business requester.
[0016] Thirdly, an electronic device is provided, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being executed by the processor.
[0017] Receive business requests initiated by the business requester for the target business scenario;
[0018] Determine the information sensitivity index of the target content requested by the business request and the information transparency index of the target content in relation to the target business scenario, wherein the information transparency index represents the degree to which the target content is disclosed to the business requester under the business logic requirements of the target business scenario;
[0019] Based on the information sensitivity index and the information transparency index, it is determined whether the target content needs to be de-identified for the business requester in the target business scenario;
[0020] If de-identification is required, the target content will be de-identified and then provided to the requesting party; otherwise, the target content will be provided directly to the requesting party.
[0021] Fourthly, a computer-readable storage medium is provided, on which a computer program is stored, the computer program performing the following steps when executed by a processor:
[0022] Receive business requests initiated by the business requester for the target business scenario;
[0023] Determine the information sensitivity index of the target content requested by the business request and the information transparency index of the target content in relation to the target business scenario, wherein the information transparency index represents the degree to which the target content is disclosed to the business requester under the business logic requirements of the target business scenario;
[0024] Based on the information sensitivity index and the information transparency index, it is determined whether the target content needs to be de-identified for the business requester in the target business scenario;
[0025] If de-identification is required, the target content will be de-identified and then provided to the requesting party; otherwise, the target content will be provided directly to the requesting party.
[0026] When the solution of this invention receives a business request from a business requester regarding a target business scenario, if the business request requires feedback of target content to the requester, it can decide whether to perform anonymization processing on the target content by comprehensively considering the sensitivity of the target content itself and the degree of public disclosure of the target content to the business requester within the target business scenario. If anonymization processing is required, the target content is provided to the business requester after anonymization; otherwise, the target content is provided directly to the business requester. Clearly, this solution avoids the problem of disconnect between business needs and the rigid anonymization strategies of traditional enterprises, and can improve the business experience for the requester to a certain extent. Attached Figure Description
[0027] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0028] Figure 1 A flowchart illustrating the business processing method for sensitive data protection provided in this embodiment of the invention.
[0029] Figure 2 A schematic diagram of the structure of a business processing device for sensitive data protection provided in an embodiment of the present invention.
[0030] Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation
[0031] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this specification, and not all embodiments. Based on the embodiments in this specification, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of this specification.
[0032] As mentioned earlier, the most common information security measure currently used by enterprises is data anonymization. Whether or not enterprise decision-making data needs anonymization depends entirely on whether the data itself contains sensitive information. However, in practical applications, users' needs for the same sensitive data vary across different business scenarios. For example, when a user inquires about previously processed business information at a bank, the associated bank card number is sensitive data and will be anonymized before being provided to the user. However, when processing new business transactions, users expect to obtain the fully public bank card number, not the anonymized one. Clearly, the rigidity of existing enterprise sensitive data anonymization strategies can lead to a disconnect from actual business needs, and in many cases, negatively impact the user experience. Therefore, this paper aims to provide a more intelligent business processing solution for sensitive data protection, capable of comprehensively considering the sensitivity of the data itself and the current business scenario's requirements for data disclosure, to determine whether data needs to be anonymized for users.
[0033] Figure 1 This is a flowchart of a business processing method for sensitive data protection according to an embodiment of the present invention, including the following steps:
[0034] S102, Receive the business request initiated by the business requester for the target business scenario.
[0035] This article does not specify particular business scenarios. It should be understood that, based on business requests, enterprises may need to provide some data back to the requesting party. If the data provided contains sensitive information, a decision needs to be made in the subsequent process regarding whether to anonymize the sensitive data before sending it to the requesting party.
[0036] S104, determine the information sensitivity index of the target content requested by the business request and the information transparency index of the target content in relation to the target business scenario. The information transparency index represents the degree to which the target content is disclosed to the business requester under the business logic requirements of the target business scenario.
[0037] Regarding the information transparency index of target content relative to the target business scenario, embodiments of the present invention can determine it based on the security level pre-set for the target business scenario and the relevance of the target content to the target business scenario.
[0038] Among these, the security level of the target business scenario is considered prior knowledge, reflecting the security of the information involved. A higher security level indicates more secure information, corresponding to a lower need for data anonymization during the anonymization decision-making process. The relevance of the target content to the target business scenario reflects the dependence of the target business scenario's business logic on the target content. In this embodiment, the relevance table of the target content to the target business scenario is negatively correlated with the corresponding information transparency index. That is, the less the target business scenario depends on the target content, the less related the target content is to the target business scenario, resulting in a higher need for data anonymization during the anonymization decision-making process, to avoid exposing irrelevant data to the business requester.
[0039] Furthermore, the information sensitivity index of the target content can be determined based on the sensitivity of the target content to the business requester in this embodiment of the invention.
[0040] Specifically, enterprises can set different permissions for different business requesters. Higher permissions allow access to more sensitive business data. In this embodiment of the invention, if the sensitivity level of the target content exceeds the permissions of the business requester, it corresponds to a higher information sensitivity index, and thus a higher de-identification requirement during the de-identification decision-making process. Conversely, if the sensitivity level of the target content does not exceed the permissions of the business requester, it corresponds to a lower information sensitivity index, and thus a lower de-identification requirement during the de-identification decision-making process.
[0041] S106. Based on the information sensitivity index and information transparency index, determine whether the target content needs to be de-identified for the business requester in the target business scenario.
[0042] Specifically, this step uses the information sensitivity index and information transparency index as factors to calculate the anonymization requirement value for the target content. If the anonymization requirement value reaches a preset threshold, then the target content is determined to require anonymization for the querying party in the target business scenario. It should be understood that the calculation method for the anonymization requirement value is not unique, and this article does not impose a specific limitation on it.
[0043] S108 If de-identification is required, the target content shall be de-identified and provided to the business requester; otherwise, the target content shall be provided directly to the business requester.
[0044] This document does not specify the exact method of data anonymization. As an example, this step can mask some information in the target content to achieve data anonymization; alternatively, asymmetric encryption technology can be used to encrypt the target content to achieve the same effect. In the former method, the masked information cannot be exposed to the requesting party. In the latter method, the requesting party can apply for and obtain the asymmetric encryption key, and after approval, view the target content based on the obtained key.
[0045] Based on the above, it can be seen that when the method of this embodiment receives a business request initiated by a business requester for a target business scenario, if the business request requires feedback of target content to the business requester, it can decide whether to perform de-identification processing on the target content by comprehensively considering the sensitivity of the target content itself and the degree of public disclosure of the target content to the business requester in the target business scenario. If de-identification processing is required, the target content is provided to the business requester after de-identification processing; otherwise, the target content is provided directly to the business requester. Obviously, the entire solution avoids the problem of disconnection from business needs caused by the rigidity of sensitive data de-identification strategies in traditional enterprises, and can improve the business experience of the business requester to a certain extent.
[0046] The following section uses an application scenario in the field of mobile communications as an example to provide a detailed description of the desensitization decision-making process of the method in this embodiment.
[0047] In this application scenario, the business request specifically refers to a query request initiated by a mobile user (business request method), and the content to be queried corresponding to the query request is the target content described above. Correspondingly, the process for handling the query request includes:
[0048] Step 1: Obtain the SQL query statement initiated by the business requester (the business requester), perform in-depth analysis of the query statement, and determine the target content requested by the business requester and the target business scenario to which it belongs.
[0049] Step 2: Determine the information transparency index of the target content for the target business scenario.
[0050] As mentioned earlier, the information transparency index of the target content relative to the target business scenario can be determined based on the security level pre-set for the target business scenario and the relevance of the target content to the target business scenario.
[0051] The business scenarios can be categorized based on the environment of the requesting terminal. Examples include mobile service hall scenarios and VPN private network scenarios. The security level of a business scenario is a pre-set empirical value; for example, the security level of a mobile service hall scenario is 1, and the security level of a VPN private network scenario is 5. This application scenario can determine the security level based on the IP address or MAC address of the requesting terminal. For example, if the IP address indicates the requesting terminal is in a mobile service hall, the security level is set to 1; if the IP address indicates the requesting terminal is in a VPN private network, the security level is set to 5.
[0052] It should be understood that the level of security is a pre-set empirical value, and the method of measurement is not unique. This article will not provide examples or elaborate on them one by one.
[0053] The relevance of target content to the target business scenario reflects, to some extent, the dependence of that business scenario on the target content. For example, in a service hall, since customers frequently inquire about a user's service plan details, a query for plan details is considered highly relevant, while a query for a user's ID number is considered less relevant. High relevance means the current scenario has a high demand for the query information. If this information were anonymized, it would disrupt the business process. Therefore, queries with high relevance are considered to have a lower security requirement in the current scenario. Similarly, low relevance means the current scenario has a low demand for the query information, or even shouldn't query it at all. If this information isn't anonymized, it could lead to data insecurity. Therefore, queries with low relevance are considered to have a higher security requirement in the current scenario.
[0054] Specifically, the relevance of the target content to the target business scenario can be determined based on the relevance of the data table to which the target content belongs to to the business scenario.
[0055] Here, it is assumed that this application scenario has pre-configured query permissions for the data table for the business requester, where:
[0056] The relevance of any first target data table to be queried that does not exceed the query permissions to the business scenario = the security level of the target business scenario × max{the connection between the first target data table to be queried and other data tables to be queried} × the total number of data tables to be queried that have connection relationships with other data tables to be queried / (the total number of all data tables to be queried - 1).
[0057] The degree of relevance of any second target data table that exceeds the query permissions and has a connection relationship with other data tables to be queried relative to the business scenario equals the security level of the target business scenario.
[0058] The relevance of any third target data table to be queried that exceeds the query permissions and has no connection relationship with other data tables to be queried to the business scenario is calculated as follows: Security level of the target business scenario × min{number of sensitive data items in each data table to be queried for the query permissions} / ∑{number of sensitive data items in each data table to be queried for the query permissions}. The data tables are pre-set with corresponding sensitive data items for different query permissions.
[0059] The join degree between the tables to be queried is determined based on the join relationships between them. These join relationships include a first join based on foreign keys and a second join based on indirect connections with other tables (which may be tables not to be queried). It should be understood that the join degree between tables joined based on the first join is greater than the join degree between tables joined based on the second join, and the join degree of tables not joined with other tables is 0.
[0060] For example, we can think of a data table as a point. If two data tables are directly connected by a foreign key, then an edge connects the two points. The connectivity represents the vulnerability between the two data tables. When the points of the two data tables are directly connected, the vulnerability is the number of edges. When the points of the two data tables are not directly connected, the vulnerability is the product of the number of edges and the security level. In other words, if data table A and data table B are directly connected via a foreign key, then there is an edge connecting the corresponding point in data table A to the corresponding point in data table B, so the connectivity degree is 1. If data table A and data table C are connected via m other data tables, then there will be m more points in other data tables between the corresponding points in data table A and the corresponding points in data table C, and these m+2 points are connected by a straight line. The number of edges between the m+2 points is m-1. However, more edges mean a longer link and reduced security. Therefore, in this case, the connectivity degree will also include security-related parameters. Thus, if data table A and data table C are connected via m other data tables, then the connectivity degree between data table A and data table C is security degree * (m-1).
[0061] Step 3: Determine the information sensitivity index of the target content.
[0062] Specifically, the information sensitivity index of the target content can be determined based on the weight coefficient of the data table to which the target content belongs and the relevance of the target content to the query statement.
[0063] The weight coefficient of a data table reflects its importance within the target business scenario. If at least two data tables to be queried are connected through a first or second join relationship, the data table joined earlier in the query order has a higher weight coefficient than the data table joined later in the query order. Obviously, the later the data table is joined, the weaker its relevance to the business scenario, and therefore its corresponding weight coefficient is lower. When a data table to be queried has multiple join relationships, meaning its weight coefficient is not unique, the maximum weight coefficient is taken.
[0064] The weighting coefficients of the data tables can be determined through preset rules, such as:
[0065] ① If there is only one data table being queried, then the weight coefficient of that data table is 1.2.
[0066] ②If there are multiple data tables for each query, then:
[0067] If there are two tables to be queried connected by a first join relationship, such as table 1 and table 2, where the primary key of table 1 is a foreign key of table 2 (i.e., table 1 is ordered before table 2 in the join relationship), then the weight coefficient for table 1 is determined to be 1.2, and the weight coefficient for table 2 is determined to be 0.8. Similarly, if there is a table 3 after table 2 in the join relationship, then the weight coefficient for table 3 is 0.6. Here, the minimum weight coefficient is set to 0.2 to avoid 0 or negative values.
[0068] If there are two tables to be queried connected by a second join relationship, such as table 5 and table 6, and table 5 is connected to table 6 through n other tables, then the weight coefficient of table 5 is 1.2. 1 / (n+1) The weight coefficient for the data in table 6 to be queried is 0.8. 1 / (n+1) .
[0069] Furthermore, the relevance of the target content to the query statement can be calculated using a clustering algorithm.
[0070] For example, in this application scenario, clustering algorithms can be used to classify each data table according to the access address used by the query statement. Here, the clustering algorithm calculates the mathematical distance from the data table to the cluster center based on the access address corresponding to the data table in the query statement.
[0071] After determining the data table to be queried, the data distance between the data table and the corresponding cluster centers of each category is calculated based on the clustering algorithm described above. This data distance is then used to quantify the relevance of the target content to the query statement. Specifically:
[0072] If there is only one data table to be queried, then the relevance of the data table to be queried relative to the query statement is = (average mathematical distance of the data table to be queried in the same category to the center of its category / farthest mathematical distance of the data table to be queried in the same category to the center of its category) × (total number of data tables to be queried in the same category / total number of data tables).
[0073] If there is more than one data table to be queried, then the relevance of the data table to be queried relative to the query statement = the standard deviation of the mathematical distance between the data table to be queried in the same category and the center of its category × (the number of data tables to be queried in the same category / the total number of data tables).
[0074] After determining the weight coefficient of the data table to which the target content belongs and the relevance of the target content to the query statement, the information sensitivity index of the target content can be calculated using the following formula.
[0075] The information sensitivity index of the target content = K × weight coefficient of the data table to which the target content belongs × relevance of the target content to the query statement. Wherein, if the data table to which the target content belongs does not exceed the query permissions of the business requester, then K is 1.2; otherwise, K is 0.2.
[0076] Step 4: Calculate the anonymization requirement value for the target content to determine whether the target content needs to be anonymized for the business requester in the target business scenario.
[0077] Specifically, the desensitization requirement value for target content = information sensitivity index of target content × information transparency index of target content × information weight of target content, where the information weight of target content is prior knowledge and can be taken as 1 if not used.
[0078] If the requirement for desensitization of the target content exceeds a preset threshold, the queried information is determined to be sensitive information, and desensitization is then performed.
[0079] The above application scenarios are exemplary descriptions of the methods in the embodiments of the present invention. It should be understood that appropriate changes can be made without departing from the principles described above, and these changes should also be considered within the scope of protection of the embodiments of the present invention.
[0080] In addition, corresponding to Figure 1 In addition to the query method shown, this embodiment of the invention also provides a business processing device for customer information. Figure 2 This is a schematic diagram of the structure of a service processing device 200 for sensitive data protection according to an embodiment of the present invention, including:
[0081] The business request module 210 receives business requests initiated by the business requester for the target business scenario.
[0082] The data anonymization analysis module 220 determines the information sensitivity index of the target content requested by the business request and the information transparency index of the target content in relation to the target business scenario. The information transparency index represents the degree to which the target content is disclosed to the business requester under the business logic requirements of the target business scenario.
[0083] The data desensitization decision module 230 determines, based on the information sensitivity index and the information transparency index, whether the target content needs to be desensitized for the business requester in the target business scenario.
[0084] If the business request feedback module 240 requires de-identification processing, it will provide the target content to the business requester after de-identification processing; otherwise, it will provide the target content directly to the business requester.
[0085] When the device in this embodiment of the invention receives a business request from a business requester regarding a target business scenario, if the business request requires feedback of target content to the business requester, it can decide whether to perform anonymization processing on the target content by comprehensively considering the sensitivity of the target content itself and the degree of public disclosure of the target content to the business requester within the target business scenario. If anonymization processing is required, the target content is provided to the business requester after anonymization processing; otherwise, the target content is provided directly to the business requester. Clearly, this entire solution avoids the problem of disconnect between business needs and the rigid anonymization strategies of traditional enterprises, which can improve the business experience for business requesters to a certain extent.
[0086] Optionally, the data anonymization analysis module 220 is specifically used to: determine the information transparency index of the target business scenario to the target content based on the security level pre-set for the target business scenario and the relevance of the target content to the target business scenario, wherein the relevance of the target content to the target business scenario is negatively correlated with the corresponding information transparency index.
[0087] Optionally, the business request is a query request, the target content is the content to be queried corresponding to the query request, and the business requester is pre-configured with query permissions for the data table; the relevance of the target content to the target business scenario is determined based on the relevance of the data table to which the target content belongs to the business scenario, wherein: the relevance of any first target data table to be queried that does not exceed the query permission to the business scenario = the security level of the target business scenario × max{the connection degree between the first target data table to be queried and other data tables to be queried} × the total number of data tables to be queried that have a connection relationship with other data tables to be queried / (the total number of all data tables to be queried - 1), and the connection degree between the data tables to be queried is based on the connection degree between the target content and the target business scenario. The connection degree of a data table that is not connected to other data tables is 0, determined by the connection relationships between the query data tables. The relevance of any second target data table that exceeds the query permission and is connected to other data tables to the business scenario is equal to the security degree of the target business scenario. The relevance of any third target data table that exceeds the query permission and is not connected to other data tables to the business scenario is equal to the security degree of the target business scenario × min{number of sensitive data items in each data table for the query permission} / ∑{number of sensitive data items in each data table for the query permission}, where the data tables are pre-set with corresponding sensitive data items for different query permissions.
[0088] Optionally, the join relationship between the data tables to be queried includes a first join relationship based on foreign key join and a second join relationship based on indirect join of other data tables, wherein the join degree between the data tables to be queried based on the first join relationship is greater than the join degree between the data tables to be queried based on the second join relationship.
[0089] Optionally, the business request also carries a query statement for the target content; the data anonymization analysis module 220 is specifically used to: determine the information sensitivity index of the target content based on the weight coefficient of the data table to which the target content belongs and the relevance of the target content to the query statement; wherein, if there are at least two data tables to be queried connected through the first connection relationship or the second connection relationship, the weight coefficient of the data table to be queried that is connected first is greater than that of the data table to be queried that is connected later, and the maximum value is taken when the weight coefficient of each data table to be queried is not unique.
[0090] Optionally, the relevance of the target content to the query statement is determined based on the relevance of the target content to the query statement to the data table to be queried; wherein: if there is only one data table to be queried, then the relevance of the data table to be queried to the query statement = (average mathematical distance of the data table to be queried for the same category to the center of its category / farthest mathematical distance of the data table to be queried for the same category to the center of its category) × (total number of data tables to be queried for the same category / total number of data tables), wherein the category of the data table to be queried is determined based on the classification of multiple data tables using a clustering algorithm, and the clustering algorithm calculates the mathematical distance of the data table to the cluster center based on the access address of the data table in the query statement; if there is more than one data table to be queried, then the relevance of the data table to be queried to the query statement = standard deviation of the mathematical distance of the data table to be queried for the same category to the center of its category × (number of data tables to be queried for the same category / total number of data tables).
[0091] Optionally, the data anonymization decision module 230 is specifically used to: calculate the anonymization processing requirement value of the target content using the information sensitivity index and the information transparency index as factors; if the anonymization processing requirement value reaches a preset threshold, then determine that the target content needs to be anonymized for the querying party in the target business scenario.
[0092] Obviously, the embodiments of the present invention Figure 2 The business processing device shown can achieve the above. Figure 1 The steps and functions of the method shown are explained below. Since the principle is the same, they will not be repeated here.
[0093] Figure 3 This is a schematic diagram of the structure of an electronic device according to one embodiment of this specification. Please refer to it. Figure 3 At the hardware level, the electronic device includes a processor, and optionally also includes an internal bus, a network interface, and memory. The memory may include main memory, such as high-speed random-access memory (RAM), or non-volatile memory, such as at least one disk drive. Of course, the electronic device may also include other hardware required for other business operations.
[0094] The processor, network interface, and memory can be interconnected via an internal bus, which can be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus, or an EISA (Extended Industry Standard Architecture) bus, etc. This bus can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 3 The symbol is represented by a single double-headed arrow, but this does not mean that there is only one bus or one type of bus.
[0095] Memory is used to store programs. Specifically, the program may include program code, which includes computer operation instructions. Memory may include main memory and non-volatile memory, and provides instructions and data to the processor. The processor reads the corresponding computer program from the non-volatile memory into main memory and then runs it, forming a business processing device at the logical level. Correspondingly, the processor executes the program stored in the memory and specifically performs the following operations:
[0096] Receive business requests initiated by the business requester for the target business scenario.
[0097] Determine the information sensitivity index of the target content requested by the business request and the information transparency index of the target content in relation to the target business scenario, wherein the information transparency index represents the degree to which the target content is disclosed to the business requester under the business logic requirements of the target business scenario.
[0098] Based on the information sensitivity index and the information transparency index, it is determined whether the target content needs to be de-identified for the business requester in the target business scenario.
[0099] If de-identification is required, the target content will be de-identified and then provided to the requesting party; otherwise, the target content will be provided directly to the requesting party.
[0100] When the electronic device in this embodiment of the invention receives a business request from a business requester regarding a target business scenario, if the business request requires feedback of target content to the business requester, it can decide whether to perform anonymization processing on the target content by comprehensively considering the sensitivity of the target content itself and the degree of public disclosure of the target content to the business requester within the target business scenario. If anonymization processing is required, the target content is provided to the business requester after anonymization processing; otherwise, the target content is provided directly to the business requester. Clearly, this entire solution avoids the problem of disconnect between business needs and the rigid anonymization strategies of traditional enterprises, and can improve the business experience for business requesters to a certain extent.
[0101] The above is as described in this instruction manual. Figure 1 The query method disclosed in the illustrated embodiments can be applied to or implemented by a processor. The processor may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by integrated logic circuits in the processor's hardware or by instructions in software form. The processor can be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it can also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this invention. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this invention can be directly manifested as execution by a hardware decoding processor, or execution by a combination of hardware and software modules in the decoding processor. The software module can reside in a mature storage medium in the field, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, or registers. This storage medium is located in memory, and the processor reads information from the memory and, in conjunction with its hardware, completes the steps of the above method.
[0102] It should be understood that the electronic device in the embodiments of the present invention can enable the service processing device to implement corresponding Figure 1 The steps and functions of the method shown are not repeated here as the principle is the same.
[0103] Of course, in addition to software implementation, the electronic device described in this specification does not exclude other implementation methods, such as logic devices or a combination of hardware and software. In other words, the execution subject of the following processing flow is not limited to each logic unit, but can also be hardware or logic devices.
[0104] Furthermore, embodiments of the present invention also provide a computer-readable storage medium that stores one or more programs, the one or more programs including instructions.
[0105] When the aforementioned instructions are executed by a portable electronic device that includes multiple applications, they enable the portable electronic device to perform... Figure 1 The steps of the query method shown include:
[0106] Receive business requests initiated by the business requester for the target business scenario.
[0107] Determine the information sensitivity index of the target content requested by the business request and the information transparency index of the target content in relation to the target business scenario, wherein the information transparency index represents the degree to which the target content is disclosed to the business requester under the business logic requirements of the target business scenario.
[0108] Based on the information sensitivity index and the information transparency index, it is determined whether the target content needs to be de-identified for the business requester in the target business scenario.
[0109] If de-identification is required, the target content will be de-identified and then provided to the requesting party; otherwise, the target content will be provided directly to the requesting party.
[0110] Those skilled in the art will understand that the embodiments of this specification can be provided as methods, systems, or computer program products. Therefore, this specification may take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this specification may take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0111] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.
[0112] The above are merely embodiments of this specification and are not intended to limit the scope of this specification. Various modifications and variations can be made to this specification by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this specification should be included within the scope of the claims of this specification. Furthermore, all other embodiments obtained by those skilled in the art without inventive effort should fall within the protection scope of this document.
Claims
1. A method of business processing for sensitive data protection, characterized in that, include: Receive business requests initiated by the business requester for the target business scenario; Determine the information sensitivity index of the target content requested by the business request and the information transparency index of the target content relative to the target business scenario. Determining the information transparency index of the target content relative to the target business scenario includes: determining the information transparency index of the target content relative to the target business scenario based on a pre-set security level for the target business scenario and the relevance of the target content to the target business scenario. The information transparency index represents the degree to which the target content is publicly disclosed to the business requester under the business logic requirements of the target business scenario. Based on the information sensitivity index and the information transparency index, it is determined whether the target content needs to be de-identified for the business requester in the target business scenario; If de-identification is required, the target content will be de-identified and then provided to the business requester; otherwise, the target content will be provided directly to the business requester. Wherein, the business request is a query request, the target content is the content to be queried corresponding to the query request, and the business requester is pre-configured with query permissions for the data table; the relevance of the target content to the target business scenario is determined based on the relevance of the data table to which the target content belongs to the business scenario; the relevance of any first target data table to be queried that does not exceed the query permission to the business scenario is determined based on the security of the target business scenario, the connectivity between the first target data table to be queried and other data tables to be queried, and the total number of data tables to be queried that have a connection relationship with the first target data table to be queried and other data tables to be queried.
2. The method according to claim 1, characterized in that, The correlation table of the target content with the target business scenario is negatively correlated with the corresponding information transparency index.
3. The method according to claim 2, characterized in that, The relevance of any first target data table to be queried that does not exceed the query permission to the business scenario is = the security level of the target business scenario × max{the connection degree between the first target data table to be queried and other data tables to be queried} × the total number of data tables to be queried that have a connection relationship with other data tables to be queried / (the total number of all data tables to be queried - 1). The connection degree between data tables to be queried is determined based on the connection relationship between data tables to be queried. The connection degree of a data table to be queried that does not have a connection relationship with other data tables to be queried is 0. The relevance of any second target data table that exceeds the query permissions and has a connection relationship with other data tables to be queried to the business scenario is equal to the security level of the target business scenario. The relevance of any third target data table to be queried that exceeds the query permission and has no connection relationship with other data tables to be queried to the business scenario is = the security level of the target business scenario × min{the number of sensitive data items in each data table to be queried for the query permission} / ∑{the number of sensitive data items in each data table to be queried for the query permission}, where the data tables are pre-set with corresponding sensitive data items for different query permissions.
4. The method according to claim 3, characterized in that, The join relationships between the tables to be queried include a first join relationship based on foreign keys and a second join relationship based on indirect joins with other tables. The join degree between the tables to be queried based on the first join relationship is greater than the join degree between the tables to be queried based on the second join relationship.
5. The method according to claim 4, characterized in that, The service request also carries a query statement for the target content; Determining the information sensitivity index of the target content includes: Based on the weight coefficient of the data table to which the target content belongs and the relevance of the target content to the query statement, the information sensitivity index of the target content is determined; If there are at least two data tables to be queried connected through the first connection relationship or the second connection relationship, the weight coefficient of the data table to be queried that is connected first is greater than that of the data table to be queried that is connected later, and the maximum value is taken when the weight coefficient of each data table to be queried is not unique.
6. The method according to claim 5, characterized in that, The relevance of the target content to the query statement is determined based on the relevance of the target content to the query table to the query statement; wherein: If there is only one data table to be queried, then the relevance of the data table to be queried relative to the query statement is = (average mathematical distance of the data table to be queried in the same category to the center of its category / farthest mathematical distance of the data table to be queried in the same category to the center of its category) × (total number of data tables to be queried in the same category / total number of data tables). The category of the data table to be queried is determined by classifying multiple data tables based on a clustering algorithm. The clustering algorithm calculates the mathematical distance from the data table to the cluster center based on the access address of the data table in the query statement. If there is more than one data table to be queried, then the relevance of the data table to be queried relative to the query statement is equal to the standard deviation of the mathematical distance between the data table to be queried in the same category and the center of its category, multiplied by (the number of data tables to be queried in the same category / the total number of data tables).
7. The method according to claim 1, characterized in that, Based on the information sensitivity index and the information transparency index, determining whether the target content needs to be anonymized for the business requester in the target business scenario includes: Using the information sensitivity index and the information transparency index as factors, the desensitization processing requirement value of the target content is calculated; If the required de-identification value reaches a preset threshold, then it is determined that the target content needs to be de-identified for the querying party in the target business scenario.
8. A service processing apparatus for sensitive data protection, characterized by, include: The business request module is used to receive business requests initiated by the business requester for the target business scenario; The data anonymization analysis module is used to determine the information sensitivity index of the target content requested by the business request and the information transparency index of the target content relative to the target business scenario. Determining the information transparency index of the target content relative to the target business scenario includes: determining the information transparency index of the target content relative to the target business scenario based on a pre-set security level for the target business scenario and the relevance of the target content to the target business scenario; the information transparency index represents the degree to which the target content is publicly disclosed to the business requester under the business logic requirements of the target business scenario. The data desensitization decision module is used to determine, based on the information sensitivity index and the information transparency index, whether the target content needs to be desensitized for the business requester in the target business scenario. If the business request feedback module requires de-identification processing, it will provide the target content to the business requester after de-identification processing; otherwise, it will provide the target content directly to the business requester. Wherein, the business request is a query request, the target content is the content to be queried corresponding to the query request, and the business requester is pre-configured with query permissions for the data table; the relevance of the target content to the target business scenario is determined based on the relevance of the data table to which the target content belongs to the business scenario; the relevance of any first target data table to be queried that does not exceed the query permission to the business scenario is determined based on the security of the target business scenario, the connectivity between the first target data table to be queried and other data tables to be queried, and the total number of data tables to be queried that have a connection relationship with the first target data table to be queried and other data tables to be queried.
9. An electronic device comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that the computer program is executed by the processor: Receive business requests initiated by the business requester for the target business scenario; Determine the information sensitivity index of the target content requested by the business request and the information transparency index of the target content relative to the target business scenario. Determining the information transparency index of the target content relative to the target business scenario includes: determining the information transparency index of the target content relative to the target business scenario based on a pre-set security level for the target business scenario and the relevance of the target content to the target business scenario. The information transparency index represents the degree to which the target content is publicly disclosed to the business requester under the business logic requirements of the target business scenario. Based on the information sensitivity index and the information transparency index, it is determined whether the target content needs to be de-identified for the business requester in the target business scenario; If de-identification is required, the target content will be de-identified and then provided to the business requester; otherwise, the target content will be provided directly to the business requester. Wherein, the business request is a query request, the target content is the content to be queried corresponding to the query request, and the business requester is pre-configured with query permissions for the data table; the relevance of the target content to the target business scenario is determined based on the relevance of the data table to which the target content belongs to the business scenario; the relevance of any first target data table to be queried that does not exceed the query permission to the business scenario is determined based on the security of the target business scenario, the connectivity between the first target data table to be queried and other data tables to be queried, and the total number of data tables to be queried that have a connection relationship with the first target data table to be queried and other data tables to be queried.
10. A computer-readable storage medium having stored thereon a computer program, characterized in that When the computer program is executed by the processor, it performs the following steps: Receive business requests initiated by the business requester for the target business scenario; Determine the information sensitivity index of the target content requested by the business request and the information transparency index of the target content relative to the target business scenario. Determining the information transparency index of the target content relative to the target business scenario includes: determining the information transparency index of the target content relative to the target business scenario based on a pre-set security level for the target business scenario and the relevance of the target content to the target business scenario. The information transparency index represents the degree to which the target content is publicly disclosed to the business requester under the business logic requirements of the target business scenario. Based on the information sensitivity index and the information transparency index, it is determined whether the target content needs to be de-identified for the business requester in the target business scenario; If de-identification is required, the target content will be de-identified and then provided to the business requester; otherwise, the target content will be provided directly to the business requester. Wherein, the business request is a query request, the target content is the content to be queried corresponding to the query request, and the business requester is pre-configured with query permissions for the data table; the relevance of the target content to the target business scenario is determined based on the relevance of the data table to which the target content belongs to the business scenario; the relevance of any first target data table to be queried that does not exceed the query permission to the business scenario is determined based on the security of the target business scenario, the connectivity between the first target data table to be queried and other data tables to be queried, and the total number of data tables to be queried that have a connection relationship with the first target data table to be queried and other data tables to be queried.