An Adaptive Data Available but Invisible Desensitization System
Through adaptive data, invisible desensitization systems can be used to classify and cluster market entity information, and generate symbol-substituted character sets, solving the privacy leakage risks and credit product accuracy problems when financial institutions acquire data, and achieving more accurate and locally applicable data transmission and use.
Patent Information
- Application Number
- CN202311339262.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-10-17
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2043-10-17
AI Technical Summary
When obtaining information about market entities provided by government departments, financial institutions face the risk of privacy leakage caused by data sensitivity, and there is room for improvement in credit products in terms of local characteristics and credit accuracy.
Adaptive data can be used to use invisible desensitization systems. By classifying and clustering market entity information data, and using improved K-MEANS clustering algorithms and threshold deviation correction models, symbolic substitution character sets are generated, allowing financial institutions to safely obtain more accurate and locally applicable data.
It realizes that on the premise of ensuring data security, more complete market entity data is transmitted and used to financial institutions, making credit products more accurate and locally applicable, and reducing the risk of data privacy leakage.
Smart Images

Figure CN117421746B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and specifically to an adaptive data available but invisible desensitization system. Background Art
[0002] Credit products refer to various forms of borrowing services provided by financial institutions to customers, including loans, bills, trade financing, supply chain finance, etc.
[0003] The main information requirements of financial institutions for government departments are as follows:
[0004] Financial institutions need government departments to provide credit information of small, medium and micro enterprises, including information on tax payment, social insurance premiums and housing provident fund payment, water, electricity and gas, real estate, etc., in order to conduct credit evaluation, risk monitoring and loan support for small, medium and micro enterprises;
[0005] Financial institutions need government departments to provide credit information of new citizens, including information on entrepreneurship, employment, housing, education, medical care, pension, etc., in order to provide professional and diversified financial services for new citizens;
[0006] Financial institutions need government departments to provide transaction relationship information between upstream and downstream enterprises in the supply chain, so as to build an integrated, digital and intelligent information system, credit evaluation and risk management system for upstream and downstream enterprises relying on the core enterprise, and provide comprehensive solutions such as settlement, financing and financial management for small, medium and micro enterprises in the supply chain.
[0007] With the rich and detailed development of data application scenarios, the conflict between the security guarantee and value realization of various types of data has become increasingly obvious. The credit products of financial institutions have always had a strong demand for the market entity information collected by government departments. However, due to security considerations, government departments have been unable to effectively transmit data to financial institutions and make the data generate value.
[0008] Chinese Patent Application No. 201610338383.2 discloses a data desensitization method and system. This application physically separates and stores the encryption key and the encrypted desensitized data in a local area network environment, and sets strict access permissions for the encryption key and the desensitized data, so as to effectively ensure the security of data encryption or decryption.
[0009] When a financial institution makes a data request, the government department releases the data and obtains detailed data through desensitization. These data are usually the privacy data of relevant personnel. During the data request process of financial institutions, it is easy to lead to the leakage of the privacy of relevant personnel. Most credit products require local data support to be more accurate and applicable. Due to the sensitivity of the data, they are still unified products from top to bottom in each financial institution, and there is still room for improvement in the local characteristics and credit granting accuracy of various credit products. Summary of the Invention
[0010] One of the purposes of the present invention is to provide an adaptive data available but invisible desensitization system. By classifying the market entity information data, financial institutions can obtain more complete market entity data, making credit products more accurate and more applicable locally, and realizing the value of data while ensuring data security.
[0011] To achieve the above purposes, the present invention is realized through the following technical solutions: An adaptive data available but invisible desensitization system, including:
[0012] A client, used to collect data and provide services to financial institutions. The client is divided into two parts, located in the financial institution and the government department that collects the market entity information required by the financial institution respectively. The client used by the financial institution is the secondary end, and the client used by the government department is the primary end;
[0013] A server, equipped with a clustering algorithm device for specifying the K value and the center point value. The clustering algorithm processes the data transmitted by the client. The client classifies the data into multiple types of data α according to the types. Each type of data α contains multiple pieces of data. An interaction of type α data is formed between the server and the client;
[0014] The server receives the data set and the expected center point value transmitted by the client, performs clustering operations, and divides the data set transmitted by the client into K set intervals and saves the start and end values of each segment and the corresponding character substitution values;
[0015] A threshold correction model, which sets a deviation threshold. When the difference between the expected result E requested by the secondary end of the client and the actual result C exceeds the set threshold or the average expected result E of the cumulative N times of data application and the actual result C reach the set threshold, the server performs re-clustering of the center point and symbol substitution conversion on this type of data α;
[0016] A credit granting model, which accesses the secondary end of the client and receives the data transmitted by the secondary end of the client, and establishes a credit granting model according to the data transmitted by the secondary end of the client.
[0017] In one or more embodiments of the present invention, the clustering algorithm device is an improved K-MEANS. Among them, the class data α contains the expected interval value K and the center point value X, which are randomly generated when there are no K and N values. K is the number of set intervals in the class data α and X is the number of midpoint values. The values are sorted in numerical sequence, and K is generated without duplicate numerical sequences.
[0018] In one or more embodiments of the present invention, the above class data α contains multiple set intervals , and the set interval contains multiple related data, that is , each related data is identified with a separate sequence, the number of digits T in the data is counted, and the data sequence length is selected according to the number of digits T, that is:
[0019] S 位数 = 3 / T;
[0020] Among them, S 位数 is the sequence length. When selecting, S 位数 is an integer, and when the length number of digits is less than 3, s 位数 the data is the position of the data point selected by the sequence, that is, the <S 位数 |1*s 位数 |2*S 位数 > in the related data, three data points; S 位数 is an integer, and when the length number of digits is greater than 3, the first three digits of S 位数 are the positions of the data points selected by the sequence; when S 位数 is a decimal, and when the length position after the decimal point is less than 3, the value after the decimal point of S 位数 is the position of the data point selected by the sequence, that is, the <S 位数 value after the decimal point|1*S 位数 value after the decimal point|2*S 位数 value after the decimal point> in the related data; S 位数 is a decimal, and when the length number of digits after the decimal point is greater than 3, the first three digits after the decimal point of S 位数 are the positions of the data points selected by the sequence.
[0021] In one or more embodiments of the present invention, the multiple related data contained in the above set interval are all circular data. When the number of digits required to generate the sequence > the length of the related data itself, the related data is circular data, and the first digit of the related data is the sequence number of the last digit of the related data, that is: d = q*d 数据长度 +1 to complete the data closed-loop and generate a complete sequence, where d is the serial number of the first digit of the data, q is the number of times the related data is selected, and d 数据长度 is the length of the related data.
[0022] In one or more embodiments of the present invention, when the above sequences are the same, the relevant sequence selection order is adopted for modification, that is: when the sequence 1 = sequence n the sequence 1 remains unchanged. Starting from the sequence d the sequence order is carried out, and the order selection is as follows:
[0023] The sequence nr = tr + r;
[0024] wherein, the sequence nr is the (r - 1)-th repeated sequence, tr is the sequence digit number, t is 1, and the repeated sequence is subjected to multiple order selection calculations.
[0025] In one or more embodiments of the present invention, the number K selection value conversion character set of the above set interval is the English letter order, that is, [A, B, C, D, E, F, G, H,..., S, T, Z, AA, AB, AC]. The financial institution gives the weight value β according to the position point of the result value in the set. The weight value β is used as the input data of the credit granting model. The financial institution requests data from the client main end through the client sub-end for the first time, and the client will return the converted set interval and all relevant data of this type of data.
[0026] In one or more embodiments of the present invention, the above center point deviation correction calculation is divided into positive deviation correction and negative deviation correction, that is, when the cumulative expected value is greater than the cumulative actual value, it is positive deviation correction, and when the cumulative expected value is less than the cumulative actual value, it is negative deviation correction. The deviation correction calculation formula is:
[0027]
[0028] wherein, K is the current clustering interval value of the set, E is the expected value each time, C is the actual value each time, and N is the set threshold.
[0029] In one or more embodiments of the present invention, the above center point deviation correction calculation is divided into positive deviation correction and negative deviation correction, that is, when the cumulative expected value is greater than the cumulative actual value, it is positive deviation correction, and when the cumulative expected value is less than the cumulative actual value, it is negative deviation correction; during positive deviation correction, each center point is successively subtracted by the deviation correction calculation result; during negative deviation correction, each center point is successively added by the deviation correction calculation result. After the center point deviation correction calculation is completed, the data of a certain type will be re-clustered according to the priority set for each type of data on time, and the data of other types will be rotated according to the set number threshold N2; when the deviation correction calculation value exceeds the maximum threshold N3, the K value will be increased or decreased once and the re-calculation process will be carried out.
[0030] In one or more embodiments of the present invention, when the above financial institution requests data through the client secondary end, the data transmitted by the client secondary end is a character set. The financial institution does not access the original data of the client primary end, and the financial institution data cannot be restored.
[0031] In one or more embodiments of the present invention, after the above data sequence is generated, it will be distributed to different set interval character sets through deviation correction operations. When the data is retrieved, different data segments can be selected according to the differences in the sequences.
[0032] Beneficial effects
[0033] The present invention provides an adaptive data available but invisible desensitization system. Compared with the prior art, it has the following beneficial effects:
[0034] 1. By classifying and segmenting the market entity information data and performing symbol substitution, the data collected by government departments can be effectively transmitted and used by financial institutions, enabling financial institutions to obtain more complete market entity data, making credit products more accurate and more locally applicable, and realizing the value of data while ensuring data security.
[0035] 2. The present invention processes the market entity information data to be available but invisible. Through clustering operations on the data information and dividing the data into intervals, multiple operation results are formed in this data segment. The multiple operation results configured can form separate character sets. When retrieving data, financial institutions can obtain more complete data by retrieving the corresponding character sets.
[0036] 3. When processing the data, most of the data is calculated into different character sets according to the operation results. When financial institutions interact with the data, weights are given based on the position points of the result values of the financial institutions in the set. According to the weights as input data for the credit granting model, when financial institutions input data, the data is in an available state. External data request ends such as financial institutions obtain the hydrangea data and apply it to the business during this process. During this process, they do not access the original data of the client, and the received data cannot be restored.
[0037] 4. After external data demand ends such as financial institutions apply the data, actual result C and expected result E will be generated. There is a deviation between actual result C and expected result E. Through deviation correction calculations, various data operations are performed. As the data values in the data set continue to increase and a large number of application feedbacks of operation conversion results, the operation conversion results of the server will ultimately approach the accuracy expected by the application product infinitely. Brief description of the drawings
[0038] Figure 1 It is a schematic diagram of data transmission between the financial institution and the government department of the present invention;
[0039] Figure 2 This is a schematic diagram of the system of the present invention. Detailed implementation manners
[0040] The following will disclose multiple implementation manners of the present invention with the accompanying drawings. For the sake of clear illustration, many practical details will be described together in the following narrative. However, it should be understood that these practical details are not used to limit the present invention. That is to say, in some implementation manners of the present invention, these practical details are not necessary. In addition, for the purpose of simplifying the drawings, some conventional structures and elements in the prior art will be illustrated in a simple schematic manner in the drawings, and in all the drawings, the same reference numerals will be used to represent the same or similar elements. And if possible in implementation, the features of different embodiments can be applied interactively.
[0041] Unless otherwise defined, all the terms (including technical and scientific terms) used herein have their ordinary meanings, and their meanings can be understood by those skilled in this field. Further, the definitions of the above terms in commonly used dictionaries should be interpreted as having the same meaning as that in the related field of the present invention. Unless specifically defined otherwise, these terms will not be construed as idealized or overly formal meanings.
[0042] Please refer to Figure 1-2 , the present invention provides an adaptive data available but invisible desensitization system, which performs available but invisible processing on data. With the continuous increase of data values in the data set and a large number of application feedbacks of the operation conversion results, the operation conversion results of the server will ultimately approach the accuracy expected by the application product infinitely.
[0043] It includes:
[0044] A client, which is used to collect data and provide services to financial institutions. The client is divided into two parts, which are located in the financial institutions and the government departments that collect the information of market entities required by the financial institutions respectively. The client used by the financial institutions is the secondary end, and the client used by the government departments is the primary end;
[0045] A server, which loads a clustering algorithm device for specifying the K value and the center point value. The clustering algorithm processes the data transmitted by the client. The client classifies the data into multiple types of data α according to the types. Each type of data α contains multiple pieces of data. There is an interaction of type data α between the server and the client;
[0046] The server receives the data set and the expected center point value transmitted by the client, performs clustering operations, and divides the data set transmitted by the client into K set intervals , and saves the start value, end value of each segment and the corresponding character substitution value;
[0047] Threshold correction model, set deviation threshold. When the difference between the expected result E and the actual result C of the client secondary end request exceeds the set threshold, or when the difference between the average expected result E and the actual result C of the cumulative N - time data application reaches the set threshold, the server performs re - clustering of the correction center point and symbol replacement conversion on this type of data α.
[0048] Credit model, access the client secondary end, and receive the data transmitted by the client secondary end, and establish a credit model according to the data transmitted by the client secondary end.
[0049] In this embodiment, by classifying different data into type - α data, different data of the same user are classified. When a financial institution retrieves data, through detailed classification, the interaction of the data required by the financial institution is carried out. During the interaction, the financial institution is prevented from interacting with other types of the user's data.
[0050] Among them, when the server classifies the data set, a data interval is randomly generated. When the client secondary end and the client primary end interact to retrieve data, it is converted into the required retrieval set interval according to the difference of the data. , and the retrieval of data is realized by using different character sets between the interval sets, and the client primary end performs data interaction according to the character set requested by the secondary end.
[0051] In one embodiment, the clustering algorithm device is an improved K - MEANS. Among them, the type - α data contains the expected interval value K and the center point value X, which are randomly generated when there are no K and N values. K is the number of set intervals in the type - α data and the number of mid - point values. The numerical values are sorted in numerical sequence, and the generation of K does not contain repeated numerical sequences.
[0052] In this embodiment, the type - α data classified by the client is split into K set intervals and N center point values. When a financial institution retrieves data, through the feedback of the position of the required result by the client secondary end, one or more interval values can be accurately located. The interval values contain the data required by the financial institution. Through the interaction between the client primary end and the client secondary end, the client secondary end inputs the interval values into the financial institution model.
[0053] Among them, after the client generates the set interval and the center point value, it can perform fragmentation processing on the data. That is, when a financial institution retrieves data, only by generating the characters of the required data through the client secondary end, data interaction can be carried out according to the characters. During the interaction, the interaction is carried out through the client primary end and the client secondary end, and the client secondary end is directly connected to the credit model. The financial institution's processing right for data is available but not visible, which can protect user data to the greatest extent during use.
[0054] In one embodiment, the class data α contains multiple set intervals , and the set interval contains multiple pieces of related data, that is . Identify each piece of related data with a separate sequence, count the number of bits T in the data, and select the data sequence length according to the number of bits T, that is:
[0055] S 位数 = 3 / T;
[0056] where S 位数 is the sequence length. When selecting, S 位数 is an integer, and when the length number of bits is less than 3, S 位数 the data is the position of the data point selected for the sequence, that is, the <S 位数 |1*S 位数 |2*S 位数 > in the related data, three data points; S 位数 is an integer, and when the length number of bits is greater than 3, the first three digits of S 位数 are the positions of the data points selected for the sequence; when S 位数 is a decimal number, and when the length position after the decimal point is less than 3, the value after the decimal point of S 位数 is the position of the data point selected for the sequence, that is, the <S 位数 value after the decimal point|1*S 位数 value after the decimal point|2*S 位数 value after the decimal point>; S 位数 is a decimal number, and when the length number of bits after the decimal point is greater than 3, the first three digits after the decimal point of S 位数 are the positions of the data points selected for the sequence.
[0057] In this embodiment, by using the uniqueness of the data to establish a unique sequence regarding the related data, when the financial institution retrieves data, by converting the retrieval request into a related sequence, the location where the sequence is located can be retrieved more quickly, and the location of the set interval in the class data α can be retrieved more quickly . Thus, the data segment containing this data in this set interval is transmitted to the credit model of the financial institution.
[0058] Among them, the sequences contained in the class data α are all unique, and the sequences contained in different class data α can be the same sequence. When the financial institution makes a data request, it is necessary to specify the type of the class data α. The same sequence can exist in different class data α, which can reduce the difficulty of sequence generation and thus facilitate the distinction of sequences.
[0059] In one embodiment, the set interval Multiple related data included All are circular data. When the number of digits of the required generated sequence > the length of the related data itself, the related data is circular data. The first digit of the related data is the sequence position of the last digit of the related data, that is: d = q * d 数据长度 +1 to complete the data loop and generate a complete sequence. Among them, d is the serial number of the first digit of the data, q is the number of times the related data is selected, and d 数据长度 is the length of the related data.
[0060] In this embodiment, since there is a difference in the number of length digits between different related data, and when calculating the sequence selection, a large number of selected data points will cause data selection gaps. Forming the data into closed-loop data can perform complete sequence position selection to ensure the integrity of sequence selection during use.
[0061] In one embodiment, when the sequences are the same, the method of taking the related sequence selection position is used for modification, that is: sequence 1 = sequence n When, sequence 1 remains unchanged. Starting from sequence d perform sequence position. The position selection is as follows:
[0062] Sequence nr = tr + r;
[0063] Among them, sequence nr is the (r - 1)th repeated sequence, tr is the number of digits of the sequence, t is 1, and the repeated sequence is calculated for multiple position selections.
[0064] In this embodiment, the sequence is a multi-digit number. When there are multiple sequences, the first sequence remains unchanged, and the modification of the sequence starts from the second sequence. The second sequence changes the selection position in the number of digits of the sequence. Exemplarily, all five sequences are 134917 - 173. Starting from the second sequence, the r value in the second sequence is 2, so the selected value is 1, that is, the overall position of the sequence is modified by performing the sequence position tr + r starting from the first digit of the sequence.
[0065] Among them, the related sequences existing during use will affect the data positioning. By changing the position of sequence selection, the sequence can be changed into a uniquely identifiable sequence, thereby ensuring the stability of sequence selection and avoiding errors during data selection.
[0066] In one embodiment, the set interval The number K selects a value to convert the character set to the English letter sequence, i.e., [A, B, C, D, E, F, G, H,..., S, T, Z, AA, AB, AC]. The financial institution gives the weight β according to the position of the result value in the set. The weight β is used as the input data of the credit model. The financial institution requests data from the client main end through the client secondary end for the first time, and the client will return the converted set interval. and all relevant data of this type of data.
[0067] In this embodiment, the set interval is converted into a character set. The character set contains multiple pieces of relevant data and marks the data in the form of a sequence. When the financial institution selects data, the client secondary end generates a sequence for the data required by the financial institution. According to the selection of the primary sequence, when the server prepares the sequence of the set interval it will distinguish the sequence times and mark the sequence times. The repeated sequences of the primary sequence are recorded as the secondary sequence after the sequence selection, and the secondary sequence is repeatedly recorded as the tertiary sequence after the sequence selection.
[0068] When the financial institution retrieves data, the client secondary end generates a sequence for the data. This generation is the primary sequence. If there are multiple similar primary sequences, the client secondary end calculates the secondary sequence for the data. When there is no identical sequence, the first data segment of the repeated sequence of the primary sequence is retrieved. When there is an identical sequence, the data segment of that sequence is retrieved.
[0069] Exemplarily, taking the provident fund payment data as an example, the data before operation is [344, 344, 630, 344, 600, 1200, 2010, 2007, 660, 892, 819, 440, 753, 2010, 1000, 3900, 3900, 4512, 4512... 1396, 4512, 3362, 2766, 899, 872, 350, 552, 3754, 600, 308, 308...], a total of 100,000 pieces of data. The length of the data set is subject to the actual input.
[0070] According to K being 29, the operation results of the case are [{60, 301}, {302, 437}, {440, 582}, {583, 690}, {691, 806}, {807, 993}, {994, 1021}, {1023, 1212}, {1313, 1311}, {1313, 1426}, {1426, 1519}, {1520, 1656}, {1658, 1777}, {1779, 1862}, {1864, 1961}, {1962, 2080}, {2081, 2101}, {2101, 2214}, {2216, 2378}, {2380, 2494}, {2496, 2599}, {2600, 2788}, {2789, 2897}, {2897, 3090}, {3092, 3180}, {3182, 3320}, {3322, 3606}, {3608, 4000}, {4001, 4512}]. The character set converted according to the operation results is [A, B, C, D, E, F, G, H, I, J, K, L, M, N, O, P, Q, R, S, T, U, V, W, X, Y, Z, AA, AB, AC], that is, each character corresponds to an interval segment. If 10 specific center points are given during the operation, the intervals corresponding to the converted characters will change accordingly.
[0071] When the financial institution requests data from the client for the first time, the client will return the converted result value and the entire set of converted values of this type of data. For example, if the returned result value is M, the entire set of returned converted values is [A, B, C, D, E, F, G, H, I, J, K, L, M, O, P, Q, R, S, T, U... AB, AC].
[0072] The financial institution can give a weight value based on the position of the result value in the set and use the weight value as the input data for the credit model.
[0073] In one embodiment, the center point deviation correction calculation is divided into positive deviation correction and negative deviation correction, that is, when the cumulative expected value is greater than the cumulative actual value, it is positive deviation correction, and when the cumulative expected value is less than the cumulative actual value, it is negative deviation correction. The deviation correction calculation formula is:
[0074]
[0075] Among them, K is the current clustering interval value of the set, E is the expected value each time, C is the actual value each time, and N is the set threshold.
[0076] In this embodiment, the obtained results can be deviation-corrected through the deviation correction formula. With the continuous increase of data values in the data set and a large number of application feedbacks of the operation conversion results, the operation conversion results of the server will ultimately approach the accuracy expected by the application product infinitely.
[0077] Among them, when the deviation of the data requested by the financial institution exceeds the set threshold, corrective measures are taken, which can increase the accuracy of the data when the financial institution retrieves the data again, making the data retrieval more accurate and improving the accuracy of the recommendations made by the financial institution.
[0078] In one embodiment, the central point deviation correction calculation is divided into positive deviation correction and negative deviation correction, that is, when the cumulative expected value is greater than the cumulative actual value, it is positive deviation correction, and when the cumulative expected value is less than the cumulative actual value, it is negative deviation correction; during positive deviation correction, each central point is successively subtracted by the deviation correction calculation result; during negative deviation correction, each central point is successively added by the deviation correction calculation result. After the central point deviation correction calculation is completed, re-clustering operations are performed on a certain type of data according to the priority set for each type of data on time, and rotation operations are performed on other types of data according to the set number threshold N2; when the deviation correction calculation value exceeds the maximum threshold N3, the value of K will be increased or decreased once and the re-operation process will be carried out.
[0079] In this embodiment, the settings of positive deviation correction and negative deviation correction can ensure the accuracy of the data re-clustering operation. When the data shows a deviation, the choice of positive deviation correction or negative deviation correction can ensure the accuracy of the data after deviation correction, and the situation where the deviation will become larger and larger will not occur when the data deviates.
[0080] In one embodiment, when the financial institution requests data through the client secondary end, the data transmitted by the client secondary end is a character set. The financial institution does not access the original data of the client primary end, and the data of the financial institution cannot be restored.
[0081] In this embodiment, what the financial institution accesses is the converted character set. The character set is input into the credit model for the establishment of the credit model, and the data is available but invisible during the data transmission process, thus ensuring the security of user data.
[0082] In one embodiment, after the data sequence is generated, it will be distributed to different set interval character sets after deviation correction operations, and different data segments can be selected according to the differences in the sequences when the data is retrieved.
[0083] In summary, the technical solutions disclosed in the above embodiments of the present invention have at least the following advantages:
[0084] 1. By classifying and segmenting the market entity information data and performing symbol substitution, the data collected by the government department can be effectively transmitted and used by the financial institution, so that the financial institution can obtain more complete market entity data, making the credit products more accurate and more applicable locally, and realizing the value of data while ensuring data security.
[0085] 2. The present invention processes the market entity information data to be available but invisible. Through clustering operations on the data information, the data is divided into intervals, enabling multiple operation results to be formed in this data segment. The multiple operation results of this configuration can form separate character sets. When retrieving data, the financial institutions can obtain more complete data by retrieving the corresponding character sets.
[0086] 3. When the present invention processes the data, most of the data is calculated into different character sets according to the operation results. When the financial institutions conduct data interaction, weights are given based on the position points of the result values of the financial institutions in the set. The weights are used as the input data for the credit-granting model. When the financial institutions input data, the data is in an available state. External data request parties such as financial institutions obtain the hydrangea data and apply it to the business during this process. During this process, they do not come into contact with the original client data, and the received data cannot be restored.
[0087] 4. After the external data demand parties such as financial institutions apply the data, actual result C and expected result E will be generated. There is a deviation between actual result C and expected result E. Through deviation correction calculations, various data operations are carried out. With the continuous increase of data values in the data set and the large amount of application feedback of the operation conversion results, the operation conversion results of the server will ultimately approach the accuracy expected by the application product infinitely.
[0088] Although the present invention is disclosed in combination with the above embodiments, it is not intended to limit the present invention. Any person skilled in this art can make various modifications and refinements without departing from the spirit and scope of the present invention. Therefore, the protection scope of the present invention should be defined by the appended claims.
Claims
1. An adaptive data available but invisible desensitization system, characterized in that, it includes: A client, which is used to collect data and provide services to financial institutions. The client is divided into two parts, which are located in the financial institution and the government department that collects the information of market entities required by the financial institution respectively. The client used by the financial institution is the secondary end, and the client used by the government department is the primary end; A server, which loads a clustering algorithm device for specifying the K value and the center point value. The clustering algorithm processes the data transmitted by the client. The client classifies the data into multiple types of data α according to the type. Each type of data α contains multiple pieces of data. An interaction of type data α is formed between the server and the client; The server receives the data set and the expected center point value transmitted by the client, performs clustering operations, and divides the data set transmitted by the client into K set intervals. And save the start value, end value of each segment and the corresponding character substitution value. A threshold correction model, which sets a deviation threshold. When the difference between the expected result E requested by the secondary end of the client and the actual result C exceeds the set threshold, or when the difference between the average expected result E and the actual result C of the cumulative N times of data applications reaches the set threshold, the server performs re-clustering of the center point and symbol replacement conversion on this type of data α; A credit model, which accesses the secondary end of the client and receives the data transmitted by the secondary end of the client, and establishes a credit model according to the data transmitted by the secondary end of the client.
2. An adaptive data available but invisible desensitization system according to claim 1, characterized in that, The clustering algorithm device is an improved K-MEANS. Among them, the class data α contains the expected interval value K and the center point value X, which are randomly generated when there are no K and N values. K is the number of the set intervals in the class data α , and X is the number of midpoint values. The values are sorted in numerical sequence, and the generation of K does not contain duplicate numerical sequences.
3. An adaptive data available but invisible desensitization system according to claim 2, characterized in that, The class data α contains multiple set intervals Set interval contains multiple pieces of related data, namely Identify each piece of related data with a separate sequence, count the number of digits T in the data, and select the data sequence length according to the number of digits T, that is: S 位数 = 3 / T; Among them, S 位数 is the sequence length.
4. An adaptive data available but invisible desensitization system according to claim 3, characterized in that, Set interval Multiple related data included All are circular data. When the number of digits of the sequence to be generated > the length of the related data itself, the related data is circular data, and the first digit of the related data is the position of the last digit of the related data, that is: d = q * d 数据长度 +1 to complete the data closed-loop and generate a complete sequence. Among them, d is the serial number of the first digit of the data, q is the number of times the related data is selected, and d 数据长度 Is the length of the related data.
5. An adaptive data available but invisible desensitization system according to claim 4, characterized in that, When sequences are identical, the relevant sequence selection order is adopted for modification, that is: sequence 1 = sequence n When, sequence 1 remains unchanged. Starting from sequence d sequence order is carried out, and the order selection is as follows: Sequence nr = tr + r; Among them, the sequence nr is the (r - 1)-th repeated sequence, tr is the number of digits of the sequence, t is 1, and multiple sequential selection calculations are performed on the repeated sequence.
6. An adaptive data available but invisible desensitization system according to claim 5, characterized in that, Set interval The conversion character set of the selected value of the number K is the alphabetical order, i.e., [A, B, C, D, E, F, G, H,..., S, T, Z, AA, AB, AC]. The financial institution gives the weight value β according to the position point of the result value in the set. The weight value β is used as the input data of the credit model. The financial institution requests data from the client main end through the client sub-end for the first time, and the client will return the converted set interval And all relevant data of this type of data.
7. An adaptive data available but invisible desensitization system according to claim 6, characterized in that, The center point correction calculation is divided into positive correction and negative correction, that is, when the cumulative expected value is greater than the cumulative actual value, it is positive correction, and when the cumulative expected value is less than the cumulative actual value, it is negative correction. The correction calculation formula is: where K is the current clustering interval value of this set, E is the expected value each time, C is the actual value each time, and N is the set threshold.
8. An adaptive data available but invisible desensitization system according to claim 7, characterized in that, The center point correction calculation is divided into positive correction and negative correction, that is, when the cumulative expected value is greater than the cumulative actual value, it is positive correction, and when the cumulative expected value is less than the cumulative actual value, it is negative correction; in the case of positive correction, each center point is successively subtracted by the correction calculation result; in the case of negative correction, each center point is successively added by the correction calculation result. After the center point correction calculation is completed, the re-clustering operation will be performed on a certain type of data according to the priority set for each type of data on time, and the rotation operation will be performed on other types of data according to the set number threshold N2; when the correction calculation value exceeds the maximum threshold N3, the K value will be increased or decreased once and the re-operation process will be performed.
9. An adaptive data available but invisible desensitization system according to claim 8, characterized in that, When the financial institution requests data to pass through the secondary client for data request, the data transmitted by the secondary client is the character set. The financial institution does not access the original data of the primary client, and the financial institution's data cannot be restored.
10. An adaptive data available but invisible desensitization system according to claim 9, characterized in that, After the data sequence is generated, it will be distributed to different set interval character sets through deviation correction operations, and different data segments can be selected according to the differences in the sequences when the data is retrieved.
Citation Information
Patent Citations
Data anonymization methods and systems
CN105975870B
Agricultural data trusted circulation platform based on data model
CN113487443A
Abnormality detection method and device and storage medium
CN114664439A