A Method for Data Sharing in a Technology Finance Platform

By extracting key fields and metadata tags from the technology finance platform, combining the model to evaluate the sensitivity level and risk level, dynamically calculate the sensitivity score and perform desensitization processing, the secure transmission problem in the data sharing process is solved, and efficient and secure data sharing is achieved.

CN119720281BActive Publication Date: 2025-07-08GUANGZHOU COLLEGE OF COMMERCE +1
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510221028.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-07-08
Estimated Expiration
2045-02-27

AI Technical Summary

Technical Problem

In the process of data sharing of technology financial platforms, how to achieve dynamic assessment and secure transmission of data sensitivity to ensure accurate data identification, sensitivity level determination and secure sharing.

Method used

By extracting key field information and metadata tags, the sensitivity level is determined using logistic regression and decision tree model, the risk level of the data source and the recipient's permissions are calculated, and the dynamic sensitivity score is used to generate dynamic sensitivity tags, and the format checksum transmission is finally performed.

Benefits of technology

It realizes flexible protection measures based on data sensitivity and source risks, balances data security and availability, improves the security and efficiency of data sharing, and ensures traceability and compliance of the desensitization process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119720281B_ABST
    Figure CN119720281B_ABST
Patent Text Reader

Abstract

The present application provides a method for data sharing in a science and technology finance platform, including: extracting key field information and metadata tags from the data stream of the science and technology finance platform, where the key field information includes data content and source identifiers; analyzing data type characteristic values according to the metadata tags, and using a logistic regression model to determine the data category and its corresponding sensitivity level; querying the source credibility rule base based on the source identifier to determine the risk level of the data source; calculating a dynamic sensitivity score and generating a dynamic sensitivity tag according to the sensitivity level, risk level, and recipient permission identifier; determining the final desensitization method of the field according to the risk level mapping table in combination with the sensitivity level; performing format verification on the processed data stream and then transmitting it, and recording the execution information of the dynamic sensitivity tag and desensitization strategy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of information technology, and in particular to a method for sharing data on a science and technology finance platform. Background Art

[0002] In the process of sharing data on a science and technology finance platform, it is necessary to solve the technical problem of dynamically evaluating data sensitivity to achieve the secure transmission of shared data. First, it is necessary to accurately identify and extract key field information and metadata tags in the financial data streams of different science and technology finance platforms. Secondly, in the process of analyzing data type characteristics, it is necessary to ensure the accuracy of sensitive level determination. At the same time, the dynamic sensitivity scoring mechanism needs to comprehensively consider the risk of data sources, the sensitive level of data, and the permissions of the receiving party. Based on the dynamic sensitivity score, a data desensitization strategy is determined, and the shared data is desensitized according to the strategy to ensure the secure transmission of the shared data. Summary of the Invention

[0003] The present invention provides a method for sharing data on a science and technology finance platform, which mainly includes:

[0004] Extracting key field information and metadata tags from the data stream of the science and technology finance platform, where the key field information includes data content and source identifier; analyzing data type characteristic values according to the metadata tags, and using a logistic regression model to determine the data category and its corresponding sensitive level; querying the source credibility rule library based on the source identifier to determine the risk level of the data source; calculating a dynamic sensitivity score and generating a dynamic sensitivity label according to the sensitive level, risk level, and recipient permission identifier; matching a desensitization strategy from a pre-established risk level mapping table according to the dynamic sensitivity label, and performing masking, hashing, or encryption processing on sensitive fields; performing format verification on the processed data stream and then transmitting it, and recording the dynamic sensitivity label and the execution information of the desensitization strategy.

[0005] Further, the extracting key field information and metadata tags from the data stream of the science and technology finance platform includes: after obtaining the data packet transmitted from the source of the science and technology finance platform, preprocessing the data packet, using the Pandas library to remove noise and unify the format; using the DecisionTreeClassifier in Scikit-learn to establish a key field recognition model, and using the key field recognition model to extract content items and identifiers from the preprocessed data packet; for numerical data, using the Pandas library to calculate the mean, variance, and extreme values to generate numerical metadata tags; for text data, using the TfidfVectorizer in Scikit-learn to calculate the term frequency and inverse document frequency to generate text metadata tags; using the Pandas library to merge the numerical metadata tags and text metadata tags to generate the final metadata tags.

[0006] Further, analyzing the data type feature values according to the metadata tags and determining the data category and its corresponding sensitivity level by using a logistic regression model includes: extracting the table structure information of the metadata from a relational database, obtaining the tag names and data types of all fields; screening the fields with data types of integer or floating point, and storing their corresponding numerical tag values as a list; summing all the numerical tag values in the list and dividing by the quantity to obtain the arithmetic mean as the feature value; converting the feature value into an input vector acceptable to the logistic regression model in a fixed format; using the logistic regression model to judge the data category, and when the output category is equal to the sensitive data category in the predefined category list, converting the category value into an input format acceptable to the decision tree model; judging the data sensitivity level based on the decision tree model, and the output result is the sensitivity level value in the predefined sensitivity level list.

[0007] Further, querying the source credibility rule base based on the source identifier and determining the risk level of the data source includes: performing standardization processing on the source code obtained from the data source identifier to obtain a standardized source code; using the standardized source code as input to match in the pre-established credibility rule base to obtain the weight value corresponding to the source code; according to the at least one obtained weight value, using a weighted average algorithm to calculate the risk value of the data source; if the risk value is lower than the preset risk threshold, determining that the data source is a low-risk source; if the risk value is higher than the preset risk threshold, obtaining the feature values of the data source, where the feature values at least include access frequency, data integrity, and historical records; using the feature values as input and inputting them into a pre-trained risk level evaluation model for classification calculation to obtain the risk level of the data source.

[0008] Further, calculating the dynamic sensitivity score and generating a dynamic sensitivity label according to the sensitivity level, risk level, and recipient permission identifier includes: obtaining the sensitivity level weight, risk level weight, and permission identifier weight for the preset dynamic weight algorithm parameter library; using a weighted summation formula to calculate the dynamic sensitivity score according to the weight parameters obtained from the parameter library; comparing the dynamic sensitivity score result with a preset threshold; if the score is higher than the preset threshold, extracting the feature values of the sensitivity level, risk level, and permission identifier; inputting the processed feature values into a pre-trained random forest classification model to obtain the corresponding classification result; converting the classification result into the corresponding dynamic sensitivity level according to the mapping rules predefined based on business requirements; generating the corresponding dynamic sensitivity label based on the dynamic sensitivity level.

[0009] Further, matching a desensitization strategy from a pre-established risk level mapping table according to the dynamic sensitivity label and performing masking, hashing, or encryption processing on sensitive fields includes: for the dynamic sensitivity label obtained from the pre-established database, extracting the type and content of the sensitive field; determining whether the field content contains sensitive information; if the field content meets the preset masking processing conditions, replacing the sensitive part in the field with asterisks to obtain the masked field content; performing hashing processing on the masked field content using the SHA256 algorithm to obtain the hashed field content; for the hashed field content, performing encryption processing using the AES algorithm to obtain the encrypted field; extracting the character frequency and length from the encrypted field, and using the character frequency and length as feature values; inputting the feature values into a pre-trained random forest model, and outputting the sensitive level of the field through the random forest model; according to the pre-established risk level mapping table, combining the sensitive level, determining the final desensitization method for the field.

[0010] Further, determining the final desensitization method for the field according to the pre-established risk level mapping table and combining the sensitive level includes: obtaining the pre-established risk level mapping table, where the risk level mapping table contains the correspondence between risk levels and sensitive levels; obtaining the sensitive level information of the target field from the data dictionary, and the sensitive level information in the data dictionary is consistent with the sensitive level information in the risk level mapping table; according to the sensitive level information, searching for the corresponding risk level in the risk level mapping table to obtain the risk level of the target field; selecting a desensitization method matching the risk level from a preset desensitization method library, where the desensitization method library contains desensitization methods corresponding to different risk levels; if the selected desensitization method contains adjustable parameters, adjusting the parameter values according to the risk level to obtain the adjusted desensitization method parameters; binding the adjusted desensitization method parameters to the target field to generate a final desensitization plan.

[0011] Further, extracting fields one by one from the data stream, for each field, judging according to preset field type and length rules to determine whether the field meets the preset requirements; if the type and length of the field match, generating execution information according to the final desensitization plan in a preset fixed format; converting the execution information into structured log data according to the preset log format requirements; storing the converted structured log data into a pre-configured MySQL database; for the data stream that passes the verification, transmitting it to a preset target system using Kafka message queue technology; if the type and length of the field do not match, recording an error log and skipping the field to continue processing the next field.

[0012] The technical solutions provided by the embodiments of the present invention may include the following beneficial effects:

[0013] The present invention discloses a method for data sharing in a science and technology finance platform. The method extracts key fields and metadata tags from the data stream, analyzes data characteristics and determines the sensitivity level, and at the same time evaluates the credibility of the data source, and then calculates a dynamic sensitivity score. Based on this score, the present invention matches the corresponding desensitization strategy from a preset risk level mapping table and performs masking, hashing or encryption processing on sensitive fields. This dynamic desensitization method can flexibly select appropriate protection measures according to the sensitivity of the data and the source risk, effectively balancing data security and usability. By performing format verification on the processed data and recording execution information, the present invention ensures the traceability and compliance of the desensitization process. This innovative method significantly improves the security and efficiency of the science and technology finance platform in processing sensitive data, providing strong support for the secure sharing and application of financial data. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] Figure 1 It is a flowchart of the method for data sharing in the science and technology finance platform of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0015] In order to enable those skilled in the art of the present technology to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings in the embodiments of this specification. Obviously, the described embodiments are only a part of the embodiments of this specification, rather than all the embodiments. Based on the embodiments in this specification, all other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the protection scope of this specification.

[0016] Such as Figure 1 , the method for data sharing in the science and technology finance platform of this embodiment may specifically include:

[0017] Step S101, extract key field information and metadata tags from the data stream of the science and technology finance platform, and the key fields include data content and source identifiers.

[0018] Obtain the data packet transmitted from the source of the science and technology finance platform. The data packet contains structured information and unstructured information. Use the Pandas library to preprocess the data packet, remove noise and unify the format to obtain consistent data. Use the DecisionTreeClassifier in Scikit-learn to establish a keyword field recognition model and obtain the model parameters. Process the preprocessed data packet through the keyword field recognition model to extract content items and identifiers. Classify according to the data type of the content item. If it is numerical data, enter the numerical data processing process. If it is text data, enter the text data processing process. For numerical data, use the Pandas library to calculate the mean, variance and extreme values to generate numerical metadata labels. For text data, use the TfidfVectorizer in Scikit-learn to calculate the word frequency and inverse document frequency to generate text metadata labels. Use the Pandas library to merge the numerical metadata labels and text metadata labels to obtain the final metadata labels.

[0019] Exemplarily, obtain the data packet transmitted from the source of the science and technology finance platform. The data packet contains structured information such as transaction amount, user ID and unstructured information such as user comments. Use the Pandas library to preprocess the data packet, remove noise such as missing values and outliers, and unify the format such as standardizing the date field to "YYYY-MM-DD" to obtain consistent data. Use the DecisionTreeClassifier in Scikit-learn to establish a keyword field recognition model, set model parameters such as the maximum depth to 5 and the minimum sample split to 10, and obtain the model parameters. Process the preprocessed data packet through the keyword field recognition model to extract content items such as transaction amount, user comments and identifiers such as user ID. Classify according to the data type of the content item. If it is numerical data such as transaction amount, enter the numerical data processing process. If it is text data such as user comments, enter the text data processing process. For numerical data, use the Pandas library to calculate the mean such as 1000, variance such as 500 and extreme values such as the maximum value of 2000 to generate numerical metadata labels. For text data, use the TfidfVectorizer in Scikit-learn to calculate the word frequency such as "transaction" appears 3 times and the inverse document frequency such as 0.5 to generate text metadata labels. Use the Pandas library to merge the numerical metadata labels and text metadata labels to obtain the final metadata labels. Associate the extracted keyword field information such as transaction amount, user comments with the generated metadata labels to determine the keyword field information set containing the data content and source identifier such as user ID.

[0020] Step S102: Analyze the data type feature values based on the final metadata tags, and use a logistic regression model to determine the data category and its corresponding sensitivity level.

[0021] Extract the table structure information of the metadata from the relational database of the final metadata tags, and obtain the tag names and data types of all fields. Filter the fields with data types of integer or floating point, and store their corresponding numerical tag values as a list. After summing all the numerical tag values in the list and dividing by the quantity, obtain the arithmetic mean as the feature value. Convert the feature value into an input vector acceptable to the logistic regression model in a fixed format. Use a logistic regression model pre-trained with historical data to judge the data category and obtain the data category value. If the data category value is equal to the sensitive data category in the predefined category list, convert this category value into an input format acceptable to the decision tree model. Based on the decision tree model, judge the data sensitivity level and obtain the sensitivity level value in the predefined sensitivity level list. According to the sensitivity level, risk level, and recipient permission identifier, calculate the dynamic sensitivity score using a weighted summation formula. Compare the dynamic sensitivity score result with a preset threshold. If the score is higher than the preset threshold, generate the corresponding dynamic sensitivity tag.

[0022] Exemplarily, extract the table structure information of the metadata from the relational database of the final metadata tags, and obtain the tag names and data types of all fields. For example, extract from a MySQL database that the data type of the field named "user age" is integer, and the data type of the field named "user income" is floating point. Filter the fields with data types of integer or floating point, and store their corresponding numerical tag values as a list. For example, store the values [25, 30, 35] of the "user age" field and the values [5000.0, 6000.0, 7000.0] of the "user income" field as two lists respectively. After summing all the numerical tag values in the list and dividing by the quantity, obtain the arithmetic mean as the feature value. For example, calculate the average value of the "user age" field as 30, and the average value of the "user income" field as 6000.0. Convert the feature value into an input vector acceptable to the logistic regression model in a fixed format. For example, convert [30, 6000.0] into a normalized vector of [0.3, 0.6]. Use a logistic regression model pre-trained with historical data to judge the data category and obtain the data category value. For example, the model outputs the category value as "financial data". If the data category value is equal to the sensitive data category in the predefined category list, convert this category value into an input format acceptable to the decision tree model. For example, convert "financial data" into a one-hot encoded vector of [1, 0, 0]. The predefined category list contains data category items and category values, and the category value of the current category can be obtained by querying the predefined category list; based on the decision tree model, judge the data sensitivity level and obtain the sensitivity level value in the predefined sensitivity level list. For example, the model outputs the sensitivity level value as "high".

[0023] Step S103: Query the source credibility rule library based on the source identifier to determine the risk level of the data source.

[0024] Based on the source identifier, query in the pre-established credibility rule library to obtain the source code corresponding to the source identifier. The credibility rule library contains the source code corresponding to the source identifier and the weight value corresponding to the source code. For the obtained source code, perform normalization processing to obtain the normalized source code. Use the normalized source code as input, match it in the credibility rule library, and obtain the weight value corresponding to the source code. According to the obtained at least one weight value, use the weighted average algorithm to calculate the risk value of the data source, where the risk value is equal to the sum of the products of each weight value and its corresponding score divided by the total sum of the weight values. If the risk value is lower than the preset risk threshold, determine that the data source is a low-risk source and end the data source risk determination process. If the risk value is higher than the preset risk threshold, obtain the characteristic values of the data source, including access frequency, data integrity, and historical records. Use the characteristic values as input and input them into the pre-trained random forest model for classification calculation to obtain the risk level of the data source.

[0025] Exemplarily, based on the source identifier "S12345", query in the pre-established credibility rule library to obtain the corresponding source code "C001". For the source code "C001", perform normalization processing by removing special characters and unifying case to obtain the normalized source code "C001-STD". Use the normalized source code "C001-STD" as input, match it in the credibility rule library, and obtain the corresponding weight values 0.8 and 0.2. According to the weight values 0.8 and 0.2 and their corresponding scores 5 and 3, use the weighted average algorithm to calculate the risk value, and the result is (0.8×5 + 0.2×3) / (0.8 + 0.2) = 4.6. If the risk value 4.6 is lower than the preset risk threshold 5, determine that the data source is a low-risk source and end the risk determination process. If the risk value is higher than the preset risk threshold, obtain the characteristic values of the data source, including an access frequency of 120 times per day, data integrity of 95%, and a good historical record. Input the characteristic values into the pre-trained risk level evaluation model for classification calculation to obtain the risk level of "medium risk". The risk level evaluation model is obtained by unsupervised training based on the random forest model.

[0026] Step S104: Calculate the dynamic sensitivity score and generate a dynamic sensitivity label according to the sensitivity level, risk level, and recipient permission identifier.

[0027] According to the sensitivity level, risk level, and recipient permission identifier, obtain the sensitivity level weight, risk level weight, and permission identifier weight from the preset dynamic weight algorithm parameter library. Using the weighted sum formula, multiply the sensitivity level weight by the sensitivity level value, multiply the risk level weight by the risk level value, and multiply the permission identifier weight by the permission identifier value to calculate the dynamic sensitivity score. Compare the dynamic sensitivity score with the preset threshold, and the preset threshold is determined by the percentile method based on business requirements. If the dynamic sensitivity score is higher than the preset threshold, extract the eigenvalue of the sensitivity level, risk level, and permission identifier. For numerical features, directly use the original value, and for categorical features, convert them into binary vectors through one-hot encoding. Input the processed eigenvalues into the pre-trained random forest classification model to obtain the classification result. According to the mapping rule defined in advance based on business requirements, convert the classification result into the corresponding dynamic sensitivity level. Based on the dynamic sensitivity level, generate the corresponding dynamic sensitivity label.

[0028] Exemplarily, according to the sensitivity level, risk level, and recipient permission identifier, obtain the sensitivity level weight of 0.5, the risk level weight of 0.3, and the permission identifier weight of 0.2 from the preset dynamic weight algorithm parameter library. These weight values are determined by statistical analysis using the linear regression method based on historical data. Using the weighted sum formula, multiply the sensitivity level weight by the sensitivity level value of 3, multiply the risk level weight by the risk level value of 4, and multiply the permission identifier weight by the permission identifier value of 2. Calculate the dynamic sensitivity score as 3×0.5 + 4×0.3 + 2×0.2 = 3.1. Compare the dynamic sensitivity score of 3.1 with the preset threshold of 2.8, and the preset threshold is determined by the percentile method based on business requirements. If the dynamic sensitivity score is higher than the preset threshold, extract the eigenvalue of the sensitivity level of 3, the risk level of 4, and the permission identifier of 2. For numerical features, directly use the original value, and for categorical features, convert them into binary vectors [1, 0, 1] through one-hot encoding. Obtain historical data, extract data features and sample labels. Use the feature selection method to screen out important features. According to the screened features and sample labels, construct the training set. For the training set, train the random forest classification model. If the model training is completed, obtain the prediction result. According to the prediction result, judge the accuracy rate and recall rate of the model. If the model evaluation result meets the preset threshold, determine that the model can be used for actual prediction to obtain the random forest classification model. Input the processed eigenvalues into the pre-trained random forest classification model, which is trained with 100 decision trees, and obtain the classification result of "high sensitivity". According to the mapping rule defined in advance based on business requirements, convert the classification result of "high sensitivity" into the corresponding dynamic sensitivity level of "3". Based on the dynamic sensitivity level of "3", generate the corresponding dynamic sensitivity label of "high".

[0029] Step S105: Match a desensitization strategy from a pre-established risk level mapping table according to the dynamic sensitivity label, and perform masking, hashing, or encryption processing on the sensitive field.

[0030] According to the dynamic sensitivity label, match a desensitization strategy from a pre-established risk level mapping table. Determine whether the sensitive field needs to be masked. If it meets the preset masking conditions, use asterisks to replace the sensitive part. For the field content after masking, perform hashing using the SHA256 algorithm. For the field content after hashing, perform encryption using the AES algorithm. Extract the character frequency and length from the encrypted field content and use them as feature values. Input the extracted feature values into a pre-trained random forest model to output the sensitivity level of the field. According to the pre-established risk level mapping table and combined with the sensitivity level, determine the final desensitization method for the field. If the desensitization method involves parameter adjustment, adjust the desensitization parameter values according to the sensitivity level and risk level. Based on the adjusted desensitization parameters, perform the final masking, hashing, or encryption processing on the sensitive field.

[0031] Exemplarily, according to the dynamic sensitivity label "High", the desensitization strategy "Mask + Hash + Encryption" is matched from the risk level mapping table. It is determined that the sensitive field "ID number" meets the mask processing condition, and the middle 8 digits are replaced with asterisks to obtain the mask result "110***********1234". The masked field content "110***********1234" is hashed using the SHA256 algorithm to generate the hash value "e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855". The hash value is encrypted using the AES algorithm with the key "1234567890123456" to generate the ciphertext "U2FsdGVkX18z7Z6X5X6X5X6X5X6X5X6X5X6X5X". The character frequency and length feature values are extracted from the ciphertext, with the frequency being "0.05, 0.03, 0.02" and the length being "64". The feature value set in the historical data is obtained, and the feature value set is used as training data to input into the random forest model. The cross-validation method is used to optimize the parameters of the random forest model to determine the optimal model structure. According to the optimized model structure, the random forest model is trained using the training data to obtain the initial model. The performance of the initial model is evaluated using the test data set. If the accuracy rate is lower than the preset threshold, the feature value set is adjusted. The model is retrained according to the adjusted feature value set to obtain the optimized sensitive level evaluation model. The feature values are input into the pre-trained random forest model, and the model outputs the sensitive level as "Level 3". According to the risk level mapping table, combined with the sensitive level "Level 3", the final desensitization method is determined to be "Encryption parameter adjustment". The key length of the AES encryption algorithm is adjusted to "256 bits", and a new ciphertext "U2FsdGVkX18z7Z6X5X6X5X6X5X6X5X6X5X6X5X6X5X" is generated by re-encryption. Based on the adjusted desensitization parameters, the final mask, hash, and encryption processing are performed on the sensitive field "ID number" to generate the final desensitization result "U2FsdGVkX18z7Z6X5X6X5X6X5X6X5X6X5X6X5X6X5X".

[0032] Specifically, according to the pre-established risk level mapping table, combined with the sensitive level, the final desensitization method for the field is determined, which specifically includes the following steps:

[0033] Obtain a pre-established risk level mapping table and extract the corresponding relationship between the risk level and the sensitivity level therein. The risk level mapping table contains the corresponding relationship between the risk level and the sensitivity level, which is used for the lookup operation in the subsequent steps. For the target field, obtain its sensitivity level information from the data dictionary. The sensitivity level information in the data dictionary is consistent with the sensitivity level information in the risk level mapping table, which is used for the lookup operation in the subsequent steps. According to the sensitivity level information, look up the corresponding risk level in the risk level mapping table. The lookup operation determines the risk level of the target field by matching the sensitivity level information in the data dictionary with the sensitivity level information in the risk level mapping table. Select a desensitization method that matches the risk level from the preset desensitization method library. The desensitization method library contains desensitization methods corresponding to different risk levels. The selection operation determines the applicable desensitization method by matching the risk level of the target field with the risk levels in the desensitization method library. If the selected desensitization method contains adjustable parameters, adjust the parameter values according to the risk level. The parameter adjustment operation determines the specific numerical value of the parameter value through the risk level to ensure the applicability of the desensitization method. Bind the adjusted desensitization method parameters to the target field to generate the final desensitization plan. The binding operation generates the final desensitization plan by applying the adjusted desensitization method parameters to the target field.

[0034] Exemplarily, the sensitivity level can be divided into three levels: low, medium, and high, and the corresponding risk levels may be levels 1, 2, and 3. This mapping relationship enables quick judgment of the risk level based on the sensitivity of the data, so as to take corresponding protection measures. As the core component of metadata management, the data dictionary not only stores the structural information of the data but also contains important attributes such as the sensitivity level. By querying the data dictionary, the sensitivity level information of the target field can be quickly obtained. After determining the sensitivity level of the target field, look up the corresponding risk level in the risk level mapping table. For example, if the sensitivity level of a field is marked as "high", the corresponding risk level may be level 3. The desensitization method library is a pre-defined set of desensitization strategies that provide corresponding desensitization methods for different risk levels. For example, for data with a risk level of 1, simple masking may be used; while for data with a risk level of 3, more complex encryption algorithms may be required. This hierarchical processing ensures the matching of the desensitization intensity with the data risk level, protecting sensitive information while avoiding the reduction of data availability caused by over-desensitization. The parameter adjustment of the desensitization method can be carried out according to the risk level. For example, for masking, the number of digits of the mask can be adjusted according to the risk level. Low-risk data may only need to mask the last four digits, while high-risk data may need to mask all characters except the first two. This flexible parameter adjustment mechanism makes the desensitization process more accurate and effective. The generation of the final desensitization plan binds the adjusted desensitization parameters to the target field to form a complete desensitization plan. This plan not only includes the specific desensitization method but also the specific parameters applied to this method. For example, for a field containing a personal ID number, the final desensitization plan may be: use masking, retain the first six and the last four digits, and replace the middle eight digits with asterisks. This systematic desensitization strategy formulation process not only improves the efficiency of data protection but also enhances the traceability and consistency of the whole process. By integrating risk assessment, desensitization method selection, and parameter adjustment into a unified framework, the needs of data security and data availability can be better balanced, providing strong support for the data management practice of the organization.

[0035] Step S106, perform format verification on the processed data stream and then transmit it, and record the execution information of the dynamic sensitivity label and the desensitization strategy.

[0036] Extract fields from the data stream one by one. For each of the said fields, judge according to the preset field type and length rules to determine whether the field meets the preset requirements; if the type and length of the field match, generate execution information according to the final desensitization scheme in a preset fixed format; convert the execution information into structured log data according to the preset log format requirements; store the converted structured log data into a pre-configured MySQL database; for the data stream that passes the verification, use Kafka message queue technology to transmit it to a preset target system; if the type and length of the field do not match, record an error log and skip the field, and continue to process the next field.

[0037] Exemplarily, it is first necessary to extract fields one by one and perform type and length matching. For example, for a data stream containing user information, the name field may be set as a character type with a length not exceeding 20 characters; the age field is an integer type with a range between 0 and 150. Such preset rules help to quickly identify abnormal data. When the field passes the verification, converting the execution information into structured log data according to the final desensitization scheme is for convenient storage and analysis. Structured logs usually include information such as timestamps, operation types, field names, and processing results. Such formatted logs can help quickly locate problems and perform data statistics and analysis. MySQL provides powerful query functions and good performance, which can meet the needs of most log storage and retrieval. A dedicated log table can be created, containing various relevant fields, such as operation time, field name, sensitivity, etc. For the data that passes the verification, choose to use Kafka message queue for transmission. Different topics can be set for different types of data to ensure that the data can be accurately transmitted to the corresponding target system. When encountering a field that does not meet the preset rules, record an error log and skip the processing. This can ensure that the processing of the entire data stream will not be interrupted due to individual abnormal data. The error log should contain sufficient information, such as field name, actual value, expected type and length, etc., for subsequent problem troubleshooting and data correction. The design of this entire process aims to improve the efficiency and security of data processing. By verifying with preset rules, abnormal data can be discovered and processed in a timely manner. The application of dynamic sensitivity labels and desensitization rules helps to protect user privacy. The use of structured logs and message queues improves the maintainability and scalability of the system.

[0038] As mentioned above, it is only the preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution of the present invention and its inventive concept, makes equivalent substitutions or changes, and all should be covered within the protection scope of the present invention.

Claims

1. A method for data sharing in a technology finance platform, characterized in that, Including: Extract keyword field information and metadata tags from the data stream of the technology finance platform, where the keyword field information includes data content and source identifiers; Analyze the data type characteristic values according to the metadata tags, and use a logistic regression model to determine the data category and its corresponding sensitivity level; Query the source credibility rule library based on the source identifier to determine the risk level of the data source; Calculate the dynamic sensitivity score and generate a dynamic sensitivity tag according to the sensitivity level, risk level and recipient permission identifier; Match the desensitization strategy from the pre-established risk level mapping table according to the dynamic sensitivity tag, and perform masking, hashing or encryption processing on the sensitive fields; Among them, it also includes: obtaining the dynamic sensitivity tag, extracting the type and content of the sensitive field; if the field content meets the preset masking processing condition, use asterisks to replace the sensitive part in the field to obtain the masked field content; for the masked field content, perform hashing processing using the SHA256 algorithm to obtain the hashed field content; for the hashed field content, perform encryption processing using the AES algorithm to obtain the encrypted field; extract the character frequency and length from the encrypted field, and use the character frequency and length as characteristic values; input the characteristic values into the pre-trained random forest model, and output the sensitivity level of the field through the random forest model; according to the pre-established risk level mapping table, combine the sensitivity level to determine the final desensitization method of the field; Perform format verification on the processed data stream and then transmit it, and record the execution information of the dynamic sensitivity tag and the desensitization strategy.

2. The method for sharing data of the science and technology finance platform according to claim 1, wherein, The extracting keyword field information and metadata tags from the data stream of the technology finance platform includes: After obtaining the data packet transmitted from the technology finance platform source, preprocess the data packet, and use the Pandas library to remove noise and unify the format; Use the DecisionTreeClassifier in Scikit-learn to establish a keyword field recognition model, and use the keyword field recognition model to extract content items and identifiers from the preprocessed data packet; For numerical data, use the Pandas library to calculate the mean, variance and extreme values, and generate numerical metadata tags; For text data, use the TfidfVectorizer in Scikit-learn to calculate the word frequency and inverse document frequency, and generate text metadata tags; Use the Pandas library to merge the numerical metadata tags and text metadata tags to generate the final metadata tags.

3. The method for sharing data of a science and technology finance platform according to claim 2, wherein The analyzing the data type characteristic values according to the metadata tags and using a logistic regression model to determine the data category and its corresponding sensitivity level includes: Extract the table structure information of the metadata from the relational database of the final metadata tags, and obtain the label names and data types of all fields; Filter the fields with data types of integer or floating point, and store their corresponding numerical label values as a list; Sum all the numerical tag values in the list and divide by the quantity to obtain the arithmetic mean as the feature value; Convert the feature value into an input vector acceptable to the logistic regression model in a fixed format; Use a logistic regression model pre-trained with historical data to judge the data category. When the output category is equal to the sensitive data category in the predefined category list, convert the output category value into an input format acceptable to the decision tree model; Judge the data sensitivity level based on the decision tree model, and the output result is the sensitivity level value in the predefined sensitivity level list.

4. The method for sharing data of a science and technology finance platform according to claim 3, wherein Query the source credibility rule library based on the source identifier to determine the risk level of the data source, including: Perform standardization processing on the source code obtained from the data source identifier to obtain a standardized source code; Use the standardized source code as input, match it in the pre-established credibility rule library, and obtain the weight value corresponding to the source code; According to the at least one obtained weight value, use the weighted average algorithm to calculate the risk value of the data source; If the risk value is lower than the preset risk threshold, determine that the data source is a low-risk source; If the risk value is higher than the preset risk threshold, obtain the feature values of the data source, and the feature values include access frequency, data integrity, and historical records; Use the feature values as input and input them into a pre-trained risk level assessment model for classification calculation to obtain the risk level of the data source.

5. The method for sharing data of a science and technology finance platform according to claim 4, wherein Calculate the dynamic sensitivity score and generate a dynamic sensitivity label according to the sensitivity level, risk level, and recipient permission identifier, including: Obtain the sensitivity level weight, risk level weight, and permission identifier weight for the preset dynamic weight algorithm parameter library; According to the weight parameters obtained from the parameter library, use the weighted summation formula to calculate the dynamic sensitivity score; Compare the dynamic sensitivity score result with a preset threshold; If the score is higher than the preset threshold, extract the feature values of the sensitivity level, risk level, and permission identifier; Input the processed feature values into a pre-trained random forest classification model to obtain the corresponding classification result; According to the mapping rules predefined based on business requirements, convert the classification result into the corresponding dynamic sensitivity level; Generate the corresponding dynamic sensitivity label based on the dynamic sensitivity level.

6. The method for sharing data of a science and technology finance platform according to claim 1, wherein Determine the final desensitization method of the field according to the pre-established risk level mapping table in combination with the sensitivity level, including: Obtain the pre-established risk level mapping table, which contains the corresponding relationship between the risk level and the sensitivity level; Obtain the sensitivity level information of the target field from the data dictionary, and the sensitivity level information in the data dictionary is consistent with the sensitivity level information in the risk level mapping table; According to the sensitivity level information, find the corresponding risk level in the risk level mapping table to obtain the risk level of the target field; Select a desensitization method matching the risk level from the preset desensitization method library, and the desensitization method library contains desensitization methods corresponding to different risk levels; If the selected desensitization method includes adjustable parameters, adjust the parameter values according to the risk level to obtain the adjusted desensitization method parameters; Bind the adjusted desensitization method parameters to the target field to generate the final desensitization plan.

7. The method for sharing data of a technology finance platform according to claim 6, characterized in that, After performing format verification on the processed data stream, transmit it and record the execution information of the dynamic sensitivity label and desensitization policy, including: Extract fields from the data stream one by one. For each field, use the preset field type and length rules to determine whether the field meets the preset requirements; If the type and length of the field match, generate execution information according to the final desensitization plan in a preset fixed format; Convert the execution information into structured log data according to the preset log format requirements; Store the converted structured log data in a pre-configured MySQL database; For the data stream that passes the verification, use the Kafka message queue technology to transmit it to a preset target system; If the type and length of the field do not match, record an error log and skip the field to continue processing the next field.

Citation Information

Patent Citations

  • Network table column type detection method based on probabilistic graph model

    CN114417885A

  • Scenarized data dynamic authorization management and control method and system

    CN115396109A

  • Intelligent hoisting mechanism load analysis method

    CN118656605A

  • Data processing method and device, equipment and storage medium

    CN118748614A

  • Computer sensitive data intelligent identification method based on artificial intelligence

    CN119312165A