Data sharing method and system based on data desensitization

The text data is predicted by the sequence labeling model and evaluated the importance of the text data, and dynamically adjusting the desensitization granularity in terminal permissions, solving the problem of difficult to balance data security and availability in traditional data desensitization methods, and achieving both security and availability of data sharing.

CN120509057AActive Publication Date: 2025-08-19BEIJING BORUIXIANGLUN SCI TECH DEV CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510999489.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-21
Publication Date
2025-08-19
Estimated Expiration
2045-07-21

AI Technical Summary

Technical Problem

Traditional data desensitization methods cannot accurately distinguish between ordinary fields in text and key sensitive fields, resulting in excessive or insufficient data desensitization, making it difficult to balance data privacy protection and data sharing availability, and making it difficult to flexibly adjust the desensitization strategy according to different data scenarios and terminal permission levels.

Method used

The sequence labeling model is used to predict the field category of the initial word sequence, and the initial field sequence is analyzed in combination with preset fields. The importance of the field is quantified and the desensitization granularity is dynamically adjusted according to the terminal permission level, and priority is set through word elements.

Benefits of technology

It realizes the security of data sharing without affecting data availability, ensuring that high-permitted terminals obtain complete data, and low-permitted terminals only obtain non-sensitive words, protecting data privacy while ensuring data availability of different terminals.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120509057A_ABST
    Figure CN120509057A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of computers, in particular to a data sharing method and system based on data desensitization, and the method comprises the steps: carrying out the field category prediction of an initial lexical element sequence through a sequence labeling model, carrying out the analysis of the initial field sequence through combining with a preset field, and achieving the quantitative evaluation of the importance degree of an initial field. The blindness of a traditional regularized desensitization mode is avoided, data security and data availability are both considered, priority setting is carried out with the initial lexical elements as units, the desensitization granularity is dynamically adjusted based on the terminal preset permission level, a high-permission terminal can obtain complete data, and a low-permission terminal can only obtain non-sensitive initial lexical elements, so that the user experience is improved. And the data availability of different terminals is ensured while the data privacy is protected.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data transmission, and in particular to a data sharing method and system based on data desensitization. Background Art

[0002] Currently, in data sharing and transmission scenarios, text data often contains key information such as personal privacy information and commercially sensitive data. Personal privacy information includes ID card numbers, bank accounts, etc., and commercially sensitive data includes financial statements, customer information and other key information.

[0003] Traditional data desensitization methods mostly use regularized processing, such as direct replacement and simple masking. However, regularized processing methods often cannot accurately distinguish between ordinary fields and key sensitive fields in the text, that is, the field recognition accuracy is insufficient, resulting in excessive or insufficient data desensitization, reducing data availability and security.

[0004] In addition, the rule-based processing method makes it difficult to flexibly adjust the desensitization strategy according to different data scenarios and terminal authority levels. It has poor generalization ability, that is, the ability to dynamically adapt to scenarios is poor, and it is difficult to achieve a balance between data privacy protection and data sharing availability.

[0005] Therefore, how to improve the security of data sharing without affecting data availability has become an urgent problem to be solved. Summary of the Invention

[0006] In response to the above technical problems, the technical solution adopted by the present invention is a data sharing method based on data desensitization, which includes the following steps: S101 , performing word-unit splitting on the acquired text data to be transmitted to obtain an initial word-unit sequence including M initial word-units, wherein M is a positive integer, and the priority corresponding to each initial word-unit is set to a first preset value.

[0007] S102: Input the initial word-gram sequence into a trained sequence labeling model to obtain first predicted field categories and first predicted probabilities corresponding to the M initial word-grams.

[0008] S103 : According to the first predicted field categories respectively corresponding to the M initial word-grams, an initial field sequence including N initial fields is formed from the M initial word-grams, where N is a positive integer smaller than M.

[0009] S104: Obtaining importance evaluation values corresponding to respective initial fields according to the preset fields and the initial field sequence.

[0010] S105 , according to the importance evaluation value corresponding to each initial field, a plurality of temporary key fields are screened from all the initial fields.

[0011] S106 , determining a plurality of target key fields according to each temporary key field and a plurality of reference key fields in a preset reference key field set.

[0012] S107: setting the priority corresponding to each initial word in each target key field to a second preset value.

[0013] S108 , when any terminal requests to obtain the text data to be transmitted, each initial word corresponding to a priority level less than or equal to the preset authority level is sent to the terminal according to the preset authority level corresponding to the terminal.

[0014] The present invention also provides a data sharing system based on data desensitization, the data sharing system based on data desensitization comprising: The word unit splitting module is used to split the acquired text data to be transmitted into word units to obtain an initial word unit sequence containing M initial word units, wherein M is a positive integer and the priority corresponding to each initial word unit is set to a first preset value.

[0015] The category prediction module is used to input the initial word-gram sequence into the trained sequence labeling model to obtain the first prediction field category and the first prediction probability corresponding to the M initial word-grams.

[0016] The field forming module is used to form an initial field sequence including N initial fields from the M initial word-grams according to the first predicted field categories corresponding to the M initial word-grams, wherein N is a positive integer less than M.

[0017] The field evaluation module is used to obtain the importance evaluation value corresponding to each initial field according to the preset field and the initial field sequence.

[0018] The field screening module is used to screen out several temporary key fields from all the initial fields according to the importance evaluation values corresponding to each initial field.

[0019] The field determination module is used to determine a plurality of target key fields according to each temporary key field and a plurality of reference key fields in a preset reference key field set.

[0020] The priority updating module is used to set the priority corresponding to each initial word in each target key field to a second preset value.

[0021] The data sharing module is used to send each initial word corresponding to a priority level less than or equal to the preset authority level to any terminal when the terminal requests to obtain the text data to be transmitted according to the preset authority level corresponding to the terminal.

[0022] The present invention has at least the following beneficial effects: the field category of the initial word-gram sequence is predicted through the sequence labeling model, the initial field sequence is analyzed in combination with the preset fields, and a quantitative evaluation of the importance of the initial field is achieved, thereby avoiding the blindness of the traditional regular desensitization method, taking into account data security and data availability, setting priorities based on the initial word-gram, and dynamically adjusting the desensitization granularity based on the preset authority level of the terminal. High-authority terminals can obtain complete data, and low-authority terminals can only obtain non-sensitive initial word-grams, thereby protecting data privacy while ensuring data availability for different terminals. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0024] Figure 1 A flowchart of a data sharing method based on data desensitization provided in the first embodiment of the present invention; Figure 2 A structural diagram of a data sharing system based on data desensitization provided in Example 2 of the present invention. DETAILED DESCRIPTION

[0025] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making any creative efforts shall fall within the scope of protection of the present invention.

[0026] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It is understood that, where appropriate, the above-mentioned terms used to distinguish similar objects can be interchanged so that the present invention can also implement other embodiments other than the above-mentioned illustrated embodiments or described embodiments. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or server that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0027] Example 1 This embodiment provides a data sharing method based on data desensitization, such as Figure 1 FIG. 1 is a flow chart of a data sharing method based on data desensitization provided in Embodiment 1 of the present invention. The data sharing method based on data desensitization includes the following steps: S101, performing word-unit splitting on the acquired text data to be transmitted to obtain an initial word-unit sequence comprising M initial word-units, where M is a positive integer and the priority corresponding to each initial word-unit is set to a first preset value; S102, inputting the initial word-gram sequence into a trained sequence labeling model to obtain first predicted field categories and first predicted probabilities corresponding to the M initial word-grams respectively; S103, forming an initial field sequence including N initial fields from the M initial word-grams according to the first predicted field categories corresponding to the M initial word-grams, where N is a positive integer less than M; S104, obtaining importance evaluation values corresponding to respective initial fields according to the preset fields and the initial field sequence; S105, according to the importance evaluation value corresponding to each initial field, a number of temporary key fields are screened from all the initial fields; S106, determining a plurality of target key fields based on each temporary key field and a plurality of reference key fields in a preset reference key field set; S107, setting the priority corresponding to each initial word in each target key field to a second preset value; S108 , when any terminal requests to obtain the text data to be transmitted, each initial word corresponding to a priority level less than or equal to the preset authority level is sent to the terminal according to the preset authority level corresponding to the terminal.

[0028] Among them, the text data to be transmitted may refer to the data to be shared with multiple terminals, the word unit may refer to the token, and the word unit splitting may adopt the N-gram splitting method, the dictionary-based splitting method, the statistical model-based splitting method, etc., which are not limited here.

[0029] The M initial word units that have been split out are sorted according to their respective positions in the text data to be transmitted to obtain an initial word unit sequence.

[0030] The priority of an initial word can be used to indicate the degree of desensitization requirement for the initial word, and can also be understood as the semantic importance of the initial word.

[0031] The sequence labeling model can adopt a recurrent neural network model, a long short-term memory network model, a Transformer model, etc., without limitation here, and the training process of the sequence labeling model will not be described in detail here.

[0032] The first prediction field category may belong to a preset field category set, which may include several preset field categories. The preset field category may be location, amount, name, etc. The first prediction probability may represent the possibility that the corresponding initial word belongs to its corresponding first prediction field category.

[0033] A single initial field may include multiple initial word-grams corresponding to the same first prediction field category. Similarly, the N initial fields are sorted according to their respective positions in the text data to be transmitted to obtain an initial field sequence.

[0034] The preset fields can be used to mask the initial fields in the initial field sequence, and the importance evaluation values corresponding to the respective initial fields are determined by masking and reconstructing the analysis.

[0035] The temporary key field may refer to a key field obtained according to occlusion reconstruction analysis, and the reference key field may refer to a key field set by the implementer according to prior information.

[0036] The target key fields may refer to key fields determined by comprehensive analysis and a priori.

[0037] The terminal may refer to a terminal that has access authority when sharing text data to be transmitted. According to the preset authority level and the priority of each initial word, the data desensitization result of the corresponding terminal may be determined, and then the data desensitization result may be sent to the corresponding terminal.

[0038] In a specific embodiment, forming an initial field sequence including N initial fields from the M initial word-grams according to the first predicted field categories respectively corresponding to the M initial word-grams includes: Initialize the word element identifier i=1, and initialize the field identifier j=1; When the first prediction field category corresponding to the i-th initial word is the same as the first prediction field category corresponding to the i-1-th initial word, add the i-th initial word to the initial field corresponding to the i-1-th initial word; When the first predicted field category corresponding to the i-th initial word is different from the first predicted field category corresponding to the i-1-th initial word, or when i=1, the j-th initial field is constructed from the i-th initial word, and j=j+1 is updated; Update i = i + 1, and return to the step of adding the i-th initial word to the initial field corresponding to the i-1-th initial word when the first predicted field category corresponding to the i-th initial word is the same as the first predicted field category corresponding to the i-1-th initial word, until i = M, determine the value of j-1 to be N, and obtain N initial fields; The initial field sequence is formed by the N initial fields.

[0039] In a specific implementation, obtaining the importance evaluation value corresponding to each initial field according to the preset field and the initial field sequence includes: For any initial field, replace the initial field with the preset field in the initial field sequence to obtain an intermediate field sequence; Inputting the intermediate field sequence into the trained reconstruction model to obtain a reconstructed field sequence; performing a difference calculation based on the reconstructed field sequence and the initial field sequence to obtain a difference calculation result; The difference calculation result is mapped to the importance evaluation value corresponding to the initial field.

[0040] Compared with the field masking method of directly deleting the initial field, setting the preset field is intended to make the trained reconstruction model aware that the preset field needs to be reconstructed.

[0041] The reconstruction model can include a convolution module and a deconvolution module. The convolution module is used to extract the feature information of the input intermediate field sequence, and the deconvolution module is used to reconstruct the reconstructed field sequence based on the feature information. The reconstruction model can adopt a Transformer model, a BERT model, a BART model, etc. The training process of the reconstruction model will not be repeated here.

[0042] Specifically, the difference calculation between the reconstructed field sequence and the initial field sequence can be performed using the cosine distance, and the mapping of the difference calculation results can be performed using the softmax function according to the difference calculation results corresponding to each initial field to obtain the importance evaluation value corresponding to each initial field.

[0043] In a specific embodiment, the determining of the target key fields according to the temporary key fields and the reference key fields in the preset reference key field set includes: Obtaining a plurality of basic key field sets, wherein the basic key field sets correspond to basic field category sequences; Forming a prediction field category sequence according to the first prediction field categories corresponding to the M initial word-grams; Determine a basic key field set corresponding to a basic field category sequence that is the same as the predicted field category sequence as the reference key field set; A plurality of target key fields are determined according to each temporary key field and a plurality of reference key fields in the reference key field set.

[0044] The first prediction field categories corresponding to the M initial word units are sorted according to their positions in the text data to be transmitted, to obtain a prediction field category sequence.

[0045] Specifically, this embodiment determines a reference key field set based on the comparison of the predicted field category sequence and the basic field category sequence, thereby determining the reference key fields that conform to the field category order, which can adapt to the expression of the text, streamline the number of reference key fields, and make the subsequently determined target key fields more reliable.

[0046] In a specific embodiment, determining a plurality of target key fields based on each temporary key field and a plurality of reference key fields in the reference key field set includes: A temporary key field set is formed by each temporary key field; A union of the temporary key field set and the reference key field set is calculated, and each temporary key field included in the union is used as a target key field.

[0047] Among them, any temporary key field in the union belongs to both the temporary key field set and the reference key field set.

[0048] In a specific implementation, the second preset value is greater than the first preset value, and the preset authority level is the first preset value or the second preset value.

[0049] The first preset value may be 1, and the second preset value may be 2.

[0050] In a specific embodiment, after step S107 and before step S108, the following steps are further included: For any target keyword field, a reference word is randomly determined from the initial word-grams contained in the target keyword field, and the initial word-gram corresponding to the reference word-gram is updated with a preset word-gram in the initial word-gram sequence, and then input into the trained sequence labeling model to obtain the second prediction field category and the second prediction probability corresponding to each of the M initial word-grams; Determining a first field category and a first classification probability corresponding to the initial field to which the reference word belongs; Determine a second field category and a second classification probability corresponding to an initial field to which a preset word-gram corresponding to the reference word-gram belongs; If the first field category and the second field category are different, setting the priority corresponding to the reference word to a third preset value; If the first field category and the second field category are the same, and the absolute value of the difference between the first classification probability and the second classification probability is greater than a preset probability threshold, setting the priority corresponding to the reference word to a third preset value; Otherwise, the priority corresponding to the reference word is not updated.

[0051] Among them, the first field category corresponding to the initial field to which the reference word belongs is the first predicted field category corresponding to the reference word, and the first classification probability can be calculated based on the average of the first predicted probabilities corresponding to all initial word units in the initial field to which the reference word belongs.

[0052] The second field category corresponding to the initial field to which the preset word corresponding to the reference word belongs is the second predicted field category corresponding to the preset word, and the first classification probability can be calculated based on the average of the second predicted probabilities corresponding to all initial word elements in the initial field to which the preset word belongs.

[0053] Specifically, when the first field category and the second field category are different, it indicates that the replacement of the reference word causes a greater semantic change, and therefore the reference word is considered to be more important, and the priority corresponding to the reference word is set to the third preset value.

[0054] When the first field category and the second field category are the same, and the absolute value of the difference between the first classification probability and the second classification probability is greater than the preset probability threshold, it also indicates that the replacement of the reference word causes a large semantic change. Therefore, it is considered that the importance of the reference word is higher, and the priority corresponding to the reference word is set to the third preset value.

[0055] In a specific implementation, the third preset value is greater than the second preset value, and the second preset value is greater than the first preset value.

[0056] Among them, the first preset value may be 1, the second preset value may be 2, and the third preset value may be 3.

[0057] In a specific implementation, the preset authority level is the first preset value, the second preset value, or the third preset value; The sending, according to a preset authority level corresponding to the terminal, each initial word corresponding to a priority level less than or equal to the preset authority level to the terminal includes: According to a preset authority level corresponding to the terminal, each initial word corresponding to a priority greater than the preset authority level in the initial word sequence is replaced with a preset word to obtain a target word sequence; The target word sequence is sent to the terminal.

[0058] Among them, the target word sequence is the data desensitization result of the corresponding terminal.

[0059] The first embodiment of this invention predicts the field category of the initial word-gram sequence through a sequence labeling model, analyzes the initial field sequence in combination with preset fields, and realizes a quantitative assessment of the importance of the initial fields, thereby avoiding the blindness of the traditional regular desensitization method, taking into account both data security and data availability, setting priorities based on the initial word-gram, and dynamically adjusting the desensitization granularity based on the preset authority level of the terminal. High-authority terminals can obtain complete data, while low-authority terminals can only obtain non-sensitive initial word-grams, thereby protecting data privacy while ensuring data availability for different terminals.

[0060] Example 2 This embodiment 2 provides a data sharing system based on data desensitization, such as Figure 2 FIG. 1 is a schematic diagram of a data sharing system based on data desensitization according to a second embodiment of the present invention. The data sharing system based on data desensitization includes: The word unit splitting module 201 is used to split the acquired text data to be transmitted into word units to obtain an initial word unit sequence including M initial word units, where M is a positive integer and the priority corresponding to each initial word unit is set to a first preset value; A category prediction module 202 is configured to input the initial word-gram sequence into a trained sequence labeling model to obtain first predicted field categories and first predicted probabilities corresponding to the M initial word-grams. a field forming module 203 configured to form an initial field sequence including N initial fields from the M initial word-grams according to the first predicted field categories corresponding to the M initial word-grams, where N is a positive integer less than M; A field evaluation module 204 is configured to obtain importance evaluation values corresponding to respective initial fields based on preset fields and the initial field sequence; The field screening module 205 is used to screen all the initial fields to obtain a number of temporary key fields according to the importance evaluation values corresponding to the initial fields; The field determination module 206 is configured to determine a plurality of target key fields based on each temporary key field and a plurality of reference key fields in a preset reference key field set; The priority updating module 207 is configured to set the priority corresponding to each initial word in each target key field to a second preset value; The data sharing module 208 is configured to send each initial word corresponding to a priority level less than or equal to the preset authority level to any terminal when the terminal requests to obtain the text data to be transmitted, according to the preset authority level corresponding to the terminal.

[0061] It should be noted that the specific limitations of the data sharing system based on data desensitization can be found in the limitations of the data sharing method based on data desensitization mentioned above, which will not be repeated here. The information interaction, execution process, etc. between the above modules are based on the same concept as the embodiment of the method of the present invention. Their specific functions and technical effects can be found in the embodiment of the method, which will not be repeated here.

[0062] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention in any form. Although the present invention has been disclosed as above in terms of preferred embodiments, they are not intended to limit the present invention. Any technician familiar with this profession can make some changes or modifications to equivalent embodiments of equivalent changes using the technical contents disclosed above without departing from the scope of the technical solution of the present invention. However, any simple modifications, equivalent changes and modifications made to the above embodiments based on the technical essence of the present invention without departing from the content of the technical solution of the present invention are still within the scope of the technical solution of the present invention.

Claims

1. A data sharing method based on data desensitization, characterized in that: The data sharing method based on data desensitization includes the following steps: S101, performing word-unit splitting on the acquired text data to be transmitted to obtain an initial word-unit sequence comprising M initial word-units, where M is a positive integer and the priority corresponding to each initial word-unit is set to a first preset value; S102, inputting the initial word-gram sequence into a trained sequence labeling model to obtain first predicted field categories and first predicted probabilities corresponding to the M initial word-grams respectively; S103, forming an initial field sequence including N initial fields from the M initial word-grams according to the first predicted field categories corresponding to the M initial word-grams, where N is a positive integer less than M; S104, obtaining importance evaluation values corresponding to respective initial fields according to the preset fields and the initial field sequence; S105, according to the importance evaluation value corresponding to each initial field, a number of temporary key fields are screened from all the initial fields; S106, determining a plurality of target key fields based on each temporary key field and a plurality of reference key fields in a preset reference key field set; S107, setting the priority corresponding to each initial word in each target key field to a second preset value; S108 , when any terminal requests to obtain the text data to be transmitted, each initial word corresponding to a priority level less than or equal to the preset authority level is sent to the terminal according to the preset authority level corresponding to the terminal.

2. The data sharing method based on data desensitization according to claim 1 is characterized in that: The forming of an initial field sequence including N initial fields from the M initial word-grams according to the first predicted field categories respectively corresponding to the M initial word-grams includes: Initialize the word element identifier i=1, and initialize the field identifier j=1; When the first prediction field category corresponding to the i-th initial word is the same as the first prediction field category corresponding to the i-1-th initial word, add the i-th initial word to the initial field corresponding to the i-1-th initial word; When the first predicted field category corresponding to the i-th initial word is different from the first predicted field category corresponding to the i-1-th initial word, or when i=1, the j-th initial field is constructed from the i-th initial word, and j=j+1 is updated; Update i = i + 1, and return to the step of adding the i-th initial word to the initial field corresponding to the i-1-th initial word when the first predicted field category corresponding to the i-th initial word is the same as the first predicted field category corresponding to the i-1-th initial word, until i = M, determine the value of j-1 to be N, and obtain N initial fields; The initial field sequence is formed by the N initial fields.

3. The data sharing method based on data desensitization according to claim 1 is characterized in that: Obtaining the importance evaluation value corresponding to each initial field according to the preset field and the initial field sequence includes: For any initial field, replace the initial field with the preset field in the initial field sequence to obtain an intermediate field sequence; Inputting the intermediate field sequence into the trained reconstruction model to obtain a reconstructed field sequence; performing a difference calculation based on the reconstructed field sequence and the initial field sequence to obtain a difference calculation result; The difference calculation result is mapped to the importance evaluation value corresponding to the initial field.

4. The data sharing method based on data desensitization according to claim 1 is characterized in that: The step of determining a plurality of target key fields based on the temporary key fields and a plurality of reference key fields in a preset reference key field set includes: Obtaining a plurality of basic key field sets, wherein the basic key field sets correspond to basic field category sequences; Forming a prediction field category sequence according to the first prediction field categories corresponding to the M initial word-grams; Determine a basic key field set corresponding to a basic field category sequence that is the same as the predicted field category sequence as the reference key field set; A plurality of target key fields are determined according to each temporary key field and a plurality of reference key fields in the reference key field set.

5. The data sharing method based on data desensitization according to claim 4 is characterized in that: The determining of a plurality of target key fields according to each temporary key field and a plurality of reference key fields in the reference key field set includes: A temporary key field set is formed by each temporary key field; A union of the temporary key field set and the reference key field set is calculated, and each temporary key field included in the union is used as a target key field.

6. The data sharing method based on data desensitization according to claim 1 is characterized in that: The second preset value is greater than the first preset value, and the preset authority level is the first preset value or the second preset value.

7. The data sharing method based on data desensitization according to claim 1 is characterized in that: After step S107 and before step S108, the following steps are further included: For any target keyword field, a reference word is randomly determined from the initial word-grams contained in the target keyword field, and the initial word-gram corresponding to the reference word-gram is updated with a preset word-gram in the initial word-gram sequence, and then input into the trained sequence labeling model to obtain the second prediction field category and the second prediction probability corresponding to each of the M initial word-grams; If the first prediction field category corresponding to the reference word-gram is different from the second prediction field category corresponding to the preset word-gram, setting the priority corresponding to the reference word-gram to a third preset value; If the first prediction field category corresponding to the reference word-gram and the second prediction field category corresponding to the preset word-gram are the same, and the absolute value of the difference between the first prediction probability corresponding to the reference word-gram and the second prediction probability corresponding to the preset word-gram is greater than a preset probability threshold, then setting the priority corresponding to the reference word-gram to a third preset value; Otherwise, the priority corresponding to the reference word is not updated.

8. The data sharing method based on data desensitization according to claim 7 is characterized in that: The third preset value is greater than the second preset value, and the second preset value is greater than the first preset value.

9. The data sharing method based on data desensitization according to claim 8 is characterized in that: The preset authority level is the first preset value, the second preset value or the third preset value; The sending, according to a preset authority level corresponding to the terminal, each initial word corresponding to a priority level less than or equal to the preset authority level to the terminal includes: According to a preset authority level corresponding to the terminal, each initial word corresponding to a priority greater than the preset authority level in the initial word sequence is replaced with a preset word to obtain a target word sequence; The target word sequence is sent to the terminal.

10. A data sharing system based on data desensitization, characterized in that: The data sharing system based on data desensitization includes: a word unit splitting module, configured to split the acquired text data to be transmitted into word units to obtain an initial word unit sequence comprising M initial word units, wherein M is a positive integer and the priority corresponding to each initial word unit is set to a first preset value; A category prediction module, configured to input the initial word-gram sequence into a trained sequence labeling model to obtain first predicted field categories and first predicted probabilities corresponding to the M initial word-grams; a field forming module, configured to form an initial field sequence including N initial fields from the M initial word-grams according to the first predicted field categories corresponding to the M initial word-grams, where N is a positive integer less than M; A field evaluation module, configured to obtain an importance evaluation value corresponding to each initial field according to a preset field and the initial field sequence; The field screening module is used to screen out several temporary key fields from all initial fields according to the importance evaluation values corresponding to each initial field; A field determination module is used to determine a plurality of target key fields based on each temporary key field and a plurality of reference key fields in a preset reference key field set; a priority updating module, configured to set the priority corresponding to each initial word in each target key field to a second preset value; The data sharing module is used to send each initial word corresponding to a priority level less than or equal to the preset authority level to any terminal when the terminal requests to obtain the text data to be transmitted according to the preset authority level corresponding to the terminal.

Citation Information

Patent Citations

  • Desensitization method, desensitization device, electronic equipment and storage medium

    CN114626097A

  • Data desensitization method, system and equipment and storage medium

    CN116702212A

  • Field desensitization mode determination method and device, electronic equipment and storage medium

    CN117272372A

  • Data desensitization method and device, electronic equipment and storage medium

    CN118797713A

  • Scientific and technological financial platform data sharing method

    CN119720281A