Data processing method and device and related equipment

Through the data processing device dynamically updates the policy set, the key words in the data are automatically identified and removed, which solves the problems of high labor costs and high delays in the prior art, and achieves efficient and accurate data security guarantees.

CN120234332APending Publication Date: 2025-07-01HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311868241.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-29
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

In the prior art, removing key information from data requires manual definition of annotation fields, resulting in high labor costs and high delays, making it difficult to adapt to changes in business demand.

Method used

Dynamically update the policy set by the data processing device, automatically identify and remove keyword entries in the data, reduce labor costs and improve removal efficiency.

Benefits of technology

It can efficiently and accurately adapt to changes in data demand in a short period of time, reduce the cost and time of removing keyword entries in the data, and improve the removal efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120234332A_ABST
    Figure CN120234332A_ABST
Patent Text Reader

Abstract

The invention discloses a data processing method and device and related equipment, and relates to the technical field of data processing. A second strategy set is obtained, the second strategy set is obtained by updating the first strategy set based on the first data and the first processing result, the first processing result is obtained by removing keyword items in the first data based on the first strategy set, and the second strategy set comprises at least one processing strategy; moreover, the data processing device also obtains second data, and removes keyword items in the second data by using a second strategy set to obtain a second processing result. The processing strategy in the strategy set can be automatically updated, and the processing strategy in the strategy set does not need to be manually adjusted, so that the labor cost of updating the processing strategy can be effectively reduced. Moreover, the dynamically updated processing strategy can adapt to the demand change of removing the keyword item in the data in an actual application scene, so that the accuracy of removing the keyword item in the data can be kept at a relatively high level.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of data processing, and in particular, to a data processing method, apparatus, and related equipment. Background Art

[0002] With the development of technologies such as big data and cloudification, the requirements for data security are getting higher and higher. For example, when a user uploads data to a cloud database, or during the process of querying data, the user usually expects to remove the key information in the uploaded / queried data, that is, to perform transformation processing on the key information in the data through corresponding rules, such as replacing or deleting the key information. In this way, even if the database is attacked by the network or the data is leaked due to other reasons, the original key information in the data will not be obtained by illegal users, thereby improving the security of the data.

[0003] Currently, it is usually to use manually defined annotation fields to first identify the key information that matches the annotation fields in an analysis object such as a text or a database, and further remove the above key information. However, in actual application scenarios, the requirements for removing keyword entries may change. For example, as the annotation fields change with business requirements, the types and specific contents of the analysis objects to be updated may change. At this time, it is often necessary to manually add newly defined annotation fields to achieve the purpose of removing key information.

[0004] However, this method of removing key information in data by manually adding newly defined annotation fields requires a high labor cost, and moreover, the latency for removing new keyword entries in the data is relatively high. Summary of the Invention

[0005] This application provides a data processing method to reduce the labor cost of removing key information in data and improve the efficiency of removing new key information in data. In addition, this application also provides a processor, a computing device, a computer-readable storage medium, and a computer program product.

[0006] In a first aspect, this application provides a data processing method, which can be executed by a corresponding data processing apparatus. Specifically, the data processing apparatus obtains a second policy set, which is obtained by updating a first policy set based on first data and a first processing result. The first data may be, for example, log data, or data in a database, etc. The first processing result is obtained by removing keyword entries in the first data based on the first policy set. The keyword entries may be words, phrases, numerical values, letters, etc. with key content. The second policy set includes at least one processing policy; and the data processing apparatus also obtains second data, and uses the second policy set to remove keyword entries in the second data to obtain a second processing result.

[0007] Since the data processing device can dynamically update the processing policies in the policy set according to the processing results of partial data without manually adjusting the processing policies in the policy set, this can not only effectively reduce the cost of updating the processing policies, that is, reduce the labor cost of removing keyword entries from the data, but also the data processing device can complete the update of the processing policies in a short period of time, thereby improving the efficiency of removing new keyword entries from the data. Moreover, in the process of removing keyword entries from the data, the data processing device will dynamically update the processing policies according to the first data and the first processing result, which enables the processing policies in the updated second policy set to identify the remaining keyword entries in the first processing result that have not been removed. Thus, when the second data also includes this keyword entry, the keyword entry in the second data can be removed by using the processing policies in the updated second policy set. In this way, the dynamically updated processing policies can adapt to the demand changes in the actual application scenario for removing keyword entries from the data, so as to keep the accuracy of removing keyword entries from the data at a high level.

[0008] In a possible implementation manner, the data processing device can also measure the removal degree corresponding to the first processing result according to the first data and the first processing result, where the removal degree is used to measure the degree of removing keyword entries from the first processing result; when the removal degree is lower than the first threshold, the data processing device uses a classifier to identify the first keyword entries in the first processing result and generates target processing policies for the first keyword entries, so that the data processing device can update the first policy set by using the target processing policies. In this way, by measuring the removal effect of the first policy set on the keyword entries in Data 1, the processing policies in the first policy set can be updated when the removal effect is poor, so as to adapt to the demand changes in the actual application scenario for removing keyword entries from the data, and thus the accuracy of removing keyword entries from the data can be kept at a high level.

[0009] In a possible implementation, when the data processing device measures the removal degree corresponding to the first processing result according to the first data and the first processing result, specifically, it can determine the usability index, relevance index, authenticity index, and reproducibility index according to the first data and the first processing result. Among them, the usability index can be the number of analysis requirements of application 100 that the data after removing the keyword entries can meet. The relevance index can be the number of associations maintained between the data and other data after the keyword entries are removed. The authenticity index can be the degree to which the characteristics of the processing result are consistent with the characteristics of the data before the keyword entries are removed. The reproducibility index can be the consistency of the processing results obtained by removing the same keyword entries. At the same time, the data processing device can also use the differential privacy algorithm to determine the privacy budget according to the first data and the first processing result. Thus, the data processing device can calculate the removal degree corresponding to the first processing result according to the usability index, relevance index, authenticity index, reproducibility index, and privacy budget. In this way, the data processing device can implement the measurement of the high or low removal degree based on the above process, so as to update the processing strategy in the first policy set in time when it is determined that the removal degree is low.

[0010] In a possible implementation, the data processing device can also identify the second keyword entries in the third data, classify the second keyword entries in the second data using a classification model to obtain the category to which the second keyword entries belong, and grade the second keyword entries in the second data using a grading model to obtain the key level corresponding to the second keyword entries. Thus, the data processing device can determine the identifier of the processing algorithm corresponding to the second keyword entries according to the category to which the second keyword entries belong and the key level corresponding to the second keyword entries, and generate a processing strategy in the first policy set. The processing strategy includes the second keyword entries, the recognition rule, and the identifier of the processing algorithm. In this way, the data processing device can automatically generate a processing strategy to use the processing strategy to remove the keyword entries in the data.

[0011] In a possible implementation, when the data processing device identifies the second keyword entries in the third data, specifically, it can match the words in the third data with the keyword entries in the thesaurus to obtain the second keyword entries in the third data, and the similarity between the word vectors of the second keyword entries and the word vectors of the keyword entries is higher than the second threshold; or, the data processing device can identify the named entities in the third data to obtain the second keyword entries in the third data, and the second keyword entries are the named entities in the third data that meet the verification rules. In this way, the data processing device can identify the keyword entries in the data to automatically generate a processing strategy based on the identified keyword entries.

[0012] In a possible implementation, when the data processing device obtains the second data, specifically, it may obtain a first database statement. The first database statement includes the second data, and the first database statement is used to instruct to write the second data into the database. Thus, after the data processing device obtains the second processing result, it can also use the second processing result to rewrite the first database statement to obtain a second database statement. The second database statement is used to instruct to write the second processing result into the database. Then, the data processing device can execute the second database statement to implement writing the second processing result into the database. In this way, the data processing device can write the data with the keyword strips removed into the database by rewriting the database statement, so as to improve the security of data storage.

[0013] In a possible implementation, when the data processing device obtains the second data, specifically, it may obtain a third database statement. The third database statement is used to instruct to query the second data in the database, and according to the third database statement, access the database to obtain the second data. Thus, the data storage device can also present an interactive interface, and the interactive interface includes the second processing result. In this way, the data processing device can present the data with the keyword strips removed to the user, so as to improve the security of data query.

[0014] In a possible implementation, the data processing device can also use the second policy set to remove the keyword strips in the first processing result to obtain a third processing result. In this way, through further processing of removing keyword strips from the first data (that is, processing of removing keyword strips from the first processing result), the first data with better removal effect can be obtained.

[0015] In a possible implementation, the first data is the first log, the second data is the second log, both the first log and the second log belong to the first log category, and both the first policy set and the second policy set correspond to the first log category. In this way, the data processing device can dynamically update the processing policies in the policy set according to the processing results of some logs, without manually adjusting the processing policies in the policy set. This can not only effectively reduce the cost of updating the processing policies, that is, reduce the labor cost of removing keyword entries from the logs, but also the data processing device can complete the update of the processing policies in a short period of time, so as to improve the efficiency of removing new keyword entries from the logs. Moreover, since the data processing device dynamically updates the processing policies according to the logs and their corresponding first processing results during the process of removing keyword entries from the logs, the processing policies in the updated second policy set can identify the remaining keyword entries that have not been removed in the first processing results. Thus, when the newly obtained log also includes this keyword entry, the data processing device can use the processing policies in the updated second policy set to remove the keyword entry in the second log. In this way, the dynamically updated processing policies can adapt to the changing requirements for removing keyword entries from logs in the actual application scenario, so as to keep the accuracy of removing keyword entries from logs at a high level.

[0016] In a possible implementation, the data processing device can also obtain a fourth policy set corresponding to the second log category. The fourth policy set is obtained by updating the third policy set based on the third log and the fourth processing result. The third processing result is obtained by removing the keyword entries from the third log based on the third policy set corresponding to the second log category. The third policy set includes at least one processing policy, and the third log belongs to the second log category. And the data processing device will also obtain a fourth log, which belongs to the second log category. Thus, the data processing device can use the fourth policy set to remove the keyword entries from the fourth log to obtain the fifth processing result. In this way, for logs of different log categories, different processing policies in the policy set can be used to remove the keyword entries from the log, which helps to improve the effect of removing keyword entries for logs of a specific category. For example, for the operation behavior data in the log belonging to log category A, based on the processing policies in the first policy set, the operation behavior data in the log may not be removed. For example, in some actual application scenarios, the operation behavior of the administrator may not be considered as a keyword entry. For the operation behavior data in the log belonging to log category B, based on the processing policies in the second policy set, the operation behavior data in the log can be removed. For example, in some actual application scenarios, the operation behavior of the user may be considered as a keyword entry. In this way, differential removal of keyword entries for different categories of logs can be realized, and the flexibility of removing keyword entries from logs can be improved.

[0017] In a possible implementation, the second log includes multiple first data records. When the data processing device uses the second policy set to remove keyword entries from the second log to obtain a second processing result, specifically, it can first obtain a keyword library, where the keyword library includes at least one keyword, and use the keyword library to filter the multiple first data records to obtain at least one second data record, and the at least one second data record includes at least one keyword in the keyword library. Thus, the data processing device can use the second policy set to remove keyword entries from at least one second data record to obtain a second processing result. In this way, the data processing device filters the multiple first data records in the log by using the keyword library, reducing the number of data records traversed by the processing policy for keyword entry matching. Thereby, the number of data records from which the data processing device removes keyword entries from the log can be reduced, and resource consumption can be reduced.

[0018] In a second aspect, the present application provides a data processing device, which includes various modules for executing the data processing method in the first aspect or any possible implementation manner of the first aspect.

[0019] In a third aspect, the present application provides a computing device, which includes a processor and a memory. The processor and the memory communicate with each other. The processor is used to execute instructions stored in the memory so that the computing device executes the data processing method in the first aspect or any implementation manner of the first aspect. It should be noted that the memory can be integrated into the processor or can be independent of the processor. The computing device may further include a bus. Among them, the processor is connected to the memory through the bus. Among them, the memory may include a readable memory and a random access memory.

[0020] In a fourth aspect, the present application provides a computer-readable storage medium, in which instructions are stored. When it runs on a computing device, it causes the computing device to execute the operation steps of the data processing method described in the first aspect or any implementation manner of the first aspect.

[0021] In a fifth aspect, the present application provides a computer program product containing instructions. When it runs on a computing device, it causes the computing device to execute the operation steps of the data processing method described in the first aspect or any implementation manner of the first aspect.

[0022] Based on the implementation manners provided in the above aspects, the present application can be further combined to provide more implementation manners. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1 It is a schematic structural diagram of an exemplary data processing system provided by the present application;

[0024] Figure 2 A flowchart showing a data processing method provided by this application;

[0025] Figure 3 A flowchart showing another data processing method provided by this application;

[0026] Figure 4 A schematic structural diagram of a data processing device provided by this application;

[0027] Figure 5 A schematic hardware structure diagram of a computing device provided by this application. Detailed implementation manners

[0028] In order to improve the processing efficiency of removing key information (which can also be referred to as keyword entries), this application provides a data processing method. During the process of removing keyword entries from multiple data, the data processing system can first use the existing processing strategy to remove keyword entries from some of the data, obtain the corresponding processing results, and use this part of the data and the processing results to update the processing strategy. Thus, the updated processing strategy can be used to remove keyword entries from the remaining part of the data, so as to adapt to the changing requirements for removing keyword entries in the actual application scenario (such as words in the data changing from non-keyword entries to keyword entries) by dynamically updating the processing strategy, and achieve the reduction of the labor cost for removing keyword entries from the data and the improvement of the efficiency of removing new keyword entries from the data.

[0029] The technical solutions in this application will be described below with reference to the accompanying drawings provided by this application.

[0030] Refer to Figure 1 , which shows a schematic structural diagram of a data processing system. As Figure 1 shown, the data processing system 10 includes an application 100, a data processing device 200, and a storage device 300.

[0031] The application 100, which is a software program for implementing specific functions, generates new data during operation and writes the new data into the storage device 300. Moreover, the application 100 can also be used to access the data in the storage device 300. Among them, the data generated / accessed by the application 100 may contain keyword entries. Herein, the keyword entry refers to an entry that needs to be concerned, such as for explanation, attribute identification, or prompting, etc. in the actual application scenario, and can be a character, word, numerical value, letter, etc. The present application does not limit the specific manifestation form of the entry. For example, if the data generated by the application 100 is specifically "The identifier (ID) number of Zhang San is 111111111111111111", then "111111111111111111" in this data is the keyword entry. Exemplarily, the application 100 can be an application running on a user device, or can be a database application / database driver program, or can be an application running on an application server, such as a big data application, etc., and this is not limited.

[0032] The data processing device 200 is used to remove the keyword entries in the data generated / accessed by the application 100, such as removing keyword entries such as "111111111111111111" in the data "The identifier number of Zhang San is 111111111111111111". After removing the keyword entries, the obtained data is "The identifier number of Zhang San is XXXX", etc. Specifically, for the data generated by the application 100, the data processing device 200 can remove the keyword entries in the data and write the obtained processing result (i.e., the data after removing the keyword entries) into the storage device 300 to improve the security of data storage in the storage device 300, that is, to prevent the keyword entry from having low data security due to data leakage during storage in the storage device 300. Or, for the data requested by the application 100 to be accessed from the storage device 300 (the keyword entries in the data are not removed when stored in the storage device 300), the data processing device 200 can remove the keyword entries in the data, obtain the corresponding processing result, and feedback the processing result to the application 100 so that the application 100 can subsequently present the processing result to the user to improve the security when the user queries the data, that is, to prevent the keyword entry from being obtained by the user's query and resulting in low data security.

[0033] A storage device 300 is used to store data, such as the data generated during the operation of the application 100. Exemplarily, the storage device 300, for example, can be a hardware device with data storage capabilities such as a memory or a server. And, a database can be constructed based on the storage device 300. This database can be, for example, a MySQL database, a MongoDB database, an Elasticsearch (ES) database, etc., or it can be other types of databases.

[0034] In actual application scenarios, as the business requirements in the application 100 change, the need to remove keyword entries from the data also changes. For example, assume that the content of the data specifically records the operation behavior of user "Li Si", such as "Li Si browsed webpage A at 18:00:00", etc. And, as the business requirements of the application 100 change, the personal behavior data of the user "browsing webpage A" will become a keyword entry. At this time, the data processing device 200 can adapt to the change in the need to remove keyword entries from the log by dynamically updating the processing strategy, and can also ensure that the accuracy of removing keyword entries from the data can be maintained at a relatively high level.

[0035] Specifically, when the data processing device 200 processes data, it will obtain a second policy set. This second policy set is obtained by updating the first policy set in advance based on the first data and the first processing result. The first processing result is obtained by removing keyword entries from the first data based on the processing policies in the first policy set. The second policy set includes at least one processing policy. And, the data processing device 200 will also obtain the second data to be processed. In this way, after obtaining the second data, the data processing device 200 can use the updated second policy set to remove keyword entries from the second data to obtain the corresponding second processing result.

[0036] Since the data processing device 200 can dynamically update the processing policies in the policy set according to the processing results of partial data, without manual adjustment of the processing policies in the policy set, this can not only effectively reduce the cost of updating the processing policy, that is, reduce the labor cost of removing keyword entries from the data, but also the data processing device 200 can complete the update of the processing policy in a short period of time, thereby improving the efficiency of removing new keyword entries from the data.

[0037] Since, in the process of removing keyword entries from data, the data processing device 200 dynamically updates the processing strategy according to the first data and the first processing result, the processing strategies in the updated second strategy set can identify the remaining keyword entries that have not been removed in the first processing result. Thus, when the second data also includes such keyword entries, the keyword entries in the second data can be removed by using the processing strategies in the updated second strategy set. In this way, the dynamically updated processing strategy can adapt to the changing requirements for removing keyword entries from data in the actual application scenario, and thus can keep the accuracy of removing keyword entries from data at a relatively high level.

[0038] Exemplarily, the data processing device 200 can be implemented by software or hardware.

[0039] In the first example, when implemented by software, the data processing device 200 can be implemented, for example, by at least one of a virtual machine, a container, a component, and a plugin. For example, the data processing device 200 can be specifically implemented in the form of a software component such as a component. Moreover, the application 100 can load and run this component to achieve data removal, thereby reducing or even avoiding the need to modify the application 100 and lowering the development difficulty of the application 100. At this time, the operations performed by the data processing device 200 can be regarded as the application 100 performing these operations.

[0040] In the second example, when implemented by hardware, the data processing device 200 can be implemented by at least one physical device including a processor, such as a server. Among them, the processor can be a central processing unit (CPU), or can be an application-specific integrated circuit (ASIC), a programmable logic device (PLD), a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), a system on chip (SoC), a software-defined infrastructure (SDI) chip, an artificial intelligence (AI) chip, a data processing unit (DPU), or any combination of any of these processors. Moreover, the number of processors included in the data processing device 200 can be any number, and the types of processors included can be one or more. Specifically, the number and types of processors can be set according to the business requirements of the actual application, and the present application does not limit this.

[0041] In addition, the application 100 and the data processing device 200 can be integrally deployed. For example, when the data processing device 200 is specifically a component, the application 100 and the data processing device 200 can be deployed in the same application server, or the same terminal, or the same processor, etc. Or, the application 100 and the data processing device 200 can be separately deployed. For example, when the data processing device 200 is implemented by hardware, the application 100 can be deployed on the application server, and the application 100 can communicate with the data processing device 200 implemented by hardware through this application server.

[0042] It should be noted that Figure 1The data processing system 10 shown is only for illustrative purposes and is not intended to be limiting. For example, in an actual application scenario, the data processing system 10 may include multiple applications, that is, the data processing device 200 may provide data removal services for multiple applications; or, the data processing system 10 may further include a log database for storing logs, so that after removing the data in the logs, the data processing device 200 can write the logs into the log database, etc. The present application does not limit the specific architecture of the data processing system 10.

[0043] For ease of understanding, the data processing method provided in the present application will be described in detail below with reference to the accompanying drawings.

[0044] See Figure 2 , Figure 2 is a schematic flowchart of a data processing method provided in the present application. This method can be applied to Figure 1 the data processing system 10 described above, or can be applied to other applicable data processing systems. For ease of description, in the present application, it is exemplified by being applied to Figure 1 the data processing system 10 shown.

[0045] Among them, Figure 2 the data processing method shown can specifically include:

[0046] S201: Application 100 generates data 1 during operation and provides the data 1 to the data processing device 200.

[0047] In this embodiment, during operation, application 100 can generate data 1, and this data 1 needs to be saved to the storage device 300. The data 1 generated by application 100 may include keyword entries. Exemplarily, the keyword entries may be personal information, such as information about a user's height, home address, etc.; or, the keyword entries may be business information, such as account numbers, passwords, etc.; or, the keyword entries may be specific types of key information, such as information about a user's cooperation manufacturers, etc. This embodiment does not limit this.

[0048] Moreover, the data generated by application 100 may be log data, or may be business data that needs to be written into the storage device 300, etc., and this is not limited.

[0049] S202: The data processing device 200 uses the policy set 1 to remove the keyword entries in the data 1 to obtain a processing result 1, where the policy set 1 includes at least one processing policy.

[0050] Among them, a processing policy refers to a policy adopted when removing keyword entries included in data.

[0051] In Implementation Mode 1, the processing strategy may include a keyword, an identification rule, and an identifier of a processing algorithm. Among them, the identification rule is used to identify whether the keyword is included in the data, and the processing algorithm is used to remove the keyword from the data after it is identified that the keyword is included in the data. For example, the processing strategy may be "username-keyword identification rule-anonymization processing algorithm identifier". Then, when the data includes a username, the keyword "username" included in the data can be identified through the "keyword identification rule" in this processing strategy, and some or all characters in the username can be replaced with specified characters, such as replaced with "******", etc. Another example is that the processing strategy may be "transaction amount-regular expression identification rule-offset rounding processing algorithm identifier". Then, when the transaction amount is included in the data, the keyword "transaction amount" included in the data can be identified through the "regular expression identification rule" in this processing strategy, and the transaction amount can be adjusted by using the offset rounding processing algorithm indicated by this processing strategy. For example, the transaction amount "1234" is offset to 12.34, and the offset transaction amount is adjusted to "12", etc.

[0052] Among them, different processing strategies may include different keywords, or different identification rules, or identifiers of different processing algorithms, and this is not limited.

[0053] In Implementation Mode 2, the processing strategy may only include a keyword and an identifier of a processing algorithm, and this processing algorithm is used to remove the keyword identified from the data. For example, the processing strategy may be "username-anonymization processing algorithm identifier", "transaction amount-offset rounding processing algorithm identifier", etc. At this time, the data processing device 200 can identify the keyword in the data based on the built-in identification rule.

[0054] In this embodiment, it is illustrated by taking Data 1 including a keyword as an example. In actual application, when Data 1 does not include a keyword, the subsequent process of removing the keyword does not need to be executed. After obtaining Data 1, the data processing device 200 can use the identification rule in the processing strategy or the built-in identification rule of the data processing device 200 to identify the keyword included in Data 1.

[0055] Some exemplary identification rules are introduced below.

[0056] In specific implementation, the data processing device 200 can use one or more pre-configured identification rules to identify keywords in Data 1. Some exemplary identification rules are introduced below.

[0057] Example 1: The keyword recognition rule means that data 1 is first segmented into multiple words. For example, data 1 can be segmented into multiple words based on a word dictionary, etc. Then, the multiple words obtained by segmentation are respectively matched with the keywords in a preset thesaurus. The matching methods include exact matching, fuzzy matching, etc. Moreover, when some words in data 1 match the keywords in the thesaurus successfully, the data processing device 200 can determine that the word is a keyword entry to be removed.

[0058] Example 2: The regular expression recognition rule means that data 1 is matched with a preset regular expression. Among them, a regular expression, which can also be called a rule expression, is used to describe and match a series of strings with the same syntactic rules using a single string. The task processing device 200 can determine whether the character arrangement in the data conforms to the character arrangement requirements corresponding to the regular expression. If so, it is determined that the data matches the regular expression. For example, assume the data is "tel:086-0666-88810009999, quite easy to remember", and the regular expression can be "^tel:[0-9]{1,3}-[0][0-9]{2,3}-[0-9]{8,11}$". Since the character arrangement of "tel:086-0666-88810009999" in the data conforms to the requirements of the regular expression for character arrangement, the task processing device 200 can determine that the data matches the regular expression. When the character arrangement in data 1 matches the regular expression, the data processing device 200 can determine the matched character content as the keyword entry to be removed.

[0059] Example 3: The semantic recognition rule means that the words in data 1 are semantically matched with the keywords in a preset thesaurus. Moreover, when the semantics of some words in data 1 match the semantics of the keywords in the thesaurus successfully, the data processing device 200 can determine that the word is a keyword entry to be removed. For example, the data processing device 200 can calculate the word vectors (word embedding) of each word in data 1. Moreover, for each word in data 1, the similarity between the word vector of the word and the word vectors of each keyword in the database can be calculated. When the similarity between the word vector of a keyword and the word vector of a word in data 1 is greater than the threshold, the word is determined as the keyword entry to be removed. In this way, keyword entries with different characters but the same semantics included in data 1 can be recognized, which helps to improve the subsequent removal effect of keyword entries in data 1.

[0060] Example 4: The named entity recognition rule refers to using the named entity recognition (NER) technology to recognize the named entities that are used as keyword entries in Data 1. Among them, a named entity refers to an entity with a specific meaning, such as a person's name, a place name, an organization name, a proper noun, etc. Moreover, for the named entities recognized from Data 1, the data processing device 200 can further determine whether the named entity meets the preset verification rules, and determine the named entity that meets the verification rules as the keyword entry to be removed. For example, when Data 1 includes a proper noun, the data processing device 200 can further determine whether the proper noun is a drug name, and when it is determined that the proper noun is a drug name, determine that the proper noun is a keyword entry. In this way, through the named entity recognition technology and the verification rules, the named entities that belong to the keyword entries in Data 1 can be accurately recognized, which helps to improve the removal effect of the keyword entries in Data 1 subsequently.

[0061] In actual application, the data processing device 200 can also use other recognition rules to recognize the keyword entries in Data 1, and this embodiment does not limit this.

[0062] In the process of recognizing the keyword entries in Data 1, when the processing strategy includes a recognition rule, the data processing device 200 can use the recognition rules indicated by each processing strategy included in the strategy set 1 to recognize whether Data 1 contains the keyword entries included in the processing strategy. Specifically, it can be to recognize whether there are characters in Data 1 that match the keyword entries included in the processing strategy. When it is determined according to the recognition rule that there are characters in Data 1 that match the keyword entries included in the processing strategy, the data processing device 200 can determine that Data 1 includes keyword entries. When there are no characters in Data 1 that match the keyword entries included in the processing strategy, the data processing device 200 can determine that Data 1 does not include the keyword entries in the processing strategy. In this way, the data processing device 200 can recognize the existing keyword entries in Data 1 based on each processing strategy in the strategy set 1.

[0063] When the processing strategy does not include a recognition rule, the data processing device 200 can recognize the keyword entries included in Data 1 based on one or more built-in recognition rules. Among them, the recognition rules built in the data processing device 200 can be any one or more of the 4 algorithms in the above example.

[0064] Further, before identifying the keyword entries in Data 1, the data processing device 200 may also preprocess Data 1, such as removing useless data or invalid data in the log, which may be deleting the error content in Data 1, deleting consecutive duplicate content in Data 1, etc. In this way, the data processing device 100 processes the preprocessed Data 1 to remove keyword entries, which can improve the quality of Data 1 and reduce the amount of resources required for subsequent data removal.

[0065] After identifying the keyword entries included in Data 1, the data processing device 200 may, according to at least one processing policy included in Policy Set 1, adopt the processing algorithm indicated by the processing policy for various keyword entries in Data 1 to remove the keyword entries in Data 1, thereby obtaining the corresponding Processing Result 1. Exemplarily, the processing algorithm indicated by the processing policy may be, for example, an "anonymization" processing algorithm for replacing the entries in the data with specified characters. Or, the processing algorithm indicated by the processing policy may be, for example, an "offset and rounding" algorithm for offsetting and rounding the magnitude of a numerical value. For example, assume that Data 1 is specifically "Li Si purchased a financial product worth 1560 yuan with his ID number 111111111111111111", where the keyword entries include "111111111111111111" and "1560". Then, based on the processing policy in Policy Set 1, the data processing device 200 uses the "anonymization" processing algorithm to replace the keyword entry "111111111111111111" with "111111********1111" (i.e., replacing some characters with the specified character "*"), uses the offset and rounding algorithm to offset the keyword entry "1560" to "15.60", and further rounds "15.60" to "15". Thus, the obtained Processing Result 1 is "Li Si purchased a financial product worth 15 yuan with his ID number 111111********1111".

[0066] After generating Processing Result 1, the data processing device 200 may also write this Processing Result 1 (replacing Data 1) into the storage device 300, as Figure 2 shown. In this way, the Data 1 written into the storage device 300 is the Data 1 after removing the keyword entries, thereby reducing or avoiding the risk that keyword entries are obtained by illegal users when the data in the storage device 300 is leaked.

[0067] Among them, the data processing device 200 may obtain Policy Set 1 based on the following two non-limiting implementation examples.

[0068] Example 1. Policy set 1 can be pre-configured in the data processing device 200 by a technician. For example, based on work experience, the technician can configure a processing policy for a possible keyword entry in the data processing device 200.

[0069] Example 2. During the process of processing data, the data processing device 200 can automatically generate a policy set including Policy set 1 according to the data being processed. In specific implementation, before removing the keyword entry in Data 1 by the data processing device 200, other data can be obtained in advance. Taking the acquisition of Data A as an example. Then, the data processing device 200 can use the built-in recognition rules to recognize the keyword entries in Data A and classify the keyword entries in Data A to obtain the categories to which the keyword entries in Data A belong. Among them, the categories to which the keyword entries belong can be, for example, personal data categories, business data categories, or key data categories. Among them, the personal data category refers to the data category related to personal information, such as ID number, home address, etc.; the business data category refers to the data category related to business, such as transaction amount, encryption algorithm name, key, etc.; the key data category refers to specific types of key data, such as the name of the user's cooperation manufacturer, etc.

[0070] Exemplarily, the data processing device 200 can input the keyword entries in Data A into a pre-trained classification model. The classification model can be, for example, a model constructed based on a neural network, or a model constructed based on a classifier, etc. Then, the classification model can output the probabilities of the keyword entry belonging to each category respectively, so that the category corresponding to the maximum probability can be used as the category of the keyword entry. It should be noted that when there are multiple different keyword entries in Data A, the data processing device 200 can input different keyword entries into the classification model in turn to obtain the categories to which each keyword entry belongs.

[0071] Moreover, the data processing device 200 can also grade the keyword entries in Data A to obtain the key levels corresponding to the keyword entries in Data A. Different key levels correspond to different degrees of criticality. For example, four key levels, S1, S2, S3, and S4, can be defined, and the degree of criticality decreases in turn. For the keyword entry of the transaction amount type, it can be determined that the level of this keyword entry is the highest level S1, that is, the degree of criticality is the highest, and the corresponding security risk of data leakage is the highest; while for the keyword entry of the operation behavior type, it can be determined that the level of this keyword entry is the lowest level S4, that is, the keyword entry is the lowest, and the corresponding security risk of data leakage is the lowest.

[0072] Taking the specific case where the keyword entries in data A are multiple keywords as an example, the data processing device 200 can first calculate the word vectors of each keyword in data A. For example, the word vectors of each keyword can be calculated through word embedding technology, etc., and use an aggregation algorithm to cluster multiple keywords. For example, the k-means clustering algorithm is used to cluster multiple ID numbers in data A into the same type of data, and multiple transaction amounts are clustered into the same type of data, etc. Then, the data processing device 200 can determine the corresponding key levels for each keyword according to the pre-configured mapping relationship between the data type and the key level (such as pre-configured by technicians, etc.). In actual application, when the data type obtained by clustering does not belong to the data type in the mapping relationship, the data processing device 200 can input the keywords under this data type (as well as other keywords whose key levels have been determined) into the pre-trained classification model, and obtain the key level inferred by the classification model for the keywords under this data type. Among them, the classification model can be, for example, a model constructed based on a neural network, etc. In actual application, the data processing device 200 can also directly determine the key levels corresponding to each keyword according to the semantics of multiple keywords in data A. For example, based on the pre-configured mapping relationship between semantics and key levels, the key levels corresponding to each keyword can be determined, etc.

[0073] Then, the data processing device 200 can determine the processing algorithm corresponding to the keyword entry according to the category to which the keyword entry in data A belongs and the key level corresponding to the keyword entry. The processing algorithm refers to the algorithm used to remove the keyword entry from the data. For example, it can be an "anonymization" processing algorithm or an offset rounding algorithm, etc. Exemplarily, the data processing device 200 can be pre-configured with a mapping relationship between the category, the key level, and the processing algorithm. Thus, the data processing device 200 can determine the processing algorithm applicable to the keyword entry in data A by looking up this mapping relationship. For example, for the keyword entry of the personal data category at the S1 key level, it can be determined that the applicable processing algorithm is the "anonymization" processing algorithm, that is, using other characters to replace some / all of the characters in the keyword entry; for the amount transaction information of the business data category at the S2 key level, it can be determined that the applicable processing algorithm is the offset rounding algorithm, that is, performing an offset process on the numerical value of the transaction amount and then rounding it. For example, if the transaction amount in data A is 1234, the amount obtained after offsetting the transaction amount can be 12.34, and then the amount obtained after rounding it is 12. In actual application, the data processing device 200 can determine to use other types of processing algorithms to remove the keyword entries in data A, and this is not limited.

[0074] In this way, the data processing device 200 can generate at least one corresponding processing policy according to the keyword entries in Data A, the processing algorithm corresponding to the keyword entry, and the recognition rule for identifying the keyword entry. For example, it can splice the three pieces of information: the keyword entry, the identifier of the processing algorithm, and the recognition rule to obtain an information string as the processing policy, thereby generating the processing policy. At this time, each processing policy can include the keyword entry, the recognition rule, and the identifier of the processing algorithm. Different processing policies include different keyword entries, and thus a policy set 1 including the at least one processing policy can be obtained.

[0075] Further, after generating the processing result 1, the data processing device 200 can write the processing result 1 (replacing Data 1) into the storage device 300, as Figure 2 shown. In this way, the data written into the storage device 300 is the processed Data 1, thereby reducing or avoiding the risk that the keyword entry is obtained by an illegal user when the data in the storage device 300 is leaked.

[0076] S203: The data processing device 200 updates the policy set 1 according to Data 1 and the processing result 1 to obtain a policy set 2, and the policy set 2 includes at least one processing policy.

[0077] Among them, the updated policy set 2 can remove the keyword entries in the processing result 1.

[0078] In an actual application scenario, after generating the policy set 1, the data processing device 200 can use the processing policies in the policy set 1 to remove other newly received data. Or, the data processing device 200 can refer to the above similar method. For the newly received data, it can generate new processing policies to remove the keyword entries in the data.

[0079] Since the processing policies in the policy set 1 may not be able to completely remove all the keyword entries in Data 1. For example, assume that the processing result 1 is specifically the above "Li Si purchased a financial product worth 15 yuan with his ID number 111111********1111". However, in some scenarios, "financial product" can also be used as a keyword entry, but based on the processing policies in the policy set 1, the "financial product" in Data 1 cannot be removed.

[0080] Therefore, the data processing device 200 can also dynamically update the policy set 1 according to the removal situation of Data 1, so as to use the updated new policy set to remove Data 1 or other data, thereby improving the effect of the data processing device 200 in removing the keyword entries in the data.

[0081] As an implementation example, the data processing device 200 can measure the degree of removal corresponding to the processing result 1 based on the data 1 and the processing result 1. The degree of removal refers to the degree of removing the keyword entries (or key content) in the data 1. Among them, the higher the degree of removal, the better the removal effect, that is, the fewer the keyword entries (or the less the key content) included in the processing result 1.

[0082] Specifically, the data processing device 200 can calculate the usability index, the relevance index, the authenticity index, and the reproducibility index based on the data 1 and the processing result 1. For ease of understanding, in this embodiment, it is described by taking the data 1 as an example that can include multiple data.

[0083] Among them, the usability index can specifically be the number of analysis requirements of the application 100 that can be satisfied by the data after removing the keyword entries. That is, when the data can still execute the corresponding analysis process based on the data after removing the keyword entries, such as the analysis process of data recognition, etc., at this time, the data processing device 200 can increment the number of the usability index by 1. In this way, the data processing device 200 can count the multiple data after removing the keyword entries to obtain the value of the usability index.

[0084] The relevance index can specifically be the number of associations maintained between the data and other data after the keyword entries are removed. For example, if field A in the same data table is removed and still can be associated with field B, the value of the relevance index can be incremented by 1. In this way, the data processing device 200 can analyze and count the multiple data after removing the keyword entries to obtain the value of the relevance index.

[0085] The authenticity index can specifically be the degree of consistency between the characteristics of the processing result and the characteristics of the data before removing the keyword entries. For example, the numerical distribution and change trend in the data before removing the keyword entries, and the degree of consistency between the numerical distribution and change trend in the processing result, etc. In this way, the data processing device 200 can calculate the degree of difference between the numerical distribution and change trend before and after removing the keyword entries to calculate the value of the authenticity index.

[0086] The reproducibility index can specifically be the consistency of the processing results obtained by removing the keyword entries for the same keyword entries. In this way, the data processing device 200 can count and calculate the degree of consistency between the removal results of the same keyword entries in multiple data to obtain the value of the reproducibility index.

[0087] Moreover, the data processing device 200 can determine a privacy budget according to the differential privacy algorithm based on Data 1 and Processing Result 1. Thus, the data processing device 200 can calculate the degree of removal corresponding to Processing Result 1 according to the privacy budget and the calculated usability index, relevance index, authenticity index, and reproducibility index.

[0088] For example, the data processing device 200 can measure the degree of removal corresponding to Processing Result 1 through the model shown in the following formula (1).

[0089] E = a * ε + x1 * b + x2 * c + x3 * d + x4 * e Formula (1)

[0090] Among them, E is the degree of removal corresponding to Processing Result 1; a is the number of keyword entries (such as keyword entries, etc.) that have been removed in Data 1; b, c, d, and e are the usability index, relevance index, authenticity index, and reproducibility index in sequence; x1, x2, x3, and x4 are the weights of each index respectively, and the sum of x1, x2, x3, and x4 is 1. ε is the privacy budget, and the data processing device 200 can calculate the value of ε according to the differential privacy algorithm based on Data 1 and Processing Result 1. Among them, the smaller the value of ε, the more noise is added. Correspondingly, the risk of keyword entry leakage is lower, the protection intensity of data security is higher, and the data removal effect is better. Exemplarily, the value of ε can be 0.1, and the values of x1, x2, x3, and x4 can be 0.25, 0.25, 0.25, 0.25, etc.

[0091] When the calculated value of E is greater than or equal to Threshold 1, it indicates that the degree of removal corresponding to Processing Result 1 is relatively high, and the removal effect of Policy Set 1 on Data 1 is better. When Data 1 includes 100 pieces of data, Threshold 1 for judging the level of removal effect can be, for example, 500. Among them, Threshold 1 can be pre-configured in the data processing device 200 by technical personnel. For example, technical personnel can determine the value of Threshold 1 according to the technical experience of removing keyword entries in the actual application scenario, or Threshold 1 in the data processing device 200 can also be determined by other means, such as being automatically determined by the data processing device 200 through self-learning, etc., and this is not limited.

[0092] When the calculated value of E is less than threshold 1, it indicates that the removal degree corresponding to processing result 1 is relatively low, and the removal effect of policy set 1 on data 1 is poor. For the convenience of description, in this embodiment, it is set that the removal degree corresponding to processing result 1 is lower than threshold 1. At this time, data processing device 200 can execute identifying the keyword entries still included in processing result 1 by using a classifier. For example, it can use a naive bayes (NB) classifier or other classifiers to identify the keyword entries still included in processing result 1. Thus, data processing device 200 can generate a target processing policy for this keyword entry, and use this target processing policy to update policy set 1, including adding the target processing policy to policy set 1, or using this target processing policy to replace the existing processing policy in policy set 1, etc.

[0093] Taking the keyword entries included in processing result 1 being specific keyword entries as an example, data processing device 200 can split processing result 1 into multiple words. For example, it can split processing result 1 into multiple words based on a word dictionary, etc., and filter the non-keyword entries among the multiple split words. For example, filter words such as "le" and "de" in processing result 1 that obviously do not belong to keyword entries. Then, data processing device 200 can input the filtered processing result 1 into the classifier, and the classifier classifies each word in processing result 1 and outputs the classification result of each word. Among them, the classification result of each word includes whether the word belongs to a keyword entry or a non-keyword entry, and the classification result also includes a confidence level. Exemplarily, the classifier can be implemented by a neural network model. Thus, after the neural network model makes inferences based on the input multiple words, the inference result (i.e., the classification result) output not only includes the indication information for whether it belongs to a keyword entry, but also can include the confidence level used to indicate the credibility of the inference result. This confidence level can be calculated by a network layer in the neural network model.

[0094] Then, data processing device 200 can update the processing policies in the policy set according to the classification result of each word. The process of updating the processing policies in the policy set will be introduced separately for different situations below.

[0095] Case 1: When the classification results of some words indicate that the words belong to keyword entries and the confidence level is higher than threshold 2, the data processing device 200 can determine these words as keyword entries. At this time, the data processing device 200 can generate a new processing strategy for the newly determined keyword entries. For example, it can obtain an information string as the processing strategy by splicing the newly determined keyword entries, the identifier of the processing algorithm for removing the keyword entries, and the recognition rule for identifying the keyword entries, so as to generate a new processing strategy. The newly generated processing strategy can include the keyword entries, the identifier of the processing algorithm for the keyword entries, and the recognition rule for identifying the keyword entries. Then, the data processing device 200 can add the newly generated processing strategy to Policy Set 1 to update Policy Set 1. Among them, threshold 2 can be pre-configured in the data processing device 200 by technical personnel. For example, technical personnel can determine the value of threshold 2 according to the technical experience of removing keyword entries from data in the actual application scenario, or the threshold 2 in the data processing device 200 can also be determined in other ways, such as being automatically determined by the data processing device 200 through self-learning, etc., and this is not limited.

[0096] Case 2: When the classification results of some words indicate that the words belong to keyword entries and the confidence level is higher than threshold 2, the data processing device 200 can determine these words as keyword entries. At this time, the data processing device 200 can generate a new processing strategy for the newly determined keyword entries. For example, it can obtain an information string as the processing strategy by splicing the newly determined keyword entries and the identifier of the processing algorithm for removing the keyword entries, so as to generate a new processing strategy. The newly generated processing strategy can include the newly determined keyword entries and the identifier of the processing algorithm for the keyword entries (excluding the recognition rule). Then, the data processing device 200 can add the newly generated processing strategy to Policy Set 1 to update Policy Set 1, and can add the newly determined keyword entries to the thesaurus to dynamically update the thesaurus, that is, to dynamically update the keyword entries that the data processing device 200 can recognize, so as to improve the removal effect of keyword entries in the data based on the updated thesaurus subsequently.

[0097] Case 3: When the classification results of some words indicate that the words belong to keyword entries and the confidence level is higher than threshold 2, and the processing strategy includes an identification rule, the data processing device 200 can determine the words as keyword entries. Moreover, the data processing device 200 can update the policy set 1 by adjusting the identification rule in the processing strategy in the policy set 1, such as adjusting the parameters in the identification rule. In this way, subsequently, the data processing device 200 can use the adjusted identification rule in the processing strategy to identify the newly determined keyword entries.

[0098] Case 4: When the classification results of some words indicate that the words belong to keyword entries and the confidence level is lower than threshold 2, these words are very likely to be the words that still retain some key content after the keyword entry removal process. Therefore, the data processing device 200 can match these words with the keyword entries included in the processing strategy in the policy set 1. Moreover, if these words match successfully with the keyword entries in the processing strategy A in the policy set 1, it indicates that the removal effect of the keyword entries using the processing strategy A is poor, that is, the data after removing the keyword entries still retains key content. At this time, the data processing device 200 can adjust the processing strategy A for the keyword entries to obtain a new processing strategy B, and the new processing strategy B can achieve a better removal effect for the keyword entries. Thus, the data processing device 200 can use the new processing strategy B to replace the processing strategy A in the policy set 1 to update the policy set 1.

[0099] For example, assume that after the data processing device 200 uses processing strategy A to remove the keyword entries from the home address "Room ee, Street dd, District cc, City bb, Province aa" in data 1, the obtained processing result 1 is "Province aa ******** Room ee", that is, "City bb, District cc, Street dd" in the home address is hidden. When the classifier classifies "Room ee" in processing result 1 and the output classification result indicates that "Room ee" belongs to the keyword entries and the confidence level is 75% (lower than threshold 2, assuming threshold 2 is 85%), the data processing device 200 can match this keyword entry with the keyword entries included in each processing strategy in strategy set 1, determine that "Room ee" matches the keyword entry in processing strategy A, indicating that the removal effect of processing strategy A on the home address is relatively low. Therefore, the data processing device 200 can adjust the processing algorithm used for the keyword entry in processing strategy A. For example, the algorithm used to remove intermediate data for the keyword entry in processing strategy A can be adjusted to an algorithm that completely removes the keyword entry, so as to obtain a new processing strategy B. At this time, using processing strategy B can remove "Room ee, Street dd, District cc, City bb, Province aa" to "************". Finally, the data processing device 200 can use this new processing strategy B to replace processing strategy A in strategy set 1 to update strategy set 1.

[0100] It should be noted that the above implementation process of updating strategy set 1 based on data 1 and processing result 1 is only an implementation example, and this embodiment does not limit this.

[0101] S204: During the operation of application 100, data 2 is generated and provided to data processing device 200.

[0102] Among them, the specific implementation process of application 100 generating data 2 is similar to the specific implementation process of application 100 generating data 1, and can be referred to the relevant descriptions above, and will not be elaborated here.

[0103] S205: The data processing device 200 uses strategy set 2 to remove the keyword entries from data 2, and obtains processing result 2.

[0104] Among them, the implementation process of the data processing device 200 using strategy set 2 to remove the keyword entries from the data is similar to the implementation process of removing the keyword entries from the data according to strategy set 1 in step 202 above, and can be referred to the relevant descriptions above, and will not be elaborated here. In this way, when data 2 includes the keyword entries that still exist in processing result 1, the updated strategy set (that is, strategy set 2) can be used to remove this part of the keyword entries in data 2, so as to improve the removal effect on data 2.

[0105] Further, after obtaining the processing result 2, the data processing device 200 can write the processing result 2 into the storage device 300, as Figure 2 shown. For example, when the storage device 300 is used to implement a database, the data processing device 200 can write the processing result 2 into the storage device 300 by rewriting the database statement generated by the application 100 for instructing to write the data 2 into the storage device 300, so as to implement retaining the data in the storage device 300 after the data is removed.

[0106] In actual application, the data processing device 200 can continuously update the policy set dynamically based on the above method, based on different data and different processing results. For example, the data processing device 200 can also continue to update the policy set 2 according to the data 2 and the corresponding processing result 2 of the data 2 to obtain the policy set 3, as Figure 2 shown. In this way, the subsequent data processing device 200 can use the policy set 3 to remove other newly received data, etc.

[0107] In this embodiment, since the data processing device 200 dynamically updates the processing policy according to the data 1 and the processing result 1 during the process of removing the keyword entries in the data, the updated processing policy (i.e., the processing policy in the policy set 2) can identify the remaining keyword entries that have not been removed in the processing result 1. Therefore, when the data 2 also includes this part of the keyword entries, this updated processing policy can be used to remove this part of the keyword entries in the data 2. In this way, the dynamically updated processing policy can adapt to the changing requirements for removing keyword entries in the actual application scenario, so as to keep the accuracy of removing keyword entries in the data at a relatively high level.

[0108] It should be noted that the method flow shown in the above Figure 2 embodiment is only for illustrative purposes and is not used for limitation. Based on the method embodiment shown in Figure 2 there may be other embodiments. An exemplary description is given below.

[0109] Example 1, Figure 2In the illustrated embodiment, the data processing device 200 dynamically updates the policy set by using the processing result 1 corresponding to data 1, and uses the updated policy set to remove keyword entries from other newly received data. At this time, some keyword entries still remain in the processing result 1 corresponding to data 1. Therefore, when the storage device 300 stores this processing result 1, there may be a risk of keyword entry leakage. Thus, in other possible embodiments, when Application 100 has a relatively low response latency requirement for removing keyword entries from data, after the data processing device 200 uses the policy set 1 to remove the keyword entries from data 1, it may not write the processing result 1 corresponding to data 1 into the storage device 300. Instead, it first updates the policy result 1 to the policy set 2 according to data 1 and the processing result 1, and uses the policy set 2 to continue removing the existing keyword entries in this processing result 1, thereby obtaining the processing result 3. Then, the data processing device 200 can write this processing result 3 into the storage device 300. In this way, the data stored in the storage device 300 can be made to not include keyword entries, thereby effectively ensuring the security of data storage, reducing the risk of keyword entry leakage, and improving the user experience.

[0110] Example two, Figure 2 In the illustrated embodiment, the data processing device 200 updates the policy set 1 by using a set of data (i.e., data 1) and the processing result corresponding to this set of data (i.e., processing result 1). In other possible embodiments, the data processing device 200 can measure the removal effect of the current policy set according to multiple sets of data and the processing result corresponding to each set of data, and determine whether to update the processing policy in the policy set according to this measurement effect. Alternatively, when the quantity of data 1 exceeds a preset quantity, the data processing device 200 updates the policy set 1 according to data 1 and its corresponding processing result 1. In this way, updating the policy set by using a sufficient quantity of data and processing results can improve the reliability and accuracy of updating the policy set.

[0111] Example three, Figure 2 In the illustrated embodiment, the data writing scenario is used as an example for illustration. That is, the data processing device 200 removes keyword entries from data before writing the data into the storage device 300, so as to reduce the risk of keyword entry leakage when the data is stored in the storage device 300. In other possible embodiments, the data processing device 200 can also be applicable to removing keyword entries from data in the data query scenario. That is, the data processing device 200 can also support removing keyword entries from the data queried by Application 100 and then presenting it to the user, so as to reduce the risk of data leakage during the process of Application 100 querying data.

[0112] In specific implementation, the data stored in the storage device 300 may include keyword entries. For example, the data is not processed to remove keyword entries before being written into the storage device 300. When the application 100 needs to query the data 3 in the storage device 300, the application 100 can generate a database statement 3 for querying the data 3 in the storage device 300. At this time, the data processing device 200 can intercept and execute the database statement 3, access the storage device 300 to obtain the data 3. Assuming that the data 3 contains keyword entries, such as keyword entries including ID card numbers or transaction amounts, etc., the data processing device 200 can use the policy set 1 to remove the keyword entries from the data 3 and obtain the corresponding processing result 3. Thus, the data processing device 300 can feedback the processing result 3 to the application 100. In this way, the data queried by the application 100 is the processing result 3 obtained after removing the keyword entries. In actual application, after obtaining the processing result 3, the application 100 can present an interactive interface, such as presenting the interactive interface to the user, and the application 100 can display the processing result 3 on the interactive interface. In this way, it can be avoided that the data queried and displayed by the application 100 contains keyword entries, thereby ensuring the security when the application 100 queries data and reducing the risk of keyword entry leakage. Further, the data processing device 300 can also update the policy set 1 according to the data 3 and the processing result 3 to obtain the policy set 2. Thus, for the data to be queried by the application 100 subsequently, the data processing device 200 can use the policy set 2 to remove the newly queried data, thereby improving the effect of the data processing device 200 using the policy set to remove data.

[0113] In an actual application scenario, the user can pre-configure the data processing device 200 according to the actual application requirements to determine whether the data processing device 200 removes the data before it is written into the storage device 300 or removes the queried data during the process of the application 100 querying data, and this is not limited.

[0114] It should be noted that the above Figure 2 described embodiments can be applied to the database read-write scenario, or can be applied to the log read-write scenario. The following respectively gives exemplary descriptions of these two scenarios.

[0115] The first scenario is for the data reading and writing scenario targeting a database. In a possible implementation manner, taking writing data into the database as an example, the application 100 can generate a database statement 1 for the data 1. The database statement 1 can be, for example, a structured query language (SQL) statement, etc., so as to use the database statement 1 to instruct to write the data 1 into the storage device 300. Among them, the database statement 1 can be used to instruct to perform data addition, or data modification, or data replacement, etc. on the storage device 300 using the data 1. Among them, the database statement 1 generated by the application 100 can include the data 1. For example, the database statement 1 can be, for example, "insert into table1(id,detail) values(1, \"Li Si browsed webpage A at 18:00:00\")", instructing to write "Li Si browsed webpage A at 18:00:00" (i.e., the data 1) into the column "detail" in the table table1.

[0116] At this time, the data processing device 200 can be, for example, a component, and the application 100 can pre-load and run this component. Taking the data processing device 200 as a component developed based on the Java programming language as an example, the data processing device 200 can define the database connection method based on the mechanism of Java JDK (Java Development Kit) SPI (Service Provider Interface), and can implement the Java Database Connectivity (JDBC) standard interface to connect to the storage device 300. Specifically, a Driver class can be defined in the application 100, and this Driver class can inherit NonRegisteringDriver to implement the java.sql.Driver interface. At the same time, a java.sql.Driver file for implementing this component can be created in the META-INF / services directory, and its content is the fully qualified class name of the custom Driver class (the fully qualified class name refers to the complete qualified name of a Java class, including the package name and the class name).

[0117] During operation, the data processing device 200 can obtain the data 1 generated by the application 100, and can intercept the database statement 1 generated by the application 100, so that the data processing device 200 can subsequently remove the key words in the data 1 according to the database statement 1. Then, the data processing device 200 parses the database statement 1 to obtain the data 1 that needs to be written into the storage device 300. In other implementations, the data processing device 200 can also obtain the data 1 generated by the application 100 based on other methods. For example, when the data processing device 200 is implemented by hardware, the application 100 can directly send the data 1 to the data processing device 200, etc. This embodiment does not limit the implementation method of the data processing device 200 obtaining the data 1.

[0118] Furthermore, after intercepting the database statement 1 generated by the application 100 , the data processing device 200 can rewrite the database statement 1 according to the processing result 1 to obtain the database statement 2, so that the data processing device 200 can execute the database statement 2 and write the processing result 1 into the storage device 300 . For example, assuming that database statement 1 is specifically "insert in to table 1 (detail) values ​​("Li Si used his ID number 111111111111111111 to buy a financial product worth 1,560 yuan")", then after the data processing device 200 rewrites the database statement 1 according to the processing result 1, the generated database statement 2 is specifically "insert into table 1 (detail) values ​​("Li Si used his ID number 111111********1111 to buy a financial product worth 15 yuan")", so that the data processing device 200 executes the database statement 2 to write the data "Li Si used his ID number 111111********1111 to buy a financial product worth 15 yuan" (that is, the processing result 1) after removing the key words into the table table 1 in the storage device 300.

[0119] The second scenario is to read and write logs for the storage device 300. The data processing device 200 can obtain the logs generated by the application 100 during operation, and write the obtained logs to the storage device 300 after removing the key words in the logs; or, when responding to the request of the application 100 to read the logs in the storage device 300, the data processing device 200 can remove the key words from the logs read from the storage device 300, and then feed the obtained logs back to the application 100.

[0120] In actual application, when removing key words from a log, the data processing device 200 may use different processing strategies in different strategy sets for different types of logs to remove key words from the log.Figure 3 Describe this in detail.

[0121] Refer to Figure 3 , Figure 3 which is a schematic flowchart of another data processing method provided by this application. This method can be applied to Figure 1 the data processing system 10 described above, or can be applied to other applicable data processing systems. For the sake of convenience of description, in this embodiment, taking the application to Figure 1 the data processing system 10 shown as an example for exemplary description.

[0122] Among them, Figure 3 the data processing method shown specifically may include:

[0123] S301: During the operation of the application 100, a log 1 of the first log category is generated, and the log 1 is provided to the data processing device 200.

[0124] In this embodiment, during the operation of the application 100, a log 1 of the first log category can be generated, and moreover, the log 1 needs to be saved to the storage device 300. The log 1 generated by the application 100 may contain keyword entries.

[0125] Exemplarily, the log category may include operation logs, access logs, program logs, etc. Among them, the operation log refers to a log used to record the corresponding operations performed on a certain object, such as a log used to record the operation behaviors of an administrator for configuring and naming the application 100, etc.; at this time, the event type, the name of the accessed resource, the access initiation end address or identifier, the access result, etc. recorded in the log can be keyword entries. The access log refers to a log used to record the corresponding requests triggered or executed by a user or the application 100, such as a log used to record the application 100's request to access a specific file, etc.; at this time, the information such as the request object and request conditions recorded in the log can be keyword entries. The program log refers to a log used to record the events that occur during the operation of the application 100, such as a log used to record the calculation events executed by the application 100 in response to user operations, etc.; at this time, some of the event content recorded in the log can be keyword entries. In this embodiment, the specific implementation of the log category and the keyword entries included in the log is not limited.

[0126] Exemplarily, the data processing device 200 can be a component, and moreover, the application 100 can pre-load and run this component.

[0127] At this time, during the operation of the data processing device 200, it can monitor whether the application 100 generates new logs, and moreover, when it detects that the application 100 generates the log 1, it can trigger the execution of subsequent steps to remove the keyword entries of the log 1.

[0128] In other embodiments, the data processing device 200 may also obtain the log 1 generated by the application 100 based on other methods. For example, when the data processing device 200 is implemented by hardware, the application 100 may directly send the log 1 to the data processing device 200. The present application does not limit the implementation manner for the data processing device 200 to obtain the log 1.

[0129] S302: The data processing device 200 uses the policy set 1 corresponding to the first log category to remove the keyword entries in the log 1, and obtains the processing result 1, where the policy set 1 includes at least one processing policy.

[0130] Among them, the processing policy refers to a policy for removing the keyword entries included in the log.

[0131] Implementation method 1: The processing policy may include a keyword entry, an identification rule, and an identifier of a processing algorithm. Among them, the identification rule is used to identify whether the keyword entry is included in the log, and the processing algorithm is used to remove the keyword entry in the log after it is identified that the keyword entry is included in the log. For example, the processing policy may be "username-keyword identification rule-anonymization processing algorithm identifier". Then, when the log includes a username, the keyword entry of the username included in the log can be identified through the "keyword identification rule" in the processing policy, and some or all characters in the username can be replaced with specified characters, such as replaced with "******", etc. Another example is that the processing policy may be "transaction amount-regular expression identification rule-offset rounding processing algorithm identifier". Then, when the transaction amount is included in the log, the keyword entry of the transaction amount included in the log can be identified through the "regular expression identification rule" in the processing policy, and the offset rounding processing algorithm indicated by the processing policy can adjust the transaction amount. For example, the transaction amount "1234" is offset to 12.34, and the offset transaction amount is adjusted to "12", etc.

[0132] Among them, different processing policies may include different keyword entries, or different identification rules, or different identifiers of processing algorithms, and this is not limited.

[0133] Implementation method 2: The processing policy may only include a keyword entry and an identifier of a processing algorithm, and the processing algorithm is used to remove the keyword entry identified from the log. For example, the processing policy may be "username-anonymization processing algorithm identifier", "transaction amount-offset rounding processing algorithm identifier", etc. At this time, the data processing device 200 may identify the keyword entries in the log based on the built-in identification rules.

[0134] In this embodiment, the description is given by taking the example that Log 1 contains keyword entries. In actual application, when Log 1 does not include keyword entries, the subsequent processing of removing keyword entries may not need to be performed. After obtaining Log 1, the data processing device 200 may use the recognition rules in the processing policy or the recognition rules built into the data processing device 200 to recognize the keyword entries contained in Log 1. In this embodiment, for the specific implementation manner of the recognition rules, reference may be made to the relevant descriptions in the above Figure 2 embodiment shown, which will not be elaborated here.

[0135] Among them, the data processing device 200 may obtain the policy set 1 based on the following two non-limiting implementation examples.

[0136] Example 1, the policy set 1 corresponding to the first log category may be pre-configured in the data processing device 200 by a technician. For example, the technician may configure the processing policy for the keyword entry for the logs of the first log category that may have keyword entries in the data processing device 200.

[0137] Example 2, the data processing device 200 may automatically generate the policy set 1 including the policy set 1 according to the operation records in one or more logs of the first log category during the process of processing logs.

[0138] In specific implementation, before the data processing device 200 removes the keyword entries in Log 1, the log H belonging to the first log category may be obtained in advance (the log H may refer to one or more logs). Then, the data processing device 200 may use the built-in recognition rules to recognize the keyword entries in the log H and classify the keyword entries in the log H to obtain the categories to which the keyword entries in the log H belong. Exemplarily, the categories to which the keyword entries included in the log H of the first log category belong may be personal data category, business data category, or key data category. Among them, the personal data category refers to the data category related to personal information, such as ID number, home address, etc.; the business data category refers to the data category related to business, such as transaction amount, encryption algorithm name, key, etc.; the key data category refers to a specific type of key data, such as the name of the user's cooperation manufacturer, etc.

[0139] Then, the data processing device 200 can input the keyword entries in log H into a pre-trained classification model. This classification model can be, for example, a model constructed based on a neural network, or a model constructed based on a classifier, etc. Then, the classification model can output the probabilities that the keyword entry belongs to each category, so that the category corresponding to the maximum probability can be used as the category of the keyword entry. It should be noted that when there are multiple different keyword entries in log H, the data processing device 200 can sequentially input different keyword entries into the classification model to obtain the category to which each keyword entry belongs.

[0140] Moreover, the data processing device 200 can also rank the keyword entries in log H to obtain the key levels corresponding to the keyword entries in log A. Different key levels correspond to different degrees of criticality. Then, the data processing device 200 can determine the processing algorithm corresponding to the keyword entry according to the category to which the keyword entry in log H belongs and the key level corresponding to the keyword entry. Exemplarily, the data processing device 200 can be pre-configured with the mapping relationship between the category, key level, and processing algorithm. Thus, the data processing device 200 can determine the processing algorithm applicable to the keyword entry in log H by looking up this mapping relationship. In this way, the data processing device 200 can generate at least one corresponding processing strategy according to the processing algorithm corresponding to the keyword entry in log H. Each processing strategy can include the keyword entry (such as a keyword) and the identifier of the processing algorithm, and can also include recognition rules, so as to obtain a policy set 1 including the at least one processing strategy. In this embodiment, for the specific implementation process of the data processing device 200 ranking the keyword entries in log H and determining the processing algorithm, reference can be made to the relevant descriptions in the foregoing Figure 2 described in the embodiments shown, and details will not be elaborated here.

[0141] S303: The data processing device 200 updates the policy set 1 according to log 1 and the processing result 1 to obtain a policy set 2. The policy set 2 includes at least one processing strategy, and the policy set 2 can remove the keyword entries in the processing result 1.

[0142] In this embodiment, for the specific implementation manner of step S303, reference can be made to the relevant descriptions of step S203 in the foregoing Figure 2 described in the embodiments shown, and details will not be elaborated here.

[0143] S304: Application 100 generates log 2 of the first log category during operation and provides the log 2 to the data processing device 200.

[0144] Among them, the specific implementation process of Application 100 for generating Log 2 is similar to that of Application 100 for generating Log 1, and can be referred to the relevant descriptions above, which will not be elaborated here.

[0145] S305: The data processing device 200 uses the policy set 2 to remove the keyword entries in Log 2, and obtains the processing result 2.

[0146] Furthermore, after obtaining the processing result 2, the data processing device 200 can write the processing result 2 into the storage device 300.

[0147] In practical applications, the data processing device 200 can, based on the above method, continuously update the policy set dynamically for the policy sets corresponding to different log categories by using the logs and processing results under the log category. In this way, subsequently, for each policy set corresponding to each log category, the data processing device 200 can use the updated policy set to remove the keyword entries in the logs under the log category.

[0148] It should be noted that Figure 3 The method flow shown in the embodiment is only for illustrative purposes and is not used for limitation. Based on Figure 3 the method embodiment shown, there may be other embodiments. An exemplary description is given below.

[0149] Example 1 Figure 3 In the shown embodiment, the data processing device 200 uses the processing result 1 corresponding to Log 1 to dynamically update the policy set, and uses the updated policy set to remove the keyword entries in other newly received data. At this time, some keyword entries or some key contents are still retained in the processing result 1 corresponding to Log 1. Therefore, when the storage device 300 stores the processing result 1, there may be a risk of leakage of key contents. Therefore, in other possible embodiments, when the response delay requirement of Application 100 for data removal is relatively low, after the data processing device 200 uses the policy set 1 to remove Log 1, it may not write the processing result 1 corresponding to Log 1 into the storage device 300. Instead, it first updates the policy result 1 to the policy set 2 according to Log 1 and the processing result 1, and continues to use the policy set 2 to remove the processing result 1 to remove the keyword entries included in the processing result 1, so as to obtain the processing result 3, and then the data processing device 200 can write the processing result 3 into the storage device 300. In this way, the data stored in the storage device 300 can be ensured not to include keyword entries, thereby effectively ensuring the security of data storage, reducing the risk of leakage of keyword entries, and improving the user experience.

[0150] Example 2 Figure 3In the illustrated embodiment, the data processing device 200 updates the policy set 1 by using log 1 and the corresponding processing result 1 of log 1. In other possible embodiments, the data processing device 200 may measure the removal effect of the current policy set according to multiple logs and the corresponding processing result of each log, and determine whether to update the processing policy in the policy set according to the measured effect. Alternatively, when the quantity of log 1 exceeds a preset quantity, the data processing device 200 updates the policy set 1 according to log 1 and its corresponding processing result 1. In this way, updating the policy set by using a sufficient quantity of logs and processing results can improve the reliability and accuracy of updating the policy set.

[0151] Example 3. In an actual application scenario, since logs usually include a relatively large number of data records (the amount of data recorded in the logs is large), and not all data records include keyword entries, in other embodiments, when the data processing device 200 uses the second policy set to remove keyword entries from the logs, it may perform a preliminary screening on multiple data records in the logs, and then use the second policy set to process the multiple screened data records.

[0152] In specific implementation, the data processing device 200 may create a keyword library, which includes at least one keyword, and each keyword may be a keyword entry included in a processing policy in the first policy set. In this way, when removing keyword entries from the logs, the data processing device 200 may first use at least one keyword in the keyword library to screen multiple first data records in the logs to obtain at least one second data record, and each screened second data record includes at least one keyword in the keyword library. Thus, the data processing device 200 may use the processing policy in the second policy set to remove the keyword entries in at least one second data record to obtain the corresponding processing result. In this way, the data processing device 200 filters multiple first data records in the logs by using the keyword library, reduces the number of data records traversed by the processing policy for keyword entry matching, thereby reducing the number of data records from which the data processing device 200 removes keyword entries in the logs and reducing resource consumption. For example, assuming that the keyword in the keyword library has the structure of "ID card: number", then, through keyword matching, the first data records that only include "ID card" but do not include "number" in multiple first data records can be filtered out. At this time, this first data record does not need to execute the process of removing keyword entries (does not include key content), thereby reducing the number of data records from which the data processing device 200 subsequently removes keyword entries in the logs and reducing resource consumption.

[0153] Example 4 Figure 3In the illustrated embodiment, the storage removal scenario is taken as an example for illustration. That is, the data processing device 200 removes the log before writing it into the storage device 300 to reduce the risk of keyword leakage when the log is stored in the storage device 300. In other possible embodiments, the data processing device 200 can also be applied to data removal in the log query scenario. That is, the data processing device 200 can also support removing the log queried by the application 100 to reduce the risk of data leakage during the log query process of the application 100.

[0154] In specific implementation, the log stored in the storage device 300 may contain keyword entries. For example, the log is not removed before being written into the storage device 300. When the application 100 requests to query the log 3 in the storage device 300, the data processing device 200 can intercept the request, access the storage device 300 to obtain the log 3. Assuming that the log 3 contains keyword entries, such as keyword entries including user personal data or transaction amounts, etc., the data processing device 200 can use the policy set 1 to remove the log 3 and obtain the corresponding processing result 3. Thus, the data processing device 300 can feedback the processing result 3 to the application 100. In this way, the log queried by the application 100 is the processing result 3 obtained after removal. In practical applications, after obtaining the processing result 3, the application 100 can present an interactive interface, such as presenting the interactive interface to the user, and the application 100 can display the processing result 3 on the interactive interface. In this way, it can be avoided that the log content queried and displayed by the application 100 contains keyword entries, thereby ensuring the security of the application 100 when querying the log and reducing the risk of keyword leakage. Further, the data processing device 300 can also update the policy set 1 according to the log 3 and the processing result 3 to obtain the policy set 2. Thus, for the log that the application 100 will query subsequently, the data processing device 200 can use the policy set 2 to remove the newly queried log, thereby improving the effect of the data processing device 200 using the policy set to remove the log.

[0155] Taking the application 100's request to print the info-level logs in Log4j2 (a Java logging framework) as an example, the log levels in Log4j2 are divided into 8 levels, and the priorities from high to low are off, fatal, error, warn, info, debug, trace, and all. Then, after the data processing device 200 accesses the logs from the storage device 300, it can filter the logs using the global filter of the "all" level and the logger filter of the "trace" level in sequence. Among them, the global filter of the "all" level allows the output of all levels of logs, and the logger filter of the "trace" level allows the output of logs above the "trace" level. Then, the data processing device 200 can output the log content in the form of placeholders and can generate custom log content by modifying the MessageFactory component. This custom log content is also the processing result obtained by removing the output log content according to the policy set 1. Next, the data processing device 200 can generate a log event based on the removed log content and filter the log event using the logger filter of the "debug" level and the Appender filter of the "info" level, so as to output the removed log record when it is determined that the log does not need to be rolled back. Or, the data processing device 200 can not modify the MessageFactory component, but can use the PatternLayout component in the log configuration file to remove the keywords in the log record using the policy set 1 and output the removed log record. In this way, the data processing device 200 can achieve the output after removing the keyword entries in the log.

[0156] In the actual application scenario, the user can pre-configure the data processing device 200 according to the actual application requirements to determine whether the data processing device 200 removes the keyword entries in the log before the log is written to the storage device 300 or removes the keyword entries in the currently queried log during the process of the application 100 querying the log, and no limitation is imposed on this.

[0157] In this embodiment, for the logs of different log categories, the data processing device 200 can adopt the processing strategies in different policy sets to remove the keyword entries in the log, which helps to improve the effect of removing keyword entries for specific categories of logs. In this way, it is possible to achieve differential removal of keyword entries for different categories of logs and improve the flexibility of removing keyword entries in the log.

[0158] It should be noted that other reasonable combinations of steps that can be conceived by those skilled in the art based on the above description also fall within the protection scope of this application. Secondly, those skilled in the art should also be familiar that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily essential for this application.

[0159] The above combination Figures 1 to 3 introduces the data processing method provided by this application. Next, the structures of the data processing device and the computing device provided by this application will be introduced with reference to the accompanying drawings.

[0160] See Figure 4 , which shows a schematic structural diagram of a data processing device 400. The data processing device 400 includes:

[0161] An acquisition module 401, configured to acquire a second policy set, where the second policy set is obtained by updating a first policy set based on first data and a first processing result, the first processing result is obtained by removing keyword entries in the first data based on the first policy set, and the second policy set includes at least one processing policy; acquire second data;

[0162] A removal module 402, configured to use the second policy set to remove keyword entries in the second data to obtain a second processing result.

[0163] In a possible implementation manner, the data processing device 400 may further include an update module 403. The update module 403 is configured to:

[0164] Measure the removal degree corresponding to the first processing result according to the first data and the first processing result, where the removal degree is used to measure the degree of removing keyword entries in the first processing result;

[0165] When the removal degree is lower than a first threshold, use a classifier to identify first keyword entries in the first processing result;

[0166] A generation module generates a target processing policy for the first keyword entries;

[0167] Update the first policy set using the target processing policy.

[0168] In a possible implementation manner, the update module 403 is specifically configured to:

[0169] Determine an availability index, a relevance index, a authenticity index, and a reproducibility index according to the first data and the first processing result;

[0170] Use the differential privacy algorithm to determine a privacy budget according to the first data and the first processing result

[0171] Calculate the removal degree corresponding to the first processing result according to the availability index, the relevance index, the authenticity index, the reproducibility index, and the privacy budget.

[0172] In a possible implementation manner, the data processing device 400 may further include a generation module 404, and the generation module 404 is configured to:

[0173] Identify the second keyword entry in the third data;

[0174] Classify the second keyword entry in the second data by using a classification model to obtain the category to which the second keyword entry belongs;

[0175] Grade the second keyword entry in the second data by using a grading model to obtain the key level corresponding to the second keyword entry;

[0176] Determine the identifier of the processing algorithm corresponding to the second keyword entry according to the category to which the second keyword entry belongs and the key level corresponding to the second keyword entry;

[0177] Generate a processing policy in the first policy set, where the processing policy includes the second keyword entry, the recognition rule, and the identifier of the processing algorithm.

[0178] In a possible implementation manner, the generation module 404 is configured to:

[0179] Match the words in the third data with the keywords in the thesaurus to obtain the second keyword entry in the third data, where the similarity between the word vectors of the second keyword entry and the keywords is higher than a second threshold;

[0180] Alternatively, identify the named entities in the third data to obtain the second keyword entry in the third data, where the second keyword entry is the named entity in the third data that satisfies the verification rule.

[0181] In a possible implementation manner, the obtaining module 401 is specifically configured to obtain a first database statement, where the first database statement includes the second data, and the first database statement is used to instruct to write the second data into the database;

[0182] The data processing device 400 may further include an execution module 405, which is configured to:

[0183] After obtaining the second processing result, rewrite the first database statement by using the second processing result to obtain a second database statement, where the second database statement is used to instruct to write the second processing result into the database;

[0184] Execute the second database statement.

[0185] In a possible implementation, the obtaining module 401 is specifically configured to:

[0186] Obtain a third database statement, where the third database statement is used to instruct to query the second data in the database;

[0187] Access the database according to the third database statement to obtain the second data;

[0188] The data processing device 400 may further include a presenting module 406 for presenting an interaction interface, where the interaction interface includes the second processing result.

[0189] In a possible implementation, the removing module 402 is further configured to remove keyword entries in the first processing result by using the second policy set to obtain a third processing result.

[0190] In a possible implementation, the first data is a first log, the second data is a second log, both the first log and the second log belong to the first log category, and both the first policy set and the second policy set correspond to the first log category.

[0191] In a possible implementation, the obtaining module 401 is further configured to obtain a fourth policy set corresponding to the second log category, where the fourth policy set is obtained by updating a third policy set based on a third log and a fourth processing result, the fourth processing result is obtained by removing keyword entries in the third log by using the third policy set corresponding to the second log category, the third policy set includes at least one processing policy, and the third log belongs to the second log category; and obtain a fourth log, where the fourth log belongs to the second log category;

[0192] The removing module 402 is further configured to remove keyword entries in the fourth log by using the fourth policy set to obtain a fifth processing result.

[0193] In a possible implementation, the second log includes multiple first data records, and the removing module 402 is specifically configured to:

[0194] Obtain a keyword library, where the keyword library includes at least one keyword;

[0195] Use the keyword library to screen the multiple first data records to obtain at least one second data record, where the at least one second data record includes at least one keyword in the keyword library;

[0196] Use the second policy set to remove keyword entries from the at least one second data record to obtain the second processing result.

[0197] Since Figure 4 the data processing device 400 shown corresponds to the data processing device 200 in the above Figure 2 or Figure 3 the data processing device 200 in the illustrated embodiment, so Figure 4 For the specific implementation manner of the data processing device 400 shown and the technical effects it has, refer to the relevant descriptions in the above Figure 2 or Figure 3 illustrated embodiment, which will not be elaborated here.

[0198] Figure 5 FIG. 500 is a schematic hardware structure diagram of a computing device 500 provided by the present application. The computing device 500 can implement, for example, the data processing device 200 in the above Figure 2 or Figure 3 illustrated embodiment.

[0199] As Figure 5 shown, the computing device 500 includes a processor 501, a memory 502, and a communication interface 503. Among them, the processor 501, the memory 502, and the communication interface 503 communicate through a bus 504, and can also achieve communication through other means such as wireless transmission. The memory 502 is used to store instructions, and the processor 501 is used to execute the instructions stored in the memory 502. Further, the computing device 500 may further include a memory unit 505, and the memory unit 505 can be connected to the processor 501, the storage medium 502, and the communication interface 503 through the bus 504. Among them, the memory 502 stores program codes, and the processor 501 can call the program codes stored in the memory 502 to perform the following operations:

[0200] Obtain a second policy set, where the second policy set is obtained by updating a first policy set based on first data and a first processing result, the first processing result is obtained by removing keyword entries from the first data based on the first policy set, and the second policy set includes at least one processing policy;

[0201] Obtain second data;

[0202] Use the second policy set to remove keyword entries from the second data to obtain a second processing result.

[0203] It should be understood that in this embodiment, the processor 501 may be a CPU, and the processor 501 may also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete device components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc.

[0204] The memory 502 may include a read-only memory and a random access memory, and provide instructions and data to the processor 501. The memory 502 may also include a non-volatile random access memory.

[0205] The memory 502 may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate synchronous DRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchlink DRAM (SLDRAM), and direct rambus RAM (DR RAM).

[0206] The communication interface 503 is used to communicate with other devices connected to the computing device 500. In addition to the data bus, the bus 504 may also include a power bus, a control bus, a status signal bus, etc. However, for the sake of clarity, all kinds of buses are labeled as the bus 504 in the figure.

[0207] It should be understood that the computing device 500 according to the present application may correspond to the processor data processing device 400 in the present application, and may correspond to the execution according to the present application Figure 2 OrFigure 3 The method executed by the data processing device 200 in the method shown, and the above and other operations and / or functions implemented by the computing device 500 are respectively for implementing Figure 2 Or Figure 3 the flow of the corresponding method in

[0208] This application also protects a computing cluster, which includes a plurality of computing devices 500 as shown in Figure 5 The plurality of computing devices are jointly used to execute the method executed by the data processing device 200 in the method shown in Figure 2 Or Figure 3 Among them, the above and other operations and / or functions implemented by the plurality of computing devices 500 are respectively for implementing Figure 2 Or Figure 3 The flow of the corresponding method in can adopt a distributed processing or centralized processing method. For the sake of brevity, it will not be elaborated here.

[0209] This application also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that a computing device can store or a data storage device such as a data center containing one or more available media. The available medium can be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a solid-state drive), etc. The computer-readable storage medium includes instructions that instruct the computing device to execute the above data processing method.

[0210] This application also provides a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on the computing device, the processes or functions according to this application are fully or partially generated.

[0211] The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, or data center to another website, computer, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.).

[0212] The computer program product can be a software installation package. In any case where the above data processing method needs to be used, the computer program product can be downloaded and executed on the computing device.

[0213] The above embodiments can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded or executed on a computer, the processes or functions described in this application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (such as coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that contains one or more collections of available media. The available media can be magnetic media (such as floppy disks, hard disks, magnetic tapes), optical media (such as DVDs), or semiconductor media. The semiconductor media can be a solid-state drive.

[0214] The terms used in the above embodiments are only for the purpose of describing specific embodiments and are not intended to limit this application. As used in the specification and claims of this application, the singular forms "a", "an", "the", "above", "said", "this", and "such" are also intended to include expressions such as "one or more", unless the context clearly indicates otherwise. It should also be understood that in this application, "one or more" means one, two, or more than two; the character " / " generally indicates an "or" relationship between the associated objects before and after. In this application, "simultaneously" means within the same time period, including the case of being at the same moment. The terms "first", "second", etc. in the specification, claims, and drawings of this application are used to distinguish similar objects and do not necessarily have to be used to describe a specific order or sequence. It should be understood that such terms can be interchanged under appropriate circumstances, which is only a way of distinguishing objects with the same attributes when describing the embodiments of this application.

[0215] References to "one embodiment" or "some embodiments" etc. described in this specification mean that a particular feature, structure, or characteristic described in connection with the embodiment is included in one or more embodiments of the present application. Thus, statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc. that appear in different places in this specification are not necessarily all referring to the same embodiment, but rather mean "one or more but not all embodiments", unless otherwise specifically emphasized. The terms "comprising", "including", "having" and their variants all mean "including but not limited to", unless otherwise specifically emphasized.

[0216] As described above, the above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.

Claims

1. A data processing method, characterized in that, The method includes: Obtaining a second policy set, which is obtained by updating a first policy set based on first data and a first processing result, where the first processing result is obtained by removing keyword entries from the first data based on the first policy set, and the second policy set includes at least one processing policy; Obtaining second data; Using the second policy set to remove keyword entries from the second data, obtaining a second processing result.

2. The method according to claim 1, characterized in that, The method further includes: According to the first data and the first processing result, measuring the removal degree corresponding to the first processing result, where the removal degree is used to measure the degree of removing keyword entries from the first processing result; When the removal degree is lower than a first threshold, using a classifier to identify first keyword entries in the first processing result; Generating a target processing policy for the first keyword entries; Using the target processing policy to update the first policy set.

3. The method according to claim 2, wherein The measuring the removal degree corresponding to the first processing result according to the first data and the first processing result includes: According to the first data and the first processing result, determining an availability index, a relevance index, a authenticity index, and a reproducibility index; Using a differential privacy algorithm to determine a privacy budget according to the first data and the first processing result; Calculating the removal degree corresponding to the first processing result according to the availability index, the relevance index, the authenticity index, the reproducibility index, and the privacy budget.

4. The method according to any one of claims 1 to 3, characterized in that, The method further includes: Identifying second keyword entries in the third data; Using a classification model to classify the second keyword entries in the second data to obtain the category to which the second keyword entries belong; Using a grading model to grade the second keyword entries in the second data to obtain the key level corresponding to the second keyword entries; According to the category to which the second keyword entries belong and the key level corresponding to the second keyword entries, determining the identifier of the processing algorithm corresponding to the second keyword entries; Generating a processing policy in the first policy set, where the processing policy includes the second keyword entries, the recognition rule, and the identifier of the processing algorithm.

5. The method according to claim 4, wherein The identifying the second keyword entries in the third data includes: Matching the words in the third data with the keyword entries in a keyword library to obtain the second keyword entries in the third data, where the similarity between the word vectors of the second keyword entries and the word vectors of the keyword entries is higher than a second threshold; Alternatively, identifying named entities in the third data to obtain the second keyword entries in the third data, where the second keyword entries are named entities in the third data that meet the verification rules.

6. The method according to any one of claims 1 to 5, characterized in that, The obtaining the second data includes: Obtaining a first database statement, where the first database statement includes the second data and is used to indicate writing the second data into a database; The method further includes: After obtaining the second processing result, rewrite the first database statement by using the second processing result to obtain a second database statement, where the second database statement is used to indicate writing the second processing result into the database; Execute the second database statement.

7. The method according to any one of claims 1 to 5, characterized in that The obtaining the second data includes: Obtain a third database statement, where the third database statement is used to indicate querying the second data in the database; Access the database according to the third database statement to obtain the second data; The method further includes: Present an interactive interface, where the interactive interface includes the second processing result.

8. The method according to any one of claims 1 to 7, characterized in that, The method further includes: Use the second policy set to remove keyword entries from the first processing result to obtain a third processing result.

9. The method according to any one of claims 1 to 5, characterized in that The first data is a first log, the second data is a second log, both the first log and the second log belong to a first log category, and both the first policy set and the second policy set correspond to the first log category.

10. The method according to claim 9, wherein The method further includes: Obtain a fourth policy set corresponding to the second log category, where the fourth policy set is obtained by updating a third policy set based on a third log and a fourth processing result, the fourth processing result is obtained by removing keyword entries from the third log by using a third policy set corresponding to the second log category, the third policy set includes at least one processing policy, and the third log belongs to the second log category; Obtain a fourth log, where the fourth log belongs to the second log category; Use the fourth policy set to remove keyword entries from the fourth log to obtain a fifth processing result.

11. The method according to claim 9 or 10, characterized in that, The second log includes multiple first data records. The using the second policy set to remove keyword entries from the second log to obtain a second processing result includes: Obtain a keyword library, where the keyword library includes at least one keyword; Use the keyword library to screen the multiple first data records to obtain at least one second data record, where the at least one second data record includes at least one keyword in the keyword library; Use the second policy set to remove keyword entries from the at least one second data record to obtain the second processing result.

12. A data processing device, characterized in that, The device includes: An obtaining module, configured to obtain a second policy set, where the second policy set is obtained by updating a first policy set based on first data and a first processing result, the first processing result is obtained by removing keyword entries from the first data by using the first policy set, and the second policy set includes at least one processing policy; obtain second data; A removing module, configured to use the second policy set to remove keyword entries from the second data to obtain a second processing result.

13. The device according to claim 12, characterized in that, The first data is a first log, the second data is a second log, both the first log and the second log belong to a first log category, and both the first policy set and the second policy set correspond to the first log category.

14. A computing device, characterized in that, Includes a processor and a memory; The processor is configured to execute the instructions stored in the memory to cause the computing device to perform the steps of the method according to any one of claims 1 to 11.

15. A computer-readable storage medium, characterized in that, Comprising instructions which, when run on a computing device, cause the computing device to perform the steps of the method according to any one of claims 1 to 11.