Data mistaken deletion prevention method and device, electronic equipment and storage medium

By receiving data write requests in big data processing and determining whether it is core data, combined with the path protection and permission management of the Ranger interface, the challenge of preventing data from being accidentally deleted under different big data engines is solved, and efficient anti-error deletion and secure backup of core data is achieved.

CN120068140APending Publication Date: 2025-05-30DUXIAOMAN TECH (BEIJING) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411985347.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

In big data processing, it is difficult for the prior art to design a general and efficient method of anti-error deletion of data, especially when the operation types, parameter settings and execution modes of different big data engines are different, resulting in the design of unified anti-error deletion strategy facing great challenges.

Method used

By receiving the data write request, it is determined whether the written data is preset core data, and the data is disabled after the write is completed, and the data is synchronized to the second storage space. This method uses the Ranger interface for path protection and permission management to ensure the security of core data.

Benefits of technology

It realizes effective prevention of error deletion of core data, ensuring that the core data will not be deleted at will to the greatest extent, and improves data security through backup storage. It is suitable for scenarios of different big data engines without targeted configuration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120068140A_ABST
    Figure CN120068140A_ABST
Patent Text Reader

Abstract

The invention provides a data mistaken deletion prevention method and device, electronic equipment and a storage medium, and the method comprises the steps: judging whether data needing to be written is core data or not under the condition that a write-in request is received, if the data is core data, writing the data into a first storage space, after the data is written, forbidding the change permission of the data, and if the data is not core data, forbidding the change permission of the data; according to the method, the data is written into the second storage space, the data is synchronized to the second storage space, it is guaranteed that the core data cannot be deleted at will to the maximum extent by forbidding permission change of the data after data writing is completed, the data is synchronized to the second storage space and serves as core data backup storage, and the core data storage safety is further improved; the core data storage security can be protected in the storage dimension without targeted configuration for different big data, and the implementation is relatively convenient.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of big data technology, and in particular, to a method, device, electronic device, and storage medium for preventing accidental data deletion. Background Art

[0002] In today's big data era, enterprises and institutions increasingly rely on massive amounts of data for business decision-making, market analysis, and product innovation. With the explosive growth of data volume, the complexity of data management has also increased accordingly, especially the protection of core data faces unprecedented challenges. Core data generally refers to those information assets that are crucial for enterprise operations, have high commercial value, or are required by law and compliance, such as customer information, financial records, intellectual property rights, etc. Once this type of data is lost or leaked due to misoperation, malicious attack, or system failure, it may not only cause serious economic losses but also damage the enterprise's reputation and even violate laws and regulations.

[0003] Therefore, it is particularly urgent to develop a method for preventing accidental deletion of core data that can adapt to the characteristics of the big data field. In related technologies, usually, refined control strategies are implemented for specific operation types of different data processing engines. This method focuses on deeply understanding and utilizing the characteristics and operation syntax of each data processing engine, and designing specific preventive measures for high-risk operations such as INSERT OVERWRITE, DELETE, UPDATE, etc.

[0004] There are a wide variety of big data processing engines, and each engine has its unique data model, operation syntax, and execution engine. For example, Hive is mainly used for batch processing and supports SQL-like query languages; while Spark provides richer data processing capabilities, including batch processing, stream processing, and interactive query, etc. When these engines were designed initially, more attention was paid to performance and function expansion, and relatively less consideration was given to data security and operation risk control. Therefore, the operation types, parameter settings, and execution modes supported by different engines are different, which brings great challenges to the design of a unified anti-accidental deletion strategy. For example, although INSERT OVERWRITE in Hive and write.mode("overwrite") in Spark DataFrame both involve data overwriting, there are differences in operation details, execution environment, and scope of influence, making it difficult to comprehensively cover with a set of general solutions. Summary of the Invention

[0005] In view of this, embodiments of the present invention provide a method, device, electronic device, and storage medium for preventing accidental data deletion to improve the applicable range of the method for preventing accidental deletion of core data, and further improve the efficiency of preventing accidental deletion.

[0006] According to one aspect of the present invention, a method for preventing accidental data deletion is provided, and the method includes:

[0007] Receive a data writing request, where the data writing request includes an identifier of the data to be written;

[0008] When the data to be written is preset core data, write the data to be written into the first storage space;

[0009] Determine whether the data writing is completed. When the status of the data to be written is written completed, prohibit the change permission of the data to be written; and synchronize the data to be written to the second storage space.

[0010] In a possible embodiment, the data writing request further includes: a target storage location of the data to be written in the first storage space; the method further includes:

[0011] Obtain metadata information of a target storage table corresponding to the target storage location, where the metadata information includes core data attributes of the storage table;

[0012] When the core data attribute of the target storage table is true, determine that the data to be written is preset core data.

[0013] In a possible embodiment, the determining whether the data writing is completed includes:

[0014] Monitor whether a file with a preset completion identifier is generated in a target storage partition corresponding to the target storage location, where the file with the preset completion identifier is a file generated by a big data engine after the production data is completed;

[0015] When the file with the preset completion identifier is generated in the target storage partition, determine that the data writing is completed.

[0016] In a possible embodiment, the method further includes:

[0017] Receive a path protection request, where the path protection request includes a path identifier that needs to be protected by user-defined permissions;

[0018] Call the Ranger interface to prohibit the change permission of the path corresponding to the path identifier.

[0019] In a possible embodiment, the method further includes:

[0020] Obtain a data change request, where the data change request includes a change operation command, changed data, and user information;

[0021] Send the data change request to a preset client for approval. When receiving the message that the approval of the data change request is passed, call the Ranger interface to add write permission to the change directory to which the changed data belongs, and prohibit the user's permission to delete the changed data;

[0022] Monitor the change status of the changed data. When the change of the changed data is completed, cancel the write permission of the directory to which the changed data belongs, and synchronize the changed data to the second storage space.

[0023] In a possible embodiment, the method further includes:

[0024] Receive a table creation request, where the table creation request includes the core data attribute information of the table to be created and the target location path;

[0025] When the core data attribute information is true, determine whether there is an existing created table with a location path the same as the target location path;

[0026] If there is an existing created table with a location path the same as the target location path, reject the creation of the table to be created and return a table creation failure message.

[0027] According to another aspect of the present invention, there is provided a data anti-deletion device, the device includes:

[0028] A receiving module, configured to receive a data writing request, where the data writing request includes an identifier of the data to be written;

[0029] A writing module, configured to write the data to be written into the first storage space when the data to be written is preset core data;

[0030] A permission limiting module, configured to determine whether the data to be written is written completely. When the status of the data to be written is written completely, prohibit the change permission of the data to be written; and synchronize the data to be written to the second storage space.

[0031] In a possible embodiment, the data writing request further includes: the target storage location of the data to be written in the first storage space; the device further includes:

[0032] A judgment module, configured to obtain the metadata information of the target storage table corresponding to the target storage location, where the metadata information includes the core data attribute of the storage table; when the core data attribute of the target storage table is true, determine that the data to be written is preset core data;

[0033] The determination of whether the data to be written is written completely includes:

[0034] Monitor whether a file with a preset completion identifier is generated in the target storage partition corresponding to the target storage location, where the file with the preset completion identifier is a file generated by the big data engine after the production data is completed;

[0035] In the case where the file with the preset completion identifier is generated in the target storage partition, determine that the writing of the data is completed.

[0036] The receiving module is further configured to receive a path protection request, where the path protection request includes a path identifier that needs to be protected by permissions defined by the user;

[0037] The permission limiting module is configured to call the Ranger interface to prohibit the change permission of the path corresponding to the path identifier;

[0038] The receiving module is further configured to obtain a data change request, where the data change request includes a change operation command, changed data, and user information;

[0039] The device further includes:

[0040] An approval module, configured to send the data change request to a preset client for approval, and in the case of receiving a message indicating that the data change request is approved, call the Ranger interface to add write permission to the change directory to which the changed data belongs, and prohibit the user's permission to delete the changed data;

[0041] A monitoring module, configured to monitor the change status of the changed data, and in the case where the changed data is changed, cancel the write permission of the directory to which the changed data belongs, and synchronize the changed data to the second storage space.

[0042] The receiving module is further configured to receive a table creation request, where the table creation request includes core data attribute information of the table to be created and a target location path;

[0043] A table creation module, configured to determine whether there is a created table with a location path the same as the target location path in the case where the core data attribute information is true; in the case where there is a created table with a storage location path the same as the target location path, reject the creation of the table to be created and return a table creation failure message.

[0044] According to another aspect of the present invention, there is provided an electronic device, including:

[0045] A processor; and

[0046] A memory storing a program

[0047] Among them, the program includes instructions that, when executed by the processor, cause the processor to execute any one of the above data anti-deletion methods.

[0048] According to another aspect of the present invention, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to cause a computer to execute any one of the above data anti-deletion methods.

[0049] In one or more technical solutions provided in the embodiments of the present invention, when a write request is received, it is determined whether the data to be written is core data. If it is core data, the data is written to the first storage space. After the data is written, the change permission of the data is prohibited, and the data is synchronized to the second storage space. By prohibiting the change permission of the data after the data is written, it is ensured to the greatest extent that the core data will not be deleted casually, and the data is synchronized to the second storage space as a backup storage of the core data, further improving the storage security of the core data. In addition, it is more convenient to protect the storage security of the core data in the storage dimension without the need for targeted configuration for different big data. Description of the Drawings

[0050] In the following description of exemplary embodiments with reference to the accompanying drawings, more details, features, and advantages of the present invention are disclosed. In the drawings:

[0051] Figure 1 It is a schematic flowchart of a data anti-deletion method provided by an embodiment of the present invention;

[0052] Figure 2 It is a schematic flowchart of data synchronization in the data anti-deletion method provided by an embodiment of the present invention;

[0053] Figure 3 It is another schematic flowchart of the data anti-deletion method provided by an embodiment of the present invention;

[0054] Figure 4 It is a schematic flowchart of changing core data in the data anti-deletion method provided by an embodiment of the present invention;

[0055] Figure 5 It is a schematic flowchart of creating a core table in the data anti-deletion method provided by an embodiment of the present invention;

[0056] Figure 6 It is a schematic architecture diagram of the data anti-deletion method provided by an embodiment of the present invention

[0057] Figure 7 It is a schematic structural diagram of a data anti-deletion device provided by an embodiment of the present invention;

[0058] Figure 8 The structural block diagram of an exemplary electronic device that can be used to implement the embodiments of the present invention is shown. Detailed implementation manners

[0059] Embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Instead, these embodiments are provided to more thoroughly and completely understand the present invention. It should be understood that the drawings and embodiments of the present invention are only for exemplary purposes and are not used to limit the protection scope of the present invention.

[0060] It should be understood that the various steps recited in the method embodiments of the present invention can be executed in a different order and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present invention is not limited in this regard.

[0061] As used herein, the term "including" and its variants are open-ended, that is, "including but not limited to". The term "based on" is "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description. It should be noted that the concepts such as "first" and "second" mentioned in the present invention are only used to distinguish different devices, modules or units, and are not used to limit the order of the functions performed by these devices, modules or units or the interdependent relationship.

[0062] It should be noted that the modifications of "one" and "plural" mentioned in the present invention are illustrative rather than restrictive. Those skilled in the art should understand that unless otherwise clearly specified in the context, it should be understood as "one or more".

[0063] The names of the messages or information exchanged between multiple devices in the embodiments of the present invention are only for illustrative purposes and are not used to limit the scope of these messages or information.

[0064] To prevent accidental deletion of core data, the industry has provided various methods. In addition to the methods described in the background art, the following methods are also included:

[0065] 1. Deploy an intelligent monitoring and warning system, that is, deploy advanced monitoring tools to track data access and modification behaviors in real time, use machine learning algorithms to analyze operation patterns, and issue warnings in advance for abnormal or potentially dangerous deletion activities. This includes intelligent analysis of user behaviors, identifying behaviors that deviate from the regular operation patterns, and intervening in a timely manner to prevent data loss caused by misoperations.

[0066] 2. Introduce a logical deletion and auditing mechanism, that is, add a flag field in the data table to mark whether the data is logically deleted rather than physically deleted. Combine the audit logs of data operations to record the detailed information of each deletion request, including the operator, time, operation content, etc. For high-risk operations, set up a mandatory secondary confirmation or approval process to ensure that each deletion is carefully considered.

[0067] However, the implementation of anti-deletion-by-mistake strategies has great limitations. The implementation of existing anti-deletion-by-mistake strategies usually focuses on known and direct harmful operations. For emerging, complex, or highly customized data processing scenarios, their adaptability and effectiveness are limited and cannot provide sufficient protection. This requires that anti-deletion-by-mistake strategies not only continuously evolve to cope with technological development, but also have a high degree of flexibility and scalability to cover more diverse operation risks.

[0068] Moreover, the impact on engine performance and security cannot be balanced. When current anti-deletion-by-mistake and data protection measures strengthen the security control level, they often inevitably drag down the running efficiency of the data processing engine, slow down the data processing speed, and increase the processing delay, which is particularly disadvantageous for application scenarios that pursue real-time or near-real-time data analysis. On the contrary, if engine performance is given priority, the weakening of security protection becomes a non-negligible hidden danger. The balance point between the two has not been effectively established so far, forming a major pain point in technology implementation.

[0069] Based on this, the embodiments of the present invention provide a data anti-deletion-by-mistake method, device, electronic device, and storage medium. The data anti-deletion-by-mistake method provided by the embodiments of the present invention can be applied to any electronic device with a data anti-deletion-by-mistake function. This electronic device can be a computer, a server, a mobile terminal, etc. In a possible embodiment, the data anti-deletion-by-mistake method provided by the embodiments of the present invention can be applied to a distributed system. This distributed system can include a variety of big data engines, such as presto, spark, hive, etc. The following describes the solution of the present invention with reference to the drawings:

[0070] Figure 1 A flowchart of the data anti-deletion-by-mistake method provided by the embodiments of the present invention may include the following steps:

[0071] S101. Receive a data write request, where the data write request contains an identifier of the written data;

[0072] S102. When the data to be written is preset core data, write the data to be written into the first storage space;

[0073] S103. Determine whether the data to be written is written completely. When the status of the data to be written is written completely, prohibit the change permission of the data to be written; and synchronize the data to be written to the second storage space.

[0074] Applying the embodiment of the present invention, when a write request is received, it is judged whether the data to be written is core data. If it is core data, the data is written into the first storage space. After the data is written completely, the change permission of the data is prohibited, and the data is synchronized to the second storage space. By prohibiting the change permission of the data after the data is written completely, it is ensured to the greatest extent that the core data will not be deleted randomly, and the data is synchronized to the second storage space as a backup storage of the core data, further improving the storage security of the core data. In addition, it is not necessary to perform targeted configuration for different big data to protect the storage security of the core data in the storage dimension, and the implementation is relatively convenient.

[0075] The following is an exemplary description of the above S101 - S103:

[0076] In S101, the above data write request may be sent by a big data engine. The big data engine may be any big data engine in a distributed system, and each big data engine may store the calculation results obtained by online calculation or offline calculation in a database. The data write request may include a write data identifier to be written, a big data engine identifier, etc. Among them, the write data identifier may be composed of the business line to which the write data belongs, the type of the write data, the production time, etc. The type may be determined according to specific application scenarios. For example, it may be user data, system operation and maintenance data, etc. The big data engine identifier may be the name of the big data engine.

[0077] After that, it may be determined whether the data to be written is core data based on the write data identifier. In a possible embodiment, a core data type may be preset, and it is compared whether the type of the data to be written is consistent with the core data type. If they are consistent, it may be determined that the data to be written is core data.

[0078] In a possible embodiment, the above write data identifier may further include the target storage path where the write data needs to be written. In the field of data storage, data is usually stored in a sharded and partitioned manner. Among them, sharded storage means creating multiple tables in a database, each corresponding to a different storage space for storing different data. Partitioned storage is to create different storage areas in a data table or database, and different storage areas are used to store different data. Here, different data can be data belonging to different business lines, data with production times in different time periods, and so on. The above target storage path is the table information, partition information, etc. that the write data needs to be written to. Among them, the table information and partition information may include the storage location. In a possible embodiment, the above target storage path may be the path of the first storage space in HDFS (Hadoop Distributed File System), and the first storage space can be flexibly set according to the actual application scenario. The target storage path may include target data table information and target partition information.

[0079] In the embodiments of the present invention, a core data storage space may be preset. As a possible implementation manner, a table for storing core data may be preset. Specifically, after creating the table, the field of "whether it is core data" in the metadata information of the table may be set to true. Exemplarily, a table for storing core data may be created through the CREATE TABLE example_table(id INT,data STRING)PARTITIONED BY(date STRING)LOCATION ' / user / hive / example_table'TBLPROPERTIES('is_core_data'='TRUE') table creation statement.

[0080] In a possible embodiment, it is also possible to determine whether the write data is core data based on the above target storage path, that is, the above data anti-deletion method may further include: obtaining the metadata information of the target storage table corresponding to the target storage location, and the metadata information includes the core data attribute of the storage table;

[0081] When the core data attribute of the target storage table is true, it is determined that the write data is preset core data.

[0082] Exemplarily, the HiveMetastore API can be called to query the metadata information of the tables included in the target storage path, and determine whether the is_core_data attribute in the table metadata TBLPROPERTIES is true. If it is true, it can be determined that the written data is core data. If it is not true, it can be determined that the written data is not core data. Therefore, the subsequent steps can be skipped for this written data.

[0083] In the case where the written data is core data, the writing status of the written data can be monitored. The writing status can include incomplete writing and complete writing. When the writing status of the written data is incomplete writing, the writing status of the written data can be continuously monitored. When the writing status of the written data is complete writing, the subsequent write permission for the written data can be prohibited, that is, it is not allowed to modify the written data. At the same time, the written data is synchronized to the second storage space. The second storage space can be determined according to actual needs, and the embodiments of the present invention do not make specific limitations thereto.

[0084] In a possible embodiment, the big data engine can be modified so that the big data engine writes a file with a preset completion identifier after the data writing is completed. When the file with the preset completion identifier is detected, it can be determined that the written data is written completed. Exemplarily, it can be set that after the data writing is completed, the big data engine generates a.complete identifier file in the written partition directory. When the.complete identifier file is generated in the partition directory corresponding to the target storage path, it can be determined that the writing status of the written data is completed.

[0085] In a possible embodiment, the writing status of the written data can be determined through the following steps:

[0086] Monitor whether a file with a preset completion identifier is generated in the target storage partition corresponding to the target storage location, where the file with the preset completion identifier is a file generated by the big data engine after the production data is completed;

[0087] In the case where the file with the preset completion identifier is generated in the target storage partition, determine that the written data is written completed.

[0088] In a possible embodiment, the writing status of the written data can be determined by a data writing status judge. The data writing status judge can specifically be a pre-written code module. Based on the above example, during the data processing process, the data writing status judge can continuously monitor the generation status of the.complete identifier file in the target partition directory included in the target storage path. When the file is produced, the data writing status judge can output a prompt message indicating that the writing is completed, such as "Data writing has been completed".

[0089] By determining whether the written data of the job is in the final state and prohibiting the write permission of the written data after the written data is written, it is possible to prevent the job from failing due to the execution of the write prohibition before the data is updated completely, thereby improving the stability of data storage.

[0090] In a possible embodiment, the Ranger component can be used to control data permissions for fine-grained read and write management of data. Among them, the Ranger component is a centralized security management framework in the field of big data, which realizes the centralized security management of Hadoop (a distributed system infrastructure) ecosystem components. Users can achieve secure access to data in the cluster through Ranger. It mainly supervises Hadoop platform components, starts services, and controls resource access. When it is determined that the data writing of the job is completed, that is, the corresponding partition directory is completely generated, the Ranger API interface can be called to create a policy of "disable-writes-on-directory" for a target storage path, that is, only "Read" and "Execute" permissions are assigned to the data in the target storage path, and "Write" permission is not assigned.

[0091] After the configuration is completed, it can take effect on the entire HDFS ACL (Access Control List). Among them, the HDFS ACL is a file containing the access permission information of each data in HDFS.

[0092] In a possible embodiment, when prohibiting the write permission of the written data and it is necessary to synchronize the written data to the second storage space, it can be executed according to the flowchart as Figure 2 shown. Specifically, after prohibiting the write permission of the target storage path (that is, the source path in Figure 2 ), the data in the target storage path can be stored in the target path of the second storage space by using a synchronization tool such as Hadoop distcp. The target path can be the same as or different from the target storage path. After the data synchronization is completed, consistency verification can be started to verify the data consistency between the source path and the target path. As a possible implementation method, the file size, the number of files, etc. of the source data and the target path can be verified. After determining that the data is consistent, this synchronization is completed. If the synchronization fails, an alarm message can be sent to remind manual intervention.

[0093] In a possible embodiment, the above method may further include receiving a path protection request, which includes a path identifier for which user-defined permission protection is required, that is, the storage path of the data that needs to be processed as core data. After receiving this request, the Ranger interface can also be called to cancel the write permission of this storage path, that is, cancel the change permission of this path, and only grant the read permission of this path.

[0094] In a possible embodiment, a permission management platform can be preset. Users can configure data permissions in this permission management platform. Correspondingly, users can send the above path protection request based on this permission management platform. Exemplarily, users can fill in information such as the data identifier and data storage path that need path protection in a preset page of this permission management platform, and check the "permission protection" option. The permission management platform can generate a path protection request based on the user input.

[0095] As Figure 3 shown, Figure 3 FIG. is another schematic flowchart of the data anti-deletion method provided by the embodiment of the present invention, which may include the following steps:

[0096] S1. Obtain a production data write request sent by a big data engine, which includes Hive, Spark, Presto, etc. The production data write request includes the target storage path of the generated data.

[0097] S2. Call the HiveMetastore API to determine whether the production table corresponding to the target storage path is core data. If it is core data, execute S3; if it is not core data, there is no need to execute subsequent steps.

[0098] S3. Write the data and determine whether the production data has been written completely. If it has been written completely, execute S4.

[0099] S4. Call the Ranger interface to set a policy to cancel the write permission of this directory.

[0100] S5. After the policy is set, start the data synchronization service to synchronize the data, and cold-backup the data in the target storage path to the second storage space.

[0101] S6. The user customizes path protection in the permission platform and executes S4.

[0102] In practical applications, in specific scenarios, users still have the need to backtrack jobs to update data. For example, when data production is incorrect and core data needs to be updated, users need to use the core data for re-jobbing. To support users' backtracking jobs, the data anti-deletion method provided by the embodiment of the present invention may further include the following steps:

[0103] S201. Obtain the data change request sent by the user. The data change request includes the data identifier to be changed, the data storage path, user information, etc. Specifically, it can include operation SQL and operation commands, etc. The above user information can include the user identifier.

[0104] S202. Send the data change request to a preset client for approval. When receiving the message that the data change request is approved, call the Ranger interface to add write permission to the change directory to which the changed data belongs, and prohibit the user's permission to delete the changed data.

[0105] Specifically, the rm and put-f permissions of the user can be prohibited, so that the user cannot log in to other terminals to execute rm and put-f operations. In this way, preventing the user from using other terminals to execute operations such as rm and put-f can avoid data being accidentally deleted. Among them, the rm command and the put-f command are used to delete files and forcibly overwrite files in the Linux system respectively.

[0106] S203. Monitor the change status of the changed data. When the changed data is changed and completed, cancel the write permission of the directory to which the changed data belongs, and synchronize the changed data to the second storage space.

[0107] In this step, the above data write status judge can be used to judge whether the data is updated and completed based on the file with a preset completion identifier. If the update is completed, the Ranger API interface can be called to cancel the write permission of the partition directory to which the changed data belongs, and at the same time start the data synchronizer to update the data at this path in the second storage space to maintain data consistency.

[0108] As Figure 4 shown, Figure 4 is another process schematic diagram of the data anti-accidental deletion method provided by the embodiment of the present invention. Specifically, it can include the following steps:

[0109] S41. The user applies for core data change.

[0110] S42. The data administrator approves. If the approval is passed, call the Ranger interface to set the policy to add write permission to the directory to which the changed data belongs. At the same time, cancel the functions such as deletion (such as rm and put-f) on the user terminal.

[0111] S43. The user runs a job to update the data and monitors the update status of the updated data.

[0112] S44. After the update is completed, call the Ranger interface to set the policy to cancel the write permission of the directory to which the changed data belongs.

[0113] S45. Start the data synchronization service to update the full amount of data in this directory in the second storage space.

[0114] As described above, Hive Metastore is used to manage the metadata of Hive and provide services. Specifically, the storage location of the file will be recorded in hivemetastore. The actual storage location of the file in HDFS usually refers to the location storage path. If multiple tables are stored in the same location storage path, the update of one table may affect the actual storage of another table in the same location storage path. Therefore, in order to prevent the data tables storing core data from being affected by other tables, in one possible embodiment, the above data anti-deletion method may further include the following steps:

[0115] S301. Receive a table creation request, which includes core data attribute information and target location path information.

[0116] This table creation request can be sent by the user through any big data engine, specifically an SQL table creation statement.

[0117] S302. When the core data attribute information is true, determine whether there is an already created table with a location path the same as the target location path.

[0118] In this step, a listener can be preset, and this listener is used to monitor the table creation operation. When a table creation operation is captured, determine whether the core data attribute information included in the table creation statement is true. Based on the above example, it can be judged whether the is_core_data attribute in TBLPROPERTIES in the table creation statement is true. If it is true, it is identified as a core table.

[0119] If the table to be created is a core table, it can be determined whether there are other tables in the target location path. If so, prohibit table creation and report an error.

[0120] S303. When there is an already created table with a storage location path the same as the target location path, reject the creation of the to-be-created table and return a table creation failure message.

[0121] As Figure 5 shown, Figure 5 is a schematic flow diagram for creating a table in an embodiment of the present invention, which may specifically include the following steps:

[0122] S51. The user initiates a table creation operation through computing engines such as Hive, Spark, and Presto.

[0123] S52. The HiveMetastore Hook listener captures CREATE_TABLE type operations at all times.

[0124] S53. After capturing this operation, it is judged whether the is_core_data attribute in TBLPROPERTIES in the table creation statement is true. If it is true, it is identified as a core table. If it is not a core table, the subsequent steps do not need to be executed.

[0125] S54. Identify the location of the table creation, and call a third-party service to judge whether this location is the same as the storage paths of other tables / libraries. If they are the same, table creation is prohibited and an error is reported.

[0126] As Figure 6 shown, Figure 6 is a schematic architecture diagram of the data anti-deletion method provided by the embodiment of the present invention, which can be divided into two parts: new data writing and old data reading and writing;

[0127] The new data writing process includes: the user initiates a data job through the big data engine, and the core data recognizer judges whether the data to be written is core protected data. If it is core protected data, the job starts and data writing begins. During the data writing process, its data writing completion status is judged through the data writing status. After writing is completed, the write prohibition policy for this written data is enabled through the data permission scheduler, and the core data synchronization is backed up by the data synchronizer.

[0128] The old data reading and writing process includes: the user initiates a big data job through the big data engine, and the data permission scheduler authenticates the data that the user needs to read and write. When the user has the data reading and writing permission, the user is allowed to access the data stored in HDFS.

[0129] The embodiment of the present invention constructs an intelligent and multi-level protection system, and through the "intelligent permission weaving network" technology, realizes fine-grained access control for different users and roles. Secondly, the "operation behavior analysis service" is introduced, and the service dynamically monitors and analyzes the user operation mode, and warns and blocks abnormal or potentially dangerous requests. Greatly reduce the risk of core data loss. The organic combination of this series of measures provides a strong guarantee for the core data security in the big data environment, effectively solving the problems of human errors and security vulnerabilities existing in traditional data protection means. And it can achieve the protection purpose of core data in different dimensions without invasive adaptation of the computing engine, completely avoiding data loss.

[0130] Based on the same inventive concept, the embodiment of the present invention also provides a data anti-deletion device. As Figure 7 shown, the device 700 may include:

[0131] A receiving module 701, configured to receive a data writing request, where the data writing request includes an identifier of the data to be written;

[0132] A writing module 702, configured to write the data to be written into a first storage space when the data to be written is preset core data;

[0133] An authority limiting module 703, configured to determine whether the data to be written is written completely, and prohibit the change authority of the data to be written when the status of the data to be written is written completely; and synchronize the data to be written to a second storage space.

[0134] In a possible embodiment, the data writing request further includes: a target storage location of the data to be written in the first storage space; the apparatus further includes:

[0135] A judgment module, configured to obtain metadata information of a target storage table corresponding to the target storage location, where the metadata information includes core data attributes of the storage table; and determine that the data to be written is preset core data when the core data attributes of the target storage table are true;

[0136] The determination of whether the data to be written is written completely includes:

[0137] Monitoring whether a file with a preset completion identifier is generated in a target storage partition corresponding to the target storage location, where the file with the preset completion identifier is a file generated by a big data engine after finishing generating production data;

[0138] Determining that the data to be written is written completely when the file with the preset completion identifier is generated in the target storage partition.

[0139] The receiving module is further configured to receive a path protection request, where the path protection request includes a path identifier that needs to be protected by user definition;

[0140] The authority limiting module is configured to call a Ranger interface to prohibit the change authority of a path corresponding to the path identifier;

[0141] The receiving module is further configured to obtain a data change request, where the data change request includes a change operation command, changed data, and user information;

[0142] The apparatus further includes:

[0143] An approval module, configured to send the data change request to a preset client for approval, and in the case of receiving a message indicating that the data change request is approved, call the Ranger interface to add write permission to the change directory to which the changed data belongs, and prohibit the user's permission to delete the changed data;

[0144] A monitoring module, configured to monitor the change status of the changed data, and in the case of the completion of the change of the changed data, cancel the write permission of the directory to which the changed data belongs, and synchronize the changed data to the second storage space.

[0145] The receiving module is further configured to receive a table creation request, where the table creation request includes core data attribute information of the table to be created and a target location path;

[0146] A table creation module, configured to determine whether there is an already created table with a location path identical to the target location path in the case where the core data attribute information is true; in the case of storing an already created table with a location path identical to the target location path, reject the creation of the table to be created and return a table creation failure message.

[0147] Wherein, in the present invention, the processing of collection, storage, use, processing, transmission, provision, and disclosure of user personal information complies with the provisions of relevant laws and regulations and does not violate public order and good customs.

[0148] An exemplary embodiment of the present invention further provides an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor. The memory stores a computer program capable of being executed by the at least one processor, and the computer program, when executed by the at least one processor, is configured to cause the electronic device to execute the method according to the embodiment of the present invention.

[0149] An exemplary embodiment of the present invention further provides a non-transitory computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor of a computer, is configured to cause the computer to execute the method according to the embodiment of the present invention.

[0150] An exemplary embodiment of the present invention further provides a computer program product, including a computer program, wherein the computer program, when executed by a processor of a computer, is configured to cause the computer to execute the method according to the embodiment of the present invention.

[0151] Reference Figure 8, the structural block diagram of the electronic device 800 that can be used as the server or client of the present invention will now be described. It is an example of a hardware device that can be applied to various aspects of the present invention. The electronic device is intended to represent various forms of digital electronic computer devices, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital processors, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described herein and / or claimed.

[0152] As Figure 8 shown, the electronic device 800 includes a computing unit 801, which can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 802 or the computer program loaded from the storage unit 808 into the random access memory (RAM) 803. In the RAM 803, various programs and data required for the operation of the electronic device 800 can also be stored. The computing unit 801, the ROM 802, and the RAM 803 are connected to each other through a bus 804. The input / output (I / O) interface 805 is also connected to the bus 804.

[0153] Multiple components in the electronic device 800 are connected to the I / O interface 805, including: an input unit 806, an output unit 807, a storage unit 808, and a communication unit 809. The input unit 806 can be any type of device that can input information into the electronic device 800. The input unit 806 can receive input digital or character information, and generate key signal inputs related to the user settings and / or function controls of the electronic device. The output unit 807 can be any type of device that can present information, and can include but is not limited to a display, a speaker, a video / audio output terminal, a vibrator, and / or a printer. The storage unit 808 can include but is not limited to magnetic disks, optical disks. The communication unit 809 allows the electronic device 800 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks, and can include but is not limited to a modem, a network card, an infrared communication device, a wireless communication transceiver, and / or a chipset, such as a BluetoothTM device, a WiFi device, a WiMax device, a cellular communication device, and / or the like.

[0154] The computing unit 801 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 801 executes the various methods and processes described above. For example, in some embodiments, the above data anti-deletion method can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as the storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 800 via the ROM 802 and / or the communication unit 809. In some embodiments, the computing unit 801 can be configured to execute the above data anti-deletion method in any other suitable way (e.g., by means of firmware).

[0155] The program code for implementing the method of the present invention can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing devices, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The program codes can be executed entirely on the machine, partially on the machine, executed as an independent software package partially on the machine and partially on a remote machine, or executed entirely on a remote machine or server.

[0156] In the context of the present invention, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0157] As used in this invention, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, apparatus, and / or device (e.g., magnetic disks, optical disks, memory, programmable logic devices (PLD)) used to provide machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal used to provide machine instructions and / or data to a programmable processor.

[0158] To provide for interaction with a user, the systems and techniques described here can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic, speech, or tactile input).

[0159] The systems and techniques described here can be implemented in a computing system that includes a back-end component (e.g., as a data server), or a computing system that includes a middleware component (e.g., an application server), or a computing system that includes a front-end component (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described here), or a computing system that includes any combination of such back-end, middleware, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.

[0160] A computer system can include clients and servers. Clients and servers are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other.

Claims

1. A method for preventing accidental deletion of data, characterized in that: The method comprises: receiving a data writing request, wherein the data writing request includes an identifier of the written data; When the written data is preset core data, writing the written data into the first storage space; Determine whether the write data is written completely, and if the write data status is that the write data is written completely, prohibit the change permission of the write data; and synchronize the write data to the second storage space.

2. The method according to claim 1, characterized in that The data writing request also includes: the target storage location of the written data in the first storage space; the method also includes: Obtaining metadata information of a target storage table corresponding to the target storage location, wherein the metadata information includes core data attributes of the storage table; When the core data attribute of the target storage table is a true value, the written data is determined to be preset core data.

3. The method according to claim 2, characterized in that The step of determining whether the writing of the written data is completed comprises: Monitoring whether a file with a preset completion mark is generated in the target storage partition corresponding to the target storage location, wherein the file with the preset completion mark is a file generated by the big data engine after the data production is completed; In the case where the file with the preset completion mark is generated in the target storage partition, it is determined that the writing of the write data is completed.

4. The method according to claim 1, characterized in that: The method further comprises: Receiving a path protection request, wherein the path protection request includes a user-defined path identifier that needs to be protected by permissions; Call the Ranger interface to prohibit the change permission of the path corresponding to the path identifier.

5. The method according to claim 1, characterized in that The method further comprises: Obtaining a data change request, wherein the data change request includes a change operation command, change data, and user information; Send the data change request to the preset client for approval. When receiving the approval message of the data change request, call the Ranger interface to add write permission to the change directory to which the changed data belongs, and prohibit the user's change data deletion permission; The change status of the changed data is monitored, and when the changed data is changed, the write permission of the directory to which the changed data belongs is cancelled, and the changed data is synchronized to the second storage space.

6. The method according to claim 1, characterized in that The method further comprises: Receive a table creation request, wherein the table creation request includes core data attribute information of the table to be created and a target location path; In the case where the core data attribute information is a true value, determining whether there is a created table with a location path that is the same as the target location path; In the case of a created table whose storage location path is the same as the target location path, the creation of the table to be created is rejected and table creation failure information is returned.

7. A device for preventing accidental deletion of data, characterized in that: The device comprises: A receiving module, used for receiving a data writing request, wherein the data writing request includes an identification of the written data; A writing module, used for writing the written data into the first storage space when the written data is preset core data; The permission limiting module is used to determine whether the writing of the data is completed, and when the writing state of the data is completed, prohibit the change permission of the writing data; and synchronize the writing data to the second storage space.

8. The device according to claim 7, characterized in that The data writing request also includes: the target storage location of the written data in the first storage space; the device also includes: A judgment module, used for obtaining metadata information of a target storage table corresponding to the target storage location, wherein the metadata information includes a core data attribute of the storage table; and determining that the written data is preset core data when the core data attribute of the target storage table is a true value; The step of determining whether the writing of the written data is completed comprises: Monitoring whether a file with a preset completion mark is generated in the target storage partition corresponding to the target storage location, wherein the file with the preset completion mark is a file generated by the big data engine after the data production is completed; In the case where the file with the preset completion mark is generated in the target storage partition, it is determined that the writing of the write data is completed. The receiving module is further used to receive a path protection request, wherein the path protection request includes a user-defined path identifier that needs to be protected by permissions; The permission limiting module is used to call the Ranger interface to prohibit the change permission of the path corresponding to the path identifier; The receiving module is further used to obtain a data change request, wherein the data change request includes a change operation command, change data and user information; The device also includes: An approval module is used to send the data change request to a preset client for approval. When receiving a message indicating that the data change request has been approved, the Ranger interface is called to add write permission to the change directory to which the changed data belongs, and to prohibit the user from deleting the changed data. The monitoring module is used to monitor the change status of the changed data, and when the changed data is changed, cancel the write permission of the directory to which the changed data belongs, and synchronize the changed data to the second storage space. The receiving module is further used to receive a table creation request, wherein the table creation request includes core data attribute information of the table to be created and a target location path; A table creation module is used to determine whether there is a created table with a location path identical to the target location path when the core data attribute information is a true value; if there is a created table with a location path identical to the target location path, refuse to create the table to be created and return table creation failure information.

9. An electronic device, comprising: processor; as well as Memory for storing programs, The program includes instructions, which, when executed by the processor, cause the processor to perform the method according to any one of claims 1 to 6.

10. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to make a computer execute the method according to any one of claims 1-6.