Data collaboration system and anonymization control system
By using a data collaboration system and an anonymization control system, the problem of data storage before anonymization is solved, ensuring that data is properly deleted after anonymization and in case of errors, thus protecting personal information security.
Patent Information
- Application Number
- CN202110303286.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-03-25
- Filing Date
- 2021-03-22
- Publication Date
- 2026-01-13
- Estimated Expiration
- 2041-03-22
AI Technical Summary
In existing technologies, data before anonymization may be temporarily stored after anonymization is completed, leading to the risk of personal information leakage.
Through data collaboration and anonymization control systems, it is ensured that after anonymization, the party storing the unanonymized data in the data collection and storage system deletes that data, including re-execution in case of errors during anonymization and data deletion if errors are not eliminated.
It enables the appropriate deletion of data before anonymization after anonymization, thus protecting personal information and ensuring that the privacy of data users and providers is not leaked.
Smart Images

Figure CN113448508B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a data collaboration system and an anonymization control system for collecting and storing data held by an information system. Background Technology
[0002] Previously, techniques for anonymizing data were known.
[0003] However, the following problem exists: the unanonymized data temporarily stored in order to generate the anonymized data may remain stored after the anonymized data is generated. Summary of the Invention
[0004] The purpose of this invention is to provide a data collaboration system and an anonymization control system capable of appropriately deleting data before anonymization.
[0005] The present invention discloses a data collaboration system, characterized in that it comprises: a data collection system for collecting data held by an information system; a data storage system for storing the data collected by the data collection system; and an anonymization control system for controlling the anonymization of the data stored by the data storage system, wherein the data storage system performs anonymization processing on the data, at least one of the data collection system and the data storage system stores the data before anonymization through the anonymization processing, and after performing the anonymization processing, the anonymization control system instructs the data collection system and the data storage system storing the data before anonymization to delete the data before anonymization.
[0006] The present invention discloses an anonymization control system for controlling the anonymization of data stored by a data storage system, wherein the data storage system stores data collected by a data collection system, and the data collection system collects data held by an information system. The anonymization control system is characterized in that, in the anonymization control system, the data storage system performs anonymization processing on the data; at least one of the data collection system and the data storage system stores the data before anonymization; and after performing the anonymization processing, the anonymization control system instructs the data collection system and the data storage system storing the data before anonymization to delete the data before anonymization.
[0007] The data collaboration system and anonymization control system of the present invention can appropriately delete data before anonymization. Attached Figure Description
[0008] Figure 1 This is a block diagram of a system according to one embodiment of the present invention.
[0009] Figure 2This refers to the situation where the data maintained by the information system is anonymized by the pipeline and stored by the big data analytics department. Figure 1 The diagram shows the sequence of actions of the system.
[0010] Figure 3 yes Figure 2 The sequence diagram for "data deletion processing" is shown.
[0011] Figure 4 This refers to the situation where an error occurs during the anonymization process of data held by the information system, which is then anonymized in pipelines and stored by the big data analytics department. Figure 1 The diagram shows the sequence of actions of the system.
[0012] Figure 5 In the process of data anonymization held by the information system and stored by the big data analytics department, if an error occurs during the anonymization process and the error is not eliminated... Figure 1 The diagram shows the sequence of actions of the system. Detailed Implementation
[0013] Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings.
[0014] First, the structure of the system involved in one embodiment of the present invention will be described.
[0015] Figure 1 This is a block diagram of system 10 according to this embodiment.
[0016] like Figure 1 As shown, system 10 includes a data source unit 20 that generates data and a data collaboration system 30 that enables the data generated by the data source unit 20 to collaborate.
[0017] The data source unit 20 has an information system 21 that generates data. The information system 21 has a structure management server 21a that stores the structure and settings of the information system 21. In addition to the information system 21, the data source unit 20 may also have at least one other information system. Examples of information systems include IoT (Internet of Things) systems such as remote management systems that remotely manage image forming devices such as MFPs (Multifunction Peripherals) and printers, and internal company systems such as ERP (Enterprise Resource Planning) and production management systems. Each information system may consist of one computer or multiple computers. Information systems can maintain files containing structured data. Information systems can also maintain files containing unstructured data. Information systems can also maintain databases containing structured data.
[0018] The data source unit 20 has a POST interface 22, which serves as a data collection system. The POST interface 22 acquires files of structured or unstructured data held by the information system and sends the acquired files to the pipeline described later in the data collaboration system 30. In addition to the POST interface 22, the data source unit 20 may also have at least one other POST interface with the same structure as the POST interface 22. The POST interface may also be constructed by a computer, which constitutes the information system for the POST interface itself to acquire files. Furthermore, the POST interface is also a component of the data collaboration system 30.
[0019] The data source unit 20 includes a POST agent 23, which serves as a data collection system. The POST agent 23 retrieves structured data from a database of structured data maintained by the information system and sends the retrieved structured data to the pipeline described later in the data collaboration system 30. In addition to the POST agent 23, the data source unit 20 may also include at least one other POST agent with the same structure as the POST agent 23. The POST agent can be configured by a computer, which constitutes the information system through which the POST agent retrieves structured data. Furthermore, the POST agent is also a component of the data collaboration system 30.
[0020] The data source unit 20 includes a GET target agent 24, which serves as a data collection system. The GET target agent 24 generates structured data for collaboration based on data held by the information system. In addition to the GET target agent 24, the data source unit 20 may also include at least one other GET target agent with the same structure. The GET target agents can be configured by a computer that forms the information system that holds the data used as the source for generating the structured data for collaboration. Furthermore, the GET target agent is also part of the structure of the data collaboration system 30.
[0021] The data collaboration system 30 includes: a data storage system 40 that stores data generated by the data source unit 20; an application unit 50 that utilizes the data stored in the data storage system 40; and a control service unit 60 that performs various controls on the data storage system 40 and the application unit 50.
[0022] The data storage system 40 includes a conduit 41 for storing data generated by the data source unit 20. In addition to the conduit 41, the data storage system 40 may also include at least one other conduit. Since the structure of data in an information system can vary for each information system, the data storage system 40 essentially includes conduits for each information system. Each conduit may be composed of one computer or multiple computers.
[0023] The data storage system 40 has a GET interface 42, which serves as a data collection system. The GET interface 42 retrieves files containing structured or unstructured data stored by the information system and enables the retrieved files to interact with a pipeline. In addition to the GET interface 42, the data storage system 40 may also have at least one other GET interface with the same structure as the GET interface 42. The GET interface can be configured by a computer, which acts as a pipeline for file interaction within the GET interface itself.
[0024] Furthermore, for information systems that do not correspond to files that obtain structured or unstructured data from the data storage system 40, system 10 provides a POST interface in the data source unit 20. On the other hand, for information systems that correspond to files that obtain structured or unstructured data from the data storage system 40, system 10 provides a GET interface in the data storage system 40.
[0025] The data storage system 40 includes a GET agent 43, which acts as a data collection system. The GET agent 43 acquires structured data generated by the GET target agent and enables the acquired structured data to collaborate with the pipeline. In addition to the GET agent 43, the data storage system 40 may also include at least one other GET agent with the same structure as the GET agent 43. The GET agent can be configured by a computer, which forms the pipeline for the GET agent itself to collaborate with the structured data.
[0026] Furthermore, for information systems that do not obtain structured data from the data storage system 40, system 10 includes a POST proxy in the data source section 20. On the other hand, for information systems that obtain structured data from the data storage system 40, system 10 includes a GET target proxy in the data source section 20 and a GET proxy in the data storage system 40.
[0027] The data storage system 40 includes a big data analysis unit 44, which serves as a data transformation system. This big data analysis unit 44 performs a final transformation process, converting data stored in multiple pipelines into a form that can be retrieved and aggregated using query languages such as SQL. The big data analysis unit 44 can also perform retrieval and aggregation on the data after the final transformation process based on retrieval and aggregation requests from the application unit 50. The big data analysis unit 44 can be composed of one computer or multiple computers.
[0028] The final conversion process may include data aggregation processing, which integrates data from multiple information systems, as a data conversion process. In the case where System 10 includes three remote management systems—one configured in Asia for remote management of multiple image forming devices configured in Asia, one configured in Europe for remote management of multiple image forming devices configured in Europe, and one configured in the United States for remote management of multiple image forming devices configured in the United States—each of these three remote management systems has an equipment management table that manages the image forming devices it manages. The equipment management table contains various information about the image forming devices, associated with an ID assigned to each device. Since each of the three remote management systems has its own equipment management table, it is possible that the same ID may be assigned to each image forming device across the three systems. Therefore, when the Big Data Analysis Unit 44 merges the equipment management tables from the three remote management systems to generate a single equipment management table, it needs to reassign the IDs of the image forming devices to avoid duplication.
[0029] The application unit 50 includes an application service 51, which uses data managed by the big data analytics unit 44 to perform specific actions instructed by the user, such as data display and data analysis. In addition to application service 51, the application unit 50 may also include at least one other application service. Each application service may be configured using one computer or multiple computers.
[0030] The application unit 50 includes an API platform 52, which provides an API (Application Programming Interface) for performing specific actions using data managed by the big data analytics unit 44. The API platform 52 can be configured as a single computer or multiple computers. For example, APIs provided by the API platform 52 may include: an API that sends data on the remaining quantity of consumables collected from the image forming apparatus by a remote management system to a consumables ordering system outside the system 10, which orders consumables such as toners for the image forming apparatus when the remaining quantity is below a certain amount; and an API that sends various data collected from the image forming apparatus by a remote management system to a fault prediction system outside the system 10 that predicts faults in the image forming apparatus.
[0031] The control service unit 60 includes a pipeline coordinator 61, which serves as a processing monitoring system. The pipeline coordinator 61 monitors the processing of data at various stages in the data source unit 20, data storage system 40, and application program unit 50. The pipeline coordinator 61 can be configured as a single computer or multiple computers. The pipeline coordinator 61 controls the anonymization of data stored in the data storage system 40, constituting the anonymization control system of this invention.
[0032] The control service unit 60 includes a structure management server 62, which stores the structure and settings of the data storage system 40 and automatically performs deployments as needed. The structure management server 62 can consist of one computer or multiple computers. The structure management server 62 constitutes a structure change system for the data change collaboration system 30.
[0033] The control service unit 60 includes a structure management gateway 63, which is connected to the structure management server of the information system and collects information for detecting structural changes related to databases and unstructured data in the information system, i.e., changes in the structure of data in the information system. The structure management gateway 63 can be composed of one computer or multiple computers.
[0034] The control service unit 60 has a key management service 64, which encrypts and stores security information such as key information and connection strings required for collaboration between information systems and other systems. The key management service 64 can be composed of one computer or multiple computers.
[0035] The control service unit 60 has a management API 65 that accepts requests from the data storage system 40 and the application unit 50. The management API 65 can be composed of one computer or multiple computers.
[0036] The control service unit 60 includes an authentication and licensing service 66, which performs authentication and licensing of the application services of the application unit 50. The authentication and licensing service 66 can be configured using one computer or multiple computers. For example, the authentication and licensing service 66 can verify whether an application service is authorized to request data updates from the information system stored in the data storage system 40.
[0037] Next, the operation of system 10 will be explained.
[0038] First, the operation of system 10 is explained when the data held by information system 21 is anonymized by pipeline 41 and stored by big data analysis department 44.
[0039] Figure 2This is a sequence diagram of the actions of system 10 when the data held by information system 21 is anonymized by pipeline 41 and stored by big data analysis department 44. Figure 3 yes Figure 2 The sequence diagram for "data deletion processing" is shown.
[0040] like Figure 2 and Figure 3 As shown, POST interface 22 obtains the data held by information system 21 from information system 21 (S101) and sends the data obtained in S101 to pipe 41 (S102). That is, POST interface 22 transmits the data held by information system 21 from information system 21 to pipe 41.
[0041] When the data transmission from information system 21 to pipeline 41 is finished, POST interface 22 notifies pipeline coordinator 61 that the data transmission from information system 21 to pipeline 41 has ended (S103).
[0042] When data is sent from POST interface 22 in S102, pipe 41 stores the data sent from POST interface 22 in S102 (S104).
[0043] When the data storage in S104 is finished, pipe 41 notifies pipe coordinator 61 that the storage of the data sent from POST interface 22 has ended (S105).
[0044] After the data storage in pipe 41 in S104 is completed, a conversion process is performed to convert the data transmitted from information system 21 in S102 into a form for storage in big data analysis unit 44 (S106).
[0045] When the conversion process in S106 is completed, pipe 41 notifies pipe coordinator 61 that the conversion process has been completed (S107).
[0046] When the conversion process in S106 is completed, pipe 41 performs anonymization processing (S108) on the data for which the conversion process has been performed in S106, anonymizing the personal information. Here, the anonymization process is, for example, as follows: for personal information that would prevent other processes from executing normally if deleted, the personal information is converted into information that makes the personal information difficult to determine; and for personal information that would still allow other processes to execute normally even after deletion, the personal information is deleted.
[0047] When the anonymization process in S108 finishes, pipe 41 notifies pipe coordinator 61 that the anonymization process has finished (S109).
[0048] When the anonymization process in S108 is completed, pipe 41 will send the data that has undergone anonymization in S108 to the big data analysis department 44 (S110).
[0049] When data is sent from pipe 41 in S110, the big data analysis unit 44 stores the data sent from pipe 41 in S110 (S111).
[0050] When the storage of data in S111 is finished, the big data analysis department 44 notifies the pipeline coordinator 61 that the storage of data sent from pipeline 41 has ended (S112).
[0051] Upon receiving the notification in S112, the pipeline coordinator 61 instructs the POST interface 22 to delete the data temporarily stored in the POST interface 22 for storing the anonymized data in the big data analysis unit 44 (S113).
[0052] Upon receiving the instruction in S113, POST interface 22 deletes all data that was temporarily stored in itself for storing the anonymized data in the big data analysis unit 44 (S114). For example, if POST interface 22 itself stores data transmitted from information system 21 to pipeline 41, POST interface 22 deletes the data.
[0053] After processing in S113, the pipeline coordinator 61 instructs the pipeline 41 to delete the temporary data stored in the pipeline 41 for storing data to the big data analysis unit 44 (S115).
[0054] Upon receiving the instruction in S115, pipe 41 will delete all the data temporarily stored within itself for the purpose of storing the anonymized data in the big data analysis department 44 (S116). For example, pipe 41 will delete the data stored in S104, the data that underwent transformation processing in S106, and the data that underwent anonymization processing in S108.
[0055] Next, the actions of system 10 when an error occurs during the anonymization process in a series of processes in which the data held in information system 21 is anonymized by pipeline 41 and stored by big data analysis department 44.
[0056] Figure 4 This is a sequence diagram of the actions of system 10 when an error occurs in the anonymization process during a series of processes in which the data held by information system 21 is anonymized by pipeline 41 and stored by big data analysis department 44.
[0057] Pipeline coordinator 61 detects Figure 2When an error occurs during the anonymization process shown in S108 (S121), pipe 41 is instructed to perform the anonymization process again (S122).
[0058] Therefore, system 10 executes Figure 2 and Figure 3 The processing steps S108 to S116 are shown.
[0059] The above describes the situation where an error occurs during the anonymization process in a series of processes where data held in information system 21 is anonymized by pipeline 41 and stored by big data analysis unit 44. However, when an error occurs in a specific process within a series of processes where data held in information system 21 is anonymized by pipeline 41 and stored by big data analysis unit 44, similar to the situation where an error occurs during anonymization, pipeline coordinator 61 instructs the structural component that experienced the error to restart from the beginning of the process where the error occurred.
[0060] Next, the actions of system 10 will be explained when an error occurs during the anonymization process of data held in information system 21 being anonymized by pipeline 41 and stored by big data analysis department 44, and the error is not eliminated.
[0061] Figure 5 This is a sequence diagram of the actions of system 10 when an error occurs during the anonymization process of data held by information system 21, which is then anonymized by pipeline 41 and stored by big data analysis department 44, and the error is not eliminated.
[0062] like Figure 5 As shown, after the pipeline coordinator 61 performs the action of instructing the pipeline 41 to perform the anonymization process again after a certain number of times and detecting an error in the anonymization process (S141), it detects an error in the anonymization process and instructs the POST interface 22 to delete the data temporarily stored in the POST interface 22 in order to store the anonymized data in the big data analysis unit 44 (S113).
[0063] Upon receiving the instruction in S113, POST interface 22 deletes all data temporarily stored within itself for storing the anonymized data in the big data analysis unit 44 (S114). For example, if POST interface 22 itself stores data transmitted from information system 21 to pipeline 41, POST interface 22 deletes that data.
[0064] Upon receiving the notification in S114, the pipeline coordinator 61 instructs the pipeline 41 to delete the temporary data stored in the pipeline 41 for storing data to the big data analysis unit 44 (S115).
[0065] Upon receiving the instruction in S115, pipe 41 will delete all the data temporarily stored within itself for the purpose of storing the anonymized data in the big data analysis unit 44 (S116). For example, pipe 41 will delete the data stored in S104 and the data that underwent transformation processing in S106.
[0066] The above explains the situation where errors in the anonymization process are not eliminated during a series of processes in which data held by information system 21 is anonymized by pipeline 41 and stored by big data analysis unit 44. However, when errors in a specific process in a series of processes in which data held by information system 21 is anonymized by pipeline 41 and stored by big data analysis unit 44 are not eliminated, similar to the situation where errors are not eliminated during anonymization, pipeline coordinator 61 instructs the structural component storing the data to delete the data temporarily stored for storing the anonymized data in big data analysis unit 44.
[0067] The above describes the situation where data held by information system 21 is transmitted to pipe 41 via POST interface 22, anonymized by pipe 41, and stored by big data analysis department 44. However, the following situations are handled in the same way: data held by information system is transmitted to pipe via GET interface 42, anonymized by the pipe, and stored by big data analysis department 44; data held by information system is transmitted to pipe via POST proxy 23, anonymized by the pipe, and stored by big data analysis department 44; and data held by information system is transmitted to pipe via GET target proxy 24 and GET proxy 43, anonymized by the pipe, and stored by big data analysis department 44.
[0068] As mentioned above, data collaboration system 30, such as Figure 2 and Figure 3 As shown, after the anonymization process is performed by the data storage system 40, the pipeline coordinator 61 instructs the data collection system and the data storage system 40 that store the data before anonymization to delete the data before anonymization (S113 and S115), thus enabling the appropriate deletion of the data before anonymization.
[0069] When an error occurs in the processing prior to anonymization, the pipeline coordinator 61 instructs the data collection system 30, or the data storage system 40, which performed the processing, to restart the processing from the beginning (S122). Thus, the processing can be restarted using the data before anonymization stored in at least one of the data collection system and the data storage system 40, resulting in efficient re-execution of the processing.
[0070] Data collaboration system 30 Figure 5 and Figure 3As shown, if an error occurs in the processing prior to anonymization, and the error is not eliminated after re-executing the process a certain number of times from the beginning, the pipeline coordinator 61 instructs the data collection system and the data storage system 40, on the side storing the data before anonymization, to delete the data before anonymization (S113 and S115), thus enabling the appropriate deletion of the data before anonymization.
[0071] Because the data stored by the big data analysis department 44 is anonymized, the data collaboration system 30 ensures that users of the data stored by the big data analysis department 44 are protected from personal information. Furthermore, when the big data analysis department 44 stores anonymized data, the data collaboration system 30 appropriately deletes the unanonymized data, ensuring that the data provider's personal information is anonymized and that no unanonymized data remains.
Claims
1. A data collaboration system, characterized in that, include: Data collection systems collect data held by information systems; Data storage system for storing data collected by the data collection system; as well as An anonymization control system controls the anonymization of data stored by the data storage system. The data storage system performs anonymization processing on the data. At least one of the data collection system and the data storage system stores the data before anonymization, processed by the anonymization procedure. When an error occurs in the processing prior to the anonymization process, the anonymization control system instructs the data collection system and the data storage system that executed the processing in which the error occurred to re-execute the processing from the processing in which the error occurred. If no errors occurred in the processing prior to the anonymization process, after the anonymization process is performed, the data collection system and the data storage system that store the data before anonymization are instructed to delete the data before anonymization.
2. The data collaboration system according to claim 1, characterized in that, If the anonymization control system fails to eliminate the error after a certain number of re-executions, it instructs the data collection system and the data storage system that store the data before anonymization to delete the data before anonymization.
3. An anonymization control system for controlling the anonymization of data stored by a data storage system, said data storage system storing data collected by a data collection system, said data collection system collecting data held by an information system, characterized in that... The data storage system performs anonymization processing on the data. At least one of the data collection system and the data storage system stores the data before anonymization, processed by the anonymization procedure. When an error occurs in the processing prior to the anonymization process, the anonymization control system instructs the data collection system and the data storage system that executed the processing in which the error occurred to restart the processing from the point where the error occurred; when no error occurs in the processing prior to the anonymization process, after the anonymization process is executed, the system instructs the data collection system and the data storage system that stores the data before anonymization to delete the data before anonymization.
Citation Information
Patent Citations
Anonymization processor, anonymization processing method, and program
JP2016139261A