Data processing system for access-restricted data
The data processing system automates personal data anonymization and manages release keys to provide transparent and secure access, addressing the challenge of manual data evaluation and compliance issues.
Patent Information
- Authority / Receiving Office
- DE · DE
- Patent Type
- Utility models
- Current Assignee / Owner
- MUNICH INNOVATION LABS GMBH
- Filing Date
- 2020-09-30
- Publication Date
- 2026-06-03
AI Technical Summary
Existing systems lack transparent and efficient methods for accessing personal data while ensuring data protection compliance, leading to potential violations of users' fundamental rights during manual data evaluation.
A data processing system with an automated anonymization module that identifies and removes personal identifiers, generating anonymized datasets, and an access module that manages release keys for selective deanonymization based on justifications, ensuring data security and compliance.
Enables transparent and targeted access to personal data, minimizing infringement on user privacy by automating anonymization and requiring case-by-case justification for deanonymization, thus improving data security and user interaction.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
AREA OF INVENTION
[0001] The present invention lies in the field of data security and relates to systems for restricted access to information from a data source, in particular for use by authorities and organizations with security tasks. BACKGROUND
[0002] With the increasing use of digital services, the publicly available data profile of each user also grows, for example, through public statements on social networks. However, access to this publicly available data profile by authorities and organizations with security responsibilities, for example, for the purpose of law enforcement, such as in cases of hate crime, is countered by fundamental ethical concerns and legal rights to exclude access, such as the right to informational self-determination.
[0003] For the work of authorities and organizations with security responsibilities, this means specifically that, without initial suspicion of a criminal offense, personal data should generally not be reviewed and analyzed. This could potentially lead to a situation where, in the area of law enforcement, digital content on social networks can only be examined manually and without documentation when there is suspicion, creating legal uncertainty for both users and authorities. This is because the manual review of data can already constitute a violation of users' fundamental rights. OVERVIEW OF THE INVENTION
[0004] Currently, no transparent access methods to personal data are known that allow for the efficient evaluation of personal digital information in a data protection-compliant manner. At the same time, the manual evaluation of personal data is counterproductive to preventing infringements of access restrictions and cannot replace such a technical solution.
[0005] The object of the invention is therefore to provide a technical solution for restricted access to personal data, while at the same time preventing non-specific viewing or evaluation of personal data.
[0006] This task is solved by a data processing system and a computer program according to the independent claims. The dependent claims relate to preferred embodiments.
[0007] According to a first aspect, the invention relates to a data processing system. The system comprises an access-restricted database and an automated anonymization module implemented in at least one computing unit. The access-restricted database is configured to receive personal data from an external data source and to store the personal data. The automated anonymization module has access to the access-restricted database and is configured to identify personal identifiers in the personal data based on a reference data structure and / or a reference data pattern, wherein the reference data pattern includes biometric features, to automatically remove the personal identifiers from the personal data in order to generate an anonymized data set, and to provide the anonymized data set.The anonymization module is configured to replace the personal identifiers identified on the basis of the reference data structure with a non-personal placeholder and / or to obscure the biometric features in the personal data.
[0008] In some embodiments, the system further comprises an access module implemented in at least one computing unit, which is configured to receive a request and a justification for the release of personal data relating to a personal identifier to be released in the anonymized dataset, and to generate a release key based on the request and the justification relating to the personal identifier to be released. The automated anonymization module is further configured to receive the release key and provide a de-anonymized dataset, wherein the anonymization module is configured to identify the personal identifier to be released in the personal data and to generate the de-anonymized dataset containing the personal identifier to be released.
[0009] The access-restricted database can temporarily store personal data and can automatically delete the personal data after a deletion period.
[0010] Using anonymized data, users can access personally identifiable information transparently, as personal content can be automatically anonymized and any increase in the level of access is subject to case-by-case justification. Accordingly, users can utilize the technical capabilities of the access system in a targeted manner, while minimizing any infringement on the data security of users of the external data source. Such a structured process also improves user interaction with a suitably programmed machine, as the inherent procedural flow of purely mental activity is reversed; that is, anonymization and the associated abstraction of data content can occur before the user views the content, and not the other way around.
[0011] The external data source can be a publicly accessible digital data source that allows users to store personal data. For example, the external data source could be a social network channel, such as a YouTube channel or a Facebook group, or a web-based blog, and could include digital contributions from a user group, such as comments or photos, which are associated with the external data source. The external data source can be automatically imported and temporarily stored in a computer system responsible for anonymizing the content of the external data source during processing. For example, the external data source can initially be provided manually, such as by providing a link, and the content of the external data source can then be (temporarily) stored by a computer system in an access-restricted database.
[0012] Preferably, personal data is stored in the access-restricted database without manual access, and an anonymized and / or deanonymized copy of the personal data is stored for evaluation or review of the content of the external data source, whereby manual access to the deanonymized copy may be enabled.
[0013] Manual access can refer to accessing the data in human-readable form at a terminal. For example, the data in the access-restricted database may be encrypted or compressed, and access to the database may depend on providing an authorization key.
[0014] The system can provide a client system and a server system, which are connected via a data line, whereby in particular only the server system is equipped with the authorization key to automatically provide deanonymized data to the client system based on personal data.
[0015] In some implementations, an automated system can access the personal data in the access-restricted database to evaluate the personal data according to predetermined criteria. For example, the content of the personal data can be evaluated by a neural network to generate an anonymized analysis of the content typical of an observed phenomenon.
[0016] In some implementations, the contents of the external data source are only temporarily stored for anonymization and / or deleted after a deletion period has expired.
[0017] The deanonymized data set can be – depending in particular on the release key – a completely deanonymized data set in which all personal identifiers are deanonymized.
[0018] Alternatively, the de-anonymized dataset can be a partially anonymized dataset in which a subset of the personal identifiers are de-anonymized. Other personal identifiers, however, can remain anonymized, depending on the release key.
[0019] In one implementation, personal data can also be successively deanonymized in order to selectively deanonymize personal identifiers across multiple release levels. In particular, the procedure described above can be iterated, with each release level requiring a corresponding request and justification.
[0020] Gradual deanonymization can be carried out, culminating in the complete deanonymization of personal data. In some implementations, deanonymization can mean the complete removal of personal data.
[0021] The de-anonymization of personal identifiers can be performed on an identifier-specific and / or content-specific basis. This means that specific personal identifiers can be selectively de-anonymized, and / or specific content within the personal data containing various personal identifiers can be selectively de-anonymized. For example, the request might include the release of all personal identifiers within a portion of the personal data, such as a comment or image, which could also be released within other parts of the personal data. Furthermore, the request could include the release / de-anonymization of a specific personal identifier within a portion of the personal data or within the entire personal data.
[0022] If deanonymized datasets with additional personal identifiers are requested, a computer system can re-collect the personal data from the external data source and regenerate the deanonymized dataset.
[0023] In preferred embodiments, the anonymization module is configured to remove the personal identifiers, except for the personal identifier to be released, from the personal data in order to generate the deanonymized data set.
[0024] For example, the computer system can anonymize the personal data upon receipt and only exclude predetermined (to be released) personal identifiers from anonymization before the contents of the external data source are provided and / or stored in deanonymized form.
[0025] The system can automatically identify and remove personal identifiers from personal data.
[0026] Personal data and / or identifying characteristics are, in one form, digital and / or digitized content that, with reasonable effort, allows conclusions to be drawn about the identity of the user. Personal data can thus include, for example, biometric data, full names, usernames, email addresses, or locations. Furthermore, complete statements can also constitute personal data if these statements can be directly correlated or associated with a specific person through an internet search.
[0027] To protect data security, the system should therefore ideally initially provide completely anonymized data, e.g., content without any specific personal reference. This can be achieved by anonymizing the content, whereby personal identifiers can be removed and optionally replaced by a non-personal placeholder (so-called pseudonymization). Personal identifiers can be automatically identified by pattern matching with the reference data pattern, e.g., based on typical patterns of biometric characteristics, typical patterns of (email) addresses and usernames (such as a prefixed "@"), or based on a standardized comment structure.
[0028] In preferred embodiments, the reference data pattern includes biometric features, in particular facial features, wherein the anonymization module is configured to make the biometric features in the personal data unrecognizable.
[0029] Computer systems can automatically recognize relevant content within graphical representations. In one embodiment, a neural network can automatically remove personal characteristics from the data, for example, by identifying facial features and then automatically and selectively obscuring them by pixelating or overlaying them with an anonymized image structure (e.g., blacking out facial features).
[0030] Furthermore, personal data published on a website or via an API can be identified based on its data structure. For example, usernames may be included in separate data fields and can be verified by analyzing the data hierarchy within the publication structure of the website content.
[0031] In preferred embodiments, the reference data structure comprises an information hierarchy of the external data source, wherein the anonymization module is configured to identify the contents of a predetermined hierarchy level of the information hierarchy as personal identifiers.
[0032] Removing personally identifiable information creates an anonymized dataset that can be analyzed without reference to any individual. Personally identifiable information can be further abstracted or anonymized by providing it in a context-independent format.
[0033] In preferred embodiments, the anonymization module is configured to generate an aggregated subset of data, wherein the subset of data particularly includes an intersection between the personal data and reference data and / or wherein the subset of data particularly includes a context-independent listing of sub-contents.
[0034] The aggregated subset of data may, for example, include content that partially matches predetermined reference data or that was selected by a neural network based on a pattern comparison with reference data.
[0035] In some embodiments, the aggregated subset includes a stochastic analysis of the content of the personal data, such as a frequency analysis of specific content, particularly phenomenon-related content, e.g., the absolute or relative frequency of using the names of known individuals from a database in the context of a statement glorifying violence. The context of the statement can be automatically determined by a computer system, for example, using a neural network, through automated language processing of the content. Furthermore, the computer system can, for example, automatically detect firearms in digital images and thus automatically provide context-independent indications of criminally relevant content. In this case, an anonymized subset can include filtered anonymized content in which graphical content is filtered through automatic object recognition.
[0036] In preferred embodiments, the access module is configured to store the request and / or the justification, and in particular to store the request and / or the justification irreversibly.
[0037] Storing the request and the justification can make it possible to consistently depend on a justification for increasing the depth of intervention and to automatically document the procedure.
[0038] The justification can be entered by the user, for example in a form, and can include, for example, a text-based explanation or a transaction number.
[0039] Preferably, the system is set up to make the granting of the release key dependent on the submission of a justification.
[0040] Furthermore, the granting of the release key may depend on authorization, such as a username-password combination, a signature card and / or a match between the user's biometric characteristics and authorized biometric characteristics (fingerprint, iris scan, facial features, etc.).
[0041] Irreversible storage can be storage in an access-restricted database with insert rights, where, for example, deletion or overwriting of the data can only be done with additional access rights, or it can be storage in a blockchain.
[0042] In preferred embodiments, the access module is further configured to store the release key for released personal identifiers with reference to a case identifier, to authenticate a user with respect to the case identifier, and to generate a deanonymized data set containing the released personal identifiers, in particular by automatically removing the personal identifiers except for the released personal identifiers from the personal data.
[0043] The stored release keys can be kept independently of the personal data, for example, to restore and / or expand a de-anonymized dataset after the personal data has been removed from a cache. Different users can access uniform release keys using the case identifier. Storing the release key can thus enable more consistent system operation.
[0044] In preferred embodiments, the system is configured to obtain first personal data from a first external data source, to obtain second personal data from a second external data source, and to map the first personal data and the second personal data to a uniform data format for generating the personal data.
[0045] The first and second external data sources may be of different formats or may not have a uniform data structure. A uniform data format can facilitate automatic anonymization of the data, as mapping to this format can ignore source-specific format properties. Simultaneously, mapping can transfer source-specific personal identifiers, such as the formatting or positioning of usernames, into a unified database to remove these identifiers consistently and across all sources. The resulting uniform data structure with identified personal identifiers can then be used across sources for anonymization and analysis of the personal data, employing standardized analysis and anonymization modules.
[0046] In preferred embodiments, the system is further configured to retrieve personal data from a related external data source. The system can be configured to automatically evaluate the related external data source and determine a similarity score between the contents of the external data source and the related external data source, and / or a relevance score by evaluating the content of the related external data source against a relevance pattern. If the similarity score and / or the relevance score exceeds a threshold, the system can display the related external data source for the inclusion of the personal data it contains and, upon inclusion, store an inclusion event with an inclusion justification.
[0047] Based on content relationships, further data sources can be automatically identified, for example, in the case of simultaneously subscribed information channels by registered users of the external data source. The similarity score can be determined based on user matches or on relationships between the external data sources. The threshold for this determined similarity score can be a function of the relative and / or absolute match between users. The relevance score can be phenomenon-related and can be determined through a semantic analysis of the related external data source.
[0048] If the similarity score and / or the relevance score exceeds a threshold, the data source can be suggested for inclusion. In this way, the selection of multiple related external data sources can be automatically filtered to improve guided human-machine interaction in an anonymized manner.
[0049] According to a second aspect, the invention relates to a computer program or computer program product with machine-readable instructions which, when executed on a processing unit, implement a system according to the second aspect. DESCRIPTION OF THE DRAWINGS
[0050] The properties according to the invention and the various advantages of the devices are best understood from a detailed description of preferred embodiments with reference to the accompanying drawings, wherein: Fig. 1. A computer system connected to an external data source is illustrated by an example; Fig. 2. A procedure for providing anonymized data is illustrated by an example; Fig. 3 illustrates a procedure for providing partially anonymized data according to an example; and Fig. Figure 4 illustrates another example of a computer system for restricted access to personal data.
[0051] Fig. Figure 1 illustrates a schematic computer system 10 connected to an external data source 12 according to an example. The system comprises a server 14, a client 16, and a database 18, wherein the server 14 is connected to the external data source 12 and provides anonymized and / or partially anonymized data records 20 to the client 16.
[0052] The external data source 12 can be connected to the server 14 via a data connection, such as the internet, and provide content through an interface. The external data source 12 can consist of a portion of the content published via the interface, such as a specific publication channel of a particular user or user group. The interface can be an API and / or a browser interface to provide content that may contain personal data. The server 14 can access the content of the external data source 12 and, among other things, read personal data, such as public statements made by individual users on a social network platform. Access to the content can occur via an automated browser and / or via requests to the API interface.
[0053] The personal data that server 14 reads from external data source 12 can be temporarily stored on server 14 or stored in database 18, which can be an access-restricted database 18.
[0054] Server 14 can automatically analyze the personal data to identify personal identifiers and perform automatic anonymization of the personal data. An anonymized data set 20 can then be provided to Client 16, which can have an interface for displaying the anonymized data set 20 in a human-readable format.
[0055] Fig. Section 2 illustrates a procedure for providing anonymized data according to an example. The procedure includes obtaining personal data from an external data source (S10) and automatically identifying personal identifiers in the personal data based on a reference data structure and / or a reference data pattern (S12). The procedure further includes automatically generating an anonymized dataset 20, including automatically removing the personal identifiers from the personal data (S14), and providing the anonymized dataset 20 (S16).
[0056] A Server 14 can identify personally identifiable information within personal data based on source-specific data structures or patterns. For example, a Server 14 can receive a response to an API request and determine usernames based on the data structure of the response, and / or identify corresponding usernames based on the position and / or formatting of entries in a publication format, such as an HTML-based website. Furthermore, a Server 14 can automatically identify personally identifiable information, such as biometric characteristics (speech patterns, facial features, etc.), in audiovisual content.
[0057] The identified personal identifiers can then be removed from the corresponding fields of the unified data structure and can also be used to anonymize the content. For example, Server 14 can anonymize personal data based on the identified personal identifiers by removing them from the personal data, e.g., by overlaying and / or replacing them, such as by pseudonymizing usernames and redacting biometric facial features in images.
[0058] The anonymized data set 20 can be stored in the database 18 and / or can be transmitted to a client 16 for the provision of the anonymized data set 20.
[0059] In some embodiments, personal data from the external data source 12 is temporarily stored by the server 14, and only anonymized and / or partially anonymized data records 20 are stored in the database 18. Furthermore, aggregated sub-content of the anonymized and / or partially anonymized data records 20 stored in the database 18 can be transferred to the client 16 to provide the client 16 with an additionally anonymized and abstracted data record 20. For this purpose, the server 14 can aggregate the anonymized and / or partially anonymized content, for example, by providing a statistical synthesis of the content according to activity periods and / or automatically captured semantic content, and provide an aggregated data record 20 with context-independent data.
[0060] In some embodiments, the personal data from the external data source 12 in the database 18 are automatically deleted after a predetermined deletion period and / or stored with different access restrictions, such as anonymized and / or partially anonymized data records 20. If a user requests partially anonymized data records 20 containing additional personal identifiers, this data can be subsequently collected from the external data source 12 and / or from the (access-restricted) database 18.
[0061] Fig. Section 3 illustrates a procedure for providing partially anonymized data by way of an example. The procedure involves receiving a request and justification for the release of personal data relating to a personal identifier to be released in the anonymized dataset (S18). The procedure further involves generating a release key based on the request and justification relating to the personal identifier to be released (S20) and automatically generating a partially anonymized dataset 20 based on the release key, wherein the partially anonymized dataset contains the personal identifier to be released (S22). Finally, the procedure involves providing the partially anonymized dataset 20 (S24).
[0062] Client 16 can, for example, request access to complete message content within a specific activity period and justify the request in a query form by indicating its proximity to a particular event. In response to this request, Server 14 can create a partially anonymized data set 20 containing complete message content for the requested activity period, while personal identifiers such as usernames remain anonymized in the partially anonymized data set 20. For example, Server 14 can retrieve anonymized data from Database 18 and select complete message content to transmit to Client 16.
[0063] Consequently, a user of Client 16 can request a specific personal identifier, such as a username, that is associated with a message substantiating suspicion of a serious crime. The request may be justified, for example, by a court order.
[0064] Server 14 can generate a release key and access personal data in the access-restricted database 18 and / or request personal data from the external data source 12 in order to directly read the personal identifier to be released, or to identify the personal identifier to be released within the personal data and remove all personal identifiers except the personal identifier to be released from the personal data. The correspondingly partially anonymized data record 20 can then be transmitted to Client 16.
[0065] The evaluation and anonymization of personal data can be performed server-side to prevent client 16 from interfering with the personal data. Preferably, the server 14 is modular, allowing subtasks for evaluation and anonymization to be assigned context-independently and executed in an internal or external server cloud.
[0066] Fig.Figure 4 illustrates a schematic computer system 10 for restricted access to personal data according to another example. A server 14 accesses an external data source 12 and provides anonymized and / or partially anonymized datasets 20 to a client 16. The server includes an input module 22 for receiving the personal data from the external data source 12. The personal data can be provided to an anonymization module 24 of the server 14, which identifies personal identifiers in the personal data and removes them to create anonymized content. The anonymized content can be provided to an access module 26, which can provide it to the client as part of anonymized and / or partially anonymized datasets 20.Furthermore, an aggregation module 28 can automatically evaluate personal data or anonymized and / or partially anonymized content according to predetermined criteria in order to provide context-independent aggregated content and / or analysis results to the access module 26. The aggregated content and / or analysis results can be transmitted to the client 16 as part of the anonymized and / or partially anonymized datasets 20.
[0067] Preferably, the modules and / or their functions are implemented as independent containers that can be run in an internal and / or external cloud. Containerizing the tasks can allow for machine-side, context-independent evaluation and anonymization of personal data, thus further improving data security. Accordingly, Server 14 is not limited to a single computing unit, but can, in various embodiments, consist of multiple computing units.
[0068] The input module can access the external data source 12 via an access authorization, such as a username-password combination. Multiple access authorizations can be stored in the computer system 10 and can be used to retrieve personal data from multiple different external data sources 12.
[0069] The input module 22 can map the personal data from the external data source 12 to a uniform data format, for example, by performing a source-specific mapping. For instance, an API interface can be provided for a specific data source, such as a particular social network, which maps the content of the external data source 12 to the uniform data format. Furthermore, the input module 22 can automatically evaluate information from a published website using an automated browser and map it to the uniform data format (so-called web scraping). Preferably, the uniform data format includes a personal data field, such as a username associated with the published content.
[0070] The personal data, structured into the standardized data format, can then be forwarded to the anonymization module 24, which automatically removes personal identifiers from the personal data. The anonymization module can delete the personal identifiers identified by the input module 22 and can also remove corresponding content from the personal data or replace it with placeholders. Furthermore, the anonymization module 24 can use pattern recognition to identify other personal identifiers, such as email addresses, biometric features, addresses, etc., and remove them to generate an anonymized dataset 20. The anonymization module 24 can also exclude previously released personal identifiers from removal.For example, the server can store personal identifiers to be released or released in a database 18, and the anonymization module 24 can exclude the released or released personal identifiers from removal in order to provide partially anonymized datasets 20.
[0071] In some embodiments, server 14 can cache the identified personal identifiers that are not released or to be released as a function argument of an anonymizing filter and apply the anonymizing filter to the personal data. The filter, as a container in an internal or external server cloud, can anonymize the personal data context-independently and return the anonymization result to an internal data manager of the anonymization module 24.
[0072] In some embodiments, the anonymization module 24 can receive an image of a person as a personal identifier to be released and can subsequently exclude the person from anonymization in the personal data. For example, the anonymization module 24 can automatically identify biometric features in images and can generate feature vectors for the identified biometric features in the personal data.
[0073] Anonymization module 24 can, for example, use a distance metric to determine whether the distance between the feature vector of the person's image and the feature vectors of the biometric features in the personal data is below a predetermined threshold. If such a match is found, anonymization module 24 can exclude the corresponding biometric feature from anonymization in the personal data, while other identified biometric features, for example in the same image, can be anonymized (e.g., pixelated, blacked out, etc.).
[0074] Access module 26 can receive the anonymized and / or partially anonymized data records 20 from anonymization module 24 and, for example, with appropriate justification, forward them to client 16. Furthermore, access module 26 can receive requests from client 16 regarding personal identifiers and, based on the request and the justification, generate release keys for the release of personal identifiers. For example, access module 26 can store personal identifiers to be released in database 18 with a justification or authorization as a release key, so that the released personal identifiers are consistently and / or across sources exempted from anonymization by anonymization module 24.
[0075] Preferably, the anonymized and / or partially anonymized data sets 20 are not transmitted completely to the client 16 without appropriate justification, but aggregated data sets 20 and anonymized and / or partially anonymized sub-data sets 20, which contain requested parts of the anonymized and / or partially anonymized data sets 20, are transmitted.
[0076] To provide aggregated datasets 20, the access module 26 can receive aggregated content from the aggregation module 28. The aggregation module 28 can analyze the anonymized and / or partially anonymized datasets 20 according to predetermined criteria and provide analysis results to the access module 26. In some embodiments, the aggregation module 28 can receive personal data prior to anonymization and can generate and provide analysis results without personal reference based on the personal data.
[0077] In some embodiments, the aggregation module 28 can execute text classification modules or object recognition modules to classify social media content based on reference data (e.g., into violence-affirming statements) or to automatically identify specific image features, such as the presence of weapons or unconstitutional symbols. The aggregation module 28 can then provide text classification metrics or incidence rates for specific objects to be identified to the access module 26 as a context-independent synthesis. In some embodiments, object recognition can also include letters, and the aggregation module 28 can incorporate semantic content from images into the text classification.
[0078] Based on the analysis, aggregated datasets 20 can be created, which can contain multiple parallel analyses, such as text classification and object recognition. The aggregated datasets 20 can then be transmitted to the client 16, and a user can use the context-independent aggregated datasets 20 to evaluate the anonymized content of the external data source 12.
[0079] A user can request anonymized subsets of data 20 based on context-independent aggregated datasets 20 and identify personal identifiers to be released based on these subsets. Based on the request and a justification, the access module 26 can select partially anonymized datasets 20, which are generated by the anonymization module 24 and can be provided to the client 16 by the access module 26, either fully or partially.
[0080] The requests that necessitate an increase in the level of intervention, and their justification, can be stored in an access-restricted database 18 to provide a transparent and traceable access procedure and to allow the user, within the framework of supported machine interaction, a logged and minimal intervention in the data security of the users of the external data source 12. In this process, the user can be provided with a procedure that, compared to a manual procedure, results in a reverse order of the level of intervention; that is, the anonymization of personal data is system-based and precedes any intervention, rather than being implemented subsequently.
[0081] The preceding description of preferred embodiments, examples, and drawings is intended only to illustrate the invention and its associated advantages and should not be construed as limiting the scope of protection. Rather, the scope of protection of the invention shall be determined solely on the basis of the accompanying claims. REFERENCE MARK LIST 10 computer systems 12 external data sources 14 servers 16 Client 18 database 20 anonymized / partially anonymized / aggregated datasets 22 Input module 24 Anonymization module 26 Access module 28 Aggregation module
Claims
Data processing system (10), wherein the system (10) comprises: an access-restricted database (18) for obtaining personal data from an external data source (12) and storing the personal data;an automated anonymization module (24) implemented in at least one computing unit with access to the access-restricted database (18), which is configured to: - identify personal identifiers in the personal data on the basis of a reference data structure and / or a reference data pattern, wherein the reference data pattern includes biometric features; - automatically remove the personal identifiers in the personal data in order to generate an anonymized data set (20), and wherein the anonymization module (24) is configured to replace the personal identifiers identified on the basis of the reference data structure with a non-personal placeholder and / or to obscure the biometric features in the personal data; and - provide the anonymized data set (20). System according to claim 1, which further comprises an access module (26) implemented in at least one computing unit, which is configured to: - receive a request and a justification for the release of personal data relating to a personal identifier to be released in the anonymized data set (20), - generate a release key based on the request and the justification relating to the personal identifier to be released, wherein the automated anonymization module (24) is further configured to receive the release key and to provide a de-anonymized data set (20), wherein the anonymization module (24) is configured to: - identify the personal identifier to be released in the personal data, and - generate the de-anonymized data set (20) which contains the personal identifier to be released. System (10) according to claim 2, wherein the anonymization module (24) is configured to remove the personal identifiers except for the personal identifier to be released from the personal data in order to generate the deanonymized data set (20). System (10) according to claim 2 or 3, wherein the access module (26) is configured to store the request and / or the justification, wherein the request and / or the justification is stored irreversibly. System (10) according to one of claims 2 to 4, wherein the access module (26) is further configured to: store the release key for released personal identifiers with reference to a case identifier, authenticate a user with regard to the case identifier, generate a deanonymized data record (20) which contains the released personal identifiers, in particular by automatically removing the personal identifiers except for the released personal identifiers from the personal data. System (10) according to one of claims 1 to 5, wherein the reference data structure comprises an information hierarchy of the external data source (12) and wherein the anonymization module (24) is configured to identify the contents of a predetermined hierarchy level of the information hierarchy as personal identifiers. System (10) according to one of claims 1 to 6, wherein the system (10) is configured to: obtain first personal data from a first external data source (12), obtain second personal data from a second external data source (12), and map the first personal data and the second personal data to a uniform data format for generating the personal data. System (10) according to one of claims 1 to 7, wherein the anonymization module (24) is configured to generate an aggregated sub-dataset (20), wherein the sub-dataset (20) in particular comprises an intersection between the personal data and reference data and / or wherein the sub-dataset (20) in particular comprises a context-independent list of sub-contents. System (10) according to any one of claims 1 to 8, wherein system (10) is further configured to obtain personal data from a related external data source (12), to automatically evaluate the related external data source (12) and to determine a similarity value between the contents of the external data source (12) and the related external data source (12) and / or a relevance value by evaluating the content of the related external data source (12) in relation to a relevance pattern, and, if the similarity value and / or the relevance value is above a threshold, to display the related external data source (12) for the inclusion of the contained personal data and, in the event of inclusion, to store an inclusion event with an inclusion justification. Computer program or computer program product with machine-readable instructions which, when executed on a processing unit, implement a system (10) according to any one of claims 1 to 9.