Device for selecting a data processing model

The computer system addresses the limitations of rule-based data processing models by selecting data processing models based on user preferences and regulatory compliance, ensuring efficient and secure document processing.

DE202026101603U1Active Publication Date: 2026-05-07FINPENSION AG
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
DE · DE
Patent Type
Utility models
Current Assignee / Owner
FINPENSION AG
Filing Date
2026-03-23
Publication Date
2026-05-07

AI Technical Summary

Technical Problem

Existing data processing models, particularly rule-based systems, struggle with handling diverse document formats and face challenges in compliance with data protection regulations when using powerful AI language models, especially those operated externally.

Method used

A computer system that selects a data processing model based on user preferences and identification data, allowing for the choice between locally hosted and cloud-based models, ensuring compliance with data sovereignty and privacy regulations by integrating user-specific data processing instructions.

Benefits of technology

Enables efficient and compliant processing of electronic documents by selecting appropriate data processing models that align with user preferences and regulatory requirements, minimizing latency and security risks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

A computer (1) comprising a communication interface (11), a processor (12) and a data storage device (13), wherein the data storage device (13) has software configured such that the processor (12) performs the following steps: Receiving (S1) an electronic document (2) assigned to a user via the communication interface (11); Extracting (S2) user identification data of the user from the electronic document (2); Matching (S3) the extracted user identification data with user-related information of the user stored in a database (3), wherein the database (3) links user-related information from a plurality of users with the data processing instructions agreed to by the respective user for the processing of user-related data, wherein the data processing instructions define properties of data processing models; Determine (S4) the data processing instructions linked to the user; Selecting (S5) a data processing model from a plurality of data processing models (4A, 4B) based on the specified data processing instructions; and Transmitting (S6) a message indicating the selection to a data processing system so that the data processing system can process the electronic document (2) in accordance with the data processing instructions.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL AREA

[0001] The invention relates to a device for selecting a data processing model. BACKGROUND OF THE INVENTION

[0002] Automated document processing is essential for efficient processes across a wide range of industries. This processing can include identifying document content, for example, through optical character recognition (OCR), identifying the document's origin based on its content or metadata, and determining the sources from which the document was received. Furthermore, processing can include extracting information from the content or metadata, as well as downstream tasks or processes. These might include identifying specific actions, workflows, or other processes to be triggered or initiated. This can occur upon receipt of the document, the presence of certain content, or the document's origin or association with a specific source, where that source is an entity such as...It can be a specific computer, a specific network, but also a natural or legal person.

[0003] However, well-known data processing models that are strictly rule-based are limited in their capabilities and may not be able to handle all eventualities, e.g., unknown document types, or documents that are formatted or designed differently than those for which the data processing models were originally designed.

[0004] Language models have proven promising in recent years for supplementing or replacing rule-based data processing models, as they have demonstrated flexibility with regard to input types and specific requirements. However, language models, especially powerful language models trained or equipped using artificial intelligence techniques (so-called AI language models), are often opaque and, due to their complexity and the associated demands on computing power and storage capacity, are frequently operated externally, i.e., outside the computer environment of the entity that wants to process a particular document.

[0005] The use of such powerful AI language models therefore remains a key challenge, particularly with regard to compliance with regulatory requirements concerning data protection and data sovereignty. This applies especially to compliance with regulations such as the GDPR in Europe. PRESENTATION OF THE INVENTION

[0006] The object of the invention is to provide a computer that overcomes one or more of the disadvantages of the prior art. In particular, the computer aims to enable the selection of a data processing model for processing electronic documents, whereby the computer can take user preferences into account when selecting the data processing model without requiring the documents to be received with separate metadata that uniquely identifies a user. These user preferences can, for example, include the choice between a locally hosted, privacy-compliant language model and a high-performance, cloud-based language model.

[0007] A computer is revealed, comprising a communication interface, a processor, and data storage. The data storage contains software configured such that the processor performs the following steps: First, an electronic document associated with a user is received via the communication interface. This electronic document can take various digital forms, such as PDFs, text documents, or image files, and contains user identification data that could potentially identify a user, as well as other information related to the user. Subsequently, the user's identification data is extracted from the electronic document. Therefore, user identification data is data extracted from the electronic document that could potentially identify the user.

[0008] The extracted user identification data is then compared with user-related information stored in a database. This user-related information is the verified data stored in the database that allows for a unique identification.

[0009] The database links user-related information from multiple users with the data processing instructions agreed upon by each user for processing their data. These data processing instructions define the characteristics of data processing models, such as the operating location (local, cloud), security standards, and the type of data to be processed. Based on this, the data processing instructions associated with the user are determined. A data processing model is then selected from a pool of available models based on these specified data processing instructions. Finally, a message indicating the selection is transmitted to a data processing system. The data processing system is an external or internal entity responsible for the actual processing of the electronic document after the appropriate data processing model has been selected.

[0010] The computer can also be configured so that the processor performs the following steps: The received electronic document is checked for machine readability, and the received electronic document is processed using character recognition if it is not machine readable, so that the electronic document becomes machine readable.

[0011] Furthermore, the computer's processor may be configured to extract the user's user identification data from the electronic document using a language model.

[0012] Another configuration of the computer provides that the processor is also set up to divide the extracted user identification data into groups. These groups comprise a first group, which includes one or more identifiers (e.g., social security number, customer ID, internal employee number) that uniquely identify the user, provided they conform to a predefined format, and a second group, which includes personal data (e.g., first name, last name, date of birth) that can (indirectly) uniquely identify the user, provided it meets predefined criteria. The extracted user identification data is only compared with user-related information if the extracted user identification data uniquely identifies the user.

[0013] The computer can also be configured so that the processor assigns the combination of properties defined by the data processing instructions to exactly one of the available data processing models.

[0014] Furthermore, a computer-implemented method for selecting a data processing model is disclosed, comprising the following steps: First, an electronic document assigned to a user is received via a communication interface. The user's identification data is then extracted from the electronic document. This extracted user identification data is then compared with user-related information stored in a database. This database links user-related information from multiple users with the data processing instructions agreed upon by each user for processing their data. These data processing instructions define the properties of data processing models. Based on these properties, the data processing instructions associated with the user are determined, and a data processing model is selected from a pool of available models based on the specified data processing instructions.Finally, a message indicating the selection is transmitted to a data processing system so that the data processing system can process the electronic document in accordance with the data processing instructions.

[0015] The computer-implemented procedure may further include the following steps: It is checked whether the received electronic document is machine-readable, and if it is not machine-readable, the received electronic document is processed using character recognition so that the electronic document becomes machine-readable.

[0016] In a further embodiment of the computer-implemented procedure, user identification data is extracted from the electronic document using a language model.

[0017] The computer-implemented process can also be designed to divide the user identification data into groups. These groups comprise a first group containing one or more identifiers (e.g., social security number, customer ID, internal employee number) that uniquely identify the user, provided they conform to a predefined format, and a second group containing personal data (e.g., first name, last name, date of birth) that can (indirectly) uniquely identify the user, provided they meet predefined criteria. The extracted user identification data is only compared with user-related information if the extracted user identification data uniquely identifies the user.

[0018] Finally, the computer-implemented procedure can be designed such that the combination of properties defined by the data processing instructions is assigned to exactly one of the available data processing models. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Aspects of the invention are explained in more detail with reference to the exemplary embodiments shown in the following figures and the accompanying description. Fig. Figure 1 shows a diagram that schematically represents a system architecture of the computer system; Fig. 2 shows a block diagram that schematically represents a computer; Fig. Figure 3 shows a flowchart illustrating a procedure for selecting a data processing model. DETAILED DESCRIPTION OF THE EXECUTION EXAMPLES

[0020] Fig.Figure 1 schematically shows a computing system R for selecting a data processing model. The computing system R comprises a computer 1, which is configured to process electronic documents.

[0021] Computer 1 is connected to database 3 via a bidirectional communication link V1. Database 3 is configured to store user-related data for multiple users and to link this data with the data processing instructions agreed to by each user. The bidirectional link V1 enables computer 1 to compare the user identification data extracted from a received electronic document 2 with the user-related data stored in database 3 and to retrieve the corresponding data processing instructions.

[0022] Database 3 can be implemented as a relational database, such as an SQL database. Furthermore, database 3 is preferably a dedicated database and can be provided by a separate vendor or service. However, database 3 can also be implemented on computer 1.

[0023] Furthermore, computer 1 exhibits two exemplary bidirectional communication connections V2, V3 to AI language models 4A, 4B, which are examples of the majority of data processing models. The AI ​​language models 4A, 4B can, for example, be so-called large language models trained using artificial intelligence techniques.

[0024] For example, the first AI language model 4A could be a so-called large language model, which is operated and made available in a dedicated data center. Examples of AI language models 4A include Gemini from Google, ChatGPT from OpenAI, or Claude from Anthropic.

[0025] The bidirectional communication link V2 encompasses data networks located between computer 1 and the AI ​​language model 4A, such as the internet. Interaction with the AI ​​language model 4A is possible via a standardized interface, such as an Application Programming Interface (API). These APIs allow computer 1 to access the functionalities of the language model 4A. Since the data center operating the AI ​​language model 4A is sometimes located in remote countries, data must be transmitted over long distances, increasing latency and exposing the use of this AI language model 4A to certain risks, as it requires a stable and functioning communication link V2.

[0026] The second AI language model 4B, in turn, can be another AI language model 4B with different properties than the first AI language model 4A. For example, AI language model 4B can be implemented by Llama from Meta. It can be run locally, either on computer 1 itself or, as shown, in a shared computing system S that includes the computer, database 3, and AI language model 4B. AI language model 4B can be implemented on separate hardware and exchange data with computer 1 via a communication link V3. This communication link V3 could, for example, be a local network.

[0027] As in Fig. As shown in Figure 2, the computer 1 comprises a communication interface 11, a processor 12 and a data storage device 13.

[0028] Depending on the embodiment, the processor 12 comprises a system-on-a-chip (SoC), a central processing unit (CPU), and / or other more specific processing units such as a graphics processing unit (GPU), application-specific integrated circuits (ASICs), reprogrammable processing units such as field-programmable gate arrays (FPGAs), and processing units specifically configured to accelerate certain applications. These include, for example, AI accelerators for describing processes in neural networks and / or machine learning methods.

[0029] Communication interface 11 is configured to initiate data exchange with other computers and devices, as described above, via data connections V1, V2, and V3. Data exchange preferably takes place via a wired broadband connection, although certain intermediate networks also transmit data wirelessly.

[0030] The data storage device 13 consists of one or more volatile and / or non-volatile memory components. The memory components can be replaceable and / or non-replaceable and can also be wholly or partially integrated into computer 1. Examples of memory components are RAM (Random Access Memory), flash memory, hard disks, data storage devices, and / or tape storage. The data storage device 13 comprises a non-volatile, computer-readable medium on which computer program code is stored, configured to control the processor 12 so that computer 1 performs one or more steps and / or functions as described herein. Depending on the embodiment, the computer program code is compiled or uncompiled program logic and / or machine code. Accordingly, computer 1 is configured to perform one or more steps and / or functions. The computer program code defines a discrete software application and / or is part of one.The person skilled in the art understands that the computer program code can also be distributed across a multitude of software applications. In one embodiment, the computer program code further provides interfaces, such as APIs, so that software applications, functionalities, and / or data can be accessed remotely. The data storage 13 can comprise one or more databases.

[0031] Fig. Figure 3 shows a computer-implemented method 100 for automatically selecting a data processing model based on user preferences or usage requirements, comprising an exemplary sequence of steps S1-S6 that are executed upon receipt of an electronic document. The computer-implemented method 100 can, in particular, be executed by a processor of a computing system R as described herein.

[0032] In the first step, S1, an electronic document assigned to a user is received via a communication interface. The electronic document can be a reproduction (e.g., a scan) of a physical document. However, the electronic document may also have originally existed only as an electronic document.

[0033] The electronic document can be, for example, a PDF file containing comprehensive information relating to the user. This information can relate to registrations, questionnaires, forms, identification documents, invoices, declarations, or other notifications connected to the user. The electronic document also includes user identification data that allows the user to be identified. This user identification data can include, for example, a surname, first name, date of birth, home address, telephone number, identification number, and insurance number (e.g., a social security number).

[0034] The received electronic document is first checked for machine readability. If it is not machine readable, it is processed using character recognition to create a machine-readable document. The computer may include an optical character recognition (OCR) software module for this purpose. Modern software modules, such as those from LightOn, convert documents into cleanly structured data regardless of the file format (e.g., PDF, scan, image). By consistently pre-processing all documents using such models, an explicit machine readability check can be eliminated.

[0035] In one embodiment, a separate language model is used for character recognition, which is configured to receive one or more images of the electronic document and to provide the characters of the document.

[0036] This language model is a software module and can preferably be implemented on the computer itself or in a local network environment of the computer, thus enabling processing with minimal latency. In one embodiment, the language model is an AI language model, for example, an open-source language model such as Llama from Meta. The electronic document or a portion thereof is fed into the language model along with appropriate instructions (a so-called "prompt"), so that the language model returns the characters contained in the document. Preferably, the language model is configured to return the characters along with their markup, so that the structure of the information in the document is preserved, i.e., the layout, paragraphs, tables, etc., are extracted.

[0037] In the second step, S2, user identification data is extracted from the machine-readable document. This step can also be performed using a language model trained to recognize and extract relevant user identification data from unstructured text. The language model can be the same one described above in relation to step S1.

[0038] In another embodiment, document properties are determined. Certain document properties can be determined directly or immediately via the electronic document, such as document metadata. Other document properties can be determined via the language model, for example, a document type, a document class, and / or specific document content.

[0039] In step S3, the extracted user identification data is compared with user-related information stored in a database. This comparison is made by matching the extracted user identification data with the user-related information stored in the database. If the extracted user identification data is faulty or incomplete, a standard data processing model can be selected.

[0040] In a further embodiment of the procedure, the extracted user identification data can be divided into groups, as shown in the table below. User identification data Example group AHV number. 756,XXX,XXXX.XX A Customer ID 9218745016 A First name Max B Cash on delivery Mustermann B birth date 07.07.1977 B Place of residence 8001, Zurich B Tel. No. +41123456789 B

[0041] A first group 'A' contains one or more identifiers (e.g., social security numbers, customer ID, internal employee number) that uniquely identify the user. If the user identification data contains one or more of these identifiers, the user identification data can uniquely identify the user.

[0042] A second group, 'B', contains personal data (e.g., first name, last name, date of birth) that can only uniquely identify the user if it meets predefined criteria. These predefined criteria might include, for example, that the combination of existing user identification data only matches the personal data of a single user in the database.

[0043] During the comparison, it is determined whether the user identification data extracted from the document matches certain user-related information, i.e., whether they are identical or similar enough to assume a match.

[0044] The user-related information in the database can include, for example, one or more of the following: surname, first name, date of birth, residential address, telephone number, ID number, or insurance number (e.g., a social security number). If the extracted user identification data cannot be assigned to a single database entry, a predefined standard data processing model is selected, as described above.

[0045] Otherwise, in the fourth step S4, the data processing instructions are determined based on the unique match. The data processing instructions comprise user-defined preferences, criteria, or other instructions for processing documents assigned to them.

[0046] The data processing instructions define the properties of the data processing models (e.g., AI language models) required for processing the respective document. As described, the data processing instructions are linked to the user via user-related information. The user is therefore able to determine which properties the data processing models must possess in order to process the electronic documents relating to them.

[0047] In one embodiment, however, further requirements can also determine the necessary properties of the data processing model, for example by specifying a minimum performance level or defining a certain security level. These further requirements can, in particular, be determined dynamically, depending on the document properties.

[0048] The characteristics of a data processing model can encompass a wide range of aspects, such as: security level, performance requirements, specific algorithms or model types, data residency requirements, cost preferences, interoperability, and model providers. The security level might include requirements for specific encryption standards, access controls, audit capabilities, or compliance with specific security certifications (e.g., ISO 27001). It could also determine whether the data processing model needs to operate in an isolated environment (sandbox) or whether certain data masking or anonymization is required before processing. Performance requirements relate to criteria such as maximum processing latency, expected throughput (number of documents per unit of time), and service availability (e.g., 99.9% uptime).Data residency requirements can specify in which geographical region or jurisdiction the data must be processed and stored (e.g., within the EU, Switzerland, locally on the computer).

[0049] With increased caution, data processing instructions can specify the selection of a data processing model that is operated locally, meaning directly on the computer or within a specific jurisdiction controlled by the user (e.g., EU, Switzerland), and meets high security levels and strict data residency requirements. If the focus is on efficiency and performance, the data processing instructions can permit the use of cloud-based language models such as Google Gemini or OpenAI ChatGPT, which may offer higher performance but have different data residency or security requirements.

[0050] Subsequently, in step S5, a data processing model is uniquely assigned from a plurality of available data processing models based on the specific data processing instructions. For example, a first AI language model, such as Gemini, might be selected for processing the electronic document because the data processing instructions do not define any special data residency requirements, or because high performance is required according to the document properties. In another example, a second AI language model, such as Llama, might be selected, which runs locally on the computer, because the data processing instructions define high data residency requirements, or again, because no special performance is required.

[0051] In a sixth step, S6, a message specifying the selected data processing model is transmitted to a data processing system. This data processing system then processes the electronic document using the assigned data processing model. For example, the data processing system can transmit or provide the electronic document to the assigned data processing model. For instance, the data processing system might call an API to do this.

[0052] The data processing system is a system configured to process the electronic document using the selected data processing model (e.g., Gemini or Llama). Processing the electronic document can include extracting information contained within the document. This processing can also include classifying the document or initiating downstream processes.

[0053] The data processing system can be a separate, external system, for example, running on another computer or server. However, in one embodiment, the data processing system can also be implemented directly on the computer using one or more software modules.

Claims

[1] A computer (1) comprising a communication interface (11), a processor (12) and a data storage (13), wherein the data storage (13) includes software configured such that the processor (12) performs the following steps: Receiving (S1) an electronic document (2) assigned to a user via the communication interface (11); Extracting (S2) user identification data of the user from the electronic document (2); Matching (S3) the extracted user identification data with user-related information of the user stored in a database (3), wherein the database (3) links user-related information from a plurality of users with the data processing instructions agreed to by the respective user for the processing of user-related data, wherein the data processing instructions define properties of data processing models; Determine (S4) the data processing instructions linked to the user; Selecting (S5) a data processing model from a plurality of data processing models (4A, 4B) based on the specified data processing instructions; and Transmitting (S6) a message indicating the selection to a data processing system so that the data processing system can process the electronic document (2) in accordance with the data processing instructions. [2] The computer (1) according to claim 1, wherein the processor (12) is configured to process the received electronic document (2) by means of character recognition if it is not machine-readable, so that the electronic document (2) becomes machine-readable. [3] The computer (1) according to one of claims 1 or 2, wherein the processor (12) is configured to perform the extraction of user identification data of the user from the electronic document (2) by means of a language model. [4] The computer (1) according to any one of claims 1 to 3, wherein the processor (12) is configured to divide the extracted user identification data into groups, comprising: a first group comprising one or more identifiers, which uniquely identify the user if the one or more identifiers conform to a predefined format, and a second group comprising personal data that can uniquely identify the user, provided it meets predefined criteria; The extracted user identification data is only compared with user-related information provided by the user if the extracted user identification data uniquely identifies the user. [5] The computer (1) according to any one of claims 1 to 4, wherein the processor (12) is configured to assign the combination of properties specified by the data processing instructions to exactly one of the available data processing models.