Circuit and computer program

A dynamic processing architecture with specialized stages and LLMs addresses the lack of configurability in digital document analysis, enabling structured annotation and complex analysis patterns.

DE202025105688U1Active Publication Date: 2025-11-27CODEFY GMBH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
DE202025105688
Authority / Receiving Office
DE · DE
Patent Type
Utility models
Current Assignee / Owner
Filing Date
2025-09-22
Publication Date
2025-11-27
Estimated Expiration
2035-09-30

AI Technical Summary

Technical Problem

Existing technologies lack a highly configurable framework for automated, multi-stage analysis and generation of annotations for digital documents, limiting the ability to support complex analysis patterns and diverse analytical goals.

Method used

A dynamic, sequential processing architecture with functionally specialized stages that operate on data segments, configured through declarative methods, including Reinforcement Learning from Human Feedback (RLHF) and Large Language Models (LLMs), enabling sophisticated, multi-layered analysis of digital documents.

Benefits of technology

Enables the transformation of raw digital documents into structured, queryable annotation objects, supporting complex analysis patterns and diverse analytical goals, including precise pattern matching and context-aware semantic filtering.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000024_0000
    Figure 00000024_0000
  • Figure 00000024_0001
    Figure 00000024_0001
  • Figure 00000025_0000
    Figure 00000025_0000
Patent Text Reader

Abstract

Circuit for processing the content of a digital document, whereby the circuit is set up for: Receiving (11) configuration data for the content preparation of a digital document, wherein the configuration data indicates a processing chain with processing modules from a predefined set of processing modules for the content preparation of a digital document; Generating (12) a processing chain according to the configuration data; Receiving (16) a digital document; and Applying (18) the processing chain to the digital document to prepare the digital document.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] The present invention relates to a circuit and a computer program.

[0002] Digital documents are generally known. Digital documents can, for example, be analyzed by machines.

[0003] The object of the present invention is to provide an improved circuit and an improved computer program.

[0004] This problem is solved by the subject matter of claims 1 and 20.

[0005] Further aspects and features of the present invention will become apparent from the dependent claims, the accompanying drawing and the following description of preferred embodiments.

[0006] Exemplary embodiments of the invention will now be described by way of example and with reference to the accompanying drawing, in which: Fig. 1 An example of a circuit for processing the content of a digital document is illustrated; Fig. 2. An exemplary embodiment of a method for the content processing of a digital document is illustrated; Fig. 3 illustrates an embodiment of a method with a first user operation in a graphical user interface; Fig. 4 illustrates an embodiment of a method with a second user operation in a graphical user interface; Fig. 5 illustrates an embodiment of a method for detecting an unfulfilled dependency; Fig. 6 illustrates an embodiment of a method for generating configuration data using a generative language model; Fig. 7 illustrates an exemplary implementation of a method for providing sample documents; Fig. Figure 8 illustrates an exemplary implementation of processing in one processing module from a set of processing modules; Fig. Figure 9 illustrates an exemplary embodiment of a processing chain with two branches; Fig. 10 illustrates an example implementation of a pipeline flow with stage categories and an interaction with an external intelligence; Fig. Figure 11 illustrates an exemplary embodiment of a sequential execution of a processing chain; Fig. 12 illustrates an exemplary implementation of a processing chain with simple term search with context expansion; Fig. Figure 13 illustrates an embodiment of a processing chain for locating, extending and refining pattern matching; Fig. Figure 14 illustrates an embodiment of a processing chain with LLM-based semantic filtering; and Fig. Figure 15 illustrates an exemplary embodiment of a multi-purpose computer.

[0007] Before describing exemplary embodiments using the figures, general explanations of the present technology are given.

[0008] It has been recognized that it may be desirable to provide a highly configurable framework for automated, multi-stage analysis and subsequent generation of annotations for digital (e.g., digitized) documents.

[0009] For example, an advanced analysis assistant can be provided that identifies potentially relevant information in digital documents based on configured criteria, while a final evaluation and interpretation can be left to a human user.

[0010] At the heart of the technology can be a dynamic, sequential processing architecture (pipeline) consisting of discrete, functionally specialized processing stages that operate on data segments derived from input documents. The composition of the pipeline and the behavior of the stages can be defined through declarative configuration, enabling the creation of highly customized workflows that can, for example, correspond to structured analysis models and / or audit checklists. Further configuration options can be based on Reinforcement Learning from Human Feedback (RLHF) or, more generally, on frequency statistics.

[0011] These workflows can support diverse analytical goals, such as detailed concept identification, precise pattern matching based on complex criteria (including contextual proximity, exclusion logic and / or mixed case sensitivity), context-aware semantic filtering using external intelligence models (e.g., Large Language Models (LLMs)), automated data categorization, and systematic identification of content for potential obfuscation.

[0012] The present technology can transform raw digital documents into structured, queryable annotation objects that, for example, can be strictly linked to their precise context origin and possibly associated with specific checklist items or aspects of analysis.

[0013] This allows for the support of complex analysis patterns, including parallel execution of diverging sub-pipelines and iterative refinement of data segments through chained application stages, thereby enabling sophisticated, multi-layered analysis approaches.

[0014] Accordingly, some embodiments of the present technology relate to a circuit for processing the content of a digital document, wherein the circuit is configured to: Receiving configuration data for the content preparation of a digital document, wherein the configuration data indicates a processing chain with processing modules from a predefined set of processing modules for the content preparation of a digital document; Generating a processing chain according to the configuration data; Receiving a digital document; and Applying the processing chain to the digital document to prepare the digital document.

[0015] The circuit may include a programmed microprocessor, an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or the like.

[0016] The circuit can include a processing section (e.g., a central processing unit (CPU)) that, based on software instructions, can perform data processing according to the available technology and / or control a function of the circuit. The processing section can also include a section (e.g., a graphics processing unit (GPU), tensor processing unit (TPU), photonics processor, quantum processor, etc.) for executing artificial intelligence (AI) such as neural networks, language models, etc., and / or simpler machine learning and / or natural language processing algorithms.

[0017] The circuit may further include a memory section (e.g., solid-state drive (SSD), flash memory, magnetic storage, optical storage, synchronous dynamic random access memory (SDRAM), electrically erasable programmable read-only memory (EEPROM), etc.) that can store the software instructions for the processing section and / or data required to perform or generated during data processing according to the present technology.

[0018] The circuit can also include a communication section that can provide a communication interface (e.g., Ethernet, IEEE 802.11 (Wi-Fi), InfiniBand, Universal Serial Bus (USB), Peripheral Component Interconnect (PCI), mobile telecommunications (e.g., 3G, 4G, 5G, 6G; e.g., Universal Mobile Telecommunications System (UMTS), Long Term Evolution (LTE), New Radio (NR)) etc.) for data exchange with other data processing devices (e.g., with a user terminal).

[0019] The circuitry can be provided by a general-purpose computer, as referenced in Fig. 15. The circuit can be implemented as a server that can be connected to other data processing devices (e.g., user terminal, network storage, front-end server, etc.) via a communication network (e.g., Internet, intranet, virtual private network (VPN), etc.). Alternatively, the circuit can be implemented as a user terminal (e.g., desktop computer, laptop, smartphone, tablet, etc.) and operated by a user via a human-machine interface (e.g., keyboard, mouse, screen, touchpad, gesture control, voice control, speakers, etc.).

[0020] The digital document can exist as a file and / or as a byte stream. The content of the digital document can include, for example, text data, images, and / or numbers. Examples of digital document formats include Portable Document Format (PDF), Word document (DOC, DOCX), Excel document (XLS, XLSX), PowerPoint document (PPT, PPTX), OpenDocument (ODT, ODS, ODP), text document (TXT), Markdown, Hypertext Markup Language (HTML, XHTML), Extensible Markup Language (XML), XRechnung, ZUGFeRD, JPEG File Interchange Format (JFIF), Portable Network Graphics (PNG), Tagged Image File Format (TIFF), etc.

[0021] Examples of digital document content include tenders, offers, contracts, invoices, applications, reports, expert opinions, building plans, notices, permits, laws, regulations, pleadings (e.g., legally relevant pleadings such as lawsuits, statements of defense, replies, duplicates, etc.), judgments, patent specifications, etc.

[0022] Content processing of a digital document can include analyzing and / or structuring its content, for example, based on approaches from computational linguistics and / or statistical, particularly neural, language processing. Content processing can encompass text recognition (e.g., when content of the digital document is available as image data and / or when text data is embedded in image data, as in blueprints or sketches), tokenization, morphological analysis, syntactic analysis, semantic analysis, classification, identification of captions, page headers and / or footers, table contents, tables of contents, text structuring (e.g., paragraphs, headings), content summarization (e.g., based on phrases and / or sentences from the (original) document or by independently generating a content summary based on an LLM application), etc.

[0023] As a result of the content processing, the circuit can, for example, generate an outline, a table of contents, a keyword index, a summary, a content representation via knowledge graphs, etc., and / or generate machine-readable annotations for query-based searching of the digital document.

[0024] A user can assemble a data processing system for the content preparation of a digital document from a set of predefined processing modules into a processing chain. Additionally and / or alternatively, the system can autonomously assemble an optimal processing chain (e.g., a pipeline). The processing modules can provide individual processing stages, processing steps, etc.

[0025] The predefined set of processing modules may include modules that are pre-installed on the circuit. In some implementations, a user can create (e.g., program) their own processing modules and add them to the predefined set. In other implementations, processing modules are offered on an internet platform (e.g., marketplace, app store, etc.) and / or through peer-to-peer exchange and can be downloaded from there and added to the predefined set.

[0026] The processing chain can display the execution sequence of the processing modules and / or a data category requested as a result of the content processing. The circuit can generate the processing chain according to the configuration data.

[0027] The configuration data can specify how the processing chain is to be created. For example, the configuration data can show which processing modules are included in the processing chain, the order in which the processing modules are executed, and / or which data category is requested as the result of the content preparation (e.g., as the result of the processing chain). The configuration data can also show a link between processing modules in the processing chain (e.g., passing output data from an earlier processing module as input data to a later processing module). The configuration data can also display parameters for processing modules in the processing chain, for example, to fine-tune the behavior of a processing module.

[0028] The circuit can generate the configuration data itself (e.g., based on user input and / or on recordings and / or logging of previous analysis runs), read it from the memory section, or receive it via the communication section from another data processing device (e.g., a user terminal, network storage, front-end server, etc.) or an external storage medium (e.g., USB stick, external hard drive, compact disc read-only memory (CD-ROM), floppy disk, magnetic tape, quick-response (QR) code, etc.).

[0029] The circuit can generate the processing chain according to the configuration data. For example, the circuit can generate executable instructions for processing the processing modules of the chain. The processing chain can be executed as a computer program containing instructions for each of its processing modules. Alternatively, the processing chain can be executed as a computer program (e.g., a script) that integrates the individual processing modules at runtime and / or calls them as subprocesses. Finally, the processing chain can be implemented as a list (e.g., XML, JSON, YAML, TOML, database, etc.) of the processing modules, which a processing engine can read and execute accordingly.

[0030] The processing chain and / or the individual processing modules can be implemented, for example, as machine code (e.g., Executable and Linkable Format (ELF), Portable Executable (PE), Common Object File Format (COFF), Mach Object File Format (Mach-O), etc.), as bytecode (e.g., for a Java runtime environment, for a .NET runtime environment, etc.), as LLVM Intermediate Representation (LLVM IR), as WebAssembly code, and / or as a program library (e.g., Shared Object (.so), Dynamic Link Library (DLL), JavaScript module, Python module, Perl module, etc.). A person skilled in the art may discover other formats for the processing chain and / or the individual processing modules.

[0031] The processing chain and / or processing modules may, for example, include a call to an artificial neural network (e.g., language model, Large Language Model (LLM)).

[0032] For example, the processing chain and / or processing modules can display a prompt for the artificial neural network, a configuration (e.g., architecture, network weights, etc.) of the artificial neural network, an endpoint (e.g., Uniform Resource Locator (URL)) for calling the artificial neural network via an application programming interface (API), a file path of the artificial neural network, etc. The artificial neural network can be configured to process at least part of the digital document.

[0033] The circuit can receive the digital document, for example, through user input, from an external storage medium (e.g., USB flash drive, external hard drive, CD-ROM, floppy disk, magnetic tape, QR code, etc.), and / or through a network request (e.g., Hypertext Transfer Protocol (HTTP), Representational State Transfer (REST), etc.). For example, the user can transfer the digital document to the circuit via a website, an HTTP application programming interface (API), and / or remote access (e.g., via Secure Shell (SSH), Remote Desktop Protocol (RDP), Independent Computing Architecture (ICA), TeamViewer, Telnet, File Transfer Protocol (FTP), etc.).

[0034] The circuit can apply the processing chain to the digital document by passing the digital document, in whole or in part, as input data to the processing chain and executing the processing chain.

[0035] The circuit receives a collection of digital documents and performs batch processing, for example applying the processing chain to each digital document in the collection.

[0036] In some implementation examples, obtaining the configuration data includes: Displaying a graphical user interface in which processing modules from the set of processing modules are represented as objects; where objects for processing modules of the processing chain are displayed in a predefined section.

[0037] The circuit can output the graphical user interface on a screen (e.g. via High Definition Multimedia Interface (HDMI), DisplayPort, Thunderbolt, Lightning, Digital Visual Interface (DVI), USB, Video Graphics Array (VGA), Miracast, AirPlay, RDP, ICA, TeamViewer, etc.) and / or provide it as a website.

[0038] The objects representing processing modules can be displayed as rectangles, circles, images, icons, text, etc. Each object can be associated with a processing module. The objects can display a name, description, origin (e.g., pre-installed, user-created, downloaded from a marketplace, developer's name / company / address, etc.), dependencies (e.g., output data from another processing module), etc., of the processing module.

[0039] The predefined section can be a section of the user interface that corresponds to the processing chain. The predefined section can be identified as a section associated with the processing chain, for example, by a heading (e.g., "Processing Chain," "Workflow," "Pipeline," etc., and / or a user-defined name for the processing chain), by a border, and / or by a background design (e.g., color, pattern, icons, etc.). Displaying an object within the predefined section can indicate that a processing module associated with that object is included in the processing chain.

[0040] The graphical user interface can also allow for the parameterization of processing modules (processing steps). For example, a processing module can accept an input parameter and execute predefined processing steps based on the value of that input parameter. The graphical user interface can display whether a processing module supports an input parameter, the function of the input parameter (e.g., which processing steps can be controlled by the input parameter), the values ​​the input parameter can take, and / or the currently assigned value. The graphical user interface can also provide an input field for entering a different value for the input parameter.

[0041] In some implementation examples, obtaining the configuration data also includes: Receiving an initial user operation indicating a dragging of one of the objects from outside the predefined section into the predefined section; and Add a processing module, corresponding to the object pulled with the first user operation, to the processing chain in the configuration data.

[0042] The user can perform the first user operation with a mouse by, for example, clicking the object, moving the mouse pointer to the predefined section while holding down the mouse button, and then releasing the mouse button. The user can also perform the first user operation with a finger or stylus on a touchscreen or touchpad by placing the finger or stylus on the touchscreen or touchpad at a position corresponding to the object, moving it to a position on the touchscreen or touchpad corresponding to the predefined section, and then lifting it from the touchscreen or touchpad. The user may discover other ways to perform the first user operation, such as voice control and / or gesture control.

[0043] The first user action can indicate that the user wants to add the processing module corresponding to the object dragged into the predefined section by the first user operation to the processing chain. Accordingly, when the circuit receives the first user operation, it can create an entry in the configuration data that displays the processing module corresponding to the object dragged into the predefined section by the first user operation.

[0044] Based on a position in the predefined section to which the user drags the object (e.g., which processing modules neighboring objects of the object dragged with the first user action correspond to), the circuit can determine at which position in the processing chain the processing module should be inserted, and can insert this position into the entry in the configuration data.

[0045] In some implementation examples, obtaining the configuration data also includes: Receiving a second user operation that indicates dragging one of the objects from a first position in the predefined section to a second position in the predefined section; and Moving a processing module of the processing chain, corresponding to the object pulled by the second user operation, to a position in the processing chain corresponding to the second position in the configuration data.

[0046] The user can perform the second user operation in a similar way to the first. The user can drag the same object in the second user operation as in the first (e.g., to adjust the position of the corresponding processing module in the processing chain). Alternatively, the user can drag a different object in the second user operation than the one in the first.

[0047] The second user operation can indicate that the user wants to adjust the position of a processing module in the processing chain that corresponds to the object selected with the second user operation. Accordingly, when the circuit receives the second user operation, it can modify an entry in the configuration data for the processing module corresponding to the object selected with the second user operation to display a position in the processing chain that corresponds to the second position. For example, the circuit can cut the entry from a position in the configuration data corresponding to the first position and insert it at a position corresponding to the second position.For example, the circuit can change a reference in the entry to a position of the corresponding processing module in the processing chain, for example by changing an index, a dependency, a position identifier, etc.

[0048] In some embodiments, the circuit is configured to: Detecting a dependency of one processing module displayed by the configuration data on output data of another processing module from the set of processing modules; Detecting the absence of the other processing module in the processing chain; and issuing a notification of an unfulfilled dependency.

[0049] A dependency between a (dependent) processing module and another processing module can exist if the dependent processing module expects and / or requires output data from the other processing module as input data. If the other processing module is missing from the processing chain or is positioned in the processing chain in such a way (e.g., at a later position or in a different branch of the processing chain) that output data from the other processing module is not available when the dependent processing module is executed, then in some cases the dependent processing module or possibly the entire processing chain may not execute correctly and / or may produce an erroneous result (e.g., inaccurate, incorrect, imprecise, etc.).

[0050] The circuit can therefore determine, when receiving configuration data, creating the processing chain, and / or applying the processing chain to the digital document, whether the output data of the other processing module is available when the dependent processing module is executed. Detecting the absence of the other processing module in the processing chain can include determining that the output data of the other processing module is not available when the dependent processing module is executed.

[0051] The unmet dependency warning can indicate the dependent and / or other processing module. By displaying this warning, a user can adjust the configuration data to fulfill the dependency (e.g., ensuring that the output data of the other processing module is available when the dependent processing module is executed). This prevents errors when applying the processing chain to the digital document.

[0052] In some implementation examples, obtaining the configuration data includes: Receiving an instruction from a user to generate the configuration data; and Executing a generative language model with the instruction as input data; where the generative language model is set up to generate the configuration data based on the instruction.

[0053] The user can enter the instruction to generate the configuration data via text input (e.g., prompt) and / or voice input. The instruction can display the digital document, content within the digital document, a category within the digital document, a desired format for the digital document, contextual information, etc. The instruction can include additional contextual information, thus enabling, for example, one-shot or few-shot learning.

[0054] The generative language model can be implemented as an artificial neural network (e.g., as a large language model (LLM)). An architecture of the generative language model can be based on, for example, transformers, encoder-decoders, encoder-only, decoder-only, mixture-of-experts (MoE), or similar structures, and can be embedded in a larger domain-specific architecture such as retrieval-augmented generation (RAG) or an agent structure. Examples of generative language models include Generative Pretrained Transformer (GPT), ChatGPT, Gemini, Grok, Claude, Mistral, Llama, etc.

[0055] The circuit can execute the generative language model on the section for running AI (e.g., GPU, TPU, etc.) and can pass the instruction or part of the instruction as input data to the generative language model.

[0056] The generative language model can be set up (e.g., via a system prompt, via training a predefined scheme, etc.) to generate and / or output the configuration data in a format that the circuit can further process (e.g., one of the formats described above).

[0057] The generative language model can be trained with training data. The training data can represent a variety of instructions. For training, the generative language model can be executed with instructions from the training data as input. One or more parameters (e.g., network weights) of the generative language model can be adjusted based on an output of the generative language model, for example, based on a user's evaluation of the output and / or based on a comparison with predefined configuration data, which are linked to the instructions in the training data as ground truth. Training can be performed, for example, until the improvement of the generative language model remains below a predefined threshold. The generative language model can also be based on fine-tuning, for example, by refining an existing (e.g.,trained) Foundation model based on data from the application domain (e.g., configuration data).

[0058] In some implementation examples, the generative language model is further configured to: Determining a characteristic for the processing chain according to the instruction; and Generating the configuration data according to the detected characteristic for the processing chain.

[0059] The feature can be any characteristic of the processing chain that can be displayed in the configuration data. Examples of the feature include a specific processing module, a sequence of processing modules, and a parameter for a processing module. A person skilled in the art can identify other suitable features for the processing chain.

[0060] The generative language model can select processing modules from the predefined set of processing modules for the processing chain, arrange them in the processing chain, and / or fine-tune them via parameters. The generative language model can display the selected processing modules, their arrangement in the processing chain, and / or the parameters in the configuration data.

[0061] In some embodiments, determining the characteristic for the processing chain includes: Determining a processing criterion according to the instructions; and Determining the characteristic for the processing chain to generate output data based on the digital document that meets the processing criterion.

[0062] The processing criterion can indicate a desired outcome of the processing chain. For example, the processing criterion can indicate that the processing chain should output an outline, a table of contents, and / or a summary of the digital document's content. Other examples of processing criteria include the content structure, language, vocabulary, subject area, category, origin (e.g., author), target audience, format, etc., of the digital document.

[0063] The user can explicitly define the processing criterion in the instruction. The user can specify a purpose for the content processing and / or an intended use of a result from the processing chain, and the generative language model can derive the processing criterion from the specified purpose or use. If the user provides the digital document without further explanation, the generative language model can determine the processing criterion based on the digital document itself, for example, based on a processing criterion that corresponds to training data of the generative language model for a similar digital document.

[0064] The generative language model can determine (e.g., based on the training data of the generative language model) which feature or features of the processing chain are needed or can be used to determine the preparation criterion, and can accordingly determine the feature for the processing chain.

[0065] In some embodiments, the circuit is further configured to: Providing a set of sample documents; Applying the processing chain to the set of sample documents; Outputting a reference to a result of the processing chain applied to the set of sample documents.

[0066] The circuit can provide the set of sample documents to test a processing chain, for example, when a user and / or a generative language model has generated configuration data.

[0067] The circuit can load the set of sample documents from the memory section and apply the processing chain to digital documents from the sample document set, e.g., to all sample documents from the sample document set, to sample documents selected by a user from the sample document set, and / or to sample documents from the sample document set that are associated with features of the processing chain. For example, the sample document set may contain a predefined sample document for a feature (e.g., a specific processing module, a specific sequence of processing modules, a specific parameter value of a processing module, etc.) of the processing chain and / or for a processing criterion, which, due to its content, may be well-suited for testing the feature and / or processing criterion.

[0068] The result of the processing chain applied to the set of sample documents can display all or part of the result for each sample document to which the processing chain was applied. The result can include data that corresponds to a processing criterion. The result can include statistical data (e.g., execution time per document, scope of the result, etc.). The result can include an assessment of the result's accuracy generated by an AI (e.g., the generative language model) and / or a mathematical calculation rule (e.g., a parameterized cost function), and / or a suggestion for improving the configuration data. The result can also indicate an error that occurred when applying the processing chain to the set of sample documents.

[0069] In some implementation examples, the configuration data is represented in a structured markup language.

[0070] The configuration data can be formatted as XML, JSON, YAML, TOML, INI, or similar. This allows the configuration data to be easily ported to other circuits.

[0071] The configuration data can contain the processing modules (e.g., as text, Base64, hexadecimal representation, etc.). The configuration data can specify an identifier (e.g., Globally Unique Identifier (GUID), Uniform Resource Identifier (URI), checksum (e.g., SHA-1, SHA-256, MD5, etc.)) of a processing module. The configuration data can specify a storage location (e.g., file path and / or web resource) from which a processing module can be loaded.

[0072] In some embodiments, generating the processing chain includes: Loading data processing instructions that correspond to processing modules of the processing chain; and Generating instructions to execute the loaded data processing instructions according to the configuration data.

[0073] Loading the data processing instructions may involve reading the data processing instructions from non-volatile memory, an external storage device and / or a network address, and writing the data processing instructions to main memory and / or CPU cache.

[0074] Generating the instructions for executing the loaded data processing instructions can involve compiling (e.g., translating into machine code) and / or generating (e.g., as a script in a scripting language such as Python, JavaScript, Perl, Lua, Bash, etc.) a program that corresponds to the processing chain. Generating the instructions for the loaded data processing instructions can also involve generating a list and / or one or more database entries, which are read by a program to execute the processing chain.

[0075] In some implementation examples, applying the processing chain to the digital document includes: Generate, according to the processing chain, an annotation object that displays output data of a processing module in the processing chain.

[0076] The annotation object can be based on a processing criterion. For example, the annotation object can display or contain output data that meets a processing criterion.

[0077] The annotation object is not limited to output data from a single processing module. For example, the annotation object can display output data from multiple processing modules in the processing chain and / or a result of the processing chain.

[0078] The circuit can insert the annotation object into the digital document or into a copy of the digital document and / or write it to a separate file or database and link it to the digital document.

[0079] Based on the annotation object, the circuit can display the digitally prepared document to the user.

[0080] In some implementation examples, the annotation object displays a section of the digital document to which the output data of the processing module refers.

[0081] The annotation object can thus be linked to the section of the digital document.

[0082] For example, the circuit can display the annotation object along with the section of the digital document.

[0083] For example, the annotation object can be indexed along with other annotation objects in a queryable index, and the circuit can display the section of the digital document if the annotation object matches a result of a query of the index.

[0084] In some implementation examples, the annotation object indicates that the processing module has not detected any data that meets a processing criterion.

[0085] For some processing criteria, it may be relevant whether the corresponding data is contained in the digital document. The annotation object can display a negative result to explicitly indicate that the digital document does not contain any data that meets the processing criterion.

[0086] In some implementation examples, applying the processing chain to the digital document includes: Receiving output data from a first processing module in the processing chain; providing the output data as input data for a second processing module in the processing chain according to a dependency of the second processing module indicated by the configuration data; and Executing the second processing module.

[0087] The second processing module may expect or require the output data of the first processing module as input data and may therefore be dependent on the first processing module.

[0088] The circuit can therefore first execute the first processing module to obtain its output data. Afterwards, the circuit can pass this output data as input to the second processing module and execute it.

[0089] In this way, the circuit can resolve a dependency between processing modules.

[0090] In some cases, the first processing module outputs all or part of the output data while the first processing module is still running, and / or the second processing module can still receive the input data while the second processing module is still running. In such cases, the circuit can begin executing the second processing module before the first processing module has finished.

[0091] In some embodiments, the circuit is further configured to: Authorizing a user connected via a user terminal device; Receipt of the digital document from the user; Receiving an instruction from the user indicating the processing chain; and Applying the processing chain to the digital document according to the user's instructions.

[0092] For example, the circuit can provide a web interface that the user can access via their end device (e.g., via a network connection such as the internet or an intranet). The end device can include a desktop computer, laptop, smartphone, tablet, or similar device.

[0093] The user can log in to the web interface using credentials such as a username, password, session cookie, smartcard, API token, or similar. Based on these credentials, the system can authorize the user to access the web interface (e.g., creating configuration data, providing the digital document, and / or applying the processing chain to the digital document).

[0094] The user can transmit the digital document and instructions to the circuit using their user device.

[0095] The instruction allows the user to specify that the processing chain should be applied to the digital document. For example, different processing chains can be configured on the system, and the system can offer the user (e.g., in the web interface) a list of available processing chains to choose from. The user can then use the instruction to specify which processing chain from the list should be applied to the digital document.

[0096] The circuit can then apply the processing chain to the digital document according to the instructions. For example, the circuit can display a result of the processing chain to the user on the web interface.

[0097] In some embodiments, a processing module from the set of processing modules is configured for: Receiving input data; Performing predefined data processing on the input data, wherein the predefined data processing is associated with the processing module; and Generating output data that displays a result of the predefined data processing.

[0098] The processing module can be executed, for example, as a program function, a script, a macro, a program, or the like.

[0099] In some embodiments, at least one processing module from the set of processing modules is configured to include at least one segmenting, filtering, expanding, merging, sorting, limiting and combining.

[0100] A person skilled in the art can find further examples of operations that the processing module can perform. For example, the processing module can perform any suitable operation or combination of operations from the fields of computational linguistics and / or information retrieval.

[0101] In some implementation examples, the configuration data displays: Splitting the processing chain into at least two branches; and Combining at least two branches with a processing module from the set of processing modules that is designed to combine branches.

[0102] The circuit can execute the branches independently of each other (e.g., concurrently, in parallel, etc.). The branches can process the digital document according to different processing criteria and / or process different sections of the digital document.

[0103] The processing module set up to combine branches can receive output data from processing modules on the branches and combine it appropriately.

[0104] Some exemplary implementations relate to a method for the content processing of a digital document, whereby the method includes: Receiving configuration data for the content preparation of a digital document, wherein the configuration data indicates a processing chain with processing modules from a predefined set of processing modules for the content preparation of a digital document; Generating a processing chain according to the configuration data; Receiving a digital document; and Applying the processing chain to the digital document to prepare the digital document.

[0105] The method can be implemented by the circuit according to any of the embodiments described above. For each embodiment of the circuit described above, there exists a corresponding embodiment of the method.

[0106] Some embodiments involve a computer program comprising instructions that, when the program is executed by a computer, cause it to execute the method according to one of the embodiments described above.

[0107] The computer can, for example, be implemented by the circuit according to one of the embodiments described above.

[0108] Returning to Fig. 1, illustrated Fig. Figure 1 shows an embodiment of a circuit 1 for processing the content of a digital document. The circuit 1 comprises a processing section 2, a storage section 3, and a communication section 4.

[0109] Processing section 2 performs data processing according to the present technology based on software instructions and controls a function of circuit 1. Processing section 2 also includes an AI section (not shown) for executing AI.

[0110] Storage section 3 stores the software instructions for processing section 2, as well as data required for or generated during the execution of data processing according to the present technology.

[0111] Communication section 4 provides a communication interface for data exchange with other data processing devices. Furthermore, communication section 4 provides an interface for communication with a user of circuit 1 (user input of data and instructions, output of information to the user, generation of a graphical user interface).

[0112] Circuit 1 is configured to perform procedures according to the exemplary embodiments of the Fig. 2 to 8, and thereby, for example, processing chains according to the exemplary embodiments of the Fig. to generate and apply 9 to 14.

[0113] Fig. Figure 2 illustrates an embodiment of a method 10 for the content processing of a digital document. Method 10 is an example of a method that is executed by circuit 1.

[0114] At point 11, circuit 1 receives configuration data for the content processing of a digital document. This configuration data specifies a processing chain with processing modules from a predefined set of modules for the content processing of a digital document. The configuration data is represented in a structured markup language.

[0115] At point 12, circuit 1 generates a processing chain according to the configuration data. To generate the processing chain at point 12, circuit 1 loads data processing instructions at point 13 that correspond to the processing modules of the processing chain, and at point 14 generates instructions to execute the loaded data processing instructions according to the configuration data.

[0116] At 15, circuit 1 authorizes a user connected via a user terminal device.

[0117] At 16, circuit 1 receives a digital document from the user.

[0118] At 17, circuit 1 receives an instruction from the user, which indicates the processing chain.

[0119] At 18, circuit 1 applies the processing chain according to the user's instructions to the digital document to process the digital document.

[0120] When the processing chain is applied to the digital document at 18, circuit 1 at 19 receives output data from a first processing module of the processing chain, provides the output data at 20 as input data for a second processing module of the processing chain according to a dependency of the second processing module shown by the configuration data, and executes the second processing module at 21.

[0121] When the processing chain is applied to the digital document at point 18, circuit 1 also generates an annotation object at point 22, according to the processing chain. This annotation object displays output data from a processing module within the chain. The annotation object indicates a section of the digital document to which the processing module's output data refers. In some cases, the annotation object indicates that the processing module has not detected any data that meets a processing criterion.

[0122] In some embodiments (for example, in the local execution of procedure 10), user authorization is omitted at 15.

[0123] In some embodiments, circuit 1 receives the instruction at 17 before the digital document at 16. In some embodiments, circuit 1 receives the digital document at 16 before receiving the configuration data at 11 and / or before generating the processing chain at 12. The person skilled in the art may find further suitable deviations from the sequence of method 10 shown here.

[0124] In some implementation examples, the configuration data is represented in a way other than in a structured markup language.

[0125] Fig. Figure 3 illustrates an embodiment of a method 30 with a first user operation in a graphical user interface 40. The method 30 is an example of obtaining the configuration data at 11 in Fig. 2.

[0126] At 31, circuit 1 displays a graphical user interface 40. The graphical user interface 40 contains predefined sections 41, 42, and 43. In the graphical user interface 40, processing modules from the set of processing modules are represented as objects 42a to 42d and 43a to 43e.

[0127] Section 41 contains buttons for saving the processing chain and for exiting the graphical user interface 40. Section 42 displays available processing modules of the processing module set as a list of objects 42a to 42d. Section 43 represents the processing chain. Objects 43a to 43e for processing modules of the processing chain are displayed in Section 43.

[0128] At 32, circuit 1 receives a first user operation, indicating the dragging of one of the objects 42b from outside section 43 into section 43. A graphical user interface 40a corresponds to the graphical user interface 40 during the first user operation. The object 42b is dragged into section 43 by the user with a mouse pointer.

[0129] At step 33, circuit 1 adds a processing module corresponding to object 42b, which was selected during the first user operation, to the processing chain in the configuration data. A graphical user interface 40b corresponds to the graphical user interface 40 after the first user operation and after adding the processing module corresponding to object 42b to the processing chain. Object 42b is displayed in section 43. The subsequent objects 43c to 43e are shifted downwards in the graphical user interface 40b, so that object 43e is moved out of a displayed section of the processing chain. Furthermore, object 42b remains displayed in section 42 so that the user can insert the processing module corresponding to object 42b at another position in the processing chain (e.g., with another user operation corresponding to the first).

[0130] Fig. Figure 4 illustrates an embodiment of a method 50 with a second user operation in a graphical user interface 60. The method 50 is an example of obtaining the configuration data at 11 in Fig. 2.

[0131] At 51, circuit 1 displays a graphical user interface 60. The graphical user interface 60 is set up similarly to the graphical user interface 40 from Fig. 3 and contains predefined sections 61, 62 and 63. In the graphical user interface 60, processing modules from the set of processing modules are represented as objects 62a to 62d and 63a to 63e.

[0132] Section 61 contains buttons for saving the processing chain and for exiting the graphical user interface 60. Section 62 displays available processing modules of the processing module set as a list of objects 62a to 62d. Section 63 represents the processing chain. Objects 63a to 63e for processing modules of the processing chain are displayed in Section 63.

[0133] At 52, circuit 1 receives a second user operation, which indicates the dragging of object 63d from a first position in section 63 to a second position in section 63. A graphical user interface 60a corresponds to the graphical user interface 60 during the second user operation. The user drags object 63d with a mouse pointer from a first position (between objects 63c and 63e) to a second position (between objects 63b and 63c).

[0134] At 53, circuit 1 moves the processing module of the processing chain, which corresponds to the object 63d drawn with the second user operation, to a position in the processing chain corresponding to the second position in the configuration data.

[0135] Fig. Figure 5 illustrates an embodiment of method 70 for detecting an unfulfilled dependency. Method 70 is an example of processing during the acquisition of configuration data at 11, the generation of processing chain 12, and / or the application of the processing chain at 18. Fig. 2.

[0136] At 71, circuit 1 detects a dependency of one processing module displayed by the configuration data on output data of another processing module from the set of processing modules.

[0137] At 72, circuit 1 detects the absence of the other processing module in the processing chain.

[0138] At 73, circuit 1 indicates an unfulfilled dependency in the processing chain.

[0139] Fig. Figure 6 illustrates an embodiment of method 80 for generating configuration data using a generative language model. Method 80 is an example of obtaining the configuration data at Figure 11 in Fig. 2.

[0140] At 81, circuit 1 receives an instruction from a user to generate the configuration data.

[0141] At 82, circuit 1 executes a generative language model 83 with the instruction as input data.

[0142] At 84, the generative language model 83 determines a feature for the processing chain according to the instruction. At 85, the generative language model also determines a processing criterion according to the instruction and determines the feature for the processing chain to generate output data based on the digital document that meets the processing criterion.

[0143] In 86, the generative language model 83 generates the configuration data based on the instruction and according to the recognized feature for the processing chain.

[0144] Fig. Figure 7 illustrates an embodiment of method 90 for providing sample documents. Circuit 1 can implement method 90 as part of method 10. Fig. Execute 2.

[0145] At 91, circuit 1 provides a set of sample documents.

[0146] At 92, circuit 1 applies the processing chain to the set of sample documents.

[0147] At 93, circuit 1 outputs a hint about a result of the processing chain applied to the set of sample documents.

[0148] Fig. Figure 8 illustrates an embodiment of a processing operation 100 in a processing module from a set of processing modules. The processing operation 100 is an example of a processing operation that is a processing module of the procedures in the Fig. 2 to 7.

[0149] At 101, the processing module receives input data.

[0150] At 102, the processing module performs predefined data processing on the input data. This predefined data processing is associated with the processing module. The associated predefined data processing of the processing module corresponds to at least one of the following: segmenting, filtering, expanding, merging, sorting, limiting, and combining. A person skilled in the art can find further examples (e.g., categories) of the associated predefined data processing of the processing module.

[0151] At 103, the processing module generates output data that displays a result of the predefined data processing.

[0152] Fig. Figure 9 illustrates an embodiment of a processing chain 110 with two branches. The processing chain 110 is an example of a processing chain that can be implemented according to the methods described in the Fig. 2 to 8 of the configuration data are displayed, generated by circuit 1 and / or applied by circuit 1 to the digital document. Arrows point in Fig. 9 a processing sequence of processing modules of the processing chain 110.

[0153] Configuration data for processing chain 110 indicates that processing chain 110 begins at processing module 111. At processing module 112, the configuration data indicates that processing chain 110 splits into at least two branches. The first branch comprises processing modules 113a to 113d. The second branch comprises processing modules 114a to 114c. The two branches are combined using a processing module 115 from the set of processing modules, which is configured to combine branches.

[0154] The circuit described herein can, for example, provide a system for the content processing of a digital document. The system can integrate functionally different, interacting subsystems. These subsystems can interact, for example, by one subsystem receiving and further processing output data from another subsystem as input data.

[0155] A document input and canonical representation subsystem can transform input documents into a complex internal data structure that preserves text, character offsets, geometric coordinates, page identifiers, font attributes, and structural classifications to enable context-aware processing. This subsystem can also perform any necessary normalizations.

[0156] A workflow configuration and interpretation subsystem can process structured, declarative configurations that describe ordered stages, parameters, nested pipelines, and global directives. This subsystem can also validate a configuration and translate it into an executable plan. This configuration can represent an executable specification of verification or analysis requirements, possibly derived from a previous analysis of workflows and data structures. This subsystem can also manage versioning.

[0157] The system can include a pipeline execution and data flow management engine. This engine can orchestrate the dynamic execution of configured stages and manage segmented flow, including parallel sub-pipelines, transfer context, apply parallelization, and / or handle errors.

[0158] The system can include an extensible library of processing steps. The library can provide a source of modular, reusable processing units (e.g., processing modules, stages) for segmenting, filtering, extending, merging, sorting, limiting, and / or combining execution paths. The library can be designed to accommodate future extensions.

[0159] The system can include an engine for advanced pattern specification, compilation, and matching. The engine can define, compile, and / or execute complex pattern searches. It can analyze specifications using nested Boolean logic (AND, OR, NOT / NOR), wildcards, phrase matching, mixed case (per term), and / or direct regular expressions, often based on sample text data collected during requirements analysis. The engine can perform normalization, compile into optimized representations (e.g., trees, automata), execute efficiently, resolve ambiguities, and / or utilize caching. Pattern matching results can include precise matching regions.

[0160] The system may include a subsystem for interaction with external intelligence. This subsystem can manage communication with external models (e.g., Large Language Models (LLMs)) for stages that require semantic understanding or generation, and / or can format inputs based on stage parameters, process model outputs, and / or translate them into actions.

[0161] The system can include a module for result synthesis and annotation generation. This module can process a final set of segments, perform configured aggregation and consolidation, synthesize annotation objects with context-rich text representations (including markup) and / or precise links (e.g., ID, area, geometry) to the source, manage the generation of default annotations for null results (e.g., to handle "negative" checklist items where the absence is significant), and apply global boundaries. Generated annotations can be linked to the specific configuration or checklist item that generated them.

[0162] The system can include an integrated configuration environment that provides a dedicated user interface for creating, editing, managing, and testing workflow configurations. The configuration environment can support structured editing that reflects the declarative configuration language (e.g., indentation for hierarchies, specific syntax for stage calls). The configuration environment can further support commenting for documentation, temporarily disabling stages or configuration blocks for isolated testing, and / or running a configured pipeline against selected test document sets directly within the environment. The configuration environment can require explicit actions to save configuration changes.

[0163] Fig. Figure 10 illustrates an example implementation of a pipeline flow with stage categories and an interaction with an external intelligence.

[0164] A configurable analysis pipeline 120 is based on a workflow configuration 121 and a configuration interpretation 122 of the workflow configuration 121. The workflow configuration 121 is an example of configuration data according to the present disclosure, e.g., for the configuration data that the circuit 1 at 11 in Fig. 2 receives. Configuration interpretation 122 is an example of generating a processing chain according to the present disclosure, e.g., generating a processing chain at 12 in Fig. 2. Analysis pipeline 120 is an example of a processing chain according to the present disclosure, e.g., the processing chain that connects circuit 1 at 12 in Fig. 2 generated.

[0165] The analysis pipeline 120 further receives a representation 123 of an input document. The input document is an example of a digital document according to the present disclosure, and the receipt of the representation 123 is an example of receiving a digital document as described in 16. Fig. 2.

[0166] At 124, a pipeline execution of the analysis pipeline 120 is started.

[0167] In the analysis pipeline 120, segmentation stages 125, filter stages 126, extension stages 127, merging stages 128, sorting stages 129 and limiting stages 130 are executed until a final segment set 131 with segments of the input document 123 is available.

[0168] Stages 125 to 130 are read from a stage library 132. Stages 125 to 130 are examples of processing modules according to the present disclosure. The stage library 132 is an example of a memory in which the set of processing modules according to the present disclosure is stored.

[0169] One stage of the segmentation stage 125 and one stage of the filter stage 126 call a pattern engine 133 for pattern matching. One stage of the segmentation stage 125 and one stage of the filter stage 126 call an external intelligence 134. Calling the pattern engine 133 and the external intelligence 134 is optional, as indicated by the dashed arrows, and is not performed in some embodiments.

[0170] At 135, a result synthesis and annotation generation take place based on a result from the analysis pipeline 120. The annotations are examples of the annotation object according to the present disclosure, e.g., for the annotation object that defines circuit 1 at 22 in Fig. 2 generated.

[0171] The annotations are displayed at 136.

[0172] It should be noted that in some embodiments, the order of stages 125 to 130 is reversed and / or stages of different categories can be executed in a mixed order. In some embodiments, one or more of the categories of stages 125 to 130 are omitted and / or the analysis pipeline 120 includes a processing module (e.g., a stage) of a category not mentioned here.

[0173] The system's input data structures can include a canonical representation of an input document. The input document is an example of the digital document as described in this disclosure. As described above, the canonical representation of the input document can include an identification identifier (ID), metadata, and / or a text stream with annotated units. The input document can also include a (technical) drawing, graphic, blueprint, etc., so that an input stream can contain the text stream in addition to graphic elements, allowing for the annotation of text elements (which can be in any orientation within the input document). The text stream with annotated units can display an offset, geometry, page ID, range, font, boundaries, and / or hash.The canonical representation of the input document may include representative test documents that can be used to validate the configuration.

[0174] The system's input data structures can include a processing workflow configuration structure. As previously described, this can be a structured format with an ordered list of stages, including a stage type and parameters derived from analysis requirements and sample data, global directives, a result aggregation strategy, and / or annotation control mechanisms, including comments on null results. Configurations can represent hierarchical checklist structures.

[0175] The analysis pipeline 120 can transform sets of processing segments.

[0176] Processing modules (e.g., stages) can be chained sequentially to implement refinement strategies.

[0177] A pipeline execution logic can execute stages sequentially according to a definition.

[0178] Fig. Figure 11 illustrates an exemplary embodiment for a sequential execution of a processing chain 140 (e.g. the analysis pipeline 120).

[0179] Upon starting the processing chain 140, the processing chain 140 receives an input segment set 141 containing segments from the digital document (e.g., according to representation 123). Logic 142 of level 1, logic 143 of level 2, and logic 144 of level n are sequentially applied to the input segment set 141. Levels 1, 2, and n are examples of processing modules according to this disclosure. After executing levels 1 through n, the processing chain 140 outputs a result segment set 145.

[0180] Combining stages can execute sub-pipelines and aggregate their results. An engine that provides the system interprets the structured configuration and processes it, for example, from top to bottom.

[0181] The use of boundaries and context can be structured as described above. Interaction with external intelligence can be structured as described above. Fig. 10 described.

[0182] The system may include the following categories of processing modules (e.g., stages), the detailed functionality of which is briefly explained below.

[0183] Segmentation levels can create new segments (e.g., drawing window, pattern-based, structural, semantic, editorially targeted, etc.). The size of the drawing window can be configurable and influence the contextual scope for subsequent pattern matching.

[0184] Filter levels can selectively allow segments to pass through (e.g., pattern-based, position-based, statistically relevant, semantic, etc.). Pattern filters can be crucial for refining results after context expansion.

[0185] Expansion levels can broaden segment boundaries (e.g., unit-based such as word or sentence expansion, fixed character count, minimum size). Word expansion can ensure that complete words are captured. Sentence expansion can provide necessary context for assessment.

[0186] Merging stages can combine adjacent and / or overlapping segments (e.g., overlap / proximity-based merging). Configurable tolerances can allow the merging of closely related but non-overlapping results.

[0187] Sorting levels can rearrange segments (e.g., position-based sorting) and thus ensure a predictable output order, for example, by document and position.

[0188] Limiting levels can reduce the number and / or size of segments (e.g., number per document, global cumulative length, etc.).

[0189] Combination stages can execute multiple workflows (e.g., sub-pipeline combination) and, for example, allow parallel analyses with different strategies or parameters.

[0190] Fig. Figure 12 illustrates an embodiment of a processing chain 150 with simple term search and context expansion. The processing chain 150 is an example of a processing chain according to the present disclosure, e.g., for the analysis pipeline 120.

[0191] Analysis pipeline 150 comprises processing modules 151 to 156, which are executed sequentially.

[0192] Processing module 151 includes a character window segmenter. Processing module 152 includes a word boundary extender. Processing module 153 includes a pattern-based segmenter that applies a simple OR pattern. Processing module 154 includes a sentence boundary extender that also considers context. Processing module 155 includes a position-based sorter. Processing module 156 includes an overlap merge stage with proximity tolerance.

[0193] Fig. Figure 13 illustrates an embodiment of a processing chain 160 for locating, extending, and refining pattern matching. The processing chain 160 is an example of a processing chain according to the present disclosure, e.g., for the analysis pipeline 120.

[0194] Analysis pipeline 160 comprises processing modules 161 to 166, which are executed sequentially.

[0195] Processing module 161 includes a character window segmenter. Processing module 162 includes a word boundary expander. Processing module 163 includes a pattern-based segmenter that creates wider segments and is based on a complex AND / OR pattern. Processing module 164 includes a sentence boundary expander. Processing module 165 includes a pattern-based filter that is more restrictive (stricter) and applies a refined pattern. Processing module 166 includes a position-based sorter.

[0196] Fig. Figure 14 illustrates an embodiment of a processing chain 170 with LLM-based semantic filtering. The processing chain 170 is an example of a processing chain according to the present disclosure, e.g., for the analysis pipeline 120.

[0197] Analysis pipeline 170 comprises processing modules 171 to 175, which are executed sequentially.

[0198] Processing module 171 includes a structural segmenter that segments, for example, by paragraph. Processing module 172 includes a semantic filter stage that uses a Large Language Model (LLM) 172a with a detailed, configurable prompt (including possible context information) and multiple passes. Processing module 173 includes a pattern-based segmenter that generates wider segments and is based on a complex AND / OR pattern. Processing module 174 includes a position-based sorter. Processing chain 170 includes optional additional stages 175.

[0199] The system can include a pattern specification and matching system. During parsing and compilation, the pattern specification and matching system can handle nested logic, mixed case, and pattern derivation from exemplary phrases, as described previously. Normalization can ignore certain characters, such as hyphens, during matching. An internal representation can have a tree structure, as described previously. Matching can be performed, as described previously, taking into account logic, sensitivity, and normalization rules. The pattern specification and matching system can be designed to find all specified variations, with an emphasis on recognition (e.g., avoiding false negatives), while subsequent stages or user review can manage precision (false positives).

[0200] During result generation and output, the system can perform contextual text synthesis, as described above. Result objects can be finalized as described above. Output objects can be designed to be clearly presented alongside source documents, e.g., linked to relevant checklist items and / or configuration locations. The system can apply proximity-based consolidation logic, as described above. The system can also perform metadata mapping as previously described, including optional comments for cases where expected patterns are not found; for example, the order of 172 and 173 in Fig. 14 will be swapped.

[0201] A configuration methodology and / or lifecycle within the system can be designed as follows. To process declarative specifications, the system can directly analyze and validate a structured configuration format. This configuration can be a central artifact that defines an analysis task or an "assistant".

[0202] A configuration development lifecycle can include the following: First, requirements can be defined. This can involve analyzing a target workflow, data, and desired outcomes. The result can be a structured set of criteria or checklist items that an automated analysis should address. A feasibility assessment can include evaluating whether defined criteria can be reliably identified based on recognizable textual regularities or semantic patterns in representative documents. A compilation of exemplary data can include a collection of various text fragments that correspond to each criterion, focusing on variation rather than quantity to enable a robust pattern specification, as well as a collection of representative test documents containing expected results.Configuration design and iterative testing can utilize an integrated configuration environment to translate requirements and examples into a declarative pipeline specification. This can include selecting and parameterizing stages, defining complex patterns, and, if necessary, hierarchically structuring the configuration to represent checklist items. Iterative testing can involve executing a configuration against test documents within the environment, reviewing the results, and refining stage parameters or patterns. Features such as commenting and disabling stages can support this iterative process. Quality assurance can involve systematic testing with a broader set of documents to evaluate performance, with the goal of minimizing missed relevant results (false negatives) and / or reducing irrelevant results (false positives).Deploying the system can include finalizing and defining a validated configuration as a reusable analysis template, which can be made available to authorized users within the system for new document sets. For continuous improvement, the system can incorporate user feedback on the performance of a template (e.g., via integrated commenting and / or rating mechanisms, separate review processes, and / or Reinforcement Learning from Human Feedback (RLHF) algorithms) to enable subsequent iterations and refinements of the configuration.

[0203] Fig. Figure 15 illustrates an embodiment of a general-purpose computer 1050. The general-purpose computer 1050 is an example of an information processing device comprising a circuit configured to execute the method according to the present technology (e.g., the methods from Fig. 2 to 8) and the user interfaces from Fig. 3 and Fig. 4 to display. The 1050 multi-purpose computer can display circuit 1 from Fig. Provide 1.

[0204] Embodiments which use software, firmware, programs or the like to carry out the procedures described herein may be installed on the computer 1050, which is then suitably configured for the corresponding embodiment.

[0205] The computer 1050 has a central processing unit (CPU) 1051, which can execute various types of procedures and processes, as described herein, for example, according to programs stored in a read-only memory (ROM) 1052; stored in a memory 1057 and loaded into a random-access memory (RAM) 1053; stored on a medium 1060 which can be inserted into a corresponding drive 1059; etc.

[0206] The Computer 1050 also includes an Artificial Intelligence (AI) Processor 1051a. The AI ​​Processor 1051a can include a Graphics Processing Unit (GPU), a Tensor Processing Unit (TPU), a Photonics Processor, and / or a Quantum Processor. The AI ​​Processor 1051a can be configured to run an AI model (e.g., an artificial neural network), such as the generative language model 83 from Fig. 6.

[0207] The CPU 1051, the ROM 1052, and the RAM 1053 are connected to a bus 1061, which in turn is connected to an input / output interface 1054. The number of CPUs, memory modules, and storage devices is only exemplary, and those skilled in the art will understand that the computer 1050 can be adapted and configured accordingly to meet specific requirements that arise when it is used as an information processing device according to the present technology.

[0208] Several components are connected to the input / output interface 1054: an input 1055, an output 1056, a memory 1057, a communication interface 1058 and the drive 1059, into which a medium 1060 (Compact Disk (CD), Digital Video Disc (DVD), Universal Serial Bus (USB) flash drive, Secure Digital (SD) card, CompactFlash (CF) memory or the like) can be inserted.

[0209] Input 1055 can include a pointing device (mouse, graphics tablet or similar), a keyboard, a microphone, a camera, a touchscreen, an eye-tracking unit, etc.

[0210] The 1056 output can have a display (liquid crystal display (LCD), cathode ray tube (CRT) display, light-emitting diode (LED) display, electronic paper, etc.; e.g., included in a touchscreen), speaker, etc.

[0211] The 1057 memory can be a hard disk drive (HDD), a solid-state drive (SSD), a flash drive, and the like.

[0212] The 1058 communication interface can be adapted to communicate via, for example, Universal Serial Bus (USB), a serial port (RS-232), a parallel port (IEEE 1284), a Local Area Network (LAN; e.g., Ethernet), Wireless Local Area Network (WLAN; e.g., Wi-Fi, IEEE 802.11), mobile telecommunications systems (3G, 4G, 5G, 6G; GSM, UMTS, LTE, NR, etc.), Bluetooth, Near-Field Communication (NFC), ZigBee, infrared, etc.

[0213] It should be noted that the above description only concerns an exemplary configuration of the Computer 1050. Alternative configurations can be implemented with additional or different sensors, storage devices, interfaces, or the like. For example, the Communication Interface 1058 can support radio access technologies other than the mentioned UMTS, LTE, and NR.

[0214] It should be noted that the embodiments described herein can be combined as appropriate. In some embodiments, the sequence of process steps differs from the sequence described herein.

[0215] Further examples of the present technology are presented below. (1) Circuit for processing the content of a digital document, wherein the circuit is set up for: Receiving configuration data for the content preparation of a digital document, wherein the configuration data indicates a processing chain with processing modules from a predefined set of processing modules for the content preparation of a digital document; Generating a processing chain according to the configuration data; Receiving a digital document; and Applying the processing chain to the digital document to prepare the digital document. (2) Circuit according to (1), comprising obtaining the configuration data: Displaying a graphical user interface in which processing modules from the set of processing modules are represented as objects; where objects for processing modules of the processing chain are displayed in a predefined section. (3) Circuit according to (2), further comprising obtaining the configuration data: Receiving an initial user operation indicating a dragging of one of the objects from outside the predefined section into the predefined section; and Add a processing module, corresponding to the object pulled with the first user operation, to the processing chain in the configuration data. (4) Circuit according to (2) or (3), further comprising obtaining the configuration data: Receiving a second user operation that indicates dragging one of the objects from a first position in the predefined section to a second position in the predefined section; and Moving a processing module of the processing chain, corresponding to the object pulled by the second user operation, to a position in the processing chain corresponding to the second position in the configuration data. (5) Circuit according to one of (1) to (4), wherein the circuit is configured to: Detecting a dependency of one processing module displayed by the configuration data on output data of another processing module from the set of processing modules; Detecting the absence of the other processing module in the processing chain; and issuing a notification of an unfulfilled dependency. (6) Circuit according to one of (1) to (5), comprising obtaining the configuration data: Receiving an instruction from a user to generate the configuration data; and Executing a generative language model with the instruction as input data; where the generative language model is set up to generate the configuration data based on the instruction. (7) Circuit according to (6), wherein the generative language model is further configured to: Determining a characteristic for the processing chain according to the instruction; and Generating the configuration data according to the detected characteristic for the processing chain. (8) Circuit according to (7), comprising determining the feature for the processing chain: Determining a processing criterion according to the instructions; and Determining the characteristic for the processing chain to generate output data based on the digital document that meets the processing criterion. (9) Circuit according to one of (1) to (8), wherein the circuit is further configured to: Providing a set of sample documents; Applying the processing chain to the set of sample documents; Outputting a reference to a result of the processing chain applied to the set of sample documents. (10) Circuit according to one of (1) to (9), where the configuration data is represented in a structured markup language. (11) Circuit according to one of (1) to (10), comprising generating the processing chain: Loading data processing instructions that correspond to processing modules of the processing chain; and Generating instructions to execute the loaded data processing instructions according to the configuration data. (12) Circuit according to one of (1) to (11), wherein the application of the processing chain to the digital document comprises: Generate, according to the processing chain, an annotation object that displays output data of a processing module in the processing chain. (13) Circuit according to (12), where the annotation object displays a section of the digital document to which the output data of the processing module refers. (14) Circuit according to (12) or (13), where the annotation object indicates that the processing module has not detected any data that meets a processing criterion. (15) Circuit according to one of (1) to (14), comprising applying the processing chain to the digital document: Receiving output data from the first processing module of the processing chain; Providing the output data as input data for a second processing module in the processing chain, according to a dependency of the second processing module indicated by the configuration data; and Executing the second processing module. (16) Circuit according to one of (1) to (15), further configured to: Authorizing a user connected via a user terminal device; Receipt of the digital document from the user; Receiving an instruction from the user indicating the processing chain; and Applying the processing chain to the digital document according to the user's instructions. (17) Circuit according to one of (1) to (16), wherein a processing module from the set of processing modules is configured to: Receiving input data; Performing predefined data processing on the input data, wherein the predefined data processing is associated with the processing module; and Generating output data that displays a result of the predefined data processing. (18) Circuit according to one of (1) to (17), wherein at least one processing module from the set of processing modules is set up to include at least one segmenting, filtering, expanding, merging, sorting, limiting and combining. (19) Circuit according to one of (1) to (18), where the configuration data is shown: Splitting the processing chain into at least two branches; and Combining at least two branches with a processing module from the set of processing modules that is designed to combine branches. (20) Methods for processing the content of a digital document, the method comprising: Receiving configuration data for the content preparation of a digital document, wherein the configuration data indicates a processing chain with processing modules from a predefined set of processing modules for the content preparation of a digital document; Generating a processing chain according to the configuration data; Receiving a digital document; and Applying the processing chain to the digital document to prepare the digital document. (21) Procedure according to (20), comprising obtaining the configuration data: Displaying a graphical user interface in which processing modules from the set of processing modules are represented as objects; where objects for processing modules of the processing chain are displayed in a predefined section. (22) Procedure according to (21), further comprising obtaining the configuration data: Receiving an initial user operation indicating a dragging of one of the objects from outside the predefined section into the predefined section; and Add a processing module, corresponding to the object pulled with the first user operation, to the processing chain in the configuration data. (23) Procedure according to (21) or (22), further comprising obtaining the configuration data: Receiving a second user operation that indicates dragging one of the objects from a first position in the predefined section to a second position in the predefined section; and Moving a processing module of the processing chain, corresponding to the object pulled by the second user operation, to a position in the processing chain corresponding to the second position in the configuration data. (24) A procedure according to one of (20) to (23), the procedure comprising: Detecting a dependency of one processing module displayed by the configuration data on output data of another processing module from the set of processing modules; Detecting the absence of the other processing module in the processing chain; and issuing a notification of an unfulfilled dependency. (25) Method according to one of (20) to (24), comprising obtaining the configuration data: Receiving an instruction from a user to generate the configuration data; and Executing a generative language model with the instruction as input data; where the generative language model is set up to generate the configuration data based on the instruction. (26) Procedure according to (25), wherein the generative language model is further configured to: Determining a characteristic for the processing chain according to the instruction; and Generating the configuration data according to the detected characteristic for the processing chain. (27) Procedure according to (26), comprising determining the characteristic for the processing chain: Determining a processing criterion according to the instructions; and Determining the characteristic for the processing chain to generate output data based on the digital document that meets the processing criterion. (28) A procedure in accordance with one of (20) to (27), the procedure further comprising: Providing a set of sample documents; Applying the processing chain to the set of sample documents; Outputting a reference to a result of the processing chain applied to the set of sample documents. (29) Procedure according to one of (20) to (28), where the configuration data is represented in a structured markup language. (30) Method according to one of (20) to (29), comprising generating the processing chain: Loading data processing instructions that correspond to processing modules of the processing chain; and Generating instructions to execute the loaded data processing instructions according to the configuration data. (31) A procedure in accordance with one of (20) to (30), comprising applying the processing chain to the digital document: Generate, according to the processing chain, an annotation object that displays output data of a processing module in the processing chain. (32) Procedure according to (31), where the annotation object displays a section of the digital document to which the output data of the processing module refers. (33) Procedure according to (31) or (32), where the annotation object indicates that the processing module has not detected any data that meets a processing criterion. (34) A procedure in accordance with one of (20) to (33), comprising applying the processing chain to the digital document: Receiving output data from the first processing module of the processing chain; Providing the output data as input data for a second processing module in the processing chain, according to a dependency of the second processing module indicated by the configuration data; and Executing the second processing module. (35) Procedures in accordance with any of (20) to (34), further comprising: Authorizing a user connected via a user terminal device; Receipt of the digital document from the user; Receiving an instruction from the user indicating the processing chain; and Applying the processing chain to the digital document according to the user's instructions. (36) A method according to one of (20) to (35), wherein a processing module from the set of processing modules is set up for: Receiving input data; Performing predefined data processing on the input data, wherein the predefined data processing is associated with the processing module; and Generating output data that displays a result of the predefined data processing. (37) Procedure according to one of (20) to (36), wherein at least one processing module from the set of processing modules is set up to include at least one segmenting, filtering, expanding, merging, sorting, limiting and combining. (38) Method according to one of (20) to (37), wherein the configuration data is displayed: Splitting the processing chain into at least two branches; and Combining at least two branches with a processing module from the set of processing modules that is designed to combine branches. (39) Computer program comprising instructions which, when the program is executed by a computer, cause it to perform the procedure according to any of (20) to (38).

Claims

[1] Circuit for processing the content of a digital document, wherein the circuit is set up to: Receiving (11) configuration data for the content preparation of a digital document, wherein the configuration data indicates a processing chain with processing modules from a predefined set of processing modules for the content preparation of a digital document; Generating (12) a processing chain according to the configuration data; Receiving (16) a digital document; and Applying (18) the processing chain to the digital document to prepare the digital document. [2] Circuit according to claim 1, wherein receiving (11) the configuration data comprises: Displays (31; 51) of a graphical user interface (40; 60) in which processing modules from the set of processing modules are represented as objects (42a, 42b, 42c, 42c, 43a, 43b, 43c, 43d, 43c; 62a, 62b, 62c, 62d, 63a, 63b, 63c, 63d, 63e); where objects (43a, 43b, 43c, 43d, 43c; 63a, 63b, 63c, 63d, 63e) for processing modules of the processing chain are displayed in a predefined section (43). [3] Circuit according to claim 2, wherein obtaining (11) the configuration data further comprises: Receiving (32) a first user operation, which indicates a dragging of one of the objects (42b) from outside the predefined section (43) into the predefined section (43); and Adding (33) a processing module corresponding to the object (42b) drawn with the first user operation to the processing chain in the configuration data. [4] Circuit according to claim 2 or 3, wherein receiving (11) the configuration data further comprises: Receiving (52) a second user operation, which indicates dragging one of the objects (63d) from a first position in the predefined section (63) to a second position in the predefined section (63); and Moving (53) a processing module of the processing chain, corresponding to the object (63d) drawn by the second user operation, to a position in the processing chain corresponding to the second position in the configuration data. [5] Circuit according to one of the preceding claims, wherein the circuit is configured to: Identifying (71) a dependency of a processing module displayed by the configuration data on output data of another processing module from the set of processing modules; Detecting (72) the absence of the other processing module in the processing chain; and Issuing (73) a reference to an unfulfilled dependency. [6] Circuit according to one of the preceding claims, wherein obtaining (11) the configuration data comprises: Receiving (81) an instruction from a user to generate the configuration data; and Executing (82) a generative language model (83) with the instruction as input data; wherein the generative language model (83) is set up to generate the configuration data based on the instruction. [7] Circuit according to claim 6, wherein the generative language model (83) is further configured to: Determine (84) a characteristic for the processing chain in accordance with the instruction; and Generating (86) the configuration data according to the detected feature for the processing chain. [8] Circuit according to claim 7, wherein determining (84) the feature for the processing chain comprises: Determine (85) a processing criterion in accordance with the instruction; and Determining the characteristic for the processing chain to generate output data based on the digital document that meets the processing criterion. [9] Circuit according to one of the preceding claims, wherein the circuit is further configured to: Providing (91) a set of sample documents; Applying (92) the processing chain to the set of sample documents; Output (93) of a reference to a result of the processing chain applied to the set of sample documents. [10] Circuit according to one of the preceding claims, wherein the configuration data is represented in a structured markup language. [11] Circuit according to one of the preceding claims, wherein generating (12) the processing chain comprises: Loading (13) data processing instructions corresponding to processing modules of the processing chain; and Generating (14) instructions to execute the loaded data processing instructions according to the configuration data. [12] Circuit according to one of the preceding claims, wherein applying (18) the processing chain to the digital document comprises: Generate (22), according to the processing chain, an annotation object that displays output data of a processing module of the processing chain. [13] Circuit according to claim 12, wherein the annotation object displays a section of the digital document to which the output data of the processing module refers. [14] Circuit according to claim 12 or 13, wherein the annotation object indicates that the processing module has not detected any data that meets a processing criterion. [15] Circuit according to any of the preceding claims, wherein applying (18) the processing chain to the digital document comprises: Receiving (19) output data from a first processing module of the processing chain; Providing (20) the output data as input data for a second processing module of the processing chain according to a dependency of the second processing module indicated by the configuration data; and Execute (21) the second processing module. [16] Circuit according to one of the preceding claims, further configured to: Authorizing (15) a user connected via a user terminal device; Receipt (16) of the digital document from the user; Receiving (17) an instruction from the user indicating the processing chain; and Applying (18) the processing chain to the digital document in accordance with the user's instructions. [17] Circuit according to one of the preceding claims, wherein a processing module from the set of processing modules is configured to: Receiving (101) input data; Performing (102) a predefined data processing on the input data, wherein the predefined data processing is associated with the processing module; and Generating (103) output data that displays a result of the predefined data processing. [18] Circuit according to one of the preceding claims, wherein at least one processing module from the set of processing modules is configured to perform at least one segmenting, filtering, expanding, merging, sorting, limiting and combining operation. [19] Circuit according to one of the preceding claims, wherein the configuration data displays: Splitting the processing chain (110) into at least two branches; and Combining at least two branches (113a, 113b, 113c, 113d, 114a, 114b, 114c) with a processing module (115) from the set of processing modules designed for combining branches. [20] Computer program comprising commands which, when the program is executed by a computer, cause it to perform a method (10) for processing the content of a digital document, wherein the method (10) comprises: Receiving (11) configuration data for the content preparation of a digital document, wherein the configuration data indicates a processing chain with processing modules from a predefined set of processing modules for the content preparation of a digital document; Generating (12) a processing chain according to the configuration data; Receiving (16) a digital document; and Applying (18) the processing chain to the digital document to prepare the digital document.