Improvements in retrieval-enhanced generation for large language models

The lossless hierarchical text processing technique addresses LLM context window limitations by structuring documents for efficient retrieval-enhanced generation, ensuring accurate and contextually relevant outputs.

DE202025101876U1Active Publication Date: 2025-06-05TIGON S L U
View PDF 0 Cites 12 Cited by

Patent Information

Application Number
DE202025101876
Authority / Receiving Office
DE · DE
Patent Type
Utility models
Current Assignee / Owner
Priority Date
2025-04-04
Filing Date
2025-04-07
Publication Date
2025-06-05
Estimated Expiration
2035-04-30

AI Technical Summary

Technical Problem

Large language models (LLMs) face limitations due to a constrained context window, leading to inefficiencies in processing large documents, increased computational resources, and loss of semantic coherence in chunking methods, particularly when dealing with multimodal content like tables and images.

Method used

A lossless hierarchical text processing technique that transforms documents into structured representations, preserving all content elements, and generates query units that fit within the LLM's context window, using iterative methods and multimodal integration to maintain contextual integrity.

Benefits of technology

Enables accurate, contextually relevant, and traceable generation by reducing processing power, storage, and bandwidth requirements, while maintaining the fidelity of the original document content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

A computer program comprising instructions which, when executed by a computer, cause the computer to perform the following operations: Receiving an input document (104), wherein the input document (104) contains content elements including at least one text element; Performing a lossless hierarchical text processing process (106) in which the input document (104) is reproduced as a hierarchical representation (108) in which all text elements from the input document (104) are preserved verbatim and organized according to a structure of the input document (104), the lossless hierarchical text processing process (106) comprising: Prompting a generative language model to generate the hierarchical representation (108), wherein the prompt instructs the generative language model to generate the hierarchical representation (108) without summarizing or omitting any of the text elements; and Performing a retrieval unit generation process (110) in which retrieval units (112) are generated based on the hierarchical representation (108) and stored in a memory accessible by a generative client language model for retrieval-enhanced generation.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELDThis disclosure relates to the field of information processing for large language models, and more particularly to systems for call extended generation.BACKGROUNDLarge Language Models (LLMs) have proven to be powerful tools for information processing and generation. These models, trained on comprehensive text corpora, can understand and generate human-like texts in various areas. As companies attempt to utilize LLMs for specific applications, Retrieval-Augmented Generation (RAG) has evolved into an important approach to support LLM outputs on reliable information sources.However, large speech models are fundamentally limited by a limited context window. This limitation places a hard limit on the amount of data that an LLM can directly process and makes it impossible to capture large documents or records without truncation or information. What is important is that the number of tokens within the context window directly affects the computational load, resulting in an increase in processing power and power consumption as the context size increases. This leads to a difficult compromise: a broader context requires exoribbitant computing resources, which limits the practical application. The smaller amount of data in modern documents encounters these limits and makes efficient processing a often insurspible challenge.To overcome these processing constraints, various approaches have been developed to disassemble documents into smaller units, called "chunks", which can be embedded, stored, and later retrieved as needed. However, conventional chunking techniques often rely on simplified length-based partitions or static rules that interfere with semantic coherence or may omit contextual structures. Recent approaches have explored somewhat more adaptive chunking techniques, such as recursive splitting on separators, semantic segmentation over embeds, or even the decision on chunk boundaries by an LLM.In a communication discussion (see https: / / communication.openai.com / t / using-gpt-4 api-to-semantically-chunk-documents / 715689), it has been proposed, for example, to request GPT-4 to analyze a document and to create a semantically separate, partial organization of the content and then to use this organization for extraction of text segments. In the discussion thread, a user describes requesting GPT-4 to create a "hierarchical" (sic) string, and proposes using the string as a guide for extraction of text chunks. However, in the discussion, it remains unclear how these general ideas could be put into practice.Another example: zChunk uses a Lama 70B model to insert special bounding characters (a user-defined character such as " ") into the text to mark semantic portion boundaries. Essentially, the model is requested to copy the document and insert interrupt characters at logical points. However, zUnk creates a flat list of chunks with limited expressiveness.Moreover, in most existing chunking techniques, the original formatting or all of the detail of the source is lost. Approaches such as hierarchical summary (as in RAPTOR; Recurrent Abstract Processing for Tree-Organized Retrieval) create multi-level summary of document clusters, which, although useful in determining the long-range context, naturally disregard details due to abstraction. While abstracts may capture the essence of a source, they are unsuitable where exact reminder of facts or formulations is required (e.g., in legal texts, technical documentation, etc.). Frameworks that make hierarchies tend to use either compressed representations at each level or exogenous skill in the art. For example, the RAPTOR system explicitly constitutes a tree of contents, but each node is an abstract summary of the underlying text. This means that RAPTOR and similar methods lose critical peculiarities, so that they are unsuitable for scenarios in which absolute accuracy is required.There is a need for a lossless chunking technique that maintains all information from the source in structured form so that the responses of an LLM can directly cite or reference the original text as needed, thereby ensuring accuracy and fidelity.In addition to text, modern information systems have increasingly been concerned with multimodal content such as tables, pictures, diagrams, code chips and the like. These elements place special requirements on the processing. Conventional RAG systems process primarily unstructured text. If non-textual elements (e.g., a complex table or explanatory diagram) are present, these are either omitted or processed in separate silo pipelines (e.g., images via image processing models, tables via simple CSV serialization), which may result in incomplete context. However, if an image embedded in a document contains important information that is not captured elsewhere, information is lost if the system ignores that image or indexes its label only.Tabular data presents a particular challenge to language model processing. Tables often contain dense information with implicit relationships between rows and columns. In a simple linear text conversion or even mark-down, some of these relationships are inevitably made indistinguishable or are completely lost (e.g., connected cells containing sums or repetitions). In addition, large tables may contain hundreds or thousands of cells, making comprehensive inclusion impractical due to the constraints of the context window and the computational effort involved. However, when tableing, critical data points that might be relevant to particular queries are likely to be omitted. Research to enable LLMs to process long tables (e.g., TableRAG, 2024) illustrates the difficulty in accommodating completeness and context length.Maintaining the original fidelity presents a further challenge in all of the above modalities. Information can be compressed with the aid of summary methods, but details are always lost in this case. In applications requiring high accuracy, such as in legal, medical or safety contexts, any loss of information could be problematic. However, obtaining the full information with simultaneous accessibility for retrieval and generation increases the memory and bandwidth requirements, which further strain the existing system boundaries. The smaller amount of data required for lossless processing represents a significant bottleneck.It is therefore an object of the present disclosure to enable speech models to access and utilize information with greater completeness and structural integrity, thereby at least partially overcoming the disadvantages of the prior art, in a manner that reduces the high computational, storage and bandwidth requirements associated with existing methods.SUMMARY OF THE DISCLOSUREThe above and other objects can be achieved by the subject matter defined in the independent claims. Advantageous modifications of embodiments of the present disclosure are defined in the dependent claims as well as in the description and the drawings.As a general overview, aspects and embodiments of this disclosure provide a lossless hierarchical text rendering technique that converts source materials to an LLM-digestible knowledge base with complete retention of information. This is the basis for an end-to-end pipeline for fetch extended generation that provides responses that are much more accurate, contextually relevant, and returnable to the original source. In essence, large language models are enabled to have detailed, trusted memory for source material of any size despite the limitations of the context window inherent to them.The techniques described below overcome the problem of bounding the context window not by lossy compression, but by intelligent re-structuring and extension of the term "context"-from a flat text string to a rich structured document. The constraints on context length and hallucination tendencies of LLMs are addressed directly by requiring LLMs no longer to inverify huge raw texts or images at the time of retrieval and they need not rate over data that has been cropped or abstracted. Instead, LLMs have access to ground truth in mouth-true chunks. The result is optimized, call extended generation, in which the responses are more accurate, contextual, and returnable to the original material.One aspect of the present disclosure relates to a method. The method may be a document preparation method which serves to prepare a document for the call-extended generation. The method may be computer-implemented. Within the scope of this disclosure, any description of a method should be understood to also disclose a corresponding computer program implementing the method and a corresponding data processing apparatus configured to execute the computer program. Likewise, any description of a method step should be understood as disclosing a corresponding operation of the computer program and a unit, module, or device of the data processing apparatusThe method may include receiving an input document. The input document may consist of content items. At least one of the content items may include at least one text item.It can be provided that the method carries out a lossless hierarchical text editing method. In lossless hierarchical text editing, the input document can be rendered as a hierarchical representation (also called "hierarchical organization" or "hierarchical representation"). In the hierarchical representation, preferably all text elements of the input document are retained word-by-word and ordered according to the structure of the input document. In other words, the lossless hierarchical text rendering method can reproduce an input document as a hierarchical representation with the aim of obtaining all content items comprising text in their original form. The term "lossless" means that the method is possible without altering or omitting the original text in the content items, but each content item of the input document is retained in the word word. The hierarchical representation reflects the actual structural organization of the input document and allows clear discrimination of the relationships between the content items whether sections, paragraphs, images or other items.The creation of a word-literal hierarchical representation supports subsequent query processes considerably. Thanks to the hierarchical representation, such query processes can operate on the entire content without truncations, so that downstream generative speech models can access a broad range of information for query-enhanced tasks. This comprehensive data access improves the ability to produce informed relevant results that closely resemble the original input document.It can be provided that the method carries out a process for generating query units. In the generation of search units, search units may be generated based on the hierarchical representation. Accordingly, during the search unit generation process, the hierarchical representation may be analyzed to derive particular search units. These query units may match contiguous segments of the hierarchical representation that may be accessed independently and that may be useful for query purposes. Each retrieval unit can encapsulate certain content items and maintain their literal qualities.The generated query units can be stored in a memory (repository), which can be accessed by a customer's generative language model for query-supported generation. Accordingly, the repository may function as a structured database that provides access to a customer's generative language model. Such models optimized for fetch-extended generation tasks may benefit from the repository by fetching relevant fetch units to inform and improve their generative capabilities. Through the use of the repository, generative language models can directly access accurate and intact content of the input document and thus support more informed, contextually relevant and more precise results of content generation.In summary, the method according to the above aspects ensures a smooth transition from raw document input to structured content presentation, resulting in a robust query framework that can improve the operational efficiency of generative language models. By preserving the integrity of the original content and organizing it for searching, the method becomes a useful tool for processing and handling documents and generative applications. In particular, the proposed method addresses the critical technical limitations of large speech models by fundamentally changing the approach of document processing for call extended generation. Rather than attempting to force whole large documents into the inherently limited context window of the LLM, resulting in loss of information and increasing computational cost, the lossless hierarchical text rendering process carefully reproduces the input document into a structured hierarchical representation with each content item maintained literally while organizing it according to the inherent structure of the document. This initial processing step can be performed once, so that the effort for processing comprehensive documents is dispensed with. Subsequently, the query unit generation process uses this hierarchical representation to create smaller contextually relevant query units that can be designed to fit comfortably within the context window of the LLM and allow efficient on-demand retrieval of specific information without requiring the model. By concentrating on these smaller units, the method can significantly reduce the processing power associated with each request and the power consumption, since the LLM processes only the necessary information.Moreover, this approach optimizes storage capacity by avoiding redundancies in the retrieved data thanks to the structured organization of the hierarchical representation, and minimizes bandwidth utilization by transmitting only the required retrieving units to the LLM.In essence, the method overcomes the constraints on context windows, processing power, storage space, and bandwidth by preprocessing the data into a hierarchical structure which is then divided into smaller query units that can be consistently deduplicated and / or combined so that the LLM only needs to handle small relevant data packets.It may be provided that the lossless hierarchical text rendering process comprises requesting a generative language model to generate the hierarchical representation. The request to the generative language model may include instructing the generative language model to generate the hierarchical representation without summarizing and / or omitting the text elements.Accordingly, this aspect takes advantage of the ability of large language models to track instructions and copy text, thereby achieving a consistent segmentation of the input document. The prompt instructs the language model what it is to do. The prompt may include instructions or cues that direct the generative language model to focus on the literal representation to ensure that each text element of the input document is retained in its original form within the hierarchical representation. The instruction to avoid pooling or omissions is part of the concept "lossless" within the hierarchical retry process. It aims to accurately reproduce each content item and maintain the full information content of the input document.It can be provided that the lossless hierarchical text rendering method is an iterative method. The iterative process may include a first iteration, a second iteration, and so forth. Each iteration may include generating a partial hierarchical representation until a processing threshold is reached. The processing threshold may be a predefined threshold. In one implementation, the processing threshold corresponds to a threshold of the context window of the generative language model responsible for processing the input document. An example of such a threshold for the context window is 10,000 tokensEach iteration may include compressing the partial hierarchical representation of the current iteration (and all partial hierarchical representations of earlier iterations).In all cases, such compression may be horizontal compression, vertical compression, or both. Horizontal compression may be to maintain the first and last portions of a text element (e.g., the first and last words, the first and last words, or the like) and replace the remainder with a place holder such as "[...].". Vertical compression may be for a sequence of text elements at the same hierarchical level by maintaining the first and last text elements and replacing the intermediate text elements with a place holder such as "[...]".Each iteration may include concatenating the compressed partial hierarchical representation of the current iteration (and all compressed partial hierarchical representations of earlier iterations) with a boundary marker followed by the unprocessed remainder of the input document. The result then serves as an input for the subsequent iteration.For illustrative purposes, the iterative process will run the first 10,000 tokens. The result is compressed horizontally and vertically. The result is concatenated with a separator and the as yet unprocessed part of the input document and used as input for the second iteration. After the process has been run through for a further 10,000 tokens, the two partial hierarchical representations (from the first and second iteration) are combined and the combination result is compressed, etc.The boundary marker that marks the boundary between the processed and unprocessed portions of the input document may be a mnemonic boundary marker. An example is "§$#NEUE CONTENTS SHALL BE SEPARATED HERE# §". This is advantageous because it can be easily understood by the generative language model that generates the hierarchical representation.In this variant of the method, the smooth manipulation of input documents that exceed the context size of the generative language model may involve an iterative approach during the lossless hierarchical text rendering process. This process supports continuity in the structural organization during iterations and provides coherent hierarchical representation even for large amounts of data. If the size of an input document exceeds the context capacity of the generative language model, which may limit its ability to store all content items simultaneously, this iterative processing divides the document into manageable portions. Each section is fed into the generative language model one after the other so that it can focus on sub-sections while the overall hierarchical framework is maintained.A key aspect of this variant of the method is to emphasize continuity in structural organization throughout iterations. This continuity can be achieved by maintaining organizational markers or depth indicators such as headers, intermediate headers, formatting effects, and levels of intervention. These markers may help the generative language model understand the relationship between the individual sections to keep the hierarchy smooth and uniform. To support the iterative process without disrupting the structural hierarchy, transition prompts or indications may assist the model in maintaining a consistent organizational depth between the sections.This capability provides the potential for efficient execution of large documents and ensures that generative language models can effectively process, retrieve, and generate content without compromising context size.It can be provided that the input document is a structured input document. In this case, the content items in the hierarchical representation may be organized according to an explicit structure in the input document.Accordingly, structured input documents may benefit from organization of content items within the hierarchical representation corresponding to the explicit structure of the document. This organization is practical to maintain inherent structural features and to improve the accuracy and relevance of the reproduced hierarchical model. Structured documents may include formats such as partitions, directories, or visibly marked sections and sub-sections. These elements can serve as an orientation aid to create a clear hierarchy for the presentation of the content. The method can use these organizational indications to create the format of the document and maintain it throughout the repeat process.The hierarchical representation may include the alignment of content items at their respective structural markers from the original document. These markers may include hierarchical indicators such as numbering systems, headers, intermediate headers, enumeration points, engagement levels, or other formatting details that suggest a particular organizational layout. These markers may be used to set relationships and priorities between content items, which may maintain the logical flow and visual alignment of the document.The structured format of the hierarchical representation supports the effective generation of retrieval units from the repository. With a well organized basis, downstream generative tasks are likely to access precisely indexed segments, enabling contextually precise and logically based results.It can be provided that the input document is an unstructured input document. In this case, the content items may be organized in the hierarchical representation according to a derived structure of the input document. The derived structure may be determined by the generative language model.Accordingly, when processing unstructured input documents, the organization of the content items within the hierarchical representation may rely on a structure proposed by the generative language model. This inference process provides an approach to creating a coherent logical ordering of documents that do not have explicit structure markers. Unstructured documents are those in which sections, sub-sections, or other organizational elements that direct the formation of hierarchies are not clearly defined. Examples of such documents are free form texts, enumerator content or mixed media documents, which may lack traditional formatting instructions. The challenge with these documents is to place them in a structured format that facilitates understanding and retrieval.In this scenario, the generative language model may utilize its ability to understand language patterns, context, and semantic hints, as well as to "think when speaking" to determine a proposed structure. The model examines the contents of the document and identifies logical breaks, thematic displacements, or natural trails that may indicate organizational paradigms.Once the proposed structure is arranged, the content items can be positioned within the hierarchy so that the document assumes an ordered format that is very similar to structured counterparts. By using a proposed organization strategy, the method adapts unstructured documents into systematic, ordered representations. This framework may improve the effectiveness of generative applications, including query-assisted synthesis, as it provides a contiguous and organized database for effective downstream operations.It may be provided that the hierarchical representation uses engagement levels to organize the content items.In this variant of the method, engagement levels, for example by tabs or spaces, are used as organization means in the hierarchical representation of the content elements of the input document. The intervention provides a visual and structural signal that indicates the placement of the content at different hierarchical levels and promotes clarity of presentation. Engagement levels may indicate the depth of the hierarchy within the document, each level representing a transition from broader categories to finer granularity. This system helps visually distinguish between primary content, sub-content and subordinate details and may reflect the organization inherent or derived from the document.The conversion of an input document to a hierarchical representation using engagement levels (e.g., by space or tab characters) provides a significant technical advantage in terms of efficient data structuring without substantially increasing file size. Hierarchies based on insertions implicitly encode structural depth, thereby reducing overhead associated with additional syntax. This lightweight structuring mechanism allows compact representation of nested contents, in which each hierarchy level is characterized by a minimum amount of void and not by detailed structure markers. Consequently, the resulting file size can remain close to that of the original input document, thereby preserving storage and transfer efficiency, which is particularly important in bandwidth constrained environments or in large data set work.The use of intrusions is also particularly advantageous in creating a hierarchical representation that is intuitive and accessible. For example, main sections with minimal engagement may be represented, while sub-sections and related elements receive progressive planes that clarify potential content relationships and dependencies. The inclusion provides a visually coherent representation that can facilitate understanding and navigation for both humans and generative language models. This hierarchy may contribute to subsequent processes, e.g., automatic generation and indexing of polling units, where coherent intervention helps efficiently retrieve contextually matching segments, as well as manual follow-up tasks such as fine tuning or reinforcement learning from human feedback (RLHF).It may be provided that the hierarchical representation uses semantic labels, such as markup elements and / or markup elements, for organization of the content elements.In this variant of the method, semantic inscriptions such as markup elements and / or markup elements are used as organization means for the hierarchical representation of the content elements of the input document. Semantic labels, including markup and markup elements, provide a framework for content characterization and cataloging. These elements can explicitly identify different sections of a document, which allows insight into functional roles, hierarchy levels and thematic divisions. Both markup and markup may provide syntactical evidence that allows content to be systematically identified and create a hierarchy that may be visually and semantically coherent.This approach allows the generative language model to identify and identify various content items using these semantic markers. For example, headings, paragraphs, lists, and highlighted sections can be clearly outlined, indicating relationships and resulting in a structured representation. The use of these markers adds an additional level of organization that can reflect both the explicit structure of documents and the derived organization for unstructured entries.Moreover, this semantic label could improve downstream processes, especially the generation of retrieval units and the storage as well as the retrieval-assisted generation. Semantic labels could allow generative models to more quickly access indexed information and provide contextual insights to effectively inform the results. This methodical improvement could promote more comprehensible, more accurate, and meaningful data usage in generative applications.It can be provided that the process for generating retrieval units comprises the generation of chunk sets as retrieval units. Each chunk set may consist of a plurality of chunks. Each chunk set may indicate a traversal path through the structure of the input document. A maximum token limit for the chunk sets may be predefined, e.g., about 500 tokens per chunk set, or the like.Accordingly, chunk sets may serve as coherent segments extracted from the hierarchical representation and including content having meaningful connections and a logical flow. Chunksets may function as entities that skip particular paths within the structured hierarchy of the document. These paths may represent an associated sequence of content items that conform to the document's organization logic whether derived or explicit. By defining traversal paths, chunk sets facilitate systematic access to document segments, thus enabling efficient search and analysis.The process for generating chunk sets includes segmenting the hierarchical representation into identifiable tracts taking into account the structure and depth defined in the document model. Each chunk set may include a sequence of related content (in the sense of chunks) while preserving word fidelity and contextual relevance. This enables the quick discovery of coherent information while preserving content integrity.Prompts could direct the generative language model to recognize and organize these chunk sets while emphasizing continuity and topic coherence. These entities can be labeled and stored in a repository that provides indexed access paths that can be used by search function language models to augment generative tasks.The use of chunksets improves the discoverability of information, which can improve the efficiency and precision of the results of generative language models. By relying on well-structured entities reflecting the complexity of the document, the models can increase their performance and provide content consistent with the original intentions and nuances of the document.It can be provided that the generation of chunk sets comprises sorting chunks from the hierarchical representation by depth and position. Generating chunk sets may include grouping together adjacent chunks at the same depth level that fit within a token boundary. Generating chunk sets may include recursively preceding parent chunks to form traversal paths. Generating chunk sets may include creating unique combinations that maintain hierarchical relationships while optimizing token efficiency.Accordingly, the process of generating chunk sets may be refined by sorting the chunks from the hierarchical representation by depth and location, which allows for systematic organization and efficiency. This approach aims to improve the use of tokens and at the same time to maintain the structural integrity and hierarchical relationships of the document. In depth organization, the relative hierarchical level is established for each chunk that may point to the structural details of the document. The position organization contributes to the chunks appearing consecutively within the document according to their appearance or position context, which contributes to continuity and information flow. Adjacent chunks located at the same depth level may be collapsed as long as they are within a predefined token boundary. This token boundary may correspond to the constraints of the generative language model, e.g., the maximum input magnitude. By collapsing adjacent chunks, related content segments may be merged, thereby reducing complexity and simultaneously saving space. Recursively adding super-source chunks to form traversal paths may improve coherence and completeness of chunk sets. Parent chunks represent broader hierarchical portions that provide the context for related information. By adding these parent chunks, traversal paths may include both overlapping topics and detailed content, thereby enhancing semantic context in each chunk set. By creating unique combinations of chunksets, hierarchical relationships can be maintained while optimizing the efficiency of the token. This step promotes the development of meaningful and unique retrieval units that closely conform to the original structure of the document. By emphasizing token efficiency, the models are to be enabled to access relevant information without requiring processing capacity, which may improve generative accuracy and relevance.Collectively, these operations promote a methodological approach to generating chunksets that focuses on organization, contextual integrity, and token preservation. By employing techniques for organizing, collapsing, adding, and optimizing, this method is suitable for generating coherent, efficient chunksets suitable for advanced query-assisted generation tasks.It can be provided that each chunk set maintains pointers to the chunks associated with it.This may be an efficient practice that helps manage resources while preserving the fidelity of the hierarchical structure and its content. Pointers serve as referencing links within the chunk sets that point to particular content items or segments in the hierarchical representation. By using pointers, chunksets can conveniently access and use content without having to replicate or store redundant copies. This way of dealing with chunks allows for savings in memory and memory resources.The use of pointers can allow dynamic interaction between chunk sets and the chunks that make them. This flexible relationship may support fast retrieval when generative language models need access to certain segments during execution of retrieve-enhanced generation tasks. By referencing instead of duplication, the model may make predictive use of the data, which may speed up access and maintain contextual relevance.Maintaining pointers may help the hierarchical relationships between the content items to remain integrated and intact. These relationships may help maintain the structural logic of the document and ensure coherency and completeness of downstream generative tasks. The inclusion of pointers may allow the formation of query units that may superimpose semantic and contextual topics, thus improving the depth and richness of the generative results.Moreover, the implementation of pointers could provide the advantage that updates or changes in the original hierarchical representation are automatically accounted for in the associated chunksets. This can facilitate the comparison between the development of the documents and the accuracy of the retrieval units and enable flexible and adaptable binding of the contents.It can be provided that the storage of the chunk sets as retrieval units comprises the calculation of embeddings for each chunk set and the storage of the embeddings in the repository.This approach allows sophisticated indexing and retrieval using embeddings (embeddings) as a compact, multi-dimensional representation of the chunksets. Embeds are numerical vectors that encode semantic information about the content and structure of the chunksets. These vectors capture notable features and relationships of the content and allow a representation that effectively models semantic similarities and differences between chunksets. By computing embeds, the method translates complex hierarchical contents into quantifiable formats that are machine readable and conducive to computer-based analysis.Once computed, these embeds are suitable for storage in a repository that may provide enhanced search and retrieval functions. The repository acts as a database for indexing these vectors and allows fast and precise access to chunk sets via their embeddings. This system may improve query efficiency by utilizing vector operations such as similarity scoring and allowing generative language models to access chunksets relevant to the respective tasks based on semantic closeness.Such optional embed-controlled storage may reduce the retrieval effort by efficiently searching the space of the embeds without requiring extensive raw content to be processed at each query. Processing time is minimized by using pre-computed semantic detections, promoting rapid interaction in call extended generation tasks.Moreover, embedding storage is consistent with modern machine learning practices, where embedding spaces are useful for various applications such as pattern recognition, content recommendations, and contextual functions in generative models. Embeddings also support adaptive learning and contextual adaptation and form a basis for the enrichment of the results of the generative model with contextual and semantically informed data.It may be provided that the method comprises a table integration process. The process of table integration may include identifying table data in the input document. Such table data is, in addition to text elements, a further example of a type of content elements in the input document. The table integration process may include converting the tabular data to a tablet ree representation in the hierarchical representation that retains all data from the tabular data.Identifying tabular data may include identifying tables or table-like structures in the input document. This may include the recognition of headers, rows, columns and data cells which are components of typical tabular representations. The recognition process may use document layout analyses or content parsing techniques to reliably extract table data from different document formats so that these information structures can be accurately interpreted.After identification, the tabular data can be converted to a tablet representation. A table tree may function as a tree-like hierarchical model that reflects the structural layout and data relationships in the original tables. This representation may maintain the organization of rows and columns, where the data cells are linked together according to their relational meaning within the original table formatting.Care may be taken in the transformation to render all data items from the tabular source to maintain semantic integrity and data fidelity. Headers, column descriptors and data entries are transcribed to ensure their inclusion in the hierarchical model so that the tabular elements are accessible and meaningful in the broader context of the hierarchical representation.Tablet ree rendering may provide several advantages for generative language models. It organizes the data in a format that takes into account the logical connections within the tabular data. This conversion facilitates query operations and allows models to interpret and include structured records in the content generation processes, supporting a synthesis that uses the findings from the tables.The integration of tables into the hierarchical representation is consistent with the approach of the document fidelity method and provides generative tasks with access to both narrative and data rich components of a document. This integration may enable generative language applications with enhanced processing, understanding, and content generation capabilities that are anchored in complex tabular information.It can be provided that the tabular data are below a predefined size threshold value. In this case, the conversion of the tabular data may be by requesting a generative language model to convert the tabular data to the tablet representation.If the tabular data falls below a certain size, the conversion to a tablet representation can consist in controlling a generative language model with specific prompts. This approach may facilitate efficient handling of smaller records and offers an option for rationalized processing that maintains content integrity and improves context. The predetermined size threshold could refer to a criterion defining the extent or extent of table data suitable for direct generative transformation. With this concept, the conditions under which generative speech models may effectively operate could be optimized, thereby controlling complexity and improving resource utilization.For smaller table datasets that reach this threshold, prompts may adapt the transformation process by directing the focus of the generative language model to the accurate transformation of table structures into hierarchical table trees. These prompts may make the model recognize tabular elements such as headers, rows, columns, and relationships between data points. By this guidance, the model can comprehensively translate tabular data sets into textual hierarchical formats. The use of generative language models with prompts for the conversion of smaller tabular data may provide speed and efficiency advantages. The model could execute transformation tasks quickly, translate structured data into organization hierarchies, and thereby keep the computational effort low. This method is well suited for operations where a rapid interaction between data interpretation and presentation supports task requirements.It can be provided that the tabular data exceed a predefined size threshold value. In this case, the transformation of the tabular data may be to request a generative language model to generate executable code that transforms the tabular data to the tablet ree representation, and to execute the generated code to generate the tablet ree representation. The "generation of executable code" may also include the generation of parameters for existing code, rather than creating the actual code.If the tabular data exceeds a certain size threshold, the method may cause a generative language model to generate executable code that converts that data to a tablet ree representation. The generated code can then be executed, so that the desired table tree structure arises. This approach provides a mechanism for processing larger data sets that uses both the interpretive and computational capabilities of generative language models. The generative language model could use prompts focusing on the analysis of the structure and semantics of the table data. These prompts direct the model in generating executable code tailored to the organization of the hierarchical structure in tablet ree format. This code may include instructions for handling relationships, categorial partitions, and connections in the large tabular dataset.Execution of the generated code facilitates conversion, indicating that larger records can be systematically processed into tablet representations, while retaining content and contextual richness. In execution, the code instructions are executed and the tabular elements are converted to well organized hierarchical structures. The creation and execution of model generated code provides scalability in transformation tasks and optimizes the procedures in managing extensive tabular data challenges. This framework accommodates the complexity of larger records and aims to preserve semantic integrity and hierarchical organization while balancing system resources.It can be provided that the method comprises a process for integrating multimodal content. The process of multimodal content integration may comprise identifying a non-textual element in the input document. Such non-textual items are another example of a type of content item in the input document, besides text items and tabular data. The process of multi-modal content integration may include integrating a textual representation of the non-textual element into the hierarchical representation.Accordingly, non-textual elements in the input document may be identified and associated with the hierarchical representation by a textual representation. This approach aims to improve the scope and accessibility of the hierarchical model by effectively considering different types of data. Non-textual elements in documents may include images, charts, charts, audio clips, video segments, or other multimedia components. For identifying these elements, techniques for recognition and classification can be used, e.g. image recognition, multimedia parsing or audio-visual analysis. The use of such technological possibilities facilitates the precise extraction and subsequent linking of these elements within the hierarchical frame. Once the non-textual elements are identified, they may be converted into equivalent textual descriptions that include their semantic content and functionality. This textual representation may use descriptive language that captures key aspects such as compositional details, topic purposes, or contextual roles within the original document. Such a translation encodes non-textual data into an accessible format that inserts seamlessly into the hierarchical representation. By integration, this practice ensures that multimodal elements can become part of the coherent organization structure and extend the scope and depth of the hierarchical model. This integral approach may enable generative language models to deal with and synthesize information that includes textual and non-textual domains, thereby improving capacity for search function generative applications. The inclusion of textual representations of multimodal elements could help provide a comprehensive context and enable a more robust semantic understanding of the document. This method promotes improved navigation and search and allows downstream generative processes to obtain enriched information that matches both the textual kernel content and embedded semantic narration.It may be provided that integrating the textual representation comprises generating a textual description of the non-textual element, inserting the textual description into the hierarchical representation at a position corresponding to the position of the non-textual element in the input document, and storing the original non-textual element with a reference identifier associated with its textual description in the hierarchical representation.This may help to provide a balance between information fidelity and organized accessibility for enhanced calls and generative tasks. Storing the original non-textual element with a reference identifier allows the element to be accessed. This identifier acts as a linking mechanism that links the textual description to the source data to allow easy retrieval and referencing. By such linkage, the descriptive text inserts into an enriched organizational framework that supports tasks that require both narrative interpretations and direct access to non-textual formats.It can be provided that the creation of the textual description comprises the calling up of an imaging language model for describing an image. Accordingly, one strategy may be to use a visual language model to describe an image. These models may process visual inputs and identify patterns, objects, contexts, and topic relevance to formulate a descriptive survey. With the aid of advanced image processing, these models can generate a text which captures the most important details of the image and specifies its role and meaning within the scope of the document.It can be provided that creating the text description comprises applying an optical character recognition to extract text from the image. Accordingly, another strategy could involve the use of optical character recognition (OCR) to extract text embedded in images. The OCR technology can recognize and rewrite characters from images, thus converting visual text formats into accessible text data. By extraction of text elements, OCR helps convert visual information to a structured form corresponding to the goals of the hierarchical representation.It may be provided that creating the textual description includes extracting metadata associated with the non-textual element. Accordingly, the extraction of metadata associated with the non-textual element provides another possibility. Metadata provides extensive information about properties, context, creation, and use of non-textual elements. The use of metadata can provide insight into the role and content attributes of the item, thus supporting the creation of a comprehensive textual description.In summary, the method is focused on offering effective tools for generating precise text representations and extending document processing capabilities for enhanced synthesis and query interactions, whether relying on voice imaging models, applying OCR, or using metadata extraction. These approaches illustrate the translation potential of advanced technologies and methods in converting various non-textual input to structured text that promote improved accessibility within generative language models with query support and support richer interaction with document complexity.It can be provided that the method comprises receiving a query. The method may optionally include embedding the query. The method may include identifying chunk sets relevant to the query. Relevant chunk sets may be identified by comparing the embedded request to stored chunk set embeds and / or by comparing the request to stored chunk sets. The method may include composing a context document from chunks referenced by the identified chunk sets.Accordingly, the method may also include processing a query that enables quick and accurate access to the relevant information stored in the hierarchical representation. Upon receipt of a query, an embedding may be generated that is a vector representation reflecting the semantic and / or contextual features of the query. This query embedding may reflect the process used in embedding chunk sets, where numerical vectors are used to represent these segments in a machine readable format. Matching the query embedding to the chunkset embedding allows direct comparison in a multi-dimensional vector space. The identification process may examine similarities between the embedded query and the stored chunk set embeds. The comparison allows the system to identify chunk sets that have high relevance and semantic closeness to the query. This approach can optimize the search by focusing on the more meaningful and contextually matched chunk sets, thus enabling efficient access to relevant data.Additionally or alternatively, the method may support full text comparisons between the raw query text and the text content of the chunksets, e.g., by keyword matching, fuzzy searching, and / or syntactic similarity scoring. A hybrid approach may also combine both techniques, embedding-based similarity and full text matching, to increase query accuracy. For example, the system may use embeds to confine a selection of chunksets and then make a text-based assessment to refine the results. With this flexible query strategy, the system can efficiently access and classify information in a manner that balances semantic relevance with text specificity.Once relevant chunk sets are identified, the organization of a context document (also called a "chatsheet") may include the placement of chunks that these chunk sets reference in a coherent enumeration. This document summarizes the retrieved information in a format closely associated with the intention of the query, and provides a consolidated overview of related content that supports other generative processes or meets the needs of the user.These operations could improve search and context generation, insert seamlessly into generative language models, and support applications where fast and accurate synthesis of document data may be advantageous.By embedding alignment and referencing chunksets, the method may facilitate interaction with document memories and extend generative capabilities with enriched context inputs.It may be provided that composing the context document comprises extracting chunks referenced by the identified chunk sets, maintaining hierarchical relationships between the extracted chunks, and / or composing the extracted chunks into a coherent document while maintaining structural information such as inclusion.Accordingly, extraction of chunks referenced by the identified chunk sets may facilitate obtaining specific content matched to the semantic focus of the request. This extraction may include accessing stored chunks within the identified chunk sets and isolating them as building blocks for the context document. These chunks capture targeted information matched to the request, thus allowing the creation of a document that addresses the specific needs of the user or model. Considering hierarchical relationships between the extracted chunks helps maintain connections and logic flows originally specified in the hierarchical representation. Hierarchical viewing can improve coherence and contextual association of the extracted contents while maintaining parent-child relationships and depth of detail in document creation. These relationships promote understanding and narrative cohesion and increase the information richness compared to the original input structure. Merging the extracted chunks into a coherent context document may involve careful integration and organization, possibly using structural cues such as indentations to ensure that the document reflects the structured genetics of the source material. Intervention and structural markers direct the visual aspect of the document's hierarchy by demarcate distinctions and organize the flow so that semantic clarity and simple navigation are maintained. This assembly process can improve the synthesis capabilities of the generative language model by facilitating access to a consolidated and contextually enriched document suitable for subsequent processing. Generative tasks use such coherent documents to generate expert, relevant, and precise results that closely resemble the query and its semantic intention.It can be provided that the compiled context document exceeds a context window of the generative language model of the client. In this case, the method may further include prioritizing chunks based on the relevance to the request while maintaining hierarchical integrity.If the compiled context document goes beyond the boundaries of the context window of the customer's generative language model, the method may optionally be adapted by favouring portions based on their relevance to the query to maintain hierarchical integrity. This approach aims to effectively balance contextual depth and content capacity within technical limits. To identify when the context document may exceed the context window, the boundaries and processing capability of the client's generative language model must be understood. Exceeding this window may cause the content of the document to be compressed or refined to fit the work frame of the generative model. Prioritizing chunks based on relevance may include evaluating the meaning and contribution of each extracted chunk to answer the request. Chunks of greater semantic relevance or those directly associated with the key aspects of the query may be preferred to ensure that important content in the context document is placed in the foreground. Scoring or ranking mechanisms may be used in this assessment, assessing semantic alignment or importance of content items. At the same time, maintaining hierarchical integrity may be advantageous to preserve organization logic and the depth of relationship of the original hierarchical representation. Even with prioritization, maintaining hierarchical relationships may promote understanding and continuity and ensure that the context document maintains thematic enumerations and connections corresponding to the original input document. This hybrid process allows the generative language model to access extensive and targeted content that satisfies operational constraints while complying with the requirements of users or system queries. It promotes content efficiency and coherence, emphasizes relevance, and preserves the information structure, which can improve the generative model's ability to precise results.It may be provided that the method further comprises providing the compiled context document and the request to the client generative language model to generate a response.Accordingly, this step combines the interaction between structured data and query dynamics and aims to provide contextual and semantically informed responses or content synthesis. Access to the context document allows the generative language model to access a customized collection of information directly relevant to the query. This document contains prioritized and hierarchically aligned content, which helps the model operate on well organized data suitable for the generative goals. Together, the context document and query form a dual input system that suggests to the generative language model how the pending task may address and aligns its processes to precise and relevant results. The client generative language model may use the given information to navigate through complex meaning, semantics, and context levels and generate a response that matches the goals of the request and accesses the richness of the assembled context document. This process of response generation relies on the prestructured document elements and uses generative techniques to synthesize, analyze, or transform the content in accordance with the desired tasks.It may be provided that each chunk contains reference information indicating its source position in the input document. In this case, the client's generative language model may be instructed to include the reference information when cited from the context document.Accordingly, inclusion of reference information for each chunk could include storing metadata that references the original location in the input document. This may include details such as section headings, paragraph numbers, page numbers, or other markers indicating the original context of the content. These reference details are useful in ensuring that when content is retrieved or synthesized, the origin and location within the original document can be accessed. Giving instructions to the generative language model to include reference information may result in more responsibility and transparency in the results. When content from the context document is cited or referenced, this practice may include ziterlike elements, thereby illustrating the origin of the information and supporting the accuracy and fidelity of the synthesized responses. This process has advantages in applications where documentation standards, validation, or verification are important. For example, academic research, legal documentation, or technical analysis tasks benefit from detailed references that help users track content back to its original extent.It may be provided when it is determined that the query requires information from a particular non-textual element. In this case, the method may include retrieving the original non-textual element based on its reference identifier in the hierarchical representation, and providing the original non-textual element to the customer's generative language model.Accordingly, such a fetch may utilize the reference identifiers already introduced earlier that associate non-textual elements with their representations within the hierarchy. These identifiers serve as access keys that facilitate systematic retrieval from memory while maintaining content fidelity. The concentration of identifiers allows the retrieval mechanism to capture the relevant non-textual data in a nuanced manner. After retrieval, the original non-textual element is provided to the customer's generative language model to increase the content's engagement. By providing this element together with the structured document data, the model can comprehensively answer the request and use the relevant information items. This could improve the ability of the model to synthesize inputs by integrating narrative elements and functional findings from the non-textual data. The inclusion of non-textual elements in responses may provide advantages when visual, audio, or multimedia data provides contextual understanding that is difficult to grasp in text form. The method aims to improve responsiveness by enriching generative applications with a variety of content types and promoting an integral response strategy. This design is intended to improve the ability to answer requests requiring detailed non-textual data and promote integrative document processes while at the same time maintaining the fidelity of interaction across different content dimensions. Overall, this approach aims to optimize document processing operations while focusing on completeness and contextual accuracy in generative tasks, with the aim of proposing an effective bridge between textual and non-textual areas.It may be provided that the method further comprises detecting an update of the input document, identifying portions of the hierarchical representation affected by the update, and selectively regenerating only the affected portions of the hierarchical representation.Accordingly, recognition of updates to the input document may trigger a responsive action in which the affected portions of the hierarchical representation are identified. This identification can be associated with the selective recovery of just these segments, which promotes the efficiency and accuracy of the document management process. The detection of updates to the input document may involve monitoring and detecting changes as they occur, whether by direct modifications, revisions, or the integration of additional data. This recognition process is suitable for techniques such as document comparison algorithms, change tracking systems, or real-time monitoring configurations that derive changes. Once an update is detected, identifying the affected parts of the hierarchical representation may include determining the scope and type of changes. This may include evaluating how the changes affect particular content segments within the hierarchy, concentrating on the locations where differences occur, and tracking the branches in all of the connectable relational nodes. Selective regeneration of the affected parts may transform only the segments affected by the update, thereby saving computational resources and increasing processing efficiency. This selective regeneration process keeps the hierarchical model up-to-date and avoids comprehensive revision if not essential. The fidelity of the content is maintained while the existing organization logic remains unchanged. The ability to selectively regenerate parts enables real-time adjustment so that operation can continue across generative and polling applications without causing disturbances. This methodology may be useful in dynamic document interaction environments where the data continues to develop because it ensures that the processing models match the current document status. Overall, this method aims to propose, by detecting, identifying and selectively regenerating, processes that balance efficiency and precision in document processing, promote adaptive and responsive constraints for handling evolving document landscapes, and offer improved interaction paradigms for advanced generative language applications.It can be provided that the method further comprises the creation of a multi-document hierarchy. The multi-document hierarchy may be created by creating a hierarchical representation for each of a plurality of input documents and creating a higher level hierarchical representation that organizes the collection of hierarchical representations.Accordingly, the method can extend its scope by allowing a multi-document hierarchy in which hierarchical representations are generated for each of a plurality of input documents and a superordinate hierarchical framework is created to organize these individual representations together. This development can facilitate the management and retrieval of interconnected data across different document sources while providing more structure. The generation of hierarchical representations for each input document follows processes similar to those described with the emphasis on the word-true rendering and organized structure for each individual document. Each hierarchical model reflects the depth, organization, and contextual features inherent in the content and provides a systematic framework for interaction with the document data. After completion of the individual representations, these models are integrated into a comprehensive and uniform document structure by creating a superordinate hierarchical representation. This overlapping hierarchy may serve as a taxonomy or schema that organizes contexts and common topics or categories that overlap the span of all documents involved. This integration allows the use of inter-document semantic relationships that allow insight into the context of topics or topics from different sources. Techniques are employed to detect and characterize content intersections, thematic continuitys, or categorial partitions and form a meta-structure that mediates coherence and coherence. Such multi-document hierarchies could prove advantageous in complex information ecosystems as they allow efficient searching and synthesis across multiple documents. By organizing information within this superordinate hierarchy, models can dynamically interact with comprehensive content collections, enabling generative voice tasks to access contextual data while maintaining integrity or precision. In summary, creating a multi-document hierarchy provides a robust framework for the management, retrieval, and synthesis of comprehensive document datasets that optimizes generative model interaction by merging structured knowledge from multiple sources for enriched content generation. This approach emphasizes the networked representation of documents and the clarity of retrieval and supports shaded and coherent handling of multivariate documents.Another aspect of the present disclosure relates to a data processing system, apparatus, or apparatus. The data processing system, apparatus or apparatus may comprise means for performing any combination of steps of the methods described. The data processing system, apparatus, or device may include a memory and a processor. The memory may store instructions or be configured to store instructions that, when executed by the processor, configure the data processing system, apparatus, or device to perform any combination of steps of the described methods.Another aspect of the present disclosure relates to a computing device having a processor and a memory. The memory stores instructions that, when executed by the processor, configure the apparatus to receive an input document, the input document including content items, perform a lossless hierarchical text rendering process in which the input document is reproduced as a hierarchical representation in which all content items of the input document are word-randomized and organized according to a structure of the input document, and perform a fetch unit generation process in which fetch units are generated based on the hierarchical representation and stored in a repository accessible to a generative client language model for fetch-extended generation.Another aspect of the present disclosure relates to a computer program. Another aspect of the present disclosure relates to a non-transitory computer readable medium having a computer program stored thereon. In either case, the computer program may include instructions which, when the program is executed by a computer such as the mentioned data processing system, apparatus or apparatus, cause the computer to perform any combination of steps of the described methods. A computer program can also be referred to as a program, software, software application, app, module, software module, script or code. A computer program may be written in a programming language, including compiled or interpreted languages. A computer program may be deployed in any form, including as a stand-alone product or as a module, component, subroutine, or other device suitable for use in a computing environment, such as in the recited data processing system, apparatus, or device.Another aspect of the present disclosure relates to a computer readable medium having computer executable instructions stored thereon for implementing a method comprising receiving an input document, the input document comprising content items, performing a lossless hierarchical text rendering process in which the input document is reproduced as a hierarchical representation in which all content items are derived from the input document wordwise and organized according to a structure of the input document, and performing a fetch unit generation process in which fetch units are generated based on the hierarchical representation and stored in a repository accessible by a generative client language model for fetch-extended generation.Another aspect of the present disclosure relates to a server for providing call extended generation services. The server may include one or more processors and a memory. The memory may store instructions that, when executed by the one or more processors, cause the server to perform any combination of steps of the described methods.Another aspect of the present disclosure relates to an application programming interface (API) for providing call-extended generation functionality. The API may include a document entry endpoint configured to perform any combination of steps of the described methods.The aspects and features described above and elsewhere in this disclosure, when implemented together or individually, may produce several beneficial effects, including, without limitation:• The overflow of contexts in large speech models can be reduced.• Hallucinations can be attenuated by preserving the original text.• Large documents can be restructured into a hierarchical organization without omitting contents.• Structured data, such as tables, can be converted to textual tree formats in which each data cell is maintained.• Script-based transformations can be used for extensive tabular data.• Images, code chips and other non-text media may be integrated via textual representations for uniform multimodal indexing.• Chunksets may be formed that reflect parent-child relationships within the organization to optimize the use of tokens.• The dynamic compilation of relevant hierarchical segments can improve the generation of search results to answer queries.• In the hierarchical chunks, the exact phrase fidelity can be maintained, enabling precise citing of the source.• Incremental updates of the source documents may be selectively made to obtain the hierarchical representation.• Multiple documents may be organized in a multi-level hierarchy for corpus-wide retrieval.• Each chunk set may maintain references to its origin location to allow transparent mapping in the generated responses.The terms used herein are generally understood as understood by one of ordinary skill in the art, unless expressly stated otherwise. The following explanations can facilitate understanding:The term "artificial intelligence" (AI) refers to a branch of informative technology that aims to develop machines or software capable of intelligent behavior, typically with the aim of mimicking or exceeding human intelligence in certain tasks. AI systems are designed to perform complex tasks such as logical thinking, learning, perception, problem solving, and natural language understanding. These systems can typically adapt to new situations and improve their performance over time. The aim of the AI is to provide systems that operate autonomously and can interact with their environment in a human-like manner.The term "artificial neural network" (ANN) or, for short, "neural network" (NN) can be understood as a machine learning model or deep learning model or algorithm. Neural networks are generally modeled on the human brain and typically consist of interconnected nodes or neurons organized into layers. Neural networks may be used for data processing and learning from examples so that they can perform tasks such as image recognition, natural language processing, and many more. A neural network typically consists of an input layer, one or more hidden layers, and an output layer. Through a process referred to as training, neural networks may learn to perform certain tasks by adjusting their internal parameters or "weights" based on labeled or unlabeled data.The term "chunk" may be understood as a segmented portion of a content item or a combination of segmented portions of one or more content items created by a chunking process. A chunk typically represents a discrete unit of information derived from an input document to facilitate efficient searching and processing within an RAG system.The term "chunk set" may be understood as a collection of one or more chunks that represent a coherent portion of the hierarchical representation. A chunk set may indicate parent-child relationships. Examples are chunks grouped by header levels and provided with corresponding sub-sections, or merged text blocks that fit within a particular token boundary while containing the required context.The term "content item" may be understood as a discrete unit of information within an input document representing a semantically or structurally distinct portion of the data. A content item may be made of various types of data including, but not limited to, text, numeric data, images, audio segments, video segments, or any combination thereof.The term "context document", also called "chatsheet", may be understood as a document, preferably a text document, composed of relevant query units or chunksets, wherein structural or contextual indications are maintained to help a generative machine learning model accurately respond to a query.The term "context size" as well as variations such as "context limit", "context limit", "token limit", and similar terms, known to those skilled in the art, may be understood as the maximum number of tokens or amount of data that a generative machine learning model, such as a generative speech model or LLM, may process or maintain in its memory-like context window at once. The term "token limit" may refer to, for example, an upper limit on the number of text units (e.g., subword units or tokens) included in a chunk or chunk set to remain within the processing capabilities of a model.The term "generative language model", also referred to as a "language model", may be understood as a type of machine learning model configured for natural language processing tasks, such as language generation. Generative speech models typically have a very large number of parameters and are trained on a large amount of text.In the present disclosure, the term "generative language model" or similar terms is used as a synonym for "large language model" (LLM), unless expressly stated otherwise. An LLM is a type of machine learning model that has been trained to recognize, generate, translate, and / or summarize large amounts of written human speech and text data. LLMs are characterized by their ability to generate general purpose speech. LLMs include a large number of parameters, typically millions or often billions of parameters, that allow them to grasp a wide range of linguistic nuances, patterns, and contextsAlthough (big) language models (also referred to as generative language models) are described throughout this disclosure, it will be appreciated that the disclosed concepts may likewise be practiced through the use of other similar types of machine learning models, such as (big) action models, to name just one example. As a result, all features disclosed herein should be understood to be applicable to any type of machine learning model, unless a particular type of model is required for the particular feature, as will be appreciated from the context.The term "hierarchical representation", also called "hierarchical organization", can be understood as a structured format in which the content elements of the input document are arranged in interleaved levels according to their relative position within the input document, each text element being maintained in word form.The term "engagement level" may be understood as a means for visually or symbolically representing the hierarchical depth by surrounding, in particular prepending, the content with different degrees of engagement, thus organizing the data in a structured manner. Examples of specific engagement implementations are increasing tabs or spaces.The term "input document" may be understood as any digital, digitized, or otherwise machine readable data structure or data sequence, typically in the form of a file or collection of files. Examples of input documents include text files, PDF documents, word processing files, presentation sheets, spreadsheets, scanned documents, HTML documents, XML documents, JSON documents, image files, audio files, video files, database entries, and combinations thereof, which serve as a source of information for processing by the described methods. An input document may be structured or unstructured and may originate from local storage, network resources, or real-time data streams.The term "lossless hierarchical text rendering method" can be understood as a computer-implemented method in which the text elements, preferably the entire content elements, of the input document are reproduced and restructured into a hierarchical representation without omitting or combining content, with all text elements being retained in the word by word.The term "machine learning" may be understood as a portion of artificial intelligence that focuses on the development of algorithms and statistical models that enable computers to perform certain tasks without explicit instructions. Instead, machine learning systems learn and make predictions or decisions based on data. Machine learning algorithms build a mathematical model based on example data referred to as training data to make predictions or decisions without being explicitly programmed to perform the task. Machine learning can be used in a variety of applications, such as image and speech recognition, medical diagnosis, predictive analysis, and many other areas where systems allow learning from and self-adapt to new data.The term "machine learning algorithm" may be understood as a computational method that analyzes, learns from, and recognizes patterns or makes decisions based on the input data without being expressly programmed for this task. Machine learning algorithms use statistical methods to enable systems to improve their performance over time with more data for a particular task. Machine learning algorithms are the basis on which machine learning models are built by providing the methods or processes that convert data into useful findings. Examples of machine learning algorithms include linear regression, decision trees, support vector machines, and neural networks.The term "machine learning model" refers to the output generated when a machine learning algorithm is trained on a dataset. It represents the knowledge or understanding that the algorithm has gained from the data by encapsulating the learned patterns or predictions. In essence, a machine learning model allows predictions or decisions based on new unseen data based on the knowledge obtained from the training process. The machine learning model is typically defined by its parameters that can be adjusted during the training phase to minimize the difference between the predicted result and the actual result. Although "machine learning algorithm" and "machine learning model" have different definitions, strictly speaking, it is not uncommon for these terms to be used interchangeably in the context language. This use arises from the close relationship between algorithms and models in the workflow of machine learning projects, where the algorithm is the means for creating the model. Therefore, these terms may be used interchangeably herein unless the distinction is critical.The term "multi-modal content integration process" may be understood as a computer-implemented method for identifying and incorporating non-textual elements such as images, audio, or video into the hierarchical representation by associating a textual counterpart at a suitable structure point.The term "natural language processing" (NLP) refers to a field of informative and artificial intelligence that focuses on enabling computers to understand, interpret, and / or manipulate human speech. In this case, computer training is generally combined with statistical, machine learning and deep learning models in order to process human speech in the form of text or speech data, so that computers can understand the intention and mood of the speaker or writer. NLP typically includes tasks such as text and speech processing, natural language understanding, and text analysis. There are various applications such as machine translation, speech recognition and chatbotes for customer service, just to name a few.The term "non-text element" may be understood as a particular type of content element that does not consist primarily of human readable speech characters, words, or symbols. Non-textual elements may include, but are not limited to, images, audio segments, video segments, numerical data, graphical representations, and metadata. Non-textual elements may be embedded in or associated with a text element within an input document.The term "query" may be understood as a query for information, typically formulated as a natural language prompt, that serves to retrieve relevant content from a data source within an RAG process.The term "repository" or even storage may be understood as a data storage system configured to store and manage data such as content items, chunks, chunk sets, hierarchical representations, context documents, or representations thereof for later retrieval. An example of a repository is a vector database configured to store vector embeds and enable query based on semantic similarity using vector distance metrics.The term "query unit" can be understood as discrete segments, chunks or chunk sets of the hierarchical representation. A query unit may maintain references to the original content, but may be dimensioned and organized for efficient searching through a generative language model. Examples of this are individual paragraphs in which the parent-child relationships are maintained, groups of enumeration points grouped by headers, or combined text blocks that fit within a predefined token limit.The term "method of generating retrieval units" may be understood as a computer-implemented method that systematically generates modular units of information or "retrieval units" derived from a hierarchical representation in a manner that facilitates later retrieval and recombination. Example implementations may include code that divides an outlined document, i.e., a hierarchical representation of an input document, into multiple smaller segments ("chunks") in consideration of depth and / or size constraints.The term "retrieval-augmented generation" (RAG) may be understood as a technique that enables artificial intelligence generative models to retrieve and include new information. RAG modifies the interaction with a large generative language model (LLM) so that the model responds to queries with reference to a particular set of documents and uses this information to supplement information from its already present training data. This allows LLMs to use domain specific and / or updated information.The term "table integration process" may be understood as a computer-implemented method for converting table data from an input document into a format suitable for hierarchical representation, in which the original logical layout or data relationships are maintained.The term "tablet ree representation" or "table tree representation" can be understood as a special format which serves to acquire tabular data as nested text elements, wherein the relationships between rows and columns are maintained.The term "tabular data" may be understood to mean a particular type of content containing data organized into rows and columns. Each row typically represents a record and each column represents a particular attribute or field associated with that record. Tabular data can be presented in various formats including, but not limited to comma separated values (CSV), tab separated values (TSV), spreadsheets, database tables, and data frames. Tabular data, in turn, may include text elements, non-text elements, or a combination thereof.The term "text element" may be understood as a particular type of content element comprising a sequence of characters, words or symbols representing a human readable language. A text element may be a single word, phrase, sentence, paragraph, or other contiguous or non-contiguous portion of text data extracted from an input document.The term "training" may be understood as the process of contributing to a machine learning model to make predictions or decisions by subjecting it to data whose results are known. In the training process, a training data set is generally fed into a machine learning algorithm, which then learns the patterns or relationships in the data using statistical analyses. During training, the algorithm iteratively adjusts the parameters of the model to minimize the difference between the predicted results and the actual results in the training data. This adaptation process is usually controlled by a loss function that measures the accuracy of the predictions of the model. The aim of training is to create a model which accurately reflects the underlying structure of the data, so that it can make reliable predictions about new, unseen data. During supervised learning, a model is trained on a tagged dataset, with each example in the training data being paired with the correct output. The model learns to predict the output from the input data. In unsupervised learning, a model is trained on data without marked responses. The model attempts to independently find patterns and relationships in the data. Semi-supervised learning combines both tagged and unlabeled data during the training process, which may be advantageous when obtaining a fully tagged dataset is costly or impractical.The term "literal" means that the original text, the font, and other textual or symbolic elements have been adjusted unchanged or minimally without altering the essential content, e.g., by removing spaces, correcting apparent encoding errors, or normalizing line breaks, while preserving the essential verbatim of the original text.Particular and preferred aspects of the present disclosure are set forth in the appended independent and dependent claims. Features from the dependent claims may be combined with features from the independent claims and with features from other dependent claims, as appropriate and not just as expressly stated in the claims.The above and other features, characteristics, and advantages of the present disclosure will become apparent from the following detailed description when taken in conjunction with the accompanying drawings, which illustrate, by way of example, the principles of the disclosure. This description is for illustrative purposes only, without limiting the scope of the disclosure.BRIEF DESCRIPTION OF THE DRAWINGSThe present disclosure will be better understood by reference to the following drawings: FIG. 1 shows a schematic overview of a processing pipeline providing advanced hierarchical chunking, multi-modal integration, and fetch optimization, according to one embodiment. FIG. 2 illustrates a lossless hierarchical text rendering method according to an embodiment. FIG. 3 illustrates a table integration process according to an embodiment. FIG. 4 illustrates a process for multi-modal content integration, according to an embodiment. FIG. 5 illustrates a process for generating a fetch unit, according to an embodiment. FIG. 6 illustrates a process of composing context documents according to an embodiment. FIG. 7A illustrates a detailed example implementation, in accordance with an embodiment. FIG. 7B shows a continuation of the detailed example implementation of FIG. 7A. FIG. 8 illustrates an example of an input document according to an embodiment. FIG. 9A shows a first chunk created from the input document in FIG. 8 with LamaCloud. FIG. 9B shows a second chunk created from the input document in FIG. 8 with LamaCloud. FIG. 9C shows a third chunk created from the input document in FIG. 8 with LamaCloud. FIG. 9D shows a fourth chunk created from the input document in FIG. 8 with LamaCloud. FIG. 9E shows a fifth chunk created from the input document in FIG. 8 with LamaCloud. FIG. 9F shows a sixth chunk created from the input document in FIG. 8 with LamaCloud. FIG. 10 shows a hierarchical representation created for the example input document of FIG. 8, in accordance with one embodiment. FIG. 11A illustrates a first chunk set created for the example document of FIG. 8, in accordance with one embodiment. FIG. 11B illustrates a second chunk set created for the example document of FIG. 8, in accordance with one embodiment. FIG. 11C illustrates a third chunk set created for the example document in FIG. 8, according to one embodiment. FIG. 12 shows a chat sheet created for the example input document in FIG. 8, according to one embodiment. FIG. 13 shows a schematic block diagram of the computer hardware on which embodiments of the present disclosure may be implemented.DETAILED DESCRIPTIONHereinafter, representative embodiments illustrated in the accompanying drawings will be explained. It should be understood that the depicted embodiments and the following descriptions relate to examples that are not intended to limit the embodiments to a preferred embodiment.Conventional RAG pipelines typically suffer from fragmented contexts, semantic incoherence, token inefficiencies, hallucinations, and impaired data fidelity. Summary-based hierarchical methods (e.g., RAPTOR) result in abstraction inaccuracies and the loss of important details. Existing chunking solutions, such as zCh, provide flat segmentation without exploiting a full hierarchical semantic understanding.Various embodiments provide a solution to these deficiencies by providing a lossless, multimodal compatible and conformal augmentation methodology that provides absolute data fidelity for precision critical applications.OverviewFIG. 1 shows a schematic overview of a processing pipeline 100 that provides advanced hierarchical chunking, multi-mode integration, and fetch optimization, according to an example embodiment.In the illustrated embodiment, the processing pipeline 100 includes a document preparation unit 102 that is responsible for preparing an input document 104 so that its content can be efficiently utilized in later query processing. The input document 104 includes various content items. These content items may include text, tables, images, and other multimedia items.The document preparation unit 102 includes a lossless hierarchical text rendering process 106. In one implementation, lossless hierarchical text rendering process 106 uses explicit prompts that instruct LLMs to word-wise reproduce an input document 104 in a hierarchical representation 108 by incremental intervention, using implicit semantic understanding. Standard length input documents 104 may undergo hierarchical inclusion in a single pass, while overlong input documents 104 may use iterative, multi-level inclusion that ensures lossless hierarchical continuity with explicit delimiters.The document preparation unit 102 includes a table integration process 114. In one implementation, the table integration process 114 converts complex table data into structured hierarchical text representations, maintaining full semantic integrity and token efficiency, and generating deep hierarchical table trees. Complex tables of normal length can be processed by direct hierarchical inclusion by LLMs. Extra-long tables can be processed by LLM-generated scripts (e.g., python or pandas), which can be executed externally and generate flat, structured table trees that optimize token efficiency and fidelity.The document preparation unit 102 includes a multimodal content integration process 116. In one implementation, the multimodal content integration process 116 generates predictive, uninterlaced text place holders from multimodal content (e.g., images, code chips, formulas) via special LLM prompts. The original multimodal content may be stored, maintaining absolute fidelity and allowing optional dynamic reintegration on demand.The document preparation unit 102 includes a process for generating retrieval units 110. In one implementation, the process of generating search units 110 algorithmically generates specialized search units 112 ("chunksets") representing particular traversal paths through the hierarchical representation 108 of the input document 104, combining parent-child relationships across depth planes according to token boundaries. In one example implementation, a multi-level algorithm building process is performed that includes (1) sorting chunks by depth and location, (2) collapsing adjacent chunks at the same depth level within the token boundaries, (3) recursively preceding parent parent chunks to form complete traversal paths, and (4) building chunk sets as unique combinations that maintain hierarchical integrity while optimizing token utilization. One implementation provides for inclusion at chunkset level rather than at the individual chunks level to enable semantically complete retrieval. The chunk sets may be stored with their embeddings in a repository, each chunk set containing pointers to the associated chunks, rather than duplexing the content. This allows retrieval at semantically reasonable levels that preserve context by matching with chunk sets rather than isolated chunks.In the illustrated embodiment, the processing pipeline 100 also includes a query processing unit 118 that is responsible for processing queries 120 with respect to the content of the input document 104.The request processing unit 118 includes a process for composing context documents 122. In one implementation, the context document assembly process 122 dynamically assembles predictive, contextually optimized, and minimally redundant context documents ("Cheatsheets") from query units 112 (chunksets) per query input document pair. Optionally, the exact original multimodal elements may be reintegrated based on explicit retrieval criteria such as user preference, precision requirements, token availability or semantic / visual differences.The embodiment illustrated in FIG. 1 advantageously combines several unique features of the present disclosure to provide important advantages, including:• Explicit use of the inherent incremental hierarchical understanding of LLMs by structured literal rendering.• Lossless iterative processing method for large documents.• Novel Transformation of Tabular Data with Two Methods that Optimize Semantic Complexity and Token Efficiency.• Uniform handling of multimodal contents with external preservation of the original fidelity of the content.• Innovative hierarchical chunk sets enabling mathematically meaningful retrieval combinations.• Dynamic, query-specific chatsheet compilation with optional multimodal reintegration of contents for improved query accuracy.While a complete embodiment combining various unique features and specific implementations of the present disclosure has been described in connection with FIG. 1, it should be appreciated that various other embodiments may implement only certain subsets of these features or implementations. For example, one embodiment may provide a processing pipeline that includes only the document preparation unit 102 while the query processing unit 118 is part of another processing pipeline that may be operated by another unit. Other processing pipelines are also conceivable that provide a document preparation unit 102 that includes only lossless hierarchical text rendering process 106, only multimodal content integration process 116, only table integration process 114, or sub-combinations of these three processes.The individual processes and units illustrated in FIG. 1 are described in more detail below:Hierarchical Verbatim Organization of the TextFIG. 2 illustrates an embodiment of a lossless hierarchical text rendering process 106. In step 202, the input document 104 is reproduced as a hierarchical representation 108 ("hierarchical organization"). The hierarchical representation 108 is a version of the input document 104 in which all content items of the input document 104 are retained word by word and organized according to the structure of the input document 104.In one implementation, the hierarchical representation 108 is generated by an LLM. The LLM is expressly instructed to render the input document 104 in an engaged organization format (e.g., using headers, intermediate headers, or enumeration points to identify the hierarchy) without merging or exiting the content. Each engagement level represents an interleaved portion or subsection of the input document 104 so that the text is organized according to its text depth rather than an arbitrary number of tokens.The actual full text of each section is included (if a section is too big to be included in an LLM response, the process may continue with sophisticated compressed subsequent prompts and / or sections so that nothing is lost). This results in a losslessly edited hierarchical representation 108.For example, a 20 page header document could be output as hierarchical representation 108 where the top level entries are the main headers, followed by indented sub-points containing exactly the paragraphs or sets of the original, with the verbatim and even formatting hints (citations, lists, etc.) maintained as needed. A specific example of an input document 104 is shown in FIG. 8, and a specific example of a hierarchical representation 108 is shown in FIG. 10, discussed further below.If the input document 104 exceeds the output limit of the LLM, the process may split the task into successive rounds of request, e.g., to create the organization to a particular point and then proceed therefrom, ensuring continuity of the hierarchical structure across the split.The prompts may be designed to maintain the structure (e.g., "Their task is to continue the organization after point X while maintaining consistent engagement..."). By explicitly exploiting the ability of the LLM to track instructions and copy text, a conformal hierarchical segmentation of the entire input document 104 can be achieved.In various embodiments, a special prompt method may be used that causes the model to output the document in a structured form rather than free form text.In various embodiments, the actual prompt may contain more rules, but essentially the model is expressly instructed to maintain all text and use the intervention to characterize the inherent structure of the document.If the document is short enough, a single prompt is sufficient to fully outline the text. For longer documents, the process may be iterative: the model may be prompted to outline the first part and then proceed from a particular delimiter or marker for the next part. The prompts ensure that the organization of the second part continues at the location at which the first part has ceased and that the engagement levels remain consistent. This is basically a form of context window management in which we split the organization task into parts that the model can process but design so that the output is contiguous and complete. The result of this is a organization that may be as large as needed (and extends across multiple LLM outputs if needed), but logically is a single document representation.This iterative process may also include a hierarchically consistent compression technique, as discussed above.In various embodiments, unlike the summary, each set of the original input document 104 appears anywhere in the hierarchical representation 108. Unlike naive chunking, no arbitrary cuts are made, but is only interrupted at natural boundaries (e.g., at the end of a section). We call this hierarchical text editing because the LLM virtually outputs the output text of the input document 104 again, but in a structured manner.The structure in hierarchical representation 108 may follow the existing headers in input document 104 or, for unstructured input documents 104, the best conclusion of the LLM for logical grouping. Even if the LLM needs to derive a structure (e.g., if the input document 104 does not contain any headers, paragraphs, or only a single long line of text), in certain embodiments it is instructed to perform oversegmentation (more intervention) rather than undersegmentation. In this manner, the hierarchical representation 108 captures an implicit structure (e.g., a paragraph followed by an engaged list in the source would be represented as such in the organization).In the preferred embodiment discussed herein, indentations (e.g., tabs or spaces) are used for the hierarchy, as this is simple and readable by humans. However, other notations (e.g., XML tags, JSON tree, etc.) may be used to achieve similar results, with the concept of the hierarchical organization remaining the same.Generative Conversion of Tabular Data to "TableTre"FIG. 3 shows an embodiment of a table integration process 114. For portions of the input document 104 that contain tabular data (e.g., spreadsheets, database digests, etc.), the process provides a way to integrate them into the hierarchical representation 108 without losing its relational structure.To this end, the method identifies tabular data in the input document 104 in step 302. In step 304, the identified tabular data is converted to a tablet representation. A tablet ree representation, or "tablet ree" for short, is a tree structured text representation of the table data that maintains all the data of the table in a hierarchical list form.In one implementation, referred to as code generation mode, the process requests an LLM to generate an executable script or code (e.g., python code using a library such as Pandas) that converts the table data to the tablet representation.The script or code, when executed, outputs the table data as an engaged list in which, for example, each row is an uppermost level enumeration character and each cell in that row is listed below (e.g., as a "column name: value" undercount character). Alternatively, each column could be a branch with rows as leaves; the exact structure can be chosen to best reflect the logical organization of the table data. An exemplary output of an LLM-generated script or code is as follows:Series 1:Column A: Value1Column B: Value2Column C: Value3Series 2:Column A: Value4Column B: Value5Column C: Value6The key is that the tablet ree maintains each data cell in a logical grouping, but eliminates excess text and compresses the representation to be token efficient. For example, a 10×10 table could be represented as 10 top level elements (one per row), each with 10 sub-elements of "column:value", thereby preserving all data, but in a linearized tree form. If the table originates from a spreadsheet containing formulas, the tablet may incorporate these formulas for accurate reconstruction along with the resulting cell valuesThe creation of the conversion script by the LLM may use the understanding of the LLM for the content and structure of the table data to ensure that the format is both accurate and optimized. The use of code to perform the conversion guarantees that there are no halluccinated values, since the LLM writes a general recipe and the actual data filling in the table tree comes directly from the execution of the script on the original table. This is in contrast to the naive request to the LLM to write the table in Prosa, which can lead to errors or simply to an overflow of the context window for long input tables. An advantage of the code generation mode is that it ensures correctness (the code literally takes over each cell value) and efficiency (the LLM does not itself potentially count up thousands of cells, but delegates to the code). Moreover, this approach guarantees that if the table has a certain value (e.g., a certain number in cell X), that value will appear anywhere below the table node in hierarchical representation 108. Thus, if the user later queries this value, the retrieval system can place that entry in the context of the LLM.For smaller or moderately large table data, an alternative implementation is a so-called direct organization mode that includes requesting the LLM to output the table directly in an organization format (where the LLM essentially itself acts as a converter), resulting in a deep table tree representation that may include hierarchical grouping for table headers and the likeIn certain embodiments where the model has vision or table understanding capabilities, it is even possible to input an image of the table or CSV that can analyze the model. In various embodiments, the direct organization mode results in a deep table tree, possibly with multiple engagement levels (e.g., table>span>line>cell).The process may select between the direct division mode (direct LLM conversion) and the code generation mode (LLM generated code), depending on the size and / or complexity of the table data. For a simple table, the LLM can list its content directly in a hierarchy; for a very large table that would inflate to enormous sharing, the code approach is used to compress it (e.g., by grouping repeated patterns together or dividing the table into parts).In various embodiments, colspan and / or rowspan (connected cells) tables may be supported. The hierarchical representation 108 or the table tree representation may represent it by repeating the content in each relevant sub-entry or by structuring the hierarchy such that aggregated cells become parents of the values that fall. In this way, it is ensured that complex table structures (e.g., a table with grouped rows under a category) are represented in a manner that the LLM can still accurately interpret (e.g., a cell extending over three columns could be represented either as a higher node with three sub-values that add up somewhat - or as repeating information, a key distinction that only the AI can provide). The general principle is that the semantics of the table (relationships between cells) are preserved, not just the data.In various embodiments, tables may be compressed by cropping or sharing extremely large tables, but only if an alternative representation (such as a statistical summary) is also stored. In the truly lossless mode, no clipping is done; the entire table is always present in any text form.In various embodiments, regardless of whether it was generated directly or via a code, the generated tablet ree representation is then inserted into the hierarchical overall representation 108 at the appropriate location in step 306 (wherein the original table is replaced by this text structure).As a result, any numeric or textual entry from the tabular data is present in the hierarchical representation, so that no information is lost.Integration of Multimodal ContentsFIG. 4 shows an embodiment of a multi-modal content integration process 116. For portions of the input document 104 that contain non-textual elements (such as images, charts / graphics, or embedded code chips), the process provides a way to integrate them into the hierarchical representation 108 without losing its information.To this end, the method identifies a non-textual element in the input document 104 in step 402. In step 404, the identified non-textual element is converted to a textual representation. Such a textual representation may be implemented as a place holder (e.g., "[Image X: Description]") or as a descriptive node to ensure that it is presented in the context. In step 406, the generated text representation is then inserted into the hierarchical representation 108 at the appropriate location.In one implementation of step 404, the process invokes an OCR or annotation tool for each non-textual element (e.g., image) to obtain descriptive text, which is then passed to a text-only LLM instructing it to be appropriately incorporated. In another implementation of step 404, for each non-textual element (e.g., image) it encounters, the process invokes an LLM with image processing capabilities or an image-to-text model to create a predictive textual description of the image (and / or extracts any text present in the image via OCR).In various embodiments, the resulting description is then placed in the hierarchical representation 108 under an "image:" node at the appropriate location (where text is not indented or labeled to serve as a place holder and not as a continuation of normal text content). The description should uniquely identify the image (e.g., "FIG. 2 is a bar graph showing the quarter offsets") that is uniquely identified in hierarchical representation 108 (e.g., the row begins with "image:" or a particular token) so that it can be later determined that that node corresponds to an image.The actual non-textual element (e.g., the image file) may be stored separately (and may be referenced via the file name or an ID in the organization).Similarly, a self-contained code snippet, as another example of a non-textual element (e.g., a JSON example or a pseudo code block), may be either word-wise included in the hierarchical representation 108 (maintaining formatting in a code block) or replaced with a wildcard summary (e.g., "[code snippet: functionality XYZ]"), and the original code stored separately.A similar approach may be used for other media such as audio or video. These types of contents could be transcribed (or merged) and incorporated into the text, and the files stored.This is intended to ensure that, when the hierarchical representation 108 is used later as an LLM context, the model knows at least the content of these non-textual elements in descriptive form. Indeed, the hierarchical representation 108 becomes multimodal: it may include natural language text, structured representations of tables and descriptive text for images, images and the like. Such a multi-faceted "sheet" (cheatsheet) conformally encapsulates the original input document 104.Although the hierarchical representation 108 of the preferred embodiments actually becomes multimodal, it contains everything in text form. The advantage is that the hierarchical representation 108 can be treated uniformly during the query, e.g., the description of an image is indexed and can be retrieved when relevant to a query. If a user's question can be answered by an image (e.g., "What illustrates the diagram about system architecture?"), the system can retrieve the descriptive node for that image. The LLM may use the description to answer the question, or the system may decide to retrieve the actual image (as it has the reference) and either present it to the user or use an image model to obtain further details.In other words, embryos essentially create a text surrogate for each modality and bind it in advance. This can be viewed as a form of data extension that makes images searchable by text. In addition, the knowledge base is made future safe: If an advanced multimodal model is later available, the stored original images can be used, but if only text models are used we still have a useful representation of this content.Hierarchical Context Mounting for RAGFIG. 5 illustrates an embodiment of a process for generating query units 110. The goal is to process the hierarchical representation 108 to create chunk sets, i.e., specialized interrogators 112 that acquire meaningful traversal paths through the hierarchical structure of the input document 104.Conventional RAG systems typically operate simply with individual chunks corresponding to individual nodes in hierarchical representation 108, such as headers, paragraphs, table entries, and the like. In contrast, in certain embodiments of the query unit generation process 110, combinations of chunks, so-called chunk sets, are systematically generated. Each chunk set consists of a combination of chunks that follow parent-child relationships across depth planes. Each chunk set represents a specific traversal path from higher to lower depths in the hierarchy of the input document 104. In a chunk set, the chunks may be collapsed and / or combined according to the token boundaries. These chunk sets, not individual chunks, become the primary units for embedding and retrieval, i.e., the chunk sets are vectorized (embedded) and stored in the repository.The process of chunk set generation 110 according to the embodiment illustrated in Figure 5 begins by parsing the rendered text to determine all individual text elements (which form the future chunks) including their respective depth (by simply counting the leading tabs / spaces) and sorting the chunks by depth and position in step 502. In step 504, adjacent chunks of the same depth plane that are within the token boundaries are collapsed. In step 506, parent chunks are recursively advanced to form complete traversal paths. In step 508, unique combinations are created that maintain hierarchical integrity while optimizing the use of tokens. The resulting chunksets, along with their embeddings, may be stored in a knowledge store (e.g., a vector database or text search index).In certain embodiments, each chunk set 112 maintains references to the associated chunks rather than duplexing content, thereby providing an efficient retrieval mechanism that maintains the hierarchical structure of the input document 104.Context Document ("Chatsheet") Compilation & Query ProcessingFIG. 6 illustrates an embodiment of a process for composing context documents 122.In various embodiments, when responding to a query 120, the processing pipeline adjusts the chunk sets 112 rather than individual chunks. This ensures that the retrieval takes place at semantically meaningful levels of the hierarchy. For example, a query 120 could match a chunk set 112 that contains both a particular detail and its contextual header. The pipeline can then extract all chunk IDs from matching chunk sets and compile them into a coherent context document 124, a so-called "chatsheet", for the customer's generative language model 608, maintaining the correct hierarchical relationships.Since the ranking points contain word-random text from the original, the response of the LLM can cite exact phrases directly, improving proper accuracy and enabling precise citations. The hierarchical representation 108 makes the mapping simple because each node in the hierarchy carries an identifier that returns to its source document and section. In this way, any response generated by the LLM can be tracked back to its exact origin.In the embodiment depicted in FIG. 6, at query time, the process optionally embeds the query 120 in step 602, evaluates which chunk sets 112 best match the query 120 (either by matching the query embedding to the chunk set embeddings or by matching the raw query text to the textual content of the chunk sets), and dynamically assembles the most relevant chunks into a customized context document 124 (chat sheet) in step 606. If all the assembled content fits within the context window of the LLM, the system may concatenate the chunks in their original structural order. If too many, the system may prioritize the most relevant chunks, maintaining hierarchical integrity.The resulting chatsheet 124 is a composite document that contains (or consists of) exactly the portions of the original organization that best answer the request 120. It could contain rows such as "Section 3: Data Retention Policies - the company challenge retain data for X years..." followed by relevant subsections or image description if appropriate. This composed context 124 may then be fed into the LLM (the client's generative language model 608) along with query 120, as shown in FIG. 6. To this end, a prompt template such as the following example may be used: Answer the question from the information below...This approach ensures that the customer's generative language model 608 accurately obtains the required information with minimal redundancy, while maintaining the full hierarchical context. Even when specific details are retrieved, the proper context (e.g., parent nodes such as section headings) is automatically incorporated. Since the chat sheet 124 resorts to exactly stored membership elements, it remains in the source input document 104 tree.The hierarchical representation 108 enables a solid source assignment. Each chunk contains its reference information (e.g., "[Doc1 §3.2]"), and the LLM can be instructed to include these references in the case of a tick, which increases confidence and verification. This interrogation method is clearly better than flat approaches in which chunks are arbitrary sections. Here, chunks are logical segments whose hierarchical context is intact.Optionally, as part of this embodiment, the pipeline may reintegrate original images or other media if query 120 expressly requires them and query LLM 608 can process them. This sophisticated query strategy allows efficient retrieval of even very large knowledge databases and functions as a dynamic navigation system through a knowledge tree that exactly extracts the information needed to answer the query. If re-assembly of the original non-textual elements is desired, such as when a query requires replacement of the image with a textual description in the hierarchical representation 108, the stored image may be retrieved and forwarded to an imaging model, or the exact code snippet may be retrieved. However, in certain embodiments, the kernel representation used for retrieval is the textual hierarchical representation 108 with place holders, as this is what a standard LLM can consume. Since the descriptions are generated via LLM or KI, care is taken that they are accurate (the prompts for image description focus on the objective detail accuracy of the image). This uniform handling of modalities is different from conventional approaches, each modality standing on its own.In various embodiments, at queries that extend across multiple input documents 104, the chatsheet 124 may include portions from different hierarchical representations 108 to form a mega-outline specific to the query 120. The common hierarchical format ensures that these different sources can be coherently merged. An additional advantage is that the users can directly request the organization themselves (e.g., "show me the hierarchical summary of document X"), and the system can return the stored organization (i.e., hierarchical representation 108) as a quick, human readable reference. This dual use of hierarchical representations, for both AI context compilation and human reference, provides a significant practical value in excess.Maintenance and IterationOver time, source documents may be updated or new documents are added to the knowledge base. Various embodiments may also assist in maintaining and updating the hierarchical representation 108 over time. As an input document 104 changes or new data is added, its hierarchical representation 108 may be re-generated or incrementally updated.In an example implementation, the hierarchical representation 108 is simply re-created from reason to reason when the input document 104 changes.In another example implementation, a comparison is made with the input document 104 and the LLM is prompted to update only the affected joint sections. Another aspect is extendibility to new data types, e.g., when a new type of structured data occurs (e.g., XML or a film document), the pipeline could request the LLM to generate a new parser or conversion code for this format (similar to the handling of table data, discussed above).In another example implementation, an LLM reads the old and new versions of the input document and then describes the differences in an organization form that could then be used to update the stored hierarchical representation. Since the hierarchical representation is lossless, it can even be treated as the canonical form and, although only minor, changes can be applied directly to it. If it is a previously unknown data format (e.g., if the system is expanded to process films or spreadsheets), the pipeline may involve new conversion requests and / or scripts. For example, if a PowerPoint file is entered, the system could request the LLM to extract text from the films and treat each film as a section in the organization, possibly with a chart frameholder.The embodiments are flexible to handle such cases by utilizing the LLM's ability to follow new instructions.In various embodiments, the hierarchical representations 108 themselves may be assembled into larger structures. Multiple input documents 104 may each be converted into a hierarchical representation 108, and then a parent organization may organize the collection (e.g., organization of an entire corpus having one node per input document 104 that is each associated with its own organization tree of the document). In this way, a multi-level hierarchy of information (a forest of trees) is created.Examples of the reactionFIGS. 7A and 7B show a detailed processing process that combines various concepts of the present disclosure in accordance with an exemplary embodiment.FIG. 8 shows an example of a digest from an input document 104, according to an embodiment. The full version of the document is available at https: / / www.bopa.ad / bopa / 0266067 / Pagines / Io26067002.aspx. In this example, the input document 104 is a right text which regulates personalized motor vehicle identifiers in Andorra. As can be seen, input document 104 contains various text items 802 and a table 804 as an example of tabular data.FIGS. 9A through 9F show chunks created from the input document 104 in FIG. 8 using LamaCloud, a managed platform for parsing and capturing data using Llamalendex. To ensure comparison, LamaCloud has been entrusted with the task of deciding, in the "Automatic Mode for Accuracy" itself, how to best handle the acquisition and retrieval of data for that target. The goal was to fully answer the question, and in both systems all other results that do not contribute to the answer (i.e. are lower in the relevance ranking list than the last required result) were excluded (however, it should be noted that the automatic limitation to relevant results was better in our system). If the individual chunk quantities are summed, the LamaCloud representation uses a total of 1,542 tokens (956+277+124+40+95+50).FIG. 10 shows an excerpt from a hierarchical representation 108 that was created for the input document 104 in FIG. 8, according to one embodiment. As can be seen, text elements 802 from input document 104 have been structured with inclusion levels 1002, and all content of input document 104 is incorporated into hierarchical representation 108. The table 804 from the input document 104 has been converted to a tablet representation 1004.FIGS. 11A through 11C show three chunk sets (fetch units 112) created for the example document 104 in FIG. 8 based on the hierarchical representation 108 in FIG. 10, in accordance with one embodiment. If the individual token sets are summed, this representation uses 616 tokens in total (144+135+337), which is significantly less than in the LamaCloud representation. However, the actual savings are much higher, as not these chunk sets are provided as context, but the de-duplicated chat sheet based thereon (see below).FIG. 12 illustrates the chat sheet (context document 124) created for the example document 104 in FIG. 8, according to one embodiment. As can be seen, the chatsheet 124 uses only 337 tokens, which is much less than in the LamaCloud representationThe ratio between the used input tokens for the LLM client is to return this to the right light: 1,542 input tokens in a single call (chat question) to GPT-4.5 currently cause costs of almost 0.12$, while the corresponding 337 tokens of the described embodiment result in less than 0.03$. The ratio is also directly proportional to the network bandwidth consumed, and particularly to the energy (for both inference and transmission).With respect to storage, it should be noted that the approach according to the described embodiment basically only needs to store the original text (i.e. the one which is unavoidable in any case) and a minimum overhead of metadata (the efficiently storable hierarchical relationships and vector embeddings). Note that in the LamaCloud example, not only a text representation (text sections with 1,000 tokens and 200 token overlap) and embeddings thereof have been stored, but also a page by page representation of the HTML original and the corresponding OCR, both of which are actually redundant.It is noted that the approach according to the described embodiment could consume slightly more resources during ingestion than the alternatives (not in the case described above, but possibly in other cases): apart from the conversions necessary in each approach, the approach according to the described embodiment still "sausages" all content (once), possibly resulting in a higher one-time consumption. However, very inexpensive, efficient and simple LLMs such as gpt-4o mini can be perfectly matched to the task of re-establishing, thereby greatly reducing this cost.Even preprocessing a document at scan time using the approach described, in which only the relevant parts are fed into a complex and cost- or energy-intensive main model, most likely results in a substantially lower overall consumption.This effect is enhanced when conclusions are drawn and the model must read the context in each "round". Such agent-based interactions also allow more time and thus represent a second application besides being incorporated once into knowledge databases.Hardware ImplementationFIG. 13 shows a schematic block diagram of the computer hardware on which embodiments of the present disclosure may be implemented. As seen, a computing device 1302 includes one or more processors 1304 and a memory 1306. The one or more processors 1304 are communicatively coupled to the memory 1306. The memory 1306 stores a computer program 1308. The computer program 1308 may implement some or all aspects of the disclosed methods and functionalities.Advantages & Other ExamplesVarious advantages achieved by the concepts disclosed herein are summarized below:By explicitly requesting an LLM to render documents in a summary form (rather than summarizing them or responding questions spontaneously), any detail is maintained in a manner segmented and indexed for easy retrieval. All the processes disclosed here are loss-free from the outset.By utilizing LLMs to generate transformation code for complex data (e.g., tables), information that is otherwise difficult to embed is converted to a form that can digest the LLM without compromising accuracy.The uniform treatment of text and non-text media in a hierarchical scheme ensures that no part of the source falls through the meshes.Even if, for efficiency reasons, additional abstracts or alternative chunk sets are created, the original hierarchical representation 108 remains with all details as a base reference. By combining the generative capabilities of LLMs with deterministic data transformation, embodiments of the present disclosure provide a rich, structured context library for large language models. The LLM is used not for summary, but for re-structuring and indexing information in a manner suitable for both machines and humans (one could read the hierarchical representation 108 as a comprehensive "sheet of mail" (Cheatsheet) of the source input document 104).In summary, various embodiments provide an end-to-end pipeline for fetch-extended generation that provides significantly more accurate, contextually relevant and original source-returnable responses. In essence, the presented embodiments allow large language models to store detailed, trusted storage of source material.In this way, the problem of bounding the context window is overcome not by lossy compression, but by intelligent re-structuring and extending the meaning of "context" (from a flat text string to a rich, query dependent structured document). The constraints on context length and hallucination tendencies of LLMs are addressed directly by the disclosed embodiments because the model no longer needs to inverify huge raw texts or images at the time of retrieval and it does not need to rate over data that has been cropped or abstracted. Instead, it has the ground truth in mouth-fit pieces. The result is optimized, call extended generation, in which the responses are more accurate, contextual, and returnable to the original material.Further exemplary embodiments are described below:Example 1. a method that explicitly prompts a generative language model for lossless hierarchical literal repetition by incremental intervention, utilizing inherent structural understanding during generation.Example 2. the method of example 1 includes iterative multi-round prompt with explicit delimiters for lossless reconstruction of documents that exceed the boundaries of single-output tokens.Example 3. The method of example 1 or 2 further comprises directly hierarchic capturing complex tables that generate deep semantically-true "table trees.".Example 4 Method of Requesting a Generative Language Model to Create Executable Transformation Scripts for Token-Efficient, Flat, Hierarchical Large Table Representation.Example 5. a method comprising integrating multimodal place holders into hierarchical structures, wherein external storage of originals is maintained and the original multimodal content is optionally re-integrated based on explicit retrieval criteria.Example 6. a method comprising:• Algorithm Generation of Specialized Fetch Units (Chunks) representing particular traversal paths through the hierarchical structure of the document by (1) sorting chunks by depth and location: (1) sorting chunks by depth and location, (2) merging adjacent chunks at the same depth level within token boundaries, (3) recursively preceding parent chunks to form full hierarchical paths, and (4) creating unique combinations that maintain parent-child relationships while optimizing token efficiency.• Embed and Retrieve at chunkset level rather than at individual chunks level, thereby ensuring that matches are made with semantically complete paths rather than isolated segments.• Dynamic construction of predicted, minimally redundant chat sheets per query document pair by extraction and reassembly of constituent chunks from the best matched chunk sets.• Optional reintegration of the original multimodal content based on explicit retrieval criteria, e.g. user selection, precision requirements, availability of replacement context or semantic / visual difference thresholds.Example 7. Method comprising:• receiving an input document (104), the input document comprising content items including at least one text item;• Performing a lossless hierarchical text editorial process (106), in which the input document is displayed as a hierarchical representation (108) in which all text elements of the input document are retained word-by-word and organized according to a structure of the input document.Example 8: The method of example 7, wherein the lossless hierarchical text rendering process (106) comprises:• causing a generative language model to generate the hierarchical representation.Example 9. the method of example 8, wherein requesting the generative language model comprises:• Instruction to the generative language model to create the hierarchical representation without merging or omitting the text elements.Example 10. the method of any of Examples 7 to 9, wherein the lossless hierarchical text rendering process is an iterative process; and / or wherein each iteration comprises generating a partial hierarchical representation until a context window threshold of the generative language model is reached; and / or wherein each iteration comprises horizontally and / or vertically compressing the partial hierarchical representation of the current iteration (and all partial hierarchical representations of earlier iterations); and / or wherein each iteration comprises concatenating the compressed partial hierarchical representation of the current iteration (and all compressed partial hierarchical representations of earlier iterations) with a delimiter, followed by the unprocessed remainder of the input document serving as input for a subsequent iteration.Example 11: The method of any of Examples 7 to 10, wherein, in a structured input document, the content items are organized in the hierarchical representation according to an explicit structure in the input document.Example 12. The method of any of Examples 7 to 11, wherein for an unstructured input document, the content items are organized in the hierarchical representation according to a derived structure of the input document determined by the generative language model.Example 13.The method of any of Examples 7 to 12, wherein the hierarchical representation uses engagement levels to organize the content items.Example 14. The method of any of Examples 7 to 12, wherein the hierarchical representation uses semantic labels, such as markup elements and / or markup elements, to organize the content elements.Example 15. Method comprising:• receiving an input document (104), the input document comprising content items including at least one text item;• Performing a fetch unit generation process (110), wherein fetch units (112) are generated based on a hierarchical representation (108) of the input document (104) and stored in a repository accessible to a generative client language model (608) for fetch-assisted generation.Example 16 The method of Example 15 in combination with any one of Examples 7 to 14.Example 17. the method of Examples 15 or 16, wherein the generation process (110) of the fetch unit comprises:• Generation of chunksets as query units, each chunkset indicating a traversal path through the structure of the input document.Example 18. The method of example 17, wherein generating chunk sets comprises one or more of the following steps:• sorting chunks in the hierarchical representation by depth and position;• Collapse adjacent chunks at the same depth level that are within a token boundary;• recursive preceding of superordinate chunks to form traversal paths; and• Create unique combinations that maintain hierarchical relationships while optimizing token efficiency.Example 19: The method of Examples 17 or 18 wherein each chunk set maintains pointers to its constituent chunks.Example 20. The method of any one of Examples 17 to 19, wherein storing the chunk sets as the fetch units comprises:• Calculation of the embeds for each chunk set and storage of the embeds in the repository.Example 21: Method comprising:• receiving an input document (104), the input document comprising content items including at least one text item;• Performing a table integration process (114) comprising:◯ identification of table data in the input document; and Converting the tabular data into a hierarchical representation in the form of a table tree in which all data from the tabular data are retained.Example 22: The method of Example 21 in combination with any one of Examples 7 to 20.Example 23. the method of Example 21 or 22, wherein if the tabular data is below a predetermined size threshold, transforming the tabular data comprises:• Causing a generative language model to convert the tabular data to the tablet representation.Example 23. the method of Example 21 or 22, wherein if the tabular data exceeds a predetermined size threshold, transforming the tabular data comprises:• requesting a generative language model to generate executable code or executable code parameters that converts the tabular data to the tablet representation; and• Executing the generated code to generate the tablet ree representation.Example 24. Method comprising:• receiving an input document (104), the input document comprising content items including at least one text item;• Performing a multi-modal content integration process (116) comprising:identifying a non-textual element in the input document; andThe integration of a textual representation of the non-textual element into a hierarchical representation.Example 25 The method of Example 24 in combination with any one of Examples 7 to 23.Example 26. The method of Example 23 or 24, wherein integrating the text representation comprises:• Generation of a textual description of the non-textual element;• inserting the textual description into the hierarchical representation at a position corresponding to the position of the non-textual element in the input document; and• storing the original non-textual element with a reference identifier associated with its textual description in the hierarchical representation.Example 27. the method of example 26, wherein generating the textual description comprises one of the following:• calling a speech imaging model to describe an image;• Application of optical character recognition to extract text from the image; or• Extract metadata associated with the non-textual element.Example 28. Method comprising:• receiving a request;• optionally embedding the query;• Identification of chunk sets relevant to the query by comparing the embedded query with stored chunk sets embeds and / or by comparing the query with stored chunk sets; and• Assembly of a context document from chunks referenced by the identified chunk sets.Example 29: The method of Example 28 in combination with any one of Examples 7 to 27.Example 30. the method of example 28 or 29, wherein composing the context document comprises:• extracting chunks referenced by the identified chunk sets;• maintaining hierarchical relationships between the extracted chunks; and• Assembly of the extracted chunks into a coherent document while maintaining structural information such as the engagement.Example 31. The method of example 29 or 30, wherein if the composite context document crosses a context window of the client generative language model, the method further comprises:• Prioritizing chunks based on their relevance to the query preserving hierarchical integrity.Example 32. the method of any one of Examples 28 to 31, further comprising:• Providing the compiled context document and the request to a generative language model of the client to generate a response.Example 33. The method of example 32, wherein each chunk includes reference information indicating its source location in the input document, and wherein the client generative language model (608) is instructed to include the reference information when cited from the context document.Example 34 The method of any of Examples 28-33 further includes performing the following steps in response to determining that the query requires information from a particular non-textual element• retrieving the original non-textual element based on its reference identifier in the hierarchical representation; and• Provision of the original non-textual element for the customer's generative language model (608).Example 35. Method comprising:• detecting an update of an input document;• identifying parts of a hierarchical representation of the input document affected by the update; and• selectively regenerate only the affected parts of the hierarchical representation.Example 36: The method of Example 35 in combination with any one of Examples 7 to 35.Example 37: Method comprising:• Creation of a Multi-Document Hierarchy by:generating a hierarchical representation for each of a plurality of input documents; andcreating a parent hierarchical representation that organizes the collection of hierarchical representations.Example 38: The method of Example 37 in combination with any one of Examples 7 to 36.Example 39. A data processing apparatus, apparatus, or system comprising:• a processor; and• a memory storing instructions which, when executed by the processor, configure the apparatus to perform the method of any one of Examples 1 to 38.Example 40. A data processing apparatus, apparatus or system comprising means for performing (the steps of) the method of any one of Examples 1 to 38.Example 41: A computer readable medium having computer executable instructions stored thereon to implement the method of any one of Examples 1 to 38.Example 42. a computer program (product) including instructions that, when the program is executed by a computer, cause the computer to perform the method of any one of Examples 1 to 38.Although various aspects and embodiments have been illustrated and described in detail in the foregoing description and drawings, these illustrations and descriptions are illustrative or exemplary and not limiting. Variations from the disclosed embodiments may be understood and carried out by those skilled in the art upon application of the claimed subject matter by studying the drawings, the disclosure, and the appended claims.Although some aspects have been described in the context of a product, apparatus, apparatus, or system, these aspects also provide a description of the corresponding process, method, or use, wherein a block or component corresponds to a step or feature of a step. Analogously, aspects described in connection with a method step also represent a description of a corresponding block or component or feature of a corresponding product, apparatus, device or system.The order of execution of the operations in the described embodiments is not essential unless otherwise stated. That is, the operations may be performed in any order unless otherwise indicated, and the embodiments may include additional or fewer operations than those mentioned.In the claims, the word "comprising" does not exclude other elements or steps, and the indefinite article "a" or "an" does not exclude a plurality. A single unit may perform the functions of a plurality of units recited in the claims. The mere fact that certain measures are recited in dependent claims does not mean that a combination of these measures may not be advantageous. Any reference signs in the claims should not be understood as limiting the scope of application.Embodiments of the present disclosure may be implemented in hardware, software, or both. The implementation can be effected using a non-transitory storage medium, such as a digital storage medium, for example a floppy disk, a DVD, a Blu-ray, a CD, a ROM, a PROM, an EPROM, an EEPROM or a FLASH memory, on which electronically readable control signals are stored which cooperate (or can cooperate) with a programmable computer system, such that the respective method is carried out. Therefore, the digital storage medium may be computer readable.Embodiments of the present disclosure may be implemented on a computer system. The computer system may be a local computing device (e.g., personal computer, laptop, tablet, or cellular phone) having one or more processors and one or more storage devices, or a distributed computing system (e.g., a cloud computing system having one or more processors and one or more storage devices distributed at different locations, e.g., at a local client and / or one or more remote server farm and / or data centers). The computer system may include any circuit or combination of circuits. In one embodiment, the computer system may include one or more processors of any type. As used herein, the term "processor" may refer to any type of computing circuit, e.g., a microprocessor, a microcontroller, a complex instruction set computing (CISC) microprocessor, a reduced instruction set computing (RISC) microprocessor, a very long instruction word (VLIW) microprocessor, a graphics processor, a digital signal processor (DSP), a multi-core processor, a field programmable gate array (FPGA), or any other type of processor or processing circuit. Other types of circuitry that may be included in the computer system may be custom circuitry, an application specific integrated circuit (ASIC), or the like, such as one or more circuits (e.g., communication circuitry) for use in wireless devices such as cellular phones, tablet computers, laptop computers, two-way radios, and similar electronic systems. The computer system may include one or more storage devices that may include one or more storage elements suitable for the particular application, such as main memory in the form of random access memory (RAM), one or more hard drives, and / or one or more drives that process removable media such as compact disks (CD), flash memory cards, digital video disks (DVD), and the like. The computer system may also include a display device, one or more speakers, and a keyboard and / or controller, which may include a mouse, a trackball, a touch screen, a voice recognition device, or other device that allows a system user to input and receive information from the computer system. Some or all of the method steps may be performed by (or using) a hardware device, such as a processor, a microprocessor, a programmable computer, or an electronic circuit. In some embodiments, one or more of the described method steps may be performed by such a device. Another embodiment is an apparatus as described herein comprising a processor and a storage medium.Embodiments of the present disclosure can be implemented as a computer program (product) with a program code, wherein the program code is used to carry out one of the methods when the computer program product runs on a computer. The program code may be stored on a machine-readable carrier, for example. Other embodiments include a computer program for performing any of the methods described herein stored on a machine readable medium. A further embodiment is a computer program having a program code for carrying out one of the methods described herein when the computer program runs on a computer. A further embodiment is a storage medium (or a data carrier or a computer-readable medium) on which the computer program for carrying out one of the methods described here is stored when it is executed by a processor. The data carrier, the digital storage medium or the recorded medium are typically tangible and / or non-transitory. A further embodiment is a computer on which the computer program for carrying out one of the methods described herein or individual steps thereof is installed.Another embodiment is a data stream or a sequence of signals representing the computer program for performing one of the methods described herein. The data stream or the signal sequence can be configured, for example, such that it / it is transmitted via a data communication connection, for example via the Internet.Another embodiment is an apparatus or system configured to transmit (e.g., electronically or optically) a computer program to a receiver for performing any of the methods described herein. The receiver may be, for example, a computer, a mobile device, a storage device, or the like. The apparatus or system may, for example, comprise a file server for transmitting the computer program to the recipient.REFERENCE NUMERALS100 Processing pipeline 102 Document creation unit 104 Input document 106 Lossless hierarchical text editor method 108 Hierarchical representation 110 Method for generating fetch units 112 Fetch units 114 Table integration process 116 Multimodal content integration process 118 Query processing unit 120 Query 122 Context document assembly process 124 Context document 202 Stage 302 Step 304 Step 306 Step 402 Step 404 Step 406 Step 502 Step 504 Step 506 Step 508 Step 602 Step 604 Step 606 Step 608 Customer's generative language model 802 Text element 804 Table 1002 Intervention levels 1004 Tablet tree representation (TableTree representation) 1302 Computing device 1304 Processor 1306 Memory 1308 Computer programReferences included in the specificationThis list of documents cited by the applicant has been produced in an automated manner and is only included for the better information of the reader. The list is not part of the German patent application or utility model application. The DPMA does not take any adhesion for any faults or omissions.Cited Non-Patent Literaturehttps: / / communication.openai.com / t / using-gpt-4 api-to-secretly-chunk-documents / 715689

[0005] https: / / www.bopa.ad / bopa / 0266067 / Pagines / Io26067002.aspx

[0221]

Claims

A computer program comprising instructions which, when the program is executed by a computer, cause the computer to perform the operations of: receiving an input document (104), the input document (104) including content items including at least one text item; performing a lossless hierarchical text rendering process (106) in which the input document (104) is reproduced as a hierarchical representation (108) in which all text items are word-wise obtained from the input document (104) and organized according to a structure of the input document (104), the lossless hierarchical text rendering process (106) comprising: prompting a generative language model to generate the hierarchical representation (108), the prompt instructing the generative language model to generate the hierarchical representation (108) without merging or omitting one of the text items; and performing a fetch unit generation process (110), wherein fetch units (112) are generated based on the hierarchical representation (108) and stored in memory that can be accessed by a generative client language model for fetch-extended generation.The computer program of claim 1, wherein the lossless hierarchical text rendering process (106) is an iterative process; wherein each iteration comprises generating a partial hierarchical representation until a context window threshold of the generative language model is reached; wherein each iteration comprises horizontally and / or vertically compressing the partial hierarchical representation of the current iteration and all partial hierarchical representations of previous iterations; wherein each iteration comprises concatenating the compressed partial hierarchical representation of the current iteration and all compressed partial hierarchical representations of previous iterations with a delimiter followed by the unprocessed remainder of the input document (104) serving as input for a subsequent iteration.The computer program of claim 1 or 2, wherein for a structured input document (104), the content items in the hierarchical representation (108) are organized according to an explicit structure in the input document (104); and wherein for an unstructured input document (104), the content items in the hierarchical representation (108) are organized according to a derived structure of the input document (104) determined by the generative language model.The computer program of any preceding claim, wherein the hierarchical representation (104) uses engagement levels to organize the content items.The computer program of any preceding claim, wherein the fetch unit generation process (110) comprises: generating chunk sets (112) as fetch units; each chunk set (112) indicating a traversal path through the structure of the input document (104); each chunk set (112) maintaining pointers to its associated chunks.The computer program of claim 5, wherein generating chunks (112) comprises: sorting chunks in the hierarchical representation (108) by depth and position; merging adjacent chunks at the same depth level that are within a token boundary; recursively preceding parent chunks to form traversal paths; and generating unique combinations that maintain hierarchical relationships while optimizing token efficiency.The computer program of any preceding claim, comprising a table integration process (114) comprising: identifying tabular data in the input document (104); and converting the tabular data to a table tree representation in the hierarchical representation (108) at which all data from the tabular data is retained.The computer program of claim 7, wherein if the tabular data is below a predetermined size threshold, converting the tabular data comprises: promptting a generative language model to convert the tabular data to the table tree representation; and wherein if the tabular data exceeds a predetermined size threshold, converting the tabular data comprises: promptting a generative language model to generate executable code or executable code parameters that convert the tabular data to the table tree representation; and executing the generated code to generate the table tree representation.The computer program of any preceding claim, comprising a multi-modal content integration process (116) comprising: identifying a non-textual element in the input document (104); and integrating a text representation of the non-textual element into the hierarchical representation (108).The computer program of claim 9, wherein integrating the text representation comprises one or more of: generating a textual description of the non-textual element; inserting the textual description into the hierarchical representation (108) at a position corresponding to the position of the non-textual element in the input document (104); and storing the original non-textual element with a reference identifier associated with its textual description in the hierarchical representation (108).The computer program of claim 10, wherein generating the textual description comprises one or more of: invoking a speech imaging model to describe an image; applying optical character recognition to extract text from the image; extracting metadata associated with the non-textual element.The computer program of any preceding claim, further comprising: receiving a query; optionally embedding the query; identifying chunk sets (112) relevant to the query by comparing the embedded query to stored chunk sets embeddings and / or by comparing the query to stored chunk sets (112); composing a context document (124) from chunks referenced by the identified chunk sets (112); providing the composed context document (124) and the query to a customer's generative language model (608) to generate a response.The computer program of claim 12, wherein each chunk contains reference information indicating its source location in the input document (104), and wherein the customer's generative language model is instructed to include the reference information when cited from the context document (124).The computer program of claim 12 or 13, further comprising, in response to a determination that the query requires information from a particular non-textual element, performing the steps of: retrieving the original non-textual element based on its reference identifier in the hierarchical representation (108); and providing the original non-textual element to the customer's generative language model (608).The computer program of any preceding claim, comprising: detecting an update of the input document (104); identifying portions of the hierarchical representation (108) affected by the update; and selectively re-generating only the affected portions of the hierarchical representation (108).The computer program of any preceding claim, further comprising creating a multi-document hierarchy by: generating a hierarchical representation (108) for each of a plurality of input documents (104); and creating a parent hierarchical representation that organizes the collection of hierarchical representations (108).A data processing apparatus (1302) comprising: a processor (1304); and a memory (1306) in which the computer program (1308) of any one of claims 1 to 16 is stored.A computer readable medium having stored thereon the computer program of any one of claims 1 to 16.

Citation Information

Cited By

  • Multi-mode sound picture storage platform

    CN120639917A

  • Iterative text refining method, system and equipment based on fidelity and medium

    CN120706380A

  • Fidelity-based iterative text refinement method, system, device, and medium

    CN120706380B

  • Multi-modal database construction method and device based on visual language model collaborative routing

    CN120950741A

  • Multimodal database construction method and device based on visual language model collaborative routing

    CN120950741B