Method for unpicking of text with repeat accuracy

The method leverages a hybrid approach of generative AI on GPU servers and rule-based analysis on CPU servers to achieve precise and repeatable text separation in complex texts, addressing inefficiencies and errors in existing methods.

WO2025107012A1PCT designated stage expired Publication Date: 2025-05-30ABP SERVICES GMBH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
PCT/AT2024/060451
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-20
Filing Date
2024-11-18
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Existing methods for text separation in complex texts, such as patent descriptions, are inefficient and prone to errors, particularly in distinguishing between different feature combinations and preventing hallucination or incorrect separation.

Method used

A method utilizing a combination of large language models with generative AI running on GPU servers and rule-based analysis on CPU servers to accurately separate text into correctly defined parts, ensuring context preservation and preventing hallucination.

Benefits of technology

The method enables repeatable and precise text separation, optimizing computationally intensive steps across different hardware components to ensure fast and accurate processing, while preventing hallucination and preserving context information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure AT2024060451_30052025_PF_FP_ABST
    Figure AT2024060451_30052025_PF_FP_ABST
Patent Text Reader

Abstract

The invention relates to a method and a computer-implemented method for the unpicking, with repeat accuracy, of text into correctly separated parts based on the meaning of the words in an infrastructure or hardware for artificial intelligence, wherein the method of unpicking of texts by understanding of natural language is carried out by way of the trained large-language model on the one or more graphic processors or graphic processor servers, and uses the processing of natural language by way of stored rules for improving the unpicking of the individual text parts and individual text parts are transferred in a machine-readable format to a management program for storing and / or displaying the unpicked text parts and / or the relation to associated text sources. The computation-intensive steps are to be optimally distributed between different hardware components or servers in order to guarantee fast processing.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] METHOD FOR REPEATEDLY PRECISE TEXT SEPARATION

[0002] The invention relates to a method for the repeatable separation of text into correctly separated parts according to the literal meaning in an infrastructure or hardware for artificial intelligence.

[0003] Analyzing complex texts, particularly patent descriptions, patent claims, technical descriptions of machines or systems, or invention disclosures, requires a considerable amount of time. There have been numerous attempts in the prior art to classify and keyword long and complex texts in order to make them comparable for human readers, as well as using computer-implemented methods. The key problem is that a disclosure can contain various embodiments in one text. These embodiments do not necessarily belong together and can each be a combination of different features. While some of these features may overlap, they still have individual, different, mandatory features and therefore do not form a common combination of features with other embodiments in the disclosure.

[0004] Features and their textual linkage are crucial for determining the scope of protection of the patent or the disclosed combination of features and for comparing it with other disclosures or combinations of patent features. Features, for example, in a patent claim, must be formulated clearly and precisely to ensure differentiation from other technologies.

[0005] A feature combination is understood as the collection of various features or properties that together describe the specific technical solution or unique aspect of an invention. The feature combination is usually contained in the patent claims, which define the legally binding scope of protection of a patent.

[0006] The combination of features plays a crucial role in distinguishing the patented invention from the prior art. A clear and unambiguous formulation of the combination of features is important for determining the scope of protection of a patent and ensuring that third parties cannot use similar technologies or processes that have the defined features without the patent owner's consent. US11714839B2 shows a method for automated patent claim analysis. This involves going through the claims word by word to identify word sequences as scope concept phrases. By using synonyms from a thesaurus, additional scope concept phrases are generated, which are then automatically classified in a product description to create a graphical representation of the textual scope of each scope concept phrase.

[0007] Even newer large language models such as Chat GPT, which have an understanding of natural language, cannot reliably, repeatably and reproducibly extract related feature combinations from the analysis of a longer text and compare these with other feature combinations of other longer texts.

[0008] In particular, there is the problem that exclusively rule-based separation, for example, with commas or conjunctions or connectors such as the word "and," can lead to incorrect separation of text sections, or the separated text sections no longer make sense on their own. There is also the problem of hallucinating or inventing text that isn't present in the source text.

[0009] The object of the present invention is to overcome the disadvantages of the prior art and to provide a method by means of which a user is able to carry out simple and repeatable separation of text into parts that are correctly separated according to the literal meaning and to optimally distribute the computationally intensive steps between different hardware components or servers in order to guarantee fast processing.

[0010] This object is achieved by a method according to the claims.

[0011] The method according to the invention is used for the repeatable separation of text into correctly separated parts according to the literal meaning in an infrastructure for artificial intelligence, which includes one or more storage systems and one or more central processor servers (CPU servers) and one or more graphics processor servers (GPU servers), wherein the method comprises the following:

[0012] Calling a stored text to be separated in a management program by one of the central processor servers (CPU servers) from one of the storage systems, transferring the text to be separated in a machine-readable format to a large language model with generative artificial intelligence, which runs on one or more graphics processor servers (GPU servers),

[0013] Separating the text by understanding natural language using the trained large language model on one or more graphics processor servers (GPU servers), saving the text source to a separated text part on one or more of the storage systems,

[0014] Transfer of text parts of the text and the text source separated by the large language model in a machine-readable format to a program which, in addition to processing natural language, also uses rules stored on one of the storage systems to improve the separation of the individual text parts;

[0015] Transfer of the individual text parts in a machine-readable format to a management program running on a central processor server (CPU server) and storage and / or display of the separated text parts and / or the relationship to associated text sources.

[0016] The advantage of this approach in the first embodiment is to divide the computationally intensive steps between different servers or server types in such a way that, on the one hand, the advantages of the different server types, namely CPU servers and GPU servers, can be used - for example, for parallel processing - and, on the other hand, that the context information about the text source is not lost and hallucinating is prevented.

[0017] In a further embodiment, the computer-implemented method according to the invention serves for the repeatable separation of text into correctly separated parts according to the literal meaning in a hardware for artificial intelligence, comprising the steps:

[0018] Calling up a stored text (5) to be separated in a management program (6) by one of the central processors (22) from one of the storage systems (2),

[0019] Transferring the text (5) to be separated in a machine-readable format to a large language model (11) with generative artificial intelligence, which runs on one or more graphics processors (23),

[0020] Separating the text (5) by understanding natural language by the trained large language model (11) on the one or more graphics processors (23), storing the text source (12) to a separated text part (13) on one or more of the storage systems (2),

[0021] Transferring text parts (13) of the text (5) and the text source (12) separated by the large language model (11) in a machine-readable format to a program (15) which, in addition to processing natural language, also uses rules (16) stored on one of the storage systems (2) to improve the separation of the individual text parts (13),

[0022] Transferring the individual text parts (13) in a machine-readable format to the management program (6) which runs on a central processor (22) and

[0023] Saving and / or displaying the separated text parts (13) and / or the relation (19) to associated text sources.

[0024] Another advantage of this further process variant is that computationally intensive steps are distributed among different hardware components in such a way that, on the one hand, the advantages of the different processor types, namely CPU and GPU, can be utilized, while, on the other hand, contextual information about the text source is not lost, preventing hallucinating. Furthermore, lower hardware investment is required.

[0025] Furthermore, the invention relates to a data processing system comprising multiple processors adapted / configured to execute the steps of the method described above. The advantage here is the linking of the individual program and hardware components in order to be able to carry out the separation of text sections in a time-efficient and precise manner.

[0026] Additionally, the invention relates to a computer program product comprising instructions that, when executed by a computer, cause the computer to execute the method as described above. It is advantageous to control the individual instructions for controlling the method steps on different hardware components from a computer program product or management program and to make them further processable.

[0027] Furthermore, it can be advantageous to transfer the separated text sections of the text and the text source in a machine-readable format to a program for rule-based analysis to improve the separation. A rule-based analysis is carried out by the program according to rules stored on one of the storage systems on a central processor server in order to achieve a further improved separation result while simultaneously conserving hardware resources. For example, a rule can be used that stipulates a separation from the following text section before a certain word or after a certain word or after a certain word combination, or that certain characters such as formulas, symbols, variables, reference numbers, bracket expressions or formatting information should not be taken into account during separation by the large language model.

[0028] Another advantageous approach is to consider the stored text source in the form of a text part to text source relation to ensure that separated text parts are displayed identically upon subsequent call. This allows for repeatable query accuracy without requiring intensive computing operations on the hardware components.

[0029] Another optional approach involves further processing the separated text segments with a large language model running on one or more graphics processor servers or graphics processors to generate concepts, keywords, or generic terms, as well as storing and / or displaying them in the management program. This improves the relevance of results for subsequent automated queries. These concepts, keywords, or generic terms can also support users in formulating queries and in the further use of the large language model to create texts such as patent claims, descriptions, or responses to official notices.

[0030] It can also be advantageous to further process the separated text sections using artificial intelligence (NLP) running on one or more graphics processor servers to generate concepts, keywords, or umbrella terms, as well as store and / or display them in the management program. The use of specially trained NLP systems leads to similar yet different results, which can lead to improved quality and form the basis for automated improvement of prompts and instructions.

[0031] A further advantageous development of the method is that the large language model is trained using data accessible via the management program on one or more of the storage systems. This allows for automated improvements to prompts and instructions, without additional training time, using information in the management program that is already available in a context-related and linked manner from human users.

[0032] It is particularly efficient that the large language model is trained using data in patent publications, especially patent claims, which are accessible on storage devices via the Internet.

[0033] Another advantageous approach is to fine-tune the accuracy of the large language model with generative artificial intelligence using prompts, natural language instructions, scripts, and sample texts, using one or more GPU servers. This allows for more efficient and faster work with smaller language models. Fewer tokens are required to train a language model on a GPU server for a specific application area, such as patent claims.

[0034] It can be particularly advantageous that the separated text segments can be compared with other text segments accessible via the management program on one or more of the storage systems using the large language model with generative artificial intelligence, and stored and / or displayed in a table within the management program. This provides a machine-readable database for improving the process, but also allows for feature analyses and comparisons.

[0035] The method can be further improved by calculating the degree of similarity between two or more text fragments to be compared using one or more graphics processors or graphics processor servers, which serves as the basis for relevance-based storage or display. Parallel processing on a GPU can significantly improve training and inference speeds.

[0036] An efficient further development of the method could advantageously be to allow the separation points between the separated text sections to be changed in the management program via user input. These changes can be used to further refine the accuracy of the large language model with generative artificial intelligence, particularly via prompts, instructions, or scripts. This allows the large language model to be further trained "incidentally" through the inputs in the management program.

[0037] It is also advantageous if the management program allows a comparison result between the separated text segments to be modified or evaluated via user input, and this user input can be used to further refine the accuracy of the large language model with generative artificial intelligence, particularly via prompts, instructions, or scripts. This allows, for example, the large language model to be further trained using decision boxes, evaluation functions, or selection lists provided to the user in the management program.

[0038] Large language models (LLMs) are comprehensive language models based on machine learning that are capable of understanding, generating, and responding to natural language. Models such as GPT-3 (Generative Pretrained Transformer 3 or LLaMA) are examples of LLMs.

[0039] For the purposes of this application, the term NLP is defined as follows: Natural language processing (NLP) is a branch of computer science, artificial intelligence, and linguistics that deals with the interaction between computers and human (natural) languages. As such, NLP is related to the field of human-computer interaction, particularly with regard to natural language understanding, which enables computers to derive meaning from human or natural language input.

[0040] Many NLP systems use ontologies to support the performance of NLP tasks. An ontology is a representation of knowledge. In the case of NLP, a semantic ontology is a representation of knowledge about the relationships between semantic concepts. Ontologies are created by humans, usually by experts, and are never a perfect representation of all available knowledge. They are often highly focused on a specific subset of a particular domain and often reflect the author's level of knowledge or level of detail. Ontologies are typically task-oriented, meaning they have some utility related to managing information or physical entities, and their design reflects the task for which their terminology is needed.In the context of artificial intelligence (AI)—as in this application—the term "token" often refers to the basic units of text processed by AI models. Here are two common meanings in the context of AI:

[0041] A token is a unit such as a word or a subunit of a word. For example, the sentence "The cat runs fast" is divided into tokens, where each token is a single word: "The," "cat," "runs," "fast." This tokenization enables AI models to understand and analyze text. Different large language models can only work with a limited number of tokens. The Llama 2 language model, for example, can handle up to 4096 tokens, and the CodeLlama language model can handle up to 16384 tokens.

[0042] In the field of artificial intelligence (AI)—as here—the term "prompts" describes short instructions or text fragments given to an AI model to obtain a desired response or output. Prompts are like tasks that ask the model to generate or understand specific information.

[0043] Features are technically and / or functionally related parts of a larger combination of features, for example in a patent claim, an invention, a technical description or a product, which can also be composed of individual words, parts of sentences, conjunctions or connectors, but which cannot be further subdivided without losing the technical context or the meaning of the word - taken on their own.

[0044] Feature analysis or feature comparison refers to the examination and identification of the specific features or properties, or a combination of features, of a patented invention, an embodiment of a product, an invention report, or a technical disclosure. Feature analysis plays a crucial role in determining the scope of protection of a patent and in distinguishing it from other technologies. The analysis aims to identify the essential features or combinations of features that constitute uniqueness and inventive character. The identified features or combinations of features can be compared with the prior art to determine whether similar features are already known or whether the invention contains innovative aspects. Feature analysis helps determine the scope of protection of the patent.Feature analysis is an important step for patent examiners, attorneys, inventors, and engineers to deepen their understanding of a patented invention and clarify its legal significance. Furthermore, feature comparison provides a condensed overview of the disclosed feature combination in two or more documents under investigation.

[0045] GPU stands for "Graphics Processing Unit" and is a specialized hardware component optimized for processing graphics and parallel computing - especially for applications in the field of machine learning and artificial intelligence.

[0046] Refinement refers to fine-tuning, a method of transferring learning from a pre-trained artificial neural network. Foundation models are machine learning models that have been pre-trained on a large amount of data; these are then fine-tuned and optimized for specific tasks or use cases.

[0047] For a better understanding of the invention, it is explained in more detail using the following figures.

[0048] They show in a highly simplified, schematic representation:

[0049] Fig. 1 is a schematic representation of the infrastructure for carrying out the method according to embodiment 1;

[0050] Fig. 2 is a schematic representation of the method for carrying out the method according to embodiment 2;

[0051] Fig. 3 a schematic representation of the separated text with feature comparison and concepts in tabular form.

[0052] By way of introduction, it should be noted that in the variously described embodiments, identical parts are provided with identical reference symbols or component designations. The disclosures contained throughout the description can be applied analogously to identical parts with identical reference symbols or component designations. Furthermore, the positional information chosen in the description, such as top, bottom, side, etc., refers to the directly described and illustrated figure, and these positional information must be applied analogously to the new position in the event of a change in position.

[0053] Fig. 1 shows a schematic representation of the infrastructure for carrying out the method according to embodiment 1, on the basis of which the method steps are explained.

[0054] To ensure repeatable separation of text into correctly separated parts according to the literal meaning, an artificial intelligence infrastructure 1 is provided. This infrastructure includes one or more centralized or decentralized storage systems 2, such as long-term storage such as hard drives, cloud storage, storage systems, or similar, and / or short-term storage such as RAM. In addition, one or more central processor servers 3 (CPU servers) and one or more graphics processor servers 4 (GPU servers) are provided.

[0055] A text 5 to be separated, stored in the storage system 2, is called from one of the storage systems 2 via a management program 6, which has a database 7 with tables 8 and access to a central processor server 3. Table 8 contains data and links 9 to documents stored on the storage system 2, such as the text 5.

[0056] Transfer 10 of the text 5 to be separated in a machine-readable format to a large language model 11 with generative artificial intelligence, which runs on one or more graphics processor servers 4. Optionally, centralized or decentralized storage systems 2 can be assigned to the graphics processor server 4.

[0057] The text 5 is then separated by natural language understanding using the trained large language model 11 on one or more graphics processor servers 4. The text source 12 or a text passage of the text 5 is stored as a separated text part 13 on one or more of the storage systems.

[0058] Further transfer 14 of text portions 13 of the text 5 and the text source 12, separated by the large language model 11, in a machine-readable format to a program 15 which, in addition to processing natural language, also uses rules 16 stored on one of the storage systems 2 to improve the separation of the individual text portions. These rules 16 can, for example, include certain word sequences such as "characterized in that" or certain other separation conditions or operators such as semicolons.

[0059] Transfer 17 of the individual text parts in a machine-readable format to the management program 6 which runs on a central processor server 3.

[0060] Storage on the storage system 2 and / or display of the separated text parts 13 and / or the relation to associated text sources on the user interface 18 of the management program 6.

[0061] The large language model 11 can, for example, be based on models such as LLaMA, Granite, or another commercially available, partially customized model or model specifically trained for the inventive method. The large language model 11 runs on GPU servers. Graphics processors (GPUs) are optimized for parallel processing, which means they can perform multiple calculations simultaneously. The large language model 11 features parallel processing because many of the operations that must be performed when processing large amounts of text are parallel. This significantly accelerates the training and inference process. The large language model 11 performs extensive matrix operations. Graphics processors (GPUs) are particularly efficient at performing matrix operations, which leads to faster calculations.

[0062] The large language model 11 can be trained on large datasets, and accessing this data requires fast data processing capabilities. The high memory bandwidth and powerful computing capabilities of graphics processing units (GPUs) enable large amounts of data to be processed efficiently.

[0063] Graphics processing units (GPUs) can be easily integrated into server farms or clusters such as graphics processor servers 4, enabling high scalability for the deployment of the large language model 11 in large infrastructures 1

[0064] Deep learning models, including the large language model 11, benefit from the specialized computing capabilities of GPUs. GPUs are designed to perform the computations required for training and inference of neural networks more efficiently than conventional central processing units (CPUs). Overall, GPU servers enable more efficient and faster processing of large language models such as LLMs, which is crucial to their performance.

[0065] Rule-based text analysis can also be performed on central processing unit (CPU) servers, since unlike many machine learning models that are inherently parallelized, rule-based systems that are based on predefined rules often work serially, i.e., one after the other.

[0066] Rule-based systems like Rules 16 involve complex decision structures and require precise control flow. CPUs can handle such logical operations effectively. Compared to more sophisticated machine learning models, rule-based text analytics require less massive parallel processing and specialized computations. CPUs are powerful enough to handle these tasks.

[0067] The rules can be easily adapted to specific requirements by modifying or adding rules. CPUs offer the flexibility to efficiently execute different rules and do not rely on prior training.

[0068] Infrastructure 1 can be operated independently of the Internet in a LAN in order to meet the security requirements especially in the area of ​​new patent applications and patent searches.

[0069] The transfer 10 of the separated text parts 13 of the text 5 and the text source 12 to improve the separation can be carried out in a machine-readable format to a program 15 for rule-based analysis and a rule-based analysis can be carried out by the program 15 according to rules 16 stored on one of the storage systems 2 on a central processor server 3.

[0070] The stored text source 12 can be taken into account in the form of a text part to text source relation 19 to ensure the identical display of separated text parts 13 during a subsequent call.

[0071] Further processing of the separated text sections 13 can be carried out using a large language model 11, which can run on one or more graphics processor servers 4 to generate concepts 20, keywords, or umbrella terms. These can be stored and / or displayed on the user interface 18 in the management program 6. Alternatively, further processing of the separated text sections 13 can be carried out using artificial intelligence for natural language processing (NLP), which runs on one or more graphics processor servers 4, to generate concepts 20, keywords, or umbrella terms. These can then be stored and / or displayed on the user interface 18 in the management program 6.

[0072] The training of the large language model 11 can be carried out using data 21 which are accessible via the management program 6 or the database 7 on one or more of the storage systems 2.

[0073] However, the training of the large language model 11 can also be carried out using data in patent publications, in particular patent claims, which are accessible on storage devices via the Internet.

[0074] To refine the accuracy or fine-tune the large language model 11 with generative artificial intelligence, prompts, instructions in natural language or scripts and example texts can be used using one or more graphics processors 23 or graphics processing servers 4.

[0075] Multiple graphics processors 23 can be integrated into graphics processing servers 4.

[0076] Several central processors 22 can be integrated into central processor server 3.

[0077] The separated text parts 13 can be stored and / or displayed in a table 8 in the management program 6 using the large language model 11 with generative artificial intelligence, compared with other text parts 13 accessible via the management program 6 on one or more of the storage systems 2.

[0078] A degree of agreement between two or more text parts 13 to be compared can be calculated by using one or more graphics processors 23 or graphics processor servers 4 and can be the basis for relevance-related storage or display on the user interface 18 of the management program 6.

[0079] In the management program 6, the separation points 25 between the separated text parts 13 can be changed via user inputs and these changes can be used to further refine the accuracy of the large language model 11 with generative artificial intelligence, in particular via prompts, instructions or scripts.

[0080] In the management program 6, a comparison result 26 between the separated text parts 13 from different texts 5 can be changed or evaluated on the user interface 18 via user inputs, and these user inputs can be used to further refine the accuracy of the large language model 11 with generative artificial intelligence, in particular via prompts, instructions or scripts.

[0081] The data processing system 27 may comprise a plurality of processors adapted / configured to carry out the steps of the method explained above.

[0082] The computer program product 28, which can be operated on the infrastructure 1 or hardware 24, comprises instructions which, when the program is executed by a computer, cause the computer to carry out the method as described above.

[0083] Fig. 2 shows a further and possibly independent embodiment of Fig. 1, wherein the same reference numerals or component designations are used for the same parts as in the previous Fig. 1. To avoid unnecessary repetition, reference is made to the detailed description in the previous Fig. 1. The embodiment of Fig. 2 in the form of a computer-implemented method differs from the embodiment of Fig. 1 in that all elements and programs for the repeatable separation of text into parts that are correctly separated according to the literal meaning run in a common hardware 24 for artificial intelligence with at least one central processor 22 and a graphics processor 23. The described steps, which run on the central process server 3 in the method of Fig. 1, are carried out here on at least one central processor 22, corresponding to the steps for Fig.1 described steps are carried out accordingly.

[0084] The steps described, which run on the graphics process server 4 in the method of Fig. 1, are executed here on at least one graphics processor 23, corresponding to the steps described for Fig. 1. Fig. 3 shows a further and possibly independent embodiment, wherein the same reference numerals or component designations are used for the same parts as in the previous Fig. 1.

[0085] Fig. 3 shows a schematic representation of the separated text 13 with feature comparison and concepts 20 in table form, as it can be displayed on a user interface 18 of the management program 6.

[0086] Of course, other forms of representation such as lists or graphics can be used instead of tables

[0087] The separation points 25 are symbolized here by a line break in the table. The user can move or edit smaller parts of the text sections 13 so that they are assigned to the subsequent or previous text section 13 or feature.

[0088] Concepts 20 are displayed for individual text parts 13 or features, for example in the same table row but in an adjacent column or as an overlay or comment fields for a text part 13.

[0089] Additional functions such as evaluation and commenting on the comparison result 26 can be integrated.

[0090] Optionally, the text source 12 for a text part 13 can be displayed.

[0091] Furthermore, in a possibly independent embodiment, an input field 29 for issuing prompts and / or other instructions or commands or for entering examples for training the large language model 11 may be provided on the user interface 18 of the management program 6. Thus, the comparison results 26, the separation of the text parts 13, the concepts 20, or the rules 16 can be improved through user interaction.

[0092] Furthermore, in a possibly independent embodiment, the large language model 11 can be used for the automated summarization of texts, text parts, selected documents, or data 21. Furthermore, in a possibly independent embodiment, the large language model 11 can be used for the automated creation of texts, such as patent applications, based on a plurality of selected separated text parts 13 and / or concepts 20.

[0093] Furthermore, in a possibly independent embodiment, the management program 6 can display relationships or comparison results 26 based on separated text parts 13 to products, invention reports, documents, third-party intellectual property rights, patent publications or Internet search results created in the management program 6 and show overlaps or matches.

[0094] The embodiments show possible embodiments, whereby it should be noted at this point that the invention is not limited to the specifically illustrated embodiments thereof, but rather various combinations of the individual embodiments with each other are also possible and this possibility of variation lies within the skill of the person skilled in the art in this technical field due to the teaching of technical action by means of the objective invention.

[0095] The scope of protection is determined by the claims. However, the description and drawings must be used to interpret the claims. Individual features or combinations of features from the various embodiments shown and described may represent independent inventive solutions. The problem underlying these independent inventive solutions can be derived from the description.

[0096] All information on value ranges in this description is to be understood as including any and all sub-ranges thereof, e.g. the information 1 to 10 is to be understood as including all sub-ranges starting from the lower limit of 1 and the upper limit of 10, ie all sub-ranges begin with a lower limit of 1 or greater and end with an upper limit of 10 or less, e.g. 1 to 1.7, or 3.2 to 8.1, or 5.5 to 10.

[0097] For the sake of clarity, it should be noted that, for a better understanding of the structure, some elements have been shown not to scale and / or enlarged and / or reduced in size.

[0098] 25 Separation point

[0099] Infrastructure

[0100] 26 Comparison result

[0101] Storage system

[0102] 27 Data processing system

[0103] Central processor server

[0104] 28 Computer program product

[0105] Graphics processor server

[0106] 29 Input field

[0107] text

[0108] Management program

[0109] database

[0110] Table

[0111] Link

[0112] Transfer of large language model

[0113] Text source

[0114] Text part

[0115] handover

[0116] program

[0117] Regulate

[0118] handover

[0119] User interface

[0120] relation

[0121] concept

[0122] Data

[0123] central processor

[0124] graphics processor

[0125] Hardware

Claims

Patent claims 1. A method for the repeatable separation of text into correctly separated parts according to the literal meaning in an infrastructure (1) for artificial intelligence which includes one or more storage systems (2) and one or more central processor servers (3) and one or more graphics processor servers (4), the method comprising the following: Calling a text (5) to be separated, stored in a management program (6), by one of the central processor servers (3) from one of the storage systems (2), Transferring the text (5) to be separated in a machine-readable format to a large language model (11) with generative artificial intelligence, which runs on one or more graphics processor servers (4), Separating the text by understanding natural language through the trained large language model (11) on the one or more graphics processor servers (4), Storing a text source (12) into a separated text part (13) on one or more of the storage systems (2), Transfer of text parts separated by the large language model (11) (13) of the text (5) and the text source (12) in a machine-readable format to a program (15) which, in addition to processing natural language, also uses rules (16) stored on one of the storage systems (2) to improve the separation of the individual text parts (13), Transfer of the individual text parts (13) in a machine-readable format to the management program (6) which runs on a central processor server (3) and Storage and / or display of the separated text parts and / or the relation (19) to associated text sources.

2. A computer-implemented method for the repeatable separation of text into correctly separated parts according to the literal meaning in a hardware (24) for artificial intelligence, comprising the steps: Calling up a stored text (5) to be separated in a management program (6) by one of the central processors (22) from one of the storage systems (2), Transferring the text (5) to be separated in a machine-readable format to a large language model (11) with generative artificial intelligence, which runs on one or more graphics processors (23), Separating the text (5) by understanding natural language by the trained large language model (11) on the one or more graphics processors (23), Storing the text source (12) into a separated text part (13) on one or more of the storage systems (2), Transferring text parts (13) of the text (5) and the text source (12) separated by the large language model (11) in a machine-readable format to a program (15) which, in addition to processing natural language, also uses rules (16) stored on one of the storage systems (2) to improve the separation of the individual text parts (13), Transferring the individual text parts (13) in a machine-readable format to the management program (6) which runs on a central processor (22) and Saving and / or displaying the separated text parts (13) and / or the relation (19) to associated text sources.

3. Method according to claim 1 or 2, wherein the separated text parts (13) of the text (5) and the text source (12) are transferred to a program (15) for rule-based analysis in a machine-readable format to improve the separation, and a rule-based analysis is carried out by the program (15) according to rules (16) stored on one of the storage systems (2) on a central processor server (3) or central processor (22).

4. Method according to one of claims 1 to 3, wherein the stored text source (12) is taken into account in the form of a text part to text source relation (19) to ensure the identical display of separated text parts (13) upon further call by the management program (6).

5. Method according to one of claims 1 to 4, wherein further processing of the separated text parts (13) with a large language model (11) which runs on the one or more graphics processor servers (4) or graphics processors (23) is used to generate concepts (20), keywords or general terms and to store and / or display the same in the management program (6).

6. The method according to one of claims 1 to 4, wherein the further processing of the separated text parts (13) with artificial intelligence for natural language processing (NLP), which runs on the one or more graphics processor servers (4) or graphics processors (23), is used to generate concepts (20), keywords or umbrella terms and to store and / or display the same in the management program (6).

7. The method according to any one of claims 1 to 6, wherein the training of the large language model (11) is carried out using data (21) which are accessible via the management program (6) on one or more of the storage systems (2).

8. Method according to one of claims 1 to 7, wherein the training of the large language model (11) is carried out on the basis of patent publications, in particular patent claims, which are accessible via the Internet on Internet storage devices.

9. The method according to any one of claims 1 to 8, wherein a refinement of the accuracy of the large language model (11) is carried out with generative artificial intelligence via prompts, instructions in natural language or scripts and example texts using one or more graphics processor servers.

10. Method according to one of claims 1 to 9, wherein the separated text parts are stored and / or displayed in a table (8) in the management program (6) using the large language model (11) with generative artificial intelligence, compared with other text parts (13) accessible via the management program (6) on one or more of the storage systems (2).

11. Method according to one of claims 1 to 10, wherein a degree of correspondence between two or more text parts (13) to be compared is calculated by using one or more graphics processor servers (4) or graphics processors (23) and is the basis for relevance-related storage or representation.

12. Method according to one of claims 1 to 11, wherein in the management program (6) the separation points (25) between the separated text parts (13) can be changed via user inputs and these changes are used to further refine the accuracy of the large language model (11) with generative artificial intelligence, in particular via prompts, instructions or scripts.

13. Method according to one of claims 1 to 12, wherein in the management program (6) a comparison result (26) between the separated text parts (13) can be changed or evaluated via user inputs and these user inputs are used to further refine the accuracy of the large language model (11) with generative artificial intelligence, in particular via prompts, instructions or scripts.

14. A data processing system comprising a plurality of processors adapted / configured to carry out the steps of the method according to any one of claims 1 to 13.

15. A computer program product comprising instructions which, when executed by a computer, cause the computer to carry out the method according to any one of claims 1 to 13.

Citation Information

Patent Citations

  • Apparatus and method for automated and assisted patent claim mapping and expense planning

    US11714839B2