Federal queries for multiple data silos associated with a product
Fine-tuning a pre-trained language model with semantic metadata from ontologies and vocabularies generates federated queries, addressing the challenge of accessing multiple data silos and improving data integration and collaboration within organizations.
Patent Information
- Application Number
- DE102024205140
- Authority / Receiving Office
- DE · DE
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-06-04
- Publication Date
- 2025-12-04
AI Technical Summary
Existing techniques for generating queries are limited in their ability to access and retrieve data from multiple data silos within an organization, often requiring manual intervention and failing to provide a flexible, automated solution for data integration across disparate systems.
A computer-implemented procedure for fine-tuning a pre-trained language model using semantic metadata from ontologies and vocabularies to generate federated queries, enabling seamless access to multiple data silos by converting natural language requests into domain-specific queries like SPARQL.
Enables automated, precise data retrieval from multiple data silos, providing a unified view of product-related data without the need for centralized repositories, enhancing data integration and collaboration across organizational departments.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[0001] Several examples of the revelation relate generally to the fine-tuning of a pre-trained language model. Specifically, several examples relate to the fine-tuning of a pre-trained language model using knowledge associated with a domain-specific ontology and / or knowledge associated with one or more vocabularies. The language model can then be used to generate a federated query associated with a hardware and / or software product from a prompt.
[0002] Analyzing company-wide data supports informed decision-making and a more holistic view of hidden opportunities or threats. However, different departments within a company, such as finance, administration, human resources, marketing, and others, require access to different information to perform their tasks. These departments tend to store their data in separate locations, often referred to as data or information silos. Data silos create barriers to information sharing and collaboration between departments.
[0003] There are several techniques for retrieving data from different data silos within an organization, such as a company. These techniques vary depending on factors like the type of data silos, the structure of the data, and the integration requirements.
[0004] For example, SPARQL (a recursive acronym for SPARQL Protocol And RDF Query Language) is a Resource Description Framework (RDF) query language—that is, a semantic query language for databases—that allows users to retrieve and manipulate data stored in the Resource Description Framework (RDF) format. It was standardized by the RDF Data Access Working Group (DAWG) of the World Wide Web Consortium (W3C) and is recognized as one of the key technologies of the Semantic Web.
[0005] SPARQL is a query language used to express queries across different data sources, regardless of whether the data is stored natively as RDF or viewed as RDF via middleware. To retrieve desired data from various data sources or data silos using SPARQL, a SPARQL query must be created. In recent years, the conversion of natural language questions into SPARQL queries has gained popularity. Several techniques are known for generating SPARQL queries. For example, non-patent literature—Rony, Md Rashad Al Hasan et al. “Sgpt: A generative approach for SPARQL query generation from natural language questions.” IEEE Access 10 (2022): 70712–70723. [1]—reveals a generative approach for generating SPARQL queries from natural language questions.A new approach has been proposed, called SGPT, which combines the advantages of end-to-end and modular systems and leverages recent advances in large-scale language models.
[0006] Non-patent literature - Luz FF, Finger M. Semantic parsing natural language into SPARQL: improving target language representation with neural attention. arXiv Preprint arXiv:1803.04329. March 12, 2018 [2] - reveals semantic parsing of natural language in SPARQL. Semantic parsing is the process of mapping a sentence in natural language into a formal representation of its meaning. A neural network approach was used to transform a sentence in natural language into a query against an ontology database in the SPARQL language.
[0007] Non-patent literature - Wang R, Zhang Z, Rossetto L, Ruosch F, Bernstein A. Nlqxform: A language model-based question to sparql transformer. arXiv preprint arXiv:2311.07588. 8 Nov 2023. [3] - reveals a language model-based question to the SPARQL transformer.
[0008] Non-patent literature - Panchbhai A, Soru T, Marx E. Exploring Sequence-to-Sequence Models for SPARQL Pattern Composition. arXiv Preprint arXiv:2010.10900. October 21, 2020. [4] - reveals the use of sequence-to-sequence models as a viable and promising option for transforming long utterances into complex SPARQL queries. Non-patent literature - Nikolov, Andriy, et al. “Ephedra: Efficiently combining RDF data and services using SPARQL federation.” Knowledge Engineering and Semantic Web: 8th International Conference, KESW 2017, Szczecin, Poland, 8–10 October 2017. November 2017, Protocol 8. Springer International Publishing, 2017. [5] - reveals a SPARQL federation engine designed to handle hybrid queries, called Ephedra.Ephedra provides a flexible declarative mechanism for including hybrid services in a SPARQL federation and implements a number of optimization techniques for static and runtime queries to improve the performance of hybrid SPARQL queries.
[0009] Non-patent literature - Heling L, Acosta M. Federated SPARQL query processing over heterogeneous linked data fragments. In Proceedings of the ACM Web Conference 2022, 25 April 2022 (pp. 1047-1057).[6] - reveals an interface-aware framework for processing SPARQL queries over federations with heterogeneous linked data fragment interfaces.
[0010] The techniques disclosed in the non-patent literature demonstrate the feasibility of calling hybrid federal services from a SPARQL query, i.e., they improve the SPARQL query with hybrid federal services.
[0011] Non-patent literature – Sima, Ana Claudia, et al. “Enabling semantic queries across federated bioinformatics databases.” Database 2019 (2019): baz106. [7] – reveals an ontology-based federated approach to data integration. This approach uses natural language templates (e.g., user interface forms) for federated query construction, aiming to relieve users of the need to manually construct SPARQL queries in the biomedical field. However, with this solution, users still need to manually fill out the template, understand its content, and comprehend the semantic metadata.
[0012] In addition, classic natural language processing techniques were also used to address the construction of SPARQL queries from natural language.
[0013] For example, non-patent literature – Sander M, Waltinger U, Roshchin M, Runkler T. Ontology-based translation of natural language queries to SPARQL. In2014 AAAI Fall Symposium Series 24 Sept. 2014. [8] – reveals an implemented approach for transforming sentences in natural language into SPARQL using background knowledge from ontologies and lexicons.
[0014] US patent application 2010 / 0185643 A1 discloses techniques for the automated generation of queries for querying ontologies.
[0015] As can be seen from the above, various techniques for generating queries are known. However, all such techniques are limited in that they cannot flexibly generate queries to access a multitude of data silos. Typically, the techniques revealed above only work well when generating a query for a given data silo. Abstraction to other data silos is either not possible or only possible to a limited extent.
[0016] Therefore, there is a need for advanced techniques to access multiple data silos and retrieve desired data from them. Specifically, there is a need for advanced techniques to automatically generate a precise natural language query to retrieve desired data from multiple data silos associated with a product.
[0017] This need is met by the features of the independent claims. The features of the dependent claims define embodiments.
[0018] A computer-implemented procedure is provided for fine-tuning a pre-trained language model to generate a federated query associated with a product from a request. The procedure involves obtaining first semantic metadata associated with an ontology representing concepts of the product, and second semantic metadata associated with a vocabulary describing those concepts. The procedure further includes fine-tuning the pre-trained language model based on the first semantic metadata associated with the ontology, and again based on the second semantic metadata associated with the vocabulary.
[0019] Another computer-implemented procedure is provided. This procedure involves receiving a request describing desired data associated with a product. The procedure further involves generating, based on the request, a federated query associated with the desired data, using a pre-trained language model fine-tuned by the computer-implemented procedure described above.
[0020] A computer device comprising a processor and memory is provided. After loading and executing program code from memory, the processor is designed to perform a procedure for fine-tuning a pre-trained language model to generate a federated query associated with a product from a request. The procedure involves obtaining first semantic metadata associated with an ontology representing concepts of the product, and second semantic metadata associated with a vocabulary describing the product's concepts. The procedure further involves fine-tuning the pre-trained language model based on the first semantic metadata associated with the ontology and, further, based on the second semantic metadata associated with the vocabulary. A computer program product is provided.A computer program or a computer-readable storage medium containing program code is provided. The program code can be executed by at least one processor. Execution of the program code causes the at least one processor to perform a procedure for fine-tuning a pre-trained language model to generate a federated query associated with a product from a request. The procedure includes obtaining first semantic metadata associated with an ontology representing concepts of the product, and second semantic metadata associated with a vocabulary describing the concepts of the product. The procedure further includes fine-tuning the pre-trained language model based on the first semantic metadata associated with the ontology and, furthermore, based on the second semantic metadata associated with the vocabulary.
[0021] It is understood that the aforementioned features and those to be discussed below can be used not only in the respective combinations specified, but also in other combinations or in isolation, without deviating from the scope of protection of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS Fig. Figure 1 is a flowchart of a process according to various examples. Fig. Figure 2 schematically illustrates a processing pipeline according to various examples. Fig. Figure 3 schematically illustrates a workflow according to various examples. Fig. Figure 4 schematically illustrates a processing device according to various examples.
[0022] Some examples in this disclosure generally provide several circuits or other electrical devices. All references to the circuits and other electrical devices and the functionality they each provide are not intended to be limited to what is illustrated and described herein. While certain designations may be assigned to the various disclosed circuits or other electrical devices, such designations are not intended to limit the operational scope of the circuits and other electrical devices. Such circuits and other electrical devices may be combined with one another and / or separated in any way, based on the specific type of electrical implementation desired.It is acknowledged that any circuit or other electrical device disclosed herein may include any number of microcontrollers, a graphics processing unit (GPU), integrated circuits, memory devices (for example, FLASH, random access memory (RAM), read-only memory (ROM), electrically programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or other suitable variants thereof), and software, which interact to perform the operation(s) disclosed herein. Furthermore, any one or more of the electrical devices may be configured to execute program code embodied in a non-volatile, computer-readable medium programmed to perform any number of the disclosed functions.
[0023] The following sections describe embodiments of the disclosure in detail with reference to the accompanying drawings. It is understood that the following description of embodiments is not to be understood in a limiting sense. The scope of protection of the disclosure is not to be limited by the embodiments described below or by the drawings, which are intended only to be illustrative.
[0024] The drawings are to be regarded as schematic representations, and elements illustrated in the drawings are not necessarily shown to scale. Rather, the various elements are depicted in such a way as to make their function and general purpose obvious to those skilled in the art. Any connection or coupling between functional blocks, devices, components, or other physical or functional units shown in the drawings or described herein may also be implemented by an indirect connection or coupling. Coupling between components may also be established via a wireless connection. Functional blocks may be implemented in hardware, firmware, software, or a combination thereof.
[0025] The following techniques for accessing multiple data silos and retrieving desired data using a federated query generated from a request are revealed. In each of multiple data silos, data is stored in separate, isolated repositories within an organization, and these repositories are managed and accessed independently. Data silos can arise for a variety of reasons, such as the use of disparate, non-interoperable technology systems, organizational structures that restrict data access to specific groups, or the historical growth of an organization.
[0026] For example, multiple data silos can be associated with different entities within a large organization. For instance, multiple data silos can be associated with different entities involved in a product's production process. For example, a first entity might participate in R&D activities to plan and develop a product; a second entity might be responsible for sourcing parts to build the product; a third entity might be responsible for validating the product's source parts; a fourth entity might be responsible for manufacturing the product; and a fifth entity might be responsible for quality control of the manufactured product. This is just one example. Other examples are possible. For instance, different data silos could be associated with different manufacturing machines on an assembly line.For example, different data silos might be associated with different departments within a hospital that collaborate to provide a diagnosis for a patient. In another example, multiple data silos might be associated with different components of a product, such as an MRI scanner, a CT scanner, and so on. Typically, different components of a product are developed by different people within an organization and / or manufactured on different production lines. Consequently, there is a tendency for these different entities to maintain isolated data silos that must be accessed using a federated query.
[0027] A product can be a hardware product or a software product. Examples of products include hardware-software products. A technical product may be subject to the techniques disclosed herein. A medical imaging device is an example of a product. Products include, for example, transportation or mobility products such as vehicles, trains, airplanes, etc. Medical devices, laboratory equipment, and testing machines are further examples. Energy conversion devices such as wind turbines, gas turbines, generators, power plants, nuclear power plants, coal-fired power plants, etc., are further examples. Products using green technology such as solar cells, fuel cells, and batteries are further examples.
[0028] Such products can contain multiple components. All such products can involve multi-stage manufacturing processes. All such products can be developed by multiple companies and / or multiple entities within a company.
[0029] In general, a federated query, as used in this disclosure, is a query that spans multiple data silos, sources, or repositories distributed across different locations or systems. Instead of querying a single centralized database, a federated query allows a user to retrieve data from multiple sources / data silos in a unified manner. Federalized queries enable organizations to access and integrate product-related data from heterogeneous sources in a unified way, providing a comprehensive view of the data landscape without the need for a centralized data repository.
[0030] According to W3C SPARQL standards, SPARQL 1.1, for example, defines the syntax and semantics for executing queries distributed across different SPARQL endpoints, i.e., federated queries. Here, federated queries involve a Protocol and Resource Framework Query Language query. Non-patent literature - Rakhmawati NA, Umbrich J, Karnstedt M, Hasnain A, Hausenblas M. Querying over Federated SPARQL Endpoints—A State of the Art Survey. arXiv Preprint arXiv:1306.1723. June 7, 2013. - Reveals a summary of techniques for querying over federated SPARQL endpoints.
[0031] According to this revelation, a request is a text in natural language that describes the task for an artificial intelligence (AI) to perform. For example, a request for a text-to-text language model could be a query, a command, or a longer statement that includes context, instructions, and a conversational history.
[0032] According to this revelation, the federated query can be generated using a pre-trained language model. The pre-trained language model includes large language models and small language models.
[0033] In general, large language models are highly sophisticated AI models capable of understanding and generating human-like text across various languages and topics. These models are built using deep learning techniques, such as transformer architectures, and are trained on massive datasets consisting of billions or even trillions of words. Large language models have millions or even billions of parameters, which are the internal variables the model learns during training. These parameters enable the model to grasp complex relationships between words and generate coherent and contextually relevant text. Examples of large language models include GPT (Generative Pre-Trained Transformer) models developed by OpenAI, BERT (Bidirectional Encoder Representations from Transformers) developed by Google, T5 (Text-to-Text Transfer Transformer) developed by Google, and others.Small language models, on the other hand, are less complex versions of large language models, typically with fewer parameters and trained on smaller datasets. For example, small language models may have fewer than one million parameters, often in the range of thousands to hundreds of thousands. Small language models can be trained on smaller datasets taken from subsets of larger datasets or curated to focus on specific domains or topics. Examples of small language models include Phi-2, as revealed in non-patent literature - Javaheripi, Mojan, et al. “Phi-2: The surprising power of small language models.” Microsoft Research Blog (2023). [9], and Orca 2, as described in non-patent literature - Mitra A, Del Corro L, Mahajan S, Codas A, Simoes C, Agarwal S, Chen X, Razdaibiedina A, Jones E, Aggarwal K, Palangi H. Orca 2: Teaching small language models how to reason. arXiv preprint arXiv:2311.11045. 18 Nov 2023.
[11] reveals, include. As a general rule, language models typically use a deep learning architecture that includes multiple layers, each designed to process and transform input data through a series of mathematical operations. These language models typically use transformer architectures. A transformer architecture uses so-called attention mechanisms to weight the importance of different words in a sequence; this enables context processing. Layers are used that perform linear transformations followed by nonlinear activations. Each layer is associated with a set of weights, which are adjustable parameters that the model optimizes during training through backpropagation and gradient descent in machine learning.The learning process involves adjusting the network's weights to minimize a loss function that quantifies the difference between the model's predictions and the actual data.
[0034] To enable the generation of a federated query, such as a federated SPARQL query, for a specific use case or domain, the pretrained language model may need to be fine-tuned based on use case-specific or domain-specific information or knowledge. The following reveals techniques for fine-tuning a pretrained language model to generate a federated query associated with a product. The federated query is generated from a request, such as one describing desired data / information associated with the product. The pretrained language model is fine-tuned based on first semantic metadata associated with an ontology representing concepts of the product, and second semantic metadata associated with a vocabulary describing those concepts.
[0035] In general, the product can be any industrial or consumer product. For example, it could be an electrical appliance, a car, or a bus. More specifically, it could be a projection radiography scanner, a magnetic resonance imaging (MRI) scanner, a computed tomography (CT) scanner, a positron emission tomography (PET) scanner, a single-photon emission tomography (SPET) scanner, or an ultrasound scanner.
[0036] According to this revelation, semantic metadata can refer to descriptive information about data encoded using semantic technologies and standards. Semantic metadata can include structured knowledge representations that enable automated inference and reasoning. For example, semantic metadata can be associated with Semantic Web technologies such as RDF, OWL (Web Ontology Language), or SPARQL, which can provide the foundations for encoding, publishing, and querying semantic metadata on the web.
[0037] In general, an ontology is a formal, explicit specification of a conceptualization. It is a way of representing knowledge about a particular domain by defining the types of entities that exist within it and the relationships between them. Ontologies can be used to structure and organize knowledge systematically and in a machine-readable way. An ontology representing the concepts of a product can provide a specification of the concepts, entities, and relationships within a domain associated with the product. Such an ontology can use vocabulary and / or syntax to describe knowledge associated with the product. Generally, ontologies define different types of entities within a domain. An ontology includes, for each domain (or specifically for each product), classes, instances, attributes, and relationships.Classes represent categories or types of things, instances are individual members of those categories, attributes describe properties or characteristics of entities, and relationships specify connections between entities. An example ontology for an "MRI scanner" product might include several classes such as "bias field magnet system," "gradient coils," "RF coils," and "control and imaging software." Attributes of the "bias field magnet system" class, for example, might include properties such as "field strength," "coolant type," and "magnet architecture." Attributes of the "gradient coil" class might include properties such as "maximum gradient strength," "rise rate," or "design." The relationships between these classes are generally complex, but to give some examples, the "control and imaging software" must interact with each of the other classes to implement an imaging protocol.Furthermore, the "gradient coils" are used to encode a magnetic field gradient according to the imaging protocol, which is tailored to the bias magnetic field applied by the "bias field magnet system," as defined by the respective relationship between the "gradient coils" class and the "bias field magnet system" class. For example, a given MRI scanner type may be compatible with multiple RF coil systems, thus specifying different instances of the "RF coils" class.
[0038] In general, an ontology can be represented as a graph data structure. For example, in a graph data structure representing the ontology of an "MRI scanner," provided above as an example, each class (such as "bias field magnet system," "gradient coils," "RF coils," and "control and imaging software") can be visualized as a node. Attributes of these classes, such as "field strength" for the "bias field magnet system" or "maximum gradient strength" for the "gradient coils," can be represented as properties associated with the respective nodes. The relationships between these classes, such as the interaction of the "control and imaging software" with other parts or the dependency of the "gradient coils" on the "bias field magnet system," are represented as edges connecting the relevant nodes.
[0039] According to this revelation, a vocabulary can comprise a structured collection of terms (words or phrases) used to describe concepts, entities, or relationships within a specific domain or subject area, such as the domain associated with the product. It can include a set of terms that provides a common understanding of the terminology used in a particular field, thereby enabling consistent and unambiguous communication between users and systems.
[0040] As another concrete example, the vocabulary for an MRI scanner might include entries such as "biased field magnet system," "field strength," "coolant type," and "magnet architecture." Each entry could include descriptive text, for example: "Field strength refers to the intensity of the magnetic field generated by the magnet, typically measured in Tesla"; and "Coolant type is the type of coolant used, often liquid helium, to maintain the temperature of the magnet system"; and "Magnet architecture describes the structural design of the magnet system, such as superconducting or permanent magnet designs." In the "gradient coils" category, entries might include "maximum gradient strength," "rise rate," and "design."The following explanation can be given here: "Maximum gradient strength measures the highest strength of the gradient field that can be achieved by the coils, given in milliteslas per meter"; "Ride rate is the speed at which gradient coils can change the magnetic field, measured in tesla per meter per second"; "Design specifies the physical configuration of the coils, either cylindrical or planar." For the class of "RF coils," terms such as "coil type," "number of channels," and "material" can be used.Some example explanations would be: “Coil type classifies the coil based on its application, such as head coil or body coil”; “Number of channels indicates the number of independent channels in the coil, which affects image quality and acquisition speed”; and “Material refers to the type of conductive material used, typically copper or silver, which affects signal sensitivity and performance.”
[0041] It is understood that the entries in the ontology can correspond to the entries in the vocabulary. Thus, the same concepts represented by the ontology are also described by the vocabulary.
[0042] According to various examples, fine-tuning of the pre-trained language model based on the first set of semantic metadata associated with the ontology, and further based on the second set of semantic metadata associated with the vocabulary, can be performed using standard neural network training procedures such as semi-supervised, unsupervised, or supervised learning. Such techniques are generally known to experts. The specific training technique used is not relevant to the disclosure, and existing techniques can be readily employed.
[0043] More generally, the fine-tuning of a language model builds upon pre-training. Initially, during pre-training, the language model learns general language patterns from general text data—that is, text data not restricted to a specific domain or product. This is achieved by adjusting the model's internal parameters, or weights, through a process called backpropagation. Backpropagation iteratively minimizes a function called the loss, which measures the discrepancy between the model's predictions and the actual data. This training helps the model develop a broad language understanding that is not specific to any particular domain. In the subsequent fine-tuning phase, the pre-trained model is further refined on a smaller, domain-specific dataset; in this case, the first and second semantic metadata.Fine-tuning allows the language model to adjust its previously learned weights to perform tasks specific to this product / domain effectively. Through the continuous use of backpropagation to minimize loss with this new dataset, the model becomes more specialized, thereby improving its ability to generate or interpret texts that are more closely aligned with domain-specific requirements.
[0044] Fig. Figure 1 illustrates aspects of Procedure 1000 for fine-tuning a pre-trained language model to generate a federated query associated with a product, where dashed blocks indicate optional processing steps. The federated query is generated from a request, such as a request describing desired data / information associated with the product, from multiple data silos associated with different product-related components, manufacturing steps, and / or organizational departments. The pre-trained language model is fine-tuned based on first semantic metadata associated with an ontology representing concepts of the product, and further based on second semantic metadata associated with a vocabulary describing the product's concepts. Details of Procedure 1000 are described below.
[0045] Block 1100: Receives first semantic metadata associated with an ontology representing concepts of the product, and second semantic metadata associated with a vocabulary describing the concepts of the product.
[0046] Block 1200: Fine-tuning the pre-trained language model based on the first set of semantic metadata associated with the ontology and, furthermore, based on the second set of semantic metadata associated with the vocabulary. Optionally or additionally, the procedure at Block 1100 can also include obtaining third set of semantic metadata associated with at least one configured federated service. Each of the at least one configured federated service is mapped to a data silo that stores data associated with the product. Accordingly, the fine-tuning of the pre-trained language model at Block 1200 can also be based on the third set of semantic metadata associated with the at least one configured federated service.
[0047] In general, a federated service can be associated with multiple service endpoints, such as physical devices connected to a network system, like mobile devices, desktop computers, virtual machines, embedded devices, and servers. An endpoint can include devices located at a specific location or a specific Uniform Resource Locator (URL) where a service can be accessed or data can be queried.
[0048] For example, the multiple service endpoints might include devices of a federated system or a distributed database management system. Within a federated system, a single SPARQL query can access services or data distributed across multiple endpoints or data sources. The federated SPARQL query can have the ability to use multiple SERVICE keywords to query and merge data from different endpoints.
[0049] According to various examples, if the data stored at an endpoint is not in RDF, the endpoint can convert the data to RDF using a service, such as middleware, which can be a universal service acting as an intermediary between systems, facilitating common communication. Such a federated service can define services associated with the product, such as information retrieval, manual generation, and maintenance and / or repair services. This allows for the capture of specific use cases associated with the product, and the query can then be tailored to these use cases.
[0050] For example, a federated SPARQL query for MRI scanners can use the keyword "SERVICE", as defined by W3C standards, to invoke different services, such as an MRI Open Platform Communications (OPC) Unified Architecture (UA) server service, a Product Lifecycle Management (PLM) system service, and / or a hospital data service.
[0051] The MRI OPC UA server service can connect to the MRI system's OPC UA server to retrieve real-time temperature data from the coil system and the hourly usage rate of an MRI scanner. The MRI OPC UA server service can expose endpoints for querying temperature sensor readings and other relevant data.
[0052] The PLM system service can interact with the PLM system to retrieve technical data, such as design data (including wiring data) and simulation data, related to the MRI coil system. The PLM system service can provide endpoints for accessing this information in a structured format. The hospital data service can collect temperature data from sensors installed throughout the hospital environment. The hospital data service can provide endpoints for querying temperature readings from various locations within a specific hospital.
[0053] A federated query associated with an MRI scanner may include variables to be retrieved from one or more services, such as an MRI Open Platform Communications (OPC) Unified Architecture (UA) server service, a Product Lifecycle Management (PLM) system service, and / or a hospital data service.
[0054] Results from the federated query associated with the MRI scanner can include real-time temperature of the MRI coil system, MRI scanner design data, and / or ambient temperature within a hospital environment. All results can be integrated into a single, coherent result set.
[0055] In block 1100, procedure 1000 can further include obtaining one or more problems and one or more corresponding federated ground truth queries. This allows for the acquisition of exemplary ground truth queries associated with prompts. This, in turn, enables the creation of input-output pairs of training data for subsequent fine-tuning of the language model.
[0056] According to various examples, the fine-tuning of the pre-trained model can be based on one or more vector embeddings of the semantic metadata. For example, the one or more vector embeddings can be generated based on the first set of semantic metadata associated with the ontology, and further based on the second set of semantic metadata associated with the vocabulary. As another example, the one or more vector embeddings can be generated based on the first set of semantic metadata associated with the ontology, further based on the second set of semantic metadata associated with the vocabulary, and additionally based on the third set of semantic metadata associated with the at least one configured federal service. A vector embedding is created from input data—here, for example,a graph representation of the ontology—determined by a series of predefined, structured computational steps. Each embedding is a dense, fixed-size numerical vector. Vector embedding generation may involve preprocessing to appropriately normalize and format the input data. This can include tokenization and normalization for the text in the vocabulary. The processed data is then fed into a predefined model, such as a neural network, trained to capture the data's distinctive features. Depending on its architecture, this model uses layers of neurons to apply nonlinear transformations to the input, adjusting internal weights. The model's output layer generates the vector embedding, which represents the input data in a high-dimensional space. This vector embedding is a representation of the input data, for example,of the ontology or vocabulary, in a compact vector space enables efficient storage, processing and analysis.
[0057] Optionally or additionally, procedure 1000 at block 1300 can further include identifying at least one change associated with any ontology, vocabulary, or configured federal service. Fine-tuning of the pre-trained language model is triggered based on the identification of this at least one change.
[0058] In particular, the ontology, vocabulary, and / or other information can be monitored, i.e., repeatedly checked for changes. Upon detection of such a change, retraining can be triggered. Retraining can generally be performed in batches, allowing it to be determined whether a certain number of changes—for example, a number of changes exceeding a threshold—has occurred before retraining is triggered. By implementing such adaptive retraining, which is triggered automatically, it can be ensured that the pre-trained language model is kept up-to-date and adapts to changes in the different data silos. This ensures that the federated query remains current and enables accurate information retrieval. At the same time, the workflow of users associated with the different data silos is not disrupted.Their work style can remain unaffected.
[0059] The techniques described above involve triggering the fine-tuning of a language model based on the detection of, for example, a change in the ontology. Alternatively or additionally, such fine-tuning can also be triggered based on a predefined timing schedule. For instance, the fine-tuning can be triggered periodically or at specific trigger dates / times. This allows the language model to remain up-to-date even when processing large datasets, making it difficult to monitor changes. The required computational resources can be reduced.
[0060] According to various examples, the procedure 1000, prior to fine-tuning the pre-trained language model based on the first semantic metadata associated with the ontology and further based on the second semantic metadata associated with the vocabulary, can also include, at block 1500, fine-tuning or modifying the architecture of the pre-trained language model so that it is, for example, better suited to a specific task, use case, or domain. For example, such fine-tuning of the architecture can include adding task-specific layers and / or adjusting hyperparameters and / or freezing certain layers to prevent them from being updated during the fine-tuning process.
[0061] Optionally or additionally, after fine-tuning the pre-trained language model in block 1200, the procedure can further include validating and / or evaluating the fine-tuned model on a validation dataset or an evaluation dataset in block 1600. This can include supervised validation steps. An evaluation can include benchmarking based on a test dataset. For example, several fine-tuned language models (e.g., using statistical variations or specific variations of the training data) can be benchmarked and the benchmarks compared.
[0062] According to various examples, after fine-tuning the pre-trained language model, procedure 1000 can repeat step 1300. If at least one further change associated with one of the ontology, vocabulary, and at least one configured federal service is determined, procedure 1000 can repeat blocks 1100 and 1200.
[0063] If no further change associated with one of the ontology, vocabulary, and at least one configured federal service is determined, the procedure can perform 1000 Block 1400, i.e., provide the fine-tuned language model.
[0064] After deploying the fine-tuned language model at block 1400, a request can be obtained that describes the desired product-associated data. For example, a plain-language request can be obtained specifying the generation of a manual for a particular component or set of components of a product. The request can specify a certain product malfunction and seek help to mitigate this malfunction. Then, based on the request, a federated query associated with the desired data can be generated using the pre-trained language model, which has been fine-tuned based on the techniques revealed above. Based on this federated query, the desired data can be retrieved from the multiple data silos associated with the product, such as those associated with different product components.
[0065] Fig. Figure 2 schematically illustrates aspects of a data processing pipeline according to various examples. The data processing pipeline includes several modules 3005, 3010, 3015, 3400, and 3510, which are coupled together. Module 3005 is a natural language interface module that provides a graphical user interface 3105 to a user 3699. The user 3699 can input a natural language request 3605 (e.g., by typing or speech transcription) via the graphical user interface 3105 and receive a corresponding response 3610 (e.g., a text response). An inference application programming interface (API) 3110 is provided within the natural language application module 3005. The inference API 3110 can access the provided language model 3210, which was fine-tuned in a fine-tuning process 3205 of the fine-tuning software module 3010.Details regarding the fine-tuning process were previously discussed in connection with Procedure 1000. Fig. 1, in particular block 1500, is discussed. In particular, the fine-tuning process 3205 accesses an ontology 3305 as metadata for training, e.g., triggered by monitoring 3206 on changes in the ontology 3305 (see Fig. 1: Block 1300). The ontology 3305 is provided within an ontology module 3015, for example, as a graph data structure. The fine-tuning process 3205 then provides the updated version of the language model 3210. The ontology module 3015 also includes a SPARQL engine 3310, which communicates with the graphical user interface 3105 of the natural language module 3005. The SPARQL engine 3310 accesses one or more federated services of a federated services submodule 3315. The queries delivered to the data layer module 3500 via an optional access layer module 3400 are typically service-specific. For example, the query to automatically generate a manual may be significantly different from a query for a maintenance / repair manual. They can be determined in a service-specific manner using the language model 3210.The data layer module 3500 includes an SQL database 3505 and a REST API 3510, as well as several data silos 3515 and 3520. Based on such a query, data is returned which is used to determine the response or reaction 3610, which is then output to the user.
[0066] Fig. Figure 3 schematically illustrates a workflow according to various examples. The workflow is associated with aspects related to fine-tuning a pre-trained language model to generate a federated query associated with a product, and with aspects related to generating, based on a prompt, a federated query associated with desired data, using the pre-trained language model that has been fine-tuned. The workflow can be implemented through the several modules of Fig. 2 will be carried out.
[0067] The workflow involves a fine-tuning process for periodically refining a pre-trained language model, such as a sequence-to-sequence language model, by feeding in first semantic metadata associated with an ontology (Box 4001) representing product concepts, and second semantic metadata associated with a vocabulary (Box 4002) describing the product concepts. The ontology may include enterprise ontologies in Shapes Constraint Language (SHACL), a W3C standard language for describing resource RDF graphs. The vocabulary may include domain-specific vocabularies, each represented using the Simple Knowledge Organization System (SKOS).By using semantic metadata associated with enterprise ontologies and semantic metadata associated with domain-specific vocabularies, the generation of federated queries, such as SPARQL queries, can be tailored to the enterprise context. Specifically, in Box 4001, one or more ontologies associated with one or more products are maintained by a subject matter expert. In Box 4002, one or more vocabularies associated with the respective product are maintained. Box 4003 optionally contains descriptions of configured federated services for accessing multiple data silos. The data structures maintained in Boxes 4001, 4002, and 4003 can be used as semantic metadata for training the large language model.Accordingly, in Box 4004, a fine-tuning process for refining the language model can detect any changes in the data maintained in Boxes 4001, 4002, and / or 4003. For example, it can detect whether a new data structure becomes available or whether an existing data structure is updated. In Box 4005, the respective metadata retrieved upon detection of a change can be encoded in a vector embedding. This results in an unlabeled dataset in Box 4006, which can be used in Box 4011 as input for a fine-tuning process to refine the language model.More generally, the language model is trained on unsupervised datasets primarily through a process called unsupervised learning, in which the language model learns to predict parts of the input data from other parts of the same data (here, the vector embeddings from Box 4005) without explicit external labels or annotations. In the context of text, this often involves predicting the next word given the preceding words in a sentence. The unlabeled dataset obtained at Box 4006 can optionally be supplemented with an annotated dataset obtained at Box 4008, for example, from corresponding pairs of prompts and SPARQL queries defined—for example, manually by a data engineer—at Box 4007. The language model is obtained at Box 4010. For example, the language model can be an open-source language model.After fine-tuning the language model in Box 4012, the refined language model can be saved and evaluated and / or validated in Box 4013, e.g., by a data engineer. If the validation and / or evaluation results are positive, the model is made available in Box 4014. Then the inference API, which was previously used in conjunction with... Fig. 2: Inference API 3110 was discussed, Box 4015 accesses the provided language model, so that Box 4016 contains a request that is made via a graphical user interface of a natural language application - as previously in conjunction with Box 3105 in Fig. 2 discussed - is obtained, into which the finely tuned language model can be fed. The SPARQL query thus obtained is delivered to the SPARQL engine at box 4017 and via the access layer - box 4018 - to the data layer - box 4019.
[0068] Fig. Figure 3 illustrates a procedure that periodically fine-tunes a sequence-to-sequence language model by injecting domain-specific context derived from predefined ontologies (Box 4001) and vocabularies (Box 4002). This enhancement enables the generation of SPARQL queries tailored to the enterprise context. Furthermore, context from available federated services (Box 4003), mapped to enterprise data silos, is integrated into the fine-tuning process to generate increasingly precise federated SPARQL queries. Box 4003 allows for the periodic retrieval of extended semantic metadata from Boxes 4001, 4002, and 4003, and the automatic triggering of the language model's fine-tuning process.The methodology also includes the generation of context-sensitive vector embeddings (Box 4005) from semantic metadata in Boxes 4001, 4002, and 4003 to improve the understanding of the domain-specific context for the language model. These vector embeddings are generated using graph-walking strategies and word embeddings for the accurate representation of semantic information. Further instruction-based (Box 4007) supervised fine-tuning (Boxes 4008 and 4009) is performed on an open-source sequence-to-sequence language model (Box 4010) using a precise natural language question and the corresponding federated SPARQL query. As a result, a high-performance task-specific language model (Box 4012) can be created for query generation. The effectiveness of the finely tuned language model in translating a question in natural language into a corresponding federal SPARQL query is demonstrated, for example, by...The Bilingual Evaluation Understudy (BLEU) metric (Box 4013) is used to evaluate the models using a validation dataset. The best-performing model is then deployed to the ontology-based fine-tuning software system using an inference API (Box 4014).
[0069] Fig. Figure 4 schematically illustrates a processing device 5005 according to various examples.
[0070] The processing device 5005 is designed to implement techniques disclosed herein in conjunction with a federated query. For example, the processing device 5005 may be designed to generate such a federated query by inferring a language model and / or by refining the language model. The processing device 5005 may implement one or more of the blocks of Method 1000 of Fig. 1. Implement and / or one or more parts of the data processing pipeline of Fig. 2. Implement and / or one or more parts of the workflow of Fig. 3. Implement.
[0071] The processing device 5005 includes a processor 5010, such as a central processing unit and / or one or more graphics processing units and / or traction processing units. The processing device 5005 also includes a memory 5015, which stores program code. The program code can be loaded and executed by the processor 5010. Furthermore, the processing device 5005 includes a communication interface 5020, through which data can be received from and / or transmitted to other computing devices. For example, a cloud server database or data repository can be accessed to obtain metadata, such as semantic metadata, which is used to train a language model. The processing device 5005 also includes a human-machine interface 5025. The human-machine interface 5025 can interact with the user, for example, by...by providing a graphical user interface through which text input can be received and text output can be provided to the user, as previously described in conjunction with . Fig. 2 and a graphical user interface 3105 explained. The processor 5010 performs techniques disclosed herein when loading and executing program code stored in the memory 5015, such as: pre-training a language model; fine-tuning a pre-trained language model; inferring a language model; obtaining metadata, such as a graph data structure representing an ontology and / or text data representing a vocabulary; outputting and / or providing a fine-tuned language model after fine-tuning; etc.
[0072] In summary, techniques were revealed that enable the automation of generating federated SPARQL queries based on organization-specific user queries in natural language. Context-sensitive vector embeddings from semantic data, coupled with instruction-based supervised fine-tuning, utilize labeled federated SPARQL queries for large language models. This allows for the generation of domain-specific federated SPARQL queries from natural language for efficient information retrieval from multiple organizational data silos.
[0073] Although the disclosure has been shown and described with respect to certain preferred embodiments, equivalents and modifications will come to mind for other persons skilled in the art when reading and understanding the description. The present disclosure includes all such equivalents and modifications and is limited only by the scope of the appended claims. [1] Rony, Md Rashad Al Hasan, et al. „Sgpt: A generative approach for SPARQL query generation from natural language questions.“ IEEE Access 10 (2022): 70712-70723. [2] Luz FF, Finger M. Semantic parsing natural language into SPARQL: improving target language representation with neural attention. arXiv Vordruck arXiv:1803.04329. 12. März 2018. [3] Wang R, Zhang Z, Rossetto L, Ruosch F, Bernstein A. Nlqxform: A language model-based question to sparql transformer. arXiv Vordruck arXiv:2311.07588. 8. Nov. 2023. [4] Panchbhai A, Soru T, Marx E. Exploring Sequence-to-Sequence Models for SPARQL Pattern Composition. arXiv Vordruck arXiv:2010.10900. 21. Okt. 2020. [5] Nikolov, Andriy, et al. „Ephedra: Efficiently combining RDF data and services using SPARQL federation.“ Knowledge Engineering and Semantic Web: 8. Internationale Konferenz, KESW 2017, Szczecin, Polen, 8.-10. November 2017, Protokoll 8. Springer International Publishing, 2017. [6] Heling L, Acosta M. Federated SPARQL query processing over heterogeneous linked data fragments. In Protokoll der ACM Web-Konferenz 2022 25. Apr. 2022 (S. 1047-1057). [7] Sima, Ana Claudia, et al. „Enabling semantic queries across federated bioinformatics databases.“ Datenbank 2019 (2019): baz106. [8] Sander M, Waltinger U, Roshchin M, Runkler T. Ontology-based translation of natural language queries to SPARQL. In 2014 AAAI Fall Symposium Series 24. Sept. 2014. [9] Javaheripi, Mojan, et al. „Phi-2: The surprising power of small language models.“ Microsoft Research Blog (2023).
[10] Rakhmawati NA, Umbrich J, Karnstedt M, Hasnain A, Hausenblas M. Querying over Federated SPARQL Endpoints---A State of the Art Survey. arXiv Vordruck arXiv:1306.1723. 7. Juni 2013.
[11] Mitra A, Del Corro L, Mahajan S, Codas A, Simoes C, Agarwal S, Chen X, Razdaibiedina A, Jones E, Aggarwal K, Palangi H. Orca 2: Teaching small language models how to reason. arXiv preprint arXiv:2311.11045. Nov 18, 2023.
[0074] Regardless of the grammatical use of the term, persons with male, female or other gender identities are included in the term. QUOTES INCLUDED IN THE DESCRIPTION
[0000] This list of documents cited by the applicant was automatically generated and is included solely for the reader's convenience. The list is not part of the German patent or utility model application. The DPMA accepts no liability for any errors or omissions. Cited patent literature
[0000] US 2010 / 0185643 A1
[0014] Cited non-patent literature
[0000] Rony, Md Rashad Al Hasan et al. „Sgpt: A generative approach for SPARQL query generation from natural language questions.“ IEEE Access 10 (2022): 70712-70723
[0005] Nicht-Patent-Literatur - Luz FF, Finger M. Semantic parsing natural language into SPARQL: improving target language representation with neural attention. arXiv Vordruck arXiv:1803.04329. 12. März 2018
[0006] Wang R, Zhang Z, Rossetto L, Ruosch F, Bernstein A. Nlqxform: A language model-based question to sparql transformer. arXiv Vordruck arXiv:2311.07588. 8. Nov. 2023 [0007, 0073] Panchbhai A, Soru T, Marx E. Exploring Sequence-to-Sequence Models for SPARQL Pattern Composition. arXiv Vordruck arXiv:2010.10900. 21. Okt. 2020
[0008] Nicht-Patent-Literatur - Nikolov, Andriy, et al. „Ephedra: Efficiently combining RDF data and services using SPARQL federation.“ Knowledge Engineering and Semantic Web: 8. Internationale Konferenz, KESW 2017, Szczecin, Polen, 8.-10. November 2017, Protokoll 8. Springer International Publishing, 2017
[0008] Heling L, Acosta M. Federated SPARQL query processing over heterogeneous linked data fragments. In Protokoll der ACM Web-Konferenz 2022 25. Apr. 2022 (S. 1047-1057 [0009, 0073] Sima, Ana Claudia, et al. „Enabling semantic queries across federated bioinformatics databases.“ Datenbank 2019 (2019): baz106 [0011, 0073] Sander M, Waltinger U, Roshchin M, Runkler T. Ontology-based translation of natural language queries to SPARQL. In2014 AAAI Fall Symposium Series 24. Sept. 2014 [0013, 0073] Rakhmawati NA, Umbrich J, Karnstedt M, Hasnain A, Hausenblas M. Querying over Federated SPARQL Endpoints---A State of the Art Survey. arXiv Vordruck arXiv:1306.1723. 7. Juni 2013
[0030] Mojan, et al. „Phi-2: The surprising power of small language models.“ Microsoft Research Blog (2023
[0033] Mitra A, Del Corro L, Mahajan S, Codas A, Simoes C, Agarwal S, Chen X, Razdaibiedina A, Jones E, Aggarwal K, Palangi H. Orca 2: Teaching small language models how to reason. arXiv Vordruck arXiv:2311.11045. 18. Nov. 2023.
[0033] Rony, Md Rashad Al Hasan, et al. „Sgpt: A generative approach for SPARQL query generation from natural language questions.“ IEEE Access 10 (2022): 70712-70723
[0073] Luz FF, Finger M. Semantic parsing natural language into SPARQL: improving target language representation with neural attention. arXiv Vordruck arXiv:1803.04329. 12. März 2018
[0073] Panchbhai A, Soru T, Marx E. Exploring Sequence-to-Sequence Models for SPARQL Pattern Composition. arXiv Vordruck arXiv:2010.10900
[0073] Nikolov, Andriy, et al. „Ephedra: Efficiently combining RDF data and services using SPARQL federation.“ Knowledge Engineering and Semantic Web: 8. Internationale Konferenz, KESW 2017, Szczecin, Polen, 8.-10. November 2017, Protokoll 8. Springer International Publishing, 2017
[0073] Javaheripi, Mojan, et al. „Phi-2: The surprising power of small language models.“ Microsoft Research Blog (2023)
[0073] Rakhmawati NA, Umbrich J, Karnstedt M, Hasnain A, Hausenblas M. Querying over Federated SPARQL Endpoints---A State of the Art Survey. arXiv Vordruck arXiv:1306.1723. 7. Juni 2013.
[0073] Mitra A, Del Corro L, Mahajan S, Codas A, Simoes C, Agarwal S, Chen X, Razdaibiedina A, Jones E, Aggarwal K, Palangi H. Orca 2: Teaching small language models how to reason. arXiv Vordruck arXiv:2311.11045. 18. Nov. 2023
[0073]
Claims
[1] Computer-implemented method (1000) for fine-tuning a pre-trained language model (3210) to generate a federated query associated with a product from a request, comprising: □ Receive (1100) first semantic metadata associated with an ontology representing concepts of the product, and second semantic metadata associated with a vocabulary describing the concepts of the product; and □ Fine-tuning (1200) the pre-trained language model based on the first semantic metadata associated with the ontology and further based on the second semantic metadata associated with the vocabulary. [2] Computer-implemented method (1000) according to claim 1, further comprising: □ Obtain (1100) third semantic metadata associated with at least one configured federated service, wherein each of the at least one configured federated service is assigned a data silo that stores data associated with the product; wherein the fine-tuning of the pre-trained language model is further based on the third semantic metadata associated with the at least one configured federated service. [3] Computer-implemented method (1000) according to claim 1 or 2, wherein the fine-tuning of the pre-trained model is based on one or more vector embeddings of the semantic metadata. [4] Computer-implemented method (1000) according to any one of the preceding claims, further comprising: □ Determine (1300) at least one change associated with the ontology and / or vocabulary and / or a configured federal service; triggering the fine-tuning of the pre-trained language model based on the determination of the at least one change. [5] Computer-implemented method (1000) according to one of the preceding claims, wherein the fine-tuning of the pre-trained language model is triggered based on a predefined timing plan. [6] Computer-implemented method (1000) according to any one of the preceding claims, further comprising: □ Receiving one or more prompts and one or more corresponding federal ground truth queries; wherein the fine-tuning of the pre-trained language model is further based on the one or more prompts and the one or more corresponding federal ground truth reference queries. [7] Computer-implemented method (1000) according to any one of the preceding claims, wherein the product comprises a projection radiography scanner, a magnetic resonance tomography scanner, a computed tomography scanner, a positron emission tomography scanner, a single photon emission tomography scanner or an ultrasound scanner. [8] Computer-implemented method (1000) according to one of the preceding claims, wherein the federated query comprises a Protocol-and-Resource-Description-Framework-Query-Language query. [9] Computer-implemented method (1000) according to one of the preceding claims, wherein the federated query serves to access multiple data silos associated with different components of the product. [10] Computer-implemented method (1000) according to any one of the preceding claims, further comprising: □ after completion of the fine-tuning: Validate and / or evaluate (1600) the fine-tuned language model. [11] Computer-implemented method, comprising: □ Receiving a request describing desired data associated with a product; and □ Generating, based on the request, a federated query associated with the desired data, using a pre-trained language model fine-tuned by one of the preceding claims. [12] Computer-implemented method according to claim 11, further comprising: □ Retrieve the desired data from multiple data silos that store product-associated data, based on the generated federated query. [13] Processing device (5005) comprising a processor (5010) and a memory (5015), wherein the processing device (5005) is designed to: □ Receive (1100) first semantic metadata associated with an ontology representing concepts of a product, and second semantic metadata associated with a vocabulary describing the concepts of the product; and □ Fine-tuning (1200) the pre-trained language model based on the first semantic metadata associated with the ontology and further based on the second semantic metadata associated with the vocabulary. [14] Processing device (5005) according to claim 13, wherein the processing device (5005) is designed to carry out the method according to any one of claims 1 to 12. [15] Computer program product or computer program or computer-readable storage medium containing program code, wherein the program code is executed by at least one processor to cause the at least one processor to perform the method according to any one of claims 1 to 12.
Citation Information
Patent Citations
CN000117555985A