System and method using a large language model

A dual-language model system addresses inefficiencies in large language model systems by using a resource-reduced second model to determine if requests are sufficient, optimizing resource usage and reducing energy consumption.

WO2025114004A1PCT designated stage expired Publication Date: 2025-06-05SIEMENS AG
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2024/082204
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-27
Filing Date
2024-11-13
Publication Date
2025-06-05

AI Technical Summary

Technical Problem

Existing systems using large language models for processing natural language inputs face inefficiencies and high resource consumption, particularly when handling spam or irrelevant requests, leading to potential overload and user dissatisfaction.

Method used

Implementing a dual-language model system where a resource-reduced second language model generates a two-valued response, determining whether to forward the request to a more powerful first language model or provide a predefined warning, thereby optimizing resource usage.

Benefits of technology

This approach reduces resource consumption and energy usage by bypassing the first language model for insufficient requests, providing a quick and efficient warning while maintaining user utility without the need for extensive processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2024082204_05062025_PF_FP_ABST
    Figure EP2024082204_05062025_PF_FP_ABST
Patent Text Reader

Abstract

A second language model is added to a system for providing a result for a request formulated in natural language by means of a first large language model. The second language model is resource-reduced compared to the first large language model and is designed to generate a two-valued response from the request. For one of the two possible responses, a predetermined response is provided as a result of the request instead of a response from the first large language model.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Description

[0002] System and method with a large language model

[0003] The invention relates to a system and a method in which a result is generated from a requirement formulated in natural language using a large language model.

[0004] Large language models (LLMs) offer a range of possible applications. For example, they can be used in a complex industrial environment to select the right component for a given scenario from a range of available components using natural language input. For example, the components can be automation components.

[0005] For this purpose, a system is used that includes the following basic components: The user is provided with an input mask (graphical user interface, GUI) that contains a text input field and a text output field, as well as a confirmation or submit button (e.g., a "Submit" button). A backend that uses a large language model is invisible to the user.

[0006] A user describes their application in natural language in the text input field. After clicking the submit button, the content of the text input field is forwarded to the Large Language Model. The Large Language Model processes the input (infers from it) and provides its output in the text output field of the input mask.

[0007] The output text can generally fall into three categories. The first category indicates incorrect input. The second category includes output texts that indicate missing information in the input. The third category includes output texts that contain the desired information, for example, a selection from the available automation components that matches the input.

[0008] The input texts belonging to these categories each exhibit one of the following three properties. It may be input text that has no connection to an application description, for example, a spam request or a completely incorrect input. Furthermore, the input text may contain an overly abstract form of the application description. Finally, the input text may be designed in such a way that the application is sufficiently described.

[0009] Inferior processing, i.e., the evaluation of input text by the Large Language Model, requires considerable hardware resources on the backend and causes significant energy consumption. Depending on the load, users may have to wait a long time for a response. Spam requests also contribute to load and resource consumption.

[0010] Based on this, the object of the present invention is to create a system with a large language model for use in the described technical field, in which resource utilization and energy consumption are reduced. A further object is to specify an operating method for such a system.

[0011] This object is achieved by a system having the features specified in claim 1. A further solution consists in a method having the features of claim 8.

[0012] The inventive system for providing a result for a request formulated in natural language comprises a first large language model. The first large language model is designed to generate the result from the request. The system further comprises a second language model, which has fewer resources than the first, and is designed to generate a two-valued response from the request.

[0013] The system is further configured to forward the request to the first large language model in the case of a first of the two possible answers of the two-valued answer, and to provide a definable warning as a result of the request in the case of a second of the two possible answers of the two-valued answer.

[0014] In the inventive method for operating a system for providing a result for a request formulated in natural language, a first large language model and a second large language model, which has fewer resources than the first, are provided. The request is forwarded to the second language model, and the second language model generates a two-valued response from it.

[0015] In the case of the first of the two possible answers of the two-valued answer, the request is forwarded to the first large language model, which generates a result from the request. In the case of the second of the two possible answers of the two-valued answer, a definable answer is provided as the result of the request.

[0016] The invention recognized that, in particular, the response to requests that have no reference to an application description provides little information, even though the request was processed by a powerful large language model whose resource consumption is determined only by the number of parameters of the large language model and the length of the input texts, not by their content. This leads to inefficient use of resources and, in the case of multiple users of the system, such as in a client-server environment, to a possible overload of the backend and dissatisfied users.

[0017] Precisely for this type of requirement text, according to the invention, a result is provided by bypassing the first large language model by evaluating the requirement in the second language model. The second language model is resource-reduced, for example, by using a reduced number of neurons (layers), i.e., parameters, and is designed to provide only a two-value answer, for example, "sufficient" and "insufficient."

[0018] If the response is "not sufficient," then the request never reaches the first Large Language Model. Instead of a result that would be provided by the Large Language Model specifically for the request, a fixed response or warning is issued instead. This is just as helpful for the user but requires significantly fewer resources to create because it is predetermined.

[0019] Language models are artificial neural networks with a large number of parameters. They are designed to generate a natural language response from a natural language input by repeatedly predicting the most likely next word of the response. Large Language Models are those that have a very large number of parameters (> 10 9 ) .

[0020] Advantageous embodiments of the system according to the invention emerge from the dependent claims. The embodiment of the independent claims can be combined with the features of one of the subclaims or, preferably, with those of several subclaims. Accordingly, the following additional features can be provided:

[0021] The request can be forwarded to the second language model with a first probability different from 1 and to the first large language model with a second probability complementary to the first probability.

[0022] The second language model can have a number of parameters in the range of 10 8 ( 100 million) to 10 9 (1 billion). Preferably, a second language model with between 200 million and 400 million parameters is used. The ratio of the parameters of the first large language model to the second language model can be at least 5:1, in particular at least 10:1 or at least 20:1.

[0023] The second language model is appropriately adapted for the task of the two-valued answer in that the output layer of the already trained second language model allows only one classification, for example having two output neurons.

[0024] The system may include a device for determining a metric for the system's utilization. This is typically a software component. The metric is preferably incorporated into the probability. This can ensure that, during periods of low utilization, even potentially inappropriate requests are processed by the first large language model, thus providing responses that go beyond the specified warning.

[0025] The probability is preferably determined as the product of the utilization metric and a definable upper limit for the probability. The definable upper limit ensures that, during operation, a portion of the requests is always directed to the first large language model.

[0026] The system can have a user interface that can be output on a data processing device and includes an input option for the request and an output option for the result generated from the request. Data processing devices include, for example, computers, tablets, and smartphones. The input can be written (typed) or verbal (dictated).

[0027] The system can comprise a device for natural language processing (NLP), designed to generate a two-value classification for a result generated by the first large language model from the requirement, wherein the system is designed to feed the requirement as an input value and the classification as the correct output value to the second language model for training. The classification by the device expediently corresponds to the two-value answer to be generated by the second language model. The device can itself be a large language model. Preferably, however, it is a device that searches the result for definable pieces of text. Such pieces of text are, for example, “insufficient”, “not clear”, “cannot” and others that indicate an unusable requirement in the result.

[0028] The invention is described and explained in more detail below with reference to the exemplary embodiments shown in the figures. They show:

[0029] Figure 1 shows a system for selecting a suitable component from automation technology for an application description in natural language,

[0030] Figure 2 shows the system in a training configuration.

[0031] Figure 1 shows a system 10 for selecting a suitable component from automation technology for an application description in natural language. For example, the component can be a control system for a production line in a manufacturing plant. In another example, the component can be a suitable robot for an activity in a manufacturing plant. In each of these examples, there are a large number of such components that are, in principle, suitable for the application. The system 100 described below helps users make such a selection.

[0032] The system 10 includes a user interface 120, which is used for interaction with users and thus for entering the application description in natural language. The user interface is expediently displayed on the screen of a computing device. The computing device can be a tablet, smartphone, or a PC. The user interface 120 can be implemented as a program, i.e., run on the computing device, but can also be a web interface, i.e., simply displayed in the style of a website.

[0033] The input is conveniently made in writing in a provided input field 121. It is understood that the input can be made not only via keyboard, but also as dictation, which is converted into text.

[0034] The user interface 120 further includes a submit button 122. If the user is satisfied with the request text, they can submit the request text by pressing the button 122. This transmits the request text to a backend 100.

[0035] The backend 100 processes the request text and generates a result. The backend 100 is typically not implemented locally on the computing device, as this would require too many resources, but is connected to the computing device via the Internet. In this example, the request text is transmitted to the backend 100 via the Internet.

[0036] The backend 100 includes an initial large language model 105. Large language models are well-known and are based on very large neural networks with weights numbering in the billions. They are trained with extensive databases of language information, typically generated from the Internet and other written sources.

[0037] The backend 100 further comprises a second language model 110. The second language model 110 has reduced resource requirements compared to the first large language model 105, i.e., it is designed such that its operation places less strain on the hardware and / or results in lower energy consumption. This can be achieved, for example, by the second language model 110 using a smaller number of artificial neurons. For example, the second language model 110 can be equipped with only 80% or only 60% of the neurons of the first large language model 105.

[0038] Furthermore, the backend 100 comprises a software component that operates as a switch 115. During operation, an incoming request text is forwarded to the switch 115. The switch 115 is connected to a device 116 for determining the utilization, which is expediently also a software component. The device 116 determines the utilization and sysa definable part of the backend 100, for example the first large language model 105. The utilization u sys is usually output as a number between 0 and 1, where 0 stands for no utilization and 1 for full utilization. The value thus determined for a current utilization u sys is available as input value in switch 115.

[0039] The switch 115 determines from the value for the current load u sys and a predeterminable maximum probability r con a probability p = u sys • r con- The given request text is forwarded to the second language model 110 with probability p and to the first large language model 105 with the complementary probability 1 - p. For the specific request text, a decision regarding forwarding is then made based on the probability p. This is usually done by determining a random number between 0 and 1 and comparing it with the probability p.

[0040] Is the utilization u sys = 0, the request text is forwarded to the first Large Language Model 105 and processed. However, if the load is different from zero, a portion of the incoming request texts is forwarded to the second Language Model 110.

[0041] As already explained, the second language model 110 operates with reduced resources compared to the first large language model 105. It is also designed to generate a two-valued result, i.e. to output exactly one first or second answer option 111, 112 for each request text. The first answer option 111 is essentially “OK”. The second answer option 112 is essentially “not OK”. It is understood that the concrete output form for the two answer options is practically arbitrary, as long as it is used consistently. With an appropriately designed neural network, the output can consist of the value of a single output neuron. If the second language model 110 outputs a text, for example, a 1 can be output for “OK” and a 0 for “not OK”, or exactly the texts “OK” and “not OK”. The second language model 110 therefore works as a classifier for the request texts.Since the second Language Model 110 , in contrast to the first Large Language Model 105 , only has to generate a two-valued answer and not a factually correct answer text in natural language , it can be implemented in a resource-saving manner.

[0042] If the output of the second language model 110 corresponds to "OK", then the request text is forwarded to the first large language model 105 and processed. The processing then corresponds to that which occurs when the request text is forwarded directly from the switch 115 to the first large language model 105. The first large language model 105 prepares a response to the request text, which is displayed as text output 123 in the user interface 120.

[0043] If, however, the output of the second language model 110 corresponds to "not OK", the request text is not forwarded to the first large language model 105 and is also not processed. Instead, in this case, a defined warning 124 is displayed in the user interface 120, informing the user that the request text is not sufficient to generate a suitable response. The defined warning 124 is therefore not a response generated by a large language model, but is already defined and stored. It is therefore not adapted to the request text.

[0044] In the present example, the second language model 110 is trained as shown in Figure 2. The system 10 that is also shown in Figure 1 is used again. Corresponding elements are designated by the same reference symbols as in Figure 1. As can be seen, at least part of the training for the second language model 110 is carried out using requirement texts that are processed in the system 10. While, in principle, prefabricated requirement texts can also be used, Figure 2 shows training using real requirement texts, i.e. requirement texts created by users.

[0045] In switch 115, a value for the maximum probability r con= 0 is used. This results in a probability of p = 0 for the forwarding of request texts to the second language model 110; thus, all request texts are forwarded to the first large language model 105. The request texts are also forwarded to a software component for labeling 215.

[0046] The first large language model 105 prepares a response to the request text, which is displayed as text output 123 in the user interface 120. Furthermore, the response of the first large language model 105 is also passed on to a device 210. The device 210 is a device for natural language processing, which is designed to classify the response of the first large language model 105. The result of the classification corresponds to that carried out by the second language model 110, but here it occurs using the response of the first large language model 105 instead of the request text. The two response options 211, 212 of the device 210 therefore correspond to the response options "OK" and "not OK".

[0047] The device 210 can, for example, itself be a large language model. Alternatively, and much simpler, a device 210 searches the output of the first large language model 105 for specific keywords, such as "cannot" or "insufficient." This search and a corresponding keyword catalog require minimal resources compared to a large language model.

[0048] The answer option 211, 212 resulting from the request text from device 210 is passed on to the software component for labeling 215. There, the answer option 211, 212 is assigned to the request text. Pairs of request text and answer thus formed are transmitted to the second language model 110 as training data in a step 220. The respective request text serves as the input variable; the second language model 110 always processes the request texts as input variable. The answer option 211, 212 serves as the correct output variable because it corresponds to the correct result to be delivered by the second language model 110.

[0049] In this way, the second language model 110 can be trained both in advance and during operation with relatively little additional effort, thus reducing resource consumption.

[0050] Reference sign

[0051] 10 systems

[0052] 100 backend

[0053] 105 first Large Language Model

[0054] 110 second language model

[0055] 111 , 112 answer options

[0056] 115 Switch

[0057] 116 Facility for capacity determination

[0058] 120 User interface

[0059] 121 input field

[0060] 122 Submit button

[0061] 123 Text edition

[0062] 124 fixed warning

[0063] 210 Natural Language Processing Facility

[0064] 211 , 212 possible answers

[0065] 215 Software component for labeling

Claims

Patent claims 1. System (10) for providing a result (123, 124) for a requirement formulated in natural language, comprising - a first large language model (105) designed to Generation of the result (123) from the request, - a second language model (110) which is resource-reduced compared to the first and is designed to generate a two-valued response (111, 112) from the request, wherein the system (10) is designed, - in the case of a first of the two possible answers (111) of the two-valued answer (111, 112), to forward the request to the first Large Language Model (105), and - in the case of a second of the two possible answers (112) of the two-valued answer (111, 112), a definable answer (124) as a result of the request.

2. System (10) according to claim 1, designed such that the request is forwarded to the second language model (110) with a first probability different from 1 and is forwarded to the first large language model (105) with a second probability complementary to the first probability.

3. System (10) according to claim 2 with a device (116) for determining a characteristic value for the utilization of the system (10), wherein the characteristic value is included in the first probability.

4. System (10) according to claim 3, wherein the first probability is determined as the product of the characteristic value and a definable upper limit for the first probability.

5. System (10) according to one of the preceding claims with a user interface (120) which can be output on a data processing device and comprises an input option for the request and an output option for the result (123, 124) generated from the request.

6. System (10) according to one of the preceding claims, comprising a device (210) for processing natural language, designed to generate a two-value classification (211, 212) for a result (123) generated from the request by the first large language model (105), wherein the system (10) is designed to feed the request as an input value and the classification (211, 212) as the correct output value to the second language model (110) for training.

7. A method for operating a system (10) for providing a result (123, 124) for a request formulated in natural language, in which - a first large language model (105) and a second language model (110) with reduced resources compared to the first are provided, - the request is forwarded to the second language model (110), - the second language model (110) generates a two-valued response (111, 112) from the request, - the request is forwarded to the first large language model (105) in the case of a first of the two possible answers (111) of the two-valued answer (111, 112), wherein the first large language model (105) generates a result (123) from the request, - in the case of a second of the two possible answers (112) of the two-valued answer (111, 112), a definable answer (124) is provided as a result of the request.

8. The method according to claim 7, wherein the request is forwarded to the second language model (110) with a first probability different from 1 and is forwarded to the first large language model (105) with a second probability complementary to the first probability.

9. The method according to claim 7 or 8, wherein the first probability is determined using a characteristic value for the utilization of the system (10).

10. The method according to claim 9, wherein the first probability is determined as the product of the characteristic value and a definable upper limit for the first probability.

11. The method according to one of claims 7 to 10, wherein - a device (210) for processing natural language generates a two-value classification (211, 212) for a result (123) generated from the request by the first large language model (105), - the requirement as input value and the classification (211, 212) as correct output value to the second language Model (110) for training.

Citation Information

Patent Citations

  • Routing natural language commands to the appropriate applications

    US11152009B1