Device for responding to audio input and method executed by processing system
By using domain-specific member models and arbitrators in a large language model alliance, directing queries to the appropriate member models, solving the difficulties in using large language models on systems with limited computing resources and misunderstandings in cross-domain applications, achieving efficient and accurate response and reducing computational burden.
Patent Information
- Application Number
- CN202411904191.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-12-22
- Filing Date
- 2024-12-23
- Publication Date
- 2025-06-24
AI Technical Summary
Large language models are difficult to use on systems with limited computing resources, especially in mobile environments such as cars. Due to the comprehensiveness and cross-domain application of the model, it is easy to misunderstand the context of the prompt, resulting in delayed and inaccurate responses.
A consortium (member model) that adopts a domain-specific large language model to respond, direct queries to the appropriate task-specific member model through an arbitrator, reducing computational burden and improving response accuracy.
By constraining the response to a specific domain, the possibility of misunderstanding the context of the prompt is reduced, the accuracy and speed of the response is improved, and the computational burden is reduced, allowing for faster access to required information within the vehicle.
Smart Images

Figure CN120196882A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a device responsive to audio input and a method performed by a processing system. Background Art
[0002] Large language models provide a way for machines to simulate human behavior. Using such models, machines attempt to predict how humans will respond to prompts. To build such models, they must be trained. This involves different patterns of making machine learning consistent with how humans respond to specific prompts.
[0003] The topics required to respond to a prompt are usually not known in advance. Therefore, it is useful to train large language models using information across multiple fields of human knowledge. This promotes the model's ability to deliver meaningful responses to prompts spanning the scope of human knowledge.
[0004] The breadth resulting from training models in so many different subject areas leads to large models, hence the name "large language models". However, this scale comes at a cost. As the model becomes "large", it becomes increasingly difficult to use the model on a system with limited computational resources. This is especially difficult when the computing system is integrated into a vehicle.
[0005] This problem can be easily overcome by providing such a system with a connection to a system with sufficient computational resources, such as a remote server. In this embodiment, the system relays the prompt to the remote server and waits while the remote server performs the work of generating a response. Then, the system receives the response from the remote server.
[0006] Typically, each solution only creates new problems. In this case, the problem is latency. Since the remote server may serve other users, it is impossible to predict or control how long it will take to receive a response.
[0007] In a moving vehicle, it is important to deliver information in a timely manner. For example, if you want to know whether a particular exit on a highway leads to the expected destination, it is important to receive an answer before passing that exit. Therefore, it is desirable to avoid having to interact with a remote server, as doing so introduces unpredictable waiting times.
[0008] Ironically, another problem lies in the comprehensiveness of large language models. It turns out that natural language often uses similar words and phrases in completely different contexts. Therefore, it is very likely to misunderstand the context of a prompt. Considering that large language models span multiple fields of knowledge, it is not surprising that they occasionally deliver responses that are completely different from the intent of the prompt. Summary of the Invention
[0009] The present invention avoids relying on using a single large language model to respond to a variety of questions and instead supports using domain-specific large language models that are members of a "league" of large language models. For convenience, these domain-specific large language models will be referred to herein as "member models".
[0010] In one aspect, the present invention contemplates receiving a query and then performing an arbitration process that directs the query to an appropriate task-specific member model that has been specifically trained for the domain pointed to by the query.
[0011] Since the members of the league (i.e., task-specific member models or domain-specific members) are often executed on different hardware elements at different locations, the arbiter solves the technical problem of routing queries to those hardware elements that are considered most suitable for solving those queries. For example, in some cases, the first member of the league is executed on a remote server and the second member of the league is executed locally. In such a case, the arbiter operates as a router that routes the query to one or the other based on the query and the attributes of the league members.
[0012] In another aspect, the present invention contemplates taking a subset of a large language model, fine-tuning, and / or extracting that subset to include task-specific information. Then, the smaller large language model that has been appropriately processed as described above is incorporated into the league environment, i.e., as members of a league of large language models that are independent of each other but enter into limited collaboration through the use of an arbiter.
[0013] As an example, consider an occupant in a vehicle near dinner time. As a person unfamiliar with the local environment, the occupant asks the vehicle's car assistant to find a nearby Albanian restaurant. In response, the car assistant invokes the arbiter to direct the query to the specific cuisine member of the league, which quickly provides the address of a nearby Albanian restaurant and an image of the menu.
[0014] Since the occupant is not fluent in the Albanian language, it is difficult to understand an important part of the menu. Therefore, the occupant asks the car assistant to translate it. In response, the arbiter directs the query to a different member of the league, i.e., the member dedicated to the Albanian language.
[0015] The technical advantages of the present invention stem from the ability to constrain the response to a query to a specific domain. This avoids the possibility of the query generating answers related to a domain that shares some but incorrect vocabulary used in the target domain. Another technical advantage of the present invention is the reduction of the computational burden. In addition, using a league of member models instead of a single model allows some of the members of the league to be stored locally for faster access from within the vehicle.
[0016] In one aspect, the present invention features an automotive assistant that executes on an infotainment system of a vehicle. The automotive assistant includes an arbiter configured to receive audio input provided by an occupant and output a member selection signal at least in part based on the audio input, the member selection signal selecting a specific domain member from a coalition of specific domain members. The automotive assistant is further configured to receive content from the selected specific domain member for provision and provide an audio output in response to the audio input provided by the occupant. The audio output is at least in part based on the content of the selected specific domain member from the coalition of specific domain members.
[0017] Embodiments include those that include a first specific domain member of the coalition. The first specific domain member is a member embedded in the vehicle.
[0018] In other embodiments, the specific domain members of the coalition include members that incorporate large language models.
[0019] Also in embodiments are those in which the specific domain members of the coalition include a first specific domain member and a second specific domain member. In such embodiments, the first specific domain member is embedded in the vehicle and the second specific domain member is at a remote server.
[0020] Embodiments further include those in which the arbiter includes a large language model, those in which the arbiter includes a graph neural network, and those in which the arbiter includes a neural network.
[0021] Some embodiments include a query splitter as part of the arbiter. This is useful for processing a composite query that includes two or more atomic queries. In such embodiments, the query splitter splits the composite query into a first atomic query and a second atomic query. However, the arbiter selects a first specific domain member of the coalition to provide content in response to the first atomic query and a second specific domain member of the coalition to provide content in response to the second atomic query.
[0022] Still other embodiments include those in which the automotive assistant further includes a multiplexer configured to receive the member selection signal. The multiplexer communicates data with each of the specific domain members of the coalition. The multiplexer provides a prompt generated by the arbiter to the selected specific domain member of the coalition.
[0023] Among these are embodiments in which the automotive assistant further includes two multiplexers, both of which communicate data with each specific domain member. A first multiplexer of the two multiplexers provides a prompt generated by the arbiter to the selected specific domain member of the coalition. A second multiplexer of the two multiplexers receives content from the selected specific domain member of the coalition and provides the content to the automotive assistant.
[0024] There are other embodiments that include a response generator and a text-to-speech converter. In these embodiments, the response generator receives content from selected domain-specific members of the coalition and generates a response based on the content, and the text-to-speech converter receives the response, which will ultimately be used for audio output.
[0025] Embodiments further include those that include a vehicle and / or an infotainment system as part of the present invention.
[0026] There are other embodiments that include those in which the arbiter is an arbiter that has been trained in conjunction with domain-specific members. Among these are: embodiments in which the arbiter has been jointly trained with domain-specific members from a coalition of domain-specific members, and embodiments in which the arbiter has been individually trained with domain-specific members from a coalition of domain-specific members.
[0027] In the above embodiments, the arbiter is trained according to a first training dataset, and the members are trained according to a second training dataset, which is different from the first training dataset.
[0028] In another aspect, the present invention features a method performed by a processing system that includes an automotive assistant executing on an infotainment system of a vehicle. The method includes: receiving an audio input from an occupant in the vehicle; selecting a domain-specific member from a coalition of domain-specific members at least in part based on the audio input; receiving content from the selected domain-specific member; and using the content to provide an audio output in response to the audio input.
[0029] In yet another aspect, the present invention features a digital automotive assistant executed in a processing system. This digital assistant includes an arbiter configured to receive an audio input provided by a human and output a member selection signal at least in part based on the audio input, the member selection signal selecting a domain-specific member from a coalition of domain-specific members. The digital assistant is further configured to receive content from the selected domain-specific member for providing and providing an audio output in response to the audio input provided by the human. The audio output is at least in part based on the content of the selected domain-specific member from the coalition of domain-specific members.
[0030] Related Applications
[0031] This application claims the benefit of the priority date of U.S. Provisional Application 63 / 613,855, filed on December 22, 2023, the content of which is incorporated herein by reference. Brief Description of the Drawings
[0032] Figure 1Shows an example of a specific architecture of an automotive assistant for implementing an arbiter for a consortium of members in a specific domain;
[0033] Figure 2 Shows details of consortium members communicating with an external application from Figure 1 ;
[0034] Figure 3 Shows Figure 2 Details of the repository and memory shown in; and
[0035] Figure 4 Shows Figure 2 Details of the repository and memory bank shown in. DETAILED DESCRIPTION
[0036] Figure 1 Shows a vehicle 10 having an infotainment system 12 that executes an automotive assistant 14. The automotive assistant 14 receives an audio input 16 from an occupant 18 of the vehicle 10. It does so via a microphone 20. The automotive assistant 14 then provides an audio output 22 back to the occupant 18 via a speaker 24 that emits speech provided by a text-to-speech (TTS) converter 26.
[0037] The audio output 22 includes certain requested information. One way to generate this information is to provide a prompt 28 to a consortium 30. The consortium 30 includes a plurality of members, where Figure 1 Shows a first member 32, a second member 34, and a third member 36. Each of the consortium members 32, 34, 36 includes a large language model. As implied by the figure, the first consortium member 32 is remote from the vehicle 10, while the second consortium member 32 and the third consortium member 34 are embedded in the vehicle 10. In some cases, an external source (such as the manufacturer of the vehicle 10) provides one or more of the consortium members 32, 34, 36.
[0038] The consortium members 32, 34, 36 jointly define a body of information, based on the information in the audio output 22, on the body of information. The body of information can be divided into separate "domains". While the domains are different from each other, it is not impossible for certain information to belong to more than one domain. Thus, the domains define a plurality of sets that are not necessarily disjoint.
[0039] Each of consortium members 32, 34, and 36 can obtain information from the corresponding domain among these domains. This makes consortium members 32, 34, and 36 "specific domains". Consortium members 32, 34, and 36 do not need to be of the same size. Each member 32, 34, and 36 includes a large language model for automobiles in a specific domain, and the large language model for automobiles in this specific domain has been fine-tuned based on one or more niche datasets to respond in a way that participates in specific domain interactions.
[0040] In some embodiments, one or more of consortium members 32, 34, and 36 are configured to process tasks by interacting with an external application 38 (e.g., by providing appropriate function calls for the external application using the application programming interface 40 of the application).
[0041] In the illustrated embodiment, the first member 32 provides information from the first domain, the second member 34 provides information from the second domain, and the third member 36 provides information from the third domain. Examples of domains include the traffic information domain, the restaurant information domain, and the astrophysical information domain.
[0042] Automobile assistant 14 includes an arbitrator 42, which receives audio input 16 and determines which of consortium members 32, 34, and 36 is most likely to provide satisfactory content for audio output 22. After doing so, the arbitrator 42 provides a member selection signal 44 to the first multiplexer 46 and the second multiplexer 48, and both the first multiplexer 46 and the second multiplexer 48 communicate with each member 32, 34, and 36 of the consortium 30 for data. In response to the member selection signal 44, the first multiplexer 46 directs the prompt 28 to the selected member among consortium members 32 and also directs the second multiplexer 48 to receive content 50 from this member 32.
[0043] Since it has been provided with the member selection signal 44, the second multiplexer 48 ultimately provides the content 50 to the response generator 52 or directly to the text-to-speech converter 26. The response generator 52 is used when the content 50 needs to be further transformed to be consistent with the occupant's expectations.
[0044] The arbitrator determines which of consortium members 32, 34, and 36 is most likely to respond satisfactorily to the user's audio input 16. In some embodiments, the arbitrator 42 includes a large language model, and the large language model has been trained to select an appropriate member 32, 34, and 36 based on the clues found in the audio input 16.
[0045] In some embodiments, the arbiter 42 receives requests that require different members of the coalition members 32, 34, 36 to perform multiple tasks. In such cases, the arbiter 42 parses the requests into individual tasks and routes each task to the appropriate member 32, 34, 36. It does this by using the query splitter 54.
[0046] The query splitter 54 receives a composite query and breaks it down into multiple atomic queries. For cases where there are multiple atomic queries, the arbiter 42 has the ability to select more than one member of the coalition members 32, 34, 36. This allows different members of the coalition members 32, 34, 36 to provide content in response to different atomic queries.
[0047] In some cases, the nature of the tasks is such that there is a natural order in which they should be performed. In such cases, the arbiter 42 routes the tasks to different members of the coalition members 32, 34, 36 in an order consistent with this natural order.
[0048] A particularly useful architecture for such large language models is a graph neural network-based architecture. Other embodiments of the arbiter 42 that include neural networks include embodiments implemented as large language models or any deep neural network. The arbiter 42 can also be implemented in a way that combines an encoder and a decoder.
[0049] In some embodiments, the arbiter 42 is equivalent to an M-way classifier that selects one of the M coalition members 32, 34, 36 as the appropriate provider of the content 50 for a given audio input 16. However, in some cases, the audio input 16 is complex enough that content 50 from two or more of the members 32, 34, 36 may be required to correctly formulate the audio output 22. This often occurs when the audio input 16 includes a composite query that contains multiple atomic queries.
[0050] In the case where the arbiter 42 and the coalition members 32, 34, 36 include large language models, the arbiter 42 and each member 32, 34, 36 are trained jointly or separately.
[0051] The training process for training the arbiter 42 jointly with the members 32, 34, 36 is an iterative process in which, during each step of the iteration, the weights of the arbiter 42 and the weights of the members 32, 34, 36 are adjusted together. This is a computationally intensive process.
[0052] The process of training the arbiter 42 using the members 32, 34, 36 individually is also an iterative process. However, in this case, each step of the iteration has two different phases. In one phase, when the weights of the members 32, 34, 36 are adjusted, the weights of the arbiter 42 are frozen. In the other phase, the weights of the members 32, 34, 36 are frozen while the weights of the arbiter 42 are adjusted.
[0053] It has been found that training the arbiter 42 using the members 32, 34, 36 individually reduces the computational amount of training, but the accuracy is only slightly reduced. However, it has been found that this slight reduction in accuracy is large enough to precisely distinguish products that have been manufactured by jointly training the arbiter 42 and 32, 34, 36, products that have been manufactured by individually training the arbiter 42 and 32, 34, 36, and products that have been manufactured without either jointly or individually training the arbiter 42 and the members 32, 34, 36. Therefore, the manufacturing process steps impart unique structural characteristics to the final product (i.e., the arbiter 42). In other words, products made by one process (i.e., joint training) will be functionally and structurally different from another process (i.e., individual training, i.e., using a two-stage training method).
[0054] In the above two-stage training method, each stage includes freezing a set of weights while allowing another set of weights to vary.
[0055] In one phase, the weights of the arbiter are varying; those weights of the members 32, 34, 36 are frozen. In this phase, the arbiter 42 learns to use the mixture-of-experts method to assign specific cues to specific members among the coalition members 32, 34, 36. In the mixture-of-experts method, a given cue is provided to each member 32, 34, 36, which in turn results in a corresponding set of responses. The nature of these responses provides the basis for forming the weighted vector of the arbiter 42. This enables the arbiter 42 to gradually learn those features that define the affinity of a given cue to a specific member among the coalition members 32, 34, 36. The result of this first training phase is that the arbiter 42 has learned how to map a given cue to the correct members 32, 34, 26 with high probability.
[0056] In some embodiments, some of the coalition members 32, 34, 36 have been trained to handle ambiguous cues, i.e., cues that represent two or more user intents. Such coalition members 32, 34, 36 are assigned labels before training the arbiter 42. Thus, when receiving an ambiguous cue that represents more than one user intent, the arbiter 42 only provides the ambiguous cue to those coalition members 32, 34, 36 that have been so labeled.
[0057] In another phase, the weights of the coalition members are not frozen, while the weights of the arbiter 42 are frozen. Then, for each prompt in a set of training prompts, each member 32, 34, 36 responds with an output and assigns a "prompt score" to the prompt, which indicates the extent to which the member 32, 34, 36 can respond to the prompt. As a result, by the end of this phase, the coalition members 32, 34, 36 will have been fine-tuned to receive a prompt, respond to it, and score that response. In combination with the first phase, this causes the arbiter 42 to have learned to process the responses of the coalition members, and each of the coalition members 32, 34, 36 has learned how to interpret the requests of the arbiter.
[0058] In operation, after being trained, depending on the nature of the prompt, the arbiter 42 sends to either member 32 the prompt that it deems to be the best choice. If that member 32 can return a response, it returns a response. If it cannot return one, then it returns a prompt score, which indicates that it cannot meaningfully respond. In the latter case, the arbiter 42 then provides to member 34 the prompt that it deems to be the second-best choice, at which point the foregoing process is repeated.
[0059] Since the number of coalition members 32, 34, 36 is limited, there is a risk of not being able to respond to a user request. To avoid this, it is useful to designate one member 36 as the member of last resort. This member 36 will provide a response even if all the other members of the coalition members 32, 34 are unable to do so.
[0060] Figure 2 An example is shown where the arbiter 42 has determined that the occupant 18 is seeking information in a particular domain (i.e., the "weather domain"). Accordingly, the arbiter 42 causes the prompt 28 to be sent to member 32, which is configured to respond to queries in the weather domain.
[0061] Weather information is a type of current information that typically requires consulting an external application 38. This means communicating with the external application 38 through its API (Application Programming Interface) 40. In the illustrated embodiment, member 32 includes a large language-model (LLM) 56 that communicates with a particular agent 58. The particular agent 58 is an agent that knows the different functions and arguments to provide to the external application 38.
[0062] The large language model 56 transmits a call query 60 to the particular agent 58. The call query 60 includes information about the nature of the information sought by the prompt 28.
[0063] In response to invoking query 60, specific agent 58 provides function-call precursor 62 back to large language model 56. Function-call precursor 62 includes information about relevant function calls and arguments to application programming interface 40 that must be provided to an external application.
[0064] Using function-call precursor 62, large language model 56 generates function call 64 and provides it to application programming interface 40. This causes external application 38 to generate content 50 that contains the sought-after information.
[0065] In some cases, this content 50, although containing relevant information, is generally not in a form suitable for delivery as audio output 22 to occupant 18. In such cases, as Figure 2 shown, member 32 provides the content to responder 52.
[0066] In a specific example, a prompt 28 in the form of "What is the weather like at Pemberley?" generates a function call query 60 in the form of "‘prompt’, ‘What is the weather like at Pemberley?’". This results in a function-call precursor 62 in the form of "‘function name’, ‘get weather’, ‘arguments’, ‘Pemberley’", which in turn provides the basis for generating a function call 64 in the form of "‘get weather(Pemberley)’". External application 38 responds with content 50 in the form of "‘arguments’, ‘Pemberley’, ‘-2°C’, ‘light snow’". Since this is not suitable for providing to occupant 18, responder 52 converts it to a more meaningful equivalent: "The weather at Pemberley is -2°C with light snow".
[0067] Figure 3 Additional features for facilitating the model's ability to provide reliable output are shown. A particular useful feature is repository 66, which provides information used by model 56 to enhance its ability to retrieve information most relevant to prompt 28. For those members among the consortium that rely on memory to dynamically infer context, it is also useful to provide a memory bank 68 for storing and updating context information.
[0068] Figure 4 Further details of repository 66 and memory bank 68 are shown.
[0069] Repository 66 is characterized by a domain-specific retrieval-enhancement module 70, a person-specific retrieval-enhancement module 72, and a function-specific retrieval-enhancement module 74.
[0070] The specific domain retrieval - enhancement module 70 causes the model 56 to limit its output to those outputs suitable for a specific domain. In the context of the example given in Figure 2 this reduces the likelihood that the model 56 will inappropriately interpret "snow" as referring to other white powdery substances that have limited interest within the meteorological domain of member 32.
[0071] The specific person retrieval - enhancement module 72 causes the model 56 to provide outputs related to a specific person, such as the occupant 18 of the vehicle 10.
[0072] The specific function retrieval - enhancement module 74 is particularly useful in cases where it is expected that the model 56 will make a function call 64 to the application programming interface 40.
[0073] The memory bank 68 is characterized by one or more types of memory that allow the model 56 to update its assessment of the context as it evolves, thus allowing the model 56 to essentially learn from experience. These types of memory are characterized by a flexible context size suitable for the member domain.
[0074] These memories in the memory bank 68 are episodic memory 76. The episodic memory 76 is particularly useful for identifying specific events closely related to the processed cue 28. The use of episodic memory 76 provides a basis for experience replay and also facilitates the task of providing information to the occupant 18 using two or more channels or modalities of information transfer (e.g., by using the speaker 24 and a display (not shown)).
[0075] Among these features is also the grounded memory 78. The grounded memory 78 provides information on how queries similar to the incoming query 28 are actually processed within the context of the specific domain specified by the member 32.
[0076] Among these features is the retrieval memory 80 that comes into play when the model 56 consults an external database. The retrieval memory 80 provides a way for the model 56 to base on the context of the external information received by the model 56 during the process of consulting an external information source. The retrieval memory 80 also provides many examples to facilitate context learning.
[0077] Among these features is also the context memory 82, which enables the model 56 to modify the context based on information from earlier cues, essentially allowing the model 56 to learn from experience.
[0078] The memory bank also includes the adaptive memory 84, which provides the model 56 with information useful for adapting its behavior based on previous interactions, thereby allowing the model 56 to provide outputs not only based on the context but also based on a time - varying context.
[0079] Now refer toFigure 4 , alternative embodiments of the vehicle assistant 14 include a text server 86 that receives audio input 16 from the occupant 18 via the microphone 20 and also provides audio output 22 to the occupant 18 via the speaker 24. Additionally, the text server 86 processes advanced context management, telemetry, and recording tasks. Communication between the text server 86 and the vehicle assistant 14 is through a guard rail 88 to block certain types of content.
[0080] Communication passes from the guard rail 88 to the orchestrator agent 90. The orchestrator agent 90 de-textualizes the audio input 16 and, if necessary, breaks down the audio input 16 into separate tasks.
[0081] The orchestrator agent 90 then proceeds to select one or more coalition agents 92, 94, 96, 98, 100. These include one or more of the following: a vehicle control agent 92 that processes functions for controlling the characteristics of the vehicle 10 itself; a communication agent 94 that processes communication between the vehicle 10 and various receiving entities external to the vehicle 10; a general data exchange agent 96 that interacts with a framework that facilitates the transfer of data between different systems and applications; a media agent 98 that executes commands for playing different types of media; and a general function call agent 100 that is configured to execute a wide range of functions not handled by other agents. These coalition agents 92, 94, 96, 98, 100 generate function calls 102, which are then provided to the function executor 104 for implementation and execution. In some cases, the coalition agents are external coalition agents 106 provided by an external source. Such agents 106 generate function calls and provide them to an external executor 108.
[0082] The orchestrator agent 90 then receives content 50 from the selected coalition agents 92, 94, 96, 98, 100 and provides it to the speech expression agent 110. In the illustrated embodiment, the speech expression agent 110 maintains a dialogue history and provides the ability to summarize the content 50. The speech expression agent 110 also performs further customization functions, including providing output in different languages.
[0083] Having described the invention and its preferred embodiments, what is claimed as new and protected by patent are the appended claims.
Claims
1. A device responsive to audio input provided by an occupant of a vehicle, the device comprising: A car assistant, the car assistant being executed on an infotainment system of the vehicle, the car assistant comprising: an arbitrator configured to receive the audio input provided by the occupant and select a domain-specific member from a coalition of domain-specific members to provide information related to responding to the audio input based at least in part on the audio input, Wherein, the automotive assistant is configured to receive content from selected domain-specific members and to provide audio output in response to the audio input provided by the occupant, the audio output being based at least in part on the content from the selected domain-specific members of the domain-specific member alliance.
2. The device according to claim 1, further comprising: A first domain-specific member of the consortium, wherein the first domain-specific member is embedded in the vehicle.
3. The device according to claim 1, wherein: The domain-specific members of the consortium include large language models.
4. The device according to claim 1, wherein: The domain-specific members of the consortium include a first domain-specific member and a second domain-specific member, wherein the first domain-specific member is embedded in the vehicle, and wherein the second domain-specific member is at a remote server.
5. The device according to claim 1, wherein: The arbitrator includes a large language model.
6. The device according to claim 1, wherein: The arbitrator includes a graph neural network.
7. The device according to claim 1, wherein: The arbitrator includes a neural network.
8. The device according to claim 1, wherein: The arbitrator includes a query divider, wherein the audio input includes a compound query, the compound query includes multiple atomic queries, wherein the query divider divides the compound query into a first atomic query and a second atomic query, wherein the arbitrator selects a first domain-specific member of the alliance to provide content in response to the first atomic query, and wherein the arbitrator selects a second domain-specific member of the alliance to provide content in response to the second atomic query.
9. The device according to claim 1, wherein: The automotive assistant further includes: a multiplexer configured to receive a member selection signal indicating the selected member of the alliance, wherein each of the specific domain members of the alliance is in data communication with the multiplexer, wherein the multiplexer provides the prompt generated by the arbitrator to the selected specific domain members of the alliance.
10. The device according to claim 1, wherein: The automotive assistant further includes: a first multiplexer configured to receive the member selection signal and a second multiplexer, wherein each of the specific domain members of the alliance communicates data with the first multiplexer, wherein each of the specific domain members of the alliance communicates data with the second multiplexer, wherein the first multiplexer provides prompts generated by the arbitrator to the selected specific domain members of the alliance, and wherein the second multiplexer receives content from the selected specific domain members of the alliance and provides the content to the automotive assistant.
11. The device according to claim 1, wherein: The automotive assistant further includes a response generator and a text-to-speech converter, wherein the response generator receives the content from the selected domain-specific members of the alliance and generates a response based on the content, and wherein the text-to-speech converter receives the response.
12. The apparatus of claim 1, further comprising the vehicle.
13. The apparatus of claim 1, further comprising the infotainment system.
14. The apparatus according to claim 1, wherein: The arbitrator is an arbitrator that has been trained jointly with domain-specific members from the domain-specific member coalition.
15. The apparatus according to claim 1, wherein: The arbitrator is an arbitrator that has been trained individually with domain-specific members from the domain-specific member federation.
16. A method performed by a processing system, the processing system comprising a car assistant executing on an infotainment system of a vehicle, the method comprising causing the car assistant to perform steps comprising: receiving audio input from an occupant of the vehicle, selecting a domain-specific member from a coalition of domain-specific members based at least in part on the audio input, Receive content from selected members in specific niches, and The content is used to provide an audio output in response to the audio input.