Artificial intelligence assistance for audio, video, and control systems
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- QSC LLC
- Filing Date
- 2024-11-05
- Publication Date
- 2026-08-04
Smart Images

Figure CN122514802A_ABST
Abstract
Description
[0001] priority
[0002] This application claims priority to U.S. Provisional Patent Application No. 63 / 596,646, filed November 7, 2023, and U.S. Patent Application No. 18 / 585,587, filed February 23, 2024, the entire contents of which are incorporated herein by reference. Technical Field
[0003] This invention relates generally to, but is not limited to, audio, video, and control (“AVC”) systems, and more specifically to methods and systems for using artificial intelligence to control AVC systems, including processing cores and peripheral devices. Attached Figure Description
[0004] Figure 1 This is a block diagram illustrating an overview of an AVC processing core according to certain illustrative embodiments of this disclosure.
[0005] Figure 2 This is a block diagram of an AVC operating system using a pre-configured command set, according to certain illustrative embodiments of this disclosure.
[0006] Figure 3 This is a flowchart of a general method for performing one or more actions on a peripheral device according to an illustrative embodiment of this disclosure.
[0007] Figure 4 This is a block diagram of the AVC operating system using a user-taught command set, based on certain illustrative embodiments of this disclosure.
[0008] Figure 5 This is a flowchart of a computer-implemented method for performing actions on peripheral devices according to certain illustrative embodiments of this disclosure. Detailed Implementation
[0009] The following describes illustrative embodiments and related methods of this disclosure, which can be used to perform actions on networked peripheral devices on an AVC system using artificial intelligence. For clarity, not all features of the actual implementation or methods are described in this specification. It should be recognized that in developing any such actual embodiment, many implementation-specific decisions must be made to achieve the developer's specific goals, such as compliance with system-related and business-related constraints, which will vary from implementation to implementation. Furthermore, it should be recognized that such development work can be complex and time-consuming, but is routine work for those skilled in the art who will benefit from this disclosure. Further aspects and advantages of the various embodiments and related methods of the invention will become apparent from the following description and drawings.
[0010] More specifically, the illustrative embodiments of this disclosure allow a user to issue verbal commands to perform actions on various AVC systems. Verbal commands can be implemented using, for example, a Large Language Model (“LLM”). LLM is an artificial intelligence deep learning algorithm that performs various natural language processing tasks. As described herein, an AVC system includes a core processor and peripheral devices such as speakers, microphones, cameras, bridging devices, network switches, etc. The operating system executed by the AVC system performs all audio, video, and control processing on a single processing core. Performing all audio, video, and control processing on a single device makes configuring the AVC system much easier, as any initial configuration or subsequent changes to the configuration are performed on a single processing core. Therefore, any audio, video, or control configuration changes (e.g., changing the gain level of an audio device) are performed on a single device (the core processor), rather than requiring audio configuration changes on one processing device and video or control configuration changes on another. Furthermore, any software or firmware upgrades on the AVC system can be performed on a single processing core. Therefore, by using the embodiments currently disclosed, a user can control an AVC system (including any of the peripheral devices) using verbal commands. However, in some embodiments, the currently disclosed embodiments are not limited to performing all audio, video, or control processing on a single processing core. In some embodiments, audio, video, or control processing can occur on any number of processing cores and in any combination.
[0011] In yet another embodiment, the LLM is equipped with a set of default verbal commands for the LLM to detect / recognize. In other embodiments, the LLM is trained to detect verbal commands by receiving user input via a web browser or verbally.
[0012] An AVC system is a system configured to manage and control audio features, video features, and control features. For example, the AVC system of this disclosure can be configured for use with networked microphones, cameras, amplifiers, controllers, etc. An AVC system may also include various associated features such as acoustic echo cancellation, multimedia player and streaming capabilities, user interface, scheduling, third-party control, Voice over IP (“VoIP”) and Session Initiation Protocol (“SIP”) capabilities, scripting platform capabilities, audio and video bridging, public address functionality, and other audio and / or video output capabilities. An example of an AVC system is included in the Q-SYS® technology from QSC LLC, the assignee of this disclosure.
[0013] In the common approach disclosed herein, the AVC operating system is implemented on an AVC processing core communicatively coupled to one or more peripheral devices. The AVC processing core is configured to manage and control the audio, video, and control characteristics of the peripheral devices. The AVC processing core has many other capabilities, including, for example, playing audio files, processing audio signals, video signals, or control signals, and influencing the processing of any of these signals, such as performing acoustic echo cancellation, and other processing that may affect the sound quality and camera quality of the peripheral devices.
[0014] Using an LLM module communicatively coupled to the AVC processing core, the system detects one or more verbal commands issued by the user. Subsequently, it executes one or more actions corresponding to the verbal commands on peripheral devices and / or the AVC processing core.
[0015] Figure 1 This is a block diagram illustrating an overview of an AVC processing core according to certain illustrative embodiments of this disclosure. The AVC processing core 100 includes various hardware components, modules, etc., including an AVC operating system (“OS”) 102. This OS manages and controls various audio, video, and control features of one or more peripheral devices 104 or other applications / platforms (not shown) that may run on the peripheral devices 104 or one or more computing devices. The peripheral device 104 can be any type of device, such as a camera, microphone, bridging device, network switch, speaker, television or other AV equipment, curtains, heating or air conditioning unit, etc. The application / platform may include, for example, a calendar platform, a teleconferencing platform, etc.
[0016] The AVC processing core 100 may include one or more input devices 106 that provide input to the CPU (processor) 108 to notify it of actions. These actions may be mediated by a hardware controller that interprets signals received from the input devices and transmits information to the CPU 108 using a communication protocol. Input devices 106 may include, for example, a mouse, keyboard, touchscreen, infrared sensor, touchpad, wearable input device, camera- or image-based input device, microphone, personal computer, smart device, or other user input device.
[0017] CPU 108 may be a single processing unit, multiple processing units within a single device, or multiple processing units distributed across multiple devices. CPU 108 may be coupled to other hardware devices, for example, via a bus (such as a PCI bus or SCSI bus). CPU 108 may communicate with a hardware controller for a device (such as for display 110). Display 110 may be used to display text and graphics. In some embodiments, display 110 provides visual feedback to a user, including graphics and text. In some embodiments, display 110 includes an input device as part of the display, such as when the input device is a touchscreen or equipped with an eye orientation monitoring system. In some embodiments, the display is separate from the input device. Examples of display devices are LCD screens, LED screens, projection, holographic, or augmented reality displays (such as head-up displays or head-mounted displays). Other I / O devices 112 may also be coupled to the processor, such as network cards, video cards, audio cards, USB, FireWire or other external devices, cameras, printers, speakers, CD-ROM drives, DVD drives, disk drives, or Blu-ray devices.
[0018] In some implementations, the AVC processing core 100 also includes a communication device capable of communicating with network nodes wirelessly or via a wired connection. The communication device can communicate with another device or server over a network using protocols such as TCP / IP, Q-LAN, or others. The AVC processing core 100 can utilize the communication device to distribute operations across multiple network devices.
[0019] CPU 108 can access memory 114, which may include one or more of various hardware devices for volatile and non-volatile storage, and may include both read-only memory and writable memory. For example, memory may include random access memory (RAM), various caches, CPU registers, read-only memory (ROM), and writable non-volatile memory such as flash memory, hard disk drive, floppy disk, CD, DVD, magnetic storage devices, tape drives, etc. Memory is not a propagating signal detached from the underlying hardware; therefore, memory is non-transitory. Memory 114 may include program memory 116 storing programs and software (such as AVC operating system 102 and other application programs 118). Memory 114 may also include data memory 120, which may include data to be operated by application programs, configuration data, settings, options, or preferences, etc., which may be provided to program memory 116 or any element of AVC processing core 100.
[0020] Some implementations can operate in many other computing system environments or configurations. Examples of computing systems, environments, and / or configurations that may be suitable for use with this technology include, but are not limited to, personal computers, AV I / O systems, networked AV peripherals, video conferencing consoles, server computers, handheld or laptop computers, cellular phones, wearable electronic devices, game consoles, tablets, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframes, distributed computing environments including any of the above systems or devices, etc.
[0021] Figure 2 This is a block diagram of an AVC OS according to certain illustrative embodiments of this disclosure. In this example, AVC OS200 is... Figure 1 AVC OS 102. Figure 2 In the example, AVC OS 200 includes an LLM module 204 that executes a pre-configured set of commands to perform actions on peripheral device 104, as will be discussed below. LLM module 204 can be any kind of large language model, including, for example, ChatGPT, Google® Bard, Meta® LLaMA, etc. In alternative embodiments discussed later in this disclosure, AVC OS includes an LLM module trained to interpret and automatically listen for spoken commands. The various LLM modules described herein can be accessed as a cloud service, as a “locally deployed” service within a local network, or hosted within the AVC processing core itself (and thus by…). Figure 2 (The dashed lines in the text indicate this). Note that in this example, the solid lines around AVS OS 200 indicate modules that are part of AVC OS 200.
[0022] exist Figure 2 In a further description, AVC OS 200 includes a text-to-speech (“TTS”) module 206 that converts text from LLM module 204 into an audio signal that is sent to AVC OS 200. AVC OS 200 can then process these signals, for example, by mixing them into an output transmitted to a conference room. For example, a peripheral speaker can output audio data, such as from a remote conference (e.g., a laptop computer playing sound from a remote participant). AVC OS 200 can then mix this audio data with the “speech” audio signal transmitted from TTS module 206.
[0023] The digital signal processor 208 (also referred to herein as an audio engine) accepts audio input from peripheral device 104 transmitted to AVC OS 200 in any supported format or media form. This format or media form may include, for example, network streaming, VoIP, common old-style telephone service (“POTS”), or other formats or media forms. In this example, the audio signal is provided as a spoken / audible command from a user via peripheral device 104 (e.g., a microphone). The audio input may be processed by the digital signal processor (“DSP”) 208 to perform typical input processing (filtering, gating, automatic gain control (“AGC”), echo cancellation, etc.). In some embodiments, the audio signal may be processed to reduce the amount of data to be sent to the transcription service (labeled as Transcription Application Programming Interface (“API”) module 210—which may or may not be part of AVC OS 200, as indicated by the dashed line). For example, to reduce the amount of data, audio processing techniques such as level detection, voice activity detection, sample rate reduction, compression, etc., may be performed.
[0024] As needed, audio signals are sent to API module 210 via local inter-process communication (IPC) or over a network. In some embodiments, DSP 208 may provide a local buffer (first-in, first-out) to handle slow access to transcription API 210 or network interruptions. In other embodiments, the local buffer may be recorded to memory and saved as conference minutes.
[0025] The output of transcription API module 210 is sent to LLM module 204 via transcription interface 212, which implements software abstractions for communicating with the transcription API. For example, transcription interface 212 sends audio data to transcription API 210, which then transcribes the audio data and sends it to artificial intelligence (“AI”) interface 214, which implements software abstractions for communicating with LLM module 204. For example, interface 214 is responsible for placing the data in a network packet, making a web call to LLM module 204 to open it, or placing the data in shared storage and sending it to LLM module 204. Depending on the location of the service, there may be a direct connection between transcription API 210 and LLM module 204 (e.g., in a cloud platform), or alternatively, the output of transcription API 210 may be returned to AVC processing 100 and then sent to LLM module 204. However, once a transcription is received, the LLM module 204 identifies the speech intended to invoke a function (command detection) using various techniques (such as system prompts, fine-tuning, and / or "function calls"), as will be understood by those skilled in the art who benefit from this disclosure. In each of these techniques, the LLM module 204 annotates its output with tagging labels (e.g., in XML, Markdown, JSON, etc.), which can be parsed by the response parser 216 and sent via the control interface 220 (a software abstraction for the communication channel with the RE 218) to the runtime engine ("RE") 218 (also known as the control engine for controlling peripherals, communicating with peripherals, and audio processing), and dispatched to a function provider outside the AVC OS 200. The tagging labels can be in any language compatible with and understandable to the response parser 216.
[0026] A set of commands related to the control of audio, video, or control systems invokes control commands (also known as remote control (“RC”)), which are then used to perform actions on peripheral device 104, such as adjusting appropriate settings or other operations. For example, the volume may be increased in response to a user mentioning that the volume is too low. Other settings include adjusting curtains, turning on or changing the temperature of the air conditioner, brightening the screen, generating a meeting summary, and other commands to peripheral device 104, such as: turning the monitor on or off, changing the video display input (from a laptop or the internet), controlling the camera (placing the camera in privacy mode or changing the pan-tilt-zoom coordinates), muting the audio (not transmitting audio to the remote location), turning off AI, selecting a different audio input, hanging up a phone call, initiating a phone call, starting transcription, booking a meeting room for another meeting, generating a summary (sending the summary to participants via email), updating the participants' calendars, loading a preset from a set of presets (placing the room in a user-preset “movie mode”—camera on, lights off, audio at a specific volume), raising or lowering blinds, adjusting HVAC system settings, changing configuration settings: changing the time zone or setting the clock, etc.
[0027] An illustrative example of an RC / system prompt used to control audio muting is the AVC system of this disclosure monitoring a conference room via a microphone peripheral. The AVC system can control various aspects of the room by returning a response to the RC in the following manner: When the AVC system uses LLM module 204 to detect that audio should be muted, LLM module 204 will respond with an explicitly specified set of text / commands, such as:<control command> system_mute=on< / control_command> This can be parsed by the relatively "simple" response parser 216 and forwarded to the control command handler. When the system detects that the audio should be unmute using the LLM module 204, the LLM module 204 can respond with the following set of text / commands:<control command> system_mute=off< / control_command> ".
[0028] In some illustrative embodiments, the LLM needs to execute certain commands. In these embodiments, the command handler may use an LLM (such as LLM module 204 or another LLM (not shown)) to assist in executing commands, such as generating a conference summary; the summary may be the viewpoints of an audible, real-time discussion. In the example of generating a summary, the command handler may send transcribed text to the LLM and instruct the LLM to generate a summary.
[0029] In some illustrative embodiments, the command processor may use an LLM to assist in command execution when the LLM module used for command detection (e.g., LLM module 204) does not have sufficient computational resources available for allocation beyond those required for command detection, or when the LLM is not sophisticated or "intelligent" enough to execute the command. For example, an LLM for command detection might need to perform more frequent, less computationally intensive tasks, as the LLM might be invoked at the end of each spoken sentence during a meeting. However, generating high-quality meeting summaries, or executing more demanding commands or tasks, might require a more advanced LLM with the necessary computational resources and bandwidth. The specific LLM module used to execute commands can vary depending on the power requirements of executing a given command, as will be understood by one of ordinary skill in the art who benefits from this disclosure.
[0030] In some alternative embodiments, for certain command detection and summarization requirements, it may be infeasible to run an LLM module on the AVC processing core that can execute every command, or both command detection and detected commands. For example, while running an LLM locally (e.g., on the processing core) to execute all identified commands is envisioned within the scope of this disclosure, running a local LLM module to handle both command detection and summarization may be very difficult given the current level of technology. In this case, the LLM module can run locally as a finely tuned small model performing simple command detection, and in parallel, the command handler responsible for executing certain commands (such as summarization, transcription, and other commands) can offload the task to a cloud service for more “intelligent” processing. Another alternative is to use one or more cloud-based LLM modules for command recognition and execution (e.g., summarization or other computationally intensive tasks).
[0031] In other embodiments, the summary can be transmitted (via control interface 220) to the control system for display in a text bar on the user control interface.
[0032] In other embodiments, the command set may instruct the LLM module 204 to ignore "small talk" and "casual conversation." This can be achieved by specifying the desired functionality of the LLM module 204 using, for example, system prompts or fine-tuning techniques. In one example, ignoring small talk and casual conversation could be done by providing the system prompt with the sentence "Please ignore small talk and casual conversation and return the string instead." <chatter / > To achieve this, the exact wording of the prompts may need to be adjusted to maximize the effectiveness of the selected model.
[0033] In other embodiments, the LLM module 204 can be instructed (via a set of commands) to return tag labels to interact with other web services, such as: appending comments to a Confluence article; appending comments or other modifications to a Jira or Confluence entry; interacting with calendars and room schedules, for example, to extend the current room's booking duration if a meeting times out; or sending a summary to meeting participants via email.
[0034] In other embodiments, the LLM module 204 may listen for factual errors in the session and respond as notifications (e.g., a red light on a touchscreen, text in a user control interface (UCI) text bar, etc.). In yet another embodiment, the LLM module 204 may cross-check the calendar to determine when attendees are available for another meeting, integrate with a calendar scheduling platform, and perform online bookings, such as reserving meeting rooms and / or arranging discussions via the meeting platform. In the embodiment described here, the LLM module 204 does not provide any of these features but instead identifies when such functionality is needed and dispatches a request to the scheduling platform via a "calendar functionality provider."
[0035] In other illustrative embodiments, the LLM module 204 may use a peripheral microphone to listen for direct requests and respond in a chatbot-like manner. For example, if the LLM module 204 detects a verbal command such as, “Hey Big Q, what is the distance from the Earth to the Moon?”, the result can be tagged so that it can be sent to a speech synthesis service, and the resulting audio can be sent back to the DSP / AE 208 for playback in the room.
[0036] In other illustrative embodiments, the LLM module 204 can be initiated to perform actions without a direct request. The LLM module 204 can determine that a question is being asked and should be answered, for example, by detecting changes in tone of voice or noticing that a question has been asked but the participant has not responded for a period of time.
[0037] As previously discussed, in some illustrative embodiments, the LLM module 204 can recognize direct commands from the user in the text, such as “Mute the audio.” This direct command might begin with a wake word, such as, “Hey LLM, could you please mute the audio?” In yet another alternative embodiment, the LLM module 204 analyzes the context of the text to distinguish commands from general conversation. For example, if a user enters a room and says, “It’s a bit cold here,” the LLM module 204 can recognize that the user wants to turn up the temperature or down the air conditioning. As another example, a user might complain about severe glare on a whiteboard; the LLM module 204 could close the blinds in the room. However, if a user complains that the temperature this summer is lower than usual, the LLM module 204 can determine that this is not a command for the LLM module 204 to execute. Furthermore, if the lyrics of a song playing in the background contain phrases like “It’s getting hotter here” or “Turn up the volume,” the LLM module 204 can determine that these are lyrics, not commands for the program to execute. The LLM module 204 can provide a variety of functions, as will be understood by one of ordinary skill in the art who benefits from this disclosure.
[0038] Still referencing Figure 2 The AVC OS 200 may also include a repository 222, which is a database for storing responses. Response data can be provided to a web user interface (not shown) via a web server 224, which provides an interface displaying the feed from the LLM module 204. User System Interface (USI) details 225 are user-specific implementation details of the web server 224 and the web user interface. The AVC OS 200 may also provide responses to a debugger 226 to debug any responses that are indicated as informal or otherwise inappropriate (e.g., output that does not match the specified format, or an unusual classification of "greetings or small talk").
[0039] Further reference Figure 2 In some embodiments, the LLM module 204 may be prompted twice to remove speech inconsistencies. The first response may be used for a "cleanup and optimization" prompt, while the second response is used for command recognition. In such embodiments, the second prompt may be supplemented by the LLM module 204 based on the first response or based on the first response.
[0040] In another embodiment, AVC OS 200 provides AVC processing core 100 with the ability to configure itself based on who is attending the meeting. For example, settings (volume, brightness, etc.) for various peripheral devices 104 can be configured based on meeting participants. Within each system, presets can be provided, which have a set of settings that can configure a room for a specific purpose. For example, a meeting room can be used for meetings or presentations. AVC OS 200 can identify participants in the room using voice detection based on logical rules and match those voice signals to a table with known user voices. Alternatively, AVC OS 200 can identify meeting participants from meeting invitations based on logical rules, and if a user profile or history associated with a participant exists, the system can adjust settings and configurations based on the user profile or history.
[0041] In other embodiments, AVC OS 200 provides the ability to implement various system configuration options, diagnostic options, or debugging options. For example, if a fault exists in the system (discovered by debugger 226), response parser 216 will read the code to handle the fault and then re-execute the last few entries in the event log.
[0042] In other embodiments, the AVC OS 200 can access the audio system via audio interface 228 to record audio to a file. Here, the audio interface can receive audio data from a microphone or from a person at the other end of the laptop and transmit that data to transcription interface 212. In this case, a set of verbal command detections (and corresponding command sets) can be present to start or stop recording. In yet another embodiment, the system can be audibly instructed to send the recording via email to a desired person / email address. Alternatively, audio interface 228 can receive audio data from TTS 206 and then play that audio data through the system's speakers.
[0043] In yet another embodiment, the AVC OS 200 provides the user with the ability to issue verbal commands to control the paging system, which is part of a networked AVC system. Such verbal commands may be to initiate paging, broadcast a paging message, and then end the broadcast.
[0044] In other examples, the AVC OS 200 also provides the ability to perform voice activity detection, implemented within the DSP / AE 208, so that no blank audio is detected; only segments containing voice are present. Here, the AVC OS 200 determines whether the sound from a microphone or other source contains human speech, rather than, for example, silence, keyboard clicks, rustling paper, dog barking, car horns, music playback, etc. Only the voice signal needs to be collected and sent to a voice transcription service, saving network bandwidth, processing time, and unnecessary transcription costs.
[0045] When any of the various subsystems of AVC OS 200 processes data at a different rate or for data blocks of different sizes, system queue 215 may be needed to hold the data output from one system, which can then be processed by the next system. For example, queue 215 may be needed to allow a certain amount of LLM data to be collected and then sent to the transcription API 210.
[0046] In other illustrative embodiments, transcription data may be received via transcription API 210 from a third-party (“3P”) provider 230, such as a Teams or Zoom platform. This platform will provide third-party transcription access, meaning that the AVC OS 200 can work with another system providing voice transcription, which will be injected into that system at queue 215 and all subsequent processing steps will be applied, as described herein.
[0047] exist Figure 2 In yet another alternative embodiment, Figure 2 Each component (except RE 218 and AE 208) can run in the cloud or on another computer, instead of on the processing core 100. These and other modifications will be apparent to those skilled in the art who will benefit from this disclosure.
[0048] In view of the foregoing, Figure 3 This is a flowchart illustrating a general method for performing one or more actions on a peripheral device according to an illustrative embodiment of this disclosure. At block 302 of method 300, the AVC system, as described herein, detects a verbal command issued by a user. At block 304, the AVC system transcribes the detected verbal command / voice / utterance. At block 306, the AVC system transmits the transcribed data to LLM module 204. At block 308, the LLM module interprets the text / transcribed data to identify instructions and corresponding command sets. At block 310, AE / DSP 208 receives the identified command set from LLM module 204 and sends it to command parser 216. At block 312, the AVC system performs an action corresponding to the parsed command set. Such an action could be, for example, adjusting settings of a peripheral device (such as HVAC settings), speaker or display volume, brightness of a display or touchscreen controller, collecting factual data from a data repository, adjusting blinds, etc.
[0049] Figure 4 This is a block diagram of an AVC OS according to certain illustrative embodiments of this disclosure. In this example, the AVC OS400 is... Figure 1The AVC OS 102. AVC OS 200 employs an LLM module 204, which includes a pre-configured command set for recognition using speech-to-text transcription. However, in the example of AVC OS 400, the AVC system can learn new commands, for example, from a user. AVC OS 400 includes some of the same components as AVC OS 200, with new components used to implement the described learning functionality. In a first embodiment, AVC OS 400 can be taught or trained via a web interface 402, or in a second alternative embodiment, by listening live to audible instructions provided by the user and processed at audio engine 208 via microphone 404.
[0050] In practice, the user will define (via a web interface or verbally via a microphone) what new thing he or she wants to control, and the instructions for that prompt to adjust the control, and then feed it to the language model. The prompt instructions will be specific codes that, when executed, perform commands, such as controlling peripheral devices, processing audio, video, or control signals in a certain way (echo cancellation, gain adjustment, etc.). The LLM module 204 distinguishes specific codes by interpreting the transcription transcribed from the transcription API 210. The LLM module 204 is “taught” by system prompts or fine-tuning what to look for and how to respond. In an earlier example of teaching the system to recognize silence, the instructions for the prompt would be, for example, “respond with… when the user expects the system to be muted.” Thus, the user teaches the LLM module 204 new controls and corresponding commands for adjusting at least one of the controls in a peripheral device or platform / application. In some embodiments, system prompts are stored in the AI interface 214 and sent to the LLM module 204. A web user interface 402 is provided to allow the user to add and edit commands.
[0051] refer to Figure 4 This illustrates certain illustrative embodiments of the AVC OS 400 according to this disclosure. Note that the same reference numerals refer to... Figure 2The components described. AVC OS 400 is a block diagram illustrating how LLM module 204 learns new verbal commands that are not default in the system. In a first illustrative embodiment, AVC OS 400 learns new commands via a web interface for adding and editing commands through web server 224. Web page 402 contains a list of all named control / command sets retrieved from configuration database 406. For example, a user selects a checkbox next to a button icon via web interface 402, types the description “Enable background music”, and then presses submit. This information is sent to web server 224 using a POST request (POST verb) and is inserted into prompt / response database 222. The LLM interface (214) then uses a template to construct an appropriate system prompt based on these values. For example, the prompt could be: “Output command when user wants to select background music”.<control_command> bgm_select=true< / control_command> Note that in some embodiments, if the system designer is sufficiently specific in the naming control, the LLM module 204 can determine the description based on the name itself.
[0052] Once the AVC OS 400 has learned from the web server 224, various new prompts / responses are stored in the prompt / response database 222. When the LLM module 204 receives a verbal command (e.g., "Turn on the disco ball" or "Turn on the lights"), the response parser 216 interprets it using the prompt database 222 and passes it to the control command handler 407. The control command handler 407 receives the parsed command and identifies it as a command required by the control engine 408 to perform an action. Subsequently, the control command handler 406 transmits the command to the control engine 408, which instructs the audio engine 208 to perform the corresponding operation, such as opening blinds, muting something, adjusting the volume, or controlling the disco ball 410 (e.g., starting the ball to spin, turning the ball on / off, turning on the ball lights, retracting the ball to the ceiling, etc.).
[0053] In an alternative embodiment, the AVC OS 400 can learn by receiving verbal instructions from the user on how to interpret new verbal commands. The verbal instructions can be received via microphone 404 and processed at DSP 208. Here, for example, the user would say, “Create a new command to enable background music using bgm_select=true.” The LLM module 204 recognizes this as an instruction to create a new command and responds as follows, for example, “<create_command> <prompt> enable backgroundmusic< / prompt> <response> bgm_select=true< / response>< / create_command>This is dispatched by the response parser 216 to the new command handler 412, which inserts the prompt and response into the database 222. Therefore, in the future, whenever the LLM module 204 hears "Turn down the background music", the response parser 216 obtains the corresponding prompt / response from the database 222 and sends it to the control engine 408 to perform the operation.
[0054] Various other command handlers 413 serve as “connectors” to other systems. For example, to send a meeting transcript via email, a command handler is needed to send the email by connecting to an email server. To create a Jira entry (used to track tasks and vulnerabilities in a software development team), a command handler is needed to connect to a Jira server. This will respond to verbal commands such as “mark vulnerability 23456 as resolved.” In another example, to query the weather, a command handler is needed to connect to a weather server. In this case, the weather command handler sends the weather data to the TTS interface 420 after receiving the result from the server. In yet another embodiment, scheduling a meeting would require a command handler to connect to a calendar server. This mechanism also enables a variety of other operations, such as pushing messages to various types of chat channels (Slack, Yammer, Microsoft Teams, or Discord, etc.); sending commands via its unique API to control other equipment that cannot be controlled through standard control command handlers (lights, thermostats, locks, cameras, blinds and curtains, televisions, etc.); creating projects or adding items to a to-do list; setting alarms and timers; or controlling streaming services (such as Netflix, Spotify, etc.).
[0055] In other illustrative embodiments, when the AVC OS 400 receives a verbal command to reconfigure the system, it uses the system configuration command handler 414. For example, reconfiguration might be changing the system from conference room mode to cinema mode. Here, the response parser 216 will see these reconfiguration commands received via the LLM module 204 and send them to the handler 414. The system controller 416 (also known as the configuration manager) then reconfigures the audio engine 208 and the control engine 408. The new configuration can be stored in the configuration database 406 and retrieved by the system controller 416 once identified via the LLM module 204.
[0056] Audio command processor 418 processes user commands that request or ask questions. The LLM module 204 recognizes the request or question and generates a response designed to be played audibly in the room. For example, when a factual question such as "How tall is Mount Everest?" is asked, audio command processor 418 is invoked, and the data "Mount Everest is the highest peak in the world, with an altitude of approximately 29,032 feet (8,849 meters)" is returned. This information is converted into audio via TTS interface 420, and the audio is sent to audio engine 208, where it can be mixed and sent to speakers.
[0057] In other embodiments, in addition to adding or appending information from the meeting to Jira or Confluence, AVC OS 400 can also retrieve relevant information from Jira, Confluence, etc., to provide context for the session. In such embodiments, command processor 413 performs these operations. Command processor 413 pushes material (images, text from a page, or text transcribed from speech to text from the meeting, etc.) from Jira or Confluence to a display or web interface 402 via web server 224 (e.g., using web sockets). This can be achieved in various ways, including, for example, by retrieving augmented generation (“RAG”) programs, such as IBM’s Watsonx.
[0058] In other illustrative embodiments, where multiple AVC processing cores 100 are networked together (e.g., in the cloud), when one AVC processing core is taught a command, the taught command can be transmitted to other secondary AVC processing cores on the network. Therefore, all cores can learn from this one taught core, or similarly from many other taught processing cores.
[0059] Figure 5 This is a flowchart of a computer-implemented method 500 of the present disclosure. At block 502, an AVC operating system is implemented on an AVC processing core. The processing core is communicatively coupled to one or more peripheral devices. As described herein, the AVC processing core is configured to manage and control the audio, video, and control features of the peripheral devices. At block 504, an LLM module communicatively coupled to the AVC processing core is used to detect one or more verbal commands from the user. At block 506, the system performs an action corresponding to the verbal command on the peripheral device or the AVC processing core.
[0060] The methods and embodiments described herein further relate to any one or more of the following paragraphs: 1. A computer-implemented method comprising: implementing an AVC operating system on an AVC processing core communicatively coupled to one or more peripheral devices, the AVC processing core being configured to manage and control audio, video, and control features of the peripheral devices; using a Large Language Model (“LLM”) module communicatively coupled to the AVC processing core to detect one or more verbal commands issued by a user; and performing actions corresponding to the verbal commands on the peripheral devices or the AVC processing core.
[0061] 2. The computer-implemented method as described in paragraph 1, wherein the LLM module executes a pre-configured set of commands to perform actions on these peripheral devices.
[0062] 3. A computer-implemented method as described in paragraph 1 or 2, wherein the LLM module executes a taught set of commands to perform actions on these peripheral devices, the taught set of commands being taught by the user.
[0063] 4. A computer-implemented method as described in any of paragraphs 1 to 3, wherein the taught set of commands is obtained from the user via a web interface.
[0064] 5. A computer-implemented method as described in any of paragraphs 1 to 4, wherein the taught set of commands is obtained from the user via a listening device.
[0065] 6. A computer-implemented method as described in any of paragraphs 1 to 5, wherein the AVC processing core transmits the taught command set to one or more secondary AVC processing cores, thereby teaching those secondary AVC processing cores.
[0066] 7. A computer-implemented method as described in any of paragraphs 1 to 6, wherein the LLM module is accessed from a cloud service, a local network service, or on the AVC processing core.
[0067] 8. A system comprising: one or more peripheral devices; and an AVC processing core communicatively coupled to the peripheral devices, the AVC processing core having an AVC operating system executable thereon for managing and controlling the peripheral devices, wherein the AVC processing core is configured to perform operations as described in any of paragraphs 1 to 7.
[0068] Furthermore, the methods described herein may be embodied in a system including a processing circuitry for implementing any method, or in a non-transitory computer-readable medium including instructions that, when executed by at least one processor, cause the processor to perform any method described herein.
[0069] Although various embodiments and methods have been shown and described, this disclosure is not limited to such embodiments and methods and will be understood to include all modifications and variations as will be apparent to those skilled in the art. Therefore, it should be understood that this disclosure is not intended to be limited to the specific forms disclosed. Rather, the invention will cover all modifications, equivalents, and alternatives falling within the spirit and scope of this disclosure as defined by the appended claims.
Claims
1. A computer-implemented method, comprising: An AVC operating system is implemented on an audio, video, and control ("AVC") processing core that can be communicatively coupled to one or more peripheral devices. The AVC processing core is configured to manage and control the audio, video, and control features of these peripheral devices. The Large Language Model ("LLM") module, communicatively coupled to the AVC processing core, is used to detect one or more spoken commands issued by the user; and Perform actions corresponding to these verbal commands on these peripheral devices or the AVC processing core.
2. The computer-implemented method as described in claim 1, wherein, The LLM module executes a pre-configured set of commands to perform actions on these peripheral devices.
3. The computer-implemented method as described in claim 1, wherein, The LLM module executes a taught set of commands to perform actions on these peripheral devices, the taught set of commands being taught by the user.
4. The computer-implemented method as described in claim 3, wherein, The set of commands taught is obtained from the user via a web interface.
5. The computer-implemented method as described in claim 3, wherein, The set of commands taught was obtained from the user via a listening device.
6. The computer-implemented method as described in claim 3, wherein, The AVC processing core transmits the taught command set to one or more secondary AVC processing cores, thereby teaching the taught command set to these secondary AVC processing cores.
7. The computer-implemented method as described in claim 1, wherein, The LLM module is accessed from cloud services, local network services, or on the AVC processing core.
8. A system comprising: One or more peripheral devices; as well as An Audio, Video, and Control ("AVC") processing core, communicatively coupled to these peripheral devices, having an AVC operating system executable on it to manage and control these peripheral devices. The AVC processing core is configured to perform the following operations: The Large Language Model ("LLM") module, communicatively coupled to the AVC processing core, is used to detect one or more spoken commands issued by the user; and Perform actions corresponding to these verbal commands on these peripheral devices or the AVC processing core.
9. The system of claim 8, wherein, The LLM module executes a pre-configured set of commands to perform actions on these peripheral devices.
10. The system of claim 8, wherein, The LLM module executes a taught set of commands to perform actions on these peripheral devices, the taught set of commands being taught by the user.
11. The system of claim 10, wherein, The set of commands taught is obtained from the user via a web interface.
12. The system of claim 10, wherein, The set of commands taught was obtained from the user via a listening device.
13. The system of claim 10, wherein, The AVC processing core transmits the taught command set to one or more secondary AVC processing cores, thereby teaching the taught command set to these secondary AVC processing cores.
14. The system of claim 8, wherein, The LLM module is accessed from cloud services, local network services, or on the AVC processing core.
15. A non-transitory computer-readable storage medium storing instructions that, when executed by a computing system, cause the computing system to perform operations including: An AVC operating system is implemented on an audio, video, and control ("AVC") processing core that can be communicatively coupled to one or more peripheral devices. The AVC processing core is configured to manage and control the audio, video, and control features of these peripheral devices. The Large Language Model ("LLM") module, communicatively coupled to the AVC processing core, is used to detect one or more spoken commands issued by the user; and Perform actions corresponding to these verbal commands on these peripheral devices or the AVC processing core.
16. The computer-readable storage medium of claim 15, wherein, The LLM module executes a pre-configured set of commands to perform actions on these peripheral devices.
17. The computer-readable storage medium of claim 15, wherein, The LLM module executes a taught set of commands to perform actions on these peripheral devices, the taught set of commands being taught by the user.
18. The computer-readable storage medium of claim 17, wherein, The set of instructions taught is obtained from the user via a web interface or via a listening device.
19. The computer-readable storage medium of claim 17, wherein, The AVC processing core transmits the taught command set to one or more secondary AVC processing cores, thereby teaching the taught command set to these secondary AVC processing cores.
20. The computer-readable storage medium of claim 17, wherein, The LLM module is accessed from cloud services, local network services, or on the AVC processing core.