Device control system and control method supporting multimodal input
The device control system addresses the limitations of current voice assistants by supporting multimodal inputs to simplify and enhance the operation of multifunction printers, reducing errors and complexity through semantic parsing and text vector retrieval.
Patent Information
- Application Number
- TW114120904
- Authority / Receiving Office
- TW · TW
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2025-06-04
- Publication Date
- 2026-07-11
- Estimated Expiration
- 2045-06-03
AI Technical Summary
Current voice control technologies, such as Google Assistant and Apple Siri, are inadequate for complex tasks in enterprise environments, particularly with multifunction printers, due to limited processing power and lack of multimodal input capabilities, leading to inefficient and error-prone operation.
A device control system supporting multimodal input, including a multimodal input interface, a multimodal large-scale language device, and a database, that interprets and executes user commands through semantic parsing and text vector retrieval, enabling control of devices like multifunction printers without manual operation.
Facilitates intuitive and efficient device operation through multimodal inputs, reducing operational errors and complexity, especially for multifunction printers, by allowing voice, text, and image inputs, and eliminating the need for manual button presses.
Smart Images

Figure IMG-2_DRAW_114120904-A0305-14-0001-1 
Figure IMG-2_DRAW_114120904-A0305-14-0001-2 
Figure IMG-2_DRAW_114120904-A0305-14-0002-3
Abstract
Description
Technical Field
[0001] This case relates to a device control system and control method, and more particularly to a device control system and control method that supports multimodal inputs. Prior Technology
[0002] Current voice control technologies, such as Google Assistant, Apple Siri, and Google Home, are widely used in home environments and can perform basic commands such as turning lights on and off and adjusting air conditioning. However, they are insufficient to meet the actual needs of device operation in enterprise or professional work environments. These voice systems can be implemented through language models and are built on a single voice input architecture. However, they can only support simple command parsing and lack sufficient processing power for tasks that require advanced parameter settings, process judgment, and data understanding, especially on multifunction printers most commonly used in business environments.
[0003] Even the most common "copying" function on a multifunction printer can require dozens of button presses in practice. For example, to copy five color copies of a manual double-sided, a user needs to manually select the work mode, set the input source, and configure various settings such as paper size, color selection, double-sided page turning method, page layout, number of copies, image processing settings, scaling ratio, and advanced function controls. Each setting requires clicking several buttons and performing multiple operations on the multifunction printer's interface to complete the process. This process is quite complex and requires precise mastery of the device's operating procedures. For casual or inexperienced users, it may require consulting the user manual and often takes considerable time or multiple attempts to successfully copy the desired document. Furthermore, the operating methods for different multifunction printer models vary. Users need to spend additional time adapting to each model, and even then, differences in usage habits can lead to errors and require repeated attempts. In addition, some operation settings are logically related, and incorrect operation selection may cause the entire document to be corrupted, and also cause unnecessary waste in the printing process.
[0004] Current voice assistants perform poorly when handling complex tasks involving multifunction printers. For example, if given the command "Please photocopy five copies of this book double-sided in color," the existing voice system, limited by its natural language understanding capabilities, cannot correctly parse the sentence and translate it into the corresponding device operation command. Furthermore, multifunction printer operation often involves multi-layered settings and conditional logic; for voice assistants installed on the ground for security reasons, such complex tasks far exceed their processing capabilities. In addition, the lack of an effective control link between current voice assistants and multifunction printers prevents voice commands from being translated into specific machine operations, resulting in command execution failures. Moreover, the operating procedures differ between various models of multifunction printers, making it difficult to control different models with current voice assistants. Moreover, voice assistants cannot handle non-voice-related input content, thus limiting their operation control to a single voice input method.
[0005] Therefore, developing a device control system and control method that supports multimodal input and overcomes the above-mentioned shortcomings is an urgent need at present. Summary of the Invention
[0006] This invention provides a device control system supporting multimodal input and a suitable control method thereof. The device control system supporting multimodal input includes a multimodal input interface, a multimodal large-scale speech device, a database, and at least one device. The multimodal input interface receives multimodal input from a user and generates user commands. The multimodal large-scale speech device receives the user commands and performs semantic interpretation, parsing, and text vector retrieval to select and generate operation codes corresponding to the user commands from at least one database. The device receives the operation codes to execute the corresponding tasks. Therefore, users can intuitively and easily control various devices through multimodal input such as voice, text, or images without reading a user manual or operating complex buttons. Furthermore, the device control system and control method supporting multimodal input of this invention are particularly suitable for controlling the operation of multifunction printers.
[0007] To achieve the above objectives, this application provides a device control system supporting multimodal input. The device control system supporting multimodal input includes a multimodal input interface, a multimodal large-scale language device, at least one database, and at least one device. The multimodal input interface receives multimodal input from a user and generates user commands. The at least one database contains a plurality of opcodes and a plurality of text vectors. Each opcode corresponds to a text vector. The multimodal large-scale language device is connected to the multimodal input interface to receive user commands and perform semantic interpretation, parsing, and text vector retrieval to select and generate an opcode corresponding to the user command from the at least one database. The at least one device is connected to the multimodal input interface and / or the multimodal large-scale language device to receive opcodes. The text vector corresponding to each opcode is parsed from the contents of the operation manual of the corresponding at least one device. The at least one device executes the opcode according to its original opcode format to perform a corresponding task.
[0008] To achieve the above objectives, this invention provides a device control method that supports multimodal input. The device control method supporting multimodal input includes the following steps: (S1) providing a device control system supporting multimodal input, including a multimodal input interface, a multimodal large-scale language device, at least one database, and at least one device, wherein the at least one device is connected to the multimodal input interface and / or the multimodal large-scale language device, wherein the database includes a plurality of opcodes and a plurality of text vectors, each opcode corresponding to a text vector, and the text vector corresponding to each opcode is parsed from the content of the operation manual of the corresponding at least one device; (S2) the user performs multimodal input on the multimodal input interface and generates a user command; (S3) the multimodal input interface transmits the user command to the multimodal large-scale language device; (S4) after receiving the user command, the multimodal large-scale language device performs semantic interpretation, parsing, and text vector retrieval to select and generate an opcode corresponding to the user command from the at least one database; (S5) the at least one device receives the opcode and executes the opcode according to its original opcode format to perform the corresponding work. Simple Explanation of the Diagram
[0009] Figure 1 is a schematic diagram of the system architecture of the device control system supporting multimodal input in the first embodiment of this case.
[0010] Figure 2 is a schematic diagram of the system architecture of the device control system supporting multimodal input in the second embodiment of this case.
[0011] Figure 3 is a schematic diagram of the system architecture of the device control system supporting multimodal input in the third embodiment of this case.
[0012] Figure 4 is a schematic diagram of the system architecture of the device control system supporting multimodal input in the fourth embodiment of this case.
[0013] Figure 5 is a schematic diagram of the system architecture of the device control system supporting multimodal input in the fifth embodiment of this case.
[0014] Figure 6 is a schematic diagram of the system architecture of the device control system supporting multimodal input in the sixth embodiment of this case.
[0015] Figure 7 is a flowchart of a control method for a device control system supporting multimodal input according to an embodiment of this case.
[0016] Figure 8 is a flowchart of possible subsequent steps (S41) to be executed after step S4 in the control method shown in Figure 7.
[0017] Figure 9 is a detailed flowchart of step S4 in an embodiment of the control method shown in Figure 7.
[0018] Figure 10 shows a detailed flowchart of step S4 in another embodiment of the control method shown in Figure 9.
[0019] Figure 11 is a schematic diagram of the operation interface of a prior art multi-functional transaction machine. Implementation
[0020] Some typical embodiments that embody the features and advantages of this case will be described in detail in the following description. It should be understood that this case can have various variations in different forms, all of which do not depart from the scope of this case, and the descriptions and illustrations therein are essentially for illustrative purposes and not for limiting this case.
[0021] Please refer to Figures 1 through 6. Figures 1 through 6 are schematic diagrams of the system architecture of the device control system supporting multimodal input according to six embodiments of this invention. Please refer to Figures 1, 2, 3, and 4 first. As shown in the figures, the device control system 100 supporting multimodal input of this invention includes a multimodal input interface 1, a multimodal large language device 2, at least one device 3, and a database 4. The multimodal input interface 1 receives multimodal input from user 10 and generates user commands. The at least one database 4 contains a plurality of opcodes and a plurality of text vectors, wherein each opcode corresponds to a text vector. The multimodal large language device 2 is connected to the multimodal input interface 1 to receive user commands and perform semantic interpretation, parsing, and text vector retrieval to select and generate an opcode corresponding to the user command from the at least one database 4. The at least one device 3 is connected to the multimodal input interface 1 and / or the multimodal large language device 2 to receive the opcode. Next, at least one device 3 executes the opcode according to its original opcode format to perform a corresponding task. The text vector corresponding to each opcode is parsed from the operation manual of the corresponding device 3; that is, the operation manual of device 3 is parsed into text vectors of opcodes beforehand to facilitate text vector retrieval by the multimodal large-scale language device 2. Therefore, the multimodal large-scale language device 2 only needs to retrieve the text vector according to semantics to accurately correspond to the original opcode of device 3, ensuring the accuracy and timeliness of the instruction. Furthermore, in some embodiments, the original opcode format of device 3 refers to the opcode and its operational format generated each time a button is pressed to control device 3, and the opcode presented through this opcode format allows device 3 to operate accordingly. In this case, device 3 only needs to receive the opcode and execute it through the original opcode processing method to perform the corresponding task without adding additional hardware or software, thus avoiding the burden of equipment upgrades.
[0022] In some embodiments, the number of devices 3 can be multiple, meaning that the user 10 can perform multimodal input through the multimodal input interface 1, and further control multiple devices 3 through semantic interpretation, parsing, text vector retrieval, translation, and other processing of the multimodal large language device 2, so as to quickly complete multiple multitasking tasks required by the user 10. It should be noted that the multimodal input interface 1, the multimodal large language device 2, at least one device 3, and the database 4 may also include processors (not shown), storage units (not shown), communication units (not shown), etc., to maintain the operation and architecture of the device control system 100 supporting multimodal input, and the present invention does not limit this. The device control system supporting multimodal input of the present invention and its applicable control method can be applied to various home appliances or different models of multifunction office machines. In particular, the device control system 100 supporting multimodal input of this invention and its applicable control method are particularly suitable for controlling the operation of multifunction office machines, that is, device 3 is a multifunction office machine. The actual control method described below will be described using a multifunction office machine.
[0023] In some embodiments, the multimodal large-scale language device 2 transmits an operation code to the multimodal input interface 1, which then transmits the operation code to the device 3 to execute the corresponding task. Specifically, the operation code refers to the instruction input to the device 3 in order to perform the task. That is, the operation code is the input program generated when various tasks are performed, such as when a button is pressed. Therefore, by directly transmitting the operation code to the device 3, the user 10 can be eliminated from personally going to the device 3 to press a button, and the device 3 can be remotely and conveniently controlled.
[0024] In some embodiments, after receiving a user instruction, the multimodal large language device 2 reads and parses the user instruction, generates a parsing result, and sends the parsing result back to the multimodal input interface 1 for the user 10 to confirm and obtain a confirmation instruction. The confirmation instruction is then sent back to the multimodal large language device 2, and the operation code is further read and retrieved based on the parsing result.
[0025] In some embodiments, the multimodal large-scale speech device 2 transmits the parsing result and the corresponding opcode to the multimodal input interface 1 and stores them. When the user 10 performs the same multimodal input again, the multimodal input interface 1 selects the stored parsing result and directly transmits the corresponding opcode to at least one device 3 with which a communication connection has been established to perform the corresponding task. Alternatively, the user 10 can also select a stored parsing result on the multimodal input interface 1 to have the corresponding opcode directly transmitted to at least one device 3 with which a communication connection has been established to perform the corresponding task. This design allows the user 10 to control the device 3 to repeatedly perform similar tasks on a large scale.
[0026] In some embodiments, multimodal input includes voice input, text input, image input, or any combination thereof, or other input methods, such as sensing gaze direction or gesture direction, etc.
[0027] In some embodiments, the multimodal input interface 1 may be, but is not limited to, a web page user interface or a mobile application.
[0028] In some embodiments, as shown in Figures 1 and 2, the multimodal input interface 1 may be integrated into an electronic device 11, a server 20, or the multimodal large voice device 2, but is not limited thereto. The electronic device 11 may be, but is not limited to, a mobile phone, a tablet computer, or a computer; the server may be, but is not limited to, a local server or a cloud server.
[0029] In some embodiments, as shown in Figures 1 and 3, the multimodal large language device 2 is disposed within a server 20 or a multimodal input interface 1, wherein the server 20 may be, but is not limited to, a local server or a cloud server.
[0030] In some embodiments, as shown in Figure 4, the multimodal large-scale language device 2 performs text vector retrieval in the database 4 to retrieve text vectors corresponding to the operation codes of at least one device 3, thereby obtaining the corresponding operation codes. Specifically, the database 4 contains various operation codes, and each operation code can be used to enable the device 3 to perform a corresponding task. Therefore, by retrieving different databases 4 or the contents stored in the database 4, operation codes for controlling different devices 3 or for controlling various other tasks can be found, thereby controlling different devices 3 to perform various tasks. In other words, for the user 10, there is no need to find the operation manual corresponding to the device 3, nor is there a need to consult the operation manual, nor is there a need to perform cumbersome key operations; the user can easily control the device 3 to perform the required tasks through multimodal input such as voice, text, and images.
[0031] In some embodiments, the multimodal large language device 2 includes a conversion module 21, a retrieval-augmented generation module 22, and a multimodal large language model 23. The conversion module 21 is configured to convert and integrate the content of multimodal input, such as speech input, text input, image input, or any combination thereof, for the multimodal large language model 23 to interpret and parse, and generate the parsing result. The retrieval-augmented generation module 22 performs text vector retrieval in the database 4 based on the parsing result to obtain the corresponding opcode of at least one device 3. The architecture in this embodiment can be implemented through retrieval-augmented generation (RAG) technology. It should be noted that the present invention does not limit the architecture of the multimodal large language model 23. For example, it can be any open-source multimodal large language model, or it can be varied according to the actual implementation situation, and is not limited thereto.
[0032] In some embodiments, as shown in Figure 5, the database 4 is disposed within the multimodal large language device 2, that is, the multimodal large language device 2 simultaneously includes components such as a conversion module 21, a retrieval enhancement generation module 22, a multimodal large language model 23, and the database 4.
[0033] In some embodiments, as shown in Figure 6, the database 4 is located within the multimodal input interface 1. Thus, when a user command is transmitted from the multimodal input interface 1 to the multimodal large-scale speech device 2, the multimodal large-scale speech device 2 can reconnect to the multimodal input interface 1 and retrieve data from the database 4 therein to perform subsequent operations.
[0034] It should be noted that each database 4 can be installed and configured on different platforms, and different operation codes and corresponding text vectors for controlling different devices 3 can be stored in different databases 4 according to actual needs. Therefore, the corresponding operation code can be found from the same or different databases 4 according to the type and model of the device 3 to be operated, and is not limited to this.
[0035] Please refer to Figure 7, in conjunction with Figure 1. Figure 7 is a flowchart of a control method for a device control system supporting multimodal input according to an embodiment of this invention. The device control method supporting multimodal input of this invention includes the following steps. First, in step S1, a device control system 100 supporting multimodal input is provided. The device control system 100 includes a multimodal input interface 1, a multimodal large-scale speech device 2, at least one device 3, and at least one database 4. At least one device 3 is connected to the multimodal large-scale speech device 2 and / or the multimodal input interface 1. The database 4 includes a complex number of opcodes and a complex number of text vectors. Each opcode corresponds to a text vector, and the text vector corresponding to each opcode is parsed from the content of the operation manual of the corresponding at least one device. In some embodiments, device 3 may be, but is not limited to, a multifunction transaction machine. Next, in step S2, user 10 performs multimodal input on the multimodal input interface 1 and generates user commands. Multimodal input includes voice content, text content, image content, or any combination thereof, and is not limited thereto. In other words, user 10 can input the work content to be performed into the multimodal input interface 1 in this step.
[0036] As shown in Figures 1 and 7, in step S3, after receiving the user command, the multimodal input interface 1 transmits the user command to the multimodal large-scale speech device 2. Then, in step S4, after receiving the user command, the multimodal large-scale speech device 2 performs semantic interpretation, parsing, and text vector retrieval on the user command, selecting and generating the corresponding operation code from the database 4. Through this text vector retrieval process, the user 10 only needs to perform quick and easy modal input, without having to search for or consult thick paper manuals, or learn the operation methods of various devices or models. This allows for quick and easy control of various types of devices 3, while also avoiding operational errors caused by different usage habits when operating different devices or models. In some embodiments, the multimodal large-scale speech device 2 can use Retrieval-Augmented Generation (RAG) technology to interpret and retrieve data from the database 4 based on the parsing results, for faster and more accurate interpretation and retrieval, but this is not a limitation.
[0037] Next, in step S5, at least one device 3 receives the corresponding opcode, and can then perform the corresponding operation. Moreover, at least one device 3 only needs to run the opcode according to its original opcode format.
[0038] Please refer to Figure 8 and Figure 1 simultaneously. Figure 8 is a flowchart of a possible subsequent step (S41) executed after step S4 in the control method shown in Figure 7. As shown in Figures 1 and 8, step S41 is included after step S4. That is, after the multimodal large language device 2 generates the operation code, as described in step S41, the multimodal large language device 2 transmits the operation code to the multimodal input interface 1, and then the multimodal input interface 1 further transmits the operation code to at least one device 3.
[0039] Alternatively, in other embodiments, as shown in Figure 9, which is a detailed flowchart of step S4 in the embodiment of the control method shown in Figure 7, after the multimodal large language device 2 generates the opcode in step (S4), the following steps are further included: First, in step S42, after receiving the user's instruction, the multimodal large language device 2 reads and parses the user's instruction and generates a parsing result. Next, as described in step S43, the multimodal large language device 2 sends the parsing result back to the multimodal input interface 1 for user 10 to confirm. Then, as described in step S44, if the user 10 confirms that the parsing result is correct, it is considered "yes," and a confirmation command is generated on the multimodal input interface 1, and this confirmation command is sent back to the multimodal large language device 2. In some embodiments, the multimodal input interface 1 (or multimodal large language device 2) can generate a preview screen (not shown) based on the parsing result and display it on the display interface (not shown) of the multimodal input interface 1 (or multimodal large language device 2). This preview screen may, for example, display zoom in / out of a document or image, print preview, etc., but is not limited to this. In this way, the user 10 can quickly confirm whether the content of the previously operated multimodal input matches the content of the parsing result through this preview screen and can proceed with subsequent procedures. Then, as described in step S45, the multimodal large language device 2 further semantically interprets, parses, and retrieves the corresponding operation code from the text vector based on the confirmation command. Next, in step S46, the multimodal large language device 2 transmits the corresponding operation code to at least one device 3. In this way, through this reconfirmation step, the reliability of the device control system 100 supporting multimodal input can be improved, and printing waste caused by user 10 input errors can be reduced.
[0040] In other embodiments, in step S43, if user 10 confirms that the parsing result is inconsistent with the original multimodal input, the confirmation result is considered "no". At this time, as described in step S47, the multimodal input interface 1 sends an error message to the multimodal large language device 2. The multimodal large language device 2 re-corrects, interprets, and parses the data based on the error message to generate a new parsing result, and then executes step S43 again to confirm whether the parsing result is correct. Through this repeated confirmation process, the multimodal large language device 2 accumulates more reference data and can learn and understand user commands generated by user 10 through multimodal input more quickly and accurately, thereby enabling subsequent interpretation, parsing, and retrieval processes.
[0041] Please refer to Figures 1 and 10 simultaneously. Figure 10 is a detailed flowchart of step S4 in another embodiment of the control method shown in Figure 9. As shown in Figures 1 and 10, as described in step S44, when user 10 confirms the parsing result as yes, a confirmation command is generated on the multimodal input interface 1 and sent back to the multimodal large language device 2. Then, step S45 is executed, which further includes the following steps: as described in step S48, the multimodal large language device 2 transmits the parsing result and the corresponding operation code to the multimodal input interface 1 and stores them. Next, in step S49, when user 10 performs the same multimodal input again, the multimodal input interface 1 selects the stored parsing result and directly transmits the corresponding operation code to the at least one device 3 with which a communication connection has been established, and executes the subsequent step S5, that is, device 3 receives the pre-stored operation code and performs the corresponding work. In this way, the user 10 can quickly select and control the device 3 to perform the same work during the process, thereby reducing the complexity of operation.
[0042] Figure 11 is a schematic diagram of the operation interface of a prior art multifunction printer. Previously, to merge four pages of an original document into one page and enlarge the photocopy ratio, the user had to stand in front of the traditional multifunction printer's interface and, as shown in Figure 11, first click the merge mode, press the button twice to enter the 4-in-1 mode, and then select the arrangement order. Next, the user had to select and enter the enlargement ratio menu, choose manual mode, and press the number keys to select and adjust the enlargement ratio. During this process, the user had to press different option buttons multiple times, which was not only lengthy and complex, inconvenient to operate, but also prone to operational errors, leading to printing failures or waste. However, through the aforementioned device control system 100 and its control method supporting multimodal input, not only are numerous manual control opportunities saved and the possibility of manual operation errors greatly reduced, but the operation is also simple and can accurately execute the required tasks, thus achieving the effect of saving time and effort.
[0043] For example, through the multimodal input-supporting device control system 100 and control method of this case, user 10 can verbally instruct multimodal input interface 1: "Combine four pages into one page and enlarge it by 140%." The voice content of "Combine four pages into one page and enlarge it by 140%" constitutes multimodal input content, and multimodal input interface 1 generates a user command based on this. The user command is then transmitted to multimodal large-scale speech device 2, and converted and integrated into text content by conversion module 21. After receiving the user command, multimodal large-scale speech device 2 interprets and understands the user command and generates a parsing result. Multimodal large-scale speech device 2 sends the parsing result back to multimodal input interface 1, waiting for user 10's confirmation to receive a confirmation command. If user 10 believes the parsing result is incorrect, multimodal input interface 1 sends an error message to multimodal large-scale speech device 2 for correction, interpretation, and parsing to generate a new parsing result. Then, the multimodal large language device 2 performs semantic interpretation, parsing, and text vector retrieval on the parsing result based on the confirmation command, in order to select and generate an operation code corresponding to the user's command from the database 4. Specifically, the retrieval enhancement generation module 22 can retrieve text vectors in the database 4 based on the parsing result, find the text vector closest to the user's command or parsing result, find the corresponding operation code based on the text vector, and provide the retrieved content to the multimodal large language model 23 for semantic interpretation and analysis to generate the operation code. The operation code generated here is the control command generated by the user 10 when pressing the button, as shown in Figure 11. The operation code can be transmitted to the multimodal input interface 1 for storage, and then transmitted from the multimodal input interface 1 to the device 3 for execution, or directly transmitted from the multimodal large language device 2 to the device 3 for execution, and is not limited to this. The connection arrangement between the multimodal input interface 1, the multimodal large language device 2, and the device 3 can be implemented in different ways as shown in Figures 1 to 6 of this case, or can be varied according to the actual implementation situation, and can also extend to multiple devices 3, and is not limited to the aforementioned implementation.
[0044] In summary, this invention provides a device control system and method supporting multimodal input. The device control system includes a multimodal input interface, a multimodal large-scale speech device, and at least one device. The multimodal input interface receives multimodal input from the user and generates user commands. The multimodal large-scale speech device receives the user commands and interprets, parses, and retrieves them to generate operation codes. The device receives the operation codes to execute the corresponding tasks. Therefore, users can intuitively and easily operate various devices without reading manuals or operating complex buttons. Furthermore, the device control system and method supporting multimodal input of this invention are particularly suitable for controlling the operation of multifunctional office machines. That is, through multimodal input by the user, the flexibility and intelligence of operation can be improved, and intuitive and diverse interaction methods can be achieved, applicable to diverse workplace equipment, and improving workplace equipment management efficiency.
[0045] 100: Device Control System 10: Users 1: Multimodal Input Interface 11: Electronic devices 2: Multimodal large-scale language device 20: Server end 21: Conversion Module 22: Retrieval Enhancement Generation Module 23: Multimodal Large-Scale Language Model 3: Device 4: Database S1~S5: Steps of the device control method supporting multimodal input S41: Sub-step S42~S47: Sub-steps S48~S49: Sub-steps
Claims
1. A device control system supporting multimodal input, comprising: A multimodal input interface that receives a user's multimodal input and generates a user command; At least one database containing a complex number of opcodes and a complex number of text vectors, wherein each opcode corresponds to a text vector; a multimodal large-scale language device connected to the multimodal input interface for receiving the user command and performing semantic interpretation, parsing, and text vector retrieval to select and generate the opcode corresponding to the user command from the database; and at least one device connected to the multimodal input interface and / or the multimodal large-scale language device for receiving the opcode, wherein the text vector corresponding to each opcode is parsed from the contents of an operation manual of the corresponding at least one device, wherein the at least one device runs the opcode according to its original opcode format to perform a corresponding task.
2. A device control system supporting multimodal input as described in claim 1, wherein the multimodal large language device transmits the opcode to the multimodal input interface, and the multimodal input interface then transmits the opcode to the at least one device to perform the corresponding operation.
3. The device control system supporting multimodal input as described in claim 1, wherein after receiving the user's instruction, the multimodal large language device reads and parses the user's instruction, generates a parsing result, and sends the parsing result back to the multimodal input interface for the user to confirm, and obtains a confirmation instruction, and then sends the confirmation instruction back to the multimodal large language device, and further performs semantic interpretation and text vector retrieval based on the parsing result to obtain the opcode.
4. The device control system supporting multimodal input as described in claim 3, wherein the multimodal large language device transmits the parsing result and the corresponding operation code to the multimodal input interface and stores them; when the user performs the same multimodal input again, the multimodal input interface selects the stored parsing result and directly transmits the corresponding operation code to the at least one device with which a communication connection has been established to perform the corresponding work.
5. A device control system supporting multimodal input as described in claim 1, wherein the multimodal input includes a voice input, a text input, an image input, or any combination thereof.
6. A device control system supporting multimodal input as described in claim 1, wherein the multimodal input interface is a web user interface or a mobile application.
7. A device control system supporting multimodal input as described in claim 6, wherein the multimodal input interface is integrated into an electronic device, a server, or the multimodal large-scale speech device, wherein the electronic device is a mobile phone, a tablet computer, or a computer, and the server is a local server or a cloud server.
8. A device control system supporting multimodal input as described in claim 6, wherein the multimodal large language device is located on a server or within the multimodal input interface, the server being a local server or a cloud server.
9. A device control system supporting multimodal input as described in claim 3, wherein the multimodal large language device includes a conversion module, a retrieval enhancement generation module, and a multimodal large language model, the conversion module being configured to convert and integrate the content of one of the multimodal inputs—a voice input, a text input, an image input, or any combination thereof—for the multimodal large language model to interpret and parse, and to generate the parsing result, the retrieval enhancement generation module retrieving the parsing result from the at least one database to obtain the corresponding opcode for the at least one device.
10. A device control system supporting multimodal input as described in claim 1, wherein the database is located within the multimodal large-scale language device.
11. A device control system supporting multimodal input as described in claim 1, wherein the database is located within the multimodal input interface.
12. A device control system supporting multimodal input as described in claim 1, wherein the device is a multi-functional transaction machine.
13. A device control method supporting multimodal input, comprising the steps of: (S1) providing a device control system supporting multimodal input, comprising a multimodal input interface, a multimodal large-scale language device, at least one device and at least one database, wherein the at least one device is connected to the multimodal input interface and / or the multimodal large-scale language device, wherein the database comprises a plurality of opcodes and a plurality of text vectors, each opcode corresponding to a text vector, and the text vector corresponding to each opcode being parsed from the contents of an operation manual of the corresponding at least one device; (S2) a user performs a multimodal input on the multimodal input interface and generates a user command; (S3) the multimodal input interface transmits the user command to the multimodal large-scale language device; (S4) after receiving the user command, the multimodal large-scale language device selects and generates the opcode corresponding to the user command from the at least one database through semantic interpretation, parsing and text vector retrieval; and (S5) The at least one device receives the opcode, and the at least one device runs the opcode according to its original opcode format to perform a corresponding task.
14. The device control method supporting multimodal input as described in claim 13, wherein after step (S4), it further comprises the step of: (S41) the multimodal large language device transmitting the opcode to the multimodal input interface, and the multimodal input interface transmitting the opcode to the at least one device.
15. The device control method supporting multimodal input as described in claim 13, wherein step (S4) comprises the steps of: (S42) after receiving the user instruction, the multimodal large language device reads and parses the user instruction and generates a parsing result; (S43) the multimodal large language device sends the parsing result back to the multimodal input interface for the user to confirm; (S44) if the user confirms, a confirmation command is generated on the multimodal input interface and the confirmation command is sent back to the multimodal large language device; (S45) the multimodal large language device further performs semantic reading, parsing, and text vector retrieval of the parsing result based on the confirmation command to obtain the corresponding opcode; and (S46) the multimodal large language device transmits the corresponding opcode to the at least one device.
16. The device control method supporting multimodal input as described in claim 15, wherein after step (S43), it further includes the step: (S47) if the user confirms no, the multimodal input interface sends an error message to the multimodal large language device, the multimodal large language device re-corrects, interprets and parses, to regenerate the parsing result, and executes step (S43) again.
17. The device control method supporting multimodal input as described in claim 15, wherein step (S4) comprises the steps of: (S42) after receiving the user's instruction, the multimodal large language device reads and parses the user's instruction and generates a parsing result; (S43) the multimodal large language device sends the parsing result back to the multimodal input interface for the user to confirm; (S44) if the user confirms, a confirmation command is generated on the multimodal input interface and sent back to the multimodal large language device; (S45) the multimodal large language device further performs semantic reading, parsing, and text vector retrieval of the parsing result based on the confirmation command to obtain the corresponding opcode; (S48) the multimodal large language device transmits the parsing result and the corresponding opcode to the multimodal input interface and stores them; and (S49) When the user performs the same multimodal input again, the multimodal input interface selects the stored parsing result and directly transmits the corresponding opcode to the at least one device with which a communication connection has been established.
18. The device control method supporting multimodal input as described in claim 13, wherein the device is a multi-functional transaction machine.