A smart control method, system, and medium based on multi-protocol task scheduling
By introducing a multi-protocol task scheduling method and a local AI computing chip into the home intelligent control system, the deep integration and collaborative operation of multiple heterogeneous systems are realized, solving the interoperability problem between systems, improving user experience and autonomous decision-making capabilities, and ensuring data privacy and security.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING LIUJINSUIYUE TECH CO LTD
- Filing Date
- 2025-12-23
- Publication Date
- 2026-05-26
AI Technical Summary
Existing home intelligent control systems lack a unified task platform mechanism, resulting in a lack of interoperability between subsystems, an inability to effectively understand complex instructions, isolation during execution, and a lack of autonomous decision-making capabilities, leading to a fragmented user experience.
It adopts an intelligent control method based on multi-protocol task scheduling, realizes semantic parsing, cross-protocol calling and closed-loop feedback control through local AI computing power chip, introduces multi-protocol connection service processing core, realizes deep integration and collaborative operation of multiple heterogeneous systems, and supports refined task processing and abnormal alternative solution generation.
It achieves deep integration and collaborative operation of multiple heterogeneous systems in the home environment, improves the intelligence level of human-computer interaction and the continuity of user experience, possesses human-like intelligent characteristics, and ensures data privacy and security.
Smart Images

Figure CN121367714B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of intelligent control technology, and in particular to an intelligent control method, system and medium based on multi-protocol task scheduling. Background Technology
[0002] With the rapid development of the smart home ecosystem, a complex Internet of Things (IoT) system has gradually formed within homes, consisting of various functional modules such as multimedia systems, security monitoring, environmental control, and lighting appliances. These systems are typically provided by different manufacturers, using their own independent communication protocols (such as MQTT, HTTP / REST, Zigbee, Bluetooth, ONVIF, etc.), data models, and service interfaces, resulting in a lack of effective interoperability between subsystems. Although some manufacturers have attempted to achieve initial interconnection through cloud platforms or general voice assistants, issues such as network latency, privacy concerns, and strong service dependencies make it difficult to meet users' advanced needs for real-time performance, security, and local autonomy.
[0003] Currently, when users issue commands containing multiple objectives, traditional systems often only recognize one part of the content or refuse to respond directly because they cannot match the preset command template, resulting in high interaction failure rates and a fragmented user experience. Furthermore, due to the lack of a unified task platform mechanism, most current smart home devices operate in isolation during execution. Even if an anomaly occurs in a certain operational step, it cannot trigger subsequent state evaluation and adaptive adjustments, resulting in the system exhibiting passive execution rather than proactive decision-making. For example, when attempting to retrieve a specific video clip, if no matching content is found, the system typically only returns a static message such as "No relevant record found," without further analyzing the cause of the failure or guiding the user to provide more precise information. Similarly, when controlling home appliances, if the target device is offline, the system cannot recommend alternative solutions or ask whether to enable a backup path, leading to task interruption and no recovery mechanism. Summary of the Invention
[0004] To address the aforementioned technical problems, this application provides an intelligent control method, system, and medium based on multi-protocol task scheduling.
[0005] Firstly, this application provides an intelligent control method based on multi-protocol task scheduling, employing the following technical solution:
[0006] An intelligent control method based on multi-protocol task scheduling, the intelligent control method comprising:
[0007] Receive raw instruction data input by the user, perform semantic parsing, extract instruction keywords, and generate a structured instruction data packet containing task type identifiers;
[0008] The structured instruction data packet is converted into multi-protocol connection service call parameters and sent to the multi-protocol connection service processing core through an internal communication interface; wherein, the multi-protocol connection service call parameters include a target system identifier and an operation instruction set;
[0009] The multi-protocol connection service processing core schedules the corresponding execution unit to execute the operation instruction set according to the target system identifier, and generates operation result data corresponding to the task type identifier;
[0010] The operation result data is analyzed to generate decision instruction data containing decision type identifiers;
[0011] Based on the decision type identifier, the task is refined, anomaly alternatives are generated, or task completion feedback is processed, and the processing results are converted into natural language feedback output.
[0012] By adopting the above technical solution, deep integration and collaborative operation of multiple heterogeneous systems in the home environment have been achieved. This solution not only solves the operational fragmentation caused by protocol barriers between traditional smart devices, but also, by introducing AI-driven context awareness and autonomous decision-making capabilities, endows the home central system with human-like intelligent characteristics such as understanding complex instructions, responding to execution deviations, and proactively optimizing paths. The entire technical system relies on local AI computing chips and solid-state storage resources, achieving a new paradigm of efficient, flexible, and scalable home intelligent control while ensuring data privacy and security. It represents an advanced direction for the integrated development of edge AI and the Internet of Things.
[0013] Secondly, this application provides an intelligent control system based on multi-protocol task scheduling, which adopts the following technical solution:
[0014] An intelligent control system based on multi-protocol task scheduling, the intelligent control system comprising:
[0015] The instruction generation module is used to receive raw instruction data input by the user, perform semantic parsing, extract instruction keywords, and generate a structured instruction data package containing task type identifiers;
[0016] The multi-protocol connection service invocation module is used to convert the structured instruction data packet into multi-protocol connection service invocation parameters and send them to the multi-protocol connection service processing core through an internal communication interface; wherein, the multi-protocol connection service invocation parameters include a target system identifier and an operation instruction set;
[0017] The scheduling and execution module is used by the multi-protocol connection service processing core to schedule the corresponding execution unit to execute the operation instruction set according to the target system identifier, and generate operation result data corresponding to the task type identifier;
[0018] The result status analysis module is used to perform status analysis on the operation result data and generate decision instruction data containing decision type identifiers.
[0019] The feedback processing module is used to perform task refinement processing, abnormal alternative generation, or task completion feedback processing based on the decision type identifier, and convert the processing results into natural language feedback output.
[0020] Thirdly, this application provides a computer-readable storage medium, which adopts the following technical solution:
[0021] A computer-readable storage medium storing a computer program that can be loaded by a processor and executed as in any of the methods in the first aspect.
[0022] In summary, this application offers at least one of the following beneficial technical effects: It not only accurately understands complex user commands input in natural language and achieves precise classification and structured processing through task type identification, but also leverages a multi-protocol connection service processing core to adapt to and route different communication protocols, solving the functional silo problem caused by protocol incompatibility in traditional smart home systems. Furthermore, after the execution result is returned, the system can generate context-aware decision command data based on state analysis, supporting detailed task follow-up, alternative solution recommendations in abnormal situations, and natural language feedback output of the final result, significantly improving the intelligence level of human-computer interaction and the continuity of user experience. The entire process relies on local AI computing power to achieve closed-loop processing, achieving a new paradigm of efficient, reliable, and scalable home intelligent control while ensuring data privacy and security. Attached Figure Description
[0023] Figure 1 This is a first flowchart illustrating an intelligent control method based on multi-protocol task scheduling, which is one embodiment of this application.
[0024] Figure 2 This is a second flowchart illustrating an intelligent control method based on multi-protocol task scheduling, which is one embodiment of this application.
[0025] Figure 3 This is a third flowchart illustrating one embodiment of the intelligent control method based on multi-protocol task scheduling in this application.
[0026] Figure 4 This is a schematic diagram of the fourth process of an intelligent control method based on multi-protocol task scheduling according to one embodiment of this application.
[0027] Figure 5 This is a fifth flowchart of an intelligent control method based on multi-protocol task scheduling according to one embodiment of this application. Detailed Implementation
[0028] To make the purpose, technical solution, and advantages of this application clearer, the following description is provided in conjunction with the appendix. Figure 1-5 The present application will be further described in detail below with reference to embodiments. It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of the application.
[0029] Currently, existing home intelligent control systems have not yet established a technical architecture capable of completing the entire process locally, from semantic understanding, task routing, cross-system calls to result feedback and dynamic decision-making. Especially when faced with non-standardized, conversational, and context-sensitive natural language input, they lack an intermediate coordination mechanism that can accurately extract intent and flexibly adapt to diverse backend services. This structural deficiency has kept home intelligent systems at the "function stacking" stage for a long time, failing to truly move towards an intelligent stage with cognitive reasoning capabilities and autonomous behavior regulation, thus limiting their widespread application and continuous evolution in real-life scenarios.
[0030] Based on this, this application discloses an intelligent control method based on multi-protocol task scheduling. This technical solution focuses on realizing unified semantic parsing, cross-protocol calling and closed-loop feedback control of various heterogeneous devices and subsystems in the home environment under localized deployment conditions. It is applicable to edge intelligent hardware platforms with certain AI computing power, such as home gateways, smart set-top boxes, AI NAS storage devices and integrated home intelligent computing centers.
[0031] Reference Figure 1 A smart control method based on multi-protocol task scheduling, the smart control method includes:
[0032] Step S101: Receive the raw instruction data input by the user and perform semantic parsing, extract instruction keywords and generate a structured instruction data packet containing task type identifier;
[0033] This step is not a simple text matching or keyword retrieval, but rather relies on a lightweight Natural Language Understanding (NLU) model deployed on a local AI computing chip to achieve context-aware parsing of the original speech or text commands. This process involves multiple NLU steps, including word segmentation, part-of-speech tagging, named entity recognition, and dependency parsing. Its purpose is to strip away redundant information from spoken expressions (such as interjections like "help me" or "can you?") and accurately extract the core elements constituting a valid operational intent: verbal actions (such as "search," "open," and "switch") and noun objects (such as "image," "light," and "television"). These keywords are then mapped to a pre-defined action-object knowledge graph, and combined with common behavioral patterns in a home setting, intent inference is performed to determine the task category to which the current command belongs.
[0034] For example, in the instruction "Play the picture of the baby smiling that I took last night," "play" corresponds to the playback action, and "baby smiling" is a description of image content features. Based on this, the system can determine that it is a media file processing task and assign a corresponding task type identifier (such as TASK_TYPE=MEDIA_SEARCH). The key to this step is realizing the transformation from vague human language to a machine-executable semantic framework, which is a prerequisite for the entire closed-loop control system to start.
[0035] Step S102: Convert the structured instruction data packet into multi-protocol connection service call parameters and send them to the multi-protocol connection service processing core through the internal communication interface; wherein, the multi-protocol connection service call parameters include the target system identifier and the operation instruction set;
[0036] Specifically, although structured instruction data packets have clear task attributes and parameter sets, there are significant differences in the communication protocols, data formats and calling methods relied upon by different back-end subsystems. Direct calling will lead to interface incompatibility issues.
[0037] To address this, the system introduces a "multi-protocol connection service" as an intermediate abstraction layer, responsible for converting structured instructions in a unified format into call parameters specific to the target system. This conversion process not only includes field mapping (such as converting the generic parameter "device_id" into the entity_id required by Home Assistant), but also embeds necessary metadata, such as the target system identifier (TARGET_SYSTEM_ID) for routing and the operation instruction set (ACTION_SET) for defining specific behavior sequences. For example, when the task type identifier points to home control, the target system identifier is set to HA_CONNECTOR, and the operation instruction set encapsulates the device address, control action code, and security token; if it points to media processing, the identifier is set to AI_NAS_PROCESSOR, and it carries filtering parameters such as time range and tag conditions.
[0038] Subsequently, these call parameters are transmitted to the multi-protocol connection service processing core via the device's internal high-speed serial communication interface (such as UART, SPI, or PCIe-based IPC channel), ensuring low-latency and highly reliable data transmission. This design breaks through the isolated operation of modules in traditional smart home systems, constructing a unified scheduling hub that enables heterogeneous systems to work collaboratively within the same logical framework.
[0039] In step S103, the multi-protocol connection service processing core schedules the corresponding execution unit to execute the operation instruction set according to the target system identifier, and generates operation result data corresponding to the task type identifier;
[0040] The multi-protocol connection service processing core serves as the scheduling hub of this method, undertaking crucial routing decisions and resource coordination functions. Based on the target system identifier in the received call parameters, it dynamically selects the corresponding execution unit and triggers its execution of the operation instruction set. This identifier-based scheduling mechanism is essentially a practical application of Service-Oriented Architecture (SOA) principles in a home edge computing environment.
[0041] For example, when the target system is an AI NAS media processor, the processing core will activate the intelligent analysis engine on the local storage system, driving it to perform content-level recognition on the audio and video files stored in the solid-state drive; when the target is a Home Assistant connector, it will send control requests to the Home Assistant platform via MQTT or RESTful API; and when the target is an HDMI switching system, the processing core will generate level control signals to act on the video path switching chip at the hardware level.
[0042] It should be noted that although all these execution units have different functions and protocols, they are logically considered as "pluggable" service components, which can be connected to the system as long as they conform to the predefined interface specifications. This loosely coupled design greatly improves the system's scalability and maintenance convenience, and also reserves technical paths for adding other types of services in the future (such as security alarms and energy management).
[0043] Step S104: Perform status analysis on the operation result data to generate decision instruction data containing decision type identifier;
[0044] The operation results data are not simply success / failure indicators, but structured responses with rich contextual information. For example, after completing image recognition, the AI NAS media processor not only returns whether a matching file was found, but also includes a list of object tags for each file and its confidence score; the Home Assistant connector returns the actual status of the device (such as whether the lights are actually on, whether the air conditioner temperature is adjusted correctly); and the HDMI switching chip also provides a status code indicating whether the path switching was successful.
[0045] Understandably, this fine-grained result data provides a solid foundation for subsequent state analysis. The state analysis phase is not a static judgment, but a dynamic reasoning process combining task type and expected output. The system evaluates execution effectiveness by parsing the status code field in the returned data: if the status code is "SUCCESS," it indicates the task has been completed as expected, entering the final feedback stage; if it is "NO_MATCH" or "DEVICE_OFFLINE," it indicates partial failure or missing external dependencies, requiring further intervention; if it is "ERROR" with an error code, it means an underlying communication or logic anomaly. Based on these analysis results, the system generates decision instruction data containing decision type identifiers to guide the next step. This mechanism endows the system with preliminary cognitive capabilities, transforming it from a passive tool for executing commands into an intelligent agent capable of proactively evaluating execution effectiveness and making adaptive responses.
[0046] Step S105: Based on the decision type identifier, perform task refinement processing, abnormal alternative solution generation, or task completion feedback processing, and convert the processing results into natural language feedback output.
[0047] Specifically, when the decision type is identified as "task refinement," it indicates that some information has been obtained but is insufficient to meet the user's original intent. In this case, the system will not simply report an error, but will instead construct follow-up questions based on the existing results to guide the user to provide more precise conditions. For example, if no photos of "baby smiling" are found, the system may generate a new list of parameters (such as narrowing the time window to the most recent 24 hours and increasing the weight of the keyword "smiley face") and re-enter the semantic parsing process to form an iteratively optimized search strategy.
[0048] When the decision type is identified as "abnormal alternative generation," the system exhibits fault tolerance and adaptability. For example, if a specified light fixture is offline, the system can automatically recommend other lighting devices in the same area as alternatives and solicit the user's opinion on whether to perform the replacement operation. This context-based reasoning-based alternative mechanism significantly improves the continuity of the user experience and the robustness of the system.
[0049] When the decision type is marked as "task completed," the system converts the final operation result data into feedback information in natural language, which is then output to the user via a speech synthesis module or on-screen display. This feedback is not only a result notification but can also include explanatory content, such as "We have found 3 videos containing cats for you, the longest of which is recorded by the living room camera at 18:27 yesterday," enhancing the transparency and user-friendliness of human-computer interaction.
[0050] The above implementation achieves deep integration and collaborative operation of multiple heterogeneous systems in a home environment. This technical solution not only solves the operational fragmentation caused by protocol barriers between traditional smart devices, but also, by introducing AI-driven context awareness and autonomous decision-making capabilities, endows the home central system with human-like intelligent characteristics, enabling it to understand complex instructions, address execution deviations, and proactively optimize paths. The entire technical system relies on local AI computing chips and solid-state storage resources, achieving a new paradigm of efficient, flexible, and scalable home intelligent control while ensuring data privacy and security. It represents an advanced direction for the integrated development of edge AI and the Internet of Things.
[0051] Reference Figure 2 As one implementation of step S101, the steps of receiving raw instruction data input by the user, performing semantic parsing, extracting instruction keywords, and generating a structured instruction data packet containing task type identifiers include:
[0052] Step S201: Receive raw instruction data input by the user, including voice signals or text strings;
[0053] Step S202: Preprocess the original instruction data, including extracting acoustic features from the speech signal or segmenting the text string to generate a segmented sequence;
[0054] When a user issues a command via voice (such as "Find the video of the kitten I took last week"), the system first receives a continuous analog audio signal. This type of signal is essentially a waveform change in the time domain and cannot be directly processed by the subsequent language understanding module. Therefore, it must undergo a series of signal-level to semantic-level conversion processes.
[0055] The first step is acoustic feature extraction, the core objective of which is to extract key information that characterizes the pronunciation from complex speech waveforms. In this process, the Mel Frequency Cepstral Coefficients (MFCC) algorithm is used to process the original speech. This algorithm simulates the nonlinear frequency response characteristics of the human auditory system, converting the linear frequency scale into a "Mel" scale that better matches human perception. Based on this, parameters such as short-time energy and spectral envelope are calculated, ultimately forming a 24-dimensional acoustic feature vector for each frame. These vectors not only preserve phoneme boundary information in the speech but also effectively suppress interference from background noise and individual speaker differences. Subsequently, an acoustic model constructed from deep neural networks (DNNs) is used to classify these feature vectors, mapping them to a series of possible phonemes. Then, a Hidden Markov Model (HMM) is used to model the temporal transition relationships between phonemes, realizing a step-by-step decoding process from "sound segment → phoneme sequence → candidate word string," ultimately outputting the corresponding text string.
[0056] This entire process forms the basic framework of Automatic Speech Recognition (ASR), which is especially important in the home environment because users often use colloquial, incomplete, or even regionally accented expressions. Only through high-precision acoustic-language joint modeling can the accuracy of subsequent semantic understanding be ensured.
[0057] Furthermore, for input already existing in text form (such as "turn on the living room lights" typed through a mobile app), the speech recognition stage is skipped, and the process proceeds directly to the text preprocessing stage. The core task of this stage is to segment and standardize the natural language text, i.e., word segmentation. Because Chinese lacks natural spaces between words, word segmentation becomes a crucial preliminary step in Chinese information processing.
[0058] Specifically, the system can use word segmentation tools based on statistical language models or pre-trained word vectors (such as BERT-BiLSTM-CRF) to segment continuous character streams into lexical units with independent semantics, forming an ordered "segmentation sequence". For example, "Please play the family video taken last night" will be correctly segmented into ["please", "play", "last night", "shot", "of", "family", "video"].
[0059] It should be noted that the system pays special attention to time adverbs (such as "just now" and "last month"), spatial qualifiers (such as "master bedroom" and "balcony"), and possessive structures (such as "dad's camera") during word segmentation. Although these components are not core verbs or nouns, they are crucial for the construction of the subsequent parameter list. After word segmentation, the resulting sequence is used as input to the natural language processing model for deeper semantic analysis.
[0060] Step S203: Perform semantic analysis on the segmented sequence using a pre-configured natural language processing model, extract noun keywords as instruction object identifiers, and extract verb keywords as operation action identifiers;
[0061] Specifically, the system invokes a lightweight natural language processing model deployed on a local AI computing chip to perform context-aware semantic analysis on the segmented sequence. Unlike simple keyword matching or regular expression rule engines, the model relied upon in this solution has the ability to understand the internal grammatical structure and semantic roles of sentences.
[0062] Specifically, the system uses a bidirectional long short-term memory network (Bi-LSTM) as its basic architecture. This model can simultaneously capture the dependencies between words in the context: forward propagation captures the information flow before the current word, while backward propagation integrates the contextual clues after it, thereby generating a hidden state vector containing global contextual information for each word segment.
[0063] Based on this, the model outputs the probability value of each word segment belonging to a "valid keyword". The system sets a threshold of 0.8, retaining only words with a confidence level higher than this threshold as valid keywords for subsequent processing. This effectively filters out components with no practical operational meaning, such as function words, auxiliary words, and conjunctions, thus improving the accuracy of instruction parsing.
[0064] More importantly, the system further incorporates dependency parsing technology to construct a dependency tree structure for sentences, clarifying the grammatical dependency relationships between words. For example, in the sentence "Put out the baby's smiling photos," "photos" is the object of "put out," and "baby smiling" is a relative clause modifying "photos." Through dependency analysis, the system accurately identifies "put out" as the predicate verb and marks it as an action identifier; while "photos," as the object of the action, is marked as an instruction object identifier. This role labeling mechanism based on syntactic structure is significantly superior to coarse-grained judgment relying solely on parts of speech, and is particularly suitable for complex nested sentences or multi-layered modification scenarios.
[0065] Step S204: According to the preset category mapping rule library, the instruction object identifier and operation action identifier are matched to the task type identifier to generate a structured instruction data packet.
[0066] The category mapping rule base is essentially an experience-based decision matrix that integrates domain knowledge and behavioral patterns. This rule base does not statically enumerate all possible instruction combinations, but is organized according to a three-dimensional logic of "object set—action set—target task". For example, when an instruction object is detected to belong to the media object set (such as "image", "video", "recording"), and the corresponding operation action falls within the analysis action set (such as "search", "play", "categorize"), the system immediately determines that the instruction belongs to the "media processing system" and assigns the corresponding task type identifier (such as TASK_MEDIA_SEARCH).
[0067] Similarly, if the object is a common home appliance name ("light", "air conditioner", "curtains") and the action is a control verb ("turn on", "turn off", "adjust"), then it is mapped to the "home control system" identifier.
[0068] Crucially, the system can also be configured with specific priority rules to handle ambiguous scenarios. For example, the command "switch TV" might be misinterpreted as device power control if broken down literally. However, by explicitly defining in the rule base that "when the object is 'TV' and the action is 'switch,' it is mapped to the video path switching system identifier," the system can accurately identify that the user's intent is to switch the HDMI path between the IPTV signal source and the local camera feed, rather than power control. This refined classification mechanism based on semantic combination enables the system to make reasonable inferences in ambiguous commands, greatly improving the accuracy of intent recognition.
[0069] Ultimately, the generated structured instruction data package not only contains a clear task type identifier but also carries a complete parameter list to guide the specific behavior of the backend execution unit. For example, for the instruction "find all videos containing dogs taken last week," the task type identifier in the system-generated data package is "media retrieval," and the parameter list is automatically filled with constraints such as the time range (last 7 days), file format (mp4 / avi), and target tag ("dog"). For control instructions such as "lower the bedroom air conditioner temperature," the parameter list includes details such as the device physical address code (e.g., HA entity ID), control action code (cooling_level_down), and execution intensity value (-2℃).
[0070] The above embodiments construct a complete technical chain from raw voice or text input to structured task instruction output. By introducing a rule-based knowledge mapping mechanism, it achieves accurate classification and parameterized expression of typical user intentions in home scenarios. This application fully considers the problems of noise interference, grammatical diversity, and semantic ambiguity in practical applications, and possesses good robustness and generalization ability. In addition, by organically combining the capabilities of artificial intelligence models with domain expert knowledge, it avoids the unexplainability risks brought by pure black box models while ensuring the level of intelligence, making it suitable for home smart terminal environments with high requirements for security and reliability.
[0071] As one implementation of a pre-defined category mapping rule base, the matching logic specifically includes:
[0072] When the instruction object identifier belongs to the media object set and the operation action identifier belongs to the analysis action set, the task type identifier is mapped to the media processing type identifier.
[0073] In the media processing scenario, the system defines specific sets of objects and actions: the media object set includes nouns representing digital content such as "image," "video," "record," and "record," while the analysis action set covers behavioral verbs that semantically refer to content retrieval or intelligent analysis, such as "search," "play," "categorize," "filter," and "recognize." When elements from both sets appear simultaneously in a single instruction, the system determines that it belongs to a content-level operation requirement on local storage resources.
[0074] For example, in the instruction "Find all videos featuring cats taken last week," "videos" are identified as media object identifiers, and "find," after semantic normalization, corresponds to the "search" action, falling into the analysis action set. Therefore, the system maps it to the "media processing type identifier." Behind this mapping relationship lies a deep modeling of typical home user behavior. People often want to quickly locate specific segments from massive amounts of personal footage, but traditional file systems cannot meet such visual content-based search needs. Therefore, this rule triggers not only simple file reading but also serves as a prerequisite for activating the AI NAS module to perform a series of intelligent processing steps, including image recognition, tag generation, and timestamp comparison.
[0075] When the instruction object identifier belongs to the home appliance set and the operation action identifier belongs to the device control action set, the task type identifier is mapped to the home control type identifier.
[0076] When faced with the need to control devices in the physical world, the system activates a second mapping path: when the instruction object identifier falls into a set of home appliances (such as "lights," "air conditioners," "curtains," "sockets," "door locks," etc.), and the operation action identifier belongs to a set of device control actions (such as "open," "close," "adjust," "raise," "start," etc.), the system automatically sets the task type identifier to "home control type identifier." The key to this judgment mechanism lies in distinguishing the semantic boundary between "controlling real devices" and "accessing virtual services." For example, in "raise the living room air conditioner temperature," "air conditioner" is a typical home appliance object, while "raise" belongs to the intensity adjustment action; the combination of the two clearly points to a request for a change in the state of a physical device.
[0077] It's important to note that this rule doesn't rely solely on vocabulary matching; it also incorporates contextual disambiguation. For example, while "television" in "turn on a TV program" is literally similar, it's not misinterpreted as device on / off control because it's paired with the verb "program" rather than directly applied to the device itself. This dual verification mechanism, based on co-occurrence probability and grammatical structure, effectively improves the accuracy of intent recognition and avoids the risk of erroneous operations.
[0078] When the instruction object identifier is a display terminal and the operation action identifier is video path switching, the task type identifier is mapped to the video path switching type identifier.
[0079] The display terminal includes physical display devices or virtual display interfaces. Its identifier generation logic is as follows: if the user instruction contains the keywords {"television", "screen", "projection"}, it is marked as the display terminal identifier.
[0080] Specifically, when the instruction object identifier is recognized as a display terminal, and the operation action identifier is "switch" or its synonym (such as "switch to", "go to", "switch back"), the system does not treat it as ordinary device control, but triggers a separate task type specifically for audio and video path management—"video path switching type identifier". Here, "display terminal" is not limited to the television itself, but refers to all terminal interfaces that may carry image output, including physical devices (such as televisions and projectors) or virtual interfaces (such as HDMI IN / OUT ports).
[0081] The system uses semantic rules to determine whether an object is automatically assigned the "display terminal" attribute if the original command contains keywords such as "television," "screen," "large screen," or "projection." For example, in the command "Switch the picture to IPTV," although "picture" does not directly mention the display device, the system can infer that the actual target is still the television screen by using referential resolution techniques, based on the path shift action of "switch" and the context. In this case, although "switching" might be classified as a control action in a general context, since its target is the signal source path rather than the device's power state, it should not be mapped to a typical home control type but should instead enter a dedicated video routing channel.
[0082] Once the task type is determined to be "video path switching," the system initiates a low-level hardware coordination mechanism. This type identifier triggers a specific control command generation process, ultimately sending a level signal to the HDMI switching chip via the General Purpose Input / Output (GPIO) interface. Specifically, when the user's instruction is to watch IPTV programs, the system outputs a 3.3V high-level signal, enabling the HDMI switching chip to conduct the video output channel decoded from the IPTV set-top board; conversely, when the instruction requests to browse content generated by the local AI computing power motherboard (such as camera recordings or family photos from a NAS), a 0V low-level signal is output, switching to the local media output channel. This process occurs within milliseconds, ensuring seamless switching between different content sources without requiring manual changes to the TV input source.
[0083] More importantly, this mechanism achieves direct binding between software layer instructions and hardware layer pathways, breaking the fragmented experience of "remote control of TV and APP control of box" in traditional home entertainment systems, and truly achieving the intelligent interactive goal of "switching content sources with one sentence".
[0084] In the above implementation, the category mapping rule base is not only a static lookup table, but also a dynamic reasoning system that integrates linguistic rules, user behavior habits, and hardware architecture characteristics. Through refined classification of "object-action" combinations, it achieves accurate identification and differentiated responses to three core home tasks (media content management, smart device control, and audio / video path scheduling). In particular, for the often overlooked but frequently used interaction scenario of "display terminal path switching," a dedicated task type identifier is established, demonstrating a profound understanding of user experience details.
[0085] This application fundamentally solves the problem of misjudgment caused by semantic ambiguity in existing smart home systems, significantly improving the accuracy of command parsing and the rationality of system response, and providing solid technical support for building a unified home AI hub. Through this mechanism, users no longer need to understand complex system architectures or technical terms; they only need to express their intentions in everyday language, and the system can automatically complete the entire closed-loop control from semantic understanding to hardware execution, truly realizing the vision of a smart life where "what you think is what you get."
[0086] As one implementation of step S103, the step of scheduling the corresponding execution unit to execute the operation instruction set according to the target system identifier and generating operation result data corresponding to the task type identifier includes:
[0087] When the target system identifier is a media processing system, the intelligent network attached storage media processor is driven to perform media file analysis operations and generate structured media data.
[0088] Specifically, when the target system identifier is determined to be a "media processing system," the system activates its most powerful local AI submodule—the Intelligent Network Attached Storage Media Processor (AI NAS Media Processor)—to perform deep media content analysis. Here, "media processing" refers not only to traditional file reading or playback control but also emphasizes semantic-level understanding and structured extraction of the audio and video content itself. For example, when a user issues the command "find all video clips featuring cats taken last week," the system has already determined through prior semantic analysis that it belongs to a media retrieval task and triggers the invocation of the AINAS module. This module incorporates lightweight computer vision models (such as MobileNetV3+TimeSformer) that can directly perform frame-level sampling and feature extraction on the video stream from the local SSD at the edge, identify whether the image contains cats, and combine this with metadata such as timestamps and geographic location tags to form a semantically annotated result set.
[0089] The entire processing is completed locally, avoiding the risks of uploading private images to the cloud and meeting security and compliance requirements in a home setting. The resulting "structured media data" is not a copy of the original files, but a standardized data object containing a list of file paths, start and end times of matched segments, confidence scores, and image thumbnail links. This facilitates subsequent organization into natural language descriptions by AI models or visualization on a television interface. This intelligent retrieval capability, based on content rather than filenames or directories, significantly improves the efficiency and experience for users accessing their private digital assets.
[0090] When the target system identifier is a home control system, the driver home assistant connector sends device control commands to the target home appliances and generates device status response data.
[0091] When the target system identifier points to "Home Control System," the system's response mechanism shifts to real-time control and status awareness of external IoT devices. At this point, the system invokes the "Home Assistant Connector" as a bridge to the smart home ecosystem. This connector is essentially a protocol converter and security proxy middleware, supporting the API interfaces of mainstream home automation platforms such as Home Assistant, OpenHAB, or Mi Home. It receives structured control requests from the MCP Server (e.g., {"device": "living_room_light", "action": "turn_on", "brightness":75%}), translates them into the communication format required by the specific platform (e.g., MQTT topics / payloads, RESTful POST requests), and sends them to the target home appliance via LAN or Bluetooth Low Energy (BLE) channel.
[0092] It's important to note that this process involves more than just one-way command issuance; it establishes a two-way feedback mechanism. After a control command is issued, the home assistant connector actively polls or listens for status confirmation messages returned by the device, such as whether the lights have been successfully turned on or whether the air conditioner's current temperature has been adjusted correctly. This information is encapsulated as "device status response data," typically presented in JSON format, including fields such as execution result (success / failure), actual status value, and response latency. If an operation fails (e.g., a socket goes offline), the system can also record an error log and trigger an alarm mechanism, providing a basis for subsequent AI decisions. This closed-loop control mode is significantly superior to the traditional open-loop voice assistant approach of simply executing commands without verification, truly achieving "controllable, observable, and traceable" smart home management.
[0093] When the target system identifier is a video path switching system, a level signal is output to the target display terminal through the general input / output interface to perform video source switching and generate a switching completion status code.
[0094] When the system identifies the task type as "video path switching system," this type of task typically occurs in scenarios where users want to quickly switch between different video sources, such as switching from watching a live IPTV program to viewing footage recorded by a front-door camera. Although it appears to be simply "changing an input source," the traditional method requires users to manually switch the TV's HDMI input port using a remote control, which is cumbersome and lacks contextual understanding. This solution, however, uses AI to automatically determine the user's true intent and directly manipulates the HDMI switching chip at the hardware layer to complete the path selection.
[0095] Specifically, the system outputs a specific level signal (e.g., a high level of 3.3V represents selecting the IPTV channel, and a low level of 0V represents selecting the local AI motherboard output channel) through the general purpose input / output (GPIO) interface of the main control chip. This signal is sent to the control pin of the HDMI multiplexer chip, thereby physically connecting the corresponding video path and achieving millisecond-level seamless switching. The entire process requires no user intervention and does not rely on an external remote control or mobile app.
[0096] More importantly, after the switch is complete, the system checks the EDID handshake status of the HDMI link or confirms signal connectivity through internal feedback circuitry, and generates a "switch complete status code" as proof of successful execution. This status code is then sent back to the MCP Server and the AI big data model, enabling them to confirm task completion and continue with subsequent interactive processes, such as displaying the camera feed on the screen while announcing: "You have been switched to the front door monitoring view." This ability to precisely map high-level semantic commands to underlying hardware actions is the key difference between this application and purely software solutions.
[0097] In the above implementation, media processing focuses on local AI inference and privacy protection, home control emphasizes protocol compatibility and status feedback, and video switching focuses on direct hardware control and low-latency response. All three rely on unified task identifiers and parameter specifications, forming an organic whole under the coordination of the MCP Server. This solution not only solves the problems of fragmented functions and complex operation in existing home smart systems, but also enables the home AI hub to truly possess autonomous decision-making and collaborative execution capabilities.
[0098] Reference Figure 3 As one implementation of step S104, the step of performing state analysis on the operation result data and generating decision instruction data containing decision type identifiers includes:
[0099] Step S301: Receive operation result data returned by the execution unit, including structured media data, device status response data, or switchover completion status code;
[0100] When the task involves media content retrieval or analysis, it returns structured media data, such as a list of video clips with timestamps, file paths, identification tags (such as "cat", "face", "night"), and confidence scores. When the task is smart home control, it returns device status response data, which is usually encapsulated in JSON format to show the actual operating status of the current device (such as whether the lights are on, the air conditioner is set to temperature, and the door lock is closed). When the task is audio / video path switching, it returns a lightweight but critical switching completion status code to indicate whether the HDMI signal path has been successfully established.
[0101] Step S302: Extract the status feature parameters from the operation result data, including data integrity indicators, execution timeliness indicators, and result validity indicators;
[0102] Among these metrics, data integrity measures whether the system has fully completed its intended actions during execution. For media processing tasks, this is reflected in the ratio of the number of effectively identified target files to the total number of files to be processed. Execution timeliness focuses on the system's response speed, reflecting the real-time interactive needs of a home environment. Result effectiveness is the most fundamental quality assessment criterion, focusing on whether the execution result truly meets the user's original intent.
[0103] Step S303: Match and analyze the state feature parameters according to the preset decision rule base to generate decision instruction data containing decision type identifiers; wherein, the decision type identifiers include task refinement identifiers, exception handling identifiers, or task completion identifiers.
[0104] Specifically, the matching logic of the decision rule base includes: if the data integrity index is lower than the first threshold or the result validity index is lower than the second threshold, a task refinement identifier is generated; if the execution timeliness index exceeds the third threshold or the result validity index is zero, an exception handling identifier is generated; when the data integrity index, execution timeliness index, and result validity index all reach the corresponding thresholds, a task completion identifier is generated.
[0105] For example, when the system finds that the data integrity is below the first threshold (e.g., <70%) or the result validity is below the second threshold (e.g., <60%), it means that the current results are insufficient to support the termination of the task, and further clarification of the requirements or supplementary processing is needed. At this time, a task refinement label is generated, triggering the AI big model to initiate precise follow-up questions to the user. For example, "I only found two relevant photos, do you mean the child in the red clothes?" This type of follow-up question is not a random question, but is transformed from a structured template automatically generated based on the unmet parameter items.
[0106] On the other hand, if the execution timeliness severely exceeds the limit (e.g., exceeding the third threshold, set at 5 times the average normal response time) or the result validity is zero (the goal is not achieved at all), it is judged as a system-level anomaly, an anomaly handling flag is generated, and the fault recovery process is initiated. At this time, the system no longer relies on user intervention, but automatically calls the historical successful case database, searches for successful execution parameter combinations of similar tasks, and re-initiates the call as an alternative. For example, if a camera fails to recognize data due to excessively high resolution, the system can automatically downgrade to retry using a low-resolution mode that was successfully processed the previous day.
[0107] Only when all three indicators are met will the system generate a task completion marker, signifying the successful completion of this interactive loop and allowing the AI to output a natural language summary to the user.
[0108] In the above implementation, unified modeling and feature extraction are performed on the returned data of different types of tasks. Combined with a threshold-based logical reasoning mechanism, the system can autonomously distinguish between three typical states: "incomplete results requiring supplementation," "execution failure requiring repair," and "task successful and ready for completion," and trigger corresponding subsequent actions accordingly. This mechanism not only improves the system's robustness and fault tolerance but also endows it with human-like reflective and adaptive capabilities, transforming the home AI hub from a mechanical command executor into a true "smart butler." Especially in complex home environments, facing various uncertainties such as device fluctuations, network latency, and semantic ambiguity, this solution ensures the reliability and continuity of service delivery, significantly enhances the trust and fluency of human-computer interaction, and provides solid technical support for building a long-term sustainable smart home ecosystem.
[0109] As one implementation of step S302, extracting state feature parameters from the operation result data includes:
[0110] When the operation result data is structured media data, the ratio of the number of validly identified files to the total number of files is calculated based on the structured media data as a data integrity indicator, the file analysis and processing time is obtained as an execution timeliness indicator, and the percentage of files with an identification accuracy rate greater than the threshold is counted as a result validity indicator.
[0111] When dealing with "structured media data" as the type of output, the system focuses on the quality of the AI NAS module's output after performing image or video content analysis tasks. Structured media data refers to a collection of semantically labeled data generated after processing by a local AI model, such as a JSON-formatted list of results, where each item contains fields such as file path, detected object (e.g., face, pet), timestamp, and confidence score.
[0112] At this point, the data integrity metric reflects whether the entire analysis process covered all objects to be processed. Specifically, the system counts all successfully identified valid files and divides it by the total number of files involved in the original request to obtain a ratio. For example, when a user requests "find all footage of babies in the past week," and the system scans 100 videos with only 85 returning valid tags, the integrity rate is 85%. A rate below a set threshold (e.g., 90%) indicates that some files have not been processed, possibly due to storage access failures, decoding anomalies, or resource preemption. This metric is designed to prevent the system from misjudging the overall task completion due to partial interruptions.
[0113] Meanwhile, the execution timeliness metric measures the time delay between the AI model issuing an analysis request and receiving the first complete result, i.e., "document analysis and processing time." Considering the limited computing power of edge devices, long waiting times can affect user experience, especially in interactive scenarios (such as browsing and searching simultaneously). If the processing time exceeds a reasonable range (such as 30 seconds), even if the final result is accurate, it should be considered a decline in service quality.
[0114] Finally, the validity metric focuses on the accuracy of the analysis itself, quantified by statistically identifying the percentage of documents with a confidence level higher than a preset threshold (e.g., 0.8). For example, if 20 out of 100 labeled results have a confidence level lower than the threshold, the validity is only 80%, indicating that the model may have overfitting or training data bias.
[0115] When the operation result data is device status response data, the device status change identifier in the device status response data is parsed as a data integrity indicator, the delay of sending the instruction to the status response is calculated as an execution timeliness indicator, and the consistency between the physical state of the device and the expected state of the instruction is verified as a result validity indicator.
[0116] When the operation result data is presented as "device status response data," the corresponding scenario is the execution feedback of a smart home control task. This type of data is usually returned by Home Assistant or other IoT platforms in the form of a structured message containing fields such as device ID, current status (on / off, temperature, brightness, etc.), and response time.
[0117] In this context, data integrity metrics no longer focus on the number of files, but rather on whether the system successfully received a clear "device state change identifier." For example, when a user issues the command "turn on the living room light," ideally, a status update message should be received within a short time, showing that the light fixture's state field has changed from off to on. If no response is received, or a null value or error code is returned, it is considered a lack of integrity, meaning there is a problem with the communication link or the device is offline. It is worth noting that this metric emphasizes the observability of "state changes," rather than simply confirmation of command delivery, and therefore better reflects real-world execution.
[0118] The execution timeliness metric is defined as the network round-trip latency from when the MCP Server sends a control command to the connector to when the first valid status feedback is received. In a home LAN environment, a normal response should be completed within 500 milliseconds; if the delay is too long (e.g., more than 3 seconds), even if it is eventually successful, it may cause users to have a negative perception of "no response".
[0119] The validity of the result is the deepest level of verification. It not only checks whether the communication was successful, but also verifies whether the actual physical state is consistent with the expected command. For example, if the system sends a "turn off the air conditioner" command and receives a status receipt, but the data shows that the air conditioner is still running in cooling mode, then the execution is deemed invalid. Such inconsistencies may stem from firmware bugs, device lag, or intermediate gateway caching errors, and can only be discovered through bidirectional comparison. This "command-state consistency verification" mechanism significantly improves the robustness and reliability of the system.
[0120] When the operation result data is a switching completion status code, the path switching success identifier in the switching completion status code is parsed as a result validity indicator, the switching completion status code is checked to see if it contains predefined necessary fields as a data integrity indicator, and the delay of the screen switching completion after the level signal is output to the target display terminal is recorded as an execution timeliness indicator.
[0121] When the operation result data is a "switching complete status code," it describes the execution result of the video source switching task at the hardware level, which is the key difference between this application and a purely software solution. In home AI intelligent computing center products, users often need to quickly switch between IPTV live programs and local AI analysis footage (such as camera monitoring). Traditional methods rely on manually selecting the TV input source, resulting in a fragmented experience. This application completes the path connection by automatically scheduling the HDMI switching chip through AI, and its execution process involves underlying hardware control. Therefore, the result evaluation of such tasks must take into account both hardware and software synergy.
[0122] At this point, the validity indicator directly derives from the "path switching success flag" in the status code. This is a Boolean flag indicating that the GPIO level signal has been correctly output and the HDMI switching chip has completed the internal routing switch. However, this alone is insufficient to prove that the user has actually seen the target image, as there may be "false successes": for example, the level switch may be completed, but the target display terminal may fail to refresh the image synchronously, or the EDID handshake may fail, resulting in a black screen. Therefore, the system further introduces an external visual feedback mechanism (such as through an auxiliary camera or screen capture interface) to verify whether the content currently displayed on the TV matches the path identifier (IPTV / LOCAL). Only when the two match is the validity flag marked as 1; otherwise, even if the hardware reports success, it is still considered invalid. This dual verification mechanism greatly enhances the system's reliability.
[0123] Data integrity metrics focus on the structural compliance of the status code itself, specifically whether it contains predefined necessary fields, such as path identifiers, timestamps (recording the trigger time of the level signal), and the physical address of the target device (MAC or HDMI port number). These fields together constitute a traceable operation log for troubleshooting and auditing. The absence of any key field is considered insufficient integrity.
[0124] The execution timeliness metric records the end-to-end delay between the output of the GPIO level signal from the main control chip and the actual display of the new image on the target display terminal. This delay needs to be controlled within milliseconds (e.g., <200ms) to achieve seamless switching. The system can measure this delay using a high-precision timer combined with the display synchronization signal (VSYNC) and use this value as an important reference for service quality.
[0125] In the above embodiments, the differentiated modeling of three typical tasks—media processing, device control, and video path switching—allows the system to fully respect the technical characteristics and physical constraints of each task while maintaining a unified evaluation framework. Whether it's the probability judgment based on confidence levels, the logical verification of state changes, or the spatiotemporal alignment of hardware levels and screen display, this application demonstrates its rigor in detailed design and its engineering feasibility.
[0126] More importantly, these extracted state feature parameters do not exist in isolation, but serve as the input foundation for the subsequent decision rule base, supporting the AI model to make advanced behavioral decisions such as "continue to ask questions," "activate fault tolerance," or "confirm completion." This forms a complete closed loop from "perception → execution → feedback → re-decision," enabling the home AI hub to truly adapt autonomously to complex environments and significantly improving the fluency, security, and intelligence level of human-computer interaction.
[0127] Reference Figure 4As one implementation of step S105, the steps of performing task refinement processing, abnormal alternative generation, or task completion feedback processing based on the decision type identifier, and converting the processing results into natural language feedback output, include:
[0128] Step S401: Perform classification processing based on the decision type identifier;
[0129] Step S4011: When the decision type identifier is a task refinement identifier, extract the unmet parameter items from the operation result data and generate a follow-up instruction template.
[0130] Specifically, when the decision type identifier is a task refinement identifier, one implementation method for extracting non-compliant parameter items from the operation result data and generating follow-up instruction templates includes:
[0131] If the operation result data is structured media data, locate the media type and time range with insufficient file coverage, map the missing media parameters to the {"file format", "time range", "target tag"} field, and combine the mapped fields to generate a follow-up instruction template;
[0132] If the operation result data is device status response data, then parse the target device identifier of the unresponsive instruction, and map the missing device identifier to the {"device name", "physical location", "device type"} field, and combine the mapped fields to generate a follow-up instruction template;
[0133] If the operation result data is a switch completion status code, mark the hardware interface number with the delay exception, and map the interface exception to the {"interface type", "transmission protocol"} field. Combine the mapping fields to generate a follow-up instruction template.
[0134] For example, when insufficient media type coverage is detected, the system maps it to fields such as "file format," "time range," and "target tag," combining them into specific questions like "Do you mean MP4 format videos?" or "Do you only want to find photos from last Saturday?" For device control failures, it maps to "device name," "physical location," and "device type," forming a precise confirmation like "Are you referring to the main living room light or the reading light?" For interface anomalies, it associates "interface type" with "transmission protocol," which can be used for subsequent engineer debugging or translated into more colloquial language to inform advanced users.
[0135] Understandably, this field mapping mechanism not only improves the professionalism and accuracy of follow-up questions, but also provides a highly structured input foundation for subsequent natural language generation, so that the generated text not only conforms to grammatical norms, but also accurately reflects the system's true intent.
[0136] Step S4012: When the decision type identifier is an exception handling identifier, call the historical case database to match similar successful case parameters and generate alternative solution parameters.
[0137] In this case, when the decision type identifier is an exception handling identifier, it typically occurs when the task completely fails or key metrics return to zero, such as a camera failing to connect, an AI model failing to load, or no signal output during HDMI switching. Directly reporting the error to the user in this situation can easily lead to frustration. Therefore, this solution designs an intelligent matching mechanism based on a historical success case database, enabling the system to learn from experience. The system can retrieve historical success records with the same task attributes from the database based on the current task type identifier (e.g., media processing, home control, video switching).
[0138] For example, if the current task is "failed to retrieve bedroom camera footage," the system will search for all past successful access cases related to the "bedroom camera," focusing on parameters such as resolution settings, encoding / decoding formats, and network bandwidth configurations. As another example, in video switching tasks, if a particular HDMI port experiences frequent delays, the system can trace back to the EDID analog strategy or level drive strength parameters used in previous successful switching of the same signal source. Through this similarity-based parameter migration, the system can quickly generate a set of feasible alternative parameter sets, encapsulate them into a new multi-protocol call request, and re-initiate execution. This mechanism significantly improves the system's robustness and self-healing capabilities, reducing reliance on external intervention.
[0139] In some embodiments, a historical success case database is retrieved based on the current task type identifier. Specifically, this includes: for media processing type identifiers, matching successful cases with the same file format; for home control type identifiers, matching successful cases with the same device type; and for video path switching type identifiers, matching successful cases with the same interface number. Subsequently, the operation parameters from the successful cases are extracted as alternative solution parameters.
[0140] Step S4013: When the decision type identifier is a task completion identifier, construct the corresponding task completion feedback template;
[0141] Specifically, the task completion feedback template framework can be selected according to the data type of the operation result. Specifically, for structured media data, the corresponding task completion feedback template is filled with {"[number] files matching [tag] were found in [time range]"}; for device status response data, the corresponding task completion feedback template is filled with {"[Device Name] has been successfully controlled to execute [action]"}; for switch completion status codes, the corresponding task completion feedback template is filled with {"[Signal Source] screen has been switched to [Device Name]"}.
[0142] Step S402: Input the generated follow-up question template, alternative solution parameters, or task completion feedback template into the natural language generation model and output interactive text for the user.
[0143] The follow-up instruction templates, alternative parameters, or task completion feedback templates will all be uniformly input into a Natural Language Generation Model (NLG Model) for semantic rendering. This model can be a lightweight version of T5, BART, or a locally deployed, finely tuned version of LLM, and its function is to transform structured data into fluent, natural, and conversational Chinese sentences.
[0144] For example, a follow-up question template containing "{'missing_field':'time_range','suggestion': 'lastweekend'}" might output as "The 'photos of children playing' you mentioned, do you mean the two days from last Saturday to Sunday?" after being processed by the NLG model. This generation method not only preserves the accuracy of the original logic, but also gives the system a human-like communication style, greatly enhancing the realism and friendliness of the user experience.
[0145] In the above implementation, differentiated processing strategies are applied to different decision types. The system can accurately follow up when tasks are incomplete, autonomously repair itself when execution fails, and proactively report when tasks are successfully completed, truly achieving a leap from "mechanical response" to "intelligent dialogue." Especially in complex systems integrating multiple functions, such as home AI intelligent computing centers, this method effectively bridges the semantic gap between the underlying hardware status and the high-level human-computer interaction. This user-centric closed-loop interaction mechanism significantly improves the reliability, adaptability, and emotional warmth of intelligent services, providing solid technical support for building a long-term, trustworthy human-computer coexistence environment.
[0146] Reference Figure 5 As a further implementation of the intelligent control method, after the steps of performing task refinement processing based on decision type identification, generating abnormal alternatives, or processing task completion feedback, and converting the processing results into natural language feedback output, the method further includes:
[0147] Step S501: Obtain the user's response data to natural language feedback, and parse the parameter supplementation instructions and / or execution confirmation flags in the response data;
[0148] Specifically, in the "acquiring user response data to natural language feedback" stage, the system must deal with typical multimodal input scenarios in a home environment. Users might say "only check last Saturday's recordings" via voice remote control, click "confirm retry" in the mobile app, or select "yes" using the infrared remote control buttons. These heterogeneous input paths are captured by different hardware modules: voice signals are transmitted to the AI computing chip via the Bluetooth channel of the WIFI / BT module, APP commands are transmitted to the MCP Server via the external gigabit network port using the HTTPS protocol, and infrared buttons are decoded into key value codes (such as KEY_OK) by a dedicated infrared receiver chip on the set-top box motherboard.
[0149] Despite their diverse origins, all inputs are uniformly directed to the multimodal fusion processing engine within the AI computing chip. This engine combines Automatic Speech Recognition (ASR), Natural Language Understanding (NLU), and a protocol parser to transform the raw input into a structured semantic representation. For example, the phrase "turn that light back on" will be parsed into a JSON object containing a contextual reference (referring_expression="that light") and the action intent (action="turn_on"). This cross-modal normalization design ensures that the system accurately captures the user's true intent regardless of the interaction method used, avoiding misunderstandings caused by differences in input format.
[0150] Subsequently, the system parses parameter supplementation instructions and / or execution confirmation flags from the response data. Parameter supplementation instructions specifically refer to the specific values or constraints provided by the user after the system asks follow-up questions, such as time range, spatial location, target tags, etc. The extraction of these parameters relies on a pre-defined domain dictionary and syntactic pattern matching mechanism. For example, when the system previously asked, "What time period of photos do you want to find?" and the user answered "last summer," the NLU component calls the time inference module to standardize it into a specific time interval and verify whether it matches the parameter type required for the media query task. Similarly, if the user says, "Use the bedroom camera to check," the system maps "bedroom" to the device physical location field registered in Home Assistant, completing the logical binding.
[0151] On the other hand, execution confirmation markers are semantic labels used by users to express "agreement to execute" or "immediate action," such as "okay," "try it," or "switch immediately." Although these phrases do not contain substantial parameters, they imply a strong operational intent and need to be recognized as switching signals that trigger hardware-level rescheduling. To this end, the system uses a lightweight classification model (such as TinyBERT) to score the sentiment polarity and urgency of the statements. Once a high-confidence affirmative instruction is determined, the priority escalation mechanism is activated.
[0152] Step S502: Extract supplementary parameter values and update the operation instruction set according to the parameter supplementation instruction, and / or trigger the multi-protocol connection service processing core to reschedule the execution unit according to the execution confirmation flag;
[0153] For example, when a user adds "analyze the surveillance video of the study," the system first queries the Home Assistant device registry to confirm the existence of a device named "study camera" and its online status. If it does not exist or is offline, the system refuses to merge the parameters and provides the user with the reason. Only parameters that pass the hardware compatibility check are written into the corresponding fields of the original operation instruction set. For media processing tasks, this manifests as updating the WHERE clause in the AI NAS query request (e.g., adding location='study' AND timestamp BETWEEN'2024-04-01' AND '2024-04-05'); for home control tasks, it manifests as precisely binding the vague instruction "turn on the light" to the specific light fixture with device_id=0xEF23. At the same time, the system also generates an update log with a timestamp, persistently stores it in eMMC flash memory, and associates it with the original request through a unique task ID, forming a traceable operation chain. This fine-grained logging not only supports fault auditing but also provides high-quality samples for subsequent rule training.
[0154] In parallel, the execution unit is rescheduled based on the execution confirmation flag, and a forced priority flag is generated based on the execution confirmation flag. Call parameters with priority flags are sent to the multi-protocol connection service processing core through the internal communication interface. The current low-priority task thread is interrupted, and the operation instruction set with priority flags is executed first. A system state snapshot at the time of task interruption is recorded for execution recovery.
[0155] Specifically, upon detecting a clear execution confirmation signal, the AI computing chip immediately outputs a high-level pulse signal via its GPIO pin. This signal is transmitted to the MCP Server's scheduler via the internal serial port and embedded as a "real-time priority" marker in the metadata of the task to be executed. Upon receiving this marker, the MCP Server's built-in task scheduler initiates a preemptive scheduling strategy: pausing currently executing low-priority background tasks (such as file index reconstruction or firmware download), releasing CPU cores and memory bandwidth resources, and ensuring that new tasks receive immediate responses. To ensure the integrity of interrupted tasks, the system uses a reserved area in LPDDR memory to create a state snapshot, saving complete context information including the program counter, general-purpose registers, and stack pointer. After a high-priority task completes, the original task can seamlessly resume execution from the point of interruption, avoiding data loss or redundant calculations. For example, when a user confirms "Switch TV channel immediately," even if the AI NAS is in the process of batch image classification, the system can still complete the control signal output of the HDMI switching chip within 200 milliseconds, meeting the hard real-time requirements of TV screen switching.
[0156] Step S503: Input the updated operation instruction set and / or rescheduled operation result data into the historical case database for storage;
[0157] All updated operation instruction sets and execution result data generated by rescheduling are further encapsulated into standardized case packages and written to a historical case database for long-term storage. This database is deployed on the AI NAS's solid-state drives and organizes data using a categorized structure: media tasks generate records containing {task ID, file type, time range, recognition accuracy, processing time}; home control tasks record fields such as {task ID, device ID, control signal, response latency, execution status}.
[0158] More importantly, these cases are not isolated but synchronized across motherboards via an internal gigabit network port. The set-top box motherboard can access cached copies for displaying "recent operation history" or "recommended frequently used settings" on the TV interface. Furthermore, the system has constructed a B+ tree-based composite index mechanism, using "task type + execution time" as the primary key, significantly improving the parameter retrieval speed for subsequent similar tasks. This structured approach to experience accumulation ensures that every success or failure becomes nourishment for system evolution.
[0159] Step S504: Generate task execution efficiency optimization parameters based on the historical case database, and update the matching threshold of the decision rule base.
[0160] The specific steps for generating task execution efficiency optimization parameters include: calculating the average execution timeliness index of similar tasks in the historical case database; statistically analyzing the distribution variance of the result validity index in the operation result data; reducing the matching threshold of the execution timeliness index in the decision rule base when the average execution timeliness index exceeds the preset baseline and the distribution variance is lower than the threshold; and increasing the weight coefficient of the data integrity index when the distribution variance of the result validity index exceeds the threshold three times consecutively.
[0161] For example, the system calculates the mean time to response (MTTR) for similar media processing tasks. If it finds that the actual time taken for ten consecutive tasks is less than 15% of the historical average, it infers that the local computing load has been reduced or the model inference efficiency has improved, thus automatically lowering the alarm threshold for the "execution timeliness index" (e.g., from 1.5 seconds to 1.3 seconds). Similarly, for home control tasks, if statistics show that the control failure variance of a certain type of device (such as an old Zigbee bulb) is consistently higher than 0.25, the system will increase the data integrity weight of this type of task and force verification of the device's online status and network signal strength before future execution. These optimization parameters are injected into the set-top box motherboard's configuration manager via an internal serial port in a hot-loading manner, dynamically replacing the threshold table of the decision rule base stored in the eMMC, and taking effect without a restart.
[0162] Simultaneously, the system broadcasts the new rule version number to all execution units (AI NAS, Home Assistant connector) to ensure that the behavior of each subsystem is synchronized and consistent. This data-driven self-optimization mechanism enables the system to maintain optimal performance under real-world conditions such as device aging, network fluctuations, and environmental changes.
[0163] The above implementation overcomes the limitations of static configuration and fixed logic in traditional smart home systems, establishing a dynamic adaptive system that deeply integrates hardware control, software scheduling, and machine learning. Through precise analysis of user feedback, the system achieves context-aware updates of task parameters and real-time preemptive scheduling of key operations. Leveraging structured case storage and statistical modeling, the system possesses the ability to extract patterns from historical experience and feed them back into decision-making logic. This closed-loop mechanism of "perception-execution-feedback-evolution" not only significantly improves task success rate and response speed but also endows the home AI hub with true autonomous growth characteristics, enabling it to operate stably for a long time in complex and ever-changing home environments, providing a solid technical foundation for human-machine collaboration.
[0164] This application also discloses an intelligent control system based on multi-protocol task scheduling.
[0165] An intelligent control system based on multi-protocol task scheduling specifically includes:
[0166] The instruction generation module is used to receive raw instruction data input by the user, perform semantic parsing, extract instruction keywords, and generate a structured instruction data package containing task type identifiers;
[0167] The multi-protocol connection service call module is used to convert structured instruction data packets into multi-protocol connection service call parameters and send them to the multi-protocol connection service processing core through the internal communication interface; the multi-protocol connection service call parameters include the target system identifier and the operation instruction set;
[0168] The scheduling and execution module is used by the multi-protocol connection service processing core to schedule the corresponding execution unit to execute the operation instruction set according to the target system identifier, and generate operation result data corresponding to the task type identifier;
[0169] The result status analysis module is used to perform status analysis on the operation result data and generate decision instruction data containing decision type identifiers.
[0170] The feedback processing module is used to perform task refinement processing, generate abnormal alternatives, or process task completion feedback based on the decision type identifier, and convert the processing results into natural language feedback output.
[0171] An intelligent control system based on multi-protocol task scheduling according to an embodiment of this application can implement any of the above methods, and the specific working process of each module in the system can refer to the corresponding process in the above method embodiments.
[0172] In one embodiment of this application, the hardware configuration of the intelligent control system during actual deployment includes: a casing, an AI computing motherboard, and a set-top box motherboard. The casing measures 130 mm (length) × 130 mm (width) × 40 mm (height), facilitating placement in a TV cabinet or other areas with a high concentration of home electronic devices.
[0173] The AI computing power motherboard integrates an AI computing chip, LPDDR memory, eMMC flash memory, a WIFI / BT module, an external WIFI antenna, a built-in solid-state drive, an internal gigabit Ethernet port, an internal serial port, an internal HDMI interface, an external gigabit Ethernet port, two external USB Type-A ports, and an external USB Type-C port. The AI computing chip boasts a computing power of over 16 TOPS, supporting efficient operation of lightweight AI models. It features 16GB of LPDDR memory and 64GB of eMMC storage. The WIFI / BT module supports the WIFI 6 communication protocol, is compatible with both 2.4GHz and 5GHz dual-band, and supports Bluetooth BT 5.4 protocol, ensuring stable wireless connectivity and low latency. The external gigabit Ethernet port and two USB Type-A ports all support the USB 3.0 communication protocol, allowing for the connection of external AI computing devices to enhance the overall system processing power. The USB Type-A ports also support external audio devices, and the device is compatible with wireless audio output devices such as Bluetooth speakers, enabling multi-mode audio and video playback. The Type-C interface is used to connect the power adapter, which uses 5V DC power and supports power supply from mainstream mobile phone chargers, making it convenient for users to use existing accessories. The set-top box motherboard includes a video decoding chip that supports broadcast industry standards, an infrared receiver chip, an HDMI switching chip, an internal HDMI interface, an internal gigabit Ethernet port, an internal serial port, and an external HDMI interface, which can complete the functions of receiving, decoding, and outputting IPTV signals.
[0174] During operation, the WIFI / BT antenna is responsible for receiving IPTV video data from the network, Bluetooth control signals sent by the smart remote control, and video stream data uploaded by the camera module. At the same time, it enables two-way data interaction between this device and external IoT devices such as cameras and smart appliances. When a user issues a natural language command (such as "Play the picture of the baby laughing that I took last night") via a TV remote control that supports voice recognition, the voice information is transmitted to the AI computing chip via Bluetooth through the WIFI / BT module. The AI computing chip calls the locally deployed lightweight natural language understanding model (NLU) to perform semantic parsing on the original command, performing word segmentation, part-of-speech tagging, named entity recognition, and dependency parsing, stripping away redundant expressions in spoken language, accurately extracting core actions (such as "play" and "search") and object descriptions (such as "baby laughing" and "living room camera recording"), and combining them with a preset action-object knowledge graph to infer the task intent and generate a structured command data packet containing a task type identifier; if it is determined to be a media file processing task, it is assigned the identifier TASK_TYPE=MEDIA_SEARCH, and if it is a device control task, it is marked as TASK_TYPE=DEVICE_CONTROL.
[0175] Subsequently, the system converts the structured instruction data packet into multi-protocol connection service call parameters, which encapsulate the target system identifier (TARGET_SYSTEM_ID) and the operation instruction set (ACTION_SET). These call parameters are sent to the multi-protocol connection service processing core via the device's internal high-speed serial communication interface (such as SPI or UART). The processing core dynamically schedules the corresponding execution unit based on the target system identifier: when the target is an AI NAS media processor, the processing core activates the intelligent analysis engine on the local solid-state drive, driving it to perform content recognition and tag-based retrieval of stored audio and video files based on a lightweight CV model; when the target is a home assistant service, it sends control requests to the connected HomeAssistant platform via MQTT or RESTful API to operate smart devices such as lights and air conditioners. During this process, the real-time video stream transmitted from the camera is also sent to the AI computing chip by the WIFI / BT module. After decoding and behavior analysis, the result data can be output to the set-top box motherboard for display via the internal HDMI interface.
[0176] IPTV video data is transmitted directly to the set-top box's motherboard video decoding chip via an internal gigabit Ethernet port for decoding and then projected onto the TV screen for viewing. Meanwhile, if the AI computing chip's semantic analysis of voice commands involves TV operation control (such as switching input sources or adjusting volume), the control commands are transmitted to the set-top box motherboard via an internal serial port, achieving compatibility and coordination with traditional remote control logic. For ordinary button-type remote controls that do not support voice functionality, they continue to interact with the infrared receiver chip on the set-top box motherboard via infrared communication to complete basic functions such as power on / off and channel selection, ensuring seamless coexistence of the old and new control methods.
[0177] The built-in solid-state drive is used for local storage of multimedia data such as pictures, audio and video files. The AI computing chip can intelligently filter, classify, and tag this data based on image content analysis technology. For example, it can automatically identify people, objects, scenes, and emotional characteristics in the picture and establish a searchable time-content index, thus possessing complete AI NAS functions. All data is stored locally throughout the process and is not uploaded to the cloud, effectively protecting user privacy and security. The HDMI switching chip is located on the set-top box motherboard and is used to dynamically switch between the video signal after IPTV decoding and the video signal output by the camera or other AI-processed signals from the AI computing motherboard. The finally selected picture is projected to the TV screen for display through the external HDMI interface. The selection of this HDMI path is directly controlled by the AI computing chip through GPIO pins, and the switching logic is uniformly coordinated by the multi-protocol connection service processing core: it can automatically switch the video path according to the user's selection command on the mobile APP or TV screen operation interface, and it can also respond to the power button operation of ordinary remote control, defaulting to TV mode or security preview mode upon power-on, realizing one-click wake-up of the desired scene.
[0178] Upon completion, each execution unit returns a structured response with status codes and contextual information: the AI NAS media processor returns a list of matching files and their confidence scores; the Home Assistant connector provides feedback on actual device status changes; and the HDMI switching chip reports the status code indicating whether the channel switching was successful. The multi-protocol connection service processing core analyzes the received result data to determine the execution effectiveness. If the result is "SUCCESS," a task completion flag is generated; if it is "NO_MATCH" or "DEVICE_OFFLINE," a task refinement process is triggered, with the AI model constructing follow-up questions to guide the user to provide more precise conditions; if there is a communication anomaly or logical error, an alternative solution is generated, such as recommending that lamps in the same area can replace offline devices. Based on the decision type flag obtained from the analysis, the system determines the subsequent course of action: enter iterative search, propose alternative solutions, or terminate the task.
[0179] Finally, once the task is confirmed, the AI model converts the result into feedback in natural language, which is then output to the user via voice announcement using the speech synthesis module, or presented as on-screen text prompts. For example, "We have found 3 videos containing cats for you, the longest of which is recorded by the living room camera yesterday at 18:27," enhancing the transparency and user-friendliness of human-computer interaction.
[0180] This application boasts several technological advantages: users can remotely view and manage images and audio / video data stored on the solid-state drive via a TV screen or mobile app, while simultaneously achieving unified control of cameras and other smart appliances; it supports voice control, allowing users to freely switch between various usage modes such as watching TV programs, browsing security monitoring footage, and controlling smart home devices; the device features built-in AI NAS functionality, combined with a large-capacity solid-state drive design that supports on-demand expansion via an M.2 interface, meeting the needs of family digital asset management. Furthermore, this application possesses strong human-computer interaction capabilities, supporting speech recognition and speech synthesis technologies, enabling intelligent services such as natural language dialogue, weather inquiries, and flight and train ticket bookings.
[0181] The entire system, through deep collaboration between the AI computing chip and the set-top box motherboard, constructs a closed-loop intelligent control system with multi-protocol task scheduling at its core. It achieves localized operation of the entire process from command input, semantic parsing, cross-system calls, execution feedback to autonomous decision-making, truly realizing the organic integration of broadcast television reception function and artificial intelligence computing capabilities. Under the premise of ensuring high-performance computing and data privacy and security, it provides a safe, convenient, and intelligent comprehensive home service platform, representing the development direction of the next generation of home smart terminals.
[0182] In the several embodiments provided in this application, it should be understood that the provided methods and systems can be implemented in other ways. For example, the system embodiments described above are merely illustrative; for example, the division of a certain module is merely a logical functional division, and in actual implementation there may be other division methods, such as multiple modules can be combined or integrated into another system, or some features can be ignored or not executed.
[0183] This application also discloses a computer-readable storage medium.
[0184] A computer-readable storage medium storing a computer program that can be loaded by a processor and executed as described above in any of the intelligent control methods based on multi-protocol task scheduling.
[0185] The computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in connection with an instruction execution system, apparatus, or device; the program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to wireless, wire, optical fiber, RF, etc., or any suitable combination thereof.
[0186] In this application, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this application, "multiple" means two or more, unless otherwise explicitly specified.
[0187] Although this application has been described herein in conjunction with various embodiments, those skilled in the art, by reviewing the accompanying drawings, disclosure, and appended claims, will understand and implement other variations of the disclosed embodiments in carrying out the claimed application. In the claims, the word "comprising" does not exclude other components or steps, and "a" or "an" does not exclude a plurality. A single processor or other unit can implement several functions listed in the claims. While different dependent claims may recite certain measures, this does not mean that these measures cannot be combined to produce a good effect.
[0188] The above are all preferred embodiments of this application and are not intended to limit the scope of protection of this application. Any feature disclosed in this specification (including the abstract and drawings) may be replaced by other equivalent or similar features unless specifically stated otherwise. That is, unless specifically stated otherwise, each feature is only one example of a series of equivalent or similar features.
Claims
1. An intelligent control method based on multi-protocol task scheduling, characterized in that, The intelligent control method includes: Receive raw instruction data input by the user, perform semantic parsing, extract instruction keywords, and generate a structured instruction data packet containing task type identifiers; The structured instruction data packet is converted into multi-protocol connection service call parameters and sent to the multi-protocol connection service processing core through the internal communication interface; wherein, the multi-protocol connection service call parameters include a target system identifier and an operation instruction set, the operation instruction set including the device address, control action code and security token for home control, or the time range and tag conditions for media processing; The multi-protocol connection service processing core schedules the corresponding execution unit to execute the operation instruction set according to the target system identifier, and generates operation result data corresponding to the task type identifier; The operation result data is analyzed to generate decision instruction data containing decision type identifiers; Based on the decision type identifier, the task is refined, anomaly alternatives are generated, or task completion feedback is processed, and the processing results are converted into natural language feedback output. The steps of scheduling the corresponding execution unit to execute the operation instruction set according to the target system identifier and generating operation result data corresponding to the task type identifier include: When the target system identifier is a media processing system, the intelligent network attached storage media processor is driven to perform media file analysis operations and generate structured media data. When the target system identifier is a home control system, the driver home assistant connector sends device control commands to the target home appliances and generates device status response data. When the target system identifier is a video path switching system, a level signal is output to the target display terminal through the general input / output interface to perform video source switching and generate a switching completion status code. The steps of performing task refinement processing, abnormal alternative generation, or task completion feedback processing based on the decision type identifier, and converting the processing results into natural language feedback output, include: Perform the following classification process based on the decision type identifier: When the decision type identifier is a task refinement identifier, extract the unmet parameter items from the operation result data and generate a follow-up instruction template; When the decision type identifier is an exception handling identifier, the historical case database is called to match similar successful case parameters and generate alternative solution parameters. When the decision type identifier is a task completion identifier, a corresponding task completion feedback template is constructed; Input the generated follow-up question template, alternative solution parameters, or task completion feedback template into the natural language generation model, and output interactive text for users.
2. The intelligent control method based on multi-protocol task scheduling according to claim 1, characterized in that, The steps of receiving raw instruction data input by the user, performing semantic parsing, extracting instruction keywords, and generating a structured instruction data packet containing task type identifiers include: Receive raw instruction data from the user, including voice signals or text strings; The original instruction data is preprocessed, including acoustic feature extraction of the speech signal or word segmentation of the text string to generate a word segmentation sequence; The word segmentation sequence is semantically analyzed using a pre-configured natural language processing model to extract noun keywords as instruction object identifiers and verb keywords as operation action identifiers. Based on a pre-defined category mapping rule base, the instruction object identifier and operation action identifier are combined and matched to the task type identifier to generate a structured instruction data packet.
3. The intelligent control method based on multi-protocol task scheduling according to claim 2, characterized in that, The matching logic of the preset category mapping rule base includes: When the instruction object identifier belongs to the media object set and the operation action identifier belongs to the analysis action set, the task type identifier is mapped to the media processing type identifier; When the instruction object identifier belongs to the home appliance set and the operation action identifier belongs to the device control action set, the task type identifier is mapped to the home control type identifier; When the instruction object identifier is a display terminal and the operation action identifier is video path switching, the task type identifier is mapped to the video path switching type identifier.
4. The intelligent control method based on multi-protocol task scheduling according to claim 3, characterized in that, The steps of performing state analysis on the operation result data and generating decision instruction data containing decision type identifiers include: Receive operation result data returned by the execution unit, including structured media data, device status response data, or switchover completion status code; Extract the status feature parameters from the operation result data, including data integrity indicators, execution timeliness indicators, and result validity indicators; The state feature parameters are matched and analyzed according to a preset decision rule base to generate decision instruction data containing decision type identifiers; wherein, the decision type identifiers include task refinement identifiers, exception handling identifiers, or task completion identifiers.
5. The intelligent control method based on multi-protocol task scheduling according to claim 4, characterized in that, Extracting state feature parameters from the operation result data includes: When the operation result data is structured media data, the ratio of the number of validly identified files to the total number of files is calculated based on the structured media data as a data integrity indicator, the file analysis and processing time is obtained as an execution timeliness indicator, and the proportion of files with an identification accuracy rate greater than the threshold is counted as a result validity indicator. When the operation result data is device status response data, the device status change identifier in the device status response data is parsed as a data integrity indicator, the delay of sending the instruction to the status response is calculated as an execution timeliness indicator, and the consistency between the physical state of the device and the expected state of the instruction is verified as a result validity indicator. When the operation result data is a switching completion status code, the path switching success identifier in the switching completion status code is parsed as a result validity indicator, the switching completion status code is checked to see if it contains a predefined field as a data integrity indicator, and the delay of the screen switching completion after the level signal is output to the target display terminal is recorded as an execution timeliness indicator.
6. The intelligent control method based on multi-protocol task scheduling according to claim 4, characterized in that, After the steps of performing task refinement processing, abnormal alternative generation, or task completion feedback processing based on the decision type identifier, and converting the processing results into natural language feedback output, the method further includes: Acquire user response data to natural language feedback, and parse the parameter supplementation instructions and / or execution confirmation identifiers in the response data; Extract supplementary parameter values and update the operation instruction set according to the parameter supplementation instruction, and / or trigger the multi-protocol connection service processing core to reschedule the execution unit according to the execution confirmation identifier; The updated set of operation instructions and / or the rescheduled operation results data are entered into the historical case database for storage; Based on the historical case database, task execution efficiency optimization parameters are generated, and the matching threshold of the decision rule base is updated.
7. An intelligent control system based on multi-protocol task scheduling, characterized in that, An intelligent control system for executing a multi-protocol task scheduling-based intelligent control method according to any one of claims 1 to 6, the intelligent control system comprising: The instruction generation module is used to receive raw instruction data input by the user, perform semantic parsing, extract instruction keywords, and generate a structured instruction data package containing task type identifiers; The multi-protocol connection service invocation module is used to convert the structured instruction data packet into multi-protocol connection service invocation parameters and send them to the multi-protocol connection service processing core through an internal communication interface; wherein, the multi-protocol connection service invocation parameters include a target system identifier and an operation instruction set; The scheduling and execution module is used by the multi-protocol connection service processing core to schedule the corresponding execution unit to execute the operation instruction set according to the target system identifier, and generate operation result data corresponding to the task type identifier; The result status analysis module is used to perform status analysis on the operation result data and generate decision instruction data containing decision type identifiers. The feedback processing module is used to perform task refinement processing, abnormal alternative generation, or task completion feedback processing based on the decision type identifier, and convert the processing results into natural language feedback output.
8. A computer-readable storage medium, characterized in that: The computer program is stored that can be loaded by a processor and executed as described in any one of claims 1 to 6.