A space management natural language instruction parsing method and a space intelligent control system
Patent Information
- Application Number
- CN202610717854.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-22
- Publication Date
- 2026-09-29
AI Technical Summary
更具体地,当不同用户的指令指向同一设备时,现有方法难以结合当前的操作焦点、空间运行状态或指令本身的明确程度来进行合理的优先级判断
[0009]本申请实施例的一种空间管理自然语言指令解析方法及空间智控系统,通过实时获取指令并匹配专用指令集,系统能够快速理解用户在会议室等场景中的操作意图,防止通用语音助手因缺乏针对性而出现的误识别或响应延迟。其次,引入包含设备状态和空间运行状态的上下文信息,使得解析结果不再是静态的,而是随会议进程动态调整。例如,同一句再暗一点在准备阶段可能调节灯光,在演示阶段则调节投影仪亮度,这种智能适配大幅降低了用户的表达负担。再者,将多设备协同操作自动分解为原子操作序列,确保了幕布、投影、灯光等设备按照正确的时序和依赖关系依次执行,既防止了设备冲突或损坏,也让用户能用一句话完成原本需要多次点击或输入才能实现的复杂场景切换。
Smart Images

Figure CN122837302A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of natural language processing technology, and in particular to a method for parsing natural language instructions for space management and a space intelligent control system. Background Technology
[0002] Existing natural language command parsing methods, whether for general voice assistants or smart home systems, typically rely solely on the command itself and a simple dialogue history. When applied to offline physical spaces such as conference rooms and exhibition halls, these methods struggle to utilize the real-time operational status of the physical space—whether it's currently idle, preparing, in use, or under maintenance—as well as the real-time status of each physical device to aid command understanding. For example, the same command "open" might refer to different devices at different stages of a meeting, and existing systems cannot adaptively adjust their parsing logic based on dynamic changes in the space's status. Furthermore, when commands involve the coordinated operation of multiple physical devices, such as "enter presentation mode" requiring simultaneous control of the screen, projector, lights, and sound system, existing methods lack an orderly operational mechanism.
[0003] Secondly, existing systems typically maintain a globally unified instruction set, attempting to match all possible instructions regardless of the physical space's state. This approach not only increases ambiguity but also makes it difficult to maintain consistent instruction parsing during state transitions. Some systems employ manual switching modes, but this contradicts the principles of natural interaction. In offline physical spaces, multiple users often issue instructions simultaneously. Existing solutions mostly employ simple first-come, first-served or last-come, last-served strategies, lacking arbitration mechanisms for conflicting instructions on the same physical device. More specifically, when different users' instructions target the same device, existing methods struggle to determine reasonable priorities by considering the current operational focus, the space's operational state, or the clarity of the instruction itself. Summary of the Invention
[0004] To address one or more problems in the existing technology, the main objective of this application is to provide a natural language command parsing method for space management and a space intelligent control system.
[0005] To achieve the aforementioned objectives, this application proposes a method for parsing natural language instructions for space management, the method comprising: Real-time acquisition of natural language commands issued to the target physical space; A preset dedicated instruction set is invoked, and natural language instructions are matched using the dedicated instruction set to determine candidate instruction types; Obtain the current context information of the target physical space, the context information including the device status and / or space operation status within the physical space; Based on the context information, the matched instructions are semantically parsed to generate a complete semantic parsing result; If the complete semantic parsing result involves the collaborative operation of multiple physical entities, then the complete semantic parsing result is decomposed into a sequence of atomic operations for each physical entity. Based on the decomposition results, the atomic operation sequence is output as an executable instruction stream.
[0006] This application also provides a space intelligent control system, including: The first acquisition module is used to acquire natural language commands issued to the target physical space in real time. The calling module is used to call a preset dedicated instruction set, and to match natural language instructions through the dedicated instruction set to determine the candidate instruction type; The second acquisition module is used to acquire the current context information of the target physical space, the context information including the device status and / or space operation status within the physical space; The generation module is used to perform semantic parsing on the matched instructions based on the context information and generate a complete semantic parsing result; The analysis module is used to analyze the complete semantic parsing result. If the complete semantic parsing result involves the collaborative operation of multiple physical entities, the complete semantic parsing result is decomposed into a sequence of atomic operations for each physical entity. An output module is used to output the atomic operation sequence as an executable instruction stream based on the decomposition results.
[0007] This application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of any of the methods described above.
[0008] This application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of any of the methods described above.
[0009] This application discloses a spatial management natural language command parsing method and a spatial intelligent control system. By acquiring commands in real time and matching them with a dedicated command set, the system can quickly understand the user's operational intentions in scenarios such as meeting rooms, preventing misrecognition or response delays caused by the lack of specificity in general voice assistants. Secondly, by introducing contextual information including device status and spatial operating status, the parsing results are no longer static but dynamically adjusted according to the meeting progress. For example, the same sentence "make it a little darker" might adjust the lighting during the preparation stage and the projector brightness during the presentation stage. This intelligent adaptation significantly reduces the user's expressive burden. Furthermore, by automatically decomposing multi-device collaborative operations into atomic operation sequences, it ensures that devices such as the screen, projector, and lights execute sequentially according to the correct timing and dependencies, preventing device conflicts or damage and allowing users to complete complex scene switching that originally required multiple clicks or inputs with a single sentence. Attached Figure Description
[0010] Figure 1 This is a flowchart illustrating a spatial management natural language instruction parsing method according to an embodiment of this application; Figure 2 This is a flowchart illustrating a spatial management natural language instruction parsing method according to an embodiment of this application; Figure 3 This is a schematic block diagram of the structure of a space intelligent control system according to an embodiment of this application; Figure 4 This is a schematic block diagram of the structure of a computer device according to an embodiment of this application.
[0011] The realization of the purpose, functional features and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0012] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0013] Reference Figure 1 This application provides a method for parsing natural language instructions for space management, the method comprising: S1. Real-time acquisition of natural language commands issued to the target physical space; S2. Invoke a preset dedicated instruction set, and match natural language instructions through the dedicated instruction set to determine candidate instruction types; S3. Obtain the current context information of the target physical space, the context information including the device status and / or space operation status within the physical space; S4. Based on the context information, perform semantic parsing on the matched instructions to generate a complete semantic parsing result; S5. Analyze the complete semantic parsing result. If the complete semantic parsing result involves the collaborative operation of multiple physical entities, then decompose the complete semantic parsing result into a sequence of atomic operations for each physical entity. S6. Based on the decomposition results, output the atomic operation sequence as an executable instruction stream.
[0014] As described in steps S1-S3 above, this embodiment uses an offline conference room as a typical example of the target physical space. However, those skilled in the art will understand that this method is also applicable to other offline spaces with multiple physical devices, such as exhibition halls, offices, and classrooms. Step one: The system collects natural language commands issued by users in real time through a microphone array or text input interface deployed in the conference room. The system does not continuously perform indiscriminate parsing of all speech to avoid false triggering and wasted power. In actual deployment, the system has a built-in wake-up mechanism. Users need to first speak a preset wake-up command to activate the system from standby mode. After wake-up, the system begins to truly collect and parse subsequent natural language commands. For example, after wake-up, when the host says "Start the presentation" or a participant says "Dim the lights," the system will convert these speech signals into text commands. This acquisition process is continuous; the system listens to all speech input throughout the meeting and can recognize the voices of different users. In step two, the system pre-builds a dedicated instruction set specifically designed for offline meeting room scenarios. This set includes all possible user-used command templates, such as controlling the projector, adjusting lights, raising and lowering the screen, turning the air conditioner on and off, and switching signal sources. Upon receiving a natural language command, the system calls this dedicated instruction set to match it and determine the type of command the user wants to execute. For example, if the user says "turn on the projector," the system will search the instruction set for a matching template, identifying it as a candidate type: "device control—projector power on." The matching process can employ keyword matching, template matching, or lightweight semantic similarity calculation, but it does not rely on general natural language understanding models that require large-scale training data. In step three, after matching the candidate command type, the system does not execute it immediately but further acquires the current context information of the target physical space. This context information includes two categories: first, the real-time status of each physical device, such as whether the projector is already on, whether the screen is lowering, the current brightness percentage of the lights, and the set temperature of the air conditioner; second, the overall operational status of the space, such as whether the meeting room is currently idle, ready, in progress, in a short break, or in a cleanup state. This information is typically obtained in real time via device buses, sensor networks, or reservation systems. The role of contextual information is to provide a basis for subsequent semantic parsing, enabling the system to understand the precise meaning of user commands at the current moment.
[0015] As described in steps S4-S6 above, step four utilizes the context information obtained in the previous step to perform in-depth semantic parsing on the matched instruction, transforming it into a complete and unambiguous semantic expression. Specifically, the parsing process handles the following common cases: When the user's instruction contains a pronoun, such as "it" in "turn it off," the system determines which device "it" refers to based on the current focused physical entity. The current focused physical entity can be a device the user recently operated, a device the user is pointing to determined through sound source localization, or the device name directly stated in the instruction. When the user's instruction omits necessary information, such as only saying "dimer" without specifying whether it refers to the lights or the projector, the system completes the instruction based on the current spatial operating status. If the meeting room is in presentation mode, the system will default to the user referring to the projector brightness; if it is in discussion mode, it will default to the light brightness. The system also verifies the operability of the instruction. For example, if the user requests to turn off the projector, but the system finds that the projector is already off or cooling down, it will determine that the instruction is inoperable and generate a user-friendly error message. Through the above processing, vague, omitted, or incomplete original instructions are transformed into a complete semantic parsing result that clearly contains the target physical entity and specific operation parameters. For example, "brighten that" might be parsed as "increase the projector brightness by 20%". Step 5: When the complete semantic parsing result involves multiple physical entities that need to work together, the system does not simply package all operations together and send them out. Instead, it decomposes them into a sequence of atomic operations for each physical entity. For example, if the user says "enter presentation mode", the parsing result might include four operations: "lower the screen, turn on the projector, dim the lights, and close the curtains". However, directly executing these four operations simultaneously would cause problems: if the projector is turned on before the screen is fully lowered, the light will shine behind the screen; if the lights are dimmed instantly, the user may not have time to adjust. Therefore, the system will determine the dependencies and timing requirements between these operations based on a predefined physical constraint rule base. For example, "lower the screen" must be completed first, "turn on the projector" must wait for the screen to be in place before being executed, and "dim the lights" can be executed in parallel with the screen lowering but must be executed after the projector is turned on. The system decomposes composite instructions into sequences of atomic operations with sequential and delay parameters according to these constraints. Step six concludes by converting the decomposed atomic operation sequences into instruction formats understandable to each physical device, such as sending JSON instructions via IoT protocols, sending Modbus commands via serial ports, or calling a third-party control system via API. The output instruction stream is executed sequentially or in parallel according to the order in the sequence, ensuring the reliability and security of multi-device collaboration.
[0016] As mentioned above, firstly, by acquiring instructions in real time and matching them with a dedicated instruction set, the system can quickly understand the user's operational intentions in scenarios such as meeting rooms, preventing misrecognition or response delays caused by the lack of specificity in general voice assistants. Secondly, by introducing contextual information including device status and spatial operating status, the parsing results are no longer static but dynamically adjusted according to the progress of the meeting. For example, the same phrase "darker" might adjust the lighting during the preparation stage and the projector brightness during the presentation stage; this intelligent adaptation significantly reduces the user's expressive burden. Furthermore, by automatically decomposing multi-device collaborative operations into atomic operation sequences, the system ensures that devices such as the screen, projector, and lights execute sequentially according to the correct timing and dependencies, preventing device conflicts or damage and allowing users to complete complex scene switching that originally required multiple clicks or inputs with a single sentence.
[0017] Reference Figure 2 In one embodiment, the step of invoking a preset dedicated instruction set and matching natural language instructions using the dedicated instruction set to determine candidate instruction types includes: S21. Obtain the current spatial operating status of the target physical space; S22. Based on the current space operation state, activate one or more instruction subsets corresponding to the space operation state from the dedicated instruction set, wherein different space operation states correspond to different instruction subsets; S23. Match the natural language instruction with the instruction template in the activated instruction subset; S24. Based on the matching results, if at least one instruction template is matched in the active instruction subset, the matching results are output as candidate instruction types. S25. If no instruction template is matched in the active instruction subset, then fall back to all instruction templates in the dedicated instruction set for matching, and output the matching result as a candidate instruction type.
[0018] As described above, traditional instruction matching methods compare all possible instruction templates one by one, regardless of the current state of the meeting room. This wastes computing resources and is prone to ambiguity. To address this issue, this solution introduces a mechanism for dynamically activating instruction subsets. First, the current operational state of the target physical space is obtained. This state can be different stages such as idle, preparing, meeting in progress, short break, or cleanup after completion. For example, through a reservation system or human body sensors, the system can determine whether the meeting room is currently in an idle state (just finished cleaning and waiting to be used), a preparing state (someone has entered and is setting up equipment), or a state where the meeting has officially started. Next, based on this state, one or more corresponding instruction subsets are activated from a preset dedicated instruction set. In other words, the system only focuses on instructions related to that state. For example, when the meeting room is in a preparing state, users are most likely to need to perform operations related to equipment setup, such as turning on the projector, lowering the screen, turning on the air conditioner, and adjusting the light brightness. Therefore, the instruction subset activated by the system mainly includes these equipment control instructions. Once the meeting is in progress, user needs shift to switching signal sources, adjusting volume, extending meeting time, or ending the meeting. At this point, the system switches to a different subset of instructions specifically designed to match these meeting-related commands. The system then matches the user's natural language commands against templates in the currently active subset. If a match is successful, for example, if the user says "turn on the projector" during preparation, the system directly finds the corresponding template in the active subset and quickly outputs the candidate command type. However, if a match fails, for example, if the user says "lower the screen" during the meeting, and this command is a common operation during preparation and not in the currently active subset, the system does not immediately report an error. Instead, it uses a rollback mechanism—continuing to match the complete, dedicated instruction set. For example, while "lower the screen" is unlikely to occur during the meeting, it is still a valid device control command and can still be recognized and executed through a global rollback.
[0019] In one embodiment, the step of performing semantic parsing on the matched instruction based on the context information to generate a complete semantic parsing result includes: Extract the current spatial operating state and current device state of the target physical space from the current context information; Analyze the current focus physical entity within the target physical space, wherein the current focus physical entity is determined by the physical entity of the most recent operation in the historical operation record, the spatial pointing entity obtained from sound source localization analysis, or the entity identifier explicitly included in the natural language command. Based on the current focused physical entity, the matched instruction is dereferenced, and the pronouns in the instruction are replaced with the corresponding physical entity identifier; Based on the current spatial operating state, the resolved instructions are omitted and completed. Based on the current device status, the operability of the completed instruction is verified. Based on the verification results, if the target physical entity involved in the instruction is in an inoperable state, an error feedback is generated and the parsing is interrupted. The instructions that pass the operability check will be output as the complete semantic parsing result.
[0020] As mentioned above, after completing instruction matching and determining candidate instruction types, it is necessary to transform the user's potentially vague, omitted, or referential spoken language into clear and unambiguous executable semantics. This process relies on the previously acquired contextual information. The first step is to extract two key elements from the current contextual information: the current spatial operating state and the current device state. The spatial operating state can be in preparation, meeting, or recess; the device state includes whether the projector is powered on, whether the screen is being raised or lowered, and the current brightness percentage of the lights. This information forms the basis for all subsequent parsing actions. The second step is to analyze the current focus physical entity. There are three ways to determine this focus entity. The first is to see which device the user last operated; for example, if the projector was just turned off a few seconds ago, then the projector becomes the focus. The second is to use sound source localization technology, using a microphone array to determine the speaker's facing direction, thereby inferring the device they are referring to; for example, if a participant faces the left curtain and says "pull that up," the system can locate the left curtain. The third is the device name directly stated in the user's instruction, such as the projector in "turn off the projector." These three methods are independent, and the system will use the most suitable one to determine the focus entity based on the actual scenario. The third step is to use this focus entity for reference resolution. Users often use pronouns in daily conversations, such as "it" in "turn it off" or "that one over there." The system will directly replace these pronouns with the identifier of the physical focus entity determined in the second step. If the user explicitly says "turn off the projector," then there is no pronoun to replace; but if they say "turn it off," the system understands that it refers to the projector currently being focused on. The fourth step is to perform abbreviation completion based on the current spatial operating state. Users often omit the operation object or parameter in quick conversations, such as only saying "dimming it a little" or "brighter." The system needs to determine whether "dimming it a little" refers to the light or the projector brightness. This is where the spatial operating state comes in handy. If the current mode is a presentation mode in a meeting, the system will assume that the user is referring to the projector brightness; if the current mode is a discussion mode, it will assume that the user is referring to the ambient light. Similarly, if someone says "be quieter," the system may complete it to lower the speaker volume. The specific rules for completion can be pre-configured, with different completion results for the same omitted expression under different states. The fifth step is to perform an operability check based on the current device status. After resolution and completion, the system has obtained a clear instruction, such as "reduce the projector brightness by 10%". However, before actually executing it, the system will first check whether the projector is actually in a state where the brightness can be adjusted. For example, whether the projector is already turned on, whether it is cooling down, and whether the brightness adjustment function is locked. If the user requests to turn off an air conditioner that is already off, the system will also find that this operation is meaningless.When a validation fails, the system doesn't silently fail; instead, it generates a user-friendly error message, such as "The projector is not powered on; please turn it on first," and interrupts the current parsing process. This allows the user to identify the problem. Finally, only instructions that pass all validations are output as the final, complete semantic parsing result, available for subsequent collaborative decomposition or direct execution.
[0021] In one embodiment, if the complete semantic parsing result involves the collaborative operation of multiple physical entities, then the complete semantic parsing result is decomposed into a sequence of atomic operations for each physical entity, the steps of which include: Extract the identifiers of multiple physical entities involved and the corresponding operations to be performed for each physical entity from the complete semantic parsing result; Obtain a preset physical constraint rule base, which includes the dependency relationships, mutual exclusion relationships and timing parameters required for each operation between different physical entities; Based on the physical constraint rule base, dependency analysis is performed on the extracted physical entities and operations to be executed to determine the execution order constraints and mutual exclusion constraints between the operations. According to the execution order constraint, each operation to be executed is constructed as an atomic operation, and the atomic operations are arranged according to the execution order constraint to form an atomic operation sequence; According to the mutual exclusion constraint, the atomic operation sequence is checked. If there are mutually exclusive atomic operations in the sequence, the sequence is adjusted or the operations are merged according to the preset conflict handling strategy. The adjusted sequence of atomic operations is output as the decomposition result.
[0022] As mentioned above, once the system obtains a complete semantic parsing result through the preceding steps, this result may only involve a single operation of a single device, such as "close the left curtain." However, more often, users will issue compound commands involving the collaborative work of multiple physical entities, such as "enter presentation mode" or "start video conference." These commands often require simultaneous or sequential control of multiple devices, including the screen, projector, lights, speakers, and camera. Sending all these operations out at once would lead to unsatisfactory execution, or even damage to equipment or disruption of the meeting. To overcome this problem, the first step is to extract all involved physical entity identifiers and the specific operations each entity needs to perform from the complete semantic parsing result. For example, if a user says "prepare to start the meeting," the parsing result might include lowering the screen, turning on the projector, dimming the lights to 30%, closing the curtain closest to the screen, and turning on the air conditioner and setting it to 24 degrees Celsius. The system will break these operations down into a task list to be executed. The second step is to obtain a pre-configured physical constraint rule base. This rule base is specifically designed for the devices and their physical characteristics in the current meeting room. The rules define dependencies between different devices. For example, the projector can only be turned on after the screen is fully lowered; otherwise, the projection light will be blocked by the screen or shine on the wrong location. Mutual exclusion relationships are also defined, such as the air conditioner's cooling and ventilation modes not being activated simultaneously, or the projector's power-on and power-off commands not being sent at the same time. Additionally, timing parameters for each operation are included; for example, it takes approximately three seconds for the screen to lower completely, and five seconds for the projector to reach normal brightness during a cold start. These rules are not universal but manually configured based on the actual device model and installation environment. The third step involves performing dependency analysis on the extracted operations based on this rule base. The analysis determines the execution order constraints and mutual exclusion constraints that must be followed between each operation. For example, lowering the screen must be completed before turning on the projector, while dimming the lights can be done simultaneously with lowering the screen, but must be completed before turning on the projector to avoid glare. Air conditioning adjustment has no dependency on other operations and can be executed at any time. Regarding mutual exclusion constraints, if one operation requires the screen to rise while another requires it to fall, these two operations cannot be executed simultaneously. The fourth step involves constructing each operation to be executed into an atomic operation according to the execution order constraint. An atomic operation here refers to the smallest indivisible unit of execution, such as sending the instruction "curtain lowers". These atomic operations are then arranged in the analyzed order to form an ordered sequence. For operations that can be executed in parallel, the system marks them in the sequence as having no waiting relationship, allowing them to be sent concurrently. The fifth step involves the system verifying the formed atomic operation sequence according to mutual exclusion constraints.If two mutually exclusive atomic operations are found in the sequence, such as a request to turn on the projector followed by a request to turn it off, the system will not simply execute them sequentially, causing the projector to turn on before turning it off. Instead, it will adjust according to a preset conflict handling strategy. Possible strategies include keeping the latter instruction and deleting the former, merging the two operations into a single state-to-state instruction, or directly notifying the user of the conflict. The adjustment principle can be timestamp priority, i.e., keeping the latest instruction, or instruction priority, such as the turn-off instruction taking precedence over the turn-on instruction. In the sixth step, the system outputs the atomic operation sequence after mutual exclusion verification and adjustment as the decomposition result for use by subsequent execution modules.
[0023] In one embodiment, when the issued natural language command is from multiple users, multi-user command conflict resolution is performed, including the following steps: Within a preset time window, detect multiple natural language commands targeting the same physical entity; Based on multiple natural language commands, the current focus physical entity is obtained, and it is determined whether the target physical entity corresponding to each of the multiple natural language commands is consistent with the current focus physical entity. If at least one natural language instruction targets a physical entity that matches the currently focused physical entity, then that instruction is determined as the priority instruction to be executed. If the target physical entity of all natural language instructions is inconsistent with the current focus physical entity, or the current focus physical entity is empty, then the current spatial operation state is obtained, and the instruction to be executed is determined according to the preset conflict handling rules corresponding to the spatial operation state. If the priority instruction cannot be determined based on the spatial operating status, then the semantic precision parameters of each of the multiple natural language instructions are extracted, and the instruction with the highest semantic precision is determined as the priority instruction to be executed. The system will begin matching the identified priority natural language instructions, determine the types of candidate instructions, and output conflict warnings for other instructions.
[0024] As mentioned above, in real-world meeting room scenarios, it's rare for only one person to speak. While the host is saying "brighten the projector," a participant might simultaneously say "close the curtains," or worse, two people might issue conflicting commands to the same device almost simultaneously—one saying "turn on the projector" while the other says "turn off the projector." Existing voice control systems typically process commands based on the order in which microphones pick up the audio; whoever's voice is captured first executes the command first. However, this approach ignores the social rules and the continuity of physical operations in offline spaces, easily leading to inconsistent device statuses and disrupted meetings. To resolve these conflicts, the system first defines a time window. This window isn't fixed in length but begins when the first natural language command is received and ends when the candidate command type is determined. Within this window, the system continuously collects commands from other users. After the window closes, the system checks for multiple commands targeting the same physical entity. For example, if User A and User B almost simultaneously request to operate the projector, these two commands are identified as conflicting commands. Next, the system identifies the current focused physical entity. The focal entity could be the projector recently operated by user A, a user speaking directly to the projector determined through sound source localization, or the device name explicitly stated in the instruction. The system compares the target physical entity of each conflicting instruction with the current focal entity. If the target entity of an instruction happens to be the current focal entity, the system assumes the initiator of this instruction is the person most recently operating the device, and therefore designates this instruction as the priority instruction. This means that the presenter who was just controlling the projector will have a higher execution priority for saying "turn it up a bit" than someone else saying "turn off the projector," thus preventing the presentation from being interrupted. If the target entities of all instructions are inconsistent with the current focal entity, or if the current focal entity itself is empty, it means there is no clear most recent operator to refer to. At this point, the system switches to the second level of judgment, namely the current spatial operating state. The spatial operating state could be different stages such as preparation, meeting, or recess. Each state corresponds to a set of preset conflict handling rules. For example, in the presentation mode of a meeting, instructions like turning off the projector might be temporarily blocked, while instructions like adjusting the volume remain open. The system determines which instruction to execute first based on the rules corresponding to the current state. If the first two levels of judgment still cannot distinguish between them—that is, if there is no focal entity to refer to and no clear rules are given for the spatial operating state—the system will proceed to the third level of judgment: comparing the semantic precision of each instruction. Semantic precision refers to whether the parameter information provided in the instruction is complete and specific enough. For example, if one user says "adjust the projector brightness to 50%," and another user says "dim it a little," the former is clearly more precise.The system extracts parameter details from each instruction and prioritizes the instruction with the highest level of completeness. This indirectly encourages users to use clear and concise expressions. Finally, the system sends the prioritized instructions to the normal matching process for candidate instruction type identification and subsequent parsing. Meanwhile, for other conflicting instructions that are temporarily shelved, the system outputs user-friendly conflict alerts, such as "Multiple users detected controlling the projector simultaneously; the host's instruction has been prioritized," allowing other users to understand why their instructions have not been executed, rather than being silently ignored.
[0025] In one embodiment, the step of detecting multiple natural language instructions targeting the same physical entity within a preset time window includes: The moment the first natural language instruction is received is taken as the starting point of the time window; The moment when the candidate instruction type of the first natural language instruction is determined is taken as the end point of the time window; Within the time window, continuously receive newly arriving natural language instructions; Detect whether there are multiple natural language instructions targeting the same physical entity among all natural language instructions received within the time window.
[0026] As mentioned above, in scenarios where multiple users issue commands simultaneously, a very practical problem is how to define which commands arrive simultaneously. If a fixed time window is used, for example, stipulating that anything within half a second of receiving the first command is considered simultaneous, then when the system takes a long time to process complex commands, truly concurrent commands may be treated as arriving sequentially and processed separately because they fall outside the fixed window, causing commands that should be captured by the conflict detection mechanism to be missed. Conversely, if the fixed window is set too large, commands that are not actually issued simultaneously will be incorrectly classified into the same batch of conflicts. To resolve this contradiction, this embodiment provides a method for determining a dynamic time window. The first step is to take the moment when the first natural language command is received as the starting point of the time window. This moment is explicitly recorded by the system's clock. For example, when the microphone array captures the sound of user A saying "turn on the projector," the system immediately records the timestamp of this moment, and the window starts timing. The second step is to take the moment when the candidate command type of the first natural language command is determined as the ending point of the time window. The determination of the candidate command type refers to the process by which the system compares the first command with the currently active subset of commands and obtains a matching result. The time required for this process is not fixed. If User A says a very common command, such as turning on the light, the matching process might only take tens of milliseconds. However, if the command is more complex or requires backtracking, the time could reach hundreds of milliseconds or even longer. Using the dynamic completion time as the end point of the window means that the window length automatically adapts to the processing complexity of the first command. The third step is to continuously receive newly arriving natural language commands within the time window. That is, from the start point to the end point of the window, the system will not immediately process subsequent commands separately, but will temporarily store them all. For example, while User A's "turn on the projector" is still being matched, User B says "turn off the projector," and User C says "open the curtains." These commands will be temporarily stored and processed uniformly after the window closes. The fourth step is to check if there are multiple commands targeting the same physical entity among all received commands after the window closes. If User A and User B's commands both point to the projector, the system determines that there is a conflict and then triggers the aforementioned multi-user command conflict resolution logic. However, User C's command points to the curtains, which is unrelated to the projector, so it does not affect the conflict detection result.
[0027] It's worth noting that this dynamic time window design addresses the inherent flaws of fixed-window solutions. In offline meeting rooms, it's common for two people to speak almost simultaneously, but the timing of their voices reaching the microphone may differ slightly. The time required for the system to process the first instruction determines whether the second instruction can be included in the same conflict batch. Using a fixed-length time window results in either an excessively long window that catches irrelevant instructions when processing is fast, or an excessively short window that misses truly concurrent instructions when processing is slow. The dynamic window binds the window's end point to the actual completion time of the first instruction, thus solving this problem. More importantly, this mechanism eliminates the need for users or system administrators to pre-set a suitable window length; instead, the window length adaptively determines itself based on the matching time of the current instruction. When the instruction is simple, the window automatically shortens to avoid unnecessary data backlog; when the instruction is complex, the window automatically lengthens to ensure that concurrent instructions occurring during this period are not missed.
[0028] In one embodiment, before the step of resolving the reference of the matched instruction based on the current focus physical entity, the method further includes: When the current focus physical entity is empty and the instruction includes a pronoun that cannot point to a specific entity identifier, obtain the spatial pointing angle corresponding to the sound source localization; Obtain the pre-configured spatial visibility angle range of each physical entity within the target physical space; Calculate the angular distance between the spatial pointing angle and the center angle of the visible angle range of each physical entity; Based on the calculation results, the angular distances are converted into corresponding probability weights; Based on the probability weights, one or more candidate physical entities are determined; Based on the determined candidate physical entities, if the highest probability weight exceeds a preset threshold, the physical entity corresponding to the probability weight is taken as the target entity pointed to by the pronoun. If the highest probability weight does not exceed the preset threshold, a confirmation query including the candidate physical entity is generated and output, and the target entity is determined based on the user's response.
[0029] As mentioned above, a crucial prerequisite for the aforementioned reference resolution step is the existence of the current focus physical entity, allowing the system to replace the pronoun in the instruction with the entity's identifier. However, in practical use, situations often arise where the focus entity is empty. For example, the conference room may have just been opened and there are no historical operation records, or the user may directly utter an instruction with a pronoun without specifying the direction of the sound source. In this case, if the user says "turn that off," the system neither knows which device "that" refers to nor has a focus entity to refer to, making this one of the most troublesome scenarios in offline voice control. To address this issue, this embodiment provides an alternative positioning method based on spatial visibility angle. This method is triggered only when the current focus physical entity is empty and the user's instruction contains a pronoun, serving as a pre-processing step for reference resolution. First, the system obtains the spatial pointing angle corresponding to the sound source positioning. Using the microphone array deployed in the conference room, the system can calculate the direction angle of the user emitting the sound relative to the array. For example, when the user is sitting on the left side of the conference table and speaking towards the projection screen, the sound source positioning algorithm can output a specific angle value, such as twelve degrees to the left of the center of the microphone array. This angle information is the basis for all subsequent calculations. Next, the pre-configured visible angle ranges of each physical entity within the target physical space are obtained. The visible angle range refers to the azimuth interval occupied by a device in space when viewed from the reference position of the microphone array. Each device needs to be pre-calibrated. For example, if the center of the projection screen is at 0 degrees directly in front, the visible angle range of the screen is from -10 degrees to +10 degrees. The visible angle range of the left curtain is from 25 degrees to 45 degrees. The visible angle range of the right curtain is from -45 degrees to -25 degrees. This calibration can be completed during device installation without user intervention. After obtaining the sound source pointing angle and the visible angle ranges of each device, the angular distance between the sound source pointing angle and the central angle of each device's visible angle range is calculated. The central angle is the midpoint of the visible angle range; for example, the central angle between -10 degrees and +10 degrees is 0 degrees, and the central angle between 25 degrees and 45 degrees is 35 degrees. If the sound source pointing angle is 12 degrees, then its angular distance to the central angle of the projection screen is 12 degrees, and its angular distance to the central angle of the left curtain is 23 degrees. The smaller the angular distance, the greater the likelihood that the user is pointing to that device. Next, each angular distance is converted into a corresponding probability weight. There are several ways to do this, such as taking the inverse of the angular distance and normalizing it, or using a Gaussian function mapping. After the conversion, each device receives a probability weight between zero and one, with all weights summing to one. Devices with smaller angular distances receive higher weights. Based on these probability weights, one or more candidate physical entities are determined. The simplest approach is to directly select the device with the highest probability weight as the sole candidate. However, for safety, the system can also retain the top three weighted devices as a candidate list, allowing for user confirmation when necessary.Then, it checks whether the highest probability weight exceeds a preset threshold. This threshold can be 70% or 80%, and can be adjusted according to the actual use case. If the highest probability weight exceeds the threshold, it means the system has enough confidence to determine that the user is referring to the correct device, so it identifies that device as the target entity pointed to by the pronoun and continues with subsequent reference resolution. If the highest probability weight does not exceed the threshold, it means that the weights of multiple devices are relatively close, making it difficult to make a reliable judgment. In this case, the system will not force a guess, but will generate a confirmation query containing candidate physical entities and output it to the user. For example, the system may prompt via voice or screen, "Do you mean the projection screen or the left curtain?" The user only needs to reply simply to complete the confirmation. The system determines the final target entity based on the user's reply.
[0030] This embodiment addresses the most challenging challenge of referential resolution in offline physical spaces. When a user utters vague pronouns such as "that," "over here," or "that thing," and there are no historical operation records to refer to, general voice systems typically can only report an error saying they cannot understand, or randomly guess a device. Random guesses have serious consequences; they might shut down an important device or operate on a completely incorrect object. This solution utilizes two inherent but often overlooked features of offline physical spaces. The first feature is that each device has a definite directional range relative to a fixed reference point, which can be precisely calibrated during installation. The second feature is that while sound source localization cannot be accurate to the centimeter level, it can provide the approximate direction of the user's voice. By combining these two pieces of information and calculating angular distance and probability weights, it is possible to statistically infer the device the user is most likely pointing to.
[0031] Reference Figure 3 This application also provides a space intelligent control system, including: The first acquisition module 1 is used to acquire natural language commands issued to the target physical space in real time. Module 2 is invoked to invoke a preset dedicated instruction set, and the natural language instructions are matched using the dedicated instruction set to determine the candidate instruction type; The second acquisition module 3 is used to acquire the current context information of the target physical space, the context information including the device status and / or space operation status within the physical space; The generation module 4 is used to perform semantic parsing on the matched instructions based on the context information and generate a complete semantic parsing result; Analysis module 5 is used to analyze the complete semantic parsing result. If the complete semantic parsing result involves the collaborative operation of multiple physical entities, the complete semantic parsing result is decomposed into a sequence of atomic operations for each physical entity. Output module 6 is used to output the atomic operation sequence as an executable instruction stream based on the decomposition result.
[0032] As described above, it is understood that each component of the space intelligent control system proposed in this application can realize the function of any of the space management natural language instruction parsing methods described above, and the specific structure will not be repeated.
[0033] Reference Figure 4 This application also provides a computer device, which may be a server, and its internal structure may be as follows: Figure 4 As shown, this computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The database stores monitoring data and other data. The network interface is used for communication with external terminals via a network connection. When executed by the processor, the computer program implements a space management natural language instruction parsing method.
[0034] The processor described above executes the aforementioned space management natural language instruction parsing method, including: acquiring natural language instructions issued to a target physical space in real time; invoking a preset dedicated instruction set, matching the natural language instructions through the dedicated instruction set, and determining candidate instruction types; acquiring the current context information of the target physical space, the context information including the device status and / or space operation status within the physical space; performing semantic parsing on the matched instructions based on the context information, generating a complete semantic parsing result; analyzing the complete semantic parsing result, and if the complete semantic parsing result involves the collaborative operation of multiple physical entities, decomposing the complete semantic parsing result into atomic operation sequences for each physical entity; and outputting the atomic operation sequences as an executable instruction stream based on the decomposition result.
[0035] One embodiment of this application also provides a computer-readable storage medium storing a computer program. When executed by a processor, the computer program implements a natural language instruction parsing method for space management, comprising the steps of: acquiring natural language instructions issued to a target physical space in real time; invoking a preset dedicated instruction set, matching the natural language instructions through the dedicated instruction set, and determining candidate instruction types; acquiring the current context information of the target physical space, the context information including the device status and / or space operation status within the physical space; performing semantic parsing on the matched instructions based on the context information to generate a complete semantic parsing result; analyzing the complete semantic parsing result, and if the complete semantic parsing result involves the collaborative operation of multiple physical entities, decomposing the complete semantic parsing result into atomic operation sequences for each physical entity; and outputting the atomic operation sequences as an executable instruction stream based on the decomposition result.
[0036] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in this application and in the embodiments can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual-speed SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
[0037] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, apparatus, article, or method that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, apparatus, article, or method. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, apparatus, article, or method that includes that element.
[0038] The above description is only a preferred embodiment of this application and does not limit the patent scope of this application. Any equivalent structural or procedural changes made based on the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.
Claims
1. A method for parsing natural language instructions for space management, characterized in that, The method includes: Real-time acquisition of natural language commands issued to the target physical space; A preset dedicated instruction set is invoked, and natural language instructions are matched using the dedicated instruction set to determine candidate instruction types; Obtain the current context information of the target physical space, the context information including the device status and / or space operation status within the physical space; Based on the context information, the matched instructions are semantically parsed to generate a complete semantic parsing result; If the complete semantic parsing result involves the collaborative operation of multiple physical entities, then the complete semantic parsing result is decomposed into a sequence of atomic operations for each physical entity. Based on the decomposition results, the atomic operation sequence is output as an executable instruction stream.
2. The spatial management natural language instruction parsing method according to claim 1, characterized in that, The steps of invoking a preset dedicated instruction set, matching natural language instructions using the dedicated instruction set, and determining candidate instruction types include: Obtain the current spatial operating state of the target physical space; Based on the current space operation state, activate one or more instruction subsets corresponding to the space operation state from the dedicated instruction set, wherein different space operation states correspond to different instruction subsets; Match the natural language instruction with the instruction template in the activated instruction subset; Based on the matching results, if at least one instruction template is matched in the active instruction subset, the matching result is output as a candidate instruction type. If no instruction template is matched in the active instruction subset, the process falls back to matching all instruction templates in the dedicated instruction set, and the matching result is output as a candidate instruction type.
3. The spatial management natural language instruction parsing method according to claim 1, characterized in that, The step of performing semantic parsing on the matched instruction based on the context information to generate a complete semantic parsing result includes: Extract the current spatial operating state and current device state of the target physical space from the current context information; Analyze the current focus physical entity within the target physical space, wherein the current focus physical entity is determined by the physical entity of the most recent operation in the historical operation record, the spatial pointing entity obtained from sound source localization analysis, or the entity identifier explicitly included in the natural language command. Based on the current focused physical entity, the matched instruction is dereferenced, and the pronouns in the instruction are replaced with the corresponding physical entity identifier; Based on the current spatial operating state, the resolved instructions are omitted and completed. Based on the current device status, the operability of the completed instruction is verified. Based on the verification results, if the target physical entity involved in the instruction is in an inoperable state, an error feedback is generated and the parsing is interrupted. The instructions that pass the operability check will be output as the complete semantic parsing result.
4. The spatial management natural language instruction parsing method according to claim 1, characterized in that, If the complete semantic parsing result involves the collaborative operation of multiple physical entities, then the complete semantic parsing result is decomposed into a sequence of atomic operations for each physical entity, the steps of which include: Extract the identifiers of multiple physical entities involved and the corresponding operations to be performed for each physical entity from the complete semantic parsing result; Obtain a preset physical constraint rule base, which includes the dependency relationships, mutual exclusion relationships and timing parameters required for each operation between different physical entities; Based on the physical constraint rule base, dependency analysis is performed on the extracted physical entities and operations to be executed to determine the execution order constraints and mutual exclusion constraints between the operations. According to the execution order constraint, each operation to be executed is constructed as an atomic operation, and the atomic operations are arranged according to the execution order constraint to form an atomic operation sequence; According to the mutual exclusion constraint, the atomic operation sequence is checked. If there are mutually exclusive atomic operations in the sequence, the sequence is adjusted or the operations are merged according to the preset conflict handling strategy. The adjusted sequence of atomic operations is output as the decomposition result.
5. The spatial management natural language instruction parsing method according to claim 1, characterized in that, The method further includes, when the issued natural language command is from multiple users, performing multi-user command conflict resolution, the steps of which include: Within a preset time window, detect multiple natural language commands targeting the same physical entity; Based on multiple natural language commands, the current focus physical entity is obtained, and it is determined whether the target physical entity corresponding to each of the multiple natural language commands is consistent with the current focus physical entity. If at least one natural language instruction targets a physical entity that matches the currently focused physical entity, then that instruction is determined as the priority instruction to be executed. If the target physical entity of all natural language instructions is inconsistent with the current focus physical entity, or the current focus physical entity is empty, then the current spatial operation state is obtained, and the instruction to be executed is determined according to the preset conflict handling rules corresponding to the spatial operation state. If the priority instruction cannot be determined based on the spatial operating status, then the semantic precision parameters of each of the multiple natural language instructions are extracted, and the instruction with the highest semantic precision is determined as the priority instruction to be executed. The system will begin matching the identified priority natural language instructions, determine the types of candidate instructions, and output conflict warnings for other instructions.
6. The spatial management natural language instruction parsing method according to claim 5, characterized in that, The step of detecting multiple natural language commands targeting the same physical entity within a preset time window includes: The moment the first natural language instruction is received is taken as the starting point of the time window; The moment when the candidate instruction type of the first natural language instruction is determined is taken as the end point of the time window; Within the time window, continuously receive newly arriving natural language instructions; Detect whether there are multiple natural language instructions targeting the same physical entity among all natural language instructions received within the time window.
7. The spatial management natural language instruction parsing method according to claim 3, characterized in that, Before the step of resolving the reference of the matched instruction based on the current focused physical entity, the method further includes: When the current focus physical entity is empty and the instruction includes a pronoun that cannot point to a specific entity identifier, obtain the spatial pointing angle corresponding to the sound source localization; Obtain the pre-configured spatial visibility angle range of each physical entity within the target physical space; Calculate the angular distance between the spatial pointing angle and the center angle of the visible angle range of each physical entity; Based on the calculation results, the angular distances are converted into corresponding probability weights; Based on the probability weights, one or more candidate physical entities are determined; Based on the determined candidate physical entities, if the highest probability weight exceeds a preset threshold, the physical entity corresponding to the probability weight is taken as the target entity pointed to by the pronoun. If the highest probability weight does not exceed the preset threshold, a confirmation query including the candidate physical entity is generated and output, and the target entity is determined based on the user's response.
8. A space intelligent control system, characterized in that, The method for any one of claims 1-7 comprises: The first acquisition module is used to acquire natural language commands issued to the target physical space in real time. The calling module is used to call a preset dedicated instruction set, and to match natural language instructions through the dedicated instruction set to determine the candidate instruction type; The second acquisition module is used to acquire the current context information of the target physical space, the context information including the device status and / or space operation status within the physical space; The generation module is used to perform semantic parsing on the matched instructions based on the context information and generate a complete semantic parsing result; The analysis module is used to analyze the complete semantic parsing result. If the complete semantic parsing result involves the collaborative operation of multiple physical entities, the complete semantic parsing result is decomposed into a sequence of atomic operations for each physical entity. An output module is used to output the atomic operation sequence as an executable instruction stream based on the decomposition results.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.