Intelligent equipment control method and device, intelligent central control equipment and medium
By acquiring video streams through camera units to detect human activity and match strategy templates, the problem of smart home systems being unable to be personalized is solved, enabling efficient and low-cost personalized services.
Patent Information
- Application Number
- CN202511774644.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-28
- Publication Date
- 2026-02-27
AI Technical Summary
Existing smart home systems are limited by their sensing capabilities and cannot understand user activities, resulting in a lack of personalized control. Furthermore, the deployment of multiple sensors increases costs and complexity.
The camera unit acquires a preview video stream, detects the movement of people, generates a description of the scene's activity process, and generates device control commands based on the centralized control strategy template to achieve personalized services.
It improves the accuracy and flexibility of smart home control, reduces hardware costs and system complexity, and enhances user experience and scalability.
Smart Images

Figure CN121578664A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of intelligent control, and in particular to an intelligent device control method and apparatus, intelligent central control equipment and medium. Background Technology
[0002] Currently, smart home systems commonly use passive human presence sensors such as infrared sensors, microwave radar, or ultrasonic sensors as the basis for triggering automated scenarios. These sensors rely on detecting movement or heat changes within a specific area to determine the presence of people. While this approach achieves basic automated control—such as lights turning on when someone is present and turning off when they leave—its underlying technology has inherent limitations. Because it can only provide binary signals indicating whether someone is present, it cannot acquire any visual contextual information about the person's identity, specific behavior, or intentions. This means the system cannot distinguish between different family members, nor can it identify whether a user is sitting, standing, lying down, or engaging in other specific activities. Therefore, control systems based on these sensors can only execute pre-set, uniform, and simple commands, unable to provide personalized services based on different users or different behavioral states of the same user; their control logic is simplistic and rigid.
[0003] Furthermore, due to a lack of deep understanding of the scenarios, achieving a certain degree of personalized or complex control—such as automatically turning on the TV and dimming the lights for the elderly, or restricting TV operation and activating ambient lighting for children—usually requires deploying various types and numbers of dedicated sensors (such as door and window sensors, pressure pads, and sensors for specific areas) and complex logic programming. This approach not only significantly increases the system's hardware costs, wiring complexity, and maintenance difficulty, but also suffers from weak customization capabilities, making it difficult to flexibly adapt to the ever-changing needs of daily family activities. When scenario requirements change, it often necessitates reconfiguring the hardware or modifying complex linkage rules, resulting in a poor user experience.
[0004] Therefore, existing smart home control systems suffer from several limitations. Firstly, due to limitations in sensing methods, they cannot understand user activities, thus failing to achieve truly personalized control. Secondly, to achieve limited scenario adaptation, they rely on multi-sensor hardware stacking, resulting in high costs, complex architecture, and poor scalability. This application aims to address these shortcomings and comprehensively improve the intelligence level of smart home systems. Summary of the Invention
[0005] The primary objective of this application is to solve at least one of the above-mentioned problems by providing an intelligent device control method and apparatus, intelligent central control equipment and medium.
[0006] To achieve the various objectives of this application, the following technical solution is adopted: A smart device control method provided for one of the purposes of this application includes the following steps: Based on the preview video stream generated by the camera unit, the process of detecting human activity is performed, and the scene activity process description corresponding to the human activity in the preview video stream in the scene is determined. Based on the scene activity process description, a matching centralized control strategy template is determined. The centralized control strategy template includes a strategy conversion rule that maps the scene activity process description to device control instructions adapted to at least one smart device in the smart home network. Based on the policy conversion rules in the centralized control policy template, corresponding device control commands are generated to control the corresponding smart devices to respond.
[0007] A smart device control apparatus, proposed for a smart device control method to meet one of the purposes of this application, comprises: The activity detection module is configured to detect the activity process of a person based on the preview video stream generated by the camera unit, and determine the scene activity process description corresponding to the activity of the person in the preview video stream in their scene. The template matching module is configured to determine a matching centralized control strategy template based on the scene activity process description. The centralized control strategy template includes a strategy conversion rule that maps the scene activity process description to device control instructions adapted to at least one smart device in the smart home network. The generation control module is configured to generate corresponding device control commands based on the policy conversion rules in the centralized control policy template, and control the corresponding smart devices to respond.
[0008] On another front, an intelligent central control device provided to suit one of the purposes of this application includes a camera unit and a controller, the controller including a processor and a memory, the camera unit being used to acquire and preview video streams, and the processor calling and running a computer program in the memory to execute the steps of the intelligent device control method.
[0009] In another aspect, a computer-readable storage medium is provided to suit another purpose of this application, which stores, in the form of computer-readable instructions, a computer program implemented according to the intelligent device control method, which, when invoked by a computer, executes the steps included in the corresponding method.
[0010] Compared to traditional technologies, this application introduces a camera unit to detect human activity in the preview video stream, acquiring far richer environmental information than traditional sensors. By analyzing the video stream, it can identify the specific identity, behavior, and activities of individuals within the scene, generating a descriptive description of the scene's activities, truly understanding who is doing what. Furthermore, control commands are generated by matching a centralized control strategy template. Since the strategy template contains conversion rules that map descriptive information to specific device control commands, it can automatically trigger corresponding device linkages based on the identified user identity and their specific activities, rather than individual actions, to execute personalized scenes, unlike traditional systems that can only execute single, fixed automated responses. This significantly improves the accuracy and flexibility of smart home control and greatly enhances the user experience. Moreover, this application utilizes existing or low-cost camera units to achieve complex environmental perception and personnel identification, avoiding the increased hardware costs, system complexity, and maintenance difficulties associated with deploying multiple types of sensors. This method of achieving multi-functional recognition with a single sensing source not only ensures more powerful functions but also effectively reduces overall cost and complexity, improves scalability, and makes smart home systems more adaptable to diverse family scenarios and needs. Attached Figure Description
[0011] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, wherein: Figure 1 This is a flowchart illustrating a typical embodiment of the intelligent device control method of this application; Figure 2 This is a schematic block diagram of the intelligent device control device of this application; Figure 3 This is a schematic diagram of the structure of a computer device used in this application. Detailed Implementation
[0012] The intelligent device control method involved in this application can be implemented in the form of a computer program. This program can be installed and run on an intelligent central control device in a smart home environment, such as a smart gateway, smart speaker, smart central control screen, or dedicated home server. This intelligent central control device acts as the processing hub of the entire smart home system, maintaining communication connections with image acquisition devices distributed in one or more areas and various controlled intelligent terminal devices (which can be simply referred to as intelligent devices) through a home LAN, thereby forming a complete collaborative control smart home network.
[0013] Image acquisition devices can be smart cameras with image processing capabilities, deployed in key areas such as living rooms, bedrooms, and hallways to continuously acquire environmental video streams. These devices can be integrated into a single camera unit within a smart device or into the camera unit of a central control device. The controlled smart devices encompass various types, including lighting, home appliances, security, and entertainment equipment, such as smart lights, smart TVs, air conditioners, curtain motors, and smart door locks. They receive device control commands from the central control device via communication protocols such as Wi-Fi, Zigbee, and Bluetooth.
[0014] In practical applications, smart home systems, with the user's legal authorization, perceive human activity by analyzing real-time video data captured by camera units, such as preview video streams. For example, in a living room scenario, it can recognize the continuous actions of a specific family member entering a room, walking to the sofa, and sitting down; in a bedroom scenario, it can recognize the user getting into bed; and in a doorway scenario, it can recognize the actions of a delivery person delivering a package. By understanding these dynamic activities, corresponding device combinations can be triggered to respond, such as automatically adjusting the color temperature and brightness of lights, activating entertainment systems, regulating the ambient temperature, and broadcasting voice messages, thereby achieving truly intelligent services tailored to the user's immediate needs.
[0015] The image information used in this application, including the preview video stream captured by the camera unit and its intermediate or final products, can be processed entirely locally on the smart home network unless specifically authorized, without leaking this image information outside the smart home network. This ensures the security of users' personal information on the one hand, and allows for rapid processing of image information and obtaining relevant results by relying on the powerful performance of the local device itself.
[0016] For ease of understanding, the following explains some of the basic concepts involved in this application. The scene activity process description referred to in this application is a comprehensive, semantic description of personnel identity, behavior, and the state of the surrounding environment generated through video stream analysis. It serves as an intermediary connecting raw sensory information and device control strategies. The centralized control strategy template, on the other hand, is a predefined set of rules, essentially mapping a specific scene activity process description into a set of ordered device control instructions. The strategy conversion rules in the centralized control strategy template are the fundamental elements of the template and can be flexibly customized. For example, they can specify the target device to be controlled to achieve a certain scene effect, the target state the device should reach, and the timing relationship of instruction execution. Through the combination of the above methods, this application achieves a leap from passive sensing to active perception, and from single-switch control to personalized scene-based services.
[0017] Please see Figure 1In some embodiments, the intelligent device control method of this application can be implemented as an application program running in the processor of the intelligent central control device. The method includes: Step S3100: Detect the character's activity process based on the preview video stream generated by the camera unit, and determine the scene activity process description corresponding to the character's activity in the scene where the character is located. When the intelligent central control device is operational, its camera unit is activated to continuously capture video images of the environment, thus forming a continuous preview video stream. This preview video stream serves as the raw data source for detecting human activity, and its processing aims to extract meaningful semantic information from the dynamic video sequence. The task of human activity detection is not only to identify the presence of people in the scene, but more importantly, to understand the activities people are performing and their spatial context, thereby generating a comprehensive description of the scene's activity process. This description goes beyond simple existence judgments; it is a coherent interpretation of who, where, what, and what objects they might be interacting with.
[0018] To achieve the above detection, one embodiment includes frame-by-frame or time-window-based image analysis of the preview video stream. For example, firstly, various objects in the video frames can be identified and classified using an object detection model, distinguishing between human figures and other non-human objects. For identified human targets, they can be continuously tracked to establish their motion trajectory between different frames. Based on this, the user identity of the person can be further determined using a face recognition model or gait recognition model, while their action sequence, i.e., temporal behavior sequence, can be analyzed using a behavior recognition model or pose estimation model. The temporal behavior sequence contains behavior types corresponding to the timestamps of different image frames. For non-human objects, their object type can be identified, and their correlation with the behavior sequence of a specific person can be determined based on factors such as their spatial position and movement relationship with the person, thereby determining the target object type. After the human activity process detection is implemented, a scene activity process description can be generated based on the detection results. The specific form of generating the scene activity process description can be diverse.
[0019] In another embodiment, the structured information, consisting of the identified user identity, temporal behavior sequence, and associated target item type, is combined according to a predefined semantic template and directly output as a structured description of the scene activity process, such as JSON format, which contains explicit field identifiers for the corresponding information. In yet another embodiment, more advanced natural language processing techniques can be used to convert the aforementioned structured data into natural language prompts, which are then input into a pre-trained natural language generation model. This model outputs fluent natural language statements as a description of the scene activity process, such as "User A walks to the sofa in the living room and sits down." Both methods effectively transform visual perception information into semantic descriptions that can be understood and processed by subsequent strategy modules.
[0020] In another embodiment, the process of generating a scene activity description can be achieved using a mature graph-to-text (Graph-to-Text) model. This involves directly feeding a pre-trained multimodal large language model (Graph-to-Text) with consecutive image frames from a preview video stream, or a sequence of keyframes extracted after preprocessing. Such models have powerful built-in visual-language alignment and sequence understanding capabilities, enabling end-to-end analysis of the input image sequence. They do not require explicit step-by-step execution of intermediate tasks such as object detection, identity recognition, and behavior analysis; instead, they directly generate a coherent natural language description—the scene activity description—based on the image content.
[0021] Specifically, the graph-to-text model extracts depth features from image frames through its visual encoder and uses its language decoder to understand the temporal relationships and dynamic changes between frames, thereby comprehensively inferring a person's identity, behavioral intentions, interactions with scene objects, and the evolution of the entire activity. For example, after inputting image frames containing a series of actions such as a user walking in from the doorway, putting down an item, and walking towards the sofa, the model may directly output a natural language description such as "User A puts their briefcase on the entryway cabinet and then walks towards the living room sofa." This approach simplifies the technical process, reduces the complexity of integrating multiple independent models, and is particularly suitable for handling complex or long-tail activity scenarios not fully covered in the training data, demonstrating the intelligent advantages of high-level semantic understanding.
[0022] The generation of scene activity process descriptions has a wide range of applications. For example, in a family living room environment, it can detect the continuous movements of an elderly person slowly walking towards the sofa and sitting down; in a kitchen environment, it can identify the sequence of actions of a user approaching the refrigerator and reaching out to open the door; in the doorway area, it can capture a series of behaviors such as a visitor stopping, pressing the doorbell, and placing items. Accurate descriptions of these activity processes provide a solid information foundation for subsequent precise and personalized smart home control.
[0023] In some embodiments, the scene activity process description may further include a regional location. Since the smart device containing the camera unit typically has its regional location pre-set in an environmental map, this regional location can be directly added to the scene activity process description to enrich the semantic information provided by the description. It should be noted that the regional location is not mandatory, because under the application of natural language processing technology, the type of target object associated with the temporal behavior sequence can also provide corresponding semantic associations. For example, when a sofa image exists within a preset range around a person's image, the corresponding natural language model can understand that the person is in the living room.
[0024] In one embodiment, the detection of human activity can rely on the computing power of the local or network edge. However, in this process, it can be ensured that the original preview video stream is processed on the device side, and only the intermediate information generated after preprocessing the preview video stream, such as the user identity, temporal behavior sequence, target item type, keyframe sequence, etc. mentioned above, is submitted to the network edge for calculation. This maximizes the protection of user privacy and security, while achieving real-time or near real-time analysis and response through an efficient algorithm model.
[0025] Step S3200: Determine a matching centralized control strategy template based on the scene activity process description. The centralized control strategy template includes a strategy conversion rule that maps the scene activity process description to device control instructions adapted to at least one smart device in the smart home network. After obtaining the scenario activity process description, a matching centralized control strategy template can be determined based on this description. The centralized control strategy template is implemented as a predefined set of rules, containing one or more policy transformation rules that define the mapping logic from the scenario activity process description to specific device control commands. Determining the matching centralized control strategy template essentially involves matching the generated scenario activity process description against a large number of pre-stored centralized control strategy templates in the policy library to find the target centralized control strategy template whose conditions best fit the current scenario.
[0026] In one embodiment, the matching process is implemented through structured queries. In this approach, it is assumed that the scene activity description exists in the form of structured information, such as a JSON object containing fields like "User Identity: Father," "Core Behavior: Sit Down," "Associated Item: Sofa," and "Area: Living Room." Each centralized control policy template in the policy library also predefines rules with a similar structure. The matching operation compares the fields in the description with the template conditions. If all fields match the template conditions, the template is matched. For example, if the description information is completely consistent with a template whose conditions are "User identity is a family member, behavior is sitting down, area is living room sofa area," then that template is determined to be the matching centralized control policy template.
[0027] In another embodiment, the matching process can be based on natural language understanding and semantic similarity calculation. This approach is particularly suitable when the scene activity process is described as natural language text. The textual description of the scene activity process is input into a semantic similarity model (such as a BERT-based sentence embedding model) along with the natural language conditional descriptions of each template in the policy library (e.g., "when a family member sits down on the living room sofa"). The model calculates the semantic relevance between the current description and each template condition and selects the template with the highest semantic similarity as the matching result. This approach is highly flexible and can understand semantically similar but differently expressed descriptions.
[0028] Considering that in practical applications, situations may arise where precise matching or multiple templates may not be found, a more robust implementation employs a hybrid matching strategy. This strategy first attempts the structured query described above to obtain fast and accurate rule matching results. If the structured query fails to find a unique template, including no template matching or multiple templates matching leading to ambiguity, a backup plan is activated. For example, the current structured description of the scene activity process is converted into a natural language description, and semantic similarity matching is used as a fallback to ensure that the most suitable centralized control strategy template can always be found to respond to the scene.
[0029] Once a matching centralized control strategy template is determined, the strategy transition rules contained within it are activated. These rules specifically define the smart devices to be controlled, the target states the devices should achieve, and the logical or temporal relationships between instructions. For example, for a scenario described as "father sitting down on the living room sofa," the matching template's strategy transition rules might specifically stipulate: set the smart TV's power-on status to "on," adjust the main lighting brightness to 30%, and switch the air conditioner to cinema mode. This rule serves as the blueprint for generating subsequent specific control instructions.
[0030] Step S3300: Generate corresponding device control instructions according to the policy conversion rules in the centralized control policy template, and control the corresponding smart devices to respond.
[0031] After identifying a matching centralized control strategy template and activating its contained strategy conversion rules, corresponding device control commands can be generated based on these rules, and the corresponding smart devices can be controlled to respond. The strategy conversion rules, acting as a bridge between scene semantics and device operation, specifically define the control logic required to achieve specific scene effects, including the identifier of the target smart device, the target state to be achieved by each smart device, and the execution logic and timing relationships between multiple control commands.
[0032] In one embodiment, the process of generating device control commands involves the direct parsing and conversion of policy conversion rules. First, the rule content is parsed to identify one or more smart devices to be controlled, and the specific control parameters, i.e., target states, for each device are defined. For example, for a policy conversion rule, the parsed result may identify target devices such as a smart TV, a main light, and an air conditioner, with target states of power on, brightness adjusted to 30%, and switching to cinema mode, respectively. Subsequently, based on the communication protocols and application programming interfaces supported by the smart devices, each target state is converted into standardized control commands that the device can recognize and execute. These commands can be generated based on Wi-Fi, Zigbee, Bluetooth, or vendor-specific IoT protocols.
[0033] When controlling multiple smart devices to coordinate responses to a complex scenario, the pre-defined timing relationships in the policy transition rules can play a scheduling role. In one embodiment, after generating all device control commands, instead of simply issuing all commands simultaneously, scheduling is performed according to the timing or dependency relationships defined in the policy transition rules. For example, for a "cinema mode" scenario, the rules might specify closing the curtains first, then dimming the lights, and finally turning on the TV and speakers. The smart central control device strictly follows this order to send control commands to the corresponding devices sequentially, ensuring smooth and accurate scene switching. In another embodiment, the timing relationships can be more complex, for example, defining that some commands can be executed in parallel, while others can only be triggered after other commands have been successfully executed. The smart central control device then needs a corresponding command scheduler to manage these dependencies.
[0034] Another embodiment focuses on handling complex task flows involving multiple rounds of interaction or continuous actions. In this case, the policy transition rule can be parsed into a task flow consisting of multiple ordered subtasks. The intelligent central control device generates corresponding device control instructions for each subtask sequentially and schedules their execution according to the dependencies between subtasks. For example, for the "delivery" scenario, the corresponding policy transition rule can be parsed into three subtasks: first, the first device control instruction controls the smart speaker to announce "Please leave the package at the door"; then, the second device control instruction waits for a preset time or visually confirms that the item has been placed; finally, the third device control instruction controls the smart speaker to announce "Thank you". The intelligent central control device generates and executes the instructions for each subtask sequentially, and the triggering of subsequent subtasks strictly depends on the execution result of the previous subtask or the fulfillment of specific conditions.
[0035] The generated device control commands are sent to the corresponding smart devices via the smart home network. Upon receiving the commands, the smart devices execute the corresponding operations, thus collectively completing the scene response defined by the centralized control strategy template. The ultimate goal of the entire process is to achieve a deep integration between the smart home system and the user's activities, providing a highly personalized, scenario-based, and seamless automated experience.
[0036] As can be seen from the above embodiments, this application achieves a paradigm shift in smart home control from passive response to proactive understanding by deeply integrating visual perception with intelligent decision-making, gaining multiple technical advantages, including but not limited to: First, it significantly enhances the depth and accuracy of environmental perception. Traditional sensors can only provide binary signals of presence or absence, while this application, by analyzing continuous video streams, can accurately resolve the specific identities of people, continuous behavioral sequences, and their interactions with objects in the scene, generating semantically rich scene descriptions. This deep understanding of "who does what" provides an irreplaceable data foundation for truly personalized services, enabling smart home systems to distinguish different family members and respond to their unique preferences, achieving a fundamental shift from controlling space to serving people.
[0037] Secondly, regarding decision-making intelligence, this application constructs a highly flexible and scalable rule mapping mechanism by introducing a centralized control strategy template and the strategy conversion rules it contains. It no longer relies on simple, hard-coded trigger conditions, but instead maps complex scenario descriptions into a series of coordinated device control commands. In particular, it supports command scheduling based on timing and dependencies, enabling it to handle complex scenarios with multiple steps and timing requirements, such as activating "cinema mode" or "delivery reception," achieving smooth and natural collaborative responses between device groups and significantly improving the consistency and intelligence of the user experience.
[0038] Furthermore, this application demonstrates exceptional cost-effectiveness and scalability. Utilizing a general-purpose camera unit as the key sensing source replaces the need to deploy multiple dedicated sensors, effectively reducing hardware costs and system complexity. Simultaneously, the software-defined policy template management control logic allows for the addition of new scenes or modification of rules without altering the hardware infrastructure; only the template library needs to be updated, greatly enhancing flexibility and maintainability. This architecture enables smart home systems to iterate and upgrade at a lower cost, continuously adapting to new needs and scenarios.
[0039] Furthermore, regarding privacy and security, this solution emphasizes localized data processing. Key tasks such as visual data analysis and scene description generation can be completed locally on the smart central control device or edge gateway, eliminating the need to upload raw video streams to the cloud. This edge-processing mechanism significantly reduces the risk of exposing user privacy data, complies with increasingly stringent data security regulations, and provides crucial trust guarantees for the deployment of technology in sensitive environments such as homes.
[0040] Finally, the technical approach of this application possesses strong cutting-edge integration capabilities, seamlessly integrating advanced computer vision models and large-scale language models. Whether it utilizes mature graph-to-text technology to simplify processes or employs semantic similarity matching to enhance flexibility, it demonstrates excellent adaptability to future technological developments. This openness ensures that the smart home system can continuously absorb the latest achievements in the field of artificial intelligence, constantly improve its perception and understanding capabilities, and maintain long-term technological advancement.
[0041] Based on any embodiment of the method in this application, the detection of human activity is performed on the preview video stream generated by the camera unit, and a scene activity description corresponding to the human activity in the preview video stream in its scene is determined, including: Step S3110: Perform image target recognition on the preview video stream to obtain multiple image targets and their types; Image target recognition of the preview video stream captured by the camera unit can be implemented using a deep learning model to obtain multiple image targets and their types. The purpose of image target recognition is to automatically detect objects or regions of interest in the image frames of the preview video stream, and to locate and classify them.
[0042] In one embodiment, image target recognition can rely on deep learning models pre-trained on large-scale datasets. These models can process each frame of input video image end-to-end, directly outputting the bounding box coordinates of each object in the image and its target type. Available model architectures include, but are not limited to, the YOLO series, SSD, and Faster R-CNN. These models extract deep features from the image through their convolutional neural network backbone and generate candidate target regions using region proposal networks or similar mechanisms. Finally, a classifier determines the target type of objects within each region. The scope of target types can be defined according to the actual application requirements, typically including common object categories in a home environment such as people, vehicles, furniture, appliances, and pets. For example, in a frame of an image in a living room scene, the model might identify two image targets: one target type is classified as "people," and the other target type is classified as "sofa."
[0043] Processing the preview video stream can involve independent target recognition for each frame, or a time-window-based sampling method can be used, such as sampling one frame for recognition every fixed number of frames or at fixed time intervals, to balance processing efficiency and real-time requirements. Recognition within consecutive frames helps to stably track the target in subsequent steps.
[0044] In another embodiment, to improve the recognition accuracy of specific target categories, especially people, a cascaded or specialized recognition strategy can be adopted. For example, a general object detection model can be used first to quickly locate all possible target regions in the image. Then, for regions initially identified as "people," a specially optimized human detection model can be used for secondary verification and fine-tuning to reduce false positives and false negatives. Similarly, dedicated detectors can be equipped for other critical objects, such as televisions and refrigerators.
[0045] The output of the image target recognition process, namely multiple image targets and their types, lays a solid foundation for subsequent tracking, identity recognition, behavior analysis, and object association. It provides an initial structured understanding of the scene composition and is the first step in transforming raw pixel data into high-level semantic information.
[0046] Step S3120: When the target type includes a person type, track the corresponding person target, identify its user identity and analyze its temporal behavior sequence. After completing image target recognition and determining that the target type is a person, the corresponding person target can be tracked to identify the user and analyze their temporal behavior sequence. Person target tracking aims to assign a unique identifier to a specific person target in a video sequence and continuously locate the target in subsequent frames to form its motion trajectory.
[0047] In one embodiment, tracking can rely on a combination of appearance feature matching and motion model prediction. For example, appearance features of a person target, such as color histograms or depth feature embeddings, can be extracted, and feature similarity matching can be performed between adjacent frames to associate the same target. Simultaneously, Kalman filtering or more complex motion models can be combined to predict the target's position in the next frame, helping to solve the problem of re-identification after occlusion or brief disappearance. Commonly used multi-target tracking algorithms, including SORT and DeepSORT, can be selected.
[0048] Based on successful tracking of the target person, their user identity can be further identified. In one embodiment, user identification is achieved through biometrics. For example, when the person's face is clearly visible, a high-quality image of the face region can be captured, a pre-trained face recognition model can be used to extract facial feature vectors, and these feature vectors can be compared with a pre-registered family member face feature database. The best-matching user identity can be determined by calculating cosine similarity or Euclidean distance. In another embodiment, if facial information is unavailable, a gait recognition model can be used to identify the user by analyzing dynamic features such as the person's walking posture and rhythm.
[0049] Simultaneously, the temporal behavior sequence of the target person can also be analyzed. A temporal behavior sequence refers to the temporal combination of a series of actions performed by the target person over a period of time, i.e., multiple image frames. In one embodiment, the analysis process can be implemented using a time-series-based model. For example, multiple image frames can be sampled continuously or at certain intervals, and the pose of the person in each frame can be estimated to obtain the coordinate sequence of their body key points. This series of pose changes is then input into a preset temporal model, such as a Long Short-Term Memory network or a temporal convolutional network, which classifies or segments continuous behavior segments, such as walking, sitting, standing, waving, etc. In another embodiment, a 3D convolutional neural network or a video Transformer model can be directly used to perform end-to-end analysis on the cropped video segments of the target person, directly outputting their behavior category or the time interval of the behavior. The analysis of temporal behavior sequences enables the system to understand the user's ongoing activities, rather than just instantaneous poses.
[0050] Through the collaborative work of tracking, identity recognition, and behavioral sequence analysis, we can accurately grasp the dynamic activities of a specific user in a scene. For example, we can identify the target being tracked as user Zhang San and analyze his complete temporal behavioral sequence, from entering through the doorway, pausing at the dining table, to walking towards the sofa. This information provides key elements for subsequently generating semantically rich descriptions of the scene's activities.
[0051] Step S3130: When the target type includes an item type, determine the target item type that is associated with the temporal behavior sequence; After identifying image targets classified as objects in the image target recognition step, it is necessary to determine the types of target objects that are related to the temporal behavior sequence of the person being analyzed. This is to establish semantic connections between the person's activities and related objects in their environment, thereby gaining a more complete understanding of the context in which the activities occur. The determination of these connections can go beyond simply listing all identified objects; instead, it can be based on spatial, temporal, and semantic cues to filter out specific objects closely related to the current person's behavior, thus integrating isolated object information into the dynamic narrative of the activity. In one embodiment, the association is primarily determined through spatial location relationships. Accordingly, the relative distance and orientation between the movement trajectory or current position of the tracked person and the image coordinates of each identified object can be calculated. For example, when a sequence of user hand movements is detected, and the endpoint of the hand's movement trajectory points to an object identified as a cup in the image space, it can be determined that the cup is associated with the user's drinking behavior, thus identifying it as the target object type. In another scenario, if a person continuously approaches an object identified as a television set in an image frame sequence and eventually comes to a stop in front of it, the television set can be determined to be associated with the user's viewing behavior.
[0052] In another embodiment, the determination of the association can be based on the functional attributes of the item and the intention of the user's behavior. To this end, a knowledge base mapping item types to their common uses or functions can be established in advance. When a specific user behavior sequence is identified, items that may be related to supporting or completing that behavior are retrieved. For example, if a user is identified to walk towards an area and sit down, and an item classified as a sofa is identified in that area, even if the sitting action itself has not been fully performed, common sense can be used to infer that the sofa is an associated item for this behavior, thus identifying it as the target item type.
[0053] In other embodiments, temporal co-occurrence relationships can also serve as an auxiliary basis for judgment. If an item continuously appears near a person throughout the entire time period during which a person performs a specific sequence of actions, rather than appearing briefly, then the item is more likely to be associated with the sequence of actions. For example, in a series of actions by a user—picking up a phone, operating the phone, and putting the phone down—the phone is highly correlated with the user's hand movements throughout the entire time period.
[0054] As can be seen, by choosing one or flexibly combining the above embodiments, the target item type associated with the temporal behavior sequence can be finally determined. The process of determining the target item type is essentially the process of integrating isolated visual detection results into meaningful scene semantics, which provides key contextual information for the final generation of high-quality scene activity process descriptions. This allows the scene activity process descriptions to not only include user identity and behavior, but also reflect the objects they interact with, such as forming a complete description like user A picking up a water glass on the table or user B walking towards the sofa in the living room.
[0055] Step S3140: Based on the user identity, the temporal behavior sequence, and the determined target item type, generate the scene activity process description.
[0056] After obtaining information elements such as user identity, temporal behavior sequence, and the determined target item type, a scene activity process description is generated to integrate the aforementioned discrete and structured recognition results into a coherent and semantically rich overall description, thereby providing a clear and easy-to-understand context for subsequent strategy matching.
[0057] In one embodiment, the scene activity process description is generated in the form of a structured data object. In this approach, elements such as user identity, temporal behavior sequences, and target item types are directly mapped to data fields in a predefined format. For example, a JSON object can be generated containing key-value pairs such as "user_identity": "father", "action": "sitting_down", "related_object": "sofa", and "location": "living_room". This form of description has excellent machine readability, facilitates subsequent querying of the strategy library through precise field matching, and offers high processing efficiency and a standardized format.
[0058] In another embodiment, natural language text is used to describe the scene's activities. This approach uses natural language generation technology to convert structured information elements into fluent, human-readable sentences. Specifically, user identity, temporal behavior sequences, and target item types can be filled into a pre-defined semantic template, such as using a template sentence structure like "<user identity> is performing <behavior> in <area>, the object is <target item type>" as the framework. A more advanced approach utilizes a pre-trained large language model, inputting the above information as prompt words, and the model automatically generates natural language sentences that conform to grammatical and semantic logic, such as outputting "Father is sitting down on the sofa in the living room." This descriptive form is intuitive and easy to understand, especially advantageous when user interaction or log recording is required.
[0059] When generating descriptions, the processing of temporal behavioral sequences can reflect different levels of granularity. One approach is to extract the most representative or ultimately stable core behavior from the entire sequence as the focus of the description. For example, the sequence of walking towards the sofa and then sitting down can be simplified to a description of "sitting down." Another approach strives to retain more process details, generating a more dynamic description, such as "the user walks in from the doorway, places their briefcase on the cabinet, then walks to the sofa and sits down," which more completely reflects the entire process of the activity.
[0060] The above embodiments, by constructing a multi-layered, progressive information processing pipeline, gradually refine the raw visual data into a semantically rich scene description, demonstrating significant positive technical effects in deepening the inventiveness of this application, as embodied in: First, the above embodiments significantly enhance the depth and structure of intelligent perception. Through a series of refined processing steps, such as image target recognition, target tracking, identity recognition, behavior sequence analysis, and object association, a leap from pixel-level information to high-level semantic understanding is achieved. This enables the accurate parsing of the identities of people in a scene, continuous behavioral intentions, and objects associated with those behaviors, thereby generating a complete scene description such as "User A picks up the water glass on the table," rather than the rudimentary signal of "motion detected" provided by traditional systems.
[0061] Secondly, it demonstrates a high degree of technological integration and process optimization. By organically integrating multiple tasks in the field of computer vision, including object detection, multi-object tracking, face recognition, pose estimation, and behavior recognition, into a coherent automated process, it ensures smooth data flow from the raw video stream to the final scene description, avoiding information loss and errors during transmission between different modules.
[0062] Furthermore, regarding the feasibility of understanding complex scenes, by introducing temporal behavioral sequences and object associations, the system successfully captures the dynamism and context of activities. It is no longer limited to analyzing static poses in a single image, but can understand the complete process constituted by a series of actions and identify key objects related to the behavior. This allows the smart home system to distinguish between two distinctly different behaviors: "the user walks towards the refrigerator" and "the user walks towards the television," even though the initial walking movements may be similar. This deep understanding of the activity context is key to achieving accurate and meaningful automated responses.
[0063] Finally, the above embodiments provide solid support for the system's flexibility and scalability. Generating both structured data formats (such as JSON) and natural language text for scenario descriptions provides diverse interfaces for the subsequent policy matching module. Structured data facilitates efficient and accurate rule matching, while natural language descriptions are easier to integrate with advanced decision-making systems based on large language models. This design allows the system to select the most suitable description method according to the needs of different scenarios and the technical characteristics of subsequent modules, demonstrating good architectural adaptability and future evolution potential.
[0064] Based on any embodiment of the method in this application, the scene activity process description is generated based on the user identity, the temporal behavior sequence, and the determined target item type, including: Step S3141: Integrate the user identity, the temporal behavior sequence, and the target item type into a structured data tuple; The purpose of this step is to integrate the discrete information elements identified in the previous steps, including user identity (such as "father"), core behaviors extracted from the temporal behavior sequence (such as "sit down"), and associated target item types (such as "sofa"), into a data structure with a clear structure and uniform format, namely, a structured data tuple.
[0065] This data tuple can be represented as a dictionary, a JSON object, or a feature vector. For example, it can generate key-value pair combinations such as {"user_identity": "father", "action": "sitting_down", "related_object": "sofa"}. This process transforms information from scattered to centralized, and from heterogeneous to homogeneous, providing a machine-readable and unambiguous data foundation for subsequent automated processing, ensuring the accuracy and efficiency of information transmission.
[0066] Step S3142: Based on the preset prompt word template, convert the data tuple into natural language prompts; Furthermore, the structured data, i.e., the data tuples, from the previous steps are converted into an input format that the natural language generation model can understand. A pre-defined prompt template can be a text string containing specific placeholders. The values of each field in the data tuple are then filled into the corresponding placeholders in the template. For example, a simple prompt template might be: "Describe the following scenario: the user is <user identity>, the behavior is <time-sequence behavior sequence>, and the related object is <target item type>." After filling in the blanks, a specific natural language prompt is obtained: "Describe the following scenario: the user is the father, the behavior is sitting down, and the related object is the sofa." This conversion process cleverly bridges the gap between structured data and natural language generation, providing the model with clear task guidance and contextual information.
[0067] Step S3143: Input the natural language prompt into a pre-trained natural language generation model, and have the model output a description of the scene activity process in natural language format.
[0068] The natural language prompts generated in the previous step are input into a pre-trained natural language generation model (such as the GPT series, T5, and other large-scale language models). Based on its language knowledge and reasoning abilities gained through training on massive amounts of text data, this model understands, integrates, and reinterprets the prompts, outputting a fluent and accurate description of the scene's activities that conforms to human language habits. For example, the model might not simply repeat the prompt, but instead generate a more natural sentence: "Father sat down on the sofa." This approach can handle complex logical relationships, generating coherent descriptions that include temporal and causal information. Its output is intuitive and easy to read, greatly improving interpretability and user-friendliness.
[0069] The above embodiments creatively transform cold visual recognition data into semantically rich and context-sensitive natural language descriptions through a progressive process. Its positive technical effect lies not only in upgrading the form of information representation, but also in significantly improving the flexibility, accuracy, and naturalness of scene descriptions by utilizing advanced large language models. This enables smart home systems to understand and express user activities in a way that more closely resembles human cognition, providing a crucial semantic understanding foundation for achieving highly intelligent and user-friendly device control.
[0070] Based on any embodiment of the method in this application, determining a matching centralized control strategy template based on the scenario activity process description includes: Step S3210: Match the structured information contained in the scenario activity process description with the predefined structured rule conditions of the centralized control strategy template in the strategy library to determine the centralized control strategy template that achieves the match as the matching result. After obtaining the scenario activity process description, a matching operation can be performed, which compares the structured information contained in the description with the predefined structured rule conditions of the centralized control policy templates in the policy library. The structured information here typically refers to machine-readable data organized in the form of key-value pairs, fields, or specific data objects, such as a JSON object containing fields like user identity, behavior, and associated item type. The policy library can be viewed as a database storing a large number of predefined centralized control policy templates. Each template contains its triggering conditions, i.e., predefined structured rule conditions, which also define a series of fields and their values in a similar structured form.
[0071] In one embodiment, the matching process is achieved through database queries or precise matching by a rule engine. Specifically, the structured information describing the scene activity process is treated as a query condition and compared field-by-field with the rule conditions of each template in the strategy library. The rule conditions can be designed as strict equality matching, such as requiring the user identity field to be exactly equal to family member and the behavior field to be equal to sit down. It can also support certain range matching or logical operations, such as requiring the duration field to be greater than a certain threshold. When the information in the scene activity process description completely satisfies all the rule conditions of a template, the template is considered to have been matched, and it is determined as the matching result. For example, if the scene description is {user identity: father, behavior: sit down, associated item: sofa}, and the rule conditions of a template in the strategy library are defined as user identity ∈ [family member] AND core behavior = sit down AND associated item = sofa, then the template is successfully matched.
[0072] In another embodiment, the matching process can incorporate a fuzzy matching or priority mechanism. For example, when the information described in the scene activity partially matches the rule conditions of multiple templates or when multiple similar templates exist, a matching score can be calculated. The matching score can be calculated comprehensively based on factors such as the number of matched fields and the importance weight of the fields. Finally, the template with the highest matching score is selected as the matching result. This approach can handle some scenarios with imprecise matching and improve robustness.
[0073] Step S3220: If no unique centralized control strategy template is matched, the structured information is converted into a natural language description through a natural language generation model. If a unique centralized control strategy template cannot be matched through the preceding structured query, i.e., if no template is matched or multiple templates are matched, leading to decision ambiguity, the aforementioned structured information can be converted into a natural language description through a natural language generation model, providing an alternative decision path based on semantic understanding.
[0074] In one embodiment, the conversion process is implemented through a pre-trained natural language generation model. Specifically, the structured information contained in the scene activity description, such as user identity, temporal behavior sequence, and associated item type organized in key-value pairs, is used as input data and filled into a preset prompt template. This template is designed to guide the model to generate grammatically correct statements. For example, for the structured information {User identity: Visitor, Behavior: Standing at the door, Associated item: Package}, after filling in the template "Describe the following scenario: The user is <User identity>, the behavior is <Core behavior>, and the related object is <Associated item type>", the natural language prompt "Describe the following scenario: The user is a visitor, the behavior is standing at the door, and the related object is a package" is generated. Subsequently, this prompt is input into the natural language generation model, which, based on its language generation capabilities, outputs a fluent natural language description, such as "A visitor is standing at the door, holding a package."
[0075] In another embodiment, to improve the quality and context relevance of the generated description, the conversion process can incorporate more refined prompting engineering. For example, additional guidance on the scene context or desired description style can be added to the prompt word template, such as requiring the model to generate a "concise description of a security monitoring scene." Based on this, the natural language generation model can output more targeted text, such as "A strange visitor was detected lingering in front of the door for a long time and holding an item"; or when the model is required to generate "a specified activity duration," it can generate "A strange visitor was detected lingering in front of the door for more than 10 seconds and holding an item."
[0076] Step S3230: Perform semantic matching between the natural language description and the natural language conditions of the centralized control strategy templates in the strategy library, and select the centralized control strategy template with the highest semantic similarity as the matching result.
[0077] After obtaining the natural language description of the scene activity process through a natural language generation model, a semantic matching operation is performed, which compares this natural language description with the natural language conditions corresponding to each centralized control policy template in the policy library. Here, natural language conditions refer to scene description statements predefined in human-readable text form, used to trigger the corresponding centralized control policy template. The purpose of semantic matching is to measure the degree of semantic similarity between the currently generated description and the template conditions, rather than simple keyword matching.
[0078] In one embodiment, semantic matching is achieved by calculating the similarity between text embedding vectors. Specifically, a pre-trained language model, such as Sentence-BERT or a similar model, can be used to convert the current natural language description and the natural language conditions of each template in the policy library into vector representations in a high-dimensional space, i.e., text embedding vectors. These vectors can capture the deep semantic information of the text. Subsequently, the semantic similarity between the current description vector and each template condition vector is quantified by calculating the cosine similarity or Euclidean distance. Finally, the centralized policy template with the highest semantic similarity to the current natural language description is selected as the matching result. For example, if the current description is a visitor lingering at the door holding a package, and the natural language condition of a template in the policy library is a stranger lingering at the door for a long time, even if the literal expressions are not exactly the same, the semantics are highly related, and its similarity score will be high, thus it will be selected.
[0079] After selecting the control strategy template with the highest semantic similarity, the template is determined as the matching result for the current scenario, and the strategy conversion rules contained within it will be activated to guide the generation of subsequent device control commands.
[0080] The fallback mechanism constructed in the above embodiments automatically activates alternative paths based on natural language generation and semantic understanding when precise matching based on structured information fails to hit the unique template. This enables the application to exhibit excellent robustness and intelligence when facing complex, ambiguous, or long-tail scenarios where training data is insufficient. This mechanism effectively overcomes the rigidity problem of traditional rule systems that rely on hard-coded conditions. By converting structured information into more expressive natural language descriptions and using advanced semantic similarity calculations for matching, it significantly improves the tolerance for diverse scenario descriptions. For example, even if the rule "visitor lingers at the door" fails to match precisely, semantic understanding can successfully associate it with the strategy template "stranger lingers at the door for a long time," thus ensuring that reasonable and effective control decisions are generated in most cases. This not only significantly reduces the risk of failure due to an incomplete rule base but also fundamentally enhances the ability of smart home systems to adapt to the complexity and uncertainty of the real world, providing a key technological guarantee for achieving the leap from mechanical execution to intelligent understanding.
[0081] Based on any embodiment of the method in this application, before determining the matching centralized control strategy template based on the scenario activity process description, the method includes: Step S2100: Determine the corresponding user permission level based on the user identity carried in the scenario activity process description; Before performing centralized control policy template matching based on the scenario activity process description, a preliminary permission determination and policy library selection process can be executed first. This process begins with parsing the user identity carried in the scenario activity process description, with the aim of determining the user permission level corresponding to that user identity. The user permission level is a predefined access control label used to distinguish the scope of device control and scenario types that different user groups are allowed to trigger.
[0082] In one embodiment, the permission levels can be simply divided into family member permissions and visitor permissions. For example, when the parsed user identity is a pre-registered father or child, their permission level can be determined to be family member permissions; while when a stranger is identified or the user identity recognition fails, their permission level can be determined to be visitor permissions.
[0083] In another, more refined embodiment, the permission levels can be further subdivided, such as administrator permissions (configurable system), ordinary user permissions (can use all functions), and restricted permissions (can only use basic functions), thereby achieving tiered management of control permissions.
[0084] Step S2200: Select a target policy library from multiple policy libraries according to the user permission level; After determining the user's permission level, a target policy library is selected from multiple preset policy libraries based on that level for the current matching operation. A policy library is a collection of centralized control policy templates. By establishing multiple independent policy libraries and binding different policy libraries to different user permission levels, physical or logical isolation of control rules is achieved.
[0085] In one embodiment, two policy libraries can be set up: an advanced policy library bound to family member permissions, containing rich personalized scene templates such as cinema mode and reading mode, allowing control of various devices such as TVs, lights, and air conditioners; and a basic policy library bound to visitor permissions, containing only limited templates focused on security and basic convenience, such as simply turning on corridor lighting or triggering doorbell notifications, without involving control of privacy-sensitive or highly personalized devices. This design ensures that users with different permissions can only access the set of control rules they are authorized to use.
[0086] Step S2300: In the target policy library, match the centralized control policy template to the scenario activity process description.
[0087] Finally, within the selected target policy library, a specific centralized control policy template is matched to the activity process description for this scenario. The matching logic in this step is consistent with that described in the previous embodiments. It can employ precise matching based on structured rules within the target policy library, or, in case of ambiguity, enable semantic matching based on natural language description as a fallback solution. For example, for users with family member permissions, the target policy library will search for a centralized control policy template that matches the description of "father sitting on the sofa" and enables cinema mode. For users with only visitor permissions, the target policy library will only search for a template that matches "visitor standing at the door" and may only trigger the door light to turn on or send a notification.
[0088] The permission-based policy library routing matching mechanism constructed in the above embodiments brings significant benefits to this application. By strongly associating user identity with permission level and permission level with policy library, access control is implemented at the source of policy template matching, greatly improving security and privacy protection capabilities and effectively preventing unauthorized operations. Simultaneously, this mechanism naturally supports differentiated personalized services, providing rich and thoughtful scenarios for family members while maintaining a friendly yet restrained response to visitors. This enables the smart home system to adapt more intelligently and securely to the actual needs of multiple users coexisting in a home environment, embodying the essence of user-centricity.
[0089] Based on any embodiment of the method in this application, a corresponding device control command is generated according to the policy conversion rule in the centralized control policy template to control the corresponding smart device to respond, including: Step S3411: Parse the policy conversion rules to determine the multiple smart devices to be controlled and their corresponding target states; After identifying a matching centralized control strategy template, the strategy transition rules contained within that template can be parsed. Essentially, a strategy transition rule is a predefined, machine-readable set of instructions that details the specific control logic required to respond to a particular scenario. Therefore, the purpose of the parsing process is to extract key actionable information from these rules; the primary task is to determine the intelligent devices to be controlled and the target states that each device needs to achieve.
[0090] In one embodiment, the policy transformation rules can be represented as a structured data format, such as JSON or XML. During parsing, the list of device identifiers defined in the rules is first identified. These identifiers uniquely correspond to specific devices in the smart home network, such as "living_room_tv", "main_ceiling_light", and "air_conditioner_1". Next, for each device identifier in the list, the control parameters associated with it, i.e., the target state, are parsed. The target state is a precise description of the desired operating state of the device, and its form depends on the device type. For example, for a smart light, the target state might include on / off status, brightness value (0-100%), and color temperature (2700K-6500K); for an air conditioner, the target state might include operating mode (cooling / heating / fan), set temperature (e.g., 25℃), and fan speed (low / medium / high); for a smart TV, the target state might include power-on status, input source (HDMI 1), and volume level.
[0091] In another embodiment, the policy transition rules may be expressed in a more logical and dynamic manner. The parsing process needs not only to identify static devices and states, but also to understand the conditional logic within the rules. For example, a rule might specify: if the ambient light sensor value is below a certain threshold, then control smart light A and the target state to 50% brightness; otherwise, control smart light B and the target state to 30% brightness. In this case, the parser needs the ability to evaluate simple conditional expressions to dynamically determine the final device to be controlled and its target state based on the real-time context.
[0092] By parsing the policy transition rules, the abstract scenario intent is translated into a series of specific device control tasks. Each task clearly defines which device to control and to what state it should be controlled.
[0093] Step S3412: For each smart device, generate a device control command to drive it to reach the target state; After parsing the policy transition rules and identifying the smart devices to be controlled and their target states, device control instructions are generated. These instructions translate the abstract target state into standardized commands that the specific smart device can recognize and execute, conforming to its communication protocol and application programming interface specifications. Each device control command aims to drive a specific smart device from its current state to the target state desired by the policy.
[0094] In one embodiment, generating device control commands relies on a device driver library or protocol adaptation layer. The smart central control device maintains a device capability and command mapping library, which stores the unique identifier, device type, and supported control command sets and parameter formats for each smart device in the home. When a command needs to be generated for a device, the library is queried based on its device identifier to find the corresponding command template, and then the specific parameter values for the target state are filled into the template. For example, for a smart TV identified as living_room_tv, its target states are power on, switching to HDMI 1 input source, and adjusting the volume to 40. Querying the mapping library reveals that the TV supports a RESTful API based on the HTTP protocol; its power-on command template is PUT / devices / {device_id} / power {“state”: “on”}, and its input source switching command template is PUT / devices / {device_id} / input {“source”: “hdmi1”}. By filling the device identifier and target state parameters into the corresponding templates, two specific device control commands that can be directly processed by the TV are generated.
[0095] In another embodiment, considering the diversity of communication protocols in the smart home ecosystem, the command generation process can involve protocol conversion. For example, the smart central control device uses a unified internal command format to represent the target state. When it is necessary to send commands to devices using different protocols, the corresponding protocol converter is invoked. For example, for a smart light using the Zigbee protocol, its target state is 50% brightness. First, an internally formatted command is generated, and then the Zigbee protocol converter converts it into a specific Zigbee cluster command, such as the Move to Level command of the ZCL Level Control cluster, and maps the brightness value of 50% to the corresponding Level parameter value. This approach shields the differences in underlying protocols, providing a unified control interface for upper-layer applications.
[0096] The generated device control commands need to be executed accurately by the target device. Therefore, the commands must include not only the operation type and parameters, but also the target device's accurate addressing information, such as its network IP address, Zigbee short address, or Bluetooth MAC address. This ensures that the device control commands can be precisely routed to the correct smart device over the home network. For example, the generated control command might ultimately be encapsulated as a JSON object with a structure containing {"device_id": "light_001", "command": "set_brightness", "params":{"brightness": 50}, "protocol": "zigbee"}, and the corresponding communication module would be responsible for sending it out.
[0097] Step S3413: The generated multiple device control commands are scheduled according to the preset collaborative timing relationship in the strategy conversion rule and sent to the corresponding smart devices to execute collaborative responses.
[0098] After generating device control commands for each smart device, these commands are organized and managed according to the pre-defined collaborative timing relationships in the policy conversion rules to ensure that multiple devices can collaboratively and orderly complete scene responses. The collaborative timing relationships here define the execution order, dependencies, and possible time intervals between commands, with the aim of avoiding device operation conflicts and creating a smooth and natural scene transition experience.
[0099] In one embodiment, the cooperative timing relationship is represented as a simple sequential execution list. The policy transition rules explicitly define the sequence of instruction delivery. For example, for a scenario of activating cinema mode, the rules might specify the instruction delivery order as follows: first, turn off the curtain motor; after the curtains are closed, dim the main lighting; and finally, turn on the TV and sound system. Strictly following this order, the next instruction is sent only after the previous instruction has been confirmed as successfully executed or after a fixed delay. This approach is logically clear, easy to implement, and suitable for most scenarios with clear sequential dependencies.
[0100] In another embodiment, the coordinated timing relationship may include parallel and conditional triggering logic. Rules may allow some independent device control commands to be sent in parallel to improve efficiency, while defining that certain commands must be triggered only after specific conditions are met. For example, in a homecoming scenario, commands to turn on the entryway light and start the air conditioner can be issued simultaneously. However, the rules may further stipulate that the command to start music playback must wait for confirmation from other sensors (such as a pressure pad) that the user has entered the living room area before being triggered. Implementing this scheduling can be achieved by providing a lightweight command scheduler that manages the command queue, listens for external events or device status feedback, and determines the timing of command transmission based on the rule logic.
[0101] When managing command transmission, the scheduler can also consider the reliability and real-time performance of network communication. In one embodiment, an asynchronous transmission and status confirmation mechanism is employed. After sending a command to the target device, the scheduler does not wait indefinitely for a response but sets a timeout period. If a successful execution confirmation is received from the device within the timeout period, the next command in the sequence is executed or subsequent conditions are triggered. If a timeout occurs or an error feedback is received, the scheduler can retry, skip the command, or trigger an exception handling process according to a preset strategy, thereby enhancing robustness in the face of network fluctuations or temporary device unresponsiveness.
[0102] Ultimately, all scheduled device control commands are reliably sent to the corresponding smart devices via the home LAN. Upon receiving the commands, the devices execute the corresponding operations, enabling the entire group of smart devices to collaboratively complete a complex and coherent scene response according to a preset coordination sequence. For example, perfectly executing a series of actions such as closing curtains, dimming lights, and turning on audio-visual equipment in sequence, presenting the user with a seamless cinematic experience, rather than multiple devices acting haphazardly at the same time.
[0103] The above embodiments contribute to the inventiveness of this application through a refined and automated control chain. Their positive technical effect lies in the lossless and reliable conversion of high-level scene semantic intent into precise and orderly collaborative actions of the underlying device group. This embodiment, by parsing policy conversion rules, deconstructs abstract scene requirements such as "cinema mode" into a set of target states for specific devices, achieving explicit and configurable control logic. Furthermore, by adapting to different device protocols to generate standardized instructions, it effectively solves the technical challenge of device heterogeneity in the smart home ecosystem, ensuring broad compatibility of control. Finally, by introducing an intelligent scheduling mechanism based on preset collaborative timing relationships, through fine-grained management of instruction sending order, parallelism, and triggering conditions, it fundamentally avoids device operation conflicts, creating a smooth, natural, and highly reliable multi-device collaborative experience. This complete technical path not only significantly improves the automation level and response accuracy of the smart home system but also ensures the stability of services in complex home environments, ultimately achieving a fundamental leap for the smart home system from executing simple commands to managing complex scenarios.
[0104] Based on any embodiment of the method in this application, a corresponding device control command is generated according to the policy conversion rule in the centralized control policy template to control the corresponding smart device to respond, including: Step S3431: Parse the policy conversion rule into a task flow containing multiple ordered subtasks, wherein the triggering of at least one subsequent subtask depends on the execution result of the previous subtask. In this embodiment, during the generation of device control instructions based on the policy conversion rules in the centralized control policy template, a task flow parsing approach is used as the starting point. The macro-control logic defined by the policy conversion rules is parsed into a task flow consisting of multiple discrete and ordered subtasks. This task flow can be viewed as a working blueprint describing a series of atomic operations and their execution logic required to complete the entire scene response. A key feature is that at least one subsequent subtask in the task flow is triggered, explicitly depending on the execution result or completion status of its preceding subtask, thus introducing sequential and conditional execution.
[0105] In one embodiment, the policy transition rule itself explicitly defines the dependencies between subtasks in a structured manner. The parsing process is similar to parsing a workflow or state machine. The rule can directly declare the sequence of subtasks and the triggering conditions between them. For example, in a package delivery scenario, the rule might be parsed into three subtasks: the first subtask is defined as controlling a smart speaker to play a prompt voice; the second subtask is defined as waiting for a preset time or confirming through visual analysis that the package has been placed; and the third subtask is defined as controlling the smart speaker to play a thank-you voice. The parser needs to identify these subtasks and understand that the second subtask depends on the completion of the first subtask, while the third subtask depends on the conditions achieved by the second subtask, such as timeout or successful visual confirmation.
[0106] In another embodiment, dependency determination can be more dynamic. Policy transition rules don't need to explicitly list all dependencies; instead, they contain event- or state-based conditional expressions. The parsing process needs to translate this conditional logic into explicit dependencies between subtasks. For example, a rule might specify: when a user is detected lying in bed, turn off the main light and turn on the night light; if the user is still in bed and inactive after 10 minutes, further turn off the TV. The parser needs to infer that the subtask of turning off the TV depends not only on the initial event of the user lying down but also on the successful execution of an intermediate subtask that continuously monitors the user's state for 10 minutes without any change in state. This parsing requires understanding the temporal and conditional semantics in the rules.
[0107] The parsed task flow provides a clear execution graph for subsequent instruction generation and scheduling. This ensures that complex scenario responses can be broken down into manageable and monitorable steps, with each step triggered by a clear logical basis. This avoids chaotic instruction sending and lays the foundation for accurate and reliable multi-turn interactions. For example, based on the parsed task flow, it is clear that the instruction to broadcast a thank-you message can only be executed after waiting for the voice prompt and confirming the package placement, thus forming a coherent and intelligent interaction process.
[0108] Step S3432: Generate the corresponding device control instructions for each subtask in sequence; After parsing the policy transformation rules into a task flow containing ordered subtasks, device control instructions for the corresponding smart devices are generated sequentially for each subtask in the task flow. This translates each atomic subtask into a low-level control command that can be recognized and executed by the specific smart device and conforms to its communication protocol specifications. Each subtask is typically associated with one or more specific smart devices and defines the specific operational goals that the device needs to achieve at that task node.
[0109] In one embodiment, the process of generating device control commands relies on a preset device command mapping library. This library stores the correspondence between the unique identifier of each smart device and the control command set and parameter format it supports. When a command needs to be generated for a subtask, the mapping library is queried based on the device identifier involved in the subtask to obtain the corresponding command template. Then, the operation parameters defined in the subtask are filled into the template to generate the specific device control command. For example, for a subtask defined as playing a specific voice prompt, the associated device is a smart speaker. Querying the mapping library reveals that this model of speaker supports playing text through a specific text-to-speech interface. Its command template can be a JSON structure containing the target device ID, the operation type as TTS playback, and the text content to be played as parameters. After filling in the specified prompt text from the subtask, an executable command is generated.
[0110] In another embodiment, the definition of a subtask can be more abstract, requiring a transformation step before it can be mapped to a specific device command. For example, a subtask might be described as creating a warm atmosphere, rather than directly specifying to dim a particular light. In this case, the command generation process could involve calling a rule engine or policy mapping layer to match the abstract subtask objective with a predefined device control policy, thereby resolving it into one or more specific device control commands. For example, creating a warm atmosphere might be resolved by adjusting the color temperature of the main living room light to 2700K and its brightness to 40%, and sending an on command to the fireplace light (if present). When generating device control commands, corresponding control commands need to be generated separately for the lighting device and the fireplace device.
[0111] The generated device control commands must be accurate and executable. The commands must contain sufficient information, including the target device's accurate network address or device identifier, the specific operation command, and the required parameters. For example, for a smart curtain motor, the command to control its closing needs to include the motor's device ID, the operation command as "close," and possible speed parameters. The command format must conform to the requirements of the device communication protocol, whether it's an HTTP-based RESTful API, MQTT messages, Zigbee cluster commands, or other proprietary protocols.
[0112] By generating corresponding device control commands for each subtask sequentially, the high-level task flow is transformed into a series of low-level operation commands that can be directly executed by the device. This ensures that the response logic of complex scenarios can be accurately decomposed and transmitted to the underlying device, laying a solid foundation for subsequent dependency-based command scheduling and reliable execution. For example, in a package delivery scenario, corresponding smart speaker control commands are generated sequentially for the three subtasks of broadcasting a prompt voice, waiting for confirmation, and broadcasting a thank-you voice, thus preparing all the necessary control elements for forming a coherent automated interaction process.
[0113] Step S3433: Based on the dependency relationship, the device control commands are scheduled to be sent to the smart device for execution in sequence.
[0114] After generating device control commands for all subtasks, the commands are sent to the corresponding smart devices for execution in sequence, based on the pre-defined dependencies between subtasks in the task flow. These dependencies define the execution order and triggering conditions of the subtasks, ensuring that the execution order strictly conforms to the logical requirements of the scenario, thereby achieving a complex automated process with multiple steps and conditional triggering.
[0115] In one embodiment, the scheduling process is implemented through a state machine or workflow engine. This engine maintains the current execution state of the task flow. When a subtask needs to be executed, the engine checks whether all its prerequisites have been met. For example, in a package delivery scenario, the task flow contains three subtasks: playing a prompt message, waiting for placement confirmation, and playing a thank-you message. The scheduler first executes the first subtask, i.e., sending a command to control the smart speaker to play the prompt message. After sending this command, the scheduler does not immediately execute the next task, but instead sets the task flow state to waiting for confirmation. Only when a signal indicating successful package placement or a timeout signal is received from the visual analysis module, indicating that the prerequisites have been met, does the scheduler trigger the execution of the second subtask, i.e., sending a command to control the speaker to play a thank-you message. This explicit event-triggered mechanism ensures that the progress of the task flow strictly depends on the execution results of the prerequisite tasks or the fulfillment of specific external conditions.
[0116] In another embodiment, the scheduling process can incorporate parallel execution and synchronization mechanisms. For multiple subtasks in the task flow that do not have dependencies, the scheduler can send their corresponding device control commands in parallel to improve overall response speed. For example, in a homecoming scenario, the subtasks of turning on the foyer lights and starting the air conditioner can be performed simultaneously. However, the scheduler needs to manage subsequent subtasks that may have dependencies. For example, the subtask of broadcasting a welcome message may need to wait for confirmation signals that the lights and air conditioner have been adjusted before it can be triggered. In this case, the scheduler needs to monitor the completion status of multiple parallel commands and trigger subsequent tasks only after all preconditions are met. This approach balances efficiency and logical correctness.
[0117] When sending instructions, the scheduler can also consider the reliability and fault tolerance of instruction execution. In one embodiment, an asynchronous sending mode with acknowledgment and retry mechanisms is adopted. After sending an instruction to the device, the scheduler starts a timer to wait for the device to return a confirmation of successful execution. If an acknowledgment is received within a preset time, the subtask is marked as completed, and an attempt is made to trigger its subsequent tasks. If no acknowledgment is received within the timeout period or an error feedback is received, the scheduler can retry a limited number of times according to a preset strategy. If the retry still fails, the scheduler can choose to skip the current subtask and record an error log, or trigger a backup subtask according to rules, thereby ensuring that the task flow will not be completely interrupted due to the temporary failure of a single device, thus enhancing robustness.
[0118] Ultimately, under the scheduler's management, all device control commands are sent sequentially to the corresponding smart devices in the home network according to the dependencies defined in the task flow. After the devices execute the commands, the entire group of smart devices will collaboratively complete a complex scenario response involving multiple steps and conditional judgments. For example, it perfectly achieves a coherent interaction of first broadcasting a prompt, waiting for user confirmation, and then broadcasting gratitude, rather than mechanically executing all operations simultaneously, thus demonstrating the system's advanced intelligence and adaptability to interacting with the real world.
[0119] The above embodiments, by introducing task flow parsing and a dependency-based intelligent scheduling mechanism, elevate complex scenario responses from simple parallel instructions to intelligent workflows with temporal logic and conditional judgments, thereby achieving a qualitative leap in the complexity, reliability, and interactivity of smart home automation. This embodiment, by parsing macro-strategies into ordered sub-task flows, can accurately decompose and manage the logical steps of complex scenarios such as multi-round interactions and phased environmental adjustments; by generating specific instructions for each sub-task, it ensures accurate mapping from control intent to device operation; finally, through dependency-driven scheduling execution, it can intelligently coordinate the order of instruction sending, strictly adhering to the logic of triggering subsequent actions only after preconditions are met. This not only avoids conflicts and chaos in device operation but also achieves an automated experience deeply integrated with real-world processes. This entire mechanism works together to enable this application to handle highly complex and dynamically changing scenarios, marking a significant evolution in smart home control from executing static commands to managing dynamic processes, and comprehensively improving the intelligence level of device nodes and the entire network itself within the smart home network.
[0120] Please see Figure 2 This invention provides a smart device control apparatus to meet one of the purposes of this application. It is a functional embodiment of the smart device control method of this application. The apparatus includes an activity detection module 3100, a template matching module 3200, and a generation control module 3300. The activity detection module 3100 is configured to detect human activity processes based on a preview video stream generated by a camera unit, and determine a scene activity process description corresponding to the human activity in the preview video stream within its current scene. The template matching module 3200 is configured to determine a matching centralized control strategy template based on the scene activity process description. The centralized control strategy template includes strategy conversion rules that map the scene activity process description to device control commands adapted to at least one smart device in a smart home network. The generation control module 3300 is configured to generate corresponding device control commands according to the strategy conversion rules in the centralized control strategy template, and control the corresponding smart device to respond.
[0121] Based on any embodiment of the device in this application, the activity detection module 3100 includes: a target detection module, configured to perform image target recognition on the preview video stream to obtain multiple image targets and their target types; a tracking and recognition module, configured to track the corresponding person target when the target type includes a person type, identify its user identity, and analyze its temporal behavior sequence; an item recognition module, configured to determine the target item type that is associated with the temporal behavior sequence when the target type includes an item type; and a description construction module, configured to generate a scene activity process description based on the user identity, the temporal behavior sequence, and the determined target item type.
[0122] Based on any embodiment of the device in this application, the description construction module includes: a data combination module, configured to integrate the user identity, the temporal behavior sequence, and the target item type into a structured data tuple; a prompt construction module, configured to convert the data tuple into a natural language prompt based on a preset prompt word template; and a description conversion module, configured to input the natural language prompt into a pre-trained natural language generation model, and have the model output a description of the scene activity process in natural language format.
[0123] Based on any embodiment of the device in this application, the template matching module 3200 includes: a rule matching module, configured to match the structured information contained in the scene activity process description with the predefined structured rule conditions of the centralized control strategy template in the strategy library to determine the centralized control strategy template that achieves the match as the matching result; a missing match conversion module, configured to convert the structured information into a natural language description through a natural language generation model if no unique centralized control strategy template is matched; and a semantic matching module, configured to perform semantic matching between the natural language description and the natural language conditions of the centralized control strategy template in the strategy library, and select the centralized control strategy template with the highest semantic similarity as the matching result.
[0124] Based on any embodiment of the device in this application, prior to the template matching module 3200, this device further includes: a level determination module, configured to determine the corresponding user permission level based on the user identity carried in the scenario activity process description; a level selection library module, configured to select a target policy library from multiple policy libraries according to the user permission level; and a matching application module, configured to match the centralized control policy template for the scenario activity process description in the target policy library.
[0125] Based on any embodiment of the device in this application, the generation control module 3300 includes: a parsing and determining module, configured to parse the policy conversion rule and determine multiple smart devices to be controlled and their corresponding target states; an instruction generation module, configured to generate a device control instruction for each smart device to drive it to reach the target state; and a scheduling and execution module, configured to schedule the generated multiple device control instructions according to the preset cooperative timing relationship in the policy conversion rule and send them to the corresponding smart devices to execute cooperative responses.
[0126] Based on any embodiment of the device in this application, the generation control module 3300 includes: a task decomposition module, configured to parse the strategy conversion rule into a task flow containing multiple ordered subtasks, wherein the triggering of at least one subsequent subtask depends on the execution result of the previous subtask; an instruction construction module, configured to generate corresponding device control instructions for the smart device for each subtask in sequence; and a streaming execution module, configured to schedule the device control instructions to be sent to the smart device for execution in sequence based on the dependency relationship.
[0127] To address the aforementioned technical problems, embodiments of this application also provide a computer device for implementing the intelligent central control device of this application. For example... Figure 3 The diagram shows the internal structure of a computer device. This computer device includes a processor, a computer-readable storage medium, a memory, a network interface, and various communication components connected via a system bus. The computer-readable storage medium stores an operating system, a database, and computer-readable instructions. The database may store control information sequences. When the computer-readable instructions are executed by the processor, they enable the processor to implement an intelligent device control method. The processor of this computer device provides computing and control capabilities, supporting the operation of the entire computer device. The memory of this computer device may store computer-readable instructions. When these computer-readable instructions are executed by the processor, they enable the processor to execute the intelligent device control method of this application. The network interface of this computer device is used for communication with a terminal. Those skilled in the art will understand that… Figure 3 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0128] In this embodiment, the processor is used to execute... Figure 2 The system defines the specific functions of each module and its sub-modules. The memory stores the program code and various data required to execute these modules or sub-modules. The network interface is used for data transmission between the user terminal and the server. In this embodiment, the memory stores the program code and data required to execute all modules / sub-modules in the intelligent device control device of this application. The server can call the server's program code and data to execute the functions of all sub-modules.
[0129] This application also provides a storage medium storing computer-readable instructions, which, when executed by one or more processors, cause the one or more processors to perform the steps of the smart device control method of any embodiment of this application.
[0130] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments of this application can be implemented by a computer program instructing related hardware. This computer program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the methods described above. The aforementioned storage medium can be a magnetic disk, optical disk, read-only memory (ROM), or random access memory (RAM), etc.
[0131] Those skilled in the art will understand that the steps, measures, and solutions in the various operations, methods, and processes discussed in this application can be alternated, modified, combined, or deleted. Furthermore, other steps, measures, and solutions in the various operations, methods, and processes discussed in this application can also be alternated, modified, rearranged, decomposed, combined, or deleted. Furthermore, steps, measures, and solutions in the prior art that are similar to those in the open-source operations, methods, and processes of this application can also be alternated, modified, rearranged, decomposed, combined, or deleted.
[0132] The above description is only a partial embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.
Claims
1. A method for controlling an intelligent device, characterized in that, include: Based on the preview video stream generated by the camera unit, the process of detecting human activity is performed, and the scene activity process description corresponding to the human activity in the preview video stream in the scene is determined. Based on the scene activity process description, a matching centralized control strategy template is determined. The centralized control strategy template includes a strategy conversion rule that maps the scene activity process description to device control instructions adapted to at least one smart device in the smart home network. Based on the policy conversion rules in the centralized control policy template, corresponding device control commands are generated to control the corresponding smart devices to respond.
2. The intelligent device control method according to claim 1, characterized in that, Based on the preview video stream generated by the camera unit, the system detects the movement of people within the preview video stream and determines the scene activity process description corresponding to the activities of the people in their respective scenes, including: Image target recognition is performed on the preview video stream to obtain multiple image targets and their types; When the target type includes a person type, the corresponding person target is tracked, their user identity is identified, and their temporal behavior sequence is analyzed; When the target type includes an item type, determine the target item type that is associated with the temporal behavior sequence; Based on the user identity, the temporal behavior sequence, and the determined target item type, a description of the scene activity process is generated.
3. The intelligent device control method according to claim 2, characterized in that, Based on the user identity, the temporal behavior sequence, and the determined target item type, a description of the scene activity process is generated, including: The user identity, the temporal behavior sequence, and the target item type are integrated into a structured data tuple; Based on a preset prompt word template, the data tuples are converted into natural language prompts; The natural language prompts are input into a pre-trained natural language generation model, which outputs a description of the scene activity process in natural language format.
4. The intelligent device control method according to claim 1, characterized in that, Based on the scenario activity process description, a matching centralized control strategy template is determined, including: The structured information contained in the scenario activity process description is matched with the predefined structured rule conditions of the centralized control strategy template in the strategy library to determine the centralized control strategy template that achieves the match as the matching result. If no unique centralized control strategy template is matched, the structured information is converted into a natural language description through a natural language generation model. The natural language description is semantically matched with the natural language conditions of the centralized control policy templates in the policy library, and the centralized control policy template with the highest semantic similarity is selected as the matching result.
5. The intelligent device control method according to any one of claims 1 to 4, characterized in that, Before determining the matching centralized control strategy template based on the scenario activity process description, the following steps are included: Based on the user identity carried in the description of the scenario activity process, determine the corresponding user permission level; Based on the user's permission level, select the target policy library from multiple policy libraries; In the target policy library, the centralized control policy template is matched to the scenario activity process description.
6. The intelligent device control method according to any one of claims 1 to 4, characterized in that, Generate corresponding device control commands based on the policy conversion rules in the centralized control policy template, and control the corresponding smart devices to respond, including: The strategy conversion rules are analyzed to determine the multiple smart devices to be controlled and their corresponding target states; For each smart device, generate device control instructions to drive it to achieve the target state; The generated multiple device control commands are scheduled according to the preset collaborative timing relationship in the strategy conversion rules and sent to the corresponding smart devices to execute collaborative responses.
7. The intelligent device control method according to any one of claims 1 to 4, characterized in that, Generate corresponding device control commands based on the policy conversion rules in the centralized control policy template, and control the corresponding smart devices to respond, including: The policy conversion rule is parsed into a task flow containing multiple ordered subtasks, wherein the triggering of at least one subsequent subtask depends on the execution result of the previous subtask. Generate corresponding device control commands for each smart device in sequence for each subtask; Based on the aforementioned dependencies, the device control commands are sequentially scheduled and sent to the smart device for execution.
8. A smart device control device, characterized in that, include: The activity detection module is configured to detect the activity process of a person based on the preview video stream generated by the camera unit, and determine the scene activity process description corresponding to the activity of the person in the preview video stream in their scene. The template matching module is configured to determine a matching centralized control strategy template based on the scene activity process description. The centralized control strategy template includes a strategy conversion rule that maps the scene activity process description to device control instructions adapted to at least one smart device in the smart home network. The generation control module is configured to generate corresponding device control commands based on the policy conversion rules in the centralized control policy template, and control the corresponding smart devices to respond.
9. An intelligent central control device, comprising a camera unit and a controller, the controller including a processor and a memory, the camera unit being used to acquire and preview video streams, characterized in that, The processor invokes and runs a computer program in the memory to perform the steps of the intelligent device control method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, It stores, in the form of computer-readable instructions, a computer program implemented according to any one of claims 1 to 7, which, when invoked by a computer, performs the steps included in the corresponding method.