Apartment network and intelligent device linkage method and system
By fusing network and device data through a spatiotemporal graph attention network and a multi-agent deep deterministic policy gradient algorithm (MADDPG), dynamic collaborative management of apartment networks and smart devices is achieved, solving the problem of insufficient resource allocation in traditional apartment networks and improving resource utilization and security response speed.
Patent Information
- Application Number
- CN202511273479.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-08
- Publication Date
- 2025-11-07
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Traditional apartment networks cannot dynamically adjust resources according to the real-time needs of devices, resulting in critical equipment lag, resource waste, and security delays, failing to meet the needs of intelligent management.
A spatiotemporal graph attention network is used to fuse network operation data and device operation data. The multi-agent deep deterministic policy gradient algorithm (MADDPG) is used to make resource allocation and device control decisions, generate linkage strategies, and achieve scene adaptive optimization.
It improved resource utilization and service response speed, ensured the real-time performance and reliability of linkage strategies, solved problems such as equipment lag, resource waste and security delays, and enhanced the system's intelligence level.
Smart Images

Figure CN120915665A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of artificial intelligence, and particularly relates to an apartment network and intelligent device linkage method and system. BACKGROUND
[0002] With the rapid development of Internet of Things technology, the demand for intelligent management of apartments is increasingly prominent. In related technologies, the traditional apartment network adopts an architecture design based on fixed bandwidth allocation, which cannot dynamically adjust resources according to real-time device requirements. For example, when multiple Internet of Things devices (such as intelligent cameras, environmental sensors, and intelligent home appliances) concurrently transmit data, network bandwidth often appears to be locally congested due to the pre-set fixed quota, resulting in picture lag or data loss of critical devices (such as security cameras). This static management mode not only causes waste of bandwidth resources, but also may cause security event response failure due to delay. Therefore, there is an urgent need to design a technical solution to solve at least one of the above technical problems. SUMMARY
[0003] The main purpose of the embodiments of the present application is to provide an apartment network and intelligent device linkage method and system, aiming to solve the technical problems of key device lag, resource waste, security delay, and insufficient intelligence in related technologies.
[0004] In a first aspect, the embodiments of the present application provide an apartment network and intelligent device linkage method, comprising: for a target apartment, collecting network running data of network devices and device running data of intelligent devices deployed inside the target apartment in real time; determining candidate service scenarios that can be realized by the target apartment based on the device running data; fusing the network running data and the device running data into a joint running graph of the target apartment through a space-time graph attention network; the joint running graph is used to reflect the dynamic coupling relationship between the device running state and the network running state in real time; using a multi-agent deep deterministic policy gradient algorithm (MADDPG) to make network resource allocation and device control decisions on the joint running graph, and combining the candidate service scenarios to realize scene adaptive optimization, to obtain a linkage strategy of the target apartment; the linkage strategy at least includes: a network resource dynamic allocation scheme, an intelligent device control instruction set, a device-to-device linkage relationship parameter, a network-to-device linkage relationship parameter, a dominant scene identifier, and a scene linkage parameter; generating a decision instruction based on the linkage strategy, and using the decision instruction to realize remote control of the network devices and the intelligent devices.
[0005] In a second aspect, the embodiments of the present application provide an apartment network and intelligent device linkage system, comprising: a collection module configured to collect, for a target apartment, network operation data of a network device and device operation data of an intelligent device deployed in the target apartment; a decision module configured to determine a candidate service scenario that can be realized by the target apartment based on the device operation data; to fuse the network operation data and the device operation data into a joint operation graph of the target apartment by using a space-time graph attention network; and to use the joint operation graph to reflect a dynamic coupling relationship between a device operation state and a network operation state in real time; to use a multi-agent deep deterministic policy gradient algorithm (MADDPG) to make network resource allocation and device control decisions for the joint operation graph, and to realize scene adaptive optimization in combination with the candidate service scenario to obtain a linkage strategy of the target apartment, wherein the linkage strategy comprises at least a network resource dynamic allocation scheme, an intelligent device control instruction set, an inter-device linkage relationship parameter, a network and device linkage relationship parameter, a dominant scene identifier, and a scene linkage parameter; and an execution module configured to generate a decision instruction based on the linkage strategy, and to use the decision instruction to realize remote control of the network device and the intelligent device.
[0006] In a third aspect, the embodiments of the present application further provide a terminal device, which comprises a processor, a memory for storing a computer program, and the processor is configured to execute the computer program and implement the apartment network and intelligent device linkage method of the first aspect or any of the embodiments of the present application when executing the computer program.
[0007] In a fourth aspect, the embodiments of the present application further provide a computer readable storage medium, which stores a computer software program, and the computer software program is executed by a processor to implement the apartment network and intelligent device linkage method of the first aspect or any of the embodiments of the present application.
[0008] The embodiment of the present application provides a kind of apartment network and the linkage method and system of intelligent equipment.The method, for target apartment, real-time collection network equipment network operation data, and the equipment operation data of the intelligent equipment deployed in target apartment.Furthermore, based on equipment operation data determines the candidate service scene that target apartment can realize currently.Through space-time diagram attention network, network operation data and equipment operation data are fused into the joint operation graph of target apartment;The joint operation graph is used to reflect the dynamic coupling relationship of equipment operation state and network operation state in real time.Then, using multi-agent deep deterministic policy gradient algorithm MADDPG, network resource allocation and equipment control decision are carried out to the joint operation graph, and scene adaptive optimization is realized in combination with the candidate service scene, and the linkage strategy of target apartment is obtained;The linkage strategy at least includes: network resource dynamic allocation scheme, intelligent equipment control instruction set, inter-device linkage relationship parameter, network and device inter-linkage relationship parameter, leading scene identifier and scene linkage parameter.Finally, based on the linkage strategy generation decision instruction, and using the decision instruction realizes the remote control to network equipment and intelligent equipment.The embodiment of the present application realizes the dynamic collaborative control of apartment network and intelligent equipment by deep fusion space-time diagram attention network and multi-agent reinforcement learning.Both can build global perception joint operation graph, accurately capture the space-time correlation law of equipment and network, and can realize scene adaptive optimization through multi-agent collaborative decision, improve resource utilization and service response speed, and can form closed-loop control mechanism, guarantee the real-time and reliability of linkage strategy. BRIEF DESCRIPTION OF DRAWINGS
[0009] Figure 1 A flowchart of the linkage method of apartment network and intelligent equipment provided by the embodiment of the present application is provided. Figure 2 A module structure schematic diagram of the linkage system of apartment network and intelligent equipment provided by the embodiment of the present application is provided. Figure 3 A structure schematic block diagram of terminal equipment provided by the embodiment of the present application is provided. DETAILED DESCRIPTION
[0010] The embodiment of the present application provides a kind of apartment network and the linkage method and system of intelligent equipment. Among them, the linkage method of apartment network and intelligent equipment can be applied to terminal equipment, which can be mobile terminal, such as mobile phone, virtual reality equipment, tablet computer, notebook computer, desktop computer, wearable device and other terminal equipment.The terminal equipment can be server connected with cloud service system, or server cluster.The above connection mode can be realized by hardware circuit, or realized by communication module.
[0011] For the key device lag problem in the related art, the space-time graph attention network extracts the time sequence and spatial features such as network transmission delay and bandwidth occupation, the network resource scheduling agent dynamically allocates bandwidth resources based on the MADDPG algorithm, adjusts the routing and queue strategy based on the device priority, and avoids the response delay of the key device caused by resource preemption. For the resource waste problem in the related art, the device cooperative control agent adjusts the operation mode and optimizes the operation time sequence according to the energy consumption data and device state in the joint operation graph to realize on-demand energy supply, and the global critic network evaluates the resource matching degree to reduce invalid energy consumption and redundant resource allocation. For the security delay problem in the related art, the scene adaptation coordination agent preferentially identifies the security scene identifier, quickly activates the camera, lock and other device linkage based on the functional dependency relationship, optimizes the decision time sequence through reinforcement learning, and shortens the response cycle from state perception to instruction execution. For the lack of intelligence in the related art, the joint operation graph realizes data deep fusion, and multiple intelligent agents work together to process network and device decisions, quickly adapt to new scenes through scene transfer learning, form a self-optimizing closed loop from perception and decision to execution and feedback, improve the intelligence level of the system, and improve the collaborative management efficiency of the network and intelligent devices.
[0012] Some embodiments of the present application will be described in detail below with reference to the accompanying drawings. The following embodiments and features in the embodiments can be combined with each other without conflict. Please refer to Figure 1 , Figure 1 A flowchart of a linkage method of an apartment network and intelligent device provided by an embodiment of the present application is shown in FIG. 1. As shown in FIG. 1, the method comprises the following steps: Figure 1 Step S101, for a target apartment, real-time collection of network operation data of network devices and device operation data of intelligent devices deployed in the target apartment.
[0013] In an embodiment of the present application, the target apartment refers to a collection of residential units that need to be managed by network and intelligent device linkage, which can flexibly cover different scales of residential scenes. It can be a single independent apartment, an apartment building containing multiple floors and multiple households, and can also be extended to a small community composed of multiple apartment units. The core commonality of these scenes is the need for collaborative operation of network devices and intelligent devices, which requires efficient management through data collection and intelligent decision-making.
[0014] The network equipment in the target apartment is the infrastructure that supports data transmission and connection. For example, the router serves as the core node, responsible for connecting external networks and internal devices, implementing data forwarding and network addressing. The switch is used to expand the network interface, connecting multiple wired devices (such as smart TVs, security hosts), and optimizing the data transmission efficiency within the local area network. The wireless access point (such as Wi-Fi router, signal booster) provides wireless connection services, supporting mobile devices such as mobile phones and smart speakers to access. The gateway node is responsible for protocol conversion between different types of networks, such as connecting ZigBee protocol sensors and Wi-Fi networks, ensuring cross-protocol device communication.
[0015] The intelligent devices deployed in the target apartment include various types of intelligent terminals used to meet the needs of residential life or office business. For example, they can be divided into five categories according to their functions: home control devices (such as smart light switches, smart air conditioners, and curtain motors) for adjusting the comfort of the living environment. Environmental monitoring devices (such as temperature and humidity sensors, air quality detectors, and light sensors) that can sense indoor environmental parameters in real time. Security monitoring devices (such as smart door locks, infrared cameras, and smoke alarms) that ensure the safety of the residence. Audio and video entertainment devices (such as smart TVs, projectors, and sound systems) that provide leisure and entertainment functions. Energy management devices (such as smart meters, water meters, and sockets) that record and regulate device energy consumption.
[0016] Network operation data is key information that reflects the real-time state of network equipment. For example, real-time bandwidth occupancy rate reflects the proportion of current network resource usage, which is too high and will cause device lag. Data transmission delay refers to the time difference between data transmission and reception, which directly affects the experience of real-time scenarios such as video calls and security monitoring. Packet loss rate reflects the proportion of data lost during transmission, which is too high and will cause picture lag or command loss. The number of connected devices records the total number of devices currently connected to the network, which is too many and may cause network congestion; time series fluctuation characteristics reflect the change law of bandwidth, delay, etc. with time, which is used to predict network load peaks.
[0017] Device operation data is core information that reflects the running state of intelligent devices. For example, running state parameters such as the real-time temperature of the air conditioner, the brightness value of the light, and the focal length parameter of the camera reflect the current running indicators of the device. Energy consumption data records the real-time power and cumulative power consumption of the device, which is used for energy saving optimization. Operation response time refers to the time from receiving an instruction to executing an action, reflecting the sensitivity of the device; the interaction frequency between devices counts the number of instruction transmissions between different devices (such as the number of times the door lock and camera are triggered in conjunction), reflecting the degree of collaboration. Function status indicators such as "running", "standby", and "fault" labels directly reflect whether the device is working properly.
[0018] In step S102, the target apartment currently achievable candidate service scenarios are determined based on the device operation data.
[0019] In the embodiments of the present application, based on the apartment device operation data and the living demand, the candidate service scenarios can be divided into the following categories, and each category corresponds to a specific device linkage logic: first, daily living scenarios, including morning mode (automatic opening of curtains in the early morning, gradual brightening of lights, starting of coffee machine), sleep mode (turning off lights after 22:00, adjusting air conditioner to sleep temperature, powering off unnecessary devices), and no one mode (after detecting that all rooms are empty, turning off main lights, reducing air conditioner power, and starting security inspection), which mainly rely on human body sensors, light / air conditioner operation data and time characteristics. Second, security protection scenarios, including leaving home security (after the door lock is closed, the camera starts recording, the door and window sensors are secured, and the network priority is secured for security devices), abnormal alarm (after the smoke alarm is triggered, the light flashes, the door lock is unlocked, and the alarm is pushed to the user), and visitor reception (after the doorbell is triggered, the camera captures the picture, the indoor screen displays the visitor information, and the temporary access control permission is opened), which rely on the state data and alarm information of the door lock, camera and sensor. Third, entertainment and leisure scenarios, including home movie (after the TV is turned on, the light is dimmed, the curtains are closed, and the sound is switched to the theater mode), game mode (after the game console is started, the network bandwidth is preferentially allocated, and the air conditioner is operated in silent mode), and music relaxation (after the sound is turned on, the light changes with the rhythm of the music, and the curtains are half open), which are mainly associated with the running state of the video and audio equipment and the network resource occupation data. Fourth, energy saving optimization scenarios, including energy saving when no one is in the room (after detecting that the room is empty for more than 30 minutes, automatically turning off the lights, air conditioner and socket power supply), peak shifting (according to the smart meter data, automatically reducing the power of the water heater during the peak period of electricity price, and delaying the start of unnecessary devices), and photovoltaic cooperation (if the apartment has photovoltaic equipment, preferentially using photovoltaic power to drive high energy consumption devices), which rely on energy management device data and environmental parameters. Fifth, health care scenarios, including air optimization (after the air quality detector detects that PM2.5 exceeds the standard, the new air system is started and the doors and windows are closed), work and rest reminders (according to the user's set work and rest time, reminding to get up / sleep through light fading and curtain adjustment), and old people's care (when the motion sensor detects no data for a long time, the emergency call device is linked to prompt family members), which rely on environmental monitoring devices and health-related sensor data.
[0020] It can be understood that step S102 can accurately identify the current implementable candidate service scene from the device running data through the intelligent mechanism of multi-feature fusion, behavior sequence mining and dynamic confidence calibration, and break through the limitation of traditional fixed rule matching. This step adopts a three-level intelligent framework of multi-modal feature fusion-temporal behavior modeling-dynamic confidence learning to realize scene recognition. First, a graph attention network is used for multi-modal feature fusion. The device running data is converted into a device-feature heterogeneous graph, the nodes are intelligent devices such as lights, air conditioners and sensors, and the edge weights are dynamically adjusted based on the interaction frequency between devices, such as the number of times of linkage between the door lock and the camera. The feature vector of each node includes real-time state parameters such as air conditioner temperature, energy consumption data such as power and function labels such as security. The graph attention network automatically learns the correlation weight between devices through the attention mechanism, such as preferentially strengthening the correlation between security devices and environmental sensors, and outputs a global feature vector that fuses the cooperative relationship of devices, solving the problem of single device data isolation. Second, an improved Transformer model is used to mine the temporal behavior pattern. The device operation sequence such as 18:30 door unlocking, 18:31 living room light on and 18:32 air conditioner starting is input into the Transformer, which captures the dependency relationship in the time dimension through the multi-head self-attention mechanism, such as the strong correlation between the door unlocking and the air conditioner starting within 5 minutes, and combines the position encoding to strengthen the temporal logic. The temporal feature vector output by the model is spliced with the global feature vector of the graph attention network to form a composite feature of spatial correlation and temporal logic, which is used to match the scene feature library such as the typical temporal pattern of the home-coming scene. Finally, a deep belief network is introduced for dynamic confidence calibration. With the composite feature as the input, the multiple hidden layers of the deep belief network automatically learn the implicit rules of scene matching, and output the initial confidence of each candidate scene. At the same time, the device health data such as the camera fault state and the network delay are connected as calibration factors, and the confidence is dynamically adjusted through the gating mechanism, such as when the security camera is offline, the confidence of the away-from-home security scene is reduced by 30%. If the confidence of a certain scene is higher than the threshold value such as 0.7 for 3 consecutive time windows, it is included in the candidate list. Otherwise, it is filtered to ensure the reliability of scene recognition. In addition, an incremental contrast learning model is used to realize the self-update of the scene library. When a new device linkage mode such as the yoga mat sensor activation, the air conditioner adjustment to 26°C and the music device starting is detected, the model automatically judges whether it is a new scene by comparing the feature difference between the new mode and the existing scene, and if the difference degree exceeds the threshold value such as 0.6, it is included in the scene library and the model parameters are fine-tuned, without manual intervention, to support user-defined scenes and enhance the flexibility and adaptability of scene recognition.
[0021] Step S103, the network running data and the device running data are fused into a joint running graph of the target apartment through a spatio-temporal graph attention network.
[0022] In the embodiments of the present application, the spatio-temporal graph attention network is a deep learning model that fuses dynamic changes in the time dimension and spatial dimension correlation, and is specially used for processing the spatio-temporal coupling characteristics of network and device data. The core advantage of the spatio-temporal graph attention network is to capture time dynamics and spatial correlation at the same time. The time attention module is used to mine the change law of device status and network load over time (such as the fluctuation trend of bandwidth peak period, the timing logic of device start and stop). The spatial attention module is used to identify the correlation strength between different devices and network nodes (such as the physical connection stability of a camera and a router, and the functional dependence degree of an air conditioner and a temperature and humidity sensor). The spatio-temporal graph attention network can automatically learn the weight of key features and preferentially focus on the spatio-temporal information that has a greater impact on scene decision-making, thereby providing accurate feature support for subsequent graph construction. In the embodiments of the present application, the joint operation graph is used to reflect the dynamic coupling relationship between the device operation state and the network operation state in real time. The joint operation graph is a structured data model that intuitively reflects the dynamic association between devices and networks in the target apartment. Its core role is to convert scattered network operation data and device operation data into a visual association graph. In the joint operation graph, nodes are divided into two categories: one represents network devices (such as routers and switches), and node attributes include real-time operation characteristics such as bandwidth occupancy and transmission delay; the other represents intelligent devices (such as air conditioners and cameras), and node attributes include device characteristics such as operating state and energy consumption data. The edges between nodes are constructed based on physical connections, data interactions, and functional dependencies, and the edge weights dynamically reflect the correlation strength (such as higher edge weights between devices with frequent data interactions). Through the graph, the state linkage of devices and networks can be viewed in real time, for example, whether a camera lag is related to insufficient router bandwidth, thereby providing a global perspective for decision-making.
[0023] Step S103 fuses the data into a joint operation graph by a spatio-temporal graph attention network. Specifically, first, data preprocessing is performed on the collected network operation data (such as bandwidth, delay) and device operation data (such as state parameters, energy consumption) to clean, filter outliers, align time granularity, and standardize features, ensuring uniform data format and comparable features. Second, a basic graph structure is constructed based on device physical connection relationships (such as the access relationship between a camera and a router), data interaction records (such as the frequency of data packet transmission between devices), and functional dependency logic (such as the dependence of security devices on the network), initializing the nodes and edges of the graph, and assigning initial state features to each node and initial association weights to the edges. Finally, spatio-temporal feature fusion is performed by inputting the preprocessed data into the spatio-temporal graph attention network. The temporal attention module captures the temporal dependence of each node feature (such as the synchronous change of device energy consumption and network load at a certain time period) through a multi-head self-attention mechanism, generating temporal association features. The spatial attention module calculates the dynamic association weights between nodes, updates the edge attributes to reflect the real-time interaction intensity, and generates spatial association features. After fusing the two types of features, a joint operation graph is formed that includes node real-time state, spatio-temporal association weight, and dynamic interaction relationship, achieving global dynamic characterization of device and network state.
[0024] As an optional embodiment, in step S103, the network operation data and the device operation data are fused into a joint operation graph of the target apartment by a spatio-temporal graph attention network, including: Step 11, preprocessing the collected network operation data and device operation data, including data cleaning, outlier filtering, time granularity alignment, and feature standardization, to obtain standardized network feature data and standardized device feature data; wherein the network operation data includes real-time bandwidth occupancy rate, data transmission delay, packet loss rate, number of connected devices, and time series fluctuation characteristics of network devices, and the device operation data includes operating state parameters, energy consumption data, operation response time, interaction frequency between devices, and function status identification of intelligent devices; Step 12, constructing a basic graph structure of the target apartment, the nodes of the basic graph structure including network device nodes and intelligent device nodes, the edges of the basic graph structure being constructed based on physical connection relationships, data interaction relationships, and functional dependency relationships, and the initial edge weights being set based on historical data interaction frequency and association strength; the node attributes of each node in the joint operation graph include first real-time operation features of corresponding network device nodes or second real-time operation features of intelligent device nodes, and the edge attributes include spatio-temporal association strength and time-varying regularity features between nodes; Step 13, input the standardized network feature data and standardized device feature data into the spatio-temporal graph attention network, capture the dynamic dependency relationship of each node feature in the time dimension through the time attention module to generate the time sequence correlation feature; calculate the correlation weight between different nodes through the spatial attention module, update the edge weight of the basic graph structure based on the correlation weight, and generate the spatial correlation feature; Step 14, fuse the time sequence correlation feature and the spatial correlation feature to obtain the joint operation graph containing the real-time state of the node, the spatio-temporal correlation weight and the dynamic interaction relationship.
[0025] Specifically, in step 11, data preprocessing is a basic link for eliminating noise and unifying data format. Through data cleaning, collection errors (such as sensor false data) are removed, outlier filtering eliminates extreme values (such as network delay suddenly jumping to 1000ms) beyond the reasonable range, time granularity alignment unifies the sampling frequency of different devices (such as adjusting to 1 minute / time), feature standardization maps the data to the same order of magnitude (such as converting the bandwidth occupancy rate to 0-1 interval value) through normalization or standardization method, and ensures that the network model can effectively learn the feature rule. Exemplarily, for network operation data, the instantaneous packet loss rate abnormal peak value caused by temporary signal interference is filtered out; the temperature data (℃) of the intelligent air conditioner and the bandwidth occupancy rate (%) of the router are standardized to eliminate the dimension difference; the sampling frequency of the device operation response time is adjusted from 10 seconds / time to 1 minute / time, which is consistent with the time granularity of the network data. Thus, through preprocessing, the interference of data noise on the model is reduced, the data format and time dimension are unified, high-quality input is provided for subsequent graph structure construction and feature fusion, and the learning efficiency of the model on key features is improved.
[0026] In step 12, the base graph structure is the skeleton of the joint operation graph, and the physical devices are abstracted as graph nodes through nodal modeling. The edges are constructed based on three types of relationships: physical connection relationship (such as wired / wireless access link between devices and routers), data interaction relationship (such as instruction transmission link between devices), and functional dependency relationship (such as the dependency relationship of network bandwidth for security cameras). Initial edge weights are assigned according to historical data interaction frequency and correlation strength, so that the graph can preliminarily reflect the static association between devices and networks. For example, routers, smart door locks, and temperature and humidity sensors are modeled as network nodes and device nodes, respectively. A physical connection edge is created for the Wi-Fi connection between the router and the smart door lock, and the initial weight is based on the historical connection stability score (such as a 90% connection success rate corresponding to a weight of 0.9). A functional dependency edge is created for the state feedback link between the temperature and humidity sensor and the air conditioner, and the initial weight is based on the historical interaction frequency (30 interactions per hour corresponding to a weight of 0.7). In this way, the scattered devices and network entities are converted into a computable graph structure through structured modeling, the node attributes and edge association types are clearly defined, and a basic framework is provided for subsequent dynamic learning of spatio-temporal features, realizing the visualization and quantitative expression of the association relationship.
[0027] In step 13, the spatio-temporal graph attention network captures the temporal dynamics and spatial associations through time attention module and space attention module respectively. The time attention module uses a multi-head self-attention mechanism to calculate the dependency relationship of features at different time steps (such as the synchronous change of device energy consumption and network load at a certain time period), highlighting the influence of key time nodes. The space attention module dynamically updates the edge weight by calculating the node feature similarity, strengthening important associations (such as the association weight between the camera and the router in the security scenario) and weakening secondary associations (such as the weak association between the coffee machine and the router). For example, the time attention module captures the temporal dependency of the router bandwidth occupancy rate rising over time after the smart TV is started at 19:00-20:00 in the evening, and assigns higher attention weight to the features at this time period. The space attention module calculates and finds that when the security camera is started, its feature similarity with the router significantly improves, dynamically adjusting the edge weight between the two from 0.6 to 0.9, strengthening the support association of the network to the security device. Thus, by using the time and space modules to respectively mine the dynamic rules in the time dimension and the correlation strength in the space dimension, the limitations of traditional static graph models are overcome, and the graph can preliminarily reflect the dynamic coupling relationship between devices and networks, providing fine-grained features for subsequent fusion.
[0028] In step 14, the fusion of the time correlation feature and the space correlation feature is the core of the joint operation graph. Through splicing or weighted fusion, the dynamic changes in the time dimension (such as the fluctuation of the device state over time) and the correlation strength in the space dimension (such as the real-time interaction weight between nodes) are integrated to form a complete graph containing the real-time state of the node, the space-time correlation weight, and the dynamic interaction relationship, realizing the global dynamic description of the device and network operation state. Exemplarily, in the fused graph, the attributes of the smart TV node include real-time power (200W) and time sequence features (power surge at 19 o'clock in the evening); the edge attributes of the router include space correlation weight (0.8) and time regularity (weight continuously higher than other periods from 19 o'clock to 20 o'clock in the evening), which intuitively reflects the strong dynamic correlation between the TV and the network in the viewing scenario. Thus, through feature fusion, the information in the time and space dimensions is complementary, enabling the joint operation graph to not only reflect the current state of the device and network, but also embody the time sequence variation law of the correlation relationship, providing a global and dynamic state perception basis for subsequent multi-agent decision-making, and improving the accuracy and adaptability of decision-making.
[0029] As an optional embodiment, in step 12, the construction of the basic graph structure of the target apartment includes: Step 121, node modeling of network devices and intelligent devices in the target apartment, assigning a unique node identifier to each network device and intelligent device, the node identifier being associated with device physical location information, device type attribute and basic function parameters; wherein the network device nodes include router nodes, switch nodes, gateway nodes and wireless access point nodes, and the intelligent device nodes include home control nodes, environment monitoring nodes, security monitoring nodes, audio and video entertainment nodes and energy consumption management nodes; Step 122, obtaining the connection relationship between network devices through network topology detection technology, obtaining the access relationship between intelligent devices and network devices through device installation layout and hardware interface mapping, creating a first type of edge for node pairs with direct physical connection, and setting the initial weight of the first type of edge based on the connection stability parameter; Step 123, collecting data packet transmission records between devices within a dynamic time window, extracting data interaction frequency, data volume and interaction time length features, creating a second type of edge for node pairs with data interaction, and setting the initial weight of the second type of edge based on the weighted calculation result of data interaction frequency and average interaction time length; Step 124, obtaining the functional synergy relationship of intelligent devices in each service scenario, identifying the control instruction dependency, state feedback dependency and resource sharing dependency between devices, creating a third type of edge for node pairs with functional dependency, and setting the initial weight of the third type of edge based on the functional dependency strength level quantization value; Step 125, integrate the first type of edge, the second type of edge and the third type of edge to form a basic graph structure containing node attributes and edge attributes, and establish a mapping relationship between the graph structure and the physical identification of the device, which is used for subsequent graph dynamic updating and node association query.
[0030] Specifically, in step 121, the node modeling is the process of converting physical devices into basic units of graph structure. In implementation, a unique node identifier is assigned to each network device and intelligent device, and the identifier uses a combination of "device type-physical location-sequence number" encoding (for example, "security camera-living room-01" and "router-entrance-01"), which ensures that each node can accurately correspond to a physical device. The node identifier is associated with three types of key information: physical location information is recorded through specific coordinates or area labels on the apartment floor plan (such as "main bedroom southeast corner" and "living room center"). The device type attribute is classified according to function (network devices are labeled as "router", "switch", etc., and intelligent devices are labeled as "home control", "environmental monitoring", etc.). The basic function parameters record the inherent properties of the device (such as the maximum number of supported devices for a router, the cooling and heating power range for an air conditioner, and the night vision function parameters for a camera). Through this step, the scattered devices are abstracted into graph nodes containing multi-dimensional features, providing a basic unit for subsequent association analysis.
[0031] In step 122, the first type of edge (physical connection edge) is used to record the physical connection relationship between devices, and implementation needs to combine technical detection and structured data verification. The connection relationship between network devices is obtained through network topology detection technology, such as using SNMP protocol to scan device interface status, or automatically identifying wired connection ports of switches and routers through link layer discovery protocol (LLDP). The access relationship between intelligent devices and network devices is confirmed through device installation layout (annotating wiring direction) and hardware interface mapping table (recording the correspondence between device MAC address and access port). For node pairs that have direct physical connection (such as Wi-Fi connection between smart door lock and wireless access point, wired connection between projector and switch), the first type of edge is created. The initial weight is set according to the connection stability, for example, referring to the number of connection interruptions in the past 30 days and the signal strength fluctuation. The more stable the connection, the fewer the interruptions, and the higher the weight (such as setting the weight of a continuously stable connection link to 0.9 and the weight of a link with occasional disconnection to 0.6).
[0032] In step 123, the second type of edge (data interaction edge) is used to reflect the dynamic data exchange intensity between devices, and real-time data collection and feature analysis are implemented. First, a dynamic time window (usually 10-30 minutes) is set, and the data packet transmission records between devices within the window are collected through network traffic monitoring tools, and three core features are extracted: data interaction frequency (number of transmitted and received data packets per unit time), data volume (total number of bytes transmitted), and interaction duration (average duration of a single data transmission). For node pairs that have data interaction (such as a temperature and humidity sensor sending status data to an air conditioner, and a mobile phone APP sending control instructions to a smart lamp), a second type of edge is created. The initial weight is set according to the interaction intensity, and the more frequent and longer the single duration, the higher the weight (for example, a link that interacts 50 times per hour and has an average duration of 10 seconds per single interaction has a higher weight than a link that interacts 10 times per hour and has an average duration of 2 seconds per single interaction).
[0033] In step 124, the third type of edge (functional dependency edge) is used to express the functional coordination relationship of devices in the service scenario, and scene logic and dependency rules need to be analyzed. By analyzing the pre-set service scenario configuration files of the apartment (such as the device linkage rules of "away mode" and "movie mode") and historical operation logs (which record the instruction triggering relationship between devices), three types of functional dependencies are identified: control instruction dependency (such as "door lock closing" triggering "camera starting"), status feedback dependency (such as "smoke alarm alarm" triggering "light flashing"), and resource sharing dependency (such as multiple devices sharing the same network bandwidth). For node pairs that have functional dependencies, a third type of edge is created. The initial weight is set according to the dependency intensity, and the core dependencies (such as the dependency of the camera and the alarm platform in the security scenario) have high weights (such as 0.8-1.0), and the secondary dependencies (such as the weak dependency of the humidifier and the temperature and humidity sensor) have low weights (such as 0.3-0.5).
[0034] Step 125 integrates the three types of edges to form a complete basic graph structure, and establishes the association index between the graph and the physical devices. When integrating, the unique properties of each type of edge (physical connection edge reflects hardware link, data interaction edge reflects dynamic exchange, and functional dependency edge reflects scene logic) need to be preserved to avoid redundancy. For example, if a pair of nodes has both physical connection and data interaction, the two types of edges need to be preserved to distinguish different association types. Then, a mapping relationship between the graph structure and the physical identifier of the device is established, and the correspondence between the node ID and the device serial number, MAC address, and the edge ID and the association type (physical / interaction / dependency) is recorded through a lookup table. This mapping relationship supports subsequent operations, and when the device state changes (such as offline), the corresponding node in the graph can be quickly located. When querying the associated devices of a certain device, the relevant edges and weights can be directly retrieved through the graph, providing support for subsequent dynamic updating of the graph.
[0035] Further optionally, cooperative rule analysis based on user scenario configuration information. Smart home systems or apartment management platforms usually support user-defined scene modes (such as homecoming mode, sleep mode, and away mode), and these scene configurations directly define the cooperative logic of the devices. For example, the configuration of the sleep mode may include a device linkage sequence of turning off the main light → turning on the night light → setting the air conditioner to 26°C → closing the curtains, from which the functional dependency relationship of the main light, night light, air conditioner, and curtains can be extracted. The configuration of the security mode may set the sensor trigger → the camera starts recording → the gateway pushes the alarm information to the user's mobile phone, which explicitly defines the cooperative path of the sensor, camera, and gateway. By analyzing the scene rules configured in the system, the user-defined device function cooperative relationship can be directly obtained. Further optionally, based on the dependency relationship of device communication protocol and interface call. Smart devices usually communicate through standardized protocols (such as MQTT, HTTP, ZigBee, etc.) or private interfaces. By monitoring the protocol interaction content between devices, the technical dependency relationship of function cooperation can be identified. For example, by analyzing the communication messages of the smart speaker and the sweeping robot, if it is found that the speaker continuously sends "start cleaning" and "return to charging" control instructions, and the robot feedbacks "cleaning progress" in real time, it can be determined that there is a cooperative relationship between control and feedback. In the interface log of a network device (such as a switch), if the power supply state adjustment of a smart socket needs to synchronize data through the switch to the energy management platform, the resource sharing cooperative relationship of the socket, switch, and management platform can be identified. Through protocol analysis and interface call analysis, the bottom-level functional dependency link between devices can be accurately located. Further optionally, supplement the cooperative relationship by combining expert knowledge and scene modeling. For complex scenes or implicit cooperative relationships (such as cross-system device cooperation), expert experience needs to be combined for manual sorting and supplementing. For example, in the energy management scene, experts can define the "solar panel → energy storage battery → smart home total control" energy distribution cooperative relationship based on experience, even if this relationship is not directly reflected through logs or configurations. For cross-type cooperation between network devices and smart devices, experts can define the cooperative rule "when the load of the wireless access point is too high, automatically notify the video device to reduce the picture quality to reduce bandwidth occupation", which supplements the scenes not covered by technical documents.
[0036] Through the combination of the above methods, the functional cooperative relationship of the devices in the target apartment can be comprehensively obtained, including both the pre-set static rules of the devices and the dynamic behaviors in actual operation, providing accurate basis for subsequent construction of the third type of edge (functional dependency edge) in the basic graph structure.
[0037] As an optional embodiment, in step 13, the standardized network feature data and the standardized device feature data are input into the spatio-temporal graph attention network, and the time attention module is used to capture the dynamic dependency relationship of each node feature in the time dimension to generate time sequence correlation features, including: Step 131, time series alignment of the standardized network feature data and the standardized device feature data is performed by node, and a feature time series of each node is constructed, the feature time series containing multi-dimensional feature values in a preset time span; Step 132, the feature time series of each node is divided into continuous time windows, and each time window contains a preset number of time step features; Step 133, the feature sequence in each time window is calculated by attention through a multi-head self-attention mechanism of a time attention module, and an attention score is obtained; the feature vector of each time step is taken as a query vector, a key vector and a value vector, a query matrix, a key matrix and a value matrix are generated through linear transformation, and the attention score between different time steps is calculated, the attention score being determined based on the similarity of the query vector and the key vector and a time decay factor; Step 134, the attention score is normalized to obtain a time attention weight, and the feature vectors of each time step are weighted and aggregated based on the time attention weight to generate local time series features of each node; Step 135, the local time series features are input into a gated recurrent unit to capture long-time dependence across time windows, and time series correlation features containing short-term fluctuation features and long-term trend features are fused, the time series correlation features carrying the dynamic change law of the node in the time dimension.
[0038] Specifically, in step 131, cross-device time series unification is achieved through a device characteristic adaptive alignment mechanism, breaking through the limitations of traditional fixed timestamp alignment. In specific implementation, the alignment granularity is dynamically adjusted in combination with the sampling characteristics of network devices and intelligent devices (such as high-frequency sampling of routers and low-frequency reporting of sensors). For devices with high real-time requirements (such as cameras and routers), the time axis is calibrated with millisecond-level precision. For devices with smooth state changes (such as temperature and humidity sensors), natural alignment is achieved through feature trend matching, avoiding distortion caused by mechanical interpolation. Finally, a coherent feature sequence is constructed for each node, and the multi-dimensional state of the device (such as running mode and interaction state) is connected in time flow order to form a complete time sequence track covering a preset period, ensuring that the sequence can truly reflect the natural evolution of the device state.
[0039] In step 132, a dynamic window mechanism driven by data activity is used to replace the fixed window division. In implementation, the window size is automatically adjusted by real-time monitoring of the change intensity of the feature sequence (such as the frequency of device state transition, the frequency of network fluctuation). When the data activity is high (such as the state of the video and audio device changes frequently), a small window is enabled to focus on short-term details and capture transient fluctuations. When the data tends to be stable (such as the stable operation of the air conditioner at night), a large window is switched to aggregate long-term features and reduce redundant calculations. During the window division process, an overlapping area between windows is also preserved to ensure that the features of adjacent windows are naturally connected and to avoid information loss caused by time breakpoints, providing continuous time sequence segments for subsequent attention calculation.
[0040] Further optionally, the dynamic window mechanism driven by data activity dynamically adjusts the size and sliding strategy of the time window by real-time sensing of data change intensity, to achieve accurate capture of key features and efficient filtering of redundant information. The specific implementation starts from the quantitative evaluation of data activity, which judges the change intensity of time sequence features in real time through multiple dimensions. The frequency of device state transition is calculated, such as the number of times of switching on and off the light, the number of times of sudden increase and decrease of bandwidth per unit time, and the more frequent the transition, the higher the activity. The fluctuation amplitude of key features is calculated, such as the change range of air conditioner temperature and router delay per unit time, and the larger the amplitude, the higher the activity. At the same time, the interaction intensity between devices is calculated, such as the instruction transmission frequency and data packet volume, and the more intensive the interaction (such as the high-frequency data exchange between the television and the server during video watching), the higher the activity, and finally an activity score in the interval [0, 1] is generated to intuitively reflect the dynamic characteristics of the data. On this basis, the mechanism presets multiple activity thresholds based on device type and scene demand as the basis for adjustment. For devices with high real-time requirements and state prone to mutation (such as video and audio devices, security cameras), a relatively sensitive threshold is set (such as a score greater than 0.6 as high activity). For devices with stable state (such as temperature and humidity sensors), a relaxed threshold is set (such as a score greater than 0.4 as high activity). For example, the score of a smart TV during the startup phase rises to 0.8, triggering the high activity threshold, and the score of a stable running air conditioner at night drops to 0.2, triggering the low activity threshold. Subsequently, the window size is dynamically adjusted according to the real-time score and the threshold. In the high activity state, a small window of 5-10 time steps is enabled to focus on short-term high-frequency fluctuations, such as accurately capturing the transient changes of “TV turned on → bandwidth suddenly increased → sound synchronized” during the startup of video and audio devices. In the low activity state, a large window of 20-30 time steps is switched to aggregate long-term stable features, such as integrating the stable trend of “air conditioner constant temperature → network load stable” at night; in the medium activity state, a medium window of 10-15 time steps is used to balance details and efficiency.
[0041] To avoid feature discontinuity during window switching, the mechanism adopts a dynamic overlapping sliding strategy. The overlap rate is set to 50% during high activity periods (e.g., previous window 1-10, and subsequent window 6-15) to ensure continuous connection of high-frequency changes. The overlap rate is reduced to 20% during low activity periods (e.g., previous window 1-30, and subsequent window 25-55) to reduce redundant calculations. The overlap rate is temporarily increased to 40% during activity switching to avoid distortion caused by sudden changes in size. Through this mechanism similar to "adaptive lens", "close-up lens" is used to capture details when data changes dramatically, and "wide-angle lens" is used to aggregate trends when it is stable, ensuring that key features are not lost and reducing redundant interference, providing high-quality input for subsequent time series feature extraction.
[0042] In step 133, time distance perception and scene-oriented attention mechanisms are integrated based on multi-head self-attention. During calculation, the features of each time step not only generate basic attention scores through feature similarity, but also dynamically adjust the weights by combining time distance. The features of recent time steps (e.g., within the last 5 minutes) automatically obtain higher basic weights, and the weights of distant time steps naturally decay with distance, consistent with the rule that device state has stronger short-term correlation. At the same time, scene pre-recognition information (e.g., security mode, daily mode) is introduced to dynamically adjust the attention direction. In the security scene, attention will focus on the time steps before and after the abnormal trigger. In the video scene, it focuses on the time series features of the device startup phase, making attention calculation more suitable for scene requirements and accurately capturing feature associations at key time nodes.
[0043] In step 134, precise normalization and feature aggregation of attention weights are achieved through a double-gated screening mechanism. When implemented, the first heavy gate importance screening gate automatically filters low-contribution time steps (e.g., redundant sampling points with no state change) and only retains time steps that have a significant impact on the current state. The second heavy gate scene adaptation gate dynamically adjusts the weight distribution ratio based on device type and scene requirements, for example, in the energy-saving scene, the weight of energy-related time steps is strengthened. In the network optimization scene, the features of bandwidth fluctuation time steps are highlighted. After double screening, the features of the remaining time steps are weighted and aggregated to generate local time series features that can concentrate on key details and meet the core needs of the current scene.
[0044] In step 135, the long-time dependence across windows is captured by the memory enhancement mechanism of the gated recurrent unit (GRU). In implementation, the local timing features generated in step 134 (containing short-term fluctuation details) are input into the GRU, and the model will automatically remember the feature correlation of different time windows: for example, learning the daily routine of the morning period light turning on -> air conditioner starting -> network bandwidth rising, or the periodic trend of the night equipment gradually going into hibernation -> network load decreasing. The GRU filters out noise information through the gating unit, retains the key timing logic, and finally fuses the short-term fluctuations and long-term trends to form complete timing correlation features that contain both the instantaneous change details of the device state and the running mode rules across the period, providing dynamic feature support in the time dimension for subsequent atlas construction.
[0045] Further optionally, in step 135, in the long-time dependence capture, the gated recurrent unit (GRU) dynamically filters and retains key information through a simplified and efficient gating mechanism, solving the gradient disappearance problem that the traditional recurrent neural network (RNN) is prone to when processing long sequences, thereby accurately capturing feature correlations across long time spans. First, the history information is filtered through the reset gate. The role of the reset gate is to determine "whether to ignore the past history information". When processing new time sequence features (such as device state data at a certain time), the reset gate generates a weight value between 0 and 1 according to the current input features and the hidden state at the previous time (i.e. the memorized history information). If the weight is close to 1, it indicates that the history information needs to be retained (such as the long-term stable state of the device in "sleep mode", which needs to remember the initial settings several hours ago). If the weight is close to 0, it indicates that part of the history information should be "forgotten" (such as abnormal fluctuations caused by temporary device failure, which does not need to be included in long-term trend analysis). Through this filtering, GRU can eliminate noise interference and focus on the history association that is meaningful to the current state. Further, new information is integrated into the candidate hidden state. After the reset gate filters the effective history information, GRU combines the current input features with the filtered history information to generate a "candidate hidden state". This step is equivalent to integrating "historical effective information + new input information" to form a preliminary judgment of the current time state. For example, when analyzing the long-term running trend of an air conditioner, the candidate hidden state integrates "temperature regulation rules in the past 2 hours" and "current room temperature data" to preliminarily predict the "temperature change trend in the next 1 hour". Then, the information retention proportion is determined by the update gate. The update gate is the core of GRU for processing long-time dependence, which determines "what proportion of the newly generated candidate state and the old hidden state (history memory) should be retained", generating another weight value between 0 and 1. If the weight is close to 1, it indicates that new information is more important and the candidate hidden state should be mainly retained (such as the key state change of the device switching from "standby" to "running", which needs to be updated first). If the weight is close to 0, it indicates that the history information is more important and the old hidden state should be mainly retained (such as the continuous low-power feature of the device in "energy-saving mode", which needs to be remembered for a long time). For example, in capturing the periodicity of "light automatically turning on at 7 am on weekdays", the update gate will retain the new trigger information at around 7 am every day, while long-term remembering the historical trend of "repeating every Monday to Friday", thereby realizing long-time dependence capture across days and weeks. Then, the final hidden state is output, completing an iteration. GRU applies the decision result of the update gate to the candidate hidden state and the old hidden state, and generates the final hidden state through weighted fusion. It contains both filtered history memory and key information of new input. This final state will be passed on as "memory at the current time" to the next time step for the next round of gating calculation.Through the cycle of screening, integration and decision-making, the GRU can dynamically retain key historical associations (such as periodic operation patterns and cross-scene state dependencies) while eliminating redundant information when processing long-time series data (such as hours or days of device operation records), thereby efficiently capturing long-time dependencies and providing stable long-term trend support for the fusion of time-series association features.
[0046] As an optional embodiment, in step 13, the calculation of the association weight between different nodes by the spatial attention module, the updating of the edge weight of the basic graph structure based on the association weight, and the generation of the spatial association feature include: Step 136, extract the time-series association feature of each node processed by the time attention module, and combine the type attribute and physical location information of the node in the basic graph structure to construct a node comprehensive feature vector; Step 137, input the node comprehensive feature vector into the spatial attention module to calculate the association weight between node pairs; for each center node, take the attribute feature vector of each center node as a query vector, and take the attribute feature vector of each adjacent node of each center node as a key vector, calculate the similarity between the query vector and each key vector, combine the initial weight of the corresponding edge in the basic graph structure and the node type matching degree to generate an initial association score; Step 138, perform nonlinear transformation on the initial association score by an activation function to obtain a spatial attention weight, which reflects the influence degree of different adjacent nodes on the center node; Step 139, update the weight of the corresponding edge in the basic graph structure based on the spatial attention weight to form a dynamically updated spatial association edge attribute; Step 1310, for each center node, weight and aggregate the attribute feature vectors of the adjacent nodes according to the updated spatial attention weight to fuse the spatial association feature containing the spatial interaction relationship between nodes.
[0047] Specifically, step 136 constructs node comprehensive features through a multi-modal feature deep fusion mechanism, breaking through the limitations of a single feature dimension. In step 136, first, the time-series correlation features of the nodes after processing by the temporal attention module (including the dynamic change law of the device state) are extracted, and then two types of core static features in the basic graph structure are fused. The node type attribute (such as the function label of the security monitoring type and the network routing type) and the physical location information (such as the spatial coordinates or regional division in the apartment). For example, the comprehensive features of the security camera node not only include the time-series features of its state fluctuations, but also integrate the security label and the location information of the corner of the living room. The router node fuses the network fluctuation time-series features, the network device label, and the location features of the entrance. Through this multi-dimensional feature splicing, the comprehensive feature vector of each node can not only reflect the dynamic operation law, but also embody the static attributes and spatial distribution, providing a comprehensive basis for subsequent correlation weight calculation.
[0048] Further optionally, the multi-modal feature deep fusion mechanism aims to break through the limitations of a single feature dimension, and construct a comprehensive node feature representation by integrating dynamic running features and static attribute features. In implementation, first, the dynamic time-series features of the nodes in the time dimension are extracted, which reflect the real-time change law of the device running state, such as the bandwidth fluctuation trend of network devices, the start-stop time-series mode of smart devices, etc. At the same time, two types of core static features of the nodes are fused: one is the type attribute feature, i.e., the functional classification label of the device, such as security monitoring type, home control type, network transmission type, etc., which is used to clarify the functional positioning of the device. The other is the physical location feature, which quantifies the spatial distribution of the device through the spatial coordinates or regional division of the apartment (such as the living room, bedroom, entrance). Through deep splicing and semantic fusion of dynamic time-series features and static attributes and location features, the comprehensive feature vector of each node can not only capture the dynamic changes in the running process, but also embody the functional attributes and spatial layout, providing multi-dimensional basis for the accurate calculation of the correlation weight between nodes in the subsequent, avoiding the correlation judgment deviation caused by single features.
[0049] In step 137, a type matching degree enhancement mechanism is introduced in the spatial attention calculation to make the correlation score more consistent with the device function coordination logic. The correlation strength dynamic calibration mechanism is used to update the edge weight of the basic graph structure in real time, so that the graph can accurately reflect the current state of the correlation between nodes. In implementation, the mechanism takes the real-time attention weight generated by the spatial attention module as the core basis, and dynamically fuses it with the original initial edge weight (based on the historical data statistics of the correlation strength) in the basic graph structure. For node pairs with frequent real-time interaction and close functional dependence, such as high-frequency data transmission between the TV and the router in the viewing scene, the calibration process will give priority to the real-time attention weight, significantly increase the edge weight, and highlight the current strong correlation state. For node pairs with stable historical correlation but less real-time interaction, such as the smart socket of the long-stay machine and the router, the calibration process balances the initial weight and the real-time weight, preserving the basic correlation while avoiding excessive weakening. In addition, when the node state changes abruptly (such as device offline or scene switching), the mechanism will trigger a rapid calibration to immediately reduce the correlation weight of the abnormal node, ensuring that the edge attribute is always synchronized with the actual running correlation strength of the device, so that the graph can dynamically respond to real-time changes in device correlation within the apartment. Further optionally, for each central node (such as a smart door lock), its comprehensive feature vector is used as the query vector, and the comprehensive feature vector of the adjacent node (such as a camera or a router) is used as the key vector to generate the basic correlation score through feature similarity calculation. At the same time, dynamically adjust the score based on two types of key information: one is the initial weight of the edge in the basic graph structure (such as the historical linkage weight of the door lock and the camera), and the other is the node type matching degree, i.e. the matching degree of functionally closely related type combinations (such as "security camera" and "security door lock") is high, and the correlation score is automatically increased. The matching degree of functionally weak combinations (such as "audio and video TV" and "security door lock") is low, and the score is appropriately reduced. For example, when calculating the correlation score of the door lock and the camera, the score will be significantly higher than that of the door lock and the sound because both belong to the security category and have a high historical linkage frequency, ensuring that the correlation calculation is consistent with the scene function logic.
[0050] In step 138, a scene-oriented activation function is used to perform non-linear transformation on the initial correlation score to strengthen the weight differentiation of key correlations. In implementation, the activation strategy is dynamically selected according to the node type and the current pre-identified scene: in the security scene, a steep activation function is used to amplify the difference between high and low correlation scores, making the strong correlation weight between security devices more prominent; in the daily scene, a gentle activation function is used to retain more potential correlation information. For example, when the "leaving home and arming" scene is detected, the activation function will focus on strengthening the correlation scores of the door lock and the camera, as well as the door and window sensors, and suppressing the weak correlation scores of the lights and the sound, generating spatial attention weights that can clearly distinguish the influence of adjacent nodes on the central node, avoiding interference from irrelevant nodes.
[0051] It is worth noting that the core of the scene-oriented activation function is to dynamically adjust the steepness of the nonlinear transformation to make the discriminability of the correlation score fit the functional requirements of the current scene. In scenes that require precise focus on key correlations, strengthen the distinction, and in scenes that need to preserve potential correlations, weaken the distinction, and achieve on-demand attention weight generation. Its implementation process revolves around scene recognition results and node type features, and selects activation strategies accordingly. In security scenarios (such as leaving home and setting up alarms), the correlation between devices has strong functional focus, and the correlation of core security devices needs to be prioritized and irrelevant device interference needs to be strictly suppressed. At this time, the steep activation function is used, which has the characteristics of amplifying the gain of the initial correlation score of medium and above, and compressing and attenuating the low score, so that the gap between high and low correlations is significantly widened.
[0052] For example, in the scene of leaving home and setting up alarms, when calculating the correlation score of the intelligent door lock (central node) and the adjacent nodes, the initial correlation score of the door lock and the camera, door and window sensors is already high (because they belong to the security category and have a high historical linkage frequency), and the steep activation function will further amplify these scores, making their weights rise from high to extremely strong. The initial correlation score of the door lock and the light, sound is low (weak functional correlation), and the function will compress these scores to near zero, making their weights drop from low to negligible. Through this strong distinction process, the spatial attention weight can clearly highlight the core security links of the door lock, camera, and sensor, avoiding the distraction of non-security devices such as lights and sound, and ensuring the accuracy of correlation focus in security scenarios.
[0053] In daily scenarios (such as morning mode, energy saving without people), device correlation is more flexible and needs to preserve more potential linkage possibilities (such as weak correlation between environmental devices and home devices that may upgrade with scene changes). At this time, the gentle activation function is used, which has the characteristics of gently adjusting the initial correlation score in different intervals, preserving the advantages of high correlation, and not compressing low correlation scores too much, maintaining the diversity of correlation.
[0054] For example, in the morning mode scenario, when calculating the correlation score of the intelligent window curtain (central node) and the adjacent nodes, the initial correlation score of the window curtain and the light, temperature and humidity sensor is high (frequent linkage in the morning), and the gentle activation function will moderately increase these scores, making their weights remain strong but not over-amplified. The initial correlation score of the window curtain and the coffee machine, router is low (indirect correlation), and the function only slightly compresses these scores, making their weights remain at a weak but perceptible level. This gentle processing can preserve potential correlations such as window curtains and coffee machines (such as the user may set up a new rule "window curtain opens, coffee machine starts" later), leaving room for flexible expansion of the scene, and avoiding the loss of valuable weak correlation information due to excessive screening.
[0055] Further optionally, the scene-oriented activation function is not fixed and unchangeable, but automatically adjusts the activation strategy in real time with the scene. When the system switches from leaving home and arming (security scene) to home viewing (daily entertainment scene), the activation function will change from steep to flat, and the processing of the correlation score will change from strong differentiation to weak adjustment. At this time, in the association between the center node (such as a smart TV) and the adjacent nodes, the high correlation score of the TV and the router, the sound, will be moderately strengthened, while the low correlation score with the security camera will not be excessively compressed, ensuring that in the viewing scene, the core association of the network and the audio and video equipment is clear, while the potential interaction with other devices is preserved. Through this scene-adaptive activation strategy, the spatial attention weight can dynamically match the functional requirements of different scenes, focusing on the core association in key scenes, while preserving the diversity of associations in flexible scenes, providing a more practical basis for subsequent edge weight updates and feature aggregation.
[0056] Further, the parameter determination of the scene-oriented activation function needs to combine the scene core demand, node association characteristics and dynamic adaptation logic, and is realized through a three-order mechanism of "scene characteristic anchoring-basic parameter preset-dynamic calibration optimization", to ensure that the parameters fit the essence of the scene and flexibly respond to changes. First, anchor the parameter direction based on the core demand of the scene. The core task of different scenes determines the behavior tendency of the activation function. The security scene needs to accurately identify key linkages, and the parameter direction focuses on improving the steepness of the function. By amplifying the high segment gradient and compressing the low segment output, the weight difference between strong and weak correlations is strengthened. The daily scene needs to retain potential correlations, and the parameter direction shifts to reduce the steepness, slowing down the growth of the high segment and widening the low segment interval, to avoid excessive filtering of weak correlations. Second, preset the basic parameters combined with historical association data. By analyzing the historical interaction data in the scene, the initial range of the parameters is determined: for typical strong correlations (such as the linkage between door locks and cameras in the security scene), the high segment threshold is preset slightly lower than the average score, to ensure that strong correlations are effectively amplified. For non-critical weak correlations (such as the correlation between door locks and lights in the security scene), the low segment suppression threshold is set slightly higher than the average score, to reduce the interference of irrelevant nodes. Cross-scene universal parameters (such as the overall scaling ratio) are determined through offline verification, for example, testing the impact of different steepness coefficients on the accuracy of security anomaly warning, and selecting the value with the lowest false positive rate. Finally, realize scene adaptation optimization through dynamic calibration. The parameters are adjusted in real time as the scene changes. When the scene changes (such as from daily leisure to home defense), the corresponding scene parameter template is automatically called, and the suppression threshold is fine-tuned based on the current online device status to avoid interference from offline devices; the downstream task effect is monitored for reverse optimization, if the key correlation in the security scene is underestimated, the high segment threshold is lowered, and if the weak correlation interference is frequent, the low segment threshold is raised. For node type differentiation correction, the suppression threshold of environmental nodes is relaxed in the daily scene to retain potential correlations, and the basic parameters of control nodes are maintained to prevent misoperation. This mechanism allows the parameters to be rooted in the essence of the scene and historical laws, and to respond to real-time changes through dynamic adjustment, ultimately achieving the effect of "clearer correlation weight for highlighted associations and more thorough interference for weakened ones", providing reliable support for the accurate extraction of spatial association features.
[0057] In step 139, the edge weights of the base graph structure are updated by the correlation strength dynamic calibration mechanism, so that the graph atlas can reflect the current state of node correlation in real time. The correlation strength dynamic calibration mechanism is used to update the edge weights of the base graph structure in real time, so that the graph atlas can accurately reflect the current state of the correlation between nodes. In implementation, this mechanism takes the real-time attention weight generated by the spatial attention module as the core basis, and dynamically fuses it with the original initial edge weight (based on historical data statistics of correlation strength) in the base graph structure. For node pairs with frequent real-time interaction and close functional dependence, such as high-frequency data transmission between the TV and the router in the viewing scene, the calibration process will give priority to real-time attention weight, significantly increase the edge weight, and highlight the current strong correlation state. For node pairs with stable historical correlation but less real-time interaction, such as the smart socket of the long-leave machine and the router, the calibration process balances the initial weight and the real-time weight, preserving the basic correlation while avoiding excessive weakening. In addition, when the node state changes abruptly (such as device offline or scene switching), the mechanism will trigger a quick calibration to immediately reduce the correlation weight of the abnormal node, ensuring that the edge attribute is always synchronized with the actual running correlation strength of the device, so that the graph atlas can dynamically respond to real-time changes in device correlation in the apartment. Further optionally, the spatial attention weight generated in step 138 is fused with the initial weight of the original edge in the base graph structure: for node pairs with close functional dependence or frequent real-time interaction (such as air conditioner and temperature and humidity sensor), the attention weight is given priority to update the edge weight, strengthening the current strong correlation. For node pairs with stable historical correlation but less real-time interaction (such as the smart socket of the long-leave machine and the router), the initial weight and the attention weight are balanced to preserve the basic correlation. For example, in the viewing scene, the real-time interaction between the TV and the router is enhanced, and the edge weight is dynamically increased with the increase of the attention weight; when the TV is turned off, the edge weight gradually falls with the weakening of the correlation, so that the edge attribute of the graph structure is always synchronized with the actual correlation strength of the device.
[0058] In step 1310, the spatial correlation feature is generated by the core adjacency focused aggregation mechanism to condense the key interaction relationship between nodes. The core adjacency focused aggregation mechanism generates a spatial correlation feature that can condense spatial interaction relationships by screening key correlation nodes and aggregating their features. In implementation, the mechanism first prioritizes the adjacency nodes of each center node according to the weight values output by the spatial attention module, and screens the core adjacency nodes with high weight ranking (usually the top 30% of nodes in terms of weight proportion), ignoring redundant nodes with low weight to reduce irrelevant information interference. In the feature aggregation process, the system will strengthen local correlation by combining the physical location information of the nodes: devices in the same space area (such as lights, televisions, and cameras in the living room) will be given higher spatial correlation weights due to their close physical distance and frequent functional cooperation; cross-area devices (such as living room air conditioners and bedroom sensors) will weaken spatial influence and focus on functional correlation. Through this focused core node and space layout-considered aggregation method, the generated spatial correlation feature can clearly reflect the close interaction logic between the center node and the core adjacency node, and also preserve the coordination rules of devices in the spatial distribution, providing precise spatial dimension feature support for the subsequent construction of joint operation graph. Further optionally, for each center node, high-weight adjacency nodes (such as the top 30% of core correlation nodes) are screened according to the updated spatial attention weight, and low-weight redundant nodes are ignored, and then the features of the screened adjacency nodes are weighted and aggregated. In the aggregation process, the spatial location correlation information of the nodes is also preserved. For example, when devices in the same area (such as lights, televisions, and cameras in the living room) are aggregated, the spatial location feature will strengthen their local correlation weight. Cross-area devices (such as living room air conditioners and bedroom sensors) will weaken spatial influence and focus on functional correlation. The finally generated spatial correlation feature can not only reflect the close interaction between the node and the core adjacency node, but also preserve the spatial distribution rules, and completely depict the spatial coordination relationship between devices.
[0059] In step S104, a multi-agent deep deterministic policy gradient algorithm (MADDPG) is used to make network resource allocation and device control decisions on the joint operation graph, and to realize scene adaptive optimization in combination with the candidate service scene, to obtain a linkage strategy of the target apartment.
[0060] In the embodiments of the present application, the linkage strategy at least includes: a network resource dynamic allocation scheme, an intelligent device control instruction set, an inter-device linkage relationship parameter, a network and device linkage relationship parameter, a dominant scene identifier, and a scene linkage parameter.
[0061] Taking an apartment management system as an example, MADDPG is a reinforcement learning algorithm suitable for multi-role collaborative decision-making. The core of the algorithm is to let multiple agents collaborate to complete network resource allocation and device control tasks in a dynamic environment through a mechanism of centralized training and distributed execution. The algorithm abstracts key decision-making roles in the apartment management scenario into independent agents, such as a network resource scheduling agent (responsible for bandwidth allocation and delay optimization), a device control agent (responsible for starting and stopping devices such as lights and air conditioners, and adjusting their parameters), and a scenario coordination agent (responsible for identifying the dominant scenario and triggering the linkage logic). Each agent has an independent policy network (generating specific decisions, such as allocating 80% bandwidth to a TV or setting the air conditioner to 26°C) and a value network (evaluating the effectiveness of the decision), while sharing a centralized critic network (Critic) that evaluates the overall value of the multi-agent joint decision based on the global state, providing a global perspective for policy optimization. For example, all device operating states, network load, and current scenario identification.
[0062] MADDPG guides agents to learn the optimal linkage strategy through a reward mechanism, with the reward signal closely related to the core goals of apartment management. The reward design focuses on three types of value: resource utilization efficiency, measured by bandwidth allocation accuracy and device energy consumption reduction, with high rewards indicating that resources are not wasted; scenario adaptation effect, evaluated by dominant scenario identification accuracy and device linkage response speed, with high rewards representing a smooth scenario experience; and system stability, measured by network delay compliance duration and device fault-free operation duration, with high rewards indicating that decisions have not caused abnormalities. Agents accumulate experience through continuous interaction with the environment, the critic network evaluates the pros and cons of decisions based on global rewards, and the policy network is optimized in reverse, for example, when there is a lack of resources for high-priority scenarios, relevant agents are prompted to adjust their allocation strategies. MADDPG has strong dynamic adaptability and can accurately respond to real-time changes in the apartment scenario. When the scenario changes with user behavior, time, or environment, each agent needs to quickly adjust its decision logic. The algorithm stores historical interaction data in an experience replay pool and uses a target network to stabilize the learning process, avoiding policy shocks caused by sudden changes in the scenario. For example, when a security alarm scenario is detected, the scenario coordination agent triggers a high-priority signal, the network scheduling agent quickly tilts the bandwidth to security devices, and the device control agent synchronously starts the linkage action. These collaborative decisions do not need to be programmed by humans, but are dynamically generated by the algorithm through learning. This multi-agent collaboration mechanism integrates device control, resource allocation, and scenario linkage tasks, and outputs linkage strategies that meet the needs, including network resource allocation plans and device control instruction sets, achieving an upgrade from passive response to active optimization in intelligent management.
[0063] As an optional embodiment, in step S104, a multi-agent deep deterministic policy gradient algorithm is used to make network resource allocation and device control decisions for the joint operation graph, and scene adaptive optimization is achieved in combination with the candidate service scenarios to obtain the linkage strategy of the target apartment, including: Step 21, based on the node classification, function division, time sequence correlation characteristics and spatial correlation characteristics of the joint operation map, a distributed multi-agent decision system is constructed, which includes network resource scheduling agent, device cooperative control agent and scene adaptation coordination agent. Among them, the network resource scheduling agent is used to manage independent router nodes, switch nodes and network device node clusters constructed by multiple network device nodes, and is responsible for the bandwidth dynamic allocation, routing path optimization and priority queue scheduling of each network device node; multiple intelligent device nodes are divided into different types of intelligent device node clusters according to the device type, and the device cooperative control agent is used to manage each type of intelligent device node cluster, and is responsible for the device running mode switching, threshold parameter adjustment and operation time sequence control of each intelligent device node; the scene adaptation coordination agent is responsible for conflict resolution and target collaborative optimization of cross-agent decision based on time sequence correlation characteristics and spatial correlation characteristics.
[0064] Step 22, configure a differentiated action space for each agent, the action space of the network resource scheduling agent includes bandwidth dynamic allocation, routing path optimization and priority queue scheduling, the action space of the device cooperative control agent includes a variety of mixed actions, and the mixed actions at least include device running mode switching, threshold parameter adjustment and operation time sequence control, and the action space of the scene adaptation coordination agent includes conflict resolution and target collaborative optimization of cross-agent decision.
[0065] Step 23, construct a MADDPG network model fused with graph structure characteristics, set an actor network based on graph attention mechanism for each agent, take the subgraph features associated with each agent in the joint operation map as the input of the actor network, extract the correlation characteristics between nodes through the graph convolution layer, and output the deterministic action containing continuous control quantity and discrete operation instruction through the full connection layer. Configure a global centralized critic network, the critic network inputs the observation features, action output of all agents and the feature vector of the candidate service scene, models the interaction relationship between agents and the scene adaptation degree through the multi-head attention mechanism, and outputs the joint action value estimation. Configure a target actor network and a target critic network for each agent, set the network parameter soft update rate based on the time decay factor, and ensure the stability of the training process.
[0066] Step 24, input the joint operation graph updated by the spatio-temporal graph attention network and the feature vector of the candidate service scene in real time, and output the initial decision action through the actor network of each agent based on the node real-time state feature and the historical association rule. The scene adaptation and coordination agent filters and integrates the initial decision action according to the device function dependency relationship and the priority of the candidate service scene, and optimizes the action combination through the value evaluation of the global critic network.
[0067] Step 25, if the node state change amount or the edge weight update amplitude of the joint operation graph is not detected to exceed the preset threshold, the linkage strategy adapted to the current state is generated based on the value evaluation and optimization of the action combination. If the node state change amount or the edge weight update amplitude of the joint operation graph is detected to exceed the preset threshold, the agent redecision mechanism is triggered, and the actor network is called again to generate the linkage strategy adapted to the current state. The linkage strategy includes a network resource dynamic allocation scheme, a device control instruction set, a linkage relationship parameter and a leading scene identifier.
[0068] Specifically, step 21, based on the node attributes and associated features of the joint operation map, the decision-making tasks are functionally split and assigned to specialized agents. The network resource scheduling agent focuses on resource management of the network device cluster, the device cooperative control agent is responsible for the operation and control of the device cluster, and the scene adaptation coordination agent coordinates cross-domain decision conflicts. Through functional division, each agent focuses on core tasks. Effectively, the system can cover network resource allocation, device control, scene coordination, and other full-dimensional needs, avoiding excessive load on a single agent decision, and improving decision-making relevance and efficiency. Step 22, according to the core responsibilities of the agent, set up a special action range, so that the action and the functional demand are accurately matched. The network resource scheduling agent focuses on resource allocation and path optimization, the device cooperative control agent covers device operation adjustment, and the scene adaptation coordination agent focuses on conflict resolution and target optimization. Differentiated action space ensures that each agent's decision-making does not exceed the boundaries and is not redundant. Effectively, it reduces irrelevant action interference, makes the agent's decision-making more focused on core goals, reduces action combination complexity, and improves decision-making accuracy. Step 23, integrate graph structure features to enhance the model's perception of node associations. The actor network extracts subgraph association features through graph convolution to generate actions, the centralized critic network evaluates joint decision value based on global state and scene features, and the target network maintains stable training through soft updates. With the help of graph attention and multi-head attention mechanisms, the model can capture node interaction and agent collaboration patterns. Effectively, the model can deeply explore the temporal and spatial associations in the graph, improving the global optimization capability of the decision-making, while the stable training mechanism reduces strategy shock and ensures model convergence quality. Step 24, based on real-time updated graph and scene features, each agent independently generates an initial decision based on historical patterns and current state, and then the scene adaptation coordination agent filters and integrates the initial decisions according to functional dependencies and scene priorities, and optimizes the action combination with the critic network value evaluation. Through the "independent decision-making + global optimization" mode, the response speed and decision quality are balanced. Effectively, the initial decision quickly responds to real-time state, and the integration and optimization process eliminates action conflicts, making the final action combination more suitable for scene requirements and improving decision-making collaboration. Step 25, by monitoring the node state and edge weight change amplitude of the joint operation graph, dynamically switch the strategy generation mode. When the change is stable, generate the strategy based on the optimized action; when the change exceeds the threshold, trigger the re-decision mechanism to recalculate. Through state perception, the strategy generation is dynamically adapted. Effectively, under stable conditions, the strategy is quickly output to reduce computational overhead, and in the case of dramatic changes, the strategy is recalculated in time to avoid strategy failure, making the linkage strategy not only efficient in responding to stable scenes, but also accurately adapting to dynamic changes, improving the environmental adaptability of the system.
[0069] Further optionally, before each agent's actor network outputs an initial decision action based on the real-time state features of the nodes and historical association rules in step 24, a hierarchical reward mechanism including a basic performance reward, a scene adaptation reward, a collaborative optimization reward, and a constraint penalty term can be configured. The basic performance reward is calculated based on the network transmission delay reduction rate, the device response speed improvement rate, and the resource utilization rate standard rate; the scene adaptation reward is calculated according to the matching degree of the current device combination state and the candidate service scene, and the scene switching response time; the collaborative optimization reward is calculated based on the instruction execution synchronization rate between devices, and the network and device resource demand matching degree; the constraint penalty term is triggered when the network load is out of limit, the device energy consumption is out of limit, or the scene conflict occurs, and is used for negative reward; the hierarchical reward mechanism realizes the adaptive adjustment of each reward term and the constraint penalty term through a dynamic weight coefficient, and the weight coefficient is determined based on the priority of the candidate service scene. In the initial stage, the actor network and critic network parameters are initialized using a pre-training data set containing historical running data and artificial optimization decision samples. In the exploration stage, each agent introduces Gaussian noise in the action space for exploration, and stores a five-tuple sample containing four dimensions in an experience replay pool, where the first dimension element is the current joint running graph state, the second dimension element is the agent action set, the third dimension element is the reward value, and the fourth dimension element is the next state after the execution of the action. In the iterative optimization stage, a batch of samples are randomly sampled from the experience replay pool, the target Q value is calculated through the target critic network, and the current critic network is updated based on the time difference error; the actor network parameters are updated using the Q value gradient of the current critic network through the deterministic policy gradient algorithm; and the target network and the current network parameters are periodically synchronized through the soft update formula. In the scene migration stage, scene feature masks are constructed for different candidate service scenes, and network parameters are fine-tuned through transfer learning to adapt to scene-specific requirements. In this way, the multi-dimensional reward precisely guides the optimization direction of the agent, the basic performance reward improves resource utilization and response speed, the scene adaptation reward enhances scene matching accuracy, the collaborative optimization reward promotes the collaborative efficiency of devices and networks, and the constraint penalty term reduces the out-of-limit and conflict. The dynamic weight is adaptively adjusted based on the scene priority to realize balanced optimization of multiple objectives. The network is pre-trained using historical data and artificial samples to accelerate model convergence, reduce initial exploration cost, improve initial decision quality, and reduce the risk of blind training. Gaussian noise exploration increases action diversity and avoids local optimization. The experience replay pool stabilizes the sample distribution, ensures the richness and representativeness of the training data, and enhances the model generalization ability. The time difference error update and the deterministic policy gradient algorithm ensure the precision of the network parameter optimization direction. The soft update synchronizes the target network and the current network, reduces training shock, and improves the stability and convergence efficiency of policy iteration.The scene feature mask and the transfer learning fine-tuned parameters enable the model to quickly adapt to different scene-specific requirements, reduce cross-scene retraining costs, and improve the flexible adaptation capability of the system in various scenes.
[0070] Further optionally, in the scene migration stage, a scene feature mask is constructed for different candidate service scenes, and network parameters are fine-tuned through transfer learning to adapt to scene-specific requirements, comprising: Step 31: The candidate service scene is analyzed for features, and scene core features are extracted, including scene type identification, core functional device set, resource demand threshold, time correlation feature, and user behavior preference feature. A high-dimensional scene feature vector is constructed based on the scene core features, wherein each dimension corresponds to a key attribute parameter of the scene. Step 31 analyzes the core attributes of the candidate service scene, extracts key features such as type identification, core devices, resource requirements, time correlation, and user preferences, and converts the abstract scene into a structured high-dimensional feature vector, enabling the scene characteristics to be quantitatively represented. This provides a standardized input for subsequent scene perception and migration, enabling the model to accurately identify the essential differences between different scenes and laying a data foundation for mask construction and similarity calculation.
[0071] Step 32: For different candidate service scenes, based on the importance ranking of scene core features, high weight coefficients are assigned to strongly related dimensions in the scene feature vector, and low weight coefficients are assigned to weakly related or irrelevant dimensions, forming a binary or continuous value weight mask. The mask is used to strengthen the input weight of key scene features and weaken interfering features in the agent observation space. Step 32 assigns high weights to strongly related dimensions and low weights to weakly related or irrelevant dimensions in the feature vector based on the importance ranking of scene core features, forming a weight mask that strengthens the agent's attention to key scene features and weakens interfering information. This improves the model's perception sensitivity to scene core requirements, reduces the interference of irrelevant node features, and enables the agent to focus on the essential needs of the scene.
[0072] Step 33: Calculate the scene similarity between the new candidate service scene and the historical trained scenes based on the scene feature vector. The scene similarity is quantified by cosine similarity or Euclidean distance, and the model parameters corresponding to the highest similarity of the pre-set number of historical scenes are selected as the initial parameters for transfer. Step 33 quantifies the similarity between the new scene and the historical trained scenes through cosine similarity or Euclidean distance, selects the most similar historical scene model parameters as the initial value for transfer, and reuses the learning experience of existing scenes. The technical effect is to reduce the data dependence and time cost of new scene training, utilize historical knowledge to accelerate model adaptation, and avoid the inefficient problem of starting from scratch.
[0073] Step 34, using the parameter freezing and hierarchical fine-tuning strategy for transfer learning; wherein the bottom layer graph convolution layer and feature extraction layer parameters of the actor network and critic network are frozen to retain the general node correlation feature extraction capability; the top layer fully connected layer and scene adaptation layer parameters are unfrozen, and a mixed training set is formed by the real-time running data of the new scene and the labeled data of the similar historical scene, and small batch iterative fine-tuning is performed. Step 34 freezes the bottom layer graph convolution layer and feature extraction layer parameters to retain the general node correlation feature extraction capability, and unfreezes the top layer adaptation layer parameters, fine-tunes with the mixed data of the new scene and the similar historical scene, and balances the general ability and the scene specificity. Thus, the forgetting of the bottom layer general knowledge is avoided, the new scene characteristics are efficiently adapted, the overfitting risk is reduced, and the fine-tuning efficiency and stability are improved.
[0074] Step 35, in the fine-tuning process, a scene adaptation loss function is introduced, which includes a scene feature matching loss and a decision consistency loss; the scene feature matching loss is calculated based on the difference between the current scene feature vector and the scene perception feature output by the model, and the decision consistency loss is calculated based on the deviation between the model decision and the artificial expert example decision in the new scene. Step 35 constrains the perception accuracy of the model to the new scene feature through the scene feature matching loss, and aligns the model decision with the expert example through the decision consistency loss, and the double constraints guide the model to accurately adapt to the new scene. Thus, the capture accuracy of the model to the new scene feature and the decision rationality are improved, and it is ensured that the model decision after fine-tuning conforms to the scene logic and expert experience.
[0075] Step 36, set the fine-tuning termination condition, when the scene adaptation loss function value of the consecutive preset rounds is lower than the threshold value, or the basic performance reward and the scene adaptation reward in the new scene are stable in the preset interval, terminate the fine-tuning and save the fine-tuned network parameters; load the updated parameters to the actor network and critic network for online decision, and realize the rapid adaptation to the candidate service scene. Step 36 judges the model adaptation state by monitoring the loss function value and the reward stability, and terminates the fine-tuning when the loss is lower than the threshold value or the reward is stable, balancing the performance and computing resources. Thus, the waste of resources caused by excessive fine-tuning is avoided, the model is ensured to quickly reach a stable and effective state in the new scene, and the efficiency and reliability of scene migration are improved.
[0076] In step S105, the decision instruction is generated based on the linkage strategy, and the remote control of the network equipment and the intelligent equipment is realized by using the decision instruction. Specifically, in step S105, the linkage strategy is first parsed into a basic control layer and an intention abstraction layer to construct a double-layer instruction tree. The intention abstraction layer encapsulates the semantic target of the scene, including the scene type identifier and the core control intention. The basic control layer is disassembled into device-level atomic operation instructions, which include device identifier, operation action and parameter configuration. The high-security level instruction is encrypted by quantum key distribution (QKD), and a quantum key pair is generated and marked with an encryption identifier during instruction encapsulation. The device end realizes physical layer decryption through a pre-installed quantum receiver. Secondly, an instruction verification mechanism based on digital twinning is constructed, and the joint operation graph data is synchronized to the device digital twinning sandbox in real time to construct a virtual control environment. The to-be-executed instruction set is injected into the twinning system, and the device response process is simulated through a multi-modal physical engine. When a resource competition conflict is detected, a conflict resolution gravity model is constructed based on the criticality of the device function, and the dynamic adjustment force is calculated through the device weight and the time difference of instruction execution to adapt the parameters or optimize the time sequence of the low-weight device instruction. Thirdly, the federated learning is used to optimize the instruction transmission link, and the transmission quality model parameters are shared through the distributed edge instruction gateway to dynamically select the optimal transmission path. The quantum entangled photon synchronization mechanism is introduced for the multi-device collaborative instruction to generate a cross-device synchronization signal to ensure the concurrent execution of the instruction within a microsecond time window. The network equipment instruction is transmitted through the SSL / TLS encrypted TCP connection, and the intelligent equipment instruction is transmitted through the wireless communication link authenticated by the device unique token, and the instruction transmission state real-time monitoring mechanism is established. Then, the state feedback data of the received device after executing the instruction is uploaded after differential privacy processing by the edge gateway; the feedback data is compared with the expected target of the linkage strategy, and the instruction execution deviation rate is calculated. When the deviation rate exceeds the preset threshold, a compensation instruction containing the superposition state parameter is generated, and the device end executes the instruction according to the real-time environment state to collapse into a determined value, or generates an optimized instruction based on the joint operation graph. Finally, the data of the whole process of instruction generation, transmission, execution and feedback is encrypted by lattice-based cryptography and written into an anti-quantum blockchain, and each log block contains the instruction quantum hash value, the device quantum fingerprint and the environmental state vector. Based on the log data, a three-dimensional holographic control effect evaluation model is constructed to output the device response heat map and the resource flow trajectory, and a report on the accuracy of device control, the compliance rate of network resource configuration and the timeliness of scene response is generated regularly to support the iteration and optimization of the linkage strategy.
[0077] For example, in the family movie watching scene linkage strategy, the double-layer instruction tree construction link encapsulates the scene identification and high-definition playback, environment light adaptation core intent in the intent abstraction layer; the basic control layer is decomposed into atomic instructions, including television startup mode switching, sound volume adjustment to 50%, router allocating 80% bandwidth for television. High-security instructions involving door locks use QKD encryption to generate key pairs and mark encrypted identifiers, and the door lock decrypts through the pre-installed receiver physical layer. In the digital twin verification stage, the device state is synchronized to the twin sandbox to build a virtual environment, and the instruction set is injected to simulate execution. When a bandwidth conflict is detected, a conflict model is built based on the function criticality, with the television (weight 0.8) taking priority over the air conditioner (weight 0.3), and the air conditioner bandwidth is reduced from 50% to 20% and delayed after calculating the adjustment force. In instruction transmission optimization, federated learning selects the optimal path through the edge gateway; the television and sound instructions are synchronized using quantum entanglement to ensure concurrent execution within 10 microseconds. Network instructions are transmitted through SSL / TLS encrypted TCP, and device instructions are sent through token authentication wireless link, with transmission state synchronization monitoring. In the feedback stage, device data is received and uploaded after differential privacy processing. If the television bandwidth is actually 75% and the deviation rate is 6.25%, which exceeds the threshold, a compensation instruction is generated to add 5% bandwidth, and the device executes according to the real-time load. The whole process data is encrypted by lattice cryptography and written into the anti-quantum blockchain, and the log contains quantum hash, device fingerprint and environment vector. The three-dimensional evaluation model outputs a response heat map, showing that the television delay and sound synchronization error are both up to standard, and a control accuracy rate and resource compliance rate report is generated regularly to support policy iteration optimization.
[0078] Thus, the double-layer instruction tree improves instruction accuracy, and multi-layer encryption strengthens security barriers. Digital twin verification and conflict resolution models enhance execution reliability and reduce the risk of misoperation. Federated learning optimizes transmission paths, quantum synchronization ensures multi-device concurrency, and transmission efficiency and adaptability are improved. Feedback compensation forms a closed loop, anti-quantum blockchain logs are stored, and a three-dimensional evaluation model visually presents the results, providing data support for iterative linkage strategies, and comprehensively improving scene adaptation accuracy and system intelligence level.
[0079] In the embodiments of the present application, through deep fusion of spatio-temporal graph attention network and multi-agent reinforcement learning, dynamic coordination and control of apartment network and intelligent devices are realized. Not only can a global perception joint operation graph be constructed to accurately capture the spatio-temporal correlation of devices and networks, but also scene adaptive optimization can be achieved through multi-agent collaborative decision-making to improve resource utilization and service response speed. A closed-loop control mechanism can also be formed to ensure the real-time and reliability of the linkage strategy.
[0080] Please refer to Figure 2 , Figure 2The apartment network and intelligent device linkage system provided by the embodiment of the application comprises: a collection module, which is configured to collect network operation data of a network device and device operation data of an intelligent device deployed in a target apartment in real time; a decision module, which is configured to determine a candidate service scenario that can be realized by the target apartment based on the device operation data; a spatio-temporal graph attention network, which is configured to fuse the network operation data and the device operation data into a joint operation graph of the target apartment; the joint operation graph is configured to reflect a dynamic coupling relationship between a device operation state and a network operation state in real time; a MADDPG, which is configured to make network resource allocation and device control decisions on the joint operation graph, and realize scene adaptive optimization in combination with the candidate service scenario to obtain a linkage strategy of the target apartment; the linkage strategy at least comprises a network resource dynamic allocation scheme, an intelligent device control instruction set, an inter-device linkage relationship parameter, a network and device linkage relationship parameter, a dominant scene identifier, and a scene linkage parameter; and an execution module, which is configured to generate a decision instruction based on the linkage strategy, and realize remote control of the network device and the intelligent device by using the decision instruction. In some embodiments, the apartment network and intelligent device linkage system can be applied to a terminal device. It should be noted that, for the sake of brevity of description, the specific working process of the apartment network and intelligent device linkage system can refer to the corresponding process in the foregoing apartment network and intelligent device linkage method embodiment, and will not be described here in detail.
[0081] Please refer to Figure 3 , Figure 3 A structural schematic block diagram of a terminal device is provided by the embodiment of the application. As shown in Figure 3 , the terminal device 300 comprises a processor 301 and a memory 302, and the processor 301 and the memory 302 are connected through a bus 303, such as an I2C bus. Specifically, the processor 301 is configured to provide computing and control capabilities to support the operation of the entire terminal device. The processor 301 can be a central processing unit, and the processor 301 can also be other general-purpose processors, digital signal processors, application-specific integrated circuits, field programmable gate arrays, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Those skilled in the art can understand, Figure 3The structure shown in the figure is only a block diagram of part of the structure related to the embodiment of the present application, and does not constitute a limitation on the terminal device to which the embodiment of the present application is applied. The specific server can include more or fewer components than those shown in the figure, or combine certain components, or have a different component arrangement. The processor is configured to run a computer program stored in the memory and implement any one of the apartment network and smart device linkage methods provided by the embodiments of the present application when the computer program is executed. It should be noted that those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the terminal device described above can refer to the foregoing apartment network and smart device linkage method embodiments, and will not be described herein.
[0082] The embodiment of the present application also provides a storage medium for computer readable storage, the storage medium stores one or more programs, the one or more programs can be executed by one or more processors to implement the steps of any one of the apartment network and smart device linkage methods provided by the embodiments of the present application.
Claims
1. An apartment network and smart device linkage method, characterized in that, The method comprises: For the target apartment, real-time collection of network operation data of network equipment and equipment operation data of intelligent equipment deployed in the target apartment; Based on the equipment operation data, determining a candidate service scenario currently achievable by the target apartment; Through a spatio-temporal graph attention network, fusing the network operation data and the equipment operation data into a joint operation graph of the target apartment; the joint operation graph is used to reflect a dynamic coupling relationship between an equipment operation state and a network operation state in real time; Using a multi-agent deep deterministic policy gradient algorithm (MADDPG), making network resource allocation and equipment control decisions on the joint operation graph, and combining the candidate service scenario to achieve scene adaptive optimization, to obtain a linkage strategy of the target apartment; the linkage strategy at least includes: a network resource dynamic allocation scheme, an intelligent equipment control instruction set, an inter-equipment linkage relationship parameter, a network and equipment linkage relationship parameter, a dominant scene identifier, and a scene linkage parameter; Based on the linkage strategy, generating a decision instruction, and using the decision instruction to realize remote control of the network equipment and the intelligent equipment.
2. The method of claim 1, wherein, The joint operation graph of the target apartment is fused by a spatio-temporal graph attention network, comprising: Pretreatment of the collected network operation data and equipment operation data, the pretreatment including data cleaning, outlier filtering, time granularity alignment, and feature standardization, to obtain standardized network feature data and standardized equipment feature data; wherein the network operation data includes real-time bandwidth occupancy rate, data transmission delay, packet loss rate, number of connected devices, and time sequence fluctuation characteristics of the network equipment, and the equipment operation data includes operating state parameters, energy consumption data, operation response time, inter-device interaction frequency, and function state identifier of the intelligent equipment; Constructing a basic graph structure of the target apartment, the nodes of the basic graph structure including network equipment nodes and intelligent equipment nodes, the edges of the basic graph structure being constructed based on physical connection relationship, data interaction relationship, and functional dependency relationship, and the initial edge weight being set based on historical data interaction frequency and correlation strength; the node attributes of each node in the joint operation graph include first real-time operation features of the corresponding network equipment nodes or second real-time operation features of the intelligent equipment nodes, and the edge attributes include spatio-temporal correlation strength and time-varying law characteristics between nodes; Inputting the standardized network feature data and the standardized equipment feature data into the spatio-temporal graph attention network, capturing dynamic dependency relationships of node features in the time dimension through a time attention module to generate time sequence correlation features; calculating correlation weights between different nodes through a spatial attention module, updating edge weights of the basic graph structure based on the correlation weights, and generating spatial correlation features; Fusing the time sequence correlation features and the spatial correlation features to obtain a joint operation graph containing node real-time state, spatio-temporal correlation weight, and dynamic interaction relationship.
3. The method of claim 2, wherein, The basic graph structure of the target apartment is constructed, comprising: The network devices and intelligent devices in the target apartment are node-modeled, and each network device and intelligent device is assigned a unique node identifier associated with device physical location information, device type attributes, and basic function parameters. The network device nodes include router nodes, switch nodes, gateway nodes, and wireless access point nodes, and the intelligent device nodes include home control nodes, environment monitoring nodes, security monitoring nodes, audio and video entertainment nodes, and energy consumption management nodes. The connection relationship between network devices is obtained through network topology detection technology, the access relationship between intelligent devices and network devices is obtained through device installation layout and hardware interface mapping, and a first type of edge is created for node pairs with direct physical connection. The initial weight of the first type of edge is set based on the connection stability parameter. The data packet transmission records between devices within the dynamic time window are collected, the data interaction frequency, data volume, and interaction time length features are extracted, and a second type of edge is created for node pairs with data interaction. The initial weight of the second type of edge is set based on the weighted calculation result of the data interaction frequency and the average interaction time length. The functional synergy relationship of intelligent devices in each service scenario is obtained, the control instruction dependency, state feedback dependency, and resource sharing dependency between devices are identified, and a third type of edge is created for node pairs with functional dependency. The initial weight of the third type of edge is set based on the functional dependency strength level quantization value. The first type of edge, the second type of edge, and the third type of edge are integrated to form a basic graph structure containing node attributes and edge attributes, and a mapping relationship between the graph structure and the device physical identifier is established for subsequent graph dynamic updating and node association query.
4. The method of claim 2, wherein, The standardized network feature data and standardized device feature data are input into the spatio-temporal graph attention network, the dynamic dependency relationship of each node feature in the time dimension is captured through the time attention module, and the time sequence correlation feature is generated, including: The standardized network feature data and standardized device feature data are time-aligned by node to construct the feature time sequence of each node, which contains multi-dimensional feature values within a preset time span. The feature time sequence of each node is divided into consecutive time windows, and each time window contains a preset number of time step features. The feature sequence in each time window is calculated by the multi-head self-attention mechanism of the time attention module to obtain the attention score. The feature vectors of each time step are used as query vectors, key vectors, and value vectors to generate query matrices, key matrices, and value matrices through linear transformation. The attention score between different time steps is calculated, and the attention score is determined based on the similarity of the query vector and the key vector and the time decay factor. The attention score is normalized to obtain the time attention weight, and the feature vectors of each time step are weighted and aggregated based on the time attention weight to generate the local time sequence feature of each node. The local time sequence feature is input into a gating cycle unit to capture long-time dependence across a time window, and to fuse time sequence correlation features containing short-term fluctuation features and long-term trend features, which carry dynamic change rules of nodes in the time dimension.
5. The method of claim 4, wherein, The spatial attention module is used to calculate correlation weights between different nodes, and the edge weights of the basic graph structure are updated based on the correlation weights to generate spatial correlation features, including: The time sequence correlation features of the nodes processed by the time attention module are extracted, and the type attributes and physical location information of the nodes in the basic graph structure are combined to construct node comprehensive feature vectors; The spatial attention module is used to calculate correlation weights between node pairs with the node comprehensive feature vectors as input. For each center node, the attribute feature vector of each center node is used as a query vector, and the attribute feature vectors of the respective adjacent nodes of each center node are used as key vectors. The similarity between the query vector and each key vector is calculated, and the initial correlation score is generated by combining the initial weights of the corresponding edges in the basic graph structure and the node type matching degree. The spatial attention weights are obtained by nonlinear transformation of the initial correlation score through an activation function, and the spatial attention weights reflect the influence degree of different adjacent nodes on the center node. The weights of the corresponding edges in the basic graph structure are updated based on the spatial attention weights to form dynamically updated spatial correlation edge attributes. For each center node, the attribute feature vectors of the adjacent nodes are weighted and aggregated according to the updated spatial attention weights to fuse spatial correlation features containing spatial interaction relationships between nodes.
6. The method of claim 1, wherein, The multi-agent deep deterministic policy gradient algorithm (MADDPG) is used to make network resource allocation and device control decisions for the joint operation graph, and scene adaptive optimization is realized in combination with the candidate service scenarios to obtain the linkage strategy of the target apartment, including: Based on the node classification, function division, time sequence correlation features, and spatial correlation features of the joint operation graph, a distributed multi-agent decision system is constructed, which includes network resource scheduling agents, device collaborative control agents, and scene adaptation coordination agents. The network resource scheduling agents are used to manage independent router nodes, switch nodes, and network device node clusters constructed by multiple network device nodes, and are responsible for dynamic bandwidth allocation, routing path optimization, and priority queue scheduling of each network device node. Multiple intelligent device nodes are divided into different types of intelligent device node clusters according to device types, and the device collaborative control agents are used to manage each type of intelligent device node cluster and are responsible for device operation mode switching, threshold parameter adjustment, and operation time sequence control of each intelligent device node. The scene adaptation coordination agents are responsible for conflict resolution and target collaborative optimization across agent decisions based on the time sequence correlation features and the spatial correlation features. The network resource scheduling agent is configured with a differentiated action space, which includes bandwidth dynamic allocation, routing path optimization, and priority queue scheduling; the device collaborative control agent is configured with a differentiated action space, which includes a plurality of hybrid actions, such as device operating mode switching, threshold parameter adjustment, and operation timing control; and the scene adaptation coordination agent is configured with a differentiated action space, which includes conflict resolution and target collaborative optimization across agents. A MADDPG network model is constructed by fusing graph structure features, an actor network based on a graph attention mechanism is set for each agent, the subgraph features associated with each agent in the graph are input into the actor network for joint operation, the correlation features between nodes are extracted through a graph convolution layer, and a deterministic action including a continuous control quantity and a discrete operation instruction is output through a fully connected layer; a global centralized critic network is configured, the critic network inputs the observation features, action outputs of all agents, and feature vectors of the candidate service scenarios, models the interaction relationship between agents and the scene adaptation degree through a multi-head attention mechanism, and outputs a joint action value estimate; a target actor network and a target critic network are configured for each agent, the network parameter soft update rate is set based on a time decay factor to ensure the stability of the training process; The joint operation graph updated by the spatio-temporal graph attention network and the feature vector of the candidate service scenario are input in real time, the actor network of each agent trained is used to make online decisions based on the real-time state features and historical correlation rules of the nodes and output initial decision actions; the scene adaptation coordination agent filters and integrates the initial decision actions based on the device function dependency relationship and the priority of the candidate service scenario, and optimizes the action combination through the value evaluation of the global critic network; If the node state change quantity or the edge weight update amplitude of the joint operation graph does not exceed the preset threshold, a linkage strategy adapted to the current state is generated based on the value evaluation and optimization of the action combination; if the node state change quantity or the edge weight update amplitude of the joint operation graph exceeds the preset threshold, an agent redetermination mechanism is triggered to generate a linkage strategy adapted to the current state by re-invoking the actor network; the linkage strategy includes a network resource dynamic allocation scheme, a device control instruction set, a linkage relationship parameter, and a dominant scene identifier.
7. The method of claim 6, wherein, Before the actor network of each agent trained is used to make online decisions based on the real-time state features and historical correlation rules of the nodes and output initial decision actions, the following steps are further included: The layered reward mechanism comprises a basic performance reward, a scene adaptation reward, a collaborative optimization reward, and a constraint penalty term. The basic performance reward is calculated based on a network transmission delay reduction rate, a device response speed improvement rate, and a resource utilization rate standard rate. The scene adaptation reward is calculated according to a matching degree of a current device combination state and a candidate service scene and a scene switching response time. The collaborative optimization reward is calculated based on an instruction execution synchronization rate between devices and a network and device resource demand matching degree. The constraint penalty term is used for negative reward when the network load is out of limit, the device energy consumption is out of limit, or the scene conflicts. The layered reward mechanism realizes adaptive adjustment of each reward term and the constraint penalty term through a dynamic weight coefficient, and the weight coefficient is determined based on a priority of the candidate service scene. In the initial stage, the actor network and the critic network parameters are initialized by using a pre-training data set, and the pre-training data set comprises historical running data and artificial optimization decision samples. In the exploration stage, each agent introduces Gaussian noise in the action space for exploration, and stores a sample five-tuple comprising four dimensions in an experience replay pool. The first dimension element is a current joint running graph state, the second dimension element is an agent action set, the third dimension element is a reward value, and the fourth dimension element is a next state after the action is executed. In the iterative optimization stage, a batch of samples are randomly sampled from the experience replay pool, a target Q value is calculated by using a target critic network, and the current critic network is updated based on a time difference error. A deterministic policy gradient algorithm is used to update the actor network parameters by using the Q value gradient of the current critic network. The target network and the current network parameters are regularly synchronized by using a soft update formula. In the scene migration stage, scene feature masks are constructed for different candidate service scenes, and network parameters are fine-tuned by using transfer learning to adapt to scene specificity requirements.
8. The method of claim 7, wherein, In the scene migration stage, scene feature masks are constructed for different candidate service scenes, by using transfer learning to fine-tune network parameters to adapt to scene specificity requirements, comprising: performing feature analysis on the candidate service scene, extracting scene core features, and constructing a high-dimensional scene feature vector based on the scene core features, wherein each dimension corresponds to a key attribute parameter of the scene; for different candidate service scenes, based on the importance of the scene core features, assigning a high weight coefficient to the dimensions in the scene feature vector that are strongly related to the current scene, and assigning a low weight coefficient to the weakly related or irrelevant dimensions, to form a binary or continuous weight mask; the mask is used to strengthen the input weight of the key features of the scene in the agent observation space and weaken the interference features; based on the scene feature vector, calculating the scene similarity between the new candidate service scene and the historical trained scenes, and selecting the model parameters corresponding to the highest similarity of a preset number of historical scenes as the initial parameters for transfer. Adopt parameter freezing and hierarchical fine-tuning strategy for transfer learning; among them, freeze the bottom layer graph convolution layer and feature extraction layer parameters of actor network and critic network, retain the general node correlation feature extraction ability; unfreeze the top layer full connection layer and scene adaptation layer parameters, form a mixed training set with real-time running data of new scene and labeled data of similar historical scene, and perform small batch iterative fine-tuning; In the fine-tuning process, introduce a scene adaptation loss function, which includes scene feature matching loss and decision consistency loss; the scene feature matching loss is calculated based on the difference between the current scene feature vector and the scene perception feature output by the model, and the decision consistency loss is calculated based on the deviation between the model decision and the artificial expert example decision in the new scene; Set the fine-tuning termination condition, when the scene adaptation loss function value of the continuous preset number of rounds is lower than the threshold, or the basic performance reward and scene adaptation reward in the new scene are stable in the preset interval, terminate the fine-tuning and save the fine-tuned network parameters; load the updated parameters to the actor network and critic network for online decision, realize the rapid adaptation to the candidate service scene.
9. The method of claim 1, wherein, The decision instruction is generated based on the linkage strategy, and the remote control of network equipment and intelligent equipment is realized by using the decision instruction, which includes: The linkage strategy is parsed into a basic control layer and an intention abstraction layer to construct a double-layer instruction tree; the intention abstraction layer encapsulates the scene semantic target, including scene type identification and core control intention; the basic control layer is disassembled into device-level atomic operation instructions, which include device identification, operation action and parameter configuration; high-security level instructions are encrypted using quantum key distribution (QKD), quantum key pairs are generated during instruction encapsulation and encrypted identification is marked, and the device end realizes physical layer decryption through a preinstalled quantum receiver; An instruction verification mechanism based on digital twinning is constructed, real-time synchronization of joint operation graph data to device digital twinning sandbox is realized, and a virtual control environment is constructed; the to-be-executed instruction set is injected into the twinning system, and the device response process is simulated through a multi-modal physical engine; when resource competition conflict is detected, a conflict resolution gravity model is constructed based on device function criticality, dynamic adjustment force is calculated through device weight and instruction execution time difference, and parameter adaptation or time sequence optimization is performed on low-weight device instructions; A federal learning optimized instruction transmission link is adopted, distributed edge instruction gateways share transmission quality model parameters, and the optimal transmission path is dynamically selected; a quantum entangled photon synchronization mechanism is introduced for multi-device collaborative instructions to generate cross-device synchronization signals to ensure that instructions are executed concurrently within a microsecond time window; network equipment instructions are transmitted through SSL / TLS encrypted TCP connections, and intelligent equipment instructions are transmitted through wireless communication links authenticated by device unique tokens, and an instruction transmission state real-time monitoring mechanism is established; The state feedback data after the receiving device executes the instructions is uploaded after differential privacy processing by the edge gateway; the feedback data is compared with the expected target of the linkage strategy, and the deviation rate of instruction execution is calculated; when the deviation rate exceeds the preset threshold, a compensation instruction containing superposition state parameters is generated, and the device end collapses into a determined value according to the real-time environment state, or generates an optimized instruction based on the joint operation graph; The whole-process data of instruction generation, transmission, execution and feedback is encrypted by lattice cryptography, and is written into an anti-quantum block chain for storage; each log block contains an instruction quantum hash value, a device quantum fingerprint and an environmental state vector; a three-dimensional holographic control effect evaluation model is constructed based on the log data, and a device response thermal map and a resource flow trajectory are output; a device control accuracy rate, a network resource configuration compliance rate and a scene response timeliness report are generated regularly to support iterative optimization of the linkage strategy.
10. An apartment network and intelligent device linkage system, characterized in that, The system comprises: A collection module configured to collect, in real time, network operation data of a network device and device operation data of an intelligent device deployed in a target apartment for the target apartment; A decision module configured to determine a candidate service scenario that can be realized by the target apartment based on the device operation data; to fuse the network operation data and the device operation data into a joint operation graph of the target apartment by means of a spatiotemporal graph attention network; to reflect a dynamic coupling relationship between a device operation state and a network operation state in real time by means of the joint operation graph; to perform network resource allocation and device control decision-making on the joint operation graph by means of a multi-agent deep deterministic policy gradient algorithm (MADDPG), and to realize scene adaptive optimization in combination with the candidate service scenario to obtain a linkage strategy of the target apartment; and to determine that the linkage strategy at least includes a network resource dynamic allocation scheme, an intelligent device control instruction set, a device-to-device linkage relationship parameter, a network-to-device linkage relationship parameter, a dominant scene identifier and a scene linkage parameter. An execution module configured to generate a decision instruction based on the linkage strategy, and to realize remote control of the network device and the intelligent device by means of the decision instruction.
Citation Information
Cited By
Intelligent operation management method and system based on multi-device cooperation
CN121397059A
Intelligent operation management method and system based on multi-device cooperation
CN121397059B
Smart home cooperative control method based on Internet of Things
CN121523088A
An intelligent home cooperative control method based on Internet of Things
CN121523088B
Intelligent door lock control method and system based on deep learning
CN121564835A