Intelligent building real-time person searching cooperative monitoring method, system, device and medium

CN122554775APending Publication Date: 2026-08-11SHANGHAI XINZHUXIN ELECTROMECHANICAL INTEGRATION CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-07
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

[0004]然而,上述集中式方案在实际应用中存在明显不足:一方面,所有视频数据均需传输至中心服务器,导致网络带宽占用高、响应延迟大,尤其在多目标并发追踪时服务器负载急剧上升;另一方面,当目标在不同楼层或区域之间移动时,容易产生定位盲区或轨迹中断,影响追踪的连续性与可靠性

Benefits of technology

1.本申请能够降低响应延迟,减少网络带宽占用,通过分布式边缘计算架构,将视频处理、特征提取和多模态融合任务下放到边缘节点,避免大量数据上传中心服务器,仅上传加密后的特征向量或融合后的坐标结果,可降低系统响应延迟并使网络带宽占用相比传统方案大幅降低;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122554775A_ABST
    Figure CN122554775A_ABST
Patent Text Reader

Abstract

This application relates to a method, system, device, and medium for real-time person-finding and collaborative monitoring in intelligent buildings, belonging to the field of smart property management technology. The method includes: receiving a target person-finding request; allocating the request to at least one edge computing node as a tracking node based on the load status of the edge computing nodes; controlling the allocated node to acquire video streams, Bluetooth beacon signals, and RFID signals within the coverage area; acquiring facial features extracted locally from the video stream by the node, and fusing them with the Bluetooth beacon and RFID signals in a multimodal manner to generate the target's real-time location; adjusting the video stream acquisition resolution and frame rate based on the confidence level of the facial features and the strength of RFID signals from adjacent floors; receiving the real-time location reported by the node, and generating the target's movement trajectory accordingly. This application can reduce system response latency, reduce network bandwidth usage, and improve positioning accuracy in complex environments.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of smart property management technology, and in particular to a method, system, device and medium for real-time person search and collaborative monitoring in smart buildings. Background Technology

[0002] With the continuous advancement of smart city construction, the demand for personnel management and security monitoring in intelligent buildings is increasing daily. How to achieve rapid and accurate real-time person location tracking in large buildings or complexes has become an important research topic in the field of building intelligence. Existing solutions generally face challenges related to response latency, system load, and multi-target concurrent processing capabilities.

[0003] To address the aforementioned issues, a common solution is to employ a centralized person-finding system. This system deploys surveillance cameras throughout the building and uploads all video streams to a central server. The server then runs facial recognition algorithms to centrally process and analyze the video streams, thereby enabling the location and tracking of the target. This solution, through centralized scheduling of computing resources, can accomplish a certain degree of personnel retrieval.

[0004] However, the aforementioned centralized approach has significant shortcomings in practical applications: Firstly, all video data must be transmitted to the central server, resulting in high network bandwidth consumption and large response latency, especially when multiple targets are tracked concurrently, causing a sharp increase in server load. Secondly, when the target moves between different floors or areas, positioning blind spots or trajectory interruptions can easily occur, affecting the continuity and reliability of tracking. In addition, signal interference or obstruction in complex environments can further reduce positioning accuracy. Summary of the Invention

[0005] The purpose of this application is to provide a real-time collaborative monitoring method for finding people in intelligent buildings, which can reduce system response latency, reduce network bandwidth usage, and improve positioning accuracy in complex environments.

[0006] Firstly, this application provides a real-time person-finding and collaborative monitoring method for intelligent buildings, which adopts the following technical solution: A real-time person-finding and collaborative monitoring method for intelligent buildings includes: Receive a missing person request for a target, and allocate the missing person request to at least one edge computing node as a tracking node based on the current load status of each edge computing node; Control the assigned edge computing nodes to acquire video streams, Bluetooth beacon signals, and RFID signals within the coverage area; The facial features extracted locally from the video stream by the edge computing node are obtained, and the facial features are fused with the Bluetooth beacon signal and the radio frequency identification signal in a multimodal manner to generate the real-time location of the target. The acquisition resolution and frame rate of the video stream are adjusted based on the confidence level of the extracted facial features and the radio frequency identification signal strength of adjacent floors. The system receives the real-time location reported by the edge computing node and generates the target's movement trajectory based on the real-time location.

[0007] By adopting the above technical solutions, receiving missing person requests and allocating tasks based on the load status of edge nodes can avoid overloading a single node; control nodes acquire multi-source signals to provide a data foundation for positioning; real-time location is generated by fusing facial features with Bluetooth and radio frequency signals locally on the node, reducing data transmission latency; resolution and frame rate are adjusted based on confidence level and radio frequency signals from adjacent floors to balance computing resources and recognition accuracy; and location is received and a movement trajectory is generated to form complete path information. All of these together achieve distributed processing, real-time location calculation, and trajectory construction for missing person tasks, reducing system response latency and network bandwidth consumption.

[0008] In a preferred embodiment, this application can be further configured such that: the step of receiving a missing person request for a target and allocating the missing person request to at least one edge computing node according to the current load status of each edge computing node includes: Obtain the CPU utilization, remaining memory, and network latency of each edge computing node; The CPU utilization, remaining memory, and network latency are each assigned a weight and then weighted and scored. The missing person request is assigned to the edge computing node with the highest weighted score.

[0009] By adopting the above technical solution, the CPU utilization, remaining memory, and network latency of each edge node are obtained to quantify the real-time carrying capacity of the node; the three factors are weighted and scored to comprehensively evaluate the suitability of the node; and the missing person request is assigned to the node with the highest score to ensure that the task falls on the device with the most abundant resources. This achieves intelligent task scheduling based on multi-dimensional load indicators, effectively avoiding single-point overload and improving system processing efficiency in multi-objective concurrent scenarios.

[0010] In a preferred embodiment, this application may be further configured as follows: after the steps of obtaining the facial features extracted locally from the video stream by the edge computing node, and performing multimodal fusion of the facial features with the Bluetooth beacon signal and the RFID signal to generate the real-time location of the target, the application further includes: Obtain the radio frequency identification signal strength of the floor corresponding to the real-time location of the target and the adjacent floors; When the radio frequency identification signal strength of the adjacent floor exceeds a preset strength threshold, it is determined whether the target has entered the floor transition zone; If the target enters the floor transition zone, data synchronization between the current edge computing node and the edge computing nodes of the adjacent floors is triggered, wherein the data synchronization includes the target's facial features and historical trajectory; The edge computing nodes of the adjacent floors are used as tracking nodes to receive the real-time location of the target reported by the adjacent floors, so as to form a continuous cross-floor trajectory.

[0011] By adopting the above technical solution, after generating the real-time location, the strength of radio frequency signals on adjacent floors is monitored to detect the target approaching the floor boundary in advance; when the signal exceeds the threshold, it is determined that the target has entered the transition zone, and a cross-floor handover is triggered in a timely manner; data synchronization between the current node and adjacent nodes is triggered to transmit facial features and historical trajectories to ensure tracking continuity; the adjacent node is set as the new master tracking node and continues to receive the location, forming a complete cross-floor trajectory. This reduces positioning interruptions during floor switching and enables continuous tracking of the target across different floors.

[0012] In a preferred embodiment, this application can be further configured as follows: the step of obtaining the facial features extracted locally from the video stream by the edge computing node, and performing multimodal fusion of the facial features with the Bluetooth beacon signal and the radio frequency identification signal to generate the real-time location of the target includes: The facial feature vectors are extracted from the detected facial regions in the video stream, and the relative coordinates are obtained by combining the camera calibration parameters; The received signal strength values ​​of multiple Bluetooth beacons are sorted according to the beacon identifier, and a first signal vector of fixed length is constructed. The signal strength values ​​of multiple RFID readers are sorted according to the reader identifier, and a second signal vector of fixed length is constructed. The first signal vector and the second signal vector are normalized. The normalized first signal vector, the second signal vector, and the face feature vector are concatenated and input into a pre-trained deep regression network to output the planar coordinates and floor number of the target.

[0013] By employing the above technical solution, facial feature vectors are extracted from the video and their relative coordinates are estimated in conjunction with camera parameters to provide visual positioning information; Bluetooth beacon signal strength is sorted to construct a fixed vector, unifying the beacon data format; RFID signal strength is sorted to construct a vector, aligning multi-source data; signal vectors are normalized to eliminate dimensional differences; all features are concatenated and input into a regression network to output planar coordinates and floor numbers, thus achieving the fusion and solution of multi-source information. This improves positioning accuracy and stability in complex environments, such as those with occlusion or signal fluctuations.

[0014] In a preferred embodiment, this application can be further configured such that: the step of adjusting the acquisition resolution and frame rate of the video stream based on the confidence level extracted from the facial features and the radio frequency identification signal strength of adjacent floors includes: The network quality parameters between the edge computing node and the central server, as well as the CPU utilization rate of the current edge computing node, are obtained in real time. When the network quality parameters are lower than a preset network threshold, or the CPU utilization rate is higher than a preset load threshold, the acquisition resolution and frame rate are reduced. When the confidence level of the extracted facial features is lower than a preset confidence threshold, or when the RFID signal strength of the adjacent floors is higher than a preset strength threshold, the acquisition resolution and frame rate are increased.

[0015] By adopting the above technical solution, network quality parameters and CPU utilization are additionally acquired when adjusting resolution and frame rate, expanding the feedback dimensions. When network quality is poor or CPU utilization is high, the acquisition parameters are reduced to free up bandwidth and computing resources. When face confidence is low or radio frequency signals from adjacent floors are strong, the acquisition parameters are increased to enhance recognition details. In this way, resolution and frame rate are dynamically adapted according to network and load conditions, avoiding excessive resource consumption while ensuring recognition accuracy and improving the system's adaptive capabilities.

[0016] In a preferred embodiment, this application can be further configured as follows: the step of obtaining the facial features extracted locally from the video stream by the edge computing node, and performing multimodal fusion of the facial features with the Bluetooth beacon signal and the radio frequency identification signal to generate the real-time location of the target includes: Monitor the illumination intensity and the ratio of detectable face area in the video stream; When the lighting conditions are lower than a preset brightness threshold, or the detectable area ratio of the face is lower than a preset area ratio, the weight of the face features in multimodal fusion is reduced, and the weight of the Bluetooth beacon signal and the radio frequency identification signal is increased accordingly. When the signal-to-noise ratio of the Bluetooth beacon signal or the RFID signal is lower than a preset signal-to-noise ratio threshold, the weight of the facial features is increased.

[0017] By employing the above technical solution, the system monitors the lighting conditions and face detectability of the video stream, and perceives environmental changes. When there is insufficient lighting or face detection fails, the weight of face feature fusion is reduced and the weights of Bluetooth and radio frequency are increased, shifting the reliance to non-visual signals. When the signal-to-noise ratio of Bluetooth or radio frequency signals falls below a threshold, the weight of face features is increased, restoring visual dominance. This achieves adaptive adjustment of multimodal fusion weights, ensuring stable positioning even when any signal source deteriorates, and enhancing robustness in complex scenarios.

[0018] In a preferred embodiment, this application may be further configured such that, after the step of assigning the missing person request to the edge computing node with the highest weighted score, it further includes: Send a stop tracking command to all assigned edge computing nodes; The edge computing node is instructed to stop extracting facial features and performing multimodal fusion on the video stream, and to restore the acquisition resolution and frame rate of the video stream to their initial low-power default values.

[0019] By adopting the above technical solution, the system determines the task termination condition upon receiving a search termination command or detecting that the target has not updated its location for an extended period. It then sends a stop tracking command to the allocated edge nodes, releasing their computing resources. The nodes are instructed to stop facial feature extraction and multimodal fusion, and their resolution and frame rate are restored to their initial low-power default values, returning to standby mode. This achieves timely resource reclamation and state reset after the task is completed, avoiding unnecessary computational overhead and reducing average power consumption during long-term operation.

[0020] Secondly, this application provides a real-time person-finding collaborative monitoring system for intelligent buildings, which adopts the following technical solution: A real-time person-finding collaborative monitoring system for intelligent buildings includes: Request allocation module: used to receive a missing person request for a target, and allocate the missing person request to at least one edge computing node as a tracking node according to the current load status of each edge computing node; Signal acquisition module: Used to control the assigned edge computing nodes to acquire video streams, Bluetooth beacon signals and RFID signals within the coverage area; Location generation module: used to obtain facial features extracted locally from the video stream by the edge computing node, and perform multimodal fusion of the facial features with the Bluetooth beacon signal and the radio frequency identification signal to generate the real-time location of the target; Video adjustment module: used to adjust the acquisition resolution and frame rate of the video stream based on the confidence level of the facial features extracted and the RFID signal strength of adjacent floors; Trajectory generation module: used to receive the real-time location reported by the edge computing node, and generate the movement trajectory of the target based on the real-time location.

[0021] Thirdly, this application provides an electronic device that adopts the following technical solution: Request allocation module: used to receive a missing person request for a target, and allocate the missing person request to at least one edge computing node as a tracking node according to the current load status of each edge computing node; Signal acquisition module: Used to control the assigned edge computing nodes to acquire video streams, Bluetooth beacon signals and RFID signals within the coverage area; Location generation module: used to obtain facial features extracted locally from the video stream by the edge computing node, and perform multimodal fusion of the facial features with the Bluetooth beacon signal and the radio frequency identification signal to generate the real-time location of the target; Video adjustment module: used to adjust the acquisition resolution and frame rate of the video stream based on the confidence level of the facial features extracted and the RFID signal strength of adjacent floors; Trajectory generation module: used to receive the real-time location reported by the edge computing node, and generate the movement trajectory of the target based on the real-time location.

[0022] An electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the above-described real-time collaborative monitoring method for finding people in intelligent buildings.

[0023] Fourthly, this application provides a computer storage medium, as follows: A computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the aforementioned real-time person-finding and collaborative monitoring method for intelligent buildings.

[0024] In summary, this application has the following beneficial technical effects: 1. This application can reduce response latency and network bandwidth usage. Through a distributed edge computing architecture, video processing, feature extraction and multimodal fusion tasks are offloaded to edge nodes, avoiding the uploading of a large amount of data to the central server. Only encrypted feature vectors or fused coordinate results are uploaded, which can reduce system response latency and significantly reduce network bandwidth usage compared to traditional solutions. 2. This application can improve high-concurrency processing capabilities. The dynamic resource allocation algorithm performs weighted scoring on CPU utilization, remaining memory and network latency based on the real-time load status of each edge node, intelligently schedules missing person search tasks, realizes load balancing when searching for missing persons in multiple regions concurrently, and improves the overall throughput of the system. 3. This application can achieve seamless tracking across floors. By deploying a directional antenna array to monitor changes in RFID signal strength, when the target enters the floor transition zone, it triggers data synchronization of adjacent edge nodes, establishes a continuous location trajectory, and reduces positioning blind spots when switching floors. 4. This application can improve positioning accuracy in complex environments. The multimodal fusion positioning algorithm comprehensively utilizes video face, Bluetooth RSSI and RFID signals, and fuses face features, Bluetooth beacon signals and radio frequency identification signals, so as to output stable and accurate coordinate positions even under complex conditions. 5. This application can improve resource utilization efficiency. The adaptive resolution adjustment mechanism dynamically selects the resolution and frame rate level based on feedback from multiple dimensions such as face confidence, RFID signal strength of adjacent floors, network quality, and CPU load, thereby avoiding resource waste and saving computing resources. Attached Figure Description

[0025] Figure 1 This is a flowchart of a real-time person-finding and collaborative monitoring method for intelligent buildings, as described in one embodiment of this application.

[0026] Figure 2 This is a flowchart of a sub-step of step S1 in one embodiment of this application.

[0027] Figure 3 This is a flowchart of the steps added after step S3 in one embodiment of this application.

[0028] Figure 4 This is a sub-step of step S3 in one embodiment of this application. Figure 1 .

[0029] Figure 5 This is a flowchart of a sub-step of step S4 in one embodiment of this application.

[0030] Figure 6 This is a sub-step of step S3 in one embodiment of this application. Figure 2 .

[0031] Figure 7 This is a flowchart of the steps added after step S12 in one embodiment of this application.

[0032] Figure 8 This is a schematic diagram of the structure of a real-time collaborative monitoring system for finding people in an intelligent building, which is one embodiment of this application.

[0033] Figure 9 This is a schematic block diagram of an electronic device in one embodiment of this application.

[0034] Attached reference numerals: 1. Request allocation module; 2. Signal acquisition module; 3. Position generation module; 4. Video adjustment module; 5. Trajectory generation module. Detailed Implementation

[0035] The following is in conjunction with the appendix Figure 1-9 This application will be described in further detail.

[0036] It should be noted that all actions involving the acquisition of data or information in this application are carried out in accordance with the relevant data protection laws and policies of the country where the application is located, and with the authorization of the relevant users.

[0037] refer to Figure 1 A real-time person-finding and collaborative monitoring method for intelligent buildings, specifically including: S1. Receive a missing person request for the target, and allocate the missing person request to at least one edge computing node as a tracking node based on the current load status of each edge computing node.

[0038] Specifically, the system first receives a person search request from the user interface or an automatically triggered request, which contains key information for identifying the target. The system then acquires real-time operational load information for all edge computing nodes within the building, including comprehensive indicators such as the number of tasks processed by each node, response speed, and resource consumption.

[0039] Based on these load states, the system needs to determine which nodes are currently relatively idle or capable of handling new tasks, and then assign the missing person request to one or more of these nodes as the tracking nodes. When there are multiple missing person requests, the system can distribute different requests to different nodes, allowing each node to share different tracking tasks. Ultimately, this achieves reasonable distribution of missing person tasks, avoids single nodes being overloaded and affecting processing speed, and improves the overall system response capability in multi-target concurrent scenarios.

[0040] S2. Control the assigned edge computing nodes to acquire video streams, Bluetooth beacon signals, and RFID signals within the coverage area.

[0041] Specifically, after receiving control commands, the selected edge computing node initiates data acquisition in its designated area. The node connects to cameras within its coverage area and begins receiving real-time video streams. Simultaneously, the node activates its built-in Bluetooth receiver module, scanning the surrounding environment for signals emitted by Bluetooth beacons and recording the identifier and signal strength of each beacon. The node also controls RFID readers or antennas to capture response signals from RFID tags within its coverage area. All three types of signals are cached locally at the edge node and are not immediately uploaded to the central server.

[0042] S3. Obtain the facial features extracted from the video stream locally by the edge computing node, and perform multimodal fusion of the facial features with Bluetooth beacon signals and RFID signals to generate the real-time location of the target.

[0043] Specifically, the edge computing node runs face detection and feature extraction algorithms on the video stream locally, extracting the feature vector of the target face from each frame and estimating the approximate orientation of the target by combining the camera installation parameters. Furthermore, the node organizes multiple signal strength values ​​collected by Bluetooth beacons into feature sequences, and also organizes the signal strength detected by the RFID reader into feature sequences. Subsequently, the node performs time alignment and fusion processing on these features from different sensors, and uses a built-in computing model to comprehensively determine the target's specific coordinates within the current floor. The entire fusion and localization process is completed within the edge node itself.

[0044] S4. Adjust the acquisition resolution and frame rate of the video stream based on the confidence level of the facial feature extraction and the RFID signal strength of adjacent floors.

[0045] Specifically, during the tracking process, edge computing nodes continuously monitor the matching confidence score output by the face recognition model. This value reflects the similarity between the currently detected face and the target template. Simultaneously, nodes acquire the RFID signal strength from adjacent floors using directional antennas or signal receivers deployed in the floor transition area. When the confidence score is low, the node sends instructions to the camera to increase the resolution and frame rate of the video stream to obtain clearer facial details. When the RFID signal strength from adjacent floors significantly increases, the node similarly increases the resolution and frame rate to prepare for possible floor switching. Conversely, if the confidence score is high and the signal from adjacent floors is weak, the node lowers the resolution and frame rate to reduce resource consumption.

[0046] In summary, by dynamically adapting video acquisition parameters, the system automatically improves image quality when fine recognition is required and actively reduces resource consumption when tracking is stable, thereby balancing recognition accuracy and system load.

[0047] S5. Receive the real-time location reported by the edge computing node and generate the target's movement trajectory based on the real-time location.

[0048] Specifically, the central server continuously receives encrypted real-time location data uploaded by each edge computing node. Each data point includes the target's planar coordinates, floor number, and timestamp. The server sorts these discrete location points chronologically and uses a smoothing algorithm to remove abnormal jumps caused by signal fluctuations. When the target moves between the coverage areas of different nodes, the server stitches the location data from multiple nodes into a continuous path. Finally, the server presents the processed trajectory on a management interface or electronic map, supporting real-time viewing and historical playback. This integrates scattered location points into a complete movement path, providing clear location history and current orientation for target search, facilitating rapid location and prediction of movement direction.

[0049] refer to Figure 2Furthermore, in one embodiment, step S1 is refined into the following sub-steps: S10. Obtain the CPU utilization, remaining memory, and network latency of each edge computing node.

[0050] Specifically, the system periodically or in real-time sends status query commands to all deployed edge computing nodes within the building. Each node reports its current CPU load percentage, unused memory capacity, and network round-trip time with the central server. The system records this raw data in a status table and stores it according to node identifiers. In this embodiment, the data acquisition operation adopts a non-blocking method to avoid blocking the overall scheduling process due to waiting for individual node responses. For network latency metrics, the system continuously collects multiple data points and takes the average value to reduce errors caused by instantaneous fluctuations.

[0051] S11. Assign weights to CPU utilization, remaining memory, and network latency, and then perform a weighted scoring.

[0052] Specifically, the system assigns different weight coefficients to three indicators—CPU utilization, remaining memory, and network latency—based on the building's actual operating environment and task characteristics. CPU utilization reflects the node's computational load and has a higher weight; remaining memory reflects the node's data caching capacity and has a moderate weight; and network latency reflects the communication efficiency between the node and the center and has a relatively lower weight. The system normalizes the values ​​of each indicator for each node, multiplies them by their corresponding weights, and then sums the three products to obtain the node's overall score. A higher score indicates that the node is more suitable to undertake new missing person search tasks.

[0053] S12. Assign the missing person request to the edge computing node with the highest weighted score.

[0054] Specifically, the system identifies the node with the highest weighted score among all edge computing nodes, sends the complete information of the missing person request, including target feature data, task priority, and expected response time, to this node, and designates it as the main processing node for this tracking task. If multiple nodes have the same highest score, the system can select the one with the lowest network latency or randomly assign tasks using a round-robin method. After allocation, the system updates the task queue status of the node to prevent the same node from being assigned too many tasks repeatedly. Ultimately, this ensures that the missing person task lands on the device with the most abundant resources, minimizing task waiting time and avoiding processing delays caused by single-point overload, thus improving the overall system processing efficiency in multi-target concurrent scenarios.

[0055] In addition, refer to Figure 3Furthermore, in one embodiment, after step S3, steps S30, S31, S32, and S33 are added: S30. Obtain the RFID signal strength of the floor corresponding to the real-time location of the target and the adjacent floors.

[0056] Specifically, after the system generates the target's real-time location at the edge computing node, it immediately and continuously collects RFID tag signal strength data through directional antenna arrays or fixed RFID readers deployed in key areas of the current floor and adjacent floors, such as elevator lobbies, stairwells, and corridor connections. This data includes signal strength values ​​of the same tag detected by multiple readers on the current floor, as well as RFID signal strength information shared by adjacent floors through inter-node communication. The system then records and timestamps the collected signal strengths according to the source floor, forming a floor-specific signal strength sequence. Ultimately, this establishes an RFID signal monitoring channel between the target's floor and adjacent floors.

[0057] S31. When the RFID signal strength of an adjacent floor exceeds a preset strength threshold, determine whether the target has entered the floor transition zone.

[0058] Specifically, the system pre-sets a strength threshold for each floor transition zone. This threshold is calibrated based on the location of the RFID reader, its transmission power, and a typical path loss model in the actual deployment environment. Edge computing nodes compare the real-time acquired RFID signal strength from adjacent floors with this threshold. If the signal strength consistently exceeds the threshold for a certain time window, the system determines that the target has entered or is about to enter the floor transition zone such as an elevator, staircase, or escalator. The process of determining whether a target has entered the floor transition zone also considers the upward trend of the signal strength to avoid misjudgments due to short-term fluctuations.

[0059] S32. If the target enters the floor transition zone, data synchronization between the current edge computing node and the edge computing nodes of the adjacent floors is triggered. The data synchronization includes the target's facial features and historical trajectory.

[0060] Specifically, once a target is determined to have entered the floor transition zone, the currently tracking edge computing node immediately initiates a data synchronization process with the edge computing nodes on adjacent floors. This synchronization process is completed through local area network communication between nodes, and the transmitted data includes the target's facial feature vector, a sequence of all historical location coordinates from the start of tracking to the current moment, the timestamp of the target's last appearance, and the identification information of the current tracking task.

[0061] To ensure data transmission security, the synchronization process employs encrypted transmission. After completing data transmission, the current node continues to track the target until it confirms that a neighboring node has successfully taken over.

[0062] The above achieves seamless transfer of tracking tasks between two edge computing nodes, and enables adjacent nodes to obtain the target's identity features and movement history in advance, avoiding tracking interruptions or repeated identification due to missing information.

[0063] S33. Use the edge computing nodes of adjacent floors as tracking nodes to receive the real-time location of the target reported by the adjacent floors, so as to form a continuous cross-floor trajectory.

[0064] Specifically, after data synchronization is complete, the system officially sets up the edge computing nodes on adjacent floors as new tracking nodes, while the original current nodes are relegated to standby or auxiliary roles. Adjacent nodes begin to continuously locate the target within their assigned floor area based on the received facial features and trajectory information, and report the real-time location to the central server. When stitching the trajectory, the central server directly connects the last location reported by the original node with the first location reported by the new node in chronological order, forming an uninterrupted movement path. If the target briefly loses signal in the transition zone, the system performs dead reckoning based on the synchronized historical trajectory to fill in the gaps, forming a complete cross-floor trajectory, thus achieving continuous tracking of the target across different floors.

[0065] In addition, refer to Figure 4 Furthermore, in one embodiment, step S3 is refined into the following sub-steps: S34. Extract facial feature vectors from the detected face regions in the video stream and obtain relative coordinates by combining them with camera calibration parameters.

[0066] Specifically, after receiving the video stream, the edge computing node first runs a lightweight face detection algorithm to locate the face region in the image, and then extracts a high-dimensional face feature vector from this region. This vector uniquely represents the facial identity information of the target. Simultaneously, the node reads pre-calibrated parameters such as the camera's installation height, pitch angle, and horizontal viewing angle. Combined with the pixel position of the face bounding box in the image, it calculates the direction angle and estimated distance of the target relative to the camera's horizontal projection point through geometric projection relationships, forming relative coordinates. These relative coordinates do not depend on an absolute position reference frame and only reflect the spatial relationship between the target and the camera.

[0067] S35. Sort the received signal strength values ​​of multiple Bluetooth beacons according to the beacon identifier and construct a first signal vector of fixed length.

[0068] Specifically, edge computing nodes scan for surrounding Bluetooth beacons within their coverage area. Each beacon has a unique Media Access Control (MACC) address. The nodes sort the received signal strength values ​​of each beacon according to the lexicographical order of the beacon identifiers or a preset order. For unscanned beacon locations, a preset minimum value is used to fill in the gaps, thereby constructing a first signal vector of fixed length, consistent with the number of beacons deployed within the building. Each element in the first signal vector represents the observed signal strength value of the corresponding beacon at the current moment.

[0069] The above process converts unstructured Bluetooth signal scanning results into fixed-dimensional structured vectors, facilitating unified processing with video features and RFID features, while avoiding feature dimension mismatch issues caused by changes in the number of beacons.

[0070] S36. Sort the signal strength values ​​of the multiple RFID readers according to the reader identifier and construct a second signal vector of fixed length.

[0071] Specifically, edge computing nodes connect to multiple RFID readers fixedly installed at key locations in the building, each reader corresponding to a unique identifier. The nodes simultaneously read the signal strength values ​​of the same target RFID tag detected by all readers, arranging them according to a fixed order of reader identifiers. For readers that fail to detect the tag, a preset minimum value is used to fill the gap, forming a second signal vector with a length equal to the number of readers. This second signal vector reflects the signal attenuation distribution of the target tag across readers at different spatial locations.

[0072] The above method organizes spatially distributed RFID observation data into a feature vector of a unified format, providing signal strength distribution information of the target relative to multiple card readers for fusion positioning, which is beneficial for solving the position using the spatial distribution of the signal.

[0073] S37. Normalize the first signal vector and the second signal vector.

[0074] Specifically, due to differences in transmit power, environmental attenuation, and receive sensitivity, Bluetooth beacons and RFID readers exhibit significant differences in the original dimensions and value ranges of their signal strength values. Edge computing nodes perform normalization operations on the first and second signal vectors, typically using maximum-minimum normalization or mean-variance normalization, mapping the element values ​​of each vector to a uniform numerical range. Normalization parameters, such as maximum, minimum, mean, and variance, can be calibrated based on pre-collected offline data to eliminate dimensional and amplitude differences between different signal sources. This ensures that during subsequent feature stitching, the data from each modality are within a similar numerical range, avoiding difficulties in training the fusion network or positioning errors caused by different numerical scales.

[0075] S38. The normalized first signal vector, second signal vector and face feature vector are concatenated and input into the pre-trained deep regression network to output the plane coordinates and floor number of the target.

[0076] Specifically, the edge computing node concatenates the normalized Bluetooth signal vector, the RFID signal vector, and the facial feature vector extracted from the video sequentially to form a fused feature vector. This fused vector is then fed into a pre-trained lightweight deep regression network containing multiple fully connected layers and non-linear activation layers, resulting in a small number of network parameters suitable for edge deployment.

[0077] The network then performs forward computation and outputs two values: the x and y coordinates of the target in the current floor's planar coordinate system, and the current floor number. This network was trained using multimodal data collected within the building and real-world location labels before deployment.

[0078] In summary, by using deep neural networks to perform end-to-end nonlinear fusion and regression calculation of multi-source heterogeneous features, the system fully utilizes the identity discrimination capability of visual features and the spatial location correlation capability of wireless signals, and can still output stable and accurate three-dimensional location information even in complex environments such as changes in light, occlusion, and signal fluctuations.

[0079] In addition, refer to Figure 5 Furthermore, in one embodiment, step S4 is refined into the following sub-steps: S40: Real-time acquisition of network quality parameters between edge computing nodes and the central server, as well as the current CPU utilization rate of the edge computing nodes.

[0080] Specifically, during operation, edge computing nodes continuously monitor their network connection status with the central server. By sending heartbeat packets or querying underlying network interface statistics, they obtain key parameters reflecting network quality, such as current uplink bandwidth availability, round-trip latency, and packet loss rate. Simultaneously, the nodes read the current CPU load percentage in real time through the operating system interface, including the proportion of user mode, system mode, and idle time, and calculate short-term average utilization to smooth out instantaneous fluctuations. These parameters are collected at fixed time intervals and cached in the node's local memory for the resolution adjustment decision module to access at any time. The collection process itself consumes very little computational and network overhead and does not affect the normal execution of the main tracing task.

[0081] S41. When the network quality parameters are lower than the preset network threshold, or the CPU utilization rate is higher than the preset load threshold, reduce the acquisition resolution and frame rate.

[0082] Specifically, the system pre-sets lower thresholds for network quality parameters such as available bandwidth and maximum allowable latency, and sets upper thresholds for CPU utilization. Edge computing nodes compare the real-time collected network quality parameters with the preset network thresholds. If the available bandwidth is lower than the threshold or the latency exceeds the threshold, it indicates network link congestion. At the same time, it compares the CPU utilization with the preset load threshold. If the utilization is higher than the threshold, it indicates that the node's computing resources are strained.

[0083] When any condition is met, the node sends an adjustment command to the camera it is connected to, switching the video stream's capture resolution from the current level to the next lower level and correspondingly reducing the frame rate by one level. The adjustment process employs a gradual downgrading strategy, reducing the resolution by only one level at a time to avoid excessively drastic changes in image quality. If multiple conditions are met simultaneously, a more aggressive downgrading is executed.

[0084] In the above way, when there is network congestion or excessive node load, the video acquisition parameters are proactively reduced, thereby reducing the amount of data that needs to be processed and transmitted, freeing up valuable bandwidth and computing resources, preventing the system from crashing due to resource exhaustion, and ensuring the basic real-time performance of the tracking task.

[0085] S42. When the confidence level of facial feature extraction is lower than the preset confidence threshold, or the RFID signal strength of adjacent floors is higher than the preset strength threshold, increase the acquisition resolution and frame rate.

[0086] Specifically, the system sets a lower threshold for the confidence level of face recognition. When the confidence level of face matching in the current frame calculated by the edge nodes is lower than this threshold, it indicates that the existing image quality is insufficient to support reliable identity verification. At the same time, an upper threshold is set for the RFID signal strength of adjacent floors. When the signal strength of adjacent floors obtained through directional antennas or inter-node sharing is higher than this threshold, it indicates that the target is approaching the floor boundary and more detailed visual information is needed to prepare for cross-floor tracking.

[0087] When any of the conditions is met, the node sends a command to the camera to increase the resolution and frame rate, switching the acquisition parameters to the next higher level. The increase process also uses a gradual increase strategy, increasing only one level at a time until the highest level supported by the camera is reached. If both conditions are met simultaneously, a faster increase is performed.

[0088] In the above-mentioned scenarios, when the recognition accuracy is insufficient or the target is about to cross floors, the video acquisition parameters are automatically improved to enhance the ability to recognize facial details, providing high-quality image input for accurate matching and cross-floor handover, thereby ensuring the success rate of the search mission.

[0089] In addition, refer to Figure 6 Furthermore, in one embodiment, step S3 is refined into the following sub-steps: S39. Monitor the illumination of the video stream and the ratio of detectable face area.

[0090] Specifically, after receiving the video stream, the edge computing node extracts the overall brightness value of the image frame by frame or at intervals. It calculates the average brightness by statistically analyzing the pixel grayscale distribution or brightness histogram, serving as a quantitative indicator of lighting conditions. Simultaneously, the node performs area statistics on all face regions output by the face detection algorithm. It divides the pixel area of ​​the largest face region by the total pixel area of ​​the video frame to obtain the face detection area ratio. The face detection area ratio reflects the relative size of the target face in the image; a smaller ratio indicates that the target is farther from the camera or more severely obscured. The monitoring process continues at a fixed frequency, and the results are updated in real-time to the node's status register.

[0091] S310. When the illumination conditions are lower than the preset brightness threshold, or the detectable area ratio of the face is lower than the preset area ratio, reduce the weight of the face features in multimodal fusion, and correspondingly increase the weight of the Bluetooth beacon signal and the radio frequency identification signal.

[0092] Specifically, the system pre-calibrates the minimum acceptable threshold for illumination and the minimum effective threshold for the detectable face area ratio through experiments. Edge computing nodes compare the real-time monitored illumination values ​​with the brightness thresholds, and simultaneously compare the detectable face area ratio with the area ratio threshold. If the illumination is below the threshold, it indicates that the environment is too dark, causing blurred facial features; if the area ratio is below the threshold, it indicates that the target is too far away or partially occluded, resulting in loss of facial details.

[0093] When any condition is met, the node dynamically reduces the weighting coefficient of the face feature vector during the multimodal fusion localization calculation, while increasing the weighting coefficients of the Bluetooth signal vector and the RFID signal vector, making the final localization result more dependent on wireless signals than on visual features. The weight adjustment adopts a smooth and gradual approach to avoid abrupt changes in the localization result.

[0094] In summary, when the quality of visual input deteriorates, its impact on the positioning results is automatically reduced, and instead, Bluetooth and radio frequency signals, which are less affected by light and distance, are relied upon. This ensures that the system can still output stable position information in dark, long-distance, or partially occluded scenarios, avoiding positioning failures caused by visual impairment.

[0095] S311. When the signal-to-noise ratio of the Bluetooth beacon signal or RFID signal is lower than the preset signal-to-noise ratio threshold, increase the weight of facial features.

[0096] Specifically, while collecting Bluetooth beacon and RFID signals, the edge computing nodes estimate the signal-to-noise ratio (SNR) of each signal, i.e., the ratio of received signal strength to background noise power. The system sets lower SNR thresholds for both Bluetooth beacon and RFID signals. When the average SNR of either the Bluetooth beacon or the RFID signal is below its threshold, it indicates severe interference with the wireless signal or that the target is far from the beacon deployment area, thus reducing the reliability of wireless positioning. In multimodal fusion positioning, the nodes correspondingly increase the weight coefficient of facial feature vectors while appropriately reducing the weight of corresponding degraded signals, making the positioning results more reliant on visual information. If both wireless signals degrade simultaneously, the weight of facial features is further increased until positioning becomes entirely visual.

[0097] The above-mentioned features automatically enhance the contribution of visual features when wireless signals are interfered with or targets enter signal blind spots. The system uses facial recognition for identity identification and orientation estimation to compensate for the deficiencies of wireless positioning. It achieves bidirectional adaptive adjustment of multimodal fusion weights, ensuring that the system can switch to other relatively reliable modes regardless of the signal source degradation, thus maintaining the continuity and accuracy of positioning.

[0098] In addition, refer to Figure 7 Furthermore, in one embodiment, after step S12, steps S120 and S121 are added: S120: Send a stop tracking command to all assigned edge computing nodes.

[0099] Specifically, when the central server receives a user-initiated command to end the search, or when its built-in monitoring logic detects that no edge nodes have reported new location updates for the target within a preset time period, it determines that the current search task has been completed or has failed. Based on the list of assigned nodes stored in the task record, the server simultaneously sends a stop tracking command to each edge computing node that participated in tracking the target. This stop tracking command includes a task identifier to ensure that nodes can accurately identify which tracking task needs to be terminated, avoiding accidental termination of other ongoing tasks. A reliable transmission mechanism is used during the sending process, requiring nodes to reply with confirmation of receipt; nodes that do not confirm in time will be resent. Once the server receives confirmation from all nodes or reaches the timeout retries limit, it removes the task from the active task list.

[0100] S121, instructs the edge computing node to stop extracting facial features and performing multimodal fusion on the video stream, and restores the acquisition resolution and frame rate of the video stream to the initial low-power default values.

[0101] Specifically, upon receiving a stop tracking command, the edge computing node immediately interrupts the face feature extraction thread and multimodal fusion localization thread for that task, releasing the CPU time and memory space occupied by these threads. Simultaneously, the node sends a parameter reset command to the managed cameras, switching the video stream's acquisition resolution from the currently possible high level back to the system's preset initial low resolution level, and restoring the frame rate to its low-power default value, such as the standby scanning level before the task started. The node also clears the cache of face feature vectors, historical signal strength data, and temporarily generated trajectory points related to that task to prevent residual data from interfering with the normal operation of subsequent new tasks. After completing the above operations, the node reports back to the central server that the status has been reset and re-enters idle standby mode.

[0102] The above achieves timely recycling and state reset of system resources after task completion, avoids invalid computation, and reduces the average power consumption during long-term operation.

[0103] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0104] This application also provides a real-time collaborative monitoring system for finding people in intelligent buildings, which corresponds one-to-one with the real-time collaborative monitoring method for finding people in intelligent buildings in the embodiments.

[0105] refer to Figure 8 A real-time person-finding collaborative monitoring system for intelligent buildings includes: a request allocation module 1, a signal acquisition module 2, a location generation module 3, a video adjustment module 4, and a trajectory generation module 5. Detailed descriptions of each functional module are as follows: Request allocation module 1: Used to receive a missing person request for a target, and allocate the missing person request to at least one edge computing node as a tracking node according to the current load status of each edge computing node; Signal acquisition module 2: Used to control the assigned edge computing nodes to acquire video streams, Bluetooth beacon signals and RFID signals within the coverage area.

[0106] Location generation module 3: Used to obtain facial features extracted from the video stream locally by the edge computing node, and perform multimodal fusion of facial features with Bluetooth beacon signals and RFID signals to generate the real-time location of the target.

[0107] Video adjustment module 4: used to adjust the acquisition resolution and frame rate of the video stream based on the confidence level of facial feature extraction and the RFID signal strength of adjacent floors.

[0108] Trajectory generation module 5: Used to receive the real-time location reported by the edge computing node and generate the target's movement trajectory based on the real-time location.

[0109] The system comprises several modules: Request Allocation Module 1 allocates person-finding requests based on the load status of each edge node, preventing overload of a single node and ensuring reasonable task distribution; Signal Acquisition Module 2 controls nodes to collect video, Bluetooth, and radio frequency signals, providing multi-source data for subsequent positioning; Location Generation Module 3 fuses facial features and signals locally at the node, reducing transmission latency and outputting real-time coordinates; Video Adjustment Module 4 dynamically adjusts resolution and frame rate based on confidence level and radio frequency signals from adjacent floors, balancing resources and recognition accuracy; and Trajectory Generation Module 5 receives the location and generates a movement trajectory, forming complete path information. Through the collaborative work of these modules, system response latency and network bandwidth usage are reduced, improving positioning accuracy in complex environments.

[0110] Specific limitations regarding the real-time person-finding collaborative monitoring system for intelligent buildings can be found in the context of the limitations on the real-time person-finding collaborative monitoring method for intelligent buildings, and will not be repeated here. Each module in the aforementioned real-time person-finding collaborative monitoring system for intelligent buildings can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in an electronic device, or stored in the memory of the electronic device in software form, so that the processor can call and execute the corresponding operations of each module. In one embodiment, an electronic device is provided, which is a user terminal. (Reference) Figure 9 The electronic device includes a processor, memory, network interface, and database connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The database stores detection data tables. The network interface allows communication with external terminals via a network connection. When the computer program is executed by the processor, it implements a real-time collaborative monitoring method for finding people in intelligent buildings.

[0111] In one embodiment, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to perform the following steps: S1. Receive a missing person request for the target, and allocate the missing person request to at least one edge computing node as a tracking node based on the current load status of each edge computing node.

[0112] S2. Control the assigned edge computing nodes to acquire video streams, Bluetooth beacon signals, and RFID signals within the coverage area.

[0113] S3. Obtain the facial features extracted from the video stream locally by the edge computing node, and perform multimodal fusion of the facial features with Bluetooth beacon signals and RFID signals to generate the real-time location of the target.

[0114] S4. Adjust the acquisition resolution and frame rate of the video stream based on the confidence level of the facial feature extraction and the RFID signal strength of adjacent floors.

[0115] S5. Receive the real-time location reported by the edge computing node and generate the target's movement trajectory based on the real-time location.

[0116] In one embodiment, the refined sub-steps of step S1 include: S10. Obtain the CPU utilization, remaining memory, and network latency of each edge computing node.

[0117] S11. Assign weights to CPU utilization, remaining memory, and network latency, and then perform a weighted scoring.

[0118] S12. Assign the missing person request to the edge computing node with the highest weighted score.

[0119] In one embodiment, the additional step after step S3 includes: S30. Obtain the RFID signal strength of the floor corresponding to the real-time location of the target and the adjacent floors.

[0120] S31. When the RFID signal strength of an adjacent floor exceeds a preset strength threshold, determine whether the target has entered the floor transition zone.

[0121] S32. If the target enters the floor transition zone, data synchronization between the current edge computing node and the edge computing nodes of the adjacent floors is triggered. The data synchronization includes the target's facial features and historical trajectory.

[0122] S33. Use the edge computing nodes of adjacent floors as tracking nodes to receive the real-time location of the target reported by the adjacent floors, so as to form a continuous cross-floor trajectory.

[0123] In one embodiment, the refined sub-steps of step S3 include: S34. Extract facial feature vectors from the detected face regions in the video stream and obtain relative coordinates by combining them with camera calibration parameters.

[0124] S35. Sort the received signal strength values ​​of multiple Bluetooth beacons according to the beacon identifier and construct a first signal vector of fixed length.

[0125] S36. Sort the signal strength values ​​of the multiple RFID readers according to the reader identifier and construct a second signal vector of fixed length.

[0126] S37. Normalize the first signal vector and the second signal vector.

[0127] S38. The normalized first signal vector, second signal vector and face feature vector are concatenated and input into the pre-trained deep regression network to output the plane coordinates and floor number of the target.

[0128] In one embodiment, the refined sub-steps of step S4 include: S40: Real-time acquisition of network quality parameters between edge computing nodes and the central server, as well as the current CPU utilization rate of the edge computing nodes.

[0129] S41. When the network quality parameters are lower than the preset network threshold, or the CPU utilization rate is higher than the preset load threshold, reduce the acquisition resolution and frame rate.

[0130] S42. When the confidence level of facial feature extraction is lower than the preset confidence threshold, or the RFID signal strength of adjacent floors is higher than the preset strength threshold, increase the acquisition resolution and frame rate.

[0131] In one embodiment, the sub-steps of step S3 refinement include: S39. Monitor the illumination of the video stream and the ratio of detectable face area.

[0132] S310. When the illumination conditions are lower than the preset brightness threshold, or the detectable area ratio of the face is lower than the preset area ratio, reduce the weight of the face features in multimodal fusion, and correspondingly increase the weight of the Bluetooth beacon signal and the radio frequency identification signal.

[0133] S311. When the signal-to-noise ratio of the Bluetooth beacon signal or RFID signal is lower than the preset signal-to-noise ratio threshold, increase the weight of facial features.

[0134] In one embodiment, the additional steps following step S12 include: S120: Send a stop tracking command to all assigned edge computing nodes.

[0135] S121, instructs the edge computing node to stop extracting facial features and performing multimodal fusion on the video stream, and restores the acquisition resolution and frame rate of the video stream to the initial low-power default values.

[0136] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0137] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

Claims

1. A real-time person-finding and collaborative monitoring method for intelligent buildings, characterized in that, include: Receive a missing person request for a target, and allocate the missing person request to at least one edge computing node as a tracking node based on the current load status of each edge computing node; Control the assigned edge computing nodes to acquire video streams, Bluetooth beacon signals, and RFID signals within the coverage area; The facial features extracted locally from the video stream by the edge computing node are obtained, and the facial features are fused with the Bluetooth beacon signal and the radio frequency identification signal in a multimodal manner to generate the real-time location of the target; The acquisition resolution and frame rate of the video stream are adjusted based on the confidence level of the extracted facial features and the radio frequency identification signal strength of adjacent floors. The system receives the real-time location reported by the edge computing node and generates the target's movement trajectory based on the real-time location.

2. The method according to claim 1, characterized in that, The step of receiving a missing person request for a target and allocating the missing person request to at least one edge computing node according to the current load status of each edge computing node includes: Obtain the CPU utilization, remaining memory, and network latency of each edge computing node; The CPU utilization, remaining memory, and network latency are each assigned a weight and then weighted and scored. The missing person request is assigned to the edge computing node with the highest weighted score.

3. The method according to claim 1, characterized in that, After the steps of obtaining the facial features extracted locally from the video stream by the edge computing node, and performing multimodal fusion of the facial features with the Bluetooth beacon signal and the RFID signal to generate the real-time location of the target, the method further includes: Obtain the radio frequency identification signal strength of the floor corresponding to the real-time location of the target and the adjacent floors; When the radio frequency identification signal strength of the adjacent floor exceeds a preset strength threshold, it is determined whether the target has entered the floor transition zone; If the target enters the floor transition zone, data synchronization between the current edge computing node and the edge computing nodes of the adjacent floors is triggered, wherein the data synchronization includes the target's facial features and historical trajectory; The edge computing nodes of the adjacent floors are used as tracking nodes to receive the real-time location of the target reported by the adjacent floors, so as to form a continuous cross-floor trajectory.

4. The method according to claim 1, characterized in that, The step of obtaining the facial features extracted locally from the video stream by the edge computing node, and performing multimodal fusion of the facial features with the Bluetooth beacon signal and the RFID signal to generate the real-time location of the target includes: The facial feature vectors are extracted from the detected facial regions in the video stream, and the relative coordinates are obtained by combining the camera calibration parameters; The received signal strength values ​​of multiple Bluetooth beacons are sorted according to the beacon identifier, and a first signal vector of fixed length is constructed. The signal strength values ​​of multiple RFID readers are sorted according to the reader identifier, and a second signal vector of fixed length is constructed. The first signal vector and the second signal vector are normalized. The normalized first signal vector, the second signal vector, and the face feature vector are concatenated and input into a pre-trained deep regression network to output the planar coordinates and floor number of the target.

5. The method according to claim 1, characterized in that, The step of adjusting the acquisition resolution and frame rate of the video stream based on the confidence level extracted from the facial features and the radio frequency identification signal strength of adjacent floors includes: The network quality parameters between the edge computing node and the central server, as well as the CPU utilization rate of the current edge computing node, are obtained in real time. When the network quality parameters are lower than a preset network threshold, or the CPU utilization rate is higher than a preset load threshold, the acquisition resolution and frame rate are reduced. When the confidence level of the extracted facial features is lower than a preset confidence threshold, or when the RFID signal strength of the adjacent floors is higher than a preset strength threshold, the acquisition resolution and frame rate are increased.

6. The method according to claim 1, characterized in that, The step of obtaining the facial features extracted locally from the video stream by the edge computing node, and performing multimodal fusion of the facial features with the Bluetooth beacon signal and the RFID signal to generate the real-time location of the target includes: Monitor the illumination intensity and the ratio of detectable face area in the video stream; When the lighting conditions are lower than a preset brightness threshold, or the detectable area ratio of the face is lower than a preset area ratio, the weight of the face features in multimodal fusion is reduced, and the weight of the Bluetooth beacon signal and the radio frequency identification signal is increased accordingly. When the signal-to-noise ratio of the Bluetooth beacon signal or the RFID signal is lower than a preset signal-to-noise ratio threshold, the weight of the facial features is increased.

7. The method according to claim 2, characterized in that, After the step of assigning the missing person request to the edge computing node with the highest weighted score, the method further includes: Send a stop tracking command to all assigned edge computing nodes; The edge computing node is instructed to stop extracting facial features and performing multimodal fusion on the video stream, and to restore the acquisition resolution and frame rate of the video stream to their initial low-power default values.

8. A real-time person-finding collaborative monitoring system for intelligent buildings, characterized in that, include: Request allocation module (1): used to receive a missing person request for a target, and allocate the missing person request to at least one edge computing node as a tracking node according to the current load status of each edge computing node; Signal acquisition module (2): used to control the assigned edge computing nodes to acquire video streams, Bluetooth beacon signals and RFID signals within the coverage area; Location generation module (3): used to obtain the facial features extracted by the edge computing node from the video stream locally, and to perform multimodal fusion of the facial features with the Bluetooth beacon signal and the radio frequency identification signal to generate the real-time location of the target; Video adjustment module (4): used to adjust the acquisition resolution and frame rate of the video stream based on the confidence level of the facial features extracted and the radio frequency identification signal strength of the adjacent floors; Trajectory generation module (5): used to receive the real-time location reported by the edge computing node and generate the movement trajectory of the target based on the real-time location.

9. An electronic device, characterized in that, It includes a memory and a processor, wherein the memory stores a computer program that can be loaded by the processor and executed as claimed in any one of claims 1 to 7 for real-time person search and collaborative monitoring of intelligent buildings.

10. A computer-readable storage medium, characterized in that, The system stores a computer program that can be loaded by a processor and executed as any one of the real-time person-finding and collaborative monitoring methods for intelligent buildings as claimed in claims 1 to 7.