Safety monitoring method, device and system based on video AI and WLAN equipment

By combining video AI with WLAN equipment, the spatial and temporal synchronization and multimodal fusion of video feature data and wireless signal data is achieved, solving the problem of identification delay and misjudgment of video surveillance system in occlusion scenarios, and improving the accuracy and real-timeness of abnormal behavior detection.

CN120279463AInactive Publication Date: 2025-07-08CHENGDU SKSPRUCE TECH +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510424221.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-07-08
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing video surveillance systems are in the problem of identification delay, missed judgment or misjudgment, especially in occlusion scenarios, and WLAN positioning lacks behavioral recognition capabilities.

Method used

Combining video AI and WLAN devices, the video feature data and wireless signal data timestamps are synchronized through a linear interpolation algorithm, and the video positioning coordinates and WLAN triangular positioning coordinates are matched using the Euro-type distance. The feature vectors are extracted and weighted fusion is performed using the CNN and RNN models to generate abnormal detection results and generate early warning information.

Benefits of technology

It improves the accuracy of abnormal behavior detection, reduces the false alarm and missed alarm rates, improves the intelligence level of the monitoring system, and ensures security and real-time response capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279463A_ABST
    Figure CN120279463A_ABST
Patent Text Reader

Abstract

The invention provides a safety monitoring method, device and system based on video AI and WLAN equipment, relates to the field of wireless local area network and video monitoring safety, and is used for solving the problems of recognition delay, missed judgment or misjudgment and the like existing in video monitoring. The wireless signal data of the target area are uploaded by the WLAN device; aligning timestamps of the video feature data and the wireless signal data by adopting a linear interpolation algorithm to obtain synchronous data; based on the synchronous data, judging whether the person is the same person or not by adopting an Euclidean distance between a video positioning coordinate and a WLAN triangulation positioning coordinate, and obtaining matched data; extracting a video feature vector based on CNN, extracting a wireless signal feature vector based on RNN, and performing weighted fusion on the video feature vector and the wireless signal feature vector by adopting an attention mechanism to obtain a fused feature vector; and obtaining a detection result according to the fused feature vector, thereby improving the accuracy of anomaly detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of monitoring technologies, and provides a security monitoring method, device, and system based on video AI and WLAN devices. Background Art

[0002] In some public places such as shopping malls, office buildings, and subway stations, modern security monitoring systems usually use video cameras to monitor the personnel flow, behaviors, and abnormal situations in public places in real time, which can help improve the security of public places and provide important evidence after an incident occurs.

[0003] In addition to traditional video monitoring, some more advanced systems also combine artificial intelligence (AI) and machine learning technologies to perform real-time analysis on video streams. For example, the system can identify specific behavior patterns, such as personnel gathering, abnormal staying, rapid movement, etc., or automatically alarm for behaviors entering restricted areas. However, video monitoring is limited by factors such as camera angles, lighting, and occlusion, which may cause problems such as recognition delay, missed judgment, or misjudgment. Summary of the Invention

[0004] This application provides a security monitoring method, device, and system based on video AI and WLAN devices, which are used to solve problems such as recognition delay, missed judgment, or misjudgment existing in video monitoring.

[0005] In a first aspect, this application provides a security monitoring method based on video AI and WLAN devices, which is applied to a central management platform of a security monitoring system. The security monitoring system further includes a camera and a WLAN device; the method includes:

[0006] Receiving video feature data of a target area uploaded by the camera, and wireless signal data of the target area uploaded by the WLAN device;

[0007] Using a linear interpolation algorithm to align the timestamps of the video feature data and the wireless signal data to obtain synchronized data;

[0008] Based on the synchronized data, using the Euclidean distance between the video positioning coordinates and the WLAN triangulation positioning coordinates to determine whether it is the same person, and obtaining matching data; the matching data includes the video feature data and the wireless signal data of the same person at each timestamp;

[0009] Input the video feature data of the same person at each timestamp into the CNN of the anomaly detection model to extract video feature vectors, and input the wireless signal data of the same person at each timestamp into the RNN of the anomaly detection model to extract wireless signal feature vectors. Then, use the attention mechanism to perform weighted fusion on the video feature vectors and the wireless signal feature vectors to obtain fused feature vectors;

[0010] Input the fused feature vectors into the fully connected layer of the anomaly detection model to obtain the detection result;

[0011] If the detection result indicates that there is an abnormal behavior of a person in the target area, generate a warning message.

[0012] Optionally, the method for determining whether they are the same person by using the Euclidean distance between the video positioning coordinates and the WLAN trilateration coordinates based on the synchronized data to obtain matching data includes:

[0013] Obtain the video positioning coordinates of multiple people in the target area at each timestamp according to the video feature data at each timestamp;

[0014] Obtain the WLAN trilateration coordinates of the multiple people at each timestamp according to the wireless signal data at each timestamp;

[0015] For the same timestamp, determine whether the video positioning coordinates and the WLAN trilateration coordinates belong to the same person according to whether the Euclidean distance between the video positioning coordinates and the WLAN trilateration coordinates is less than a preset distance;

[0016] Obtain the matching data according to the video feature data and the wireless signal data of the same person at each timestamp.

[0017] Optionally, the wireless signal data includes MAC address, signal strength, and WLAN device identifier. The method for obtaining the WLAN trilateration coordinates of the multiple people at each timestamp according to the wireless signal data at each timestamp includes:

[0018] Determine multiple WLAN device identifiers and multiple signal strengths corresponding to the same MAC address according to the wireless signal data at each timestamp;

[0019] Use the trilateration algorithm to determine the WLAN trilateration coordinates of the person carrying the target device among the multiple people at each timestamp according to the multiple signal strengths and the positions of the multiple WLAN devices corresponding to the multiple WLAN device identifiers; the target device is the wireless device identified by the same MAC address.

[0020] Optionally, after obtaining the detection result according to the fusion feature vector, the method further includes:

[0021] If the detection result indicates that there is no abnormal behavior of anyone in the target area, determine the scene type to which the target area belongs according to the video feature data;

[0022] Determine the WLAN network load of the target area according to the wireless signal data; the WLAN network load includes the network load of the 2.4GHz band;

[0023] If the scene type is a meeting room and the WLAN network load is higher than the first threshold, preferentially allocate 5GHz band resources and adjust the channel width to 40MHz;

[0024] If the scene type is a shopping mall and the network load of the 2.4GHz band is higher than the second threshold, switch from the 2.4GHz band to the 5GHz band and enable the load balancing algorithm.

[0025] Optionally, the central management platform includes a database, and the database prestores the face features of multiple business personnel; if the detection result indicates that there is someone with abnormal behavior in the target area, generate a warning message, including:

[0026] If the detection result indicates that there is someone with abnormal behavior in the target area, extract the target face features of the target person with abnormal behavior from the video feature data;

[0027] Match the target face features with the face features of the multiple business personnel;

[0028] If the match is unsuccessful, generate a warning message.

[0029] Optionally, the database also prestores the business area of each business personnel; after matching the target face features with the face features of the multiple business personnel, the method further includes:

[0030] If the match is successful, determine the target business area of the target person from the database;

[0031] Determine the location information of the target person according to the wireless signal data;

[0032] If the location information of the target person belongs to the target business area, allocate access rights to the target person.

[0033] Optionally, after determining the location information of the target person according to the wireless signal data, the method further includes:

[0034] If the location information of the target person does not belong to the target business area, a warning message is generated.

[0035] Optionally, after generating a warning message if the detection result indicates that someone in the target area has abnormal behavior, the method further includes:

[0036] According to the matching data and the abnormal behavior data, a dynamic situation map of the personnel activities in the target area is generated by using a heat map algorithm; the abnormal behavior data includes the timestamp, location coordinates, and personnel where the abnormal behavior occurs.

[0037] In a second aspect, the present application provides a security monitoring device based on video AI and WLAN devices, which is set in the central management platform of the security monitoring system. The security monitoring system further includes a camera and a WLAN device; the device includes:

[0038] A data acquisition module, configured to receive the video feature data of the target area uploaded by the camera and the wireless signal data of the target area uploaded by the WLAN device;

[0039] A data fusion module, configured to align the timestamps of the video feature data and the wireless signal data by using a linear interpolation algorithm to obtain synchronized data; based on the synchronized data, determine whether it is the same person by using the Euclidean distance between the video positioning coordinates and the WLAN triangulation positioning coordinates to obtain matching data; the matching data includes the video feature data and the wireless signal data of the same person at each timestamp; input the video feature data of the same person at each timestamp into the CNN of the abnormal detection model to extract video feature vectors, and input the wireless signal data of the same person at each timestamp into the RNN of the abnormal detection model to extract wireless signal feature vectors, and use an attention mechanism to perform weighted fusion on the video feature vectors and the wireless signal feature vectors to obtain fused feature vectors;

[0040] An abnormal detection module, configured to input the fused feature vectors into the fully connected layer of the abnormal detection model to obtain a detection result;

[0041] A warning module, configured to generate a warning message if the detection result indicates that someone in the target area has abnormal behavior.

[0042] In a third aspect, the present application provides a security monitoring system, which includes a camera, a WLAN device, and a central management platform; wherein:

[0043] The camera is configured to collect a video stream of the target area, extract video feature data based on the video stream, and upload the video feature data to the central management platform;

[0044] The WLAN device is used to collect wireless signal data of the target area and upload the wireless signal data to the central management platform;

[0045] The central management platform is used to implement the security monitoring method based on video AI and WLAN devices as described in the first aspect.

[0046] Compared with the prior art, the beneficial effects of this application are as follows:

[0047] This application provides a security monitoring method based on video AI and WLAN devices. This method is applied to the central management platform of a security monitoring system, and the security monitoring system further includes a camera and a WLAN device; the method includes: receiving video feature data of the target area uploaded by the camera and wireless signal data of the target area uploaded by the WLAN device; using a linear interpolation algorithm to align the timestamps of the video feature data and the wireless signal data to obtain synchronized data; based on the synchronized data, using the Euclidean distance between the video positioning coordinates and the WLAN triangulation positioning coordinates to determine whether it is the same person to obtain matching data; the matching data includes the video feature data and the wireless signal data of the same person at each timestamp; inputting the video feature data of the same person at each timestamp into the CNN of the anomaly detection model to extract video feature vectors, and inputting the wireless signal data of the same person at each timestamp into the RNN of the anomaly detection model to extract wireless signal feature vectors, and using an attention mechanism to perform weighted fusion on the video feature vectors and the wireless signal feature vectors to obtain fused feature vectors; inputting the fused feature vectors into the fully connected layer of the anomaly detection model to obtain a detection result; if the detection result indicates that there is an abnormal behavior of a person in the target area, then generate a warning message.

[0048] Through spatio-temporal synchronization and multi-modal data fusion of video feature data and wireless signal data, the fused data can provide more multi-dimensional information, reducing the monitoring blind spots that may be brought by a single data source. Whether the video capture angle is poor or there is WLAN signal interference, it can be supplemented by another data source. Using the fused data for abnormal behavior detection can improve the accuracy of the detection result, and generate a warning message immediately when an abnormal behavior occurs, ensuring that relevant personnel can take actions quickly, avoiding or reducing potential security threats, and improving the intelligence level of the monitoring system. Brief Description of the Drawings

[0049] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the related technologies, the drawings required for use in the embodiments or the related technical descriptions are briefly introduced below. Obviously, the drawings described below are only the embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.

[0050] Figure 1 A schematic diagram of the structure of a security monitoring system provided in an embodiment of the present application;

[0051] Figure 2 A flowchart of a security monitoring method based on video AI and WLAN devices provided in an embodiment of the present application;

[0052] Figure 3 A schematic diagram of the structure of a CNN+RNN dual-stream network provided in an embodiment of the present application;

[0053] Figure 4 A schematic diagram of the structure of a security monitoring device based on video AI and WLAN equipment provided in an embodiment of the present application. DETAILED DESCRIPTION

[0054] In order to make the purpose, technical scheme and advantages of the present application clearer, the technical scheme in the embodiment of the present application will be clearly and completely described below in conjunction with the drawings in the embodiment of the present application. Obviously, the described embodiment is only a part of the embodiment of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in the field without making creative work are within the scope of protection of the present application. In the absence of conflict, the embodiments in the present application and the features in the embodiments can be combined with each other arbitrarily. In addition, although the logical order is shown in the flow chart, in some cases, the steps shown or described can be performed in an order different from that here.

[0055] Existing video surveillance systems cannot accurately locate people in obstructed scenes, and WLAN positioning lacks behavior recognition capabilities. In view of this, the present application embodiment provides a security monitoring method based on video AI and WLAN devices, which can be applied to security monitoring systems. Please refer to Figure 1 , is a schematic diagram of the structure of a security monitoring system provided in an embodiment of the present application. The security monitoring system includes a camera, a WLAN device and a central management platform.

[0056] Among them, the camera and the WLAN device can be set in the target area, which can be public places such as shopping malls and office buildings. The central management platform can be implemented through a terminal or a server. The terminal can be, for example, a mobile terminal, a fixed terminal or a portable terminal, such as a mobile phone, a site, a unit, a device, a multimedia computer, a multimedia tablet, an Internet node, a communicator, a desktop computer, a notebook computer, a tablet computer or any combination thereof, including accessories and peripherals of these devices or any combination thereof. The server can be an independent physical server, or a server cluster or a distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms, but is not limited thereto.

[0057] The camera has a video AI function, can collect the video stream of the target area, extract video feature data based on the video stream, and upload the video feature data to the central management platform. The WLAN device (Wireless Local Area Network device) refers to a hardware device that can connect to a wireless local area network (WLAN, Wireless Local Area Network) through a wireless signal, and has WiFi probe and WiFi Sensing functions. It can collect wireless signal data of the target area and upload the wireless signal data to the central management platform. The central management platform can perform security monitoring on the target area according to the video feature data and the wireless signal data. The specific security monitoring method will be introduced below.

[0058] It should be noted that Figure 1 The described security monitoring system takes one camera and one WLAN device as an example. In fact, the number of cameras and WLAN devices in the security monitoring system is not limited.

[0059] Please refer to Figure 2 , which is a schematic flowchart of a security monitoring method based on video AI and WLAN devices provided by an embodiment of this application. Based on the security monitoring system shown below Figure 1 shown, a security monitoring method based on video AI and WLAN devices shown below Figure 2 will be introduced.

[0060] S210. Receive the video feature data of the target area uploaded by the camera and the wireless signal data of the target area uploaded by the WLAN device.

[0061] Specifically, the camera can first collect the video stream of the target area, and then use the built-in video AI model (such as Qwen series algorithms) to process the collected video stream in real time to obtain video feature data, and finally upload the video feature data to the central management platform.

[0062] Among them, the video feature data includes the timestamp of each video frame, the face features, movement trajectories, and behavior features of each person in the video stream. The face features are used to uniquely identify each person in the video stream, the movement trajectories are used to indicate the position coordinates of each person in the video stream at each timestamp, and the behavior features are used to indicate the behavior patterns of each person in the video stream, such as staying, walking, jumping, falling, etc.

[0063] The WLAN device can first detect the wireless signals emitted by nearby wireless devices in real time through WiFi probes, and then analyze the wireless signals through WiFi Sensing technology to obtain wireless signal data, and finally upload the wireless signal data to the central management platform.

[0064] Among them, the wireless signal data includes signal strength, the timestamp of receiving the wireless signal, MAC address, activity status, wireless channel and wireless frequency band, and WLAN device identifier. The signal strength represents the strength of the wireless signal, the MAC address is used to uniquely identify each wireless device, the activity status indicates whether the wireless device is actively emitting wireless signals or in an idle state, the wireless channel and wireless frequency band refer to the channel and frequency band used by the wireless device, and the WLAN device identifier is used to uniquely identify each WLAN device.

[0065] Considering that signal interference may occur due to the overlapping working frequency bands of the camera and the WLAN device, in a possible embodiment, the camera and the WLAN device use different working frequency bands. For example: the camera uses the 5GHz frequency band to transmit video feature data to the central management platform, and the WLAN device uses the 2.4GHz frequency band to collect wireless signal data and transmit it to the central management platform.

[0066] S220. Use the linear interpolation algorithm to align the timestamps of the video feature data and the wireless signal data to obtain synchronized data.

[0067] Specifically, each data point of the video feature data has a timestamp, and each data point of the wireless signal data also has a timestamp. After the central management platform obtains these two types of data, it can, based on the timestamps, use the linear interpolation algorithm to interpolate the video feature data to the time points of the wireless signal data, or interpolate the wireless signal data to the time points of the video feature data, so as to obtain synchronized data. Among them, the synchronized data includes the video feature data and the wireless signal data at each timestamp. For example {timestamp, video feature data, wireless signal data}.

[0068] For example, by interpolating the wireless signal data (sampled 10 times per second) to the target time points of the video feature data (30 fps), it is possible to iterate through each target time point t of the video feature data v , find the timestamps t1 and t2 of the two wireless signal data before and after it (t1 ≤ t v ≤ t2), and then use the interpolation formula to calculate the wireless signal data RSSI(t v ) corresponding to the target time point t. The interpolation formula is as follows: v )

[0069] RSSI(t v ) = RSSI(t1) + (t2 - t1)(t v - t1)·(RSSI(t2) - RSSI(t1))

[0070] where RSSI(t2) is the wireless signal data corresponding to the timestamp t2, RSSI(t1) is the wireless signal data corresponding to the timestamp t1, t v is the target time point, and RSSI(t v ) is the wireless signal data interpolated at t v .

[0071] In the embodiments of the present application, synchronizing the time of the video feature data and the wireless signal data can avoid analysis deviations caused by asynchronous data sources, ensure that the data corresponding to each timestamp is accurate, and enable efficient fusion across data sources.

[0072] S230. Based on the synchronized data, use the Euclidean distance between the video positioning coordinates and the WLAN triangulation positioning coordinates to determine whether it is the same person, and obtain matching data.

[0073] The matching data includes the video feature data and the wireless signal data of the same person at each timestamp, such as {person, timestamp, video feature data, wireless signal data}

[0074] In a possible embodiment, according to the video feature data at each timestamp, obtain the video positioning coordinates of multiple people in the target area at each timestamp; according to the wireless signal data at each timestamp, obtain the WLAN triangulation positioning coordinates of multiple people at each timestamp; for the same timestamp, determine whether the video positioning coordinates and the WLAN triangulation positioning coordinates belong to the same person according to whether the Euclidean distance between the video positioning coordinates and the WLAN triangulation positioning coordinates is less than a preset distance; and obtain matching data according to the video feature data and the wireless signal data corresponding to the same person at each timestamp.

[0075] In the specific implementation process, first, by analyzing the motion trajectories in the video feature data, the video positioning coordinates of each person in the video stream at each timestamp can be obtained. Second, according to the signal strength and the distribution positions of all WLAN devices in the target area, using the triangulation algorithm, the WLAN triangulation coordinates of the person carrying the wireless device at each timestamp can be calculated. Then, for the same timestamp, the Euclidean distance between any video positioning coordinate and any WLAN triangulation coordinate can be calculated. The specific formula is as follows:

[0076]

[0077] where d represents the Euclidean distance between the video positioning coordinate and the WLAN triangulation coordinate, (x v ,y v ) represents the video positioning coordinate, and (x w ,y w ) represents the WLAN triangulation coordinate.

[0078] If the Euclidean distance d is greater than or equal to a preset distance (e.g., 1.5m), it is determined that the video positioning coordinate and the WLAN triangulation coordinate do not belong to the same person. If the Euclidean distance is less than the preset distance (e.g., 1.5m), it is determined that the video positioning coordinate and the WLAN triangulation coordinate belong to the same person.

[0079] In the embodiments of the present application, spatial matching of the video feature data and the wireless signal data ensures that the behaviors and positions of each person at different time points can be accurately identified and tracked, avoiding judgment errors caused by spatial misalignment or errors.

[0080] In a possible embodiment, according to the wireless signal data at each timestamp, the WLAN triangulation coordinates of multiple people at each timestamp are obtained, including:

[0081] According to the wireless signal data corresponding to each timestamp, multiple WLAN device identifiers and multiple signal strengths corresponding to the same MAC address are determined; according to the multiple signal strengths and the positions of the multiple WLAN devices corresponding to the multiple WLAN device identifiers, the WLAN triangulation coordinates of the person carrying the target device among the multiple people at each timestamp are determined by using the triangulation algorithm.

[0082] In the specific implementation process, since multiple WLAN devices can detect the same wireless device simultaneously, the central management platform can receive the wireless signal data uploaded by these WLAN devices, obtain multiple WLAN device identifiers and multiple signal strengths corresponding to the same MAC address. Then, based on the distribution locations of all WLAN devices, the positions of multiple WLAN devices corresponding to multiple WLAN device identifiers can be determined. Finally, based on multiple signal strengths and the positions of multiple WLAN devices, using the triangulation algorithm, the WLAN triangulation coordinates of the person carrying the target device at each timestamp can be estimated, where the target device is the wireless device identified by the same MAC address.

[0083] For example, calculate the distance between the target device and each WLAN device according to the signal strength, and the formula is as follows:

[0084]

[0085] Where, P0 is the reference signal strength, γ is the path loss exponent, RSSI1 is the signal strength of the first WLAN device, d1 is the distance between the target device and the first WLAN device, RSSI2 is the signal strength of the second WLAN device, d2 is the distance between the target device and the second WLAN device, RSSI3 is the signal strength of the third WLAN device, and d3 is the distance between the target device and the third WLAN device.

[0086] Based on the geometric relationship of triangulation, the position of the target device satisfies the intersection problem of the following three circles. By jointly solving the following three equations, the WLAN triangulation coordinates (x, y) of the target device can be obtained.

[0087] (x - x1) 2 +(y - y1) 2 = d1 2

[0088] (x - x2) 2 +(y - y2) 2 = d2 2

[0089] (x - x3) 2 +(y - y3) 2 = d3 2

[0090] Where, (x1, y1) is the position of the first WLAN device, (x2, y2) is the position of the second WLAN device, and (x3, y3) is the position of the third WLAN device.

[0091] S240. Input the video feature data of the target person at each timestamp into the CNN of the anomaly detection model to extract video feature vectors, and input the wireless signal data of the target person at each timestamp into the RNN of the anomaly detection model to extract wireless signal feature vectors. Then, use the attention mechanism to perform weighted fusion on the video feature vectors and wireless signal feature vectors to obtain the fused feature vectors.

[0092] The anomaly detection model in the embodiments of the present application uses a CNN+RNN dual-stream network. Please refer to Figure 3 , which is the structural schematic diagram of the CNN+RNN dual-stream network provided by the embodiments of the present application. It can be seen that the anomaly detection model includes a CNN branch, an RNN branch, a fusion layer, and a fully connected layer. Among them, the CNN branch can adopt the ResNet-50 structure, the dimension of the input video feature data is 224×224×3, and it outputs a 1024-dimensional video feature vector; the RNN branch can adopt a double-layer GRU, the sequence length T of the input wireless signal data is 10, and it outputs a 512-dimensional wireless signal feature vector; the fusion layer can generate a 1280-dimensional fused feature vector through attention weights.

[0093] In the specific implementation process, first splice the video feature vectors and wireless signal feature vectors, then perform a linear transformation through a trainable weight matrix, and obtain the attention weight α through the sigmoid activation function. The formula is as follows:

[0094] α = σ(W α ·[h v , h2])

[0095] Among them, h v is the video feature vector extracted by ResNet-50, h w is the wireless signal feature vector extracted by GRU, [h v , h w means splicing h v and h w , W α is the trainable weight matrix, α is the attention weight, σ(·) is the sigmoid activation function, and the output value is mapped to [0, 1].

[0096] Use the attention weight α to perform weighted fusion on the video feature vectors and wireless signal feature vectors. The formula is as follows:

[0097] h fusion = α·h v + (1 - α)·h w

[0098] Among them, h fusion represents the fused feature vector.

[0099] It should be noted that the entire model uses the cross - entropy loss function and is optimized using the Adam optimizer. The learning rate is set to 0.001. This design allows the model to automatically learn how to allocate weights between the video feature vectors and the wireless signal feature vectors, so as to better fuse information for subsequent anomaly detection.

[0100] S250. Input the fused feature vector into the fully - connected layer of the anomaly detection model to obtain the detection result.

[0101] Specifically, the anomaly detection model can use a multi - layer perceptron (MLP) as the fully - connected layer, and the fully - connected layer is used to output the detection result. The detection result includes the behavior categories and confidence levels of all personnel within the target area. The behavior categories include normal behavior and abnormal behavior, and abnormal behaviors such as long - term stay and abnormal aggregation. The confidence level represents the credibility of the model's judgment of the behavior category, for example, a score from 0 to 1 or a value from 0 to 100. The higher the confidence level, the more likely the judgment of the anomaly detection model is to be accurate.

[0102] For example, if there are 3 people in the target area, the detection result example is as follows:

[0103] Person 1: Behavior category: Normal behavior, Confidence level: 0.92;

[0104] Person 2: Behavior category: Abnormal behavior, Abnormal type: Long - term stay, Confidence level: 0.85;

[0105] Person 3: Behavior category: Abnormal behavior, Abnormal type: Abnormal aggregation, Confidence level: 0.88.

[0106] S260. If the detection result indicates that there are people with abnormal behaviors in the target area, generate a warning message.

[0107] Specifically, if the behavior category of any person in the detection result is abnormal behavior, an alarm will be automatically triggered, and a warning message will be generated according to the abnormal type and confidence level, and the security personnel will be notified by text message or APP.

[0108] In a possible embodiment, the central management platform includes a database, which pre - stores the face features and business areas of multiple business personnel. If the detection result indicates that there are people with abnormal behaviors in the target area, extract the target face features of the target personnel with abnormal behaviors from the video feature data; match the target face features with the face features of multiple business personnel; if the match is unsuccessful, generate a warning message.

[0109] In the specific implementation process, if the detection result indicates that there is an abnormal behavior of a person in the target area, the central management platform can extract the target face feature of the target person with the abnormal behavior according to the face feature in the video feature data, calculate the similarity between the target face feature and the face features of multiple business personnel in the database. If the similarity is less than the preset similarity, it means that the matching is unsuccessful, and a warning message is generated.

[0110] Considering that the target person with abnormal behavior may be a staff member, in the embodiment of the present application, by extracting the face feature of the target person in the video and comparing it with the personnel in the database, it is possible to quickly and accurately identify whether the target person is an authorized person and automatically complete the identity authentication. If the facial features of the target person do not match the features of the business personnel in the database, the abnormal individual can be quickly identified and a warning can be issued, which helps to detect potential security threats.

[0111] Traditional visitor management and identity authentication rely on manual review or single face recognition, which not only increases the operation complexity but also reduces the authentication efficiency and accuracy. Therefore, in a possible embodiment, after matching the target face feature with the face features of multiple business personnel, the method may further include:

[0112] If the matching is successful, determine the target business area of the target person from the database; determine the location information of the target person according to the wireless signal data; if the location information of the target person belongs to the target business area, assign access rights to the target person; if the location information of the target person does not belong to the target business area, generate a warning message.

[0113] In the specific implementation process, if the similarity is greater than or equal to the preset similarity, it means that the matching is successful, and then the business area of the business personnel corresponding to the target face feature is determined as the target business area of the target person from the database. Then, according to the signal strength in the wireless signal data and the distribution positions of all WLAN devices, the triangulation algorithm is used to determine the location information of the target person. Further, the geographical fence algorithm can be used to determine whether the location information of the target person belongs to the target business area. If so, access rights are assigned to the target person, such as opening the door, allowing entry into a specific area, etc. If not, an alarm is automatically triggered, a warning message is generated, and the security personnel are notified by text message or AP.

[0114] According to the different shapes of the target business area, the specific geographical fence algorithms are different, and they will be introduced separately below.

[0115] First, the target business area is a circular geographical fence.

[0116] The distance between the location of the target person and the center point of the fence can be calculated using the Haversine formula or the Vincenty formula. If the distance is less than or equal to the radius, the target person is inside the circular geographical fence; otherwise, the target person is outside the circular geographical fence.

[0117] Second, the target business area is a rectangular geographical fence.

[0118] The longitude and latitude coordinates (lat1, lon1) of the lower left corner and the longitude and latitude coordinates (lat2, lon2) of the upper right corner of the target business area can be determined, and it can be judged whether the longitude and latitude (lat t , lon t ) of the location of the target person is within the longitude and latitude range of the rectangular geographical fence. If lat1 ≤ lat t ≤ lat2 and lon1 ≤ lon t ≤ lon2, the target person is inside the rectangular geographical fence; otherwise, the target person is outside the rectangular geographical fence.

[0119] Third, the target business area is a polygonal geographical fence.

[0120] A ray can be drawn from the location of the target person to check whether the ray intersects each side of the polygonal geographical fence. If the number of intersection points of the ray and the polygonal geographical fence is odd, the target person is inside the polygonal geographical fence; if it is even, the target person is outside the polygonal geographical fence.

[0121] In the embodiments of the present application, by combining the location information of face recognition and wireless signals, when the system determines whether a target person has access rights, a double-verification method is adopted. It is necessary not only to confirm whether the identity of the person matches, but also to ensure that the location information of the person meets the requirements of the predetermined business area, which greatly reduces the probability of misjudgment and improves the security of permission allocation.

[0122] In some scenarios, when there are video conferences, large-scale events or security peaks, the traditional WLAN network resources are configured relatively statically and are difficult to adjust dynamically in a timely manner, resulting in a decline in network quality and affecting video transmission and data interaction. In a possible embodiment, after S250, the method may further include the following steps:

[0123] If the detection result indicates that there is no abnormal behavior in the target area, determine the scene type to which the target area belongs according to the video feature data; determine the WLAN network load of the target area according to the wireless signal data; if the scene type is a meeting room and the WLAN network load is higher than the first threshold, preferentially allocate 5GHz band resources and adjust the channel width to 40MHz; if the scene type is a shopping mall and the 2.4GHz band network load is higher than the second threshold, switch from the 2.4GHz band to the 5GHz band and enable the load balancing algorithm.

[0124] In the specific implementation process, by analyzing the video feature data, the scene type of the target area can be identified, such as a meeting room, a shopping mall, and a public transportation hub. Monitor the WLAN network load of the target area according to the wireless signal data. The WLAN network load refers to the overall network load brought by device connections in a wireless local area network (such as the number of devices, network traffic, bandwidth utilization, channel occupancy, etc.). The WLAN network load includes the 2.4GHz band network load, and the 2.4GHz band network load refers to the network burden generated by the use of multiple devices on the 2.4GHz wireless band.

[0125] If the scene type is a meeting room and the WLAN network load is higher than the first threshold, the 5GHz band resources can be preferentially allocated to the meeting devices, and at the same time, by adjusting the channel width to 40MHz, the high-bandwidth requirements can be met. If the scene type is a shopping mall and the 2.4GHz band network load is higher than the second threshold, some terminals can be switched from the 2.4GHz band to the 5GHz band with less interference, and at the same time, the load balancing algorithm is adopted to reduce the load of a single WLAN device. If the scene type is a public transportation hub (airport, station), the DHCP lease time can be shortened, roaming authentication resources can be pre-allocated, and a fast handover mechanism can be adopted to reduce the handover delay. When the detection result indicates no abnormal behavior, the QoS priority queue can also be adopted to mark the video feature data as high priority, improve the transmission efficiency of the video feature data, and thus improve the real-time performance of subsequent abnormal behavior detection.

[0126] In the embodiment of the present application, by combining video feature data, wireless signal data, scene type, and network load to dynamically adjust channel allocation and spectrum resources, the utilization rate of network resources can be significantly improved, interference can be reduced, data transmission of key services can be optimized, the user experience can be improved, and at the same time, intelligent management of the network can be realized.

[0127] In a possible embodiment, after S260, the method may further include the following steps:

[0128] Generate a dynamic situation map of personnel activities in the target area using the heat map algorithm based on the matching data and abnormal behavior data; the abnormal behavior data includes the timestamp, location coordinates, and personnel where the abnormal behavior occurs.

[0129] In the specific implementation process, the kernel density estimation algorithm is used to generate the heat map, and the formula is as follows:

[0130]

[0131] where f(x, y) is the density value at each position (x, y), K is the Gaussian kernel function, h is the bandwidth parameter, (x i , y i ) is the position coordinate of the i-th person, and n is the total number of people.

[0132] The central management platform can present the density of personnel activities through the heat map, and display the activity intensity in this area through different colors (such as gradually changing from blue to red). As time goes by, the movement trajectories of each person are shown, and the movement trajectories of different people can be distinguished using different line styles or colors. For abnormal events, the area or person within the time period of the abnormal event can be marked with a warning color, thus forming a dynamic situation map. The dynamic situation map shows the activity trajectories of personnel and the occurrence of abnormal behaviors through visualization technology.

[0133] In the embodiment of this application, the central management platform can collect, analyze, and store the data collected by each device in multiple areas to form a dynamic situation map. Users can clearly observe the dynamic changes of personnel and potential abnormal events in the target area through the dynamic situation map, and discover problems in a timely manner and respond.

[0134] In the prior art, the combination of video and WLAN positioning is mainly used for trajectory tracking, that is, to determine the position and movement path of the target through video and wireless signal data, and does not involve the feature fusion at the behavior recognition level. This application provides a security monitoring method based on video AI and WLAN devices, which combines video AI and WLAN devices, and uses the attention mechanism to perform feature-level fusion on these two types of spatio-temporally heterogeneous data, namely video feature data and wireless signal data, to achieve cross-modal data interaction, and realizes the recognition of abnormal behaviors through the fused data, so as to achieve precise security monitoring and management.

[0135] Traditional single-video detection cannot identify abnormal behaviors in occluded scenarios due to the lack of video data. Compared with traditional single-video detection, the method provided in the embodiments of the present application integrates video feature data and wireless signal data. In the shopping mall anti-theft scenario, the missed detection rate is reduced from 18% to 5%. In the abnormal gathering detection in the subway station, the false alarm rate drops from 15% to 4%. The personnel positioning error is reduced from 2.0m in single WLAN positioning to 0.5m, and the accuracy of abnormal behavior detection is increased to 95% (only 82% in single-video detection).

[0136] In summary, the present application has the following advantages:

[0137] 1. Deep data fusion: Through multi-modal data fusion technology, it breaks through the problems of traditional video surveillance and WiFi probe data islands, and realizes abnormal detection and personnel tracking with higher accuracy.

[0138] 2. Real-time and high accuracy: Utilize video AI and abnormal detection models (i.e., machine learning models) to achieve real-time behavior pattern recognition, reduce false alarm and missed detection rates, and improve the response speed of the monitoring system.

[0139] 3. Dynamic network resource optimization: Use video feature data for scene recognition, and use wireless network data to monitor the current network load and signal interference conditions, ensuring the optimization of WLAN resource configuration in high-load scenarios and improving the transmission quality of key business data.

[0140] 4. Invisible visitor authentication: Based on the invisible authentication jointly by video AI and wireless network data, it simplifies the traditional authentication process and improves the efficiency of visitor management.

[0141] 5. Centralized management and data tracking: The central management platform supports cross-regional and cross-device data integration, facilitating security situation monitoring, historical event tracking, and post-event analysis.

[0142] Based on the same inventive concept, please refer to Figure 4 , the present application also provides a security monitoring device based on video AI and WLAN devices, which is set in the central management platform of the security monitoring system. The security monitoring system also includes a camera and WLAN devices; the device includes:

[0143] A data acquisition module, configured to receive the video feature data of the target area uploaded by the camera, and the wireless signal data of the target area uploaded by the WLAN device;

[0144] A data fusion module, which is used to align the timestamps of video feature data and wireless signal data by using a linear interpolation algorithm to obtain synchronized data; based on the synchronized data, judge whether it is the same person by using the Euclidean distance between the video positioning coordinates and the WLAN triangulation positioning coordinates to obtain matching data; the matching data includes the video feature data and wireless signal data of the same person at each timestamp; input the video feature data of the same person at each timestamp into the CNN of the anomaly detection model to extract video feature vectors, and input the wireless signal data of the same person at each timestamp into the RNN of the anomaly detection model to extract wireless signal feature vectors, and use the attention mechanism to perform weighted fusion on the video feature vectors and wireless signal feature vectors to obtain fused feature vectors;

[0145] An anomaly detection module, which is used to input the fused feature vectors into the fully connected layer of the anomaly detection model to obtain detection results;

[0146] A warning module, which is used to generate a warning message if the detection result indicates that there is an abnormal behavior of a person in the target area.

[0147] It should be noted that in this embodiment, each module in the security monitoring device based on video AI and WLAN devices corresponds one by one to each step in the security monitoring method based on video AI and WLAN devices in the foregoing embodiment. Therefore, the specific implementation manner of this embodiment may refer to the implementation manner of the foregoing security monitoring method based on video AI and WLAN devices, and will not be elaborated here.

[0148] Based on the same inventive concept, the present application also provides a security monitoring system, which includes a camera, a WLAN device and a central management platform; wherein:

[0149] The camera is used to collect the video stream of the target area, extract video feature data based on the video stream, and upload the video feature data to the central management platform;

[0150] The WLAN device is used to collect wireless signal data of the target area and upload the wireless signal data to the central management platform;

[0151] The central management platform is used to implement the foregoing security monitoring method based on video AI and WLAN devices.

[0152] It should be noted that, in this document, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, such that a process, method, article or system comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or system. Without further limitation, an element defined by the statement "comprising one..." does not exclude the presence of additional identical elements in the process, method, article or system comprising such element.

[0153] The serial numbers of the embodiments of the present application above are only for description and do not represent the superiority or inferiority of the embodiments.

[0154] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. The computer software product is stored in a storage medium (such as a read-only memory / random access memory, magnetic disk, optical disk), and includes several instructions for causing a multimedia terminal device (which can be a mobile phone, a computer, a television receiver, or a network device, etc.) to execute the methods of the various embodiments of the present application.

[0155] The above are only the preferred embodiments of the present application and do not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made by using the description of the specification and the attached drawings of the present application, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present application.

Claims

1. A security monitoring method based on video AI and WLAN devices, characterized in that, A central management platform applied to a security monitoring system, the security monitoring system further including a camera and a WLAN device; the method includes: Receiving video feature data of a target area uploaded by the camera and wireless signal data of the target area uploaded by the WLAN device; Using a linear interpolation algorithm to align the timestamps of the video feature data and the wireless signal data to obtain synchronized data; Based on the synchronized data, using the Euclidean distance between the video positioning coordinates and the WLAN triangular positioning coordinates to determine whether it is the same person, and obtaining matching data; the matching data includes the video feature data and the wireless signal data of the same person at each timestamp; Inputting the video feature data of the same person at each timestamp into the CNN of the anomaly detection model to extract video feature vectors, and inputting the wireless signal data of the same person at each timestamp into the RNN of the anomaly detection model to extract wireless signal feature vectors, and using an attention mechanism to perform weighted fusion on the video feature vectors and the wireless signal feature vectors to obtain fused feature vectors; Inputting the fused feature vectors into the fully connected layer of the anomaly detection model to obtain a detection result; If the detection result indicates that there is an abnormal behavior of a person in the target area, generating a warning message.

2. The security monitoring method based on video AI and WLAN devices according to claim 1, characterized in that, The determining whether it is the same person based on the synchronized data by using the Euclidean distance between the video positioning coordinates and the WLAN triangular positioning coordinates to obtain matching data includes: Obtaining the video positioning coordinates of multiple persons in the target area at each timestamp according to the video feature data at each timestamp; Obtaining the WLAN triangular positioning coordinates of the multiple persons at each timestamp according to the wireless signal data at each timestamp; For the same timestamp, determining whether the video positioning coordinates and the WLAN triangular positioning coordinates belong to the same person according to whether the Euclidean distance between the video positioning coordinates and the WLAN triangular positioning coordinates is less than a preset distance; Obtaining matching data according to the video feature data and the wireless signal data of the same person at each timestamp.

3. The security monitoring method based on video AI and WLAN devices according to claim 2, wherein, The wireless signal data includes a MAC address, a signal strength, and a WLAN device identifier; The obtaining the WLAN triangular positioning coordinates of the multiple persons at each timestamp according to the wireless signal data at each timestamp includes: Determining multiple WLAN device identifiers and multiple signal strengths corresponding to the same MAC address according to the wireless signal data at each timestamp; Using a triangular positioning algorithm to determine the WLAN triangular positioning coordinates of the person carrying the target device among the multiple persons at each timestamp according to the multiple signal strengths and the positions of the multiple WLAN devices corresponding to the multiple WLAN device identifiers; the target device is the wireless device identified by the same MAC address.

4. The security monitoring method based on video AI and WLAN devices according to claim 1, wherein, After obtaining the detection result according to the fused feature vectors, the method further includes: If the detection result indicates that there is no abnormal behavior of a person in the target area, determining the scene type to which the target area belongs according to the video feature data; Determine the WLAN network load of the target area according to the wireless signal data; the WLAN network load includes the network load of the 2.4GHz band; If the scenario type is a meeting room and the WLAN network load is higher than the first threshold, preferentially allocate 5GHz band resources and adjust the channel width to 40MHz; If the scenario type is a shopping mall and the network load of the 2.4GHz band is higher than the second threshold, switch from the 2.4GHz band to the 5GHz band and enable the load balancing algorithm.

5. The security monitoring method based on video AI and WLAN devices according to claim 1, wherein, The central management platform includes a database, and the database pre-stores the face features of multiple business personnel; If the detection result indicates that someone in the target area has abnormal behavior, generate a warning message, including: If the detection result indicates that someone in the target area has abnormal behavior, extract the target face features of the target person with abnormal behavior from the video feature data; Match the target face features with the face features of the multiple business personnel; If the match is unsuccessful, generate a warning message.

6. The security monitoring method based on video AI and WLAN devices according to claim 5, characterized in that The database also pre-stores the business area of each business personnel; after matching the target face features with the face features of the multiple business personnel, the method further includes: If the match is successful, determine the target business area of the target person from the database; Determine the location information of the target person according to the wireless signal data; If the location information of the target person belongs to the target business area, allocate access rights to the target person.

7. The security monitoring method based on video AI and WLAN devices according to claim 6, wherein After determining the location information of the target person according to the wireless signal data, the method further includes: If the location information of the target person does not belong to the target business area, generate a warning message.

8. The security monitoring method based on video AI and WLAN devices according to claim 1, characterized in that After generating a warning message if the detection result indicates that someone in the target area has abnormal behavior, the method further includes: Generate a dynamic situation map of the activities of the people in the target area according to the matching data and the abnormal behavior data by using the heat map algorithm; the abnormal behavior data includes the timestamp, location coordinates and personnel of the abnormal behavior.

9. A security monitoring method based on video AI and WLAN devices, characterized in that, It is set in the central management platform of the security monitoring system, and the security monitoring system further includes a camera and a WLAN device; the device includes: A data acquisition module, configured to receive the video feature data of the target area uploaded by the camera and the wireless signal data of the target area uploaded by the WLAN device; A data fusion module, which is used to align the timestamps of the video feature data and the wireless signal data by using a linear interpolation algorithm to obtain synchronized data; based on the synchronized data, judge whether it is the same person by using the Euclidean distance between the video positioning coordinates and the WLAN triangular positioning coordinates to obtain matching data; the matching data includes the video feature data and the wireless signal data of the same person at each timestamp; input the video feature data of the same person at each timestamp into the CNN of the anomaly detection model to extract video feature vectors, and input the wireless signal data of the same person at each timestamp into the RNN of the anomaly detection model to extract wireless signal feature vectors, and use an attention mechanism to perform weighted fusion on the video feature vectors and the wireless signal feature vectors to obtain fused feature vectors; An anomaly detection module, which is used to input the fused feature vectors into the fully connected layer of the anomaly detection model to obtain a detection result; An early warning module, which is used to generate an early warning message if the detection result indicates that there is an abnormal behavior of a person in the target area.

10. A security monitoring system, characterized in that, The security monitoring system includes a camera, a WLAN device and a central management platform; wherein: The camera is used to collect the video stream of the target area, extract video feature data based on the video stream, and upload the video feature data to the central management platform; The WLAN device is used to collect the wireless signal data of the target area and upload the wireless signal data to the central management platform; The central management platform is used to implement the security monitoring method based on video AI and WLAN device according to any one of claims 1-8.