3D Scene Construction Method and System for Police Cloud IoT Devices Based on Digital Twin

The digital twin-based system for police-cloud IoT devices enhances scene visualization and interaction by real-time risk detection and response, addressing both environmental and human risks in crowded spaces.

CN119964346BActive Publication Date: 2025-07-15CONGWEN SOFTWARE TECHNOLOGICAL SHENZHEN CITY
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510422120.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-07-15
Estimated Expiration
2045-04-07

AI Technical Summary

Technical Problem

The existing digital twin technology has a relatively single application range in the field of police cloud IoT devices, and it has failed to effectively monitor and deal with man-made risks and environmental hazards in densely populated public places, especially risk monitoring and early warning of suspicious people.

Method used

Surveillance cameras and environmental sensors are arranged in the target area, a three-dimensional cloud-based solid model is established, scene videos and environmental data are analyzed in real time through the first inspection unit, disposal plans are generated, and multi-section video is processed through the second inspection unit to identify suspicious people and risk levels, screen high-risk personnel and generate alarms.

Benefits of technology

Real-time risk monitoring and emergency response to densely populated places have been achieved, the speed and accuracy of emergency response has been improved, false alarms and missed reports have been reduced, and safety management efficiency has been improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119964346B_ABST
    Figure CN119964346B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and system for building a three-dimensional scene of police cloud Internet of Things devices based on digital twins, belonging to the field of digital twin technology. The method includes: arranging sensing devices in a target area; establishing a three-dimensional entity model of the target area in a cloud server, arranging corresponding virtual models in the entity model based on the installation information of the sensing devices, and accessing the data from the sensing devices into the virtual models to generate digital twin models; a first troubleshooting unit analyzes the scene video and environmental data collected by the sensing devices in real time, and generates a disposal plan if environmental risks are detected; a second troubleshooting unit processes the scene video into multi-segment videos, and analyzes the multi-segment videos in combination with environmental data to determine suspicious persons and risk levels in the target area; screening the suspicious persons with risk levels greater than a first threshold as high-risk risks. The present invention helps security management personnel to timely pay attention to and handle high-risk risks, and effectively improves the emergency response speed and accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of digital twins, and particularly relates to a method and system for building a three-dimensional scene of police cloud Internet of Things devices based on digital twins. Background Art

[0002] In the field of police cloud Internet of Things devices, by associating physical entities such as police equipment and Internet of Things sensors with digital models through digital twin technology, remote monitoring and danger warning can be achieved. In addition, digital twin technology also has characteristics such as high precision, scalability, and visualization, and can support three-dimensional modeling and simulation of large-scale and complex scenarios, providing strong technical support for the management and decision-making of police cloud Internet of Things devices.

[0003] For the application of digital twins, in the prior art, for example, a Chinese patent document with the publication number CN116187105A discloses a fire evacuation planning method and system based on digital twin technology. This method obtains the current combustion range according to the fire starting position and camera position in the monitoring screen information, inputs the current combustion range into the building digital twin model, predicts the expansion speed of the combustion range according to the change of the combustion range per unit time, simulates the combustion spread trend in the building digital twin model, and then according to the current position information, escape speed information, and protection level information of personnel, uses the building digital twin model for simulation, and makes evacuation planning and guidance according to the simulation results, so that the escape route of each person is more efficient and reasonable.

[0004] However, the above digital twin technology is only applied in the field of fire escape, and its function is relatively single. For densely populated public places, in addition to considering environmental hazards such as fires, the human risks of suspicious persons also need to be considered. Therefore, it is necessary to further expand the application scope of digital twin technology in the field of police cloud Internet of Things devices and incorporate the monitoring, warning, and response to the human risks of suspicious persons. Summary of the Invention

[0005] To solve the above problems, the present invention provides a method and system for building a three-dimensional scene of police cloud Internet of Things devices based on digital twins to solve the problems existing in the prior art.

[0006] To achieve the above invention purpose, the present invention proposes a method for building a three-dimensional scene of police cloud Internet of Things devices based on digital twins, including:

[0007] Deploy sensing devices in the target area. The sensing devices include surveillance cameras and environmental sensors. The surveillance cameras are used to record scene videos, and the environmental sensors are used to detect environmental data;

[0008] Establish a three-dimensional entity model of the target area in the cloud server, deploy the corresponding virtual model in the entity model based on the installation information of the sensing device, access the data from the sensing device into the corresponding virtual model, and generate a digital twin model;

[0009] The first investigation unit analyzes the scene video and the environmental data in real time. If an environmental risk is detected, a disposal plan is generated and displayed in the digital twin model;

[0010] The second investigation unit processes the scene video into multi-segment videos, and analyzes the multi-segment videos in combination with the environmental data to determine the suspicious persons and risk levels in the target area;

[0011] Select the suspicious persons with a risk level greater than the first threshold as high-risk risks, and generate a suspicious risk alarm pointing to the high-risk risks in the digital twin model.

[0012] Furthermore, the real-time analysis of the scene video and the environmental data includes the following steps:

[0013] Define the scene video to be analyzed as the video to be analyzed. Based on the standard interval, extract the video to be analyzed into multiple frame images, locate the frame image at the current time in the video to be analyzed. If a dangerous marker is detected in the frame image, determine the corresponding environmental risk according to the dangerous marker, and generate the disposal plan;

[0014] The environmental sensors include temperature sensors, smoke sensors and current sensors. The environmental data includes temperature data, smoke data and current data. In the case where the dangerous marker is not detected, if the environmental data exceeds the corresponding set threshold, it is determined that there is the environmental risk, and the disposal plan is generated.

[0015] Furthermore, processing the scene video into the multi-segment videos includes the following steps:

[0016] The multi-segment videos include a first segment, a second segment and a third segment. Arrange the frame images in chronological order, extract the change regions between adjacent frame images based on the frame difference method, and count the number of pixel changes in the change regions;

[0017] If the number of changes is greater than or equal to the dynamic threshold, mark the two frame images that generate the number of changes as the first dynamic images. If the number of changes is less than the dynamic threshold, obtain the size of the dynamic threshold. If the dynamic threshold is greater than or equal to the second threshold, mark the two frame images as the second dynamic images. Otherwise, mark the frame images as static images, integrate the first dynamic images into the first section, the second dynamic images into the second section, and the static images into the third section.

[0018] Further, determining the dynamic threshold includes the following steps:

[0019] Obtain a historical video, process the historical video into a dynamic time series regarding the number of changes, take the preset window size as a parameter, process the dynamic time series based on the sliding window method to obtain historical window values of multiple time windows, split the recording time of the historical video into multiple first time periods, and statistically analyze the historical window values to determine the adjustment coefficient of each first time period;

[0020] Extract the average value of the number of changes within the second time period before the two frame images, determine the corresponding adjustment coefficient based on the first time period where the two frame images are located, correct the average value based on the adjustment coefficient to obtain a first value, set a base value, and correct the base value based on the first value to obtain the dynamic threshold.

[0021] Further, analyzing the multi-section video includes the following steps:

[0022] Define the first section and the second section in the multi-section video as the first analysis section and the second analysis section respectively. Perform target tracking on the first analysis section based on the target detection algorithm. If personnel movement is detected, calculate the risk degree based on the occurrence frequency, activity trajectory entropy value, and the number of face matching failures;

[0023] Obtain the second time period when the second analysis section appears, intercept the video images of the scene video captured by other monitoring cameras within the second time period, define it as the auxiliary section, split the second analysis section and the auxiliary section into frame images, combine the frame images with the same shooting time into an image group, perform mapping analysis on the image group, and obtain the coordinate positions of the same target in different frame images;

[0024] The environmental sensor includes a beacon sensor. The beacon sensor obtains the return signal strength based on the employee's work card, combines the return signal strength with reference to time into a signal group corresponding to the image group, and establishes an analysis model. The analysis model outputs the risk degree of a suspicious person based on the coordinate positions of the image group and the return signal strength in the corresponding signal group.

[0025] Further, establishing the analysis model includes the following steps:

[0026] Collect historical data, where the historical data includes the image group, the signal group, and output labels. The output labels include no risk and at risk. Numerically adjust the coordinate positions in the image group and the return signal strength in the signal group to expand the historical data and obtain augmented data. Establish the analysis model based on a neural network and train the analysis model using the augmented data. When the output result of the analysis model is at risk, use the probability value of being at risk as the risk degree.

[0027] Further, generating the disposal plan includes the following steps:

[0028] The disposal plan includes multiple predetermined escape routes. If there are multiple escape routes, determine the fire location, simulate the fire spread direction based on the fire location, and select the best route from the predetermined evacuation routes based on the spread direction.

[0029] Further, the danger markers include flames, smoke, and accumulated water.

[0030] The present invention also provides a three-dimensional scene construction system for a police cloud Internet of Things device based on digital twins. This system is used to implement the above-mentioned three-dimensional scene construction method for a police cloud Internet of Things device based on digital twins. The system includes:

[0031] Sensing devices, where the sensing devices include surveillance cameras and environmental sensors. The surveillance cameras are used to record scene videos, and the environmental sensors are used to detect environmental data;

[0032] A cloud server, which is used to establish a three-dimensional entity model of the target area, deploy corresponding virtual models in the entity model based on the installation information of the sensing devices, and connect the data from the sensing devices to the corresponding virtual models to generate a digital twin model;

[0033] A first investigation unit, which is used to analyze the scene video and the environmental data in real time. If environmental risks are detected, generate a disposal plan and display it in the digital twin model;

[0034] The second screening unit is used to process the scene video into a multi-segment video, analyze the multi-segment video in combination with the environmental data, determine the suspicious persons and risk levels in the target area, screen the suspicious persons with risk levels greater than the first threshold as high-risk risks, and generate a suspicious risk alarm pointing to the high-risk risks in the digital twin model.

[0035] Advantageous effects:

[0036] By applying digital twin technology, the present invention establishes a three-dimensional entity model of the target area in the cloud server, corresponding to the actually deployed monitoring cameras and environmental sensors, forming a digital twin model that can reflect the actual scene state in real time. This model not only enhances the visualization and interactivity of the scene, but also enables users to view various data on the site at any time. At the same time, the cloud server analyzes the scene video and environmental data in real time, can timely detect environmental risks including fires and generate corresponding disposal plans, and effectively improves the emergency response speed and accuracy by displaying the above content in the digital twin model. Description of the drawings

[0037] Figure 1 It is a flowchart of the steps of a three-dimensional scene construction method for police cloud IoT devices based on digital twins according to the present invention;

[0038] Figure 2 It is a schematic structural diagram of a three-dimensional scene construction system for police cloud IoT devices based on digital twins according to the present invention. Detailed implementation manners

[0039] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0040] It can be understood that the terms "first", "second", etc. used in the present application may be used herein to describe various elements, but unless otherwise specified, these elements are not limited by these terms. These terms are only used to distinguish the first element from another element. For example, without departing from the scope of the present application, the first xx script may be called the second xx script, and similarly, the second xx script may be called the first xx script.

[0041] As Figure 1 shown, a three-dimensional scene construction method for police cloud IoT devices based on digital twins includes:

[0042] S1: Deploy sensing devices in the target area. The sensing devices include monitoring cameras and environmental sensors. The monitoring cameras are used to record scene videos, and the environmental sensors are used to detect environmental data.

[0043] S2: Establish a three-dimensional entity model of the target area in the cloud server, deploy corresponding virtual models in the entity model based on the installation information of the sensing devices, and connect the data from the sensing devices to the corresponding virtual models to generate a digital twin model.

[0044] The target area includes various densely populated areas such as office buildings. The sensing devices include surveillance cameras and various types of environmental sensors. The environmental sensors include flame sensors, smoke sensors, temperature sensors, humidity sensors, current sensors, etc. Different types of sensors collect corresponding environmental data and upload it to the cloud server. A three-dimensional entity model corresponding to the actual scene is established in the cloud server, and virtual models corresponding to the surveillance cameras and environmental sensors are set in the three-dimensional entity model. After the cloud server receives the scene video and environmental data, it connects them to the corresponding virtual models, thereby obtaining a digital twin model that corresponds and maps to the actual scene. Users can click on the virtual models in the digital twin model to view various on-site data in real time.

[0045] S3: The first investigation unit analyzes the scene video and environmental data in real time. If an environmental risk is detected, a disposal plan is generated and displayed in the digital twin model.

[0046] After the cloud server receives the scene video and environmental data, it intercepts the frame image at the latest moment in the scene video for analysis, and obtains the latest data in the environmental data for analysis to determine whether there is an environmental risk. Environmental risks include fires, thick smoke, waterlogging, and overloaded electrical circuits. For example, when it is determined that there is a flame in the current frame image, it is determined that there is a fire risk. Or when the smoke sensor detects that the particle concentration greater than 2.5μm exceeds a certain threshold, it is determined that there is a thick smoke risk. If an environmental risk is detected, it is displayed in the digital twin model. At the same time, a corresponding disposal plan is also generated. For example, for a fire risk, an escape route is generated, and for waterlogging or overloaded short circuits, corresponding solutions are generated to prevent the disaster from expanding.

[0047] S4: After the second investigation unit processes the scene video into multi-segment videos, it combines the environmental data to analyze the multi-segment videos to determine the suspicious persons and risk levels in the target area.

[0048] For areas such as office buildings, in addition to environmental risks, the human risks of outsiders also need to be considered. However, due to reasons such as mask occlusion, relying solely on face recognition technology and using the current few frames of images for judgment may not achieve the expected effect, and it is necessary to combine historical multi-frame images for comprehensive recognition of behavioral trajectories. For this purpose, in order to improve the recognition efficiency, the scene video is processed into a multi-segment video before recognition. Specifically, the multi-segment video includes a dynamic segment and a static segment. After division, only the dynamic segment is analyzed and recognized, which can speed up the recognition efficiency. The specific calculation methods for identifying suspicious persons and the risk level will be introduced later.

[0049] S5: Select the suspicious persons with a risk level greater than the first threshold as high-risk risks, and generate a suspicious risk alarm pointing to the high-risk risks in the digital twin model.

[0050] If it is determined that there are suspicious persons in the scene video and their risk level is greater than the first threshold, for example, the first threshold is 0.8, then the suspicious persons will be regarded as high-risk risks, and a risk alarm specifying this area will be generated in the digital twin model.

[0051] By applying digital twin technology, the present invention establishes a three-dimensional entity model of the target area in the cloud server, which corresponds to the actually deployed monitoring cameras and environmental sensors, forming a digital twin model that can reflect the actual scene state in real time. This model not only enhances the visualization and interactivity of the scene, but also enables users to view various data on the site at any time. At the same time, the cloud server analyzes the scene video and environmental data in real time, can timely detect environmental risks including fires and generate corresponding disposal plans, and effectively improves the emergency response speed and accuracy by displaying the above content in the digital twin model.

[0052] In order to improve the recognition efficiency of suspicious persons, the system in the present invention processes the scene video into a multi-segment video, and only analyzes and recognizes the dynamic segment. At the same time, it combines environmental data and historical multi-frame images for comprehensive judgment. This method can more accurately identify suspicious persons and their risk levels, reducing false alarms and missed alarms. For suspicious persons with a risk level greater than the set threshold, the system can automatically screen them as high-risk risks and generate a suspicious risk alarm pointing to this area in the digital twin model, which helps security management personnel to pay attention to and handle high-risk risks in a timely manner, improving the safety management level and efficiency.

[0053] In this embodiment, the real-time analysis of the scene video and environmental data includes the following steps:

[0054] Define the scene video to be analyzed as the video to be analyzed. Based on a standard interval, extract the video to be analyzed into multiple frame images, locate the frame image at the current time in the video to be analyzed. If a dangerous marker is detected in the frame image, determine the corresponding environmental risk based on the dangerous marker and generate a disposal plan.

[0055] Dangerous markers include flames, smoke, and accumulated water.

[0056] Specifically, convolutional neural network algorithms such as the Yolo algorithm and the U-net semantic segmentation algorithm can be used to identify dangerous markers in the current frame image. Dangerous markers include flames, smoke, and accumulated water. Each marker has a corresponding preset disposal plan. For example, escape routes are set for fires and smoke, and a corresponding accumulated water cleaning plan is available for accumulated water.

[0057] Environmental sensors include temperature sensors, smoke sensors, and current sensors. Environmental data includes temperature data, smoke data, and current data. In the case where no dangerous marker is detected, if the environmental data exceeds the corresponding set threshold, it is determined that there is an environmental risk and a disposal plan is generated.

[0058] If no dangerous marker is identified in the scene video, but the temperature is detected to be too high by the temperature sensor, such as exceeding 80 °C, it is determined that there is a fire. When the smoke sensor detects that the particle concentration greater than 2.5 μm exceeds 200 μm / ㎡, it is determined that there is smoke. When the current sensor detects that the line current exceeds the specified value, such as 16 A, it is determined that there is an overload risk. Cut off the power supply for the area with current overload, and at the same time arrange professional electricians to conduct line inspections and repairs to eliminate potential safety hazards and avoid serious accidents such as electrical fires caused by current overload.

[0059] In this embodiment, processing the scene video into a multi-segment video includes the following steps:

[0060] The multi-segment video includes a first segment, a second segment, and a third segment. Arrange the frame images in chronological order, extract the changed regions between adjacent frame images based on the inter-frame difference method, and count the number of changed pixel points included in the changed regions.

[0061] Set the standard interval to 1 s. Whenever the recorded scene video reaches 1 min, split the scene video into 60 frame images, arrange the 60 frame images in the recorded chronological order, and then successively compare two adjacent frame images in time based on the inter-frame difference method to obtain the changed regions between the frame images. The inter-frame difference method subtracts the pixel values of the corresponding pixel points in two frame images to obtain the pixel difference. If the pixel difference is greater than a predetermined value, the corresponding pixel point is determined as a changed pixel point, and the number of changed pixel points determined after comparing the two frame images is counted to determine the change quantity.

[0062] If the change quantity is greater than or equal to the dynamic threshold, mark the two frame images generating the change quantity as the first dynamic images. If the change quantity is less than the dynamic threshold, obtain the size of the dynamic threshold. If the dynamic threshold is greater than or equal to the second threshold, mark the two frame images as the second dynamic images. Otherwise, mark the frame images as static images, integrate the first dynamic images into the first section, the second dynamic images into the second section, and the static images into the third section.

[0063] The dynamic threshold is a value that continuously adjusts according to the historical picture change situation. When the value of the dynamic threshold is small, it indicates that the historical picture change situation is small. For example, everyone is working at their office positions and there will be no significant changes in the picture for a long time. When the value of the dynamic threshold is large, it indicates that the historical picture change situation is large. For example, there are a large number of people walking in the picture. The specific determination method of the dynamic threshold will be introduced later.

[0064] If the change quantity is greater than the dynamic threshold, for example, the resolution of the frame image is 1920*1080, the change quantity of the two frame images is determined to be 500000, and the current dynamic threshold is 200000. Since the change quantity is greater than the dynamic threshold, set the label representing the first dynamic image in both of the compared frame images, indicating that a large scene change has occurred in the two frame images. For example, there is no person A in the first frame image, and person A walks into the second frame image. If it is less than the dynamic threshold, continue to obtain the size of the dynamic threshold at this time. For the case with a resolution of 1920*1080, the second threshold is set to 400000. When the dynamic threshold is less than the second threshold, it is determined that the current dynamic threshold is small, which also indicates that the change situation between the two frame images is small. Then set the label representing the static image in both of the compared frame images. On the contrary, when the dynamic threshold is greater than or equal to the second threshold, it is determined that the currently determined dynamic threshold is large, indicating that even if there are people walking in the image, it may be determined to be less than the dynamic threshold, that is, there may still be people walking in the second dynamic image. At this time, set the label representing the second dynamic image in both of the compared frame images.

[0065] Finally, take the frame images with the label of the first dynamic image as the first section, the frame images with the label of the second dynamic image as the second section, and the frame images with the label of the static image as the third section.

[0066] In this embodiment, determining the dynamic threshold includes the following steps:

[0067] Obtain a historical video, process the historical video into a dynamic time series regarding the number of changes, take the preset window size as a parameter, process the dynamic time series based on the sliding window method to obtain historical window values of multiple time windows, split the recording time of the historical video into multiple first time periods, and statistically analyze the historical window values to determine the adjustment coefficient for each first time period.

[0068] Specifically, the historical video includes the scene videos recorded within the past month. According to the above method, the scene videos within a month are split into multiple frame images, and then adjacent two frame images are compared to obtain the number of changes between the two frame images. Arranging the number of changes in chronological order can obtain the dynamic time series corresponding to the historical video.

[0069] For a simple introduction here, assume the dynamic time series is (11, 15, 30, 14, 16), and the window size is set to 3. Then the first historical window value is (11 + 15 + 30) / 3 = 18.7, the second historical window value is (15 + 30 + 14) / 3 = 19.7, and so on until the window slides to the end of the dynamic time series.

[0070] In this embodiment, the recording duration is split into first time periods with a duration of one hour and split by week. For example, the first first time period is from 00:00 to 01:00 on Monday, and the second first time period is from 01:00 to 02:00 on Monday. Then, the change situation of the historical window values within each first time period is statistically analyzed. In actual processing, when the window size is set to 10, there are 3591 historical window values within one first time period. This embodiment uses the following method to determine the adjustment coefficient. First, set multiple adjustment coefficient values, such as 0.1, 0.2, 0.3, etc. Each adjustment coefficient corresponds to a critical threshold. For example, the adjustment coefficients 0.1, 0.2, 0.3 correspond to the critical thresholds 1, critical threshold 2, and critical threshold 3 respectively, and a proportion standard is also set, such as 80%.

[0071] For example, after statistics, there are a total of 3591 * 4 = 14364 historical window values in the four 00:00 - 01:00 on Monday within a month. If more than 98% of the historical window values are less than the critical threshold 1, then the adjustment coefficient is set to 0.1. If more than 98% of the historical window values are between the critical threshold 1 and the critical threshold 2, then the adjustment coefficient is set to 0.2, and so on.

[0072] Extract the average value of the number of changes within the second duration before two frame images, determine the corresponding adjustment coefficient based on the first time period where the two frame images are located, correct the average value based on the adjustment coefficient to obtain a first value, set a basic value, and correct the basic value based on the first value to obtain a dynamic threshold.

[0073] For example, through statistics, it is determined that the historical window values between 12:00 and 13:00 from Monday to Friday are relatively large, and the corresponding adjustment coefficient is set to 0.9. Here, the second duration is set to 20s. For example, if the time points corresponding to two frame images are 12:00:21 and 12:00:22 on Monday respectively, then 20 frame images between 12:00:00 and 12:00:20 on Monday are obtained. Then, the number of changes in the 20 frame images is calculated. Finally, 19 values can be obtained, and the average value of the 19 values is calculated. For example, it is 1000000, and the base value can be set to 50000. By setting the base value, it can be avoided that when the picture changes slightly, the corresponding frame images are classified as the first dynamic images. Then, the dynamic threshold corresponding to the two frame images at this time is 50000 + 1000000 * 0.9 = 950000.

[0074] The analysis of the multi-segment video in this embodiment includes the following steps:

[0075] Define the first segment and the second segment in the multi-segment video as the first analysis segment and the second analysis segment respectively. Based on the object detection algorithm, perform object tracking on the first analysis segment. If personnel movement is detected, calculate the risk degree based on the occurrence frequency, activity trajectory entropy value, and the number of face matching failures.

[0076] Based on the previous introduction, the first analysis segment is the segment where the picture suddenly changes in a static environment. Therefore, when using the object detection algorithm to perform object tracking on the first analysis segment, a better tracking effect can be obtained. The object tracking algorithm can use the similarity matching algorithm of the Siamese network, the yolo algorithm, the correlation filtering algorithm, etc. During the process of tracking the object, continuously perform face recognition on the object to obtain the number of successful face matches and the number of face matching failures, and at the same time record and generate the movement trajectory of the object. This embodiment obtains the occurrence frequency and activity trajectory entropy value based on the following method. First, specify multiple fixed position points in the digital twin model. Whenever the tracked object passes through a fixed position point, the occurrence frequency is accumulated once.

[0077] When calculating the activity trajectory entropy value, first represent the object movement trajectory with multiple coordinate points, and then calculate the probability of the th coordinate point through the first formula , and the first formula is: , where is the number of coordinate points in the movement trajectory, represents the number of times the movement trajectory passes through the th coordinate point. Then, calculate the activity trajectory entropy value through the second formula, and the second formula is: , The calculation result of the second formula shows that the higher the entropy value, the more irregular the movement of the target. Generally, the behavior of office workers has certain patterns.

[0078] Finally, by performing a weighted sum of the occurrence frequency, the entropy value of the activity trajectory, and the number of face matching failures, the risk level is calculated. The specific weights can be determined according to actual experience.

[0079] Obtain the second time period that appears in the second analysis section, intercept the video images within the second time period from the scene videos captured by other surveillance cameras, define it as the auxiliary section, split the second analysis section and the auxiliary section into frame images, and combine the frame images with the same shooting time into an image group. Perform mapping analysis on the image group to obtain the coordinate positions of the same target in different frame images.

[0080] The second analysis section is a section where the picture has been changing significantly, and there may be situations where people need to move around. At this time, using the target tracking algorithm may obtain poor results due to the occlusion caused by the back-and-forth movement of the crowd. Therefore, the present invention also proposes the following method. First, in an office building, to achieve full coverage of the same area, multiple surveillance cameras need to be deployed, and there are overlapping surveillance areas between the cameras to achieve non-blind spot surveillance. For example, they are deployed at the four corners of the ceiling. Based on this hardware, assume that there are three surveillance cameras in the target area. The second analysis section is from the scene video A captured by surveillance camera 1. Determine the second time period that appears in the second analysis section. For example, 12:00:00 - 12:00:10, then obtain the video sections within 12:00:00 - 12:00:10 from the scene video B captured by surveillance camera 2 and the scene video C captured by surveillance camera 3, and define them as auxiliary section 1 and auxiliary section 2 respectively.

[0081] Split the second analysis section, auxiliary section 1, and auxiliary section 2 into frame images at 1s intervals. Each frame image has a label of the shooting time. Then, combine the frame images with the same shooting time into an image group. Each image group contains three frame images, which are from the shooting perspectives of surveillance cameras 1, 2, and 3 respectively. Then perform mutual mapping analysis on the three frame images in the image group to determine the positions of the same person in different frame images. The similarity comparison algorithm or the contour similarity comparison algorithm as above can be used. For the human body, the coordinate position can be selected as the head, the endpoints of the limbs, or other positions, etc. The coordinate position can be one or multiple. The positions of the head and limbs of the human body in the image can be determined using the bone analysis algorithm. The bone analysis algorithm is an existing technology and will not be introduced here. Through the above solution, the number of people in the target area can be obtained, avoiding double counting of a person who appears in multiple cameras.

[0082] To simplify the determination of coordinate positions, in this embodiment, each frame image is divided into a 10*10 grid, and then the coordinate position is determined according to the grid where the user's head is located. For example, for employee A, whose head is located in the grid area of the third row and third column in the frame image, then his coordinate position is (3,3). For employee A, who appears in all three frame images in the image group, his position coordinates are (3,3), (5,6), and (6,4).

[0083] The environmental sensor includes a beacon sensor. The beacon sensor obtains the return signal strength based on the employee's work card. Taking time as a reference, the return signal strengths are combined into a signal group corresponding to the image group, and an analysis model is established. The analysis model outputs the risk degree of a suspicious person based on the coordinate positions of the image group and the return signal strengths in the corresponding signal group.

[0084] In this embodiment, the environmental sensor further includes a beacon sensor, and a beacon tag is embedded in the employee card. The beacon sensor can send a signal to the beacon tag and can obtain the returned signal strength. In other embodiments, the environmental sensor can also be a Bluetooth sensor or a wireless router. The Bluetooth sensor and the wireless router obtain the return signal strength based on the employee's mobile terminal. The above devices are all devices that employees carry with them, so there will be a relatively high detection accuracy.

[0085] Based on the above introduction, through 10 image groups in the stage from 12:00:00 to 12:00:10, the coordinate changes of the people therein can be obtained, thereby constructing a position matrix. The position matrix represents the movement of people in the target area within these 10s. For example, in the position matrix, the first row and first column represent the number of people in the target area in the 1st second, and the second row and first column represent the abscissa of the coordinate position of the people in the first frame image in the image group. The other elements in the matrix are the same by analogy and will not be introduced one by one. A signal matrix is constructed based on 10 signal groups. The first row and first column therein represent the sequence of return signal strengths received by the beacon sensor from each beacon in the 1st second. The above construction method for inputting into the analysis model is not unique and can be determined according to the actual situation. Then, the position matrix and the signal matrix are input into the analysis model, and the analysis model outputs the risk degree of a suspicious person.

[0086] The core principle of the above method is that, for example, within 10s, through video analysis, it is detected that 8 people are walking in the picture, but it is determined based on the signal strength that there are only 7 people walking on the scene, which indicates that there is a suspicious person.

[0087] This embodiment establishes an analysis model including the following steps:

[0088] Collect historical data, where the historical data includes an image group, a signal group, and output labels. The output labels include risk-free and risky. Numerically adjust the coordinate positions in the image group and the return signal intensities in the signal group to augment the historical data and obtain augmented data. Establish an analysis model based on a neural network and use the augmented data to train the analysis model. When the output result of the analysis model is risky, use the probability value of being risky as the risk level.

[0089] In this embodiment, an analysis model is established based on an LSTM time-series neural network. In other embodiments, a fully connected neural network can also be used to establish the analysis model. As a type of supervised learning, the time-series neural network needs to be pre-trained with data. Therefore, historical data needs to be collected. The historical data includes input data and output data. The input data is in the same form as the actual input data described above, that is, a piece of historical data includes a position matrix and a signal matrix generated from an image group and a signal group. The output data includes risky and risk-free, and risky and risk-free are labeled in it through manual annotation. In the case of reduced data volume, the goal of augmenting data can be achieved by slightly adjusting the values of the elements in the position matrix or the image matrix, such as increasing the coordinate values or appropriately reducing the return signal intensity. Finally, by adding a Softmax function to the neural network, the probability of being risky or risk-free is output, thereby obtaining the corresponding risk level.

[0090] In this embodiment, the fire response plan includes the following steps:

[0091] The response plan includes multiple predefined escape routes. If there are multiple escape routes, determine the fire location, simulate the fire spread direction based on the fire location, and select the best route from the predefined evacuation routes based on the spread direction.

[0092] Specifically, first, real-time fire source location data is collected through smoke sensors, temperature detectors, and video monitoring systems, and the coordinates of the ignition point are accurately located in combination with a digital twin model. At the same time, building structure parameters such as the ventilation system status and material combustion characteristics are obtained. Then, using a fire smoke diffusion model based on a Gaussian distribution, predict the spread paths of smoke, high temperature, and toxic gases, generate a thermal distribution map, and highlight key areas such as stairwells and corridors that may be covered by the fire. Next, perform topological analysis on the pre-stored multiple evacuation routes, eliminate the routes passing through high-temperature areas or downwind areas through a spatial overlay algorithm, and preferentially select safe passages in the upwind direction and away from thick smoke. Finally, use the Dijkstra algorithm to comprehensively evaluate the path length, coverage of dangerous areas, and travel time, push the optimal route through a mobile terminal, and dynamically update the data every 30 seconds. When it is detected that the original route is blocked by the fire, automatically activate the alternate route. The above fire smoke diffusion model is a prior art and will not be introduced here.

[0093] As shown Figure 2 The present invention also provides a three-dimensional scene construction system for police cloud Internet of Things devices based on digital twins, which is used to implement the above-mentioned three-dimensional scene construction method for police cloud Internet of Things devices based on digital twins. The system includes:

[0094] Sensing devices, which include surveillance cameras and environmental sensors. The surveillance cameras are used to record scene videos, and the environmental sensors are used to detect environmental data;

[0095] A cloud server, which is used to establish a three-dimensional entity model of the target area, deploy corresponding virtual models in the entity model based on the installation information of the sensing devices, access the data from the sensing devices into the corresponding virtual models, and generate a digital twin model;

[0096] A first troubleshooting unit, which is used to analyze the scene video and environmental data in real time. If an environmental risk is detected, a disposal plan is generated and displayed in the digital twin model;

[0097] A second troubleshooting unit, which is used to process the scene video into multi-segment videos, analyze the multi-segment videos in combination with environmental data, determine the suspicious persons and risk levels in the target area, screen out the suspicious persons with a risk level greater than the first threshold as high-risk risks, and generate suspicious risk alerts pointing to the high-risk risks in the digital twin model.

[0098] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as these combinations of technical features do not conflict, they should be considered as the scope recorded in this specification.

[0099] The above embodiments only represent several implementation manners of the present invention, and their descriptions are relatively specific and detailed, but they should not be construed as limiting the scope of the invention patent of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the invention patent of the present invention should be subject to the appended claims.

[0100] The above is only the preferred embodiment of the present invention, and it is not used to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A three-dimensional scene construction method for police cloud IoT devices based on digital twins, characterized in that, Including: Deploy sensing devices within the target area. The sensing devices include surveillance cameras and environmental sensors. The surveillance cameras are used to record scene videos, and the environmental sensors are used to detect environmental data; Establish a three-dimensional entity model of the target area in the cloud server, deploy corresponding virtual models in the entity model based on the installation information of the sensing devices, and connect the data from the sensing devices to the corresponding virtual models to generate a digital twin model; The first troubleshooting unit analyzes the scene video and the environmental data in real time. If an environmental risk is detected, a disposal plan is generated and displayed in the digital twin model; The second troubleshooting unit processes the scene video into multi-segment videos, and analyzes the multi-segment videos in combination with the environmental data to determine the suspicious persons and risk levels in the target area; Select the suspicious persons with a risk level greater than the first threshold as high-risk risks, and generate suspicious risk alerts pointing to the high-risk risks in the digital twin model; Processing the scene video into the multi-segment videos includes the following steps: Define the scene video to be analyzed as the video to be analyzed. Based on a standard interval, extract the video to be analyzed into multiple frame images. The multi-segment videos include a first segment, a second segment, and a third segment. Arrange the frame images in chronological order, extract the changed areas between adjacent frame images based on the frame difference method, and count the number of changed pixels in the changed areas; If the number of changes is greater than or equal to the dynamic threshold, mark the two frame images that generate the number of changes as the first dynamic images. If the number of changes is less than the dynamic threshold, obtain the magnitude of the dynamic threshold. If the dynamic threshold is greater than or equal to the second threshold, mark the two frame images as the second dynamic images. Otherwise, mark the frame images as static images. Integrate the first dynamic images into the first segment, integrate the second dynamic images into the second segment, and integrate the static images into the third segment; Determining the dynamic threshold includes the following steps: Obtain historical videos, process the historical videos into dynamic time series regarding the number of changes, use the preset window size as a parameter, process the dynamic time series based on the sliding window method to obtain historical window values of multiple time windows, split the recording time of the historical videos into multiple first time periods, and statistically analyze the historical window values to determine the adjustment coefficients for each first time period; Extract the average value of the number of changes within the second time period before the two frame images, determine the corresponding adjustment coefficient based on the first time period where the two frame images are located, correct the average value based on the adjustment coefficient to obtain a first value, set a base value, and correct the base value based on the first value to obtain the dynamic threshold; Analyzing the multi-segment videos includes the following steps: The first segment and the second segment in the multi-segment video are defined as a first analysis segment and a second analysis segment, respectively, and target tracking is performed on the first analysis segment based on a target detection algorithm. If a person walking is detected, the risk level is calculated based on the occurrence frequency, the activity trajectory entropy value, and the number of face matching failures; Obtain a second time period in which the second analysis section appears, intercept a video image of the scene video shot by other surveillance cameras in the second time period, define it as an auxiliary section, split the second analysis section and the auxiliary section into the frame images, and combine the frame images with the same shooting time into an image group, perform mapping analysis on the image group, and obtain the coordinate position of the same target in different frame images; The environmental sensor includes a beacon sensor, which obtains the return signal strength based on the employee's work card, combines the return signal strength into a signal group corresponding to the image group with time as a reference, and establishes an analysis model. The analysis model is based on the coordinate position of the image group and the return signal strength output corresponding to the signal group to output the risk of a suspicious person.

2. The method according to claim 1, characterized in that Real-time analysis of the scene video and the environmental data includes the following steps: Locating the frame image at the current time in the video to be analyzed, if a dangerous marker is detected in the frame image, determining the corresponding environmental risk according to the dangerous marker, and generating the disposal plan; The environmental sensors include temperature sensors, smoke sensors and current sensors, and the environmental data include temperature data, smoke data and current data. If the environmental data exceeds the corresponding set threshold when the hazardous markers are not detected, it is determined that the environmental risk exists and the disposal plan is generated.

3. The method according to claim 1, wherein Establishing the analysis model includes the following steps: Collect historical data, the historical data includes the image group, the signal group and output labels, the output labels include no risk and risky, make numerical adjustments to the coordinate position in the image group and the return signal strength in the signal group to expand the historical data to obtain expanded data, establish the analysis model based on a neural network, and use the expanded data to train the analysis model; when the output result of the analysis model is that there is risk, use the probability value of the risk as the risk degree.

4. The method according to claim 1, wherein Generating the treatment plan includes the following steps: The disposal plan includes a plurality of predetermined escape routes. If there are a plurality of such escape routes, the fire location is determined, a fire spreading direction is simulated and generated based on the fire location, and the best route is selected from the escape routes based on the spreading direction.

5. The method according to claim 2, wherein The hazard signs include flames, smoke and stagnant water.

6. A three-dimensional scene construction system for police cloud IoT devices based on digital twins, which is used to implement the method described in any one of claims 1-5, characterized in that, include, A sensing device, wherein the sensing device includes a monitoring camera and an environmental sensor, wherein the monitoring camera is used to record scene video, and the environmental sensor is used to detect environmental data; A cloud server, which is used to establish a three-dimensional entity model of the target area, deploy corresponding virtual models in the entity model based on the installation information of the sensing devices, access the data from the sensing devices into the corresponding virtual models, and generate a digital twin model; A first troubleshooting unit, which is used to analyze the scene video and the environmental data in real time. If an environmental risk is detected, a disposal plan is generated and displayed in the digital twin model; A second troubleshooting unit, which is used to process the scene video into multi-segment videos, analyze the multi-segment videos in combination with the environmental data, determine the suspicious persons and the risk levels in the target area, screen the suspicious persons with risk levels greater than a first threshold as high-risk risks, and generate suspicious risk alarms pointing to the high-risk risks in the digital twin model.

Citation Information

Patent Citations

  • Fire evacuation planning method and system based on digital twinborn technology

    CN116187105A

  • Three-dimensional video fusion method and system based on digital twin technology

    CN118864723A

  • Digital twin monitoring system

    CN218273143U