Iot real-time monitoring engineering design method and system

By using IoT and neural network models to perform feature fusion analysis on monitoring videos of construction areas, the problem of insufficient accuracy and robustness of construction monitoring technology in complex scenarios has been solved, achieving higher safety and compliance monitoring.

CN118555370BActive Publication Date: 2025-11-11GUANGZHOU JINJIANG DECORATION ENG CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202410707994.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-03
Publication Date
2025-11-11
Estimated Expiration
2044-06-03

AI Technical Summary

Technical Problem

Existing construction monitoring technologies lack accuracy and robustness when processing video data in complex scenarios, failing to meet the high safety and compliance requirements of the construction industry.

Method used

By acquiring monitoring videos of overlapping areas collected by the first and second monitoring devices at the same time through the Internet of Things, feature fusion analysis is performed using a neural network model to generate monitoring results to indicate construction behavior.

Benefits of technology

This improves the accuracy and reliability of monitoring results, enabling it to better meet the high safety and compliance requirements of the construction industry.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118555370B_ABST
    Figure CN118555370B_ABST
Patent Text Reader

Abstract

This application provides an IoT real-time monitoring engineering design method and system, belonging to the field of data processing technology, to meet the potentially higher safety and compliance requirements that the construction industry may face in the future. The method includes: electronic devices acquiring, via the Internet of Things, a first monitoring video of a target monitored area in a first construction area collected by a first monitoring device, and a second monitoring video of the target monitored area in a second construction area collected by a second monitoring device. The first and second monitoring videos are from the same time period and have the same duration, and the target monitored area is the overlapping area of ​​the first and second construction areas. The electronic devices perform feature fusion analysis on the images of the same time frame in the first and second monitoring videos using a neural network model to obtain monitoring results. These monitoring results are used to indicate the construction behavior of the monitored object in the target monitored area.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing technology, and in particular to an engineering design method and system for real-time monitoring of the Internet of Things. Background Technology

[0002] With the rapid development of technology, the construction field is constantly innovating to improve construction quality, efficiency, and safety. In this process, the technology of using neural network models to monitor and analyze construction activities has gradually attracted widespread attention.

[0003] Currently, monitoring technology in the construction field mainly relies on camera deployment. By installing cameras at key locations within the construction area, real-time image data of the construction site can be acquired. However, traditional image recognition technologies still fall short of the accuracy and robustness required in the construction field when processing large amounts of complex video data. This necessitates a more efficient and accurate monitoring method to ensure construction safety and quality. Neural network models, as computational models based on artificial neurons, possess powerful learning and adaptive capabilities. In recent years, deep learning technology has made significant progress in the field of image recognition, especially in the development of convolutional neural networks (CNNs). These technologies can automatically extract features when processing large amounts of data, thereby achieving the recognition and classification of image content.

[0004] In the construction industry, real-time monitoring of construction activities can be achieved by setting up cameras in different construction areas and using neural network models to analyze the video streams captured by the cameras. First, the neural network model needs to learn and recognize various construction behaviors from a large amount of training data, such as the use of construction tools, the operation of construction workers, and the stacking of construction materials. Through learning from this data, the neural network model can develop an understanding and recognition ability of construction behaviors. When the cameras capture video streams from the construction site, the neural network model processes the video data and extracts key construction information. By comparing the behavioral patterns learned by the model, the neural network can analyze the construction behavior of objects within the monitored area and issue timely alerts when abnormal or non-compliant behaviors are detected. This allows construction managers to quickly take measures to ensure the safety and quality of the construction site.

[0005] However, the construction industry may place higher demands on safety and compliance in the future, and the accuracy and robustness of current neural network analysis do not yet meet these requirements. Summary of the Invention

[0006] This application provides an IoT real-time monitoring engineering design method and system to meet the higher safety and compliance requirements that the construction field may face in the future.

[0007] To achieve the above objectives, this application adopts the following technical solution:

[0008] Firstly, an IoT real-time monitoring engineering design method is provided, applied to electronic devices. The method includes: the electronic device acquiring a first monitoring video of a target monitored area in a first construction area collected by a first monitoring device, and a second monitoring video of a target monitored area in a second construction area collected by a second monitoring device, wherein the first monitoring video and the second monitoring video are monitoring videos of the same time and duration, and the target monitored area is the overlapping area of ​​the first construction area and the second construction area; the electronic device performing feature fusion analysis on the images of the same time frame in the first monitoring video and the second monitoring video through a neural network model to obtain monitoring results, and the monitoring results are used to indicate the construction behavior of the monitored object in the target monitored area.

[0009] Optionally, a first D2D side-by-side connection is established between the electronic device and the first monitoring device, and a second D2D side-by-side connection is also established between the electronic device and the second monitoring device. The electronic device acquires the first monitoring video of the target monitored area in the first construction area collected by the first monitoring device and the second monitoring video of the target monitored area in the second construction area collected by the second monitoring device through the Internet of Things. This includes: the electronic device receiving the first monitoring video of the target monitored area in the first construction area collected by the first monitoring device through the first side-by-side connection and receiving the second monitoring video of the target monitored area in the second construction area collected by the second monitoring device through the second side-by-side connection.

[0010] Optionally, the first monitoring video is a monitoring video collected by the first monitoring device monitoring the target area in a first shooting direction, and the second monitoring video is a monitoring video collected by the second monitoring device monitoring the target area in a second shooting direction, wherein the angle between the first shooting direction and the second shooting direction is greater than 0° and less than 90°.

[0011] Optionally, the electronic device performs feature fusion analysis on the images of the same time frame in the first monitoring video and the second monitoring video through a neural network model to obtain monitoring results. This includes: the electronic device performs feature fusion analysis on the images of the same time frame in the first monitoring video and the second monitoring video through a neural network model according to the safety level of the first construction area and the second construction area, with a processing ratio corresponding to the safety level, to obtain monitoring results.

[0012] Optionally, if the safety level of the first construction area is higher than that of the second construction area, the electronic equipment, based on the safety levels of the first and second construction areas and with a processing ratio corresponding to the safety levels, fuses and analyzes the features of the images of the same time frame in the first and second monitoring videos using a neural network model to obtain monitoring results. This includes: the electronic equipment dividing the first monitoring video into K segments and the second monitoring video into K segments based on the duration of the first or second monitoring video; where K is an integer greater than 1, and the value of K is positively correlated with the duration. For the i-th first monitoring video segment out of K first monitoring video segments, and the i-th second monitoring video segment out of K second monitoring video segments, where i is an integer from 1 to K, the i-th first monitoring video segment and the i-th second monitoring video segment each contain M frames of images, where M is an integer greater than 1; the electronic equipment, based on the fact that the safety level of the first construction area is greater than that of the second construction area, fuses and analyzes the features of the N frames of first images in the i-th second monitoring video segment and the M frames of second images in the i-th first monitoring video segment through a neural network model to obtain the monitoring result, where N is an integer greater than or equal to 1 and less than M.

[0013] Optionally, the electronic device, based on the fact that the safety level of the first construction area is greater than that of the second construction area, fuses and analyzes the features of N frames of first images in the i-th second monitoring video segment and M frames of second images in the i-th first monitoring video segment through a neural network model to obtain monitoring results. This includes: the electronic device, based on the fact that the safety level of the first construction area is greater than that of the second construction area, divides the M frames of first images in the i-th second monitoring video segment into N equal parts to obtain N sets of first images in the i-th second monitoring video segment; the electronic device extracts one frame of first image from each of the N sets of first images in the i-th second monitoring video segment to obtain N frames of first images in the i-th second monitoring video segment; the electronic device divides the M frames of second images in the i-th first monitoring video segment into N equal parts to obtain N sets of second images in the i-th first monitoring video segment; j is an integer from 1 to N, and the electronic device fuses and analyzes the features of the j-th frame of first images in the N frames of first images and the j-th set of second images in the N sets of second images in the i-th first monitoring video segment through a neural network model to obtain monitoring results.

[0014] Optionally, the electronic device fuses and analyzes the features of the j-th first image in the N frames of first images and the j-th second image set in the N sets of second images in the i-th first monitoring video segment through a neural network model to obtain monitoring results. This includes: the electronic device convolves the j-th first image through the convolutional layer of the neural network model to obtain a feature vector set #j, and convolves each frame of the j-th second image set through the convolutional layer of the neural network model to obtain multiple feature vector sets; the electronic device fuses the feature vector set #j into multiple feature vector sets through the feature processing layer of the neural network model to obtain multiple fused feature vector sets, and processes the fusion through the fully connected layer of the neural network model. Multiple feature vector sets are combined to obtain monitoring results. The fusion of feature vector set #j into each of the multiple feature vector sets means that at least some feature vectors in feature vector set #j are uniformly inserted between the feature vectors in each feature vector set, resulting in a fused feature vector set. For any feature vector set in the multiple feature vector sets, if the frame number / time of the second image corresponding to that feature vector set in the j-th second image set is closer to the frame number / time of the first image in the j-th frame, then feature vector set #j will have more feature vectors fused into that feature vector set. When the frame number / time is the same, all feature vectors from feature vector set #j will be fused into that feature vector set.

[0015] Optionally, if the safety level of the first construction area is equal to the safety level of the second construction area, the electronic equipment, based on the safety levels of the first and second construction areas and using a processing ratio corresponding to the safety levels, fuses and analyzes the features of the images from the same time frame in the first and second monitoring videos through a neural network model to obtain monitoring results. This includes: the electronic equipment dividing the first monitoring video into K segments and the second monitoring video into K segments based on the duration of the first or second monitoring video; where K is an integer greater than 1, and the value of K is positively correlated with the duration. For the i-th first monitoring video segment out of K first monitoring video segments, and the i-th second monitoring video segment out of K second monitoring video segments, where i is an integer from 1 to K, the i-th first monitoring video segment and the i-th second monitoring video segment each contain M frames of images, where M is an integer greater than 1; the electronic equipment, based on the fact that the safety level of the first construction area is equal to the safety level of the second construction area, uses a neural network model to fuse and analyze the features of the M first frames of the i-th second monitoring video segment and the M second frames of the i-th first monitoring video segment in a one-to-one correspondence, to obtain the monitoring result, where N is an integer greater than or equal to 1 and less than M.

[0016] Optionally, the electronic device, based on the fact that the safety level of the first construction area is equal to the safety level of the second construction area, performs feature fusion analysis on the M frames of the first image in the i-th second monitoring video segment and the M frames of the second image in the i-th first monitoring video segment through a neural network model to obtain monitoring results. This includes: the electronic device convolves the j-th frame of the first image in the M frames of the first image through the convolutional layer of the neural network model to obtain a first feature vector set #j, and convolves the j-th frame of the second image in the M frames of the second image through the convolutional layer of the neural network model to obtain a second feature vector set #j, where j is an integer traversing from 1 to M; the electronic device fuses a portion of the feature vectors in the first feature vector set #j through the feature processing layer of the neural network model. The first feature vector set #j is either merged into the second feature vector set #j, or a portion of the feature vectors in the second feature vector set #j are merged into the first feature vector set #j to obtain a fused feature vector set #j. This fused feature vector set #j is then processed by a fully connected layer of a neural network model to obtain the monitoring results. Specifically, merging a portion of the feature vectors in the first feature vector set #j into the second feature vector set #j means uniformly inserting a portion of the feature vectors in the first feature vector set #j into the various feature vectors in the second feature vector set #j. Conversely, merging a portion of the feature vectors in the second feature vector set #j into the first feature vector set #j means uniformly inserting a portion of the feature vectors in the second feature vector set #j into the various feature vectors in the first feature vector set #j.

[0017] Secondly, an IoT real-time monitoring engineering design system is provided. The system includes an electronic device configured to: acquire, via the Internet of Things, a first monitoring video of a target monitored area in a first construction area collected by a first monitoring device, and a second monitoring video of a target monitored area in a second construction area collected by a second monitoring device, wherein the target monitored area is the overlapping area of ​​the first and second construction areas; and perform feature fusion analysis on the images of the same time frame in the first and second monitoring videos through a neural network model to obtain monitoring results, which are used to indicate the construction behavior of the monitored object in the target monitored area.

[0018] Optionally, a first D2D side-by-side connection is established between the electronic device and the first monitoring device, and a second D2D side-by-side connection is also established between the electronic device and the second monitoring device. The electronic device acquires the first monitoring video of the target monitored area in the first construction area collected by the first monitoring device and the second monitoring video of the target monitored area in the second construction area collected by the second monitoring device through the Internet of Things. This includes: the electronic device receiving the first monitoring video of the target monitored area in the first construction area collected by the first monitoring device through the first side-by-side connection and receiving the second monitoring video of the target monitored area in the second construction area collected by the second monitoring device through the second side-by-side connection.

[0019] Optionally, the first monitoring video is a monitoring video collected by the first monitoring device monitoring the target area in a first shooting direction, and the second monitoring video is a monitoring video collected by the second monitoring device monitoring the target area in a second shooting direction, wherein the angle between the first shooting direction and the second shooting direction is greater than 0° and less than 90°.

[0020] Optionally, the electronic device performs feature fusion analysis on the images of the same time frame in the first monitoring video and the second monitoring video through a neural network model to obtain monitoring results. This includes: the electronic device performs feature fusion analysis on the images of the same time frame in the first monitoring video and the second monitoring video through a neural network model according to the safety level of the first construction area and the second construction area, with a processing ratio corresponding to the safety level, to obtain monitoring results.

[0021] Optionally, if the safety level of the first construction area is higher than that of the second construction area, the electronic equipment, based on the safety levels of the first and second construction areas and with a processing ratio corresponding to the safety levels, fuses and analyzes the features of the images of the same time frame in the first and second monitoring videos using a neural network model to obtain monitoring results. This includes: the electronic equipment dividing the first monitoring video into K segments and the second monitoring video into K segments based on the duration of the first or second monitoring video; where K is an integer greater than 1, and the value of K is positively correlated with the duration. For the i-th first monitoring video segment out of K first monitoring video segments, and the i-th second monitoring video segment out of K second monitoring video segments, where i is an integer from 1 to K, the i-th first monitoring video segment and the i-th second monitoring video segment each contain M frames of images, where M is an integer greater than 1; the electronic equipment, based on the fact that the safety level of the first construction area is greater than that of the second construction area, fuses and analyzes the features of the N frames of first images in the i-th second monitoring video segment and the M frames of second images in the i-th first monitoring video segment through a neural network model to obtain the monitoring result, where N is an integer greater than or equal to 1 and less than M.

[0022] Optionally, the electronic device, based on the fact that the safety level of the first construction area is greater than that of the second construction area, fuses and analyzes the features of N frames of first images in the i-th second monitoring video segment and M frames of second images in the i-th first monitoring video segment through a neural network model to obtain monitoring results. This includes: the electronic device, based on the fact that the safety level of the first construction area is greater than that of the second construction area, divides the M frames of first images in the i-th second monitoring video segment into N equal parts to obtain N sets of first images in the i-th second monitoring video segment; the electronic device extracts one frame of first image from each of the N sets of first images in the i-th second monitoring video segment to obtain N frames of first images in the i-th second monitoring video segment; the electronic device divides the M frames of second images in the i-th first monitoring video segment into N equal parts to obtain N sets of second images in the i-th first monitoring video segment; j is an integer from 1 to N, and the electronic device fuses and analyzes the features of the j-th frame of first images in the N frames of first images and the j-th set of second images in the N sets of second images in the i-th first monitoring video segment through a neural network model to obtain monitoring results.

[0023] Optionally, the electronic device fuses and analyzes the features of the j-th first image in the N frames of first images and the j-th second image set in the N sets of second images in the i-th first monitoring video segment through a neural network model to obtain monitoring results. This includes: the electronic device convolves the j-th first image through the convolutional layer of the neural network model to obtain a feature vector set #j, and convolves each frame of the j-th second image set through the convolutional layer of the neural network model to obtain multiple feature vector sets; the electronic device fuses the feature vector set #j into multiple feature vector sets through the feature processing layer of the neural network model to obtain multiple fused feature vector sets, and processes the fusion through the fully connected layer of the neural network model. Multiple feature vector sets are combined to obtain monitoring results. The fusion of feature vector set #j into each of the multiple feature vector sets means that at least some feature vectors in feature vector set #j are uniformly inserted between the feature vectors in each feature vector set, resulting in a fused feature vector set. For any feature vector set in the multiple feature vector sets, if the frame number / time of the second image corresponding to that feature vector set in the j-th second image set is closer to the frame number / time of the first image in the j-th frame, then feature vector set #j will have more feature vectors fused into that feature vector set. When the frame number / time is the same, all feature vectors from feature vector set #j will be fused into that feature vector set.

[0024] Optionally, if the safety level of the first construction area is equal to the safety level of the second construction area, the electronic equipment, based on the safety levels of the first and second construction areas and using a processing ratio corresponding to the safety levels, fuses and analyzes the features of the images from the same time frame in the first and second monitoring videos through a neural network model to obtain monitoring results. This includes: the electronic equipment dividing the first monitoring video into K segments and the second monitoring video into K segments based on the duration of the first or second monitoring video; where K is an integer greater than 1, and the value of K is positively correlated with the duration. For the i-th first monitoring video segment out of K first monitoring video segments, and the i-th second monitoring video segment out of K second monitoring video segments, where i is an integer from 1 to K, the i-th first monitoring video segment and the i-th second monitoring video segment each contain M frames of images, where M is an integer greater than 1; the electronic equipment, based on the fact that the safety level of the first construction area is equal to the safety level of the second construction area, uses a neural network model to fuse and analyze the features of the M first frames of the i-th second monitoring video segment and the M second frames of the i-th first monitoring video segment in a one-to-one correspondence, to obtain the monitoring result, where N is an integer greater than or equal to 1 and less than M.

[0025] Optionally, the electronic device, based on the fact that the safety level of the first construction area is equal to the safety level of the second construction area, performs feature fusion analysis on the M frames of the first image in the i-th second monitoring video segment and the M frames of the second image in the i-th first monitoring video segment through a neural network model to obtain monitoring results. This includes: the electronic device convolves the j-th frame of the first image in the M frames of the first image through the convolutional layer of the neural network model to obtain a first feature vector set #j, and convolves the j-th frame of the second image in the M frames of the second image through the convolutional layer of the neural network model to obtain a second feature vector set #j, where j is an integer traversing from 1 to M; the electronic device fuses a portion of the feature vectors in the first feature vector set #j through the feature processing layer of the neural network model. The first feature vector set #j is either merged into the second feature vector set #j, or a portion of the feature vectors in the second feature vector set #j are merged into the first feature vector set #j to obtain a fused feature vector set #j. This fused feature vector set #j is then processed by a fully connected layer of a neural network model to obtain the monitoring results. Specifically, merging a portion of the feature vectors in the first feature vector set #j into the second feature vector set #j means uniformly inserting a portion of the feature vectors in the first feature vector set #j into the various feature vectors in the second feature vector set #j. Conversely, merging a portion of the feature vectors in the second feature vector set #j into the first feature vector set #j means uniformly inserting a portion of the feature vectors in the second feature vector set #j into the various feature vectors in the first feature vector set #j.

[0026] In summary, the above methods and systems have the following technical effects:

[0027] When acquiring first monitoring videos of the target monitored area in a first construction area collected by a first monitoring device, and second monitoring videos of the same target monitored area in a second construction area collected by a second monitoring device, the electronic equipment can use a neural network model to perform feature fusion analysis on images from the same time frame in the first and second monitoring videos to obtain monitoring results, such as the construction behavior of the monitored object in the target monitored area, specifically whether the behavior is dangerous / safe / compliant. In this case, because the monitoring results are obtained through the fusion analysis of different videos, the accuracy and reliability of the results are higher, meeting the potentially higher safety and compliance requirements that the construction field may impose in the future. Attached Figure Description

[0028] Figure 1 A flowchart illustrating the IoT real-time monitoring engineering design method provided in this application embodiment;

[0029] Figure 2 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0030] This application will present various aspects, embodiments, or features relating to systems that may include multiple devices, components, modules, etc. It should be understood and appreciated that individual systems may include additional devices, components, modules, etc., and / or may not include all the devices, components, modules, etc. discussed in conjunction with the accompanying drawings. Furthermore, combinations of these approaches are also possible.

[0031] Furthermore, in the embodiments of this application, the words "exemplary," "for example," etc., are used to indicate that they are examples, illustrations, or descriptions. Any embodiment or design scheme described as "exemplary" in this application should not be construed as being more preferred or advantageous than other embodiments or design schemes. Specifically, the use of the term "exemplary" is intended to present the concept in a concrete manner.

[0032] In the embodiments of this application, the terms "information," "signal," "message," "channel," and "singaling" may sometimes be used interchangeably. It should be noted that, without emphasizing their distinction, their intended meanings are consistent. Similarly, "of," "corresponding (relevant)," and "corresponding" may sometimes be used interchangeably. It should be noted that, without emphasizing their distinction, their intended meanings are consistent. Furthermore, the " / " mentioned in this application can be used to indicate an "or" relationship.

[0033] The network architecture and business scenarios described in the embodiments of this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided in the embodiments of this application. As those skilled in the art will know, with the evolution of network architecture and the emergence of new business scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.

[0034] The method described in this application embodiment can be executed by an electronic device, which may be a terminal, a chip or chip system that can be disposed on the terminal. The terminal may also be referred to as user equipment (UE), access terminal, subscriber unit, user station, mobile station (MS), mobile station, remote station, remote terminal, mobile device, user terminal, terminal, wireless communication equipment, user agent, or user apparatus. The terminals in the embodiments of this application may be mobile phones, cellular phones, smartphones, tablets, wireless data cards, personal digital assistants (PDAs), wireless modems, handsets, laptop computers, machine-type communication (MTC) terminals, computers with wireless transceiver capabilities, virtual reality (VR) terminals, augmented reality (AR) terminals, wireless terminals in industrial control, wireless terminals in self-driving, wireless terminals in remote medical care, wireless terminals in smart grids, wireless terminals in transportation safety, wireless terminals in smart cities, wireless terminals in smart homes, vehicle-mounted terminals, roadside units (RSUs) with terminal functions, etc. The terminal of this application may also be an on-board module, on-board unit, on-board component, on-board chip or on-board unit that is built into a vehicle as one or more components or units.

[0035] For ease of understanding, the following will combine... Figure 1 This application provides a detailed description of the IoT real-time monitoring engineering design method provided in the embodiments.

[0036] For example, Figure 1 This application provides a flowchart illustrating the IoT real-time monitoring engineering design method. This method can be applied to the interaction of the aforementioned electronic devices.

[0037] like Figure 1 As shown, the process of this IoT real-time monitoring engineering design method is as follows:

[0038] S101, the electronic device acquires the first monitoring video of the target monitored area in the first construction area collected by the first monitoring device, and the second monitoring video of the target monitored area in the second construction area collected by the second monitoring device via the Internet of Things.

[0039] The first construction area can be adjacent to the second construction area but with different construction objects / processes. The first monitoring video and the second monitoring video are monitoring videos taken at the same time and for the same duration, with the target monitored area being the overlapping area of ​​the first and second construction areas. For example, the first monitoring video is captured by a first monitoring device (such as a camera) monitoring the target monitored area from a first shooting direction, and the second monitoring video is captured by a second monitoring device (such as a camera) monitoring the target monitored area from a second shooting direction. The angle between the first and second shooting directions is greater than 0° and less than 90°. This is to ensure that the differences between the same monitored object captured by different cameras are not too large, so as not to affect the stability and robustness of the subsequent fusion analysis results. Additionally, the shooting direction can be the direction in which the camera's image sensor points to the lens.

[0040] The Internet of Things (IoT) refers to the establishment of a first sideline D2D (electronic device to first monitoring device) connection between an electronic device and a first monitoring device, and a second sideline D2D (electronic device to second monitoring device) connection between the electronic device and a second monitoring device. In this way, the electronic device can receive the first monitoring video of the target monitored area in the first construction area collected by the first monitoring device through the first sideline connection, and receive the second monitoring video of the target monitored area in the second construction area collected by the second monitoring device through the second sideline connection.

[0041] S102, the electronic device uses a neural network model to perform feature fusion analysis on the images of the same time frame in the first monitoring video and the second monitoring video to obtain the monitoring results.

[0042] The neural network model can be a convolutional neural network (CNN), specifically a pre-trained model; the specific training process is not limited in this application's embodiments. The monitoring results are used to indicate the construction behavior of the monitored objects in the monitored area, and can be one or more construction behaviors of one or more monitored objects within the time period corresponding to the monitoring video.

[0043] The safety levels of the first and second construction areas can be the same or different. For example, if the construction process in the first construction area is more important or dangerous, then the safety level of the first construction area can be higher than that of the second construction area, and vice versa. Electronic equipment can, based on the safety levels of the first and second construction areas and with a corresponding processing ratio, fuse and analyze the features of images from the same time frame in the first and second monitoring videos using a neural network model to obtain the monitoring results. The following sections will describe different scenarios.

[0044] Scenario 1:

[0045] If the safety level of the first construction area is higher than that of the second construction area, the electronic equipment, based on the safety levels of the first and second construction areas and using a processing ratio corresponding to the safety levels, fuses and analyzes the features of images from the same time frame in the first and second monitoring videos through a neural network model to obtain monitoring results, including:

[0046] The electronic device can divide the first monitoring video into K segments and the second monitoring video into K segments based on the duration of the first or second monitoring video (e.g., even division; for cases where the division is not even, rounding up or down can be used). Here, K is an integer greater than 1, and the value of K is positively correlated with the duration; this can be achieved by establishing a specific correspondence.

[0047] For the i-th first monitoring video segment out of K first monitoring video segments, and the i-th second monitoring video segment out of K second monitoring video segments, where i is an integer from 1 to K, each i-th first monitoring video segment and the i-th second monitoring video segment each contain M frames of images, where M is an integer greater than 1. Based on this, the electronic device can, according to the fact that the safety level of the first construction area is greater than that of the second construction area, fuse and analyze the features of N frames of first images in the i-th second monitoring video segment and M frames of second images in the i-th first monitoring video segment through a neural network model to obtain the monitoring result, where N is an integer greater than or equal to 1 and less than M.

[0048] Specifically, the electronic device can divide the M frames of the first image in the i-th second monitoring video segment into N equal parts, based on the fact that the safety level of the first construction area is greater than that of the second construction area. This results in N sets of first images for the i-th second monitoring video segment. One frame is then extracted from each of these N sets (the extraction location is not limited and can be random or the median can be used), resulting in N frames of first images in the i-th second monitoring video segment. The electronic device further divides the M frames of the second image in the i-th first monitoring video segment into N equal parts. If the division is even, values ​​can be rounded up or down for non-divisible divisions; this is also considered an even division. For example, 110 frames are divided into 3 parts: 36, 36, and 37. j is an integer from 1 to N, and the value of N can be preset without restriction. The electronic device can fuse and analyze the features of the j-th first image in N frames of first images and the j-th second image set in N frames of the i-th first monitoring video segment through a neural network model to obtain the monitoring result. For example, the electronic device convolves the j-th first image through the convolutional layer of the neural network model to obtain the feature vector set #j, and convolves each frame of the j-th second image set through the convolutional layer of the neural network model to obtain multiple feature vector sets; the electronic device then fuses the feature vector set #j into multiple feature vector sets through the feature processing layer of the neural network model to obtain multiple fused feature vector sets, and finally processes the fused multiple feature vector sets through the fully connected layer of the neural network model to obtain the monitoring result.

[0049] In this context, fusing feature vector set #j into each of the multiple feature vector sets means that at least some feature vectors from feature vector set #j are evenly inserted among the feature vectors in each feature vector set, resulting in a fused feature vector set. For example, feature vector set #j has 10 feature vectors, and each feature vector set contains 30 feature vectors. The 10 feature vectors can be evenly inserted among the 30 feature vectors, with one vector inserted between every two feature vectors. In other words, fusion is the fusion of feature vectors, combining feature vectors from different sets into the same set. This allows the same feature collected from the same monitored object at different angles to be processed jointly by the neural network model, improving the accuracy and robustness of the model. For any feature vector set in the multiple feature vector sets, if the frame number / time of the second image corresponding to that feature vector set in the j-th second image set is closer to the frame number / time of the first image in the j-th frame, then feature vector set #j will have more feature vectors fused into that feature vector set. When the frame number / time is the same, all feature vectors from feature vector set #j will be fused into that feature vector set. This improves fusion efficiency; that is, the more time-separated feature vector sets are, the fewer feature vectors can be used for fusion. This is because, on a larger time scale, the correlation between features weakens, thus requiring less feature fusion to reduce overhead. Of course, the specific proportions can be preset, such as 80% of feature vectors used for fusion after 2 frames, 60% after 3 frames, and 40% after 4 frames, etc.

[0050] Scenario 2: If the safety level of the first construction area is equal to the safety level of the second construction area, the electronic equipment, based on the safety levels of the first and second construction areas and using the corresponding processing ratio, fuses and analyzes the features of the images from the same time frame in the first and second monitoring videos through a neural network model to obtain the monitoring results, including:

[0051] The electronic device can divide the first monitoring video into K first monitoring video segments and the second monitoring video into K second monitoring video segments according to the duration of the first monitoring video or the second monitoring video. Here, K is an integer greater than 1, and the value of K is positively correlated with the duration. For the i-th first monitoring video segment and the i-th second monitoring video segment among the K first monitoring video segments, i is an integer from 1 to K. Each of the i-th first and i-th second monitoring video segments contains M frames, where M is an integer greater than 1.

[0052] The electronic device can, based on the fact that the safety level of the first construction area is equal to the safety level of the second construction area, perform feature fusion analysis on the M frames of the first image in the i-th second monitoring video segment and the M frames of the second image in the i-th first monitoring video segment through a neural network model to obtain the monitoring result. N is an integer greater than or equal to 1 and less than M. For example, the electronic device performs convolution on the j-th frame of the first image in the M frames of the first image through the convolutional layer of the neural network model to obtain the first feature vector set #j, and performs convolution on the j-th frame of the second image in the M frames of the second image through the convolutional layer of the neural network model to obtain the second feature vector set #j, where j is an integer ranging from 1 to M. The electronic device then uses the feature processing layer of the neural network model to fuse some feature vectors from the first feature vector set #j into the second feature vector set #j, or fuses some feature vectors from the second feature vector set #j into the first feature vector set #j, to obtain the fused feature vector set #j. Finally, the fused feature vector set #j is processed by the fully connected layer of the neural network model to obtain the monitoring result.

[0053] Specifically, fusing some feature vectors from the first feature vector set #j into the second feature vector set #j means: uniformly inserting some feature vectors from the first feature vector set #j into the various feature vectors in the second feature vector set #j; and fusing some feature vectors from the second feature vector set #j into the first feature vector set #j means: uniformly inserting some feature vectors from the second feature vector set #j into the various feature vectors in the first feature vector set #j.

[0054] The principle behind Case 2 can be found in the relevant introduction to Case 1 above, and will not be repeated here.

[0055] In summary, when acquiring first monitoring videos of the target monitored area in a first construction area collected by a first monitoring device, and second monitoring videos of the same target monitored area in a second construction area collected by a second monitoring device, electronic devices can use neural network models to perform feature fusion analysis on images from the same time frame in the first and second monitoring videos to obtain monitoring results, such as the construction behavior of the monitored object in the target monitored area, specifically whether the behavior is dangerous / safe / compliant. In this case, because the monitoring results are obtained through the fusion analysis of different videos, the accuracy and reliability of the results are higher, which can meet the potentially higher safety and compliance requirements that the construction field may place on it in the future.

[0056] The above combination Figure 1 This application provides a detailed description of the IoT real-time monitoring engineering design method provided in its embodiments. The following details an IoT real-time monitoring engineering design system for implementing the IoT real-time monitoring engineering design method provided in this application, the system including electronic equipment.

[0057] The system is configured such that: electronic devices acquire first monitoring videos of the target monitored area in the first construction area collected by the first monitoring device and second monitoring videos of the target monitored area in the second construction area collected by the second monitoring device via the Internet of Things; the target monitored area is the overlapping area of ​​the first and second construction areas; electronic devices perform feature fusion analysis on the images of the same time frame in the first and second monitoring videos through a neural network model to obtain monitoring results, which are used to indicate the construction behavior of the monitored object in the target monitored area.

[0058] Optionally, a first D2D side-by-side connection is established between the electronic device and the first monitoring device, and a second D2D side-by-side connection is also established between the electronic device and the second monitoring device. The electronic device acquires the first monitoring video of the target monitored area in the first construction area collected by the first monitoring device and the second monitoring video of the target monitored area in the second construction area collected by the second monitoring device through the Internet of Things. This includes: the electronic device receiving the first monitoring video of the target monitored area in the first construction area collected by the first monitoring device through the first side-by-side connection and receiving the second monitoring video of the target monitored area in the second construction area collected by the second monitoring device through the second side-by-side connection.

[0059] Optionally, the first monitoring video is a monitoring video collected by the first monitoring device monitoring the target area in a first shooting direction, and the second monitoring video is a monitoring video collected by the second monitoring device monitoring the target area in a second shooting direction, wherein the angle between the first shooting direction and the second shooting direction is greater than 0° and less than 90°.

[0060] Optionally, the electronic device performs feature fusion analysis on the images of the same time frame in the first monitoring video and the second monitoring video through a neural network model to obtain monitoring results. This includes: the electronic device performs feature fusion analysis on the images of the same time frame in the first monitoring video and the second monitoring video through a neural network model according to the safety level of the first construction area and the second construction area, with a processing ratio corresponding to the safety level, to obtain monitoring results.

[0061] Optionally, if the safety level of the first construction area is higher than that of the second construction area, the electronic equipment, based on the safety levels of the first and second construction areas and with a processing ratio corresponding to the safety levels, fuses and analyzes the features of the images of the same time frame in the first and second monitoring videos using a neural network model to obtain monitoring results. This includes: the electronic equipment dividing the first monitoring video into K segments and the second monitoring video into K segments based on the duration of the first or second monitoring video; where K is an integer greater than 1, and the value of K is positively correlated with the duration. For the i-th first monitoring video segment out of K first monitoring video segments, and the i-th second monitoring video segment out of K second monitoring video segments, where i is an integer from 1 to K, the i-th first monitoring video segment and the i-th second monitoring video segment each contain M frames of images, where M is an integer greater than 1; the electronic equipment, based on the fact that the safety level of the first construction area is greater than that of the second construction area, fuses and analyzes the features of the N frames of first images in the i-th second monitoring video segment and the M frames of second images in the i-th first monitoring video segment through a neural network model to obtain the monitoring result, where N is an integer greater than or equal to 1 and less than M.

[0062] Optionally, the electronic device, based on the fact that the safety level of the first construction area is greater than that of the second construction area, fuses and analyzes the features of N frames of first images in the i-th second monitoring video segment and M frames of second images in the i-th first monitoring video segment through a neural network model to obtain monitoring results. This includes: the electronic device, based on the fact that the safety level of the first construction area is greater than that of the second construction area, divides the M frames of first images in the i-th second monitoring video segment into N equal parts to obtain N sets of first images in the i-th second monitoring video segment; the electronic device extracts one frame of first image from each of the N sets of first images in the i-th second monitoring video segment to obtain N frames of first images in the i-th second monitoring video segment; the electronic device divides the M frames of second images in the i-th first monitoring video segment into N equal parts to obtain N sets of second images in the i-th first monitoring video segment; j is an integer from 1 to N, and the electronic device fuses and analyzes the features of the j-th frame of first images in the N frames of first images and the j-th set of second images in the N sets of second images in the i-th first monitoring video segment through a neural network model to obtain monitoring results.

[0063] Optionally, the electronic device fuses and analyzes the features of the j-th first image in the N frames of first images and the j-th second image set in the N sets of second images in the i-th first monitoring video segment through a neural network model to obtain monitoring results. This includes: the electronic device convolves the j-th first image through the convolutional layer of the neural network model to obtain a feature vector set #j, and convolves each frame of the j-th second image set through the convolutional layer of the neural network model to obtain multiple feature vector sets; the electronic device fuses the feature vector set #j into multiple feature vector sets through the feature processing layer of the neural network model to obtain multiple fused feature vector sets, and processes the fusion through the fully connected layer of the neural network model. Multiple feature vector sets are combined to obtain monitoring results. The fusion of feature vector set #j into each of the multiple feature vector sets means that at least some feature vectors in feature vector set #j are uniformly inserted between the feature vectors in each feature vector set, resulting in a fused feature vector set. For any feature vector set in the multiple feature vector sets, if the frame number / time of the second image corresponding to that feature vector set in the j-th second image set is closer to the frame number / time of the first image in the j-th frame, then feature vector set #j will have more feature vectors fused into that feature vector set. When the frame number / time is the same, all feature vectors from feature vector set #j will be fused into that feature vector set.

[0064] Optionally, if the safety level of the first construction area is equal to the safety level of the second construction area, the electronic equipment, based on the safety levels of the first and second construction areas and using a processing ratio corresponding to the safety levels, fuses and analyzes the features of the images from the same time frame in the first and second monitoring videos through a neural network model to obtain monitoring results. This includes: the electronic equipment dividing the first monitoring video into K segments and the second monitoring video into K segments based on the duration of the first or second monitoring video; where K is an integer greater than 1, and the value of K is positively correlated with the duration. For the i-th first monitoring video segment out of K first monitoring video segments, and the i-th second monitoring video segment out of K second monitoring video segments, where i is an integer from 1 to K, the i-th first monitoring video segment and the i-th second monitoring video segment each contain M frames of images, where M is an integer greater than 1; the electronic equipment, based on the fact that the safety level of the first construction area is equal to the safety level of the second construction area, uses a neural network model to fuse and analyze the features of the M first frames of the i-th second monitoring video segment and the M second frames of the i-th first monitoring video segment in a one-to-one correspondence, to obtain the monitoring result, where N is an integer greater than or equal to 1 and less than M.

[0065] Optionally, the electronic device, based on the fact that the safety level of the first construction area is equal to the safety level of the second construction area, performs feature fusion analysis on the M frames of the first image in the i-th second monitoring video segment and the M frames of the second image in the i-th first monitoring video segment through a neural network model to obtain monitoring results. This includes: the electronic device convolves the j-th frame of the first image in the M frames of the first image through the convolutional layer of the neural network model to obtain a first feature vector set #j, and convolves the j-th frame of the second image in the M frames of the second image through the convolutional layer of the neural network model to obtain a second feature vector set #j, where j is an integer traversing from 1 to M; the electronic device fuses a portion of the feature vectors in the first feature vector set #j through the feature processing layer of the neural network model. The first feature vector set #j is either merged into the second feature vector set #j, or a portion of the feature vectors in the second feature vector set #j are merged into the first feature vector set #j to obtain a fused feature vector set #j. This fused feature vector set #j is then processed by a fully connected layer of a neural network model to obtain the monitoring results. Specifically, merging a portion of the feature vectors in the first feature vector set #j into the second feature vector set #j means uniformly inserting a portion of the feature vectors in the first feature vector set #j into the various feature vectors in the second feature vector set #j. Conversely, merging a portion of the feature vectors in the second feature vector set #j into the first feature vector set #j means uniformly inserting a portion of the feature vectors in the second feature vector set #j into the various feature vectors in the first feature vector set #j.

[0066] Figure 2 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Exemplarily, the electronic device may be a terminal device, or a chip (system) or other component or assembly that can be disposed in the terminal device. Figure 2 As shown, the electronic device 400 may include a processor 401. Optionally, the electronic device 400 may also include a memory 402 and / or a transceiver 403. The processor 401 is coupled to the memory 402 and the transceiver 403, for example, they can be connected via a communication bus. Alternatively, the electronic device 400 may also be a chip, such as including the processor 401; in this case, the transceiver may be the chip's input / output interface.

[0067] The following is combined with Figure 2 The various components of electronic device 400 are described in detail below:

[0068] The processor 401 is the control center of the electronic device 400. It can be a single processor or a collective term for multiple processing elements. For example, the processor 401 can be one or more central processing units (CPUs), application-specific integrated circuits (ASICs), or one or more integrated circuits configured to implement the embodiments of this application, such as one or more digital signal processors (DSPs), or one or more field-programmable gate arrays (FPGAs).

[0069] Optionally, the processor 401 can perform various functions of the electronic device 400 by running or executing software programs stored in the memory 402 and calling data stored in the memory 402, such as performing the aforementioned functions. Figure 2 The method for designing real-time monitoring projects for the Internet of Things is shown.

[0070] In a specific implementation, as one example, processor 401 may include one or more CPUs, for example... Figure 2 CPU0 and CPU1 are shown in the diagram.

[0071] In a specific implementation, as one example, the electronic device 400 may also include multiple processors. Each of these processors may be a single-core processor (single-CPU) or a multi-core processor (multi-CPU). Here, a processor may refer to one or more devices, circuits, and / or processing cores used to process data (e.g., computer programs or instructions).

[0072] The memory 402 is used to store the software program that executes the solution of this application, and is controlled by the processor 401 to execute it. The specific implementation method can be referred to the above method embodiment, and will not be repeated here.

[0073] Optionally, the memory 402 may be a read-only memory (ROM) or other type of static storage device capable of storing static information and instructions, random access memory (RAM) or other type of dynamic storage device capable of storing information and instructions, or electrically erasable programmable read-only memory (EEPROM), compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital universal optical discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but not limited thereto. The memory 402 may be integrated with the processor 401 or may exist independently and be accessible through the interface circuit of the electronic device 400. Figure 2 (Not shown in the image) is coupled to processor 401, and this embodiment of the application does not specifically limit this.

[0074] Transceiver 403 is used for communication with other electronic devices. For example, if electronic device 400 is a terminal device, transceiver 403 can be used to communicate with a network device or with another terminal device. As another example, if electronic device 400 is a network device, transceiver 403 can be used to communicate with a terminal device or with another network device.

[0075] Alternatively, transceiver 403 may include a receiver and a transmitter. Figure 2 (Not shown separately). The receiver is used to implement the receiving function, and the transmitter is used to implement the sending function.

[0076] Alternatively, the transceiver 403 can be integrated with the processor 401, or it can exist independently and be connected via the interface circuit of the electronic device 400. Figure 2 (Not shown in the image) is coupled to processor 401, and this embodiment of the application does not specifically limit this.

[0077] Understandable Figure 2 The structure of the electronic device 400 shown does not constitute a limitation on the electronic device. Actual electronic devices may include more or fewer components than shown, or combine certain components, or have different component arrangements.

[0078] Furthermore, the technical effects of the electronic device 400 can be referred to the technical effects of the methods described in the above method embodiments, and will not be repeated here.

[0079] It should be understood that the processor in the embodiments of this application can be a central processing unit (CPU), or it can be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor.

[0080] It should also be understood that the memory in the embodiments of this application can be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate synchronous DRAM (DDR SDRAM), enhanced synchronous DRAM (ESDRAM), synchronous linked DRAM (SLDRAM), and direct rambus RAM (DR RAM).

[0081] The above embodiments can be implemented, in whole or in part, by software, hardware (such as circuits), firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer program or instructions can be stored in a computer-readable storage medium or transferred from one computer-readable storage medium to another. For example, the computer program or instructions can be transferred from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more sets of available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium. A semiconductor medium can be a solid-state drive.

[0082] It should be understood that the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. A and B can be singular or plural. Additionally, the character " / " in this article generally indicates an "or" relationship between the preceding and following related objects, but it can also represent an "and / or" relationship. Please refer to the context for a more accurate understanding.

[0083] In this application, "at least one" means one or more, and "more than one" means two or more. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or multiple items. For example, at least one of a, b, or c can mean: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or multiple.

[0084] It should be understood that in the various embodiments of this application, the order of the above-mentioned processes does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0085] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0086] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0087] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0088] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0089] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0090] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0091] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A method for designing real-time monitoring engineering for the Internet of Things, characterized in that, Applied to electronic devices, the method includes: The electronic device acquires a first monitoring video of the target monitored area in the first construction area collected by the first monitoring device and a second monitoring video of the target monitored area in the second construction area collected by the second monitoring device through the Internet of Things. The first monitoring video and the second monitoring video are monitoring videos at the same time and with the same duration. The target monitored area is the overlapping area of ​​the first construction area and the second construction area. The electronic device uses a neural network model to perform feature fusion analysis on images of the same time frame in the first monitoring video and the second monitoring video to obtain monitoring results. The monitoring results are used to indicate the construction behavior of the monitored object in the target monitored area. Wherein, if the safety level of the first construction area is greater than the safety level of the second construction area, the electronic device performs feature fusion analysis on the images of the same time frame in the first monitoring video and the second monitoring video through a neural network model to obtain monitoring results, including: The electronic device divides the first monitoring video into K first monitoring video segments and the second monitoring video into K second monitoring video segments according to the duration of the first monitoring video or the second monitoring video. Wherein; K is an integer greater than 1, and the value of K is positively correlated with the duration; for the i-th first monitoring video segment among the K first monitoring video segments, and the i-th second monitoring video segment among the K second monitoring video segments, i is an integer from 1 to K; the i-th first monitoring video segment and the i-th second monitoring video segment each contain M frames of images, where M is an integer greater than 1; The electronic device, based on the fact that the safety level of the first construction area is greater than that of the second construction area, divides the M frames of the first image of the i-th second monitoring video segment into N parts to obtain N sets of first images of the i-th second monitoring video segment. It then extracts one frame of the first image from each of the N sets of the first image of the i-th second monitoring video segment to obtain N frames of the first image in the i-th second monitoring video segment; N is an integer greater than or equal to 1 and less than M. The electronic device divides the M frames of the second image of the i-th first monitoring video segment into N parts to obtain N sets of second images of the i-th first monitoring video segment; j is an integer from 1 to N. The electronic device performs convolution on the first image of the j-th frame through the convolutional layer of the neural network model to obtain a feature vector set #j, and performs convolution on each second image in the j-th second image set through the convolutional layer of the neural network model to obtain multiple feature vector sets. The electronic device uses the feature processing layer of the neural network model to fuse the feature vector set #j into multiple feature vector sets to obtain multiple fused feature vector sets. The fused feature vector sets are then processed by the fully connected layer of the neural network model to obtain the monitoring result. Wherein, the feature vector set #j is fused into each of the plurality of feature vector sets, that is, at least some feature vectors in the feature vector set #j are uniformly inserted between each feature vector in each feature vector set to obtain a fused feature vector set; for any feature vector set in the plurality of feature vector sets, if the frame number / time of the second image corresponding to the feature vector set in the j-th second image set is closer to the frame number / time of the first image in the j-th frame, then the feature vector set #j is fused into the feature vector set more feature vectors; when the frame number / time is the same, all feature vectors of the feature vector set #j are fused into the feature vector set.

2. The method according to claim 1, characterized in that, A first D2D sideline connection is established between the electronic device and the first monitoring device, and a second D2D sideline connection is also established between the electronic device and the second monitoring device. The electronic device acquires, via the Internet of Things, a first monitoring video of the target monitored area in a first construction area collected by the first monitoring device, and a second monitoring video of the target monitored area in a second construction area collected by the second monitoring device, including: The electronic device receives the first monitoring video of the target monitored area in the first construction area collected by the first monitoring device through the first side connection, and receives the second monitoring video of the target monitored area in the second construction area collected by the second monitoring device through the second side connection.

3. The method according to claim 2, characterized in that, The first monitoring video is a monitoring video collected by the first monitoring device monitoring the target monitoring area in a first shooting direction, and the second monitoring video is a monitoring video collected by the second monitoring device monitoring the target monitoring area in a second shooting direction, wherein the angle between the first shooting direction and the second shooting direction is greater than 0° and less than 90°.

4. The method according to claim 1, characterized in that, If the safety level of the first construction area is equal to the safety level of the second construction area, then the electronic device, based on the safety levels of the first and second construction areas and using a processing ratio corresponding to the safety levels, fuses and analyzes the features of images from the same time frame in the first and second monitoring videos through the neural network model to obtain the monitoring result, including: The electronic device divides the first monitoring video into K first monitoring video segments and the second monitoring video into K second monitoring video segments according to the duration of the first monitoring video or the second monitoring video. Wherein; K is an integer greater than 1, and the value of K is positively correlated with the duration; for the i-th first monitoring video segment among the K first monitoring video segments, and the i-th second monitoring video segment among the K second monitoring video segments, i is an integer from 1 to K; the i-th first monitoring video segment and the i-th second monitoring video segment each contain M frames of images, where M is an integer greater than 1; The electronic device, based on the fact that the safety level of the first construction area is equal to the safety level of the second construction area, performs feature fusion analysis by matching the M first images in the i-th second monitoring video segment with the M second images in the i-th first monitoring video segment through the neural network model to obtain the monitoring result, where N is an integer greater than or equal to 1 and less than M.

5. The method according to claim 4, characterized in that, The electronic device, based on the fact that the safety level of the first construction area is equal to the safety level of the second construction area, performs feature fusion analysis through the neural network model to obtain the monitoring results, including: M frames of the first image in the i-th second monitoring video segment and M frames of the second image in the i-th first monitoring video segment, corresponding one-to-one, to obtain the monitoring results. The electronic device performs convolution on the j-th frame of the first image in the M-frame first image through the convolutional layer of the neural network model to obtain a first feature vector set #j, and performs convolution on the j-th frame of the second image in the M-frame second image through the convolutional layer of the neural network model to obtain a second feature vector set #j, where j is an integer traversing from 1 to M; The electronic device uses the feature processing layer of the neural network model to fuse some feature vectors from the first feature vector set #j into the second feature vector set #j, or to fuse some feature vectors from the second feature vector set #j into the first feature vector set #j, to obtain a fused feature vector set #j. The fused feature vector set #j is then processed by the fully connected layer of the neural network model to obtain the monitoring result. Wherein, fusing some feature vectors from the first feature vector set #j into the second feature vector set #j means: uniformly inserting some feature vectors from the first feature vector set #j into each feature vector in the second feature vector set #j, and fusing some feature vectors from the second feature vector set #j into the first feature vector set #j means: uniformly inserting some feature vectors from the second feature vector set #j into each feature vector in the first feature vector set #j.

Citation Information

Patent Citations

  • Driver abnormal behavior detection method and device

    CN113283286A