Intelligent Supervision System and Method for Law Enforcement Behavior Based on Edge Computing

Through edge computing and blockchain technology, an intelligent supervision system for law enforcement behaviors with adaptive denoising, real-time scenario adaptation and hierarchical uploading is built, which solves the problems of noise changes and target tracking of law enforcement systems in complex scenarios, and achieves efficient data sharing and coherence in data analysis.

CN120071224BActive Publication Date: 2025-08-05SHENZHEN DACHENWEI TECH GRP CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510544406.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-28
Publication Date
2025-08-05
Estimated Expiration
2045-04-28

AI Technical Summary

Technical Problem

The existing law enforcement system is difficult to adapt to noise changes in real time in complex and changing law enforcement scenarios, has weak target tracking capabilities, unreasonable data upload methods, and there are problems with data security and sharing.

Method used

An intelligent supervision system for law enforcement behavior based on edge computing is adopted, including adaptive denoising module, scene switching detection module, target tracking module and hierarchical upload module. Combined with blockchain technology, a law enforcement data sharing platform is built to realize adaptive denoising, real-time scenario adaptation, dynamic target tracking and hierarchical upload of multimodal data.

Benefits of technology

It improves the denoising accuracy of law enforcement video data and the stability of target tracking, ensures the security and efficient sharing of data, adapts to the dynamic changes of complex law enforcement scenarios, optimizes data transmission strategies, and improves the coherence and coordination efficiency of law enforcement behavior analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120071224B_ABST
    Figure CN120071224B_ABST
Patent Text Reader

Abstract

An intelligent supervision system and method for law enforcement behavior based on edge computing, which relates to the field of edge computing technology, collects multi-modal data of each law enforcement communication terminal, adaptively denoises the video data of the law enforcement communication terminal; incrementally updates the general noise level of the video data of the law enforcement communication terminal, and re-classifies the scene according to the incremental update result; performs human feature recognition on the first frame image of the video data, performs human feature matching and occlusion frame and visible frame marking on the subsequent frame images of the video data, and constructs a dynamic graph convolutional network according to the marking result, marks the tracking target, models each tracking target as a graph node, constructs the dynamic edge weights between each node, and outputs the spatial features of the tracking target in each occlusion frame; outputs the danger warning level of each tracking target, and uploads the data in a classified manner according to the danger warning level, significantly improving the data recognition accuracy and the security of data management.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of edge computing, and specifically to an intelligent supervision system and method for law enforcement behavior based on edge computing. Background Art

[0002] A Chinese patent with the publication number CN111263094B discloses an artificial intelligence analysis method and system for public security law enforcement videos, including: equipped with a video acquisition device on law enforcement officers, which rotates 360 degrees around the law enforcement officers to collect audio and video data, stores the data in a local non-volatile storage medium, and at the same time transmits it to the law enforcement center in real time through a wireless network. Identify attack weapons in video frames by comparing with an attack weapon image library, and at the same time use a time series analysis model to predict the probability that law enforcement objects attack law enforcement officers.

[0003] A Chinese patent with the publication number CN118588094B discloses a law enforcement recorder audio processing method and a law enforcement recorder based on voice analysis, including: collecting an audio signal through a law enforcement recorder, and using the predicted coding result of the audio signal as the audio signal to be transmitted. Construct a Huffman coding table and supplementary coding according to the audio signal to be transmitted. Divide the audio signal to be transmitted according to each division length in a preset range to obtain each division method and all subsequences included therein. By calculating the preference degree of each division method, select the division method with the largest preference degree as the final division method, and supplement the coding results of all subsequences included in the final division method to make the final coding result length of each subsequence equal to the target coding length. Finally, form the final coding result of the audio signal to be transmitted by combining the final coding results of all subsequences and transmit it. The receiving end verifies the received final coding result according to the target coding length and the preset range to ensure the integrity and accuracy of the transmitted audio.

[0004] Traditional denoising methods usually adopt general filtering algorithms and do not fully consider the diversity of law enforcement scenarios. It is difficult to perform real-time and dynamic incremental updates on the noise level of video data. During law enforcement, the scenario may quickly switch, such as from indoors to outdoors, from day to night, etc., and the noise level also changes accordingly. However, the existing technologies cannot capture these changes in time and cannot adjust the denoising strategy according to the real-time change of the noise level, resulting in the denoising effect being out of touch with the actual needs after the scenario switches, reducing the adaptability of the system to complex and changing law enforcement scenarios. Most existing systems rely on simple feature extraction and classification methods during scenario recognition and cannot make full use of the rich information of multi-modal data. Moreover, during the target tracking process, the existing technologies have weak processing capabilities for occlusion situations. When the target is partially or completely occluded, traditional tracking algorithms are prone to losing the target, which makes it difficult for existing systems to continuously track in actual law enforcement scenarios, such as crowded areas, once the target is occluded, affecting the full-process supervision of law enforcement behavior.

[0005] When the existing system uploads data, it usually adopts a unified upload method without distinguishing the importance and urgency of the data. In terms of data storage and sharing, the existing system mostly adopts the traditional centralized storage method, where the data is vulnerable to attacks and tampering, and there are obstacles to data sharing between different law enforcement departments. Summary of the Invention

[0006] To solve the above technical problems, the purpose of the present invention is to provide an intelligent supervision system for law enforcement behavior based on edge computing, including a law enforcement data sharing platform, which is communicatively linked with a number of law enforcement communication terminals, an adaptive denoising module, a scene switching detection module, a target tracking module, and a hierarchical upload module;

[0007] The law enforcement communication terminals are used to collect multimodal data, and the multimodal data includes video data and sensor data;

[0008] The adaptive denoising module is used to construct a scene-noise intensity mapping model, obtain the law enforcement scene of the law enforcement communication terminal and the general noise level of the video data, and perform adaptive denoising on the video data of the law enforcement communication terminal;

[0009] The scene switching detection module is used to incrementally update the general noise level of the video data of the law enforcement communication terminal, and judge whether to reclassify the scene according to the incremental update result;

[0010] The image feature recognition module is used to perform human feature recognition on the first frame image of the video data, and obtain the spatial distribution features, pose features, appearance features, and identity identifiers of the target human body;

[0011] The target tracking module is used to perform human feature matching on the subsequent frame images of the video data, mark the occluded frames and visible frames, and construct a dynamic graph convolutional network according to the marking results, mark the tracking targets, model each tracking target as a graph node, construct the dynamic edge weights between each node, and output the spatial features of the tracking targets in each occluded frame;

[0012] The hierarchical upload module is used to construct an action label recognition model, output the danger warning levels of each tracking target, and perform data hierarchical upload according to the danger warning levels.

[0013] Furthermore, the law enforcement data sharing platform is constructed based on blockchain technology. A number of blockchain nodes are set in the law enforcement data sharing platform. The blockchain nodes are communicatively linked with the law enforcement communication terminals within a preset range and are used to receive the law enforcement data packets uploaded by the law enforcement communication terminals. Each blockchain node is linked to each other to form a blockchain network.

[0014] Furthermore, the adaptive denoising module obtains the law enforcement scenario of the law enforcement communication terminal and the general noise level of the video data. The process of adaptively denoising the video data of the law enforcement communication terminal includes:

[0015] Build an edge scene classification model and an adaptive denoising model on each law enforcement communication terminal, extract features from the multimodal data collected by the law enforcement communication terminal to generate a multimodal feature vector, input the multimodal feature vector into the edge scene classification model, and obtain the law enforcement scenario of the law enforcement communication terminal according to the output of the edge scene classification model;

[0016] Set a fixed number of frames n according to the law enforcement scenario of the law enforcement communication terminal, convert the video data collected by the law enforcement communication terminal into a continuous sequence of n frame images, and perform data format preprocessing on each frame of image to convert it into a grayscale image;

[0017] Divide each frame of grayscale image into several non-overlapping region blocks, evaluate the local variance of each pixel point in each non-overlapping region block to obtain the local variance of each pixel point, build a scene-noise intensity mapping model, and obtain the general noise level of the video data corresponding to the current law enforcement scenario according to the scene-noise intensity mapping model;

[0018] Input the local variance of each pixel point and the general noise level into the adaptive denoising model, output the Gaussian filtering parameters of each pixel point according to the adaptive denoising model, and apply the corresponding Gaussian filter to each pixel point of each frame of image according to the Gaussian filtering parameters. Furthermore, the process of outputting the Gaussian filtering parameters of each non-overlapping region block according to the adaptive denoising model includes:

[0019] For each pixel point (i, j), calculate the local variance within the non-overlapping region block to which the pixel point (i, j) belongs:

[0020] ;

[0021] ;

[0022] Among them, represents the local variance of the pixel point (i, j), represents the grayscale mean of the non-overlapping region block to which the pixel point (i, j) belongs, represents the pixel point at the gray value, W represents the set of all pixel points within the non-overlapping region block to which the pixel point (i, j) belongs, and n1 represents the total number of non-overlapping region block pixel points;

[0023] Build an adaptive denoising model of σ(i,j) according to the ratio of the local variance of the non-overlapping region block to the general noise level. The specific formula for outputting the Gaussian filtering parameters according to the adaptive denoising model is as follows:

[0024] ;

[0025] Among them, is the benchmark , represents the hyperparameter that controls the attenuation rate, represents the square of the standard deviation of the noise intensity. The above formulas are all calculated by removing the dimension and taking their numerical values. The formulas are obtained by collecting a large amount of data for software simulation to get a formula closest to the actual situation. The preset parameters and preset thresholds in the formulas are set by those skilled in the art according to the actual situation or obtained by simulating a large amount of data;

[0026] Furthermore, the process of the adaptive denoising module constructing the scene-noise intensity mapping model includes:

[0027] Obtain several sample frame images labeled with different law enforcement scenes;

[0028] Divide the sample frame images into several non-overlapping region blocks, obtain the gradient mean value of each non-overlapping region block, preset the gradient mean value threshold, mark the non-overlapping region blocks with the gradient mean value less than the gradient mean value threshold as key regions, evaluate the local variance of the key regions, and obtain the local variance of the key regions , , where N represents the number of pixel points in the key region, represents the gray value of the i-th pixel point in the key region, the gray mean value of, take the square root of the local variance to generate the noise standard deviation, and mark the noise standard deviation as the noise intensity of the sample frame image;

[0029] Perform mean processing on the noise intensities corresponding to all sample frame images of the same law enforcement scene to generate the general noise level of the video data corresponding to the law enforcement scene. The general noise level includes the noise intensity standard deviation and the noise intensity mean , and so on. Obtain the general noise levels of the video data corresponding to each law enforcement scene, construct a scene-noise intensity mapping model based on machine learning, use the general noise levels of the video data corresponding to each law enforcement scene as training data, and train the scene-noise intensity mapping model with the training data to obtain the trained scene-noise intensity mapping model;

[0030] Among them, the process of performing mean processing on the noise intensities corresponding to all sample frame images of the same law enforcement scene to generate the general noise level of the video data corresponding to the law enforcement scene and is as follows:

[0031] ;

[0032] ;

[0033] Where, M represents the total number of all sample frame images in the same law enforcement scenario, represents the noise intensity of the j-th frame image in the same law enforcement scenario, represents the average noise intensity in the same law enforcement scenario, represents the standard deviation of the noise intensity in the same law enforcement scenario;

[0034] Further, the process of the scene change detection module incrementally updating the general noise level of the video data of the law enforcement communication terminal and determining whether to perform scene reclassification based on the incremental update result includes:

[0035] Preset a scene change detection period, and mark the general noise level of the video data of the law enforcement communication terminal output by the scene-noise intensity mapping model as the judgment criterion;

[0036] Take the end timestamp of the scene change detection period as the detection time point. At the detection time point, obtain the general noise level of the video data at the end timestamp of the previous scene change detection period, and based on each frame image collected during the current scene change detection period of the law enforcement communication terminal, incrementally update the general noise level of the video data at the end timestamp of the previous scene change detection period;

[0037] Compare the incrementally updated general noise level with the judgment criterion to obtain a noise level deviation value. When the noise level deviation value is greater than the preset noise level deviation value threshold, trigger scene reclassification. The scene reclassification operation includes extracting multi-modal data during the current scene change detection period for feature extraction, generating a multi-modal feature vector, inputting the multi-modal feature vector into the edge scene classification model, and outputting the new law enforcement scene of the law enforcement communication terminal according to the edge scene classification model.

[0038] Further, the specific process of incrementally updating the general noise level of the video data at the end timestamp of the previous scene change detection period based on each frame image collected during the current scene change detection period includes:

[0039] Preset the general noise level at the end timestamp of the previous scene change detection period as , , the total number of image frames is n; the number of frames in the image sequence collected during the current scene change detection period is k. It should be noted that if the total number of image frames at the end timestamp of the previous scene change detection period is zero, that is, there is no previous scene change detection period, then take the general noise level of the law enforcement scene of the law enforcement communication terminal output by the scene-noise intensity mapping model as the general noise level at the end timestamp of the previous scene change detection period;

[0040] Segment each frame of the image collected within the current scene switching detection period into several non-overlapping regional blocks, obtain the gradient mean value of each non-overlapping regional block, preset a gradient mean value threshold, mark the non-overlapping regional blocks with a gradient mean value less than the gradient mean value threshold as key regions, perform a local variance evaluation on the key regions, obtain the local variance of the key regions, take the square root of the local variance to generate the noise standard deviation of each frame of the image, mark the noise standard deviation as the noise intensity, and obtain the incrementally updated general noise level and noise intensity mean value of the end timestamp of the current scene switching detection period based on the noise intensity of each frame of the image collected within the current scene switching detection period, the general noise level, and the noise intensity mean value of the end timestamp of the previous scene switching detection period;

[0041] ;

[0042] ;

[0043] wherein, represents the incrementally updated noise intensity standard deviation of the end timestamp of the current scene switching detection period, represents the incrementally updated noise intensity mean value of the end timestamp of the current scene switching detection period.

[0044] Furthermore, the process by which the image feature recognition module performs human feature recognition on the first frame of the video data to obtain the spatial distribution feature, pose feature, appearance feature, and identity identifier of the target human body includes:

[0045] Construct an edge pose estimation model and an appearance feature extraction model, input an n-frame image sequence into the edge pose estimation model, and according to the edge pose model, output the set of skeletal points of the target human body in each frame of the image, where the set of skeletal points includes the position coordinates of each skeletal point;

[0046] Obtain the number of skeletal points in the set of skeletal points of the target human body in the first frame of the n-frame image sequence, compare the number of skeletal points with a preset skeletal point threshold. If the number of skeletal points is less than the skeletal point threshold, then eliminate the set of skeletal points of the target human body. If the number of skeletal points is greater than or equal to the skeletal point threshold, then perform a skeletal point distribution analysis on the position coordinates of each skeletal point of the target human body to obtain the spatial distribution feature and pose feature of the target human body; the spatial distribution feature includes the aspect ratio and area of the bounding rectangle of all skeletal points, and the pose feature includes the joint angles and relative distances between each skeletal point;

[0047] The constraint threshold intervals corresponding to the preset spatial distribution features and pose features. If both the spatial distribution features and pose features of the target human body are within the corresponding constraint threshold intervals, then crop the region of interest in the first frame image according to the set of skeleton points, input the region of interest into the appearance feature extraction model, obtain the appearance features of the target human body according to the output of the appearance feature extraction model, generate an identity identifier based on the spatial distribution features, pose features, and appearance features of the target human body, and mark the identity identifier for each skeleton point in the set of skeleton points of the target human body. If either the spatial distribution features or pose features of the target human body are not within the corresponding constraint threshold intervals, then remove the set of skeleton points of the target human body.

[0048] Further, the process of the target tracking module performing human feature matching and marking occluded frames and visible frames on subsequent frame images of the video data includes:

[0049] Obtain the spatial distribution features, pose features, and appearance features of the target human body in each subsequent frame image after the first frame in the n-frame image sequence, compare the spatial distribution features, pose features, and appearance features of the target human body in each subsequent frame image with the spatial distribution features, pose features, and appearance features of the target human body in the first frame, and obtain the feature matching degree of the target human body in each subsequent frame image with the target human body in the first frame;

[0050] Preset a feature matching degree threshold. If the feature matching degree of the target human body in the subsequent frame image with the target human body in the first frame is greater than the feature matching degree threshold, then mark the identity identifier for each skeleton point in the set of skeleton points of the target human body in the subsequent frame image according to the identity identifier of the target human body in the first frame;

[0051] Preset the number of occluded frames ky. If there is no target human body in the subsequent continuous ky frame images whose feature matching degree with the target human body in the first frame is greater than the feature matching degree threshold, then perform dynamic graph convolutional target tracking operations, and mark the frames without a target human body whose feature matching degree with the target human body in the first frame is greater than the feature matching degree threshold as occluded frames, and mark the other frames in the n-frame image sequence except the occluded frames as visible frames.

[0052] Further, the process of the target tracking module marking the tracking target, constructing a dynamic graph convolutional network, modeling each tracking target as a graph node, constructing the dynamic edge weights between each node, and outputting the spatial features of the tracking target in each occluded frame includes:

[0053] Mark the target human body in the first frame of the n-frame image sequence as the tracking target, and mark the position coordinates of each skeleton point of the tracking target in each visible frame of the n-frame image sequence as spatial features , , where Represents the spatial features of the i-th tracking target in the t-th visible frame. Represents the total number of skeleton points. Perform motion feature analysis on the spatial feature sequences of the tracking targets in each visible frame to obtain motion features. , , Construct a spatio-temporal graph by connecting each node according to the spatial features and motion features of each tracking target in each visible frame as nodes.

[0054] Preset the occlusion frame number threshold, obtain the number of occlusion frames in the n-frame image sequence. If the number of occlusion frames is less than or equal to the occlusion frame number threshold, obtain the dynamic edge weights between each node according to the spatial features, motion features, and law enforcement scenarios of each tracking target in each visible frame. If the number of occlusion frames is greater than the occlusion frame number threshold, obtain the dynamic edge weights between each node according to the motion features and law enforcement scenarios of each tracking target in each visible frame.

[0055] Among them, the calculation formula for obtaining the dynamic edge weights between each node according to the spatial features, motion features, and current law enforcement scenario of each tracking target in each visible frame is:

[0056] ;

[0057] Among them, the calculation formula for obtaining the dynamic edge weights between each node according to the motion features and current law enforcement scenario of each tracking target in each visible frame is:

[0058] ;

[0059] Among them, Represents the dynamic edge weight between the i-th tracking target and the j-th tracking target within the t-th visible frame. Represents the spatial distance term, which is used to measure the proximity of the positions of targets i and j in the current visible frame. The closer the positions, the higher the edge weight. Represents the time distance term, which is used to measure the similarity of target speeds. The more consistent the speeds, the higher the edge weight. Represents the sensitivity coefficient, which is used to control the sensitivity of the weight. The larger, the higher the "tolerance" of the dynamic edge weight to the distance difference. Preset the corresponding sensitivity coefficients under different law enforcement scenarios, obtain the current law enforcement scenario, and adaptively determine the sensitivity coefficient according to the current law enforcement scenario.

[0060] Construct a spatial feature prediction model based on the graph convolutional neural network, and perform learning representation on the spatio-temporal graph, including: fuse node features through the graph convolutional layer (GCN), and the formula is:

[0061] ;

[0062] Among them, The hidden state of node u at frame t; Denote other nodes connected to node u and the node itself in historical visible frames, Denote the learnable weight matrix;

[0063] Use the output of graph convolution as the prediction of skeleton point coordinates, and the formula is:

[0064] ;

[0065] Among them, and are regression weights;

[0066] Output the spatial features of the tracking targets in each occluded frame.

[0067] Furthermore, the hierarchical upload module constructs an action label recognition model to output the danger warning levels of each tracking target. The process of data hierarchical upload according to the danger warning levels includes:

[0068] Construct an action label recognition model, input the spatial features of the tracking targets in each visible frame and occluded frame into the action label recognition model, output the action labels of each tracking target according to the action label recognition model, and set the danger warning levels of each tracking target according to the action labels and identity identifiers of each tracking target. The specific process of setting the danger warning levels of each tracking target includes: preset an action label - identity identifier - danger warning level mapping table, where the action label - identity identifier - danger warning level mapping table includes the danger warning level mapping tables corresponding to different action labels of different identity identifiers, input the action labels and identity identifiers of each tracking target into the action label - identity identifier - danger warning level mapping table, and obtain the danger warning levels of each tracking target;

[0069] Among them, the process of constructing the action label recognition model includes:

[0070] Construct an action label recognition model based on the LSTM network, preset the spatial feature sequences corresponding to several types of action labels, where the action labels include normal standing, raising a hand to signal, pushing and shoving conflicts, etc., use the spatial feature sequences corresponding to several types of action labels as the training set and the test set, input the training set into the action label recognition model for training until the loss function is trained stably, save the model parameters, test the action label recognition model through the test set until it meets the preset requirements, and output the action label recognition model;

[0071] Preset a dangerous warning level threshold. If the dangerous warning level of the tracked target is greater than or equal to the preset dangerous warning level threshold, then pack the dangerous warning level, spatial distribution characteristics, pose characteristics, and appearance characteristics of the tracked target into a law enforcement data packet and encrypt and upload it to the blockchain node;

[0072] Preset a regular upload period, and use the end timestamp of the regular upload period as the upload time point. If the dangerous warning level of the tracked target is less than the preset dangerous warning level threshold, then pack the dangerous warning level, spatial distribution characteristics, pose characteristics, and appearance characteristics of the tracked target into a law enforcement data packet and encrypt and upload it to the blockchain node at the upload time point;

[0073] The blockchain node decrypts the law enforcement data packet uploaded to the blockchain node, presets the accounting node, consensus mechanism, and smart contract of the law enforcement data sharing platform. The accounting node packs the decrypted law enforcement data packet into a new block. This block includes a block header and a transaction list. The accounting node broadcasts the new block to the entire blockchain network according to the peer-to-peer communication protocol. Other blockchain nodes in the blockchain network verify the correctness of the block data and the validity of the transaction data of the new block according to the consensus mechanism and the smart contract. After the verification of the new block passes, the new block is added to the end of the blockchain, and the block information of the new block is updated to the local blockchain copies of all blockchain nodes.

[0074] An intelligent supervision method for law enforcement behavior based on edge computing includes the following steps:

[0075] Step s1: Collect multi-modal data of each law enforcement communication terminal, construct a scene-noise intensity mapping model, obtain the law enforcement scene of the law enforcement communication terminal and the general noise level of the video data, and perform adaptive denoising on the video data of the law enforcement communication terminal;

[0076] Step s2: Incrementally update the general noise level of the video data of the law enforcement communication terminal, and judge whether to perform scene reclassification according to the incremental update result;

[0077] Step s3: Perform human feature recognition on the first frame image of the video data to obtain the spatial distribution characteristics, pose characteristics, appearance characteristics, and identity identifier of the target human body;

[0078] Step s4: Perform human feature matching on the subsequent frame images of the video data, mark the occluded frames and visible frames, and construct a dynamic graph convolutional network according to the marking results, mark the tracked targets, model each tracked target as a graph node, construct the dynamic edge weights between each node, and output the spatial features of the tracked targets in each occluded frame;

[0079] Step s5: Build an action label recognition model, output the danger warning levels of each tracking target, and upload data in levels according to the danger warning levels.

[0080] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0081] 1. Precise denoising: The adaptive denoising module can accurately analyze the general noise level of video data in different law enforcement scenarios (such as at night, in rainy days, indoors, etc.) by constructing a scene-noise intensity mapping model. This makes the denoising process of video data more targeted. Compared with traditional general denoising methods, it can better retain the effective information in the video, improve the image quality, and provide a clearer data basis for subsequent analysis tasks. For example, in a low-light environment at night, it can effectively remove Gaussian noise, making the outlines of law enforcement personnel and relevant objects clearer for identification and analysis.

[0082] 2. Real-time scene adaptation: The scene switching detection module can incrementally update the general noise level of video data and determine whether to reclassify the scene based on the update result. This feature enables the system to adapt to the dynamic changes of law enforcement scenarios in real time. For example, when quickly switching from an outdoor sunny scene to an indoor dim light scene, the system can rapidly adjust the processing method of video data, continuously maintain accurate analysis and processing of various data, and ensure the coherence and accuracy of law enforcement supervision.

[0083] 3. Efficient target tracking: The target tracking module uses a dynamic graph convolutional network to model the tracking target as a graph node and construct dynamic edge weights. It performs excellently in dealing with occlusion situations and can relatively accurately output the spatial features of the tracking target in the occluded frame through the information of visible frames. For example, in a crowded law enforcement scene, when the target is partially occluded, it can still accurately track its position and movement trend, greatly improving the stability and reliability of target tracking and providing continuous and accurate target movement data for law enforcement behavior analysis.

[0084] 4. The image feature recognition module performs human feature recognition on the first frame image of video data and can obtain the spatial distribution features, pose features, appearance features, and identity identifiers of the target human body. These multi-dimensional feature information provides rich and accurate initial data for subsequent target tracking and behavior analysis, helping to quickly and accurately locate and identify different individuals in complex law enforcement scenarios. For example, the relative position relationship between targets can be judged through spatial distribution features, and the behavior intention can be preliminarily judged through pose features, etc.

[0085] 5. The hierarchical upload module constructs an action label recognition model, sets the danger warning level according to the action labels and identity identifiers of the tracking targets, and uploads data in a hierarchical manner based on the danger warning level. This method optimizes the data transmission strategy. For data with a higher danger warning level, such as data involving emergencies like violent conflicts, it can be uploaded to the blockchain nodes in a timely and prioritized manner, facilitating the law enforcement department to respond quickly. For data with a lower danger warning level, it is uploaded within an appropriate timed upload cycle, effectively balancing the timeliness of data transmission and the rational utilization of network resources, and improving the efficiency and security of data management.

[0086] 6. Build a law enforcement data sharing platform based on blockchain technology. The blockchain nodes in the platform are in communication links with law enforcement communication terminals. The distributed ledger feature of the blockchain ensures the security and immutability of law enforcement data. Each blockchain node is interconnected to form a blockchain network, enabling the secure sharing of law enforcement data between different nodes. This not only enhances the credibility of the data but also facilitates data collaboration and supervision between different law enforcement departments. For example, in cross-regional law enforcement operations, law enforcement departments in different regions can securely and accurately obtain relevant law enforcement data through the blockchain network, improving the efficiency of law enforcement collaboration. Brief Description of the Drawings

[0087] Figure 1 It is a schematic diagram of the intelligent supervision system for law enforcement behavior based on edge computing in an embodiment of the present application.

[0088] Figure 2 It is a schematic diagram of the intelligent supervision method for law enforcement behavior based on edge computing in an embodiment of the present application. Detailed Embodiments

[0089] Next, in combination with the drawings in the embodiments of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts fall within the scope of protection of the present application.

[0090] As Figure 1 shown, the intelligent supervision system for law enforcement behavior based on edge computing includes a law enforcement data sharing platform, and the law enforcement data sharing platform is in communication links with a number of law enforcement communication terminals, an adaptive denoising module, a scene switching detection module, a target tracking module, and a hierarchical upload module;

[0091] The law enforcement communication terminals are used to collect multimodal data, and the multimodal data includes video data and sensor data;

[0092] The adaptive denoising module is used to construct a scene-noise intensity mapping model, obtain the law enforcement scene of the law enforcement communication terminal and the general noise level of the video data, and perform adaptive denoising on the video data of the law enforcement communication terminal;

[0093] The scene switching detection module is used to incrementally update the general noise level of the video data of the law enforcement communication terminal, and judge whether to reclassify the scene according to the incremental update result;

[0094] The image feature recognition module is used to perform human feature recognition on the first frame image of the video data, and obtain the spatial distribution features, pose features, appearance features and identity identifiers of the target human body;

[0095] The target tracking module is used to perform human feature matching and occlusion frame and visible frame marking on the subsequent frame images of the video data, and construct a dynamic graph convolutional network according to the marking results, mark the tracking targets, model each tracking target as a graph node, construct the dynamic edge weights between each node, and output the spatial features of the tracking targets in each occlusion frame;

[0096] The hierarchical upload module is used to construct an action label recognition model, output the danger warning levels of each tracking target, and perform data hierarchical upload according to the danger warning levels.

[0097] It should be further noted that in the specific implementation process, a law enforcement data sharing platform is constructed based on blockchain technology. There are several blockchain nodes in the law enforcement data sharing platform. The blockchain nodes are communicatively linked to law enforcement communication terminals (such as law enforcement recorders, law enforcement vehicles, law enforcement booths, etc.) within a preset range, and are used to receive law enforcement data packets uploaded by the law enforcement communication terminals. Each blockchain node is linked to each other to form a blockchain network.

[0098] It should be further noted that in the specific implementation process, the process of the adaptive denoising module obtaining the law enforcement scene of the law enforcement communication terminal and the general noise level of the video data and performing adaptive denoising on the video data of the law enforcement communication terminal includes:

[0099] Build an edge scenario classification model and an adaptive denoising model on each law enforcement communication terminal. Among them, the process of building the edge scenario classification model includes: pre-collecting a large amount of historical multi-modal data from each law enforcement communication terminal for law enforcement scenario annotation. For example, if the light intensity < 50 lux, annotate the sub-class "night - low light", and if the humidity > 80% and the time is in the rainy season, annotate the sub-class "rainy day". Build an edge scenario classification model based on Faster R-CNN, and train the edge scenario classification model with a large amount of historical multi-modal data that has completed law enforcement scenario annotation to obtain a trained edge scenario classification model. Extract features from the multi-modal data (including video data and sensor data collected by the law enforcement communication terminal, and the sensor data includes light data, acceleration data, humidity data, etc.) collected by the law enforcement communication terminal to generate a multi-modal feature vector. The multi-modal feature vector includes visual features and sensor features. For example, fuse video data, light intensity (lux), humidity (%), and acceleration (judge the motion state): {multi-modal feature vector = [{brightness mean}, {saturation}, {light}, {humidity}, {acceleration}, {timestamp}]}, and input the multi-modal feature vector into the edge scenario classification model to obtain the law enforcement scenario of the law enforcement communication terminal according to the output of the edge scenario classification model.

[0100] Set a fixed number of frames n according to the law enforcement scenario of the law enforcement communication terminal, convert the video data collected by the law enforcement communication terminal into a continuous sequence of n-frame images, and perform data format preprocessing on each frame of the image. The data format preprocessing includes adjusting each frame of the image to a unified size, normalizing the image pixel values, scaling the pixel values to a specific range (such as [0, 1] or [-1, 1]), and grayscale processing to convert it into a grayscale image.

[0101] Divide each frame of the grayscale image into several non-overlapping region blocks. For example, divide the image into several non-overlapping region blocks with a size of 32×32 pixels, and there is no pixel overlap between adjacent blocks. Starting from the upper left corner of the image, first take the block of columns 1 - 32 in the first row, then take the block of columns 33 - 64 in the first row, and so on. After each row ends, move down 32 rows to start a new row. Evaluate the local variance of each pixel point in each non-overlapping region block to obtain the local variance of each pixel point, build a scene-noise intensity mapping model, and obtain the general noise level of the video data corresponding to the current law enforcement scenario according to the scene-noise intensity mapping model.

[0102] The local variance of each pixel and the general noise level are input into an adaptive denoising model. According to the adaptive denoising model, the Gaussian filtering parameters for each pixel are output. Each pixel of each frame of the image is processed by applying the corresponding Gaussian filter according to the Gaussian filtering parameters. Specifically, the process of the adaptive denoising model outputting the Gaussian filtering parameters for each pixel is to generate a suitable standard deviation σ for each pixel in each non-overlapping region block. When the local variance is close to the general noise level, it indicates that this region is mainly composed of noise. At this time, the value of σ should be increased to enhance the smoothing effect. On the contrary, if the local variance is much larger than the general noise level, it indicates that there may be important image features such as edges here. At this time, the value of σ should be reduced to reduce the impact on these features. Due to the adoption of the adaptive strategy, even in different regions of the same picture, differential filtering results can be obtained, thus better protecting the fine structures in the image.

[0103] It should be further noted that in the specific implementation process, the process of outputting the Gaussian filtering parameters for each non-overlapping region block according to the adaptive denoising model includes:

[0104] For each pixel (i, j), calculate the local variance within the non-overlapping region block to which the pixel (i, j) belongs:

[0105] ;

[0106] ;

[0107] Among them, represents the local variance of the pixel (i, j), represents the gray-scale mean value of the non-overlapping region block to which the pixel (i, j) belongs, represents the gray-scale value at the pixel location, W represents the set of all pixels within the non-overlapping region block to which the pixel (i, j) belongs, and n1 represents the total number of pixels in the non-overlapping region block;

[0108] According to the ratio of the local variance of the non-overlapping region block to the general noise level, an adaptive denoising model of σ(i,j) is constructed. The specific formula for outputting the Gaussian filtering parameters according to the adaptive denoising model is as follows:

[0109] ;

[0110] Among them, is the benchmark , represents the hyperparameter that controls the attenuation speed, represents the square of the standard deviation of the noise intensity. All the above formulas are calculated by removing the dimension and taking the numerical value. The formula is obtained by collecting a large amount of data for software simulation to get a formula closest to the actual situation. The preset parameters and preset thresholds in the formula are set by those skilled in the art according to the actual situation or obtained through simulation of a large amount of data;

[0111] It should be further noted that in the specific implementation process, the process of the adaptive denoising module constructing the scene-noise intensity mapping model includes:

[0112] Obtain a number of sample frame images labeled with different law enforcement scenarios through channels such as COCO / OpenImages;

[0113] Divide the sample frame images into several non-overlapping region blocks, obtain the gradient mean value of each non-overlapping region block (implemented through the Sobel operator), preset the gradient mean value threshold, mark the non-overlapping region blocks with a gradient mean value less than the gradient mean value threshold as key regions, evaluate the local variance of the key regions, and obtain the local variance of the key regions , , where N represents the number of pixel points in the key region, represents the gray value of the i-th pixel point in the key region, the gray mean value, take the square root of the local variance to generate the noise standard deviation, and mark the noise standard deviation as the noise intensity of the sample frame image;

[0114] Perform mean processing on the noise intensities corresponding to all sample frame images of the same law enforcement scenario to generate the general noise level of the video data corresponding to the law enforcement scenario. The general noise level includes the standard deviation of the noise intensity and the mean value of the noise intensity , and so on, obtain the general noise levels of the video data corresponding to each law enforcement scenario, construct a scene-noise intensity mapping model based on machine learning, use the general noise levels of the video data corresponding to each law enforcement scenario as training data, and train the scene-noise intensity mapping model with the training data to obtain the trained scene-noise intensity mapping model;

[0115] Among them, the process of performing mean processing on the noise intensities corresponding to all sample frame images of the same law enforcement scenario to generate the general noise level of the video data corresponding to the law enforcement scenario and is as follows:

[0116] ;

[0117] ;

[0118] Among them, M represents the total number of all sample frame images of the same law enforcement scenario, represents the noise intensity of the j-th frame image within the same law enforcement scenario, represents the average noise intensity of the same law enforcement scenario, represents the standard deviation of the noise intensity of the same law enforcement scenario;

[0119] It should be further noted that in the specific implementation process, the process of the scene switching detection module incrementally updating the general noise level of the video data of the law enforcement communication terminal and determining whether to reclassify the scene according to the incremental update result includes:

[0120] Preset a scene switching detection period, and mark the general noise level of the video data of the law enforcement communication terminal output by the scene-noise intensity mapping model as the judgment criterion;

[0121] Take the end timestamp of the scene switching detection period as the detection time point, obtain the general noise level of the video data at the end timestamp of the previous scene switching detection period at the detection time point, and incrementally update the general noise level of the video data at the end timestamp of the previous scene switching detection period according to each frame image collected within the current scene switching detection period of the law enforcement communication terminal;

[0122] Compare the incrementally updated general noise level with the judgment criterion to obtain the noise level deviation value. When the noise level deviation value is greater than the preset noise level deviation value threshold, trigger scene reclassification. The scene reclassification operation includes extracting multi-modal data within the current scene switching detection period for feature extraction, generating a multi-modal feature vector, inputting the multi-modal feature vector into the edge scene classification model, and outputting the new law enforcement scene of the law enforcement communication terminal according to the edge scene classification model.

[0123] It should be further noted that in the specific implementation process, the specific process of incrementally updating the general noise level of the video data at the end timestamp of the previous scene switching detection period according to each frame image collected within the current scene switching detection period includes:

[0124] Preset the general noise level at the end timestamp of the previous scene switching detection period as 、 、the total number of image frames is n; the number of frames of the image sequence collected within the current scene switching detection period is k. It should be noted that if the total number of image frames at the end timestamp of the previous scene switching detection period is zero, that is, there is no previous scene switching detection period, then take the general noise level of the law enforcement scene of the law enforcement communication terminal output by the scene-noise intensity mapping model as the general noise level at the end timestamp of the previous scene switching detection period;

[0125] Each frame of image collected during the current scene switch detection period is divided into a number of non-overlapping area blocks, a gradient mean of each non-overlapping area block is obtained, a gradient mean threshold is preset, non-overlapping area blocks whose gradient mean is less than the gradient mean threshold are marked as key areas, a local variance evaluation is performed on the key areas, the local variance of the key areas is obtained, the square root of the local variance is taken, a noise standard deviation of each frame of image is generated, the noise standard deviation is marked as noise intensity, and based on the noise intensity of each frame of image collected during the current scene switch detection period and the universal noise level and noise intensity mean of the end timestamp of the previous scene switch detection period, the universal noise level and noise intensity mean after incremental update of the end timestamp of the current scene switch detection period are obtained;

[0126] ;

[0127] ;

[0128] in, Indicates the standard deviation of the noise intensity after incremental update of the end timestamp of the current scene switch detection period, Indicates the noise intensity mean after incremental update of the end timestamp of the current scene switch detection period.

[0129] It should be further explained that, in a specific implementation, the image feature recognition module performs human feature recognition on the first frame of the video data, and the process of obtaining the spatial distribution features, posture features, appearance features, and identity identifier of the target human body includes:

[0130] Construct an edge pose estimation model and an appearance feature extraction model, input a sequence of n frames of images into the edge pose estimation model, and output a set of skeleton points of the target human body in each frame of the image according to the edge pose model, wherein the skeleton point set includes the position coordinates of each skeleton point;

[0131] Due to the limited computing power of edge computing, the lightweight backbone network MobileNetV3 was selected as the architecture of the edge pose estimation model. The design of the pose estimation head chose single-stage regression to directly predict the coordinates of skeleton points (such as SimpleBaseline) to avoid multi-stage inference delays. At the same time, the convolutional layer channels were pruned and low-importance channels were removed (such as by filtering the BN layer scaling factor). For example, the number of backbone network channels was pruned to 60% of the original model, keeping the accuracy loss less than 2%. The training strategy of the edge pose estimation model was optimized, and Smooth L1 Loss was selected to directly optimize the coordinates of skeleton points. Compared with heat map loss, it is more suitable for fast inference of edge devices.

[0132] Obtain the number of skeleton points in the set of skeleton points of the target human body in the first frame of the n-frame image sequence, compare the number of skeleton points with a preset skeleton point threshold. If the number of skeleton points is less than the skeleton point threshold, then eliminate the set of skeleton points of the target human body. If the number of skeleton points is greater than or equal to the skeleton point threshold, then perform a skeleton point distribution analysis on the position coordinates of each skeleton point of the target human body to obtain the spatial distribution characteristics and pose characteristics of the target human body; the spatial distribution characteristics include the aspect ratio and area of the bounding rectangle of all skeleton points, and the pose characteristics include the joint angles and relative distances between each skeleton point.

[0133] Preset the constraint threshold intervals corresponding to the spatial distribution characteristics and pose characteristics. If both the spatial distribution characteristics and pose characteristics of the target human body are within the corresponding constraint threshold intervals, then crop the region of interest according to the set of skeleton points in the first frame image, input the region of interest into the appearance feature extraction model, and according to the output of the appearance feature extraction model, obtain the appearance features of the target human body (including color, texture features, police number encoding, etc.). Generate an identity identifier based on the spatial distribution characteristics, pose characteristics, and appearance features of the target human body, and mark the identity identifier for each skeleton point in the set of skeleton points of the target human body. If either the spatial distribution characteristics or pose characteristics of the target human body are not within the corresponding constraint threshold intervals, then eliminate the set of skeleton points of the target human body.

[0134] It should be further noted that in the specific implementation process, the process of the target tracking module performing human feature matching and marking occluded frames and visible frames on the subsequent frame images of the video data includes:

[0135] Obtain the spatial distribution characteristics, pose characteristics, and appearance features of the target human body in each subsequent frame image after the first frame in the n-frame image sequence, compare the spatial distribution characteristics, pose characteristics, and appearance features of the target human body in each subsequent frame image with the spatial distribution characteristics, pose characteristics, and appearance features of the target human body in the first frame to obtain the feature matching degree of the target human body in each subsequent frame image with the target human body in the first frame.

[0136] Preset a feature matching degree threshold. If the feature matching degree of the target human body in the subsequent frame image with the target human body in the first frame is greater than the feature matching degree threshold, then mark the identity identifier for each skeleton point in the set of skeleton points of the target human body in the subsequent frame image according to the identity identifier of the target human body in the first frame.

[0137] Preset the number of occluded frames \(k_y\). If there is no target human body in the subsequent consecutive \(k_y\) frames whose feature matching degree with the target human body in the first frame is greater than the feature matching degree threshold, perform dynamic graph convolution target tracking operation, and mark the frames without a target human body whose feature matching degree with the target human body in the first frame is greater than the feature matching degree threshold as occluded frames, and mark the other frames except the occluded frames in the \(n\)-frame image sequence as visible frames.

[0138] It should be further noted that in the specific implementation process, the process of the target tracking module marking the tracking target, constructing a dynamic graph convolution network, modeling each tracking target as a graph node, constructing the dynamic edge weights between each node, and outputting the spatial features of the tracking target in each occluded frame includes:

[0139] Mark the target human body in the first frame of the \(n\)-frame image sequence as the tracking target, and mark the position coordinates of each bone point of the tracking target in each visible frame of the \(n\)-frame image sequence as spatial features , , where represents the spatial feature of the \(i\)-th tracking target in the \(t\)-th visible frame, represents the total number of bone points, and perform motion feature analysis on the spatial feature sequence of the tracking target in each visible frame to obtain motion features , , construct a spatio-temporal graph by connecting each node with the spatial features and motion features of each tracking target in each visible frame as nodes;

[0140] Preset the occluded frame number threshold, obtain the number of occluded frames in the \(n\)-frame image sequence. If the number of occluded frames is less than or equal to the occluded frame number threshold, obtain the dynamic edge weights between each node according to the spatial features, motion features of each tracking target in each visible frame and the current law enforcement scenario. If the number of occluded frames is greater than the occluded frame number threshold, obtain the dynamic edge weights between each node according to the motion features of each tracking target in each visible frame and the current law enforcement scenario;

[0141] For example, when facing the need to identify collaborative behavior requirements such as "multiple people surrounding" and "group action", when the spatial positions of multiple targets are close and the speeds are the same (such as 3 law enforcement officers surrounding a suspect in a triangle), the dynamic edge weights will form a strongly connected subgraph, directly corresponding to the "collaborative group", providing a priori for subsequent action analysis;

[0142] Among them, the calculation formula for obtaining the dynamic edge weights between each node according to the spatial features, motion features of each tracking target in each visible frame and the current law enforcement scenario is:

[0143] ;

[0144] Among them, the calculation formula for obtaining the dynamic edge weights between each node according to the motion characteristics of each tracking target in each visible frame and the current law enforcement scenario is as follows:

[0145] ;

[0146] Among them, represents the dynamic edge weight between the i-th tracking target and the j-th tracking target in the t-th visible frame, represents the spatial distance term, which is used to measure the proximity of the positions of targets i and j in the current visible frame. The closer the positions are, the higher the edge weight. represents the time distance term, which is used to measure the similarity of target speeds. The more consistent the speeds are, the higher the edge weight. represents the sensitivity coefficient, which is used to control the sensitivity of the weight. The larger it is, the higher the "tolerance" of the dynamic edge weight to distance differences. Corresponding sensitivity coefficients are preset for different law enforcement scenarios. Obtain the current law enforcement scenario, adaptively determine the sensitivity coefficient according to the current law enforcement scenario, and adapt to law enforcement scenarios in different environments through adaptive adjustment For example, if the law enforcement scenarios include night and rain, then is smaller, and the "tolerance" of the dynamic edge weight to distance differences is lower. If the law enforcement scenarios include day and clear, then is larger, and the "tolerance" of the dynamic edge weight to distance differences is greater. The above formulas are all calculated by removing the dimension and taking their numerical values. The formula is obtained by collecting a large amount of data for software simulation to get a formula closest to the real situation. The preset parameters and preset thresholds in the formula are set by those skilled in the art according to the actual situation or obtained through a large amount of data simulation;

[0147] Construct a spatial feature prediction model based on the graph convolutional neural network to learn and represent the spatio-temporal graph, including: fusing node features through the graph convolutional layer (GCN), and the formula is:

[0148] ;

[0149] Among them, is the hidden state of node u at frame t; represents other nodes connected to node u and its own node in the historical visible frames, represents the learnable weight matrix; [[ID=^40]]

[0150] Use the graph convolution output as the prediction of the bone point coordinates, and the formula is:

[0151] ;

[0152] Among them, , are the regression weights;

[0153] Output the spatial features of the tracked targets in each occluded frame. Traditional object tracking methods (such as Kalman filtering and Hungarian algorithm) rely on heuristic rules (such as IOU overlap and speed matching), making it difficult to capture the interaction relationships (such as occlusion, cooperation, and competition) between targets in law enforcement scenarios. For example, when target A is occluded by target B, an IOU-based tracker may misjudge the identity (identity switch) due to the loss of the appearance features of A. This embodiment performs object tracking based on a dynamic graph convolutional network, treating each target as a node of the graph. The dynamic edge weights represent the "association strength" between targets, and this association is dynamically quantified through spatio-temporal distance, thereby explicitly modeling the dependence relationships between targets. For example, even if A is occluded, the speed information of the historical trajectory ( ) can still be associated with A in the previous frame through the dynamic edge weights; if the speed difference between B and A is large (such as B suddenly accelerating), the edge weight will decrease, avoiding misjudging B as A. Compared with traditional methods, the ID switch rate of this embodiment can be reduced by 25%;

[0154] At the same time, traditional methods based on matching of adjacent frames (such as matching within 3 frames) cannot handle the situation after a target is occluded for a long time. This embodiment realizes long-distance cross-frame association through the global connection of the graph (each node is connected to all other nodes), and the dynamic edge weights can trace the spatio-temporal information of multiple historical frames (such as the speed direction before the target disappears). For example, when a law enforcement officer is briefly occluded in a crowd, the spatial feature prediction model associates with the target in the same direction that appears after occlusion through its historical moving direction (speed vector), avoiding trajectory breakage.

[0155] It should be further noted that in the specific implementation process, the hierarchical upload module constructs an action label recognition model and outputs the danger warning levels of each tracked target. The process of data hierarchical upload according to the danger warning levels includes:

[0156] Construct an action label recognition model, input the spatial features of the tracked targets in each visible frame and occluded frame into the action label recognition model, output the action labels of each tracked target according to the action label recognition model, and set the danger warning levels of each tracked target according to the action labels and identity identifiers of each tracked target. The specific process of setting the danger warning levels of each tracked target includes: preset an action label - identity identifier - danger warning level mapping table, where the action label - identity identifier - danger warning level mapping table includes the danger warning level mapping tables corresponding to different action labels of different identity identifiers, input the action labels and identity identifiers of each tracked target into the action label - identity identifier - danger warning level mapping table, and obtain the danger warning levels of each tracked target;

[0157] Among them, the process of constructing the action label recognition model includes:

[0158] Build an action label recognition model based on the LSTM network, preset the spatial feature sequences corresponding to several types of action labels, where the action labels include normal standing, raising a hand to signal, pushing and shoving conflicts, etc. Use the spatial feature sequences corresponding to several types of action labels as the training set and the test set. Input the training set into the action label recognition model for training until the loss function is trained smoothly, and save the model parameters. Test the action label recognition model with the test set until it meets the preset requirements, and output the action label recognition model;

[0159] Preset a dangerous warning level threshold. If the dangerous warning level of the tracked target is greater than or equal to the preset dangerous warning level threshold, then package the dangerous warning level, spatial distribution characteristics, pose characteristics, and appearance characteristics of the tracked target into an enforcement data packet for encrypted upload to the blockchain node;

[0160] Preset a regular upload period, and use the end timestamp of the regular upload period as the upload time point. If the dangerous warning level of the tracked target is less than the preset dangerous warning level threshold, then package the dangerous warning level, spatial distribution characteristics, pose characteristics, and appearance characteristics of the tracked target into an enforcement data packet for encrypted upload to the blockchain node at the upload time point;

[0161] The blockchain node decrypts the enforcement data packet uploaded to the blockchain node, presets the accounting node, consensus mechanism, and smart contract of the law enforcement data sharing platform. The accounting node packages the decrypted enforcement data packet into a new block. This block includes a block header and a transaction list. The accounting node broadcasts the new block to the entire blockchain network according to the peer-to-peer communication protocol. Other blockchain nodes in the blockchain network verify the correctness of the block data and the validity of the transaction data of the new block according to the consensus mechanism and the smart contract. After the verification of the new block passes, the new block is added to the end of the blockchain, and the block information of the new block is updated to the local blockchain copies of all blockchain nodes.

[0162] As Figure 2 shown, the intelligent supervision method for law enforcement behavior based on edge computing includes the following steps:

[0163] Step s1: Collect the multi-modal data of each law enforcement communication terminal, build a scene-noise intensity mapping model, obtain the law enforcement scene of the law enforcement communication terminal and the general noise level of the video data, and perform adaptive denoising on the video data of the law enforcement communication terminal;

[0164] Step s2: Incrementally update the general noise level of the video data of the law enforcement communication terminal, and judge whether to re-classify the scene according to the incremental update result;

[0165] Step s3: Perform human feature recognition on the first frame image of the video data to obtain the spatial distribution features, pose features, appearance features of the target human body, and identity identifiers;

[0166] Step s4: Perform human feature matching and occlusion frame and visible frame marking on the subsequent frame images of the video data, and construct a dynamic graph convolutional network according to the marking results, mark and track the targets, model each tracked target as a graph node, construct the dynamic edge weights between each node, and output the spatial features of the tracked targets in each occlusion frame;

[0167] Step s5: Construct an action label recognition model, output the danger warning levels of each tracked target, and perform data classification and upload according to the danger warning levels.

[0168] The above embodiments are only used to illustrate the technical method of the present invention and not to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical method of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical method of the present invention.

Claims

1. The intelligent supervision system for law enforcement based on edge computing is characterized by: It includes a law enforcement data sharing platform, wherein the law enforcement data sharing platform is connected to a plurality of law enforcement communication terminals, and the law enforcement communication terminals are connected to an adaptive denoising module, a scene switching detection module, a target tracking module and a hierarchical uploading module; The law enforcement communication terminal is used to collect multimodal data, wherein the multimodal data includes video data; The adaptive denoising module is used to build a scene-noise intensity mapping model to obtain the law enforcement scene and the general noise level of the video data of the law enforcement communication terminal, and to perform adaptive denoising on the video data of the law enforcement communication terminal, including: Obtain several sample frame images labeled with different law enforcement scenes; The sample frame image is divided into a number of non-overlapping area blocks, the gradient mean of each non-overlapping area block is obtained, a gradient mean threshold is preset, the non-overlapping area blocks whose gradient mean is less than the gradient mean threshold are marked as key areas, local variance evaluation is performed on the key areas, the local variance of the key areas is obtained, the local variance is squared to generate a noise standard deviation, and the noise standard deviation is marked as the noise intensity of the sample frame image; Performing mean processing on the noise intensities corresponding to all sample frame images of the same law enforcement scene to generate a universal noise level of the video data corresponding to the law enforcement scene. Similarly, obtaining the universal noise level of the video data corresponding to each law enforcement scene, constructing a scene-noise intensity mapping model, using the universal noise level of the video data corresponding to each law enforcement scene as training data, and using the training data to train the scene-noise intensity mapping model to obtain a trained scene-noise intensity mapping model; Construct an edge scene classification model and an adaptive denoising model, extract features from multimodal data, generate multimodal feature vectors, input the multimodal feature vectors into the edge scene classification model, and output the law enforcement scene of the law enforcement communication terminal based on the edge scene classification model; A fixed frame number n is set according to the law enforcement scenario of the law enforcement communication terminal, and the video data collected by the law enforcement communication terminal is converted into a continuous n-frame image sequence. Each frame image is pre-processed in data format and converted into a grayscale image; Each frame of grayscale image is divided into several non-overlapping area blocks, and local variance evaluation is performed on each pixel in each non-overlapping area block to obtain the local variance of each pixel. The general noise level of the video data corresponding to the current law enforcement scene is obtained based on the scene-noise intensity mapping model; The local variance and the general noise level of each pixel point of each frame image are input into the adaptive denoising model, the Gaussian filter parameters of each pixel point are output according to the adaptive denoising model, and the corresponding Gaussian filter is applied to each pixel point according to the Gaussian filter parameters; The scene change detection module is used to incrementally update the prevailing noise level of the video data of the law enforcement communication terminal and determine whether to reclassify the scene based on the incremental update result; The image feature recognition module is used to perform human feature recognition on the first frame of the video data to obtain the spatial distribution features, posture features, appearance features and identity identifier of the target human body; The target tracking module is used to match human features and mark occluded and visible frames in subsequent frames of video data. Based on the marking results, a dynamic graph convolutional network is constructed to mark the tracking targets. Each tracking target is modeled as a graph node, dynamic edge weights are constructed between each node, and the spatial features of the tracking target in each occluded frame are output. The hierarchical upload module is used to build an action label recognition model, output the danger warning level of each tracked target, and upload data in a hierarchical manner according to the danger warning level.

2. The edge computing-based intelligent supervision system for law enforcement behavior according to claim 1 is characterized in that: An enforcement data sharing platform is constructed based on blockchain technology. A number of blockchain nodes are set up in the enforcement data sharing platform. The blockchain nodes are communicated with law enforcement communication terminals within a preset range and are used to receive law enforcement data packets uploaded by the law enforcement communication terminals. Each blockchain node is linked to each other to form a blockchain network.

3. The edge computing-based intelligent supervision system for law enforcement according to claim 2 is characterized in that: The scene change detection module incrementally updates the prevailing noise level of the video data from the law enforcement communication terminal. The process of determining whether to reclassify the scene based on the incremental update result includes the following: A scene switching detection period is preset, and the prevailing noise level of the video data of the law enforcement communication terminal output by the scene-noise intensity mapping model is marked as a judgment criterion; The end timestamp of the scene switch detection period is used as the detection time point, and the general noise level of the video data at the end timestamp of the previous scene switch detection period is obtained at the detection time point, and the general noise level of the video data at the end timestamp of the previous scene switch detection period is incrementally updated based on each frame of image captured by the law enforcement communication terminal during the current scene switch detection period; The incrementally updated general noise level is compared with the judgment standard to obtain a noise level deviation value. When the noise level deviation value is greater than a preset noise level deviation value threshold, the scene reclassification is triggered.

4. The edge computing-based intelligent supervision system for law enforcement behavior according to claim 3 is characterized in that: The image feature recognition module performs human feature recognition on the first frame of the video data to obtain the spatial distribution features, posture features, appearance features, and identity identifier of the target human body. The process includes: Construct an edge pose estimation model and an appearance feature extraction model, input a sequence of n frames of images into the edge pose estimation model, and output a set of skeleton points of the target human body in each frame of the image according to the edge pose model, wherein the skeleton point set includes the position coordinates of each skeleton point; Obtain the number of bone points in the bone point set of the target human body in the first frame image of the n-frame image sequence, compare the number of bone points with a preset bone point threshold, if the number of bone points is less than the bone point threshold, eliminate the bone point set of the target human body, if the number of bone points is greater than or equal to the bone point threshold, perform bone point distribution analysis on the position coordinates of each bone point of the target human body, and obtain spatial distribution characteristics and posture characteristics of the target human body; Preset constraint threshold intervals corresponding to spatial distribution features and posture features. If both the spatial distribution features and posture features of the target human body are within the corresponding constraint threshold intervals, crop a region of interest from the first frame image based on the skeleton point set, input the region of interest into an appearance feature extraction model, output appearance features of the target human body based on the appearance feature extraction model, generate an identity identifier based on the spatial distribution features, posture features, and appearance features of the target human body, and mark each skeleton point in the skeleton point set of the target human body with the identity identifier. If the spatial distribution characteristics or posture characteristics of the target body are not within the corresponding constraint threshold interval, the skeleton point set of the target body is eliminated.

5. The edge computing-based intelligent supervision system for law enforcement behavior according to claim 4 is characterized in that: The target tracking module matches human features on subsequent frames of video data and marks occluded and visible frames. The process includes: Obtaining spatial distribution features, posture features, and appearance features of a target human body in each frame image subsequent to the first frame in the n-frame image sequence, comparing the spatial distribution features, posture features, and appearance features of the target human body in each subsequent frame image with the spatial distribution features, posture features, and appearance features of the target human body in the first frame image for similarity, and obtaining a feature matching degree between the target human body in each subsequent frame image and the target human body in the first frame; A feature matching threshold is preset, and if the feature matching degree of the target human body in the subsequent frame image and the target human body in the first frame is greater than the feature matching threshold, then an identity identifier is marked on each skeleton point in the skeleton point set of the target human body in the subsequent frame image according to the identity identifier of the target human body in the first frame; The number of occluded frames ky is preset. If there is no target human body whose feature matching degree with the target human body in the first frame is greater than the feature matching degree threshold in the subsequent continuous ky frame images, a dynamic graph convolution target tracking operation is performed, and the frames without the target human body whose feature matching degree with the target human body in the first frame is greater than the feature matching degree threshold are marked as occluded frames, and the frames other than the occluded frames in the n-frame image sequence are marked as visible frames.

6. The edge computing-based intelligent supervision system for law enforcement behavior according to claim 5 is characterized in that: The target tracking module marks the target, builds a dynamic graph convolutional network, models each target as a graph node, constructs dynamic edge weights between nodes, and outputs the spatial features of the target in each occluded frame. The process includes: Mark the target human body in the first frame of an n-frame image sequence as a tracking target, mark the position coordinates of each skeletal point of the tracking target in each visible frame of the n-frame image sequence as spatial features, perform motion feature analysis on the spatial feature sequence of the tracking target in each visible frame to obtain motion features, use the spatial features and motion features of each tracking target in each visible frame as nodes, and connect each node to construct a spatiotemporal graph; A threshold for the number of occluded frames is preset, and the number of occluded frames in the n-frame image sequence is obtained. If the number of occluded frames is less than or equal to the threshold, the dynamic edge weights between each node are obtained based on the spatial features, motion features, and law enforcement scenes of each tracked target in each visible frame. If the number of occluded frames is greater than the threshold, the dynamic edge weights between each node are obtained based on the motion features and law enforcement scenes of each tracked target in each visible frame. A spatial feature prediction model is constructed based on a graph convolutional neural network to learn and represent the spatiotemporal graph and output the spatial features of the tracked target in each occluded frame.

7. The edge computing-based intelligent supervision system for law enforcement according to claim 6 is characterized in that: The hierarchical upload module builds an action label recognition model and outputs the danger warning level of each tracked target. The process of uploading data in a hierarchical manner according to the danger warning level includes: Construct an action label recognition model, input the spatial features of the tracking target in each visible frame and the occluded frame into the action label recognition model, output the action label of each tracking target based on the action label recognition model, and set the danger warning level of each tracking target based on the action label and identity identifier of each tracking target; A preset danger warning level threshold is set. If the danger warning level of the target being tracked is greater than or equal to the preset danger warning level threshold, the danger warning level, spatial distribution characteristics, posture characteristics, and appearance characteristics of the target being tracked are packaged into an enforcement data packet and encrypted and uploaded to the blockchain node. A scheduled upload cycle is preset, and the end timestamp of the scheduled upload cycle is used as the upload time point. If the danger warning level of the tracked target is lower than the preset danger warning level threshold, the danger warning level, spatial distribution characteristics, posture characteristics and appearance characteristics of the tracked target are packaged into a law enforcement data packet and encrypted and uploaded to the blockchain node at the upload time point.

8. The method for intelligent supervision of law enforcement based on edge computing is specifically applied to the intelligent supervision system for law enforcement based on edge computing according to any one of claims 1 to 7, characterized in that: The following steps are involved: Step s1: Collect multimodal data from each law enforcement communication terminal, construct a scene-noise intensity mapping model, obtain the law enforcement scene of the law enforcement communication terminal and the general noise level of the video data, and perform adaptive denoising on the video data of the law enforcement communication terminal; Step s2: incrementally updating the general noise level of the video data of the law enforcement communication terminal, and determining whether to reclassify the scene based on the incremental update result; Step s3: performing human feature recognition on the first frame of the video data to obtain the spatial distribution features, posture features, appearance features and identity identifier of the target human body; Step s4: Perform human feature matching and mark occluded and visible frames on subsequent frames of the video data. Based on the marking results, a dynamic graph convolutional network is constructed to mark the tracking targets. Each tracking target is modeled as a graph node, dynamic edge weights between each node are constructed, and the spatial features of the tracking target in each occluded frame are output. Step s5: Build an action label recognition model, output the danger warning level of each tracked target, and upload data in a graded manner according to the danger warning level.

Citation Information

Patent Citations

  • A method and system for artificial intelligence analysis of public security law enforcement videos

    CN111263094B

  • Audio processing method for law enforcement recorder based on speech analysis and law enforcement recorder

    CN118588094B

  • Fractal-wavelet self-adaption image denoising method based on multivariate statistical model

    CN104778670A

  • Skeleton detection and fall detection method based on improved space-time adaptive graph convolution

    CN117372844A