Intelligent video monitoring management method and system based on big data

By employing a big data-based intelligent video surveillance management method, combining VIBE and DECA models to generate 3D human and facial meshes, and using Z-Standard and Brotli compression algorithms to optimize data compression, multi-level anonymized videos are generated through encryption. This solves the problems of robustness, real-time performance, and privacy protection in complex scenarios for intelligent video surveillance technology, achieving efficient 3D modeling and privacy protection.

CN120976831AActive Publication Date: 2025-11-18JIANGSU FANTAXI TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202511109841.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-08
Publication Date
2025-11-18
Estimated Expiration
2045-08-08

AI Technical Summary

Technical Problem

Existing intelligent video surveillance technologies have significant problems in terms of robustness, real-time performance, and the balance between privacy protection and semantic information preservation in complex scenarios. In particular, the accuracy of 3D modeling is insufficient in scenarios with occlusion, low light, or multiple people interacting. Dynamic region segmentation and semantic feature extraction are not efficient enough, and data compression and encryption technologies are difficult to meet the real-time requirements of edge devices.

Method used

By preprocessing the raw video stream, 2D key points and confidence scores of the human body and face are extracted. 3D human body and face meshes are generated using VIBE and DECA models. Data compression is optimized by combining Z-Standard and Brotli compression algorithms. Multi-level anonymized videos are generated using XChaCha20 encryption to adapt to different privacy and resource requirements.

Benefits of technology

It improves the accuracy of key point detection in occluded and low-light scenes, optimizes the extraction of semantic features in moving areas, reduces computational and storage overhead, ensures high-precision 3D human body modeling and privacy protection, and is highly adaptable to intelligent video surveillance in complex scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120976831A_ABST
    Figure CN120976831A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent video monitoring management method and system based on big data, and relates to the technical field of monitoring videos, and the method comprises the steps: using a VIBE model to generate SMPL-X model parameters, obtaining 3D human body grid head vertex coordinates, using a DECA model to generate FLAME parameters, obtaining 3D face grid vertex coordinates, and fusing the human body grid vertex coordinates and the face grid vertex coordinates. By introducing a traditional image processing and deep learning combined key point extraction method, the accuracy of key point detection in a shielding and weak light scene is improved, and the robustness of 3D human body modeling is enhanced; semantic feature extraction of a motion area is optimized by using a dynamic area mask and a YOLACTT + + model, a high-precision 3D human body model is generated through fixed point coordinate fusion optimization, and encryption and multi-level LoP video output are combined, so that high-privacy and resource-limited scenes are flexibly adapted, and the adaptability and practicability of a monitoring system are improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of monitoring video, in particular to an intelligent video monitoring management method and system based on big data. BACKGROUND

[0002] In recent years, with the rapid development of computer vision and artificial intelligence technology, intelligent video monitoring systems have been widely applied in public safety, urban management, and commercial retail fields. However, there are still some deficiencies in the existing technology. Traditional human pose estimation and 3D modeling methods lack robustness in complex scenes (such as occlusion, weak light, or multi-person interaction), and the key point detection accuracy decreases, leading to distortion in 3D mesh reconstruction. Secondly, existing methods lack efficient motion region segmentation and semantic feature extraction mechanisms when processing dynamic regions, making it difficult to achieve fine-grained anonymization in high-privacy scenarios. In addition, the fusion process of 3D human and face meshes usually relies on complex parameter optimization, which has high computational overhead and is difficult to meet the real-time requirements of edge devices. Although the application of data compression and encryption technology improves transmission efficiency, the balance between semantic information preservation and privacy protection still needs to be optimized. Existing monitoring systems have difficulty in flexibly adapting to different privacy levels and resource constraints in different scenarios, limiting their promotion in diversified application scenarios. The existing intelligent video monitoring technology has significant problems in robustness, real-time performance, and the balance between privacy protection and semantic information preservation in complex scenarios. The present application provides an intelligent video monitoring management method based on big data, which solves the accuracy and efficiency problems of 3D human modeling in complex scenarios, optimizes the semantic feature extraction and data compression and encryption process of dynamic regions, and realizes multi-level anonymized video output to adapt to different privacy and resource requirements. The present application belongs to the technical fields of computer vision, 3D modeling, and privacy protection, and combines deep learning, image processing, and data encryption technology to provide an innovative solution for intelligent video monitoring. SUMMARY

[0003] In view of the above-mentioned existing problems, the present application is proposed.

[0004] Therefore, the present application provides an intelligent video monitoring management method and system based on big data, which solves the significant problems of existing intelligent video monitoring technology in robustness, real-time performance, and the balance between privacy protection and semantic information preservation in complex scenarios.

[0005] To solve the above technical problems, the present application provides the following technical solutions:

[0006] In the first aspect, the present application provides an intelligent video monitoring management method based on big data, which includes collecting raw video streams and preprocessing to generate standardized low-resolution images and human bounding boxes.

[0007] Based on the standardized image and the bounding box, the 2D key point coordinates and the confidence of the human body and the face are extracted by a morphological refinement algorithm and an MTCNN model, the average confidence of the face key points is calculated, and the face key points and the confidence are dynamically updated;

[0008] Based on the human key points and the confidence, the SMPL-X model parameters are generated by using a VIBE model, the 3D human mesh head vertex coordinates are obtained, based on the dynamically updated face key points and the confidence, the FLAME parameters are generated by using a DECA model, the 3D face mesh vertex coordinates are obtained, the human and face mesh vertex coordinates are fused, and the SMPL-X and FLAME parameters are compressed;

[0009] The compressed dynamic area mask is decompressed by using a Z-Standard decoding algorithm, and the 3D positioning and the semantic feature vector of the object are generated based on the dynamic area mask;

[0010] The semantic data compression and encryption data packet is generated based on the semantic feature vector;

[0011] The encrypted data packet is decompressed, the 3D model parameters and the semantic feature vector of the object are fused, and the anonymized video with semantic annotation is generated.

[0012] As a preferred scheme of the intelligent video monitoring management method based on big data, wherein: the standardized low-resolution image and the human body bounding box are generated by collecting the original video stream and pre-processing, the original frame I t is denoised to obtain a smooth frame I' t , the pixel difference between the current frame I t and the previous frame I t-1 is calculated by using an inter-frame difference method, if the total difference value is lower than a threshold k, it is marked as a static frame, only the timestamp is recorded, and the subsequent processing is skipped, if the total difference value is higher than or equal to the threshold k, it is marked as a non-static frame, I' t is scaled to 512*512 resolution to generate a standardized image I" t , the human body is detected on the standardized image I" t by using a YOLOv5s model, the human body bounding box coordinates and the confidence c b are output, the pixel difference between the current frame standardized image I" t and the previous frame I" t-1 is calculated, after the difference image is generated, the threshold A is applied to the difference image for binaryzation to generate a binary mask E t , the region with a pixel difference value greater than or equal to the threshold A is marked as 1, indicating a motion region, and the region with a pixel difference value less than the threshold A is marked as 0, indicating a static region.

[0013] As a preferred embodiment of the big data-based intelligent video surveillance management method of the present invention, the method involves: extracting 2D key point coordinates and confidence scores of the human body and face based on standardized images and bounding boxes using morphological thinning algorithms and the MTCNN model; calculating the average confidence score of facial key points; and dynamically updating the facial key points and confidence scores. The standardized image is then processed using a morphological thinning algorithm. t China B t For the corresponding region, extract the human skeleton image to represent the human centerline structure. Use Hough transform to detect straight line segments in the skeleton image to identify limb parts. Locate 25 key points through geometric constraints, calculate the confidence score for each key point, and generate the key point position x. t and confidence vector c t According to B t Trim the human body ROI region R t ;

[0014] For the trimmed R t 68 facial landmarks were extracted using the MTCNN model to generate facial landmarks. and confidence vector Calculate the average confidence score of facial key points like If the threshold Q is greater than or equal to the current frame's threshold, then the current frame's threshold is used directly. As output, input to the DECA model, if <Threshold Q, extract dynamic region mask E t For regions with a median value of 1, rerun MTCNN within the masked dynamic region to obtain new facial key points. and confidence level Use new As input to the DECA model, if new If the result is greater than or equal to the threshold Q, use the new result; otherwise, use the facial landmarks and confidence scores from the previous frame.

[0015] As a preferred embodiment of the intelligent video surveillance management method based on big data described in this invention, the following steps are performed: Based on human key points and confidence levels, SMPL-X model parameters are generated using the VIBE model to obtain the 3D human head mesh vertex coordinates; based on dynamically updated facial key points and confidence levels, FLAME parameters are generated using the DECA model to obtain the 3D face mesh vertex coordinates; the vertex coordinates of the human and face meshes are fused; and the SMPL-X and FLAME parameters are compressed. The VIBE model is then used based on human key points x... t and confidence level c t Generate the keypoint rotation parameters θ of the SMPL-X model. t Shape parameter β t Translation parameter T tThe keypoint parameters are input into the SMPL-X model, and the vertex positions of the 3D mesh are calculated based on the forward kinematics function, outputting the 3D human body mesh vertex U. t For rotation parameter θ t Perform time series optimization and calculate smoothing rotation parameters. Extracting fixed points v of the head region from the vertices of a 3D human body mesh. i Coordinates, based on the predefined head vertex index i of the SMPL-X model, generate a fixed point coordinate set, denoted as V. SMPL-X ;

[0016] Using the DECA model, based on 68 dynamically updated facial landmarks using MTCNN. and confidence level Generate FLAME model parameters, including facial pose. Facial shape Facial expression parameter k t The FLAME model outputs the 3D mesh of the face, specifically the Y-axis. t Update facial shape parameters Based on 68 facial key points output by the MTCNN model, a subset of 3D mesh vertices corresponding to the 68 key points is obtained using the predefined topology of the FLAME model. Through the key point-to-vertex index mapping provided by FLAME, core vertices of key facial regions are selected. The weighted average of the 3D coordinates of the core vertices is calculated to generate the facial vertex coordinates u. fixed ;

[0017] By using the predefined SMPL-X head vertex index and FLAME model face vertex mapping, a V... SMPL-X The correspondence between the head vertices and the FLAME mesh vertices is determined using the iterative nearest-point algorithm. The rigid transformation matrix of the FLAME mesh vertices is calculated, such that u... fixed Facial vertex and V SMPL-X The head vertices are aligned, and the aligned FLAME mesh vertices and SMPL-X head mesh vertices are weighted and fused to generate a new 3D mesh vertex V. new The merged head vertex V new The body mesh vertices are merged with those of SMPL-X, and the Laplacian smoothing algorithm is applied to optimize the mesh geometric continuity, generating a unified 3D human body model. Differential coding is applied to the SMPL-X and FLAME parameters to record inter-frame parameter changes, and the dynamic region mask E is used. t Flattened into a 1D binary sequence, the input is the Z-Standard compression algorithm. The binary sequence is compressed using dictionary compression and entropy coding mechanisms to generate compressed mask data. Huffman coding is then used to perform secondary compression on the differencing parameters and the Z-Standard compressed mask to generate smaller anonymized data packets.

[0018] As a preferred embodiment of the big data-based intelligent video surveillance management method of the present invention, the following steps are included: Decompressing the dynamic region mask using the Z-Standard decoding algorithm, generating 3D positioning and semantic feature vectors of objects based on the dynamic region mask, and using the Z-Standard decoding algorithm to decompress the dynamic region mask, recovering the binary mask, applying the YOLACT++ model only to regions with a dynamic region mask value of 1 on the standardized image, generating object masks, object categories, and confidence levels, filtering objects with confidence levels greater than or equal to a threshold G, obtaining the center point coordinates of the object mask, and using the Monodepth2 model on the standardized image I. t Perform depth calculations to generate a depth map D. t Extract the depth value of the center point of the human body, and the Y-axis of the human body 3D mesh vertex. t Using a coordinate system, the center point of the object is mapped from 2D to 3D space, generating the object's 3D positioning coordinates. A PointNet++ model is then used to perform semantic analysis on the pixel point cloud of the object mask, generating the object's semantic feature vector. This is achieved from the vertices U of the human body 3D mesh. t Extract the vertices of the hand and combine them with the 3D mesh vertices of the face (Y). t Generate local 3D scene features.

[0019] As a preferred embodiment of the big data-based intelligent video surveillance management method of the present invention, wherein: the generation of semantic data compression and encryption data packets based on semantic feature vectors refers to using the MeshLab algorithm to process new 3D mesh vertices V new Vertex simplification is performed to generate compressed mesh data. The object mask is flattened into a 1D bit string, and a compressed object mask is generated using the Brotli compression algorithm. The object's 3D positioning coordinates, local 3D scene features, and semantic feature vectors are quantized into 16-bit floating-point numbers to form a semantic parameter vector. The semantic parameter vector, compressed object mask, and object category label are packaged into Protobuf format data. The Protobuf data is encrypted using the XChaCha20 encryption algorithm to generate an encrypted data packet. The encrypted data packet is transmitted to the cloud server via a 5G network using the gRPC protocol.

[0020] As a preferred scheme of the intelligent video monitoring management method based on big data, the encrypted data packet is decrypted using an XChaCha20 decryption algorithm, the semantic parameter vector, the compressed object mask, the object category label, the object 3D positioning coordinate, and the semantic feature vector are restored, the object mask is restored using a Brotli decoding algorithm, the object 3D positioning coordinate and the semantic feature vector are mapped to a 3D scene based on a unified 3D human body model, the semantic scene representation is generated by embedding a face grid into a human body grid, and a LoP4-level anonymous video is generated using an UnrealEngine renderer combined with the 3D human body model, the object mask, the object category label, and the semantic feature, wherein the object category and the semantic description are labeled in the video, the 3D scene data based on LoP4 is used to generate a low-level LoP video, LoP1 is a joint line of the extracted 3D human body model, which is rendered as a 2D skeleton line, LoP2 is the object contour displayed by superimposing the object mask on the basis of LoP1, LoP3 is the 3D human body and object grid with removed texture and semantic annotation, and the output level is selected according to the monitoring scene requirement: LoP1 is output for a high-privacy scene, LoP2 or LoP3 is output for a medium-privacy scene, and LoP1 or LoP2 is preferred for a real-time and storage-limited scene.

[0021] In a second aspect, the present application provides an intelligent video monitoring management system based on big data, comprising a video preprocessing and human body detection module for collecting raw video streams and performing preprocessing to generate standardized low-resolution images and human body bounding boxes.

[0022] A key point extraction and dynamic updating module is used to extract human body and face 2D key points and confidence based on the standardized images and the bounding boxes.

[0023] A 3D modeling and parameter compression module is used to generate SMPL-X and FLAME parameters based on the 2D key points, output 3D human body and face grid vertices, and compress the parameters.

[0024] An object positioning and semantic analysis module is used to generate object 3D positioning and semantic features based on dynamic area masks and 3D grid vertices.

[0025] A data compression and encrypted transmission module is used to compress 3D grid vertices and semantic data and transmit them to the cloud in an encrypted manner.

[0026] In a third aspect, the present application provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, any step of the intelligent video monitoring management method based on big data according to the first aspect of the present application is implemented.

[0027] In a fourth aspect, the present application provides a computer readable storage medium having stored thereon a computer program, wherein the computer program, when executed by a processor, implements any step of the big data-based intelligent video monitoring management method according to the first aspect of the present application.

[0028] The present application has the following beneficial effects: by introducing a key point extraction method combining traditional image processing and deep learning, the accuracy of key point detection in occlusion and low-light scenes is effectively improved, and the robustness of 3D human modeling is enhanced; by using dynamic region masks and YOLACT++ model to optimize the semantic feature extraction of the motion region, combining Z-Standard and Brotli compression algorithm, the calculation and storage overhead is significantly reduced, and the real-time processing capability of the edge device is improved; by fusing and optimizing the fixed point coordinates of the SMPL-X and FLAME models, a high-precision unified 3D human model is generated, and combined with XChaCha20 encryption and multi-level LoP video output, it is flexible to adapt to high-privacy and resource-constrained scenarios, ensuring the balance between semantic information preservation and privacy protection, significantly improving the adaptability and practicality of the monitoring system, and providing an efficient and reliable solution for intelligent video monitoring in complex scenarios. BRIEF DESCRIPTION OF DRAWINGS

[0029] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the embodiment description. Obviously, the drawings in the following description are only some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained without creative labor on the basis of these drawings.

[0030] Fig. 1 The flowchart of the big data-based intelligent video monitoring management method in embodiment 1.

[0031] Fig. 2 The structural schematic diagram of the big data-based intelligent video monitoring management system in embodiment 1.

[0032] Fig. 3 The flowchart of generating multi-level anonymous video of the big data-based intelligent video monitoring management method in embodiment 1. DETAILED DESCRIPTION

[0033] In order to make the above-mentioned purposes, features and advantages of the present application more obvious and easy to understand, the specific embodiments of the present application will be described in detail below with reference to the drawings of the specification.

[0034] In the following description, numerous specific details are set forth in order to provide a thorough understanding of the present application. However, it will be apparent to one skilled in the art that the present application can be practiced without the specific details set forth in this description. In other instances, well-known methods, procedures, components, and circuits have not been described in detail so as not to unnecessarily obscure aspects of the present application.

[0035] Secondly, the "one embodiment" or "an embodiment" referred to herein means a specific feature, structure, or characteristic under discussion. Thus, "one embodiment" does not mean a single embodiment nor is it to be taken individually or selectively from other embodiments.

[0036] Embodiment 1, Reference Figs. 1-3 For the first embodiment of the present application, the embodiment provides a big data-based intelligent video monitoring management method, including the following steps:

[0037] S1, generating standardized low-resolution images and human body bounding boxes by collecting raw video streams and preprocessing;

[0038] Based on the standardized images and bounding boxes, the OpenPose model and the MTCNN model are applied to extract the 2D key point coordinates and confidence of the human body and face, calculate the average confidence of the face key points, and dynamically update the face key points and confidence;

[0039] Specifically, by collecting raw video streams and preprocessing, standardized low-resolution images and human body bounding boxes are generated, which means using Gaussian filtering (standard deviation 1.5) on the original frame I t (Obtaining the original video frame I t from the high-definition camera of the edge device) to remove noise and obtain the smooth frame I' t to reduce the interference of noise on subsequent detection, calculate the pixel difference between the current frame I t and the previous frame I t-1 , if the difference sum is lower than the threshold value k (set by empirical parameter setting method), mark it as a static frame, only record the timestamp, skip the subsequent processing, if the difference sum is higher than or equal to the threshold value k, mark it as a non-static frame, scale I' t to 512x512 resolution to generate standardized image I" t , record the scaling ratio (1080 / 512) for coordinate restoration, use the YOLOv5s model to detect the human body on the standardized image I" t , output the human body bounding box coordinates B t ={x,y,w,h} (coordinates (x,y), width w, height h) and confidence c b(b is the identification of the bounding box), locate the rough area of the human body in the image (rectangular box), do not extract specific key points, only provide position information;

[0040] Calculate the current frame standardization image I" t ∈R 512×512 The pixel difference with the previous frame I" t-1 After generating the difference image, apply threshold A (empirical value, obtained by empirical parameter tuning method, such as 0.1) to the difference image to generate a 512x512 binary mask E t The area with pixel difference greater than or equal to threshold A is marked as 1, indicating the moving area, and the area with pixel difference less than threshold A is marked as 0, indicating the static area.

[0041] By applying Gaussian filter (standard deviation 1.5) to the original video frame for denoising, smooth frame is generated, which effectively reduces the interference of noise on subsequent detection and improves the image processing stability in complex scenes (such as weak light or occlusion). The threshold value obtained by empirical parameter tuning is combined with the inter-frame difference method to accurately distinguish static frames and non-static frames, significantly reducing the computational load, optimizing the real-time performance of edge devices, using YOLOv5s model to efficiently detect human body bounding box, only output rough position information, reducing the computational overhead of key point extraction, and generating 512x512 binary mask through inter-frame difference and threshold A, accurate segmentation of dynamic area, providing a reliable foundation for subsequent semantic feature extraction. These improvements not only improve the robustness and real-time processing capability of the system in complex scenes, but also lay the foundation for efficient 3D modeling and privacy protection, and adapt to diversified monitoring needs.

[0042] Further, based on the standardization image and the bounding box, the 2D key point coordinates and confidence of the human body and face are extracted through the morphological refinement algorithm and MTCNN model, the average confidence of the face key points is calculated, and the face key points and confidence are dynamically updated. The face key points and confidence are dynamically updated by processing the standardization image I" t B tFor the corresponding region, the human skeleton image is extracted, representing the human centerline structure, the straight line segment in the skeleton image is detected using Hough transform, the limb part is identified, and 25 key points are located according to the COCO key point dataset standard (including nose, left and right eyes, left and right ears, left and right shoulders, left and right elbows, left and right wrists, left and right hips, left and right knees, left and right ankles, neck, chest, and pelvis) through geometric constraints (for example, the head key points (nose, eyes, and ears) are located by using a pre-defined head template to match the high-intensity area in the top 1 / 5 region of the bounding box, and the shoulder and neck are determined by the end points of the horizontal skeleton line segment in the upper 1 / 3 of the bounding box and the midpoint of the connecting line between the head and the center of the bounding box), the confidence of each key point is calculated (the gradient value of each key point position is extracted from the edge image generated by Canny edge detection, and is normalized to the range of 0 to 1, reflecting the reliability of the edge intensity, as the first component of the confidence, the length of the skeleton line segment from the key point to its adjacent key point (according to the COCO key point topology definition, such as the shoulder to the elbow) is calculated, and is compared with the expected length based on the human body proportion, if the length is close to the expected length, a higher connectivity score (range 0 to 1) is given as the second component of the confidence, and the key point is checked whether it is located in the dynamic region of the binary mask E t , if in the dynamic region, a fixed bonus of 0.3 is given, otherwise 0, as the third component of the confidence, the three components (edge intensity, connectivity score, dynamic region bonus) are added with weights of 0.4, 0.3, and 0.3 respectively, to obtain the confidence value of each key point, ranging from 0 to 1, ensuring the reliability of the key point and the adaptability of the dynamic scene), the key point position x t and the confidence vector c t are generated.

[0043] According to B t , the human body ROI region is cropped (the coordinates and dimensions are obtained from the bounding box B t : the upper left corner coordinates (x, y), the width w, and the height h), the rectangular region is extracted on the normalized image I t : from (x, y) to (x+w, y+h), if the bounding box exceeds the image range, the edge pixels are filled (using zero padding), and the cropped ROI region is output where H ∈ ≈h, W ∈ ≈w, and ∈ is the index of the cropped region);

[0044] On the cropped R t , the MTCNN model is used to extract 68 facial key points (the MTCNN model follows the Dlib 68-point facial key point standard, covering facial contour, eye, nose, and mouth feature points), and the facial key points and confidence vector are generated:

[0045]

[0046] wherein, is a face key point, is a confidence level;

[0047] calculating a face key point average confidence level (68 face key points and a confidence vector output by the MTCNN are obtained, and an average value is calculated), if ≥ threshold Q, directly using the current frame as the output, inputting the DECA model, if < threshold Q, extracting a dynamic region mask E t The region with a median value of 1 (indicating a motion region) is re-run within the mask dynamic region (E t value is 1) to obtain new face key points and confidence levels The new is used as the input of the DECA model, if the new ≥ threshold Q, the new result is used, otherwise, the face key points and the confidence levels of the previous frame are followed, the face key points and the confidence levels are dynamically updated to avoid detection failure caused by occlusion or low resolution.

[0048] By adopting a morphological thinning algorithm and a Hough transform, a human skeleton is extracted and 25 COCO standard key points are positioned from a standardized image and a bounding box, and by combining geometric constraints and edge intensity analysis, the accuracy and stability of key point detection in an occluded and low resolution scene are significantly improved; the MTCNN model is used to dynamically update 68 face key points, and by confidence weighting and dynamic region mask optimization, the reliability of the face key points in a complex dynamic scene is ensured, and the detection failure problem caused by occlusion or weak light is overcome; the key point confidence level calculation comprehensively considers edge intensity, skeleton connectivity and dynamic region weight, effectively improves the robustness and adaptability of key point positioning, reduces the dependence on a deep learning model, and reduces the computational complexity, not only enhances the accuracy and efficiency of key point extraction before 3D modeling, but also provides high-quality input for subsequent SMPL-X and FLAME model parameter generation, and significantly improves the practicability and performance of an intelligent video monitoring system in a diversified scene.

[0049] S2, based on the human key points and the confidence levels, using a VIBE model to generate SMPL-X model parameters, obtaining 3D human mesh head vertex coordinates, based on the dynamically updated face key points and the confidence levels, using a DECA model to generate FLAME parameters, obtaining 3D face mesh vertex coordinates, fusing the human and face mesh vertex coordinates, and compressing the SMPL-X and FLAME parameters;

[0050] decompressing the compressed dynamic region mask using a Z-Standard decoding algorithm, generating a 3D localization and semantic feature vector of the object based on the dynamic region mask;

[0051] Specifically, based on the human key points and confidence, the SMPL-X model parameters are generated using the VIBE model to obtain the 3D human mesh head vertex coordinates, based on the dynamically updated face key points and confidence, the FLAME parameters are generated using the DECA model to obtain the 3D face mesh core vertex coordinates, the human and face mesh vertex coordinates are fused, and the SMPL-X and FLAME parameters are compressed. Using the VIBE model (pre-trained on Human3.6M and 3DPW datasets) based on human key points x t and confidence c t , generate the key point rotation parameters θ t of the SMPL-X model, shape parameters β t and translation parameters T t , input the key point parameters into the SMPL-X model, calculate the 3D mesh vertex position based on the forward kinematics function, (based on the pre-defined topology structure of the SMPL-X model, obtain the correspondence table of 25 key points (such as nose, left and right eyes, left and right ears, left and right shoulders, left and right elbows, left and right wrists, left and right knees, left and right ankles) of COCO Keypoint dataset and 52 joints (including body 22, hand 20 and face 10) in SMPL-X model, which is provided by SMPL-X model, directly mapping each key point to the corresponding joint or adjacent joint combination, for example, the left shoulder key point of COCO is mapped to the left shoulder joint of SMPL-X, and the left elbow key point is mapped to the left elbow joint, for the 25 key point rotation parameters (local joint angle represented by quaternion) output by the VIBE model, through the mapping table, the rotation parameter of each key point is assigned to the corresponding SMPL-X joint, for the joints not directly corresponding to the key points (such as hand or face fine joints), the interpolation method (based on the rotation parameters of adjacent key points and the joint hierarchy structure of SMPL-X model) is used to estimate the rotation parameters, generate a complete rotation parameter set of 52 joints, through the forward kinematics function of SMPL-X model, the joint rotation parameter (quaternion) is converted into rotation matrix, combined with shape parameter and translation parameter, calculate the 3D human mesh vertex position, ensure that the generated mesh is consistent with the input key points) output 3D human mesh vertex U t , representing the geometric structure of the human body in 3D space, formula:

[0052] U t = SMPL-X(θ t , β t , T t ),

[0053] the rotation parameter θ t Time series optimization is performed, and a smooth rotation parameter is calculated based on the confidence weighting of the key points to ensure motion continuity, formula:

[0054]

[0055] wherein, is the optimized rotation parameter of the jth key point, F is the time window, which is set based on the monitoring scene motion continuity, j=25 is the number of key points of the SMPL-X model (corresponding to the 25 key points of COCO), which is defined by the SMPL-X model, θ key,j (f) is the rotation parameter of the jth key point in the fth frame, c j (f) is the confidence of the jth key point, w m (f) is the dynamic region mask weight of the fth frame, which is generated based on the sparse optical flow algorithm, if the key point j is located in the dynamic region, then w m =1, otherwise w m =0;

[0056] Fixed points v i coordinates (such as the top of the head, the lower jaw, and the ears) are extracted from the 3D human body mesh vertices, based on the head vertex index i predefined by the SMPL-X model, a fixed point coordinate set is generated, denoted as V SMPL-X ={v i ∣v i =(x i ,y i ,z i ,i=1,…,N head}, where N head is the total number of fixed point coordinates, and is used for subsequent fusion optimization;

[0057] The DECA model (pre-trained on the VGGFace2 and CelebA datasets) is used to generate the FLAME model parameters based on the 68 facial key points dynamically updated by MTCNN and the confidence The FLAME model parameters include the facial pose the facial shape and the expression parameter k t The FLAME model outputs the facial 3D mesh vertices Y t (the facial pose parameter the shape parameter and the expression parameter k t are input into the FLAME model, and the 3D mesh vertex position is calculated by the geometric transformation function (a predefined facial template) of the model, which represents the geometric structure of the face in 3D space), formula:

[0058]

[0059] updating the face shape parameters To ensure temporal continuity, the formula is:

[0060]

[0061] wherein, is the optimized face shape parameter of the previous frame, from the last iteration (initial value is the DECA model output), r t is the average confidence of the face key points e is the base of the natural logarithm (about 2.718), λ is the forgetting factor, obtained by experimental parameter tuning method, and u is the frame index in the time window;

[0062] Based on the 68 face key points output by the MTCNN model, the 3D grid vertex subset corresponding to the 68 key points is obtained by using the topological structure predefined by the FLAME model. By using the index mapping of the key points to the vertices provided by the FLAME (the mapping relationship between the key points and the FLAME grid vertices is provided by the FLAME model, and the vertex index corresponding to each key point is directly specified), the core vertices of the key face regions (such as the tip of the nose, the corners of the mouth, and the corners of the eyes) are screened out, and the weighted average value of the 3D coordinates of the core vertices is calculated, and the weight is the confidence of the corresponding MTCNN key point Generate face vertex coordinates u fixed =(x i ,y i ,z i );

[0063] By using the predefined SMPL-X head vertex index and the FLAME model face vertex mapping, the correspondence between the head vertices in V SMPL-X and the FLAME grid vertices is established, and the iterative closest point (ICP) algorithm is used to calculate the rigid transformation matrix H=[R|t] of the FLAME grid vertices, wherein R is a rotation matrix and t is a translation vector, so that the face vertices u fixed are aligned with the head vertices of V SMPL-X , and based on the average confidence of the MTCNN face key points and the confidence of the SMPL-X head key points c t , the fusion weight α is calculated by confidence weighted average, and the aligned FLAME grid vertices and the SMPL-X head grid vertices are weighted and fused to generate new 3D grid vertices V new , the fused head vertices V new are combined with the body grid vertices of the SMPL-X, the Laplacian smoothing algorithm is applied to optimize the grid geometric continuity, and a unified 3D human body model is generated to ensure the seamless connection of the head and body grids;

[0064] V new = a H(u fixed ) + (1-a) V SMPL-X ,

[0065] SMPL-X parameters [θ t , β t , T t ] and FLAME parameters Differential coding is applied to record the inter-frame parameter changes (reduce redundancy), flatten the dynamic region mask E t into a 1D binary sequence, input Z-Standard compression algorithm (configure high compression level, such as level = 9), use dictionary compression and entropy coding mechanism to compress the binary sequence, generate compressed mask data, use Huffman coding to compress the differential parameters and Z-Standard compressed mask for the second time, generate smaller anonymized data packets (<500 bytes).

[0066] In view of the defects of the existing intelligent video monitoring technology in complex scenes, such as insufficient 3D modeling accuracy, high computational overhead, and poor balance between privacy protection and semantic information preservation, by combining the VIBE model and the DECA model, based on 25 COCO standard key points and dynamically updated 68 facial key points, SMPL-X and FLAME model parameters are generated, and high-precision 3D human and facial meshes are efficiently constructed, significantly improving the modeling robustness in occlusion and dynamic scenes; the confidence weighted time series optimization method is used to smooth the key point rotation parameters, ensuring the continuity of the action and reducing the mesh distortion in complex scenes; the ICP algorithm and confidence weighted average fusion SMPL-X head vertex and FLAME face vertex are used to generate a unified 3D human model, realizing seamless connection of head and body mesh, and optimizing the geometric consistency; in addition, by using differential coding combined with Z-Standard and Huffman compression algorithm, the SMPL-X and FLAME parameters and dynamic region mask are compressed into bytes, which greatly reduces the storage and transmission overhead, adapts to the real-time requirements of edge devices, and significantly enhances the 3D modeling accuracy, computational efficiency and privacy protection capability.

[0067] Further, the use of Z-Standard decoding algorithm to decompress the compressed dynamic region mask, based on the dynamic region mask to generate the 3D positioning and semantic feature vector of the object refers to the use of Z-Standard decoding algorithm to decompress the compressed dynamic region mask, restore the 512x512 binary mask, represent the moving area, on the normalized image, only for the region with dynamic region mask value of 1, apply YOLACT++ model to generate object mask (512x512 binary matrix), object category (such as "knife" or "mobile phone") and confidence, filter the objects with confidence greater than or equal to threshold G (obtained by empirical parameter tuning method, based on object detection experiment on COCO dataset and monitoring scene dataset (such as SIAT), statistical confidence distribution of YOLACT++ model output, select threshold G that can effectively distinguish correct detection and false detection), obtain the center point coordinates of object mask, use Monodepth2 model to calculate the depth of standardized image I" t , generate depth map D t , extract the depth value of object center point, the coordinate system of human body 3D grid vertex Y t , map the object center point from 2D (pixel coordinates + depth) to 3D space to generate object 3D positioning coordinates, use PointNet++ model to perform semantic analysis on the pixel point cloud of object mask (extract pixel coordinates (x, y) with value of 1 from object mask, combine the depth value of corresponding pixel in depth map Dt to generate point cloud data (each point contains x, y, z coordinates), input point cloud data into PointNet++ model, the model performs semantic segmentation and feature extraction based on pre-trained weights (trained on ShapeNet dataset)), generate object semantic feature vector (128 dimensions), extract hand vertices (about 500) from human body 3D grid vertex U t , combine face 3D grid vertex Y t , generate local 3D scene feature (64 dimensions), enhance scene semantics.

[0068] The Z-Standard decoding algorithm is used to efficiently decompress the binary dynamic region mask, and the YOLACT++ model is applied only to the motion region for object detection and segmentation. Combined with Monodepth2 depth estimation and PointNet++ point cloud analysis, the object 3D positioning coordinates and 128-dimensional semantic feature vectors are accurately generated, which significantly improves the accuracy and efficiency of object semantic extraction in complex dynamic scenes. The 64-dimensional local scene features are generated by extracting the hand and face vertices from the unified 3D human mesh, which further enhances the richness of scene semantics and reduces the computational overhead of non-dynamic regions. This method effectively overcomes the robustness problem of traditional technology in extracting semantic information under occlusion or multi-person scenes, and optimizes the utilization of computing resources by focusing on dynamic regions, which meets the real-time requirements of edge devices. It provides high-quality input for subsequent generation of semantically annotated anonymized videos, balancing privacy protection and semantic information.

[0069] S3, generating semantic data compression and encryption data packets based on semantic feature vectors;

[0070] Decompressing the encrypted data packet, fusing the 3D model parameters and the object semantic feature vector, and generating a semantically annotated anonymized video;

[0071] Specifically, the semantic feature vector-based semantic data compression and encryption data packet generation refers to using the MeshLab algorithm to perform vertex simplification on the new 3D mesh vertex V new The vertex simplification (reduced to about 50% of the vertices) generates compressed mesh data, the object mask is flattened into a 1-dimensional bit string, the Brotli compression algorithm is used to generate a compressed object mask with a compression ratio of about 1:1500, the object 3D positioning coordinates (3D), local 3D scene features, and semantic feature vectors (128D) are quantized to 16-bit floating-point numbers to form a semantic parameter vector. The semantic parameter vector, compressed object mask, and object class label are packaged into Protobuf format data, and the XChaCha20 encryption algorithm is used to encrypt the Protobuf data to generate an encrypted data packet. The encrypted data packet is transmitted to the cloud server through the 5G network via the gRPC protocol.

[0072] The 3D mesh vertices are simplified by using the MeshLab algorithm, and the object mask is compressed by using the Brotli compression algorithm, which greatly reduces the data storage and transmission overhead and optimizes the resource utilization of the edge device; the 3D positioning coordinates of the object, the local scene features and the 128-dimensional semantic feature vector are quantized by using 16-bit floating point, an efficient semantic parameter vector is generated, and the Protobuf format is packed, and the XChaCha20 encryption algorithm is used to generate an encrypted data packet of less than 500 bytes, which ensures the high security and privacy protection capability of the data; through the gRPC protocol and the 5G network transmission to the cloud server, the real-time performance and reliability of data transmission are further improved; the problems of data redundancy and privacy leakage in complex scenes in traditional technologies are effectively solved, and an efficient and secure data processing and transmission scheme is provided for the intelligent video monitoring system, which significantly enhances the applicability of the intelligent video monitoring system in diversified scenes.

[0073] Further, the encrypted data packet is decrypted, the 3D model parameters and the semantic feature vector are fused, and an anonymized video with semantic annotation is generated. The encrypted data packet is decrypted using the XChaCha20 decryption algorithm, the semantic parameter vector, the compressed object mask, the object class label, the object 3D positioning coordinates, and the semantic feature vector are restored, the 512x512 object mask is restored using the Brotli decoding algorithm, the face grid is embedded into the human body grid based on the unified 3D human body model, the object 3D positioning coordinates and the semantic feature vector are mapped to the 3D scene, and the semantic scene representation (including human body and object position) is generated. Using the UnrealEngine renderer, combined with the 3D human body model, the object mask, the object class label, and the semantic feature, a 512x512x3 LoP4 level anonymous video is generated, with object class labels (such as "knife") and semantic descriptions annotated in the video. Based on the LoP4 3D scene data, low-level LoP videos are generated, LoP1 (skeleton) is the joint line of the extracted 3D human body model, rendered as a 2D skeleton line, LoP2 (skeleton + mask) is the object mask superimposed on LoP1, showing the object outline, LoP3 (mesh) is the 3D human body and object mesh, with textures and semantic annotations removed, and the output level is selected according to the monitoring scene requirements: high privacy scene (such as public area) outputs LoP1, medium privacy scene (such as commercial retail) outputs LoP2 or LoP3, real-time and storage limited scene prioritizes LoP1 or LoP2, reducing rendering and storage overhead, meeting different privacy and resource requirements.

[0074] The semantic parameter vector and object mask are efficiently recovered through XChaCha20 decryption and Brotli decoding algorithms, the face grid is seamlessly embedded into the human body grid combined with the unified 3D human body model, and the object 3D positioning coordinates and semantic feature vector are accurately mapped to the 3D scene, thereby generating a scene representation rich in semantics; the Unreal Engine renderer is used to generate LoP4 level anonymous video, combined with the object category and semantic description, and LoP1 (skeleton), LoP2 (skeleton + mask) and LoP3 (grid) video are generated through degradation, thereby flexibly adapting to the needs of high privacy scenes (such as public areas), medium privacy scenes (such as commercial retail) and real-time / storage limited scenes; this multi-level anonymization output effectively balances privacy protection and semantic information retention, significantly reduces rendering and storage overhead, and improves the robustness and adaptability of the system in complex scenes.

[0075] The embodiment also provides an intelligent video monitoring management system based on big data, which comprises a video preprocessing and human body detection module, a key point extraction and dynamic updating module, a 3D modeling and parameter compression module, an object positioning and semantic analysis module and a data compression and encrypted transmission module.

[0076] The key point extraction and dynamic updating module is used for extracting human body and face 2D key points and confidence based on the standardized image and the boundary box.

[0077] The 3D modeling and parameter compression module is used for generating SMPL-X and FLAME parameters based on the 2D key points, outputting 3D human body and face grid vertices, and compressing the parameters.

[0078] The object positioning and semantic analysis module is used for generating object 3D positioning and semantic features based on the dynamic area mask and the 3D grid vertices.

[0079] The data compression and encrypted transmission module is used for compressing the 3D grid vertices and semantic data and transmitting the compressed data to the cloud in an encrypted manner.

[0080] The embodiment also provides a computer device suitable for the intelligent video monitoring management method based on big data, which comprises a memory and a processor.

[0081] The computer device can be a terminal, and the computer device includes a processor, a memory, a communication interface, a display screen and an input device connected by a system bus. The processor of the computer device is configured to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for running the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is configured to perform wired or wireless communication with an external terminal. The wireless communication can be achieved by WIFI, a carrier network, NFC (Near Field Communication) or other technologies. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer overlaid on the display screen, or a key, a trackball or a touchpad arranged on the shell of the computer device, or an external keyboard, a touchpad or a mouse, etc.

[0082] The embodiment also provides a storage medium having a computer program stored thereon, and the computer program is executed by a processor to implement the method and system for intelligent video monitoring management based on big data according to the above embodiment. The storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as a static random access memory (SRAM), an electrically erasable programmable read-only memory (EEPROM), an erasable programmable read-only memory (EPROM), a programmable read-only memory (PROM), a read-only memory (ROM), a magnetic storage, a flash memory, a magnetic disk or an optical disk.

[0083] To sum up, by introducing the key point extraction method combining traditional image processing and deep learning, the accuracy of key point detection in occlusion and weak light scenes is effectively improved, and the robustness of 3D human modeling is enhanced; the dynamic area mask and YOLACT++ model are used to optimize the semantic feature extraction of the motion area, combined with Z-Standard and Brotli compression algorithm, the calculation and storage overhead are significantly reduced, and the real-time processing capability of the edge device is improved; through the fusion optimization of the fixed point coordinates of the SMPL-X and FLAME models, a high-precision unified 3D human model is generated, combined with XChaCha20 encryption and multi-level LoP video output, flexible adaptation to high privacy and resource limited scenes is realized, the balance between semantic information reservation and privacy protection is ensured, the adaptability and practicability of the monitoring system are significantly improved, and an efficient and reliable solution for intelligent video monitoring in complex scenes is provided.

[0084] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present application and not to limit the present application. Although the present application has been described in detail with reference to the preferred embodiments, it should be understood by those skilled in the art that the technical solutions of the present application can be modified or replaced equivalently without departing from the spirit and scope of the technical solutions of the present application, and they should be covered in the scope of the claims of the present application.

Claims

1. A big data-based intelligent video surveillance management method, characterized by: include, By acquiring the raw video stream and preprocessing it, standardized low-resolution images and human bounding boxes are generated. Based on standardized images and bounding boxes, 2D key point coordinates and their confidence scores of the human body and face are extracted through morphological thinning algorithms and MTCNN models. The average confidence score of facial key points is calculated, and facial key points and confidence scores are dynamically updated. Based on human key points and confidence levels, the VIBE model is used to generate SMPL-X model parameters to obtain the vertex coordinates of the head of the 3D human body mesh. Based on dynamically updated facial key points and confidence levels, the DECA model is used to generate FLAME parameters to obtain the vertex coordinates of the 3D facial mesh. The vertex coordinates of the human body and facial mesh are fused, and the SMPL-X and FLAME parameters are compressed. The Z-Standard decoding algorithm is used to decompress the dynamic region mask, and the 3D localization and semantic feature vector of the object are generated based on the dynamic region mask. Generate semantic data compression and encryption data packets based on semantic feature vectors; Decompress the encrypted data packet, fuse 3D model parameters and object semantic feature vectors, and generate anonymized videos with semantic annotations.

2. The intelligent video surveillance management method based on big data as described in claim 1, characterized in that: The process involves acquiring the original video stream and preprocessing it to generate standardized low-resolution images and human bounding boxes. This refers to using Gaussian filtering on the original frame I. t Denoising is performed to obtain a smooth frame I' t The current frame I is calculated using the inter-frame difference method. t Compared to the previous frame I t-1 If the sum of the pixel differences is less than a threshold k, it is marked as a static frame, only its timestamp is recorded, and subsequent processing is skipped. If the sum of the differences is greater than or equal to the threshold k, it is marked as a non-static frame, and I'... t Scaled to 512×512 resolution to generate a normalized image I” t Using the YOLOv5s model in normalized image I” t Detect the human body and output the human body bounding box coordinates B. t and confidence level c b Calculate the normalized image I of the current frame. t With the previous frame I” t-1 The pixel differences are used to generate a difference image. Then, a threshold A is applied to binarize the difference image to generate a binary mask E. t Regions with pixel differences greater than or equal to threshold A are marked as 1, representing moving regions, while regions with pixel differences less than threshold A are marked as 0, representing static regions.

3. The intelligent video surveillance management method based on big data as described in claim 2, characterized in that: Based on standardized images and bounding boxes, 2D keypoint coordinates and their confidence scores of the human body and face are extracted using morphological thinning algorithms and the MTCNN model. The average confidence score of facial keypoints is calculated, and the facial keypoints and confidence scores are dynamically updated. The standardized images are then processed using morphological thinning algorithms. t China B t For the corresponding region, extract the human skeleton image to represent the human centerline structure. Use Hough transform to detect straight line segments in the skeleton image to identify limb parts. Locate 25 key points through geometric constraints, calculate the confidence score for each key point, and generate the key point position x. t and confidence vector c t According to B t Trim the human body ROI region R t ; For the trimmed R t 68 facial landmarks were extracted using the MTCNN model to generate facial landmarks. and confidence vector Calculate the average confidence score of facial key points like Use the current frame directly As output, input to the DECA model, if Extracting the dynamic region mask E t For regions with a median value of 1, rerun MTCNN within the masked dynamic region to obtain new facial key points. and confidence level Use new As input to the DECA model, if new Use the new results; otherwise, use the facial landmarks and confidence scores from the previous frame.

4. The intelligent video surveillance management method based on big data as described in claim 3, characterized in that: The process involves using a VIBE model to generate SMPL-X model parameters based on human keypoints and confidence levels, obtaining the 3D human head vertex coordinates. Then, based on dynamically updated facial keypoints and confidence levels, a DECA model is used to generate FLAME parameters, obtaining the 3D face vertex coordinates. The vertex coordinates of the human and face meshes are then fused, and the SMPL-X and FLAME parameters are compressed. The VIBE model is used based on human keypoints x... t and confidence level c t Generate the keypoint rotation parameters θ of the SMPL-X model. t Shape parameter β t Translation parameter T t The keypoint parameters are input into the SMPL-X model, and the vertex positions of the 3D mesh are calculated based on the forward kinematics function, outputting the 3D human body mesh vertex U. t For rotation parameter θ t Perform time series optimization and calculate smoothing rotation parameters. Extracting fixed points v of the head region from the vertices of a 3D human body mesh. i Coordinates, based on the predefined head vertex index i of the SMPL-X model, generate a fixed point coordinate set, denoted as V. SMPL-X ; Using the DECA model, based on 68 dynamically updated facial landmarks using MTCNN. and confidence level Generate FLAME model parameters, including facial pose. Facial shape Facial expression parameter k t The FLAME model outputs the 3D mesh of the face, specifically the Y-axis. t Update facial shape parameters Based on 68 facial key points output by the MTCNN model, a subset of 3D mesh vertices corresponding to the 68 key points is obtained using the predefined topology of the FLAME model. Through the key point-to-vertex index mapping provided by FLAME, core vertices of key facial regions are selected. The weighted average of the 3D coordinates of the core vertices is calculated to generate the facial vertex coordinates u. fixed ; By using the predefined SMPL-X head vertex index and FLAME model face vertex mapping, a V... SMPL-X The correspondence between the head vertices and the FLAME mesh vertices is determined using the iterative nearest-point algorithm. The rigid transformation matrix of the FLAME mesh vertices is calculated, such that u... fixed Facial vertex and V SMPL-X The head vertices are aligned, and the aligned FLAME mesh vertices and SMPL-X head mesh vertices are weighted and fused to generate a new 3D mesh vertex V. new The merged head vertex V new The body mesh vertices are merged with those of SMPL-X, and the Laplacian smoothing algorithm is applied to optimize the mesh geometric continuity, generating a unified 3D human body model. Differential coding is applied to the SMPL-X and FLAME parameters to record inter-frame parameter changes, and the dynamic region mask E is used. t Flattened into a 1D binary sequence, the input is the Z-Standard compression algorithm. The binary sequence is compressed using dictionary compression and entropy coding mechanisms to generate compressed mask data. Huffman coding is then used to perform secondary compression on the differencing parameters and the Z-Standard compressed mask to generate smaller anonymized data packets.

5. The intelligent video surveillance management method based on big data as described in claim 4, characterized in that: The process involves using the Z-Standard decoding algorithm to decompress the dynamic region mask, generating 3D object localization and semantic feature vectors based on the dynamic region mask. This involves using the Z-Standard decoding algorithm to decompress the dynamic region mask, recovering the binary mask, applying the YOLACT++ model only to regions with a dynamic region mask value of 1 on the standardized image, generating object masks, object categories, and confidence scores, filtering objects with confidence scores greater than or equal to a threshold G, obtaining the center point coordinates of the object mask, and using the Monodepth2 model on the standardized image I. t Perform depth calculations to generate a depth map D. t Extract the depth value of the center point of the human body, and the Y-axis of the human body 3D mesh vertex. t Using a coordinate system, the center point of the object is mapped from 2D to 3D space, generating the object's 3D positioning coordinates. A PointNet++ model is then used to perform semantic analysis on the pixel point cloud of the object mask, generating the object's semantic feature vector. This is achieved from the vertices U of the human body 3D mesh. t Extract the vertices of the hand and combine them with the 3D mesh vertices of the face (Y). t Generate local 3D scene features.

6. The intelligent video surveillance management method based on big data as described in claim 5, characterized in that: The generation of semantic data compression and encryption data packets based on semantic feature vectors refers to using the MeshLab algorithm to process new 3D mesh vertices V. new Vertex simplification is performed to generate compressed mesh data. The object mask is flattened into a 1D bit string, and a compressed object mask is generated using the Brotli compression algorithm. The object's 3D positioning coordinates, local 3D scene features, and semantic feature vectors are quantized into 16-bit floating-point numbers to form a semantic parameter vector. The semantic parameter vector, compressed object mask, and object category label are packaged into Protobuf format data. The Protobuf data is encrypted using the XChaCha20 encryption algorithm to generate an encrypted data packet. The encrypted data packet is transmitted to the cloud server via a 5G network using the gRPC protocol.

7. The intelligent video surveillance management method based on big data as described in claim 6, characterized in that: The decompressed encrypted data packet is fused with 3D model parameters and semantic feature vectors to generate an anonymized video representation with semantic annotations. The encrypted data packet is then decrypted using the XChaCha20 decryption algorithm to recover the semantic parameter vectors, compressed object mask, object category labels, object 3D location coordinates, and semantic feature vectors. The object mask is then restored using the Brotli decoding algorithm. Based on a unified 3D human body model, a facial mesh is embedded into the human body mesh. The object 3D location coordinates and semantic feature vectors are mapped to the 3D scene to generate a semantic scene representation. This representation is then processed using Unreal Engine. The Engine renderer combines 3D human models, object masks, object category labels, and semantic features to generate LoP4 level anonymized videos. The videos are labeled with object categories and semantic descriptions. Based on the 3D scene data of LoP4, lower-level LoP videos are generated. LoP1 extracts the joint lines of the 3D human model and renders them as 2D skeleton lines. LoP2 overlays object masks on top of LoP1 to display the object outlines. LoP3 retains the 3D human and object meshes but removes textures and semantic annotations. The output level is selected according to the monitoring scene requirements: high privacy scenes output LoP1, medium privacy scenes output LoP2 or LoP3, and real-time and storage-constrained scenes prioritize LoP1 or LoP2.

8. A big data-based intelligent video surveillance management system, based on the big data-based intelligent video surveillance management method according to any one of claims 1 to 7, characterized in that: This includes a video preprocessing and human detection module, which is used to acquire raw video streams and preprocess them to generate standardized low-resolution images and human bounding boxes; The key point extraction and dynamic update module is used to extract 2D key points and confidence scores of the human body and face based on standardized images and bounding boxes. The 3D modeling and parameter compression module is used to generate SMPL-X and FLAME parameters based on 2D key points, output 3D human body and face mesh vertices, and compress parameters; The object localization and semantic analysis module is used to generate 3D localization and semantic features of objects based on dynamic region masks and 3D mesh vertices; The data compression and encrypted transmission module is used to compress 3D mesh vertices and semantic data and encrypt them for transmission to the cloud.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that: When the processor executes the computer program, it implements the steps of the intelligent video surveillance management method based on big data as described in any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the processor, it implements the steps of the intelligent video surveillance management method based on big data as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • 3D human body posture estimation method based on complementary enhancement of key points and grid vertexes

    CN116631064A

  • Grid recovery method and system based on feature fusion of vision-language large model

    CN119007184A

  • Anti-collision monitoring system based on depth estimation and instance segmentation fused three-dimensional model

    CN120356173A