Internet of Things video monitoring big data privacy protection and efficient retrieval method based on artificial intelligence
By constructing an adaptive acquisition architecture for multi-source heterogeneous video data and an improved YOLO object detection system, combined with graph-cut technology and dynamic hash functions, the system solves the problems of heterogeneous protocols and privacy protection in multiple scenarios of IoT video surveillance systems, and achieves efficient privacy data retrieval and system stability.
Patent Information
- Application Number
- CN202511657386.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-13
- Publication Date
- 2026-02-10
AI Technical Summary
IoT video surveillance systems face challenges such as heterogeneous protocols across multiple scenarios, low data transmission efficiency, insufficient privacy protection, low retrieval efficiency, and poor system stability. In particular, they struggle to meet the privacy data protection requirements of different sensitivity levels in scenarios such as vehicle-mounted, park, and battery swapping stations.
An AI-based approach is used to construct an adaptive acquisition architecture for multi-source heterogeneous video data. An improved YOLO object detection and Graph-Cut technique are combined to extract and encrypt privacy information. A multi-dimensional privacy level assessment model is constructed, and a dynamic hash function and adaptive indexing mechanism are used to achieve efficient retrieval and real-time monitoring.
It achieves adaptive protocol switching in multiple scenarios, maintains video quality while compressing data at a ratio of 1:1.8, accurately filters privacy targets, shortens search response time, improves search result matching accuracy, and ensures system stability and security through real-time monitoring and adaptive optimization.
Smart Images

Figure FT_1 
Figure FT_2 
Figure FT_3
Abstract
Description
Technical Field
[0001] This invention belongs to the field of interdisciplinary technology of artificial intelligence, Internet of Things and information security. Specifically, it relates to an artificial intelligence-based method for privacy protection and efficient retrieval of big data in Internet of Things video surveillance, which is particularly suitable for video surveillance data processing in various scenarios such as vehicle-mounted, park, and battery swapping stations. Background Technology
[0002] With the rapid development of IoT technology, video surveillance systems have been widely applied in various fields such as transportation, industrial parks, and energy facilities. The surveillance data exhibits characteristics of multi-source heterogeneity and massive growth. Currently, IoT video surveillance systems face numerous challenges in practical applications: Firstly, data acquisition in various scenarios suffers from protocol heterogeneity. The network environments of fixed and mobile scenarios differ significantly, making it difficult for traditional fixed-protocol acquisition methods to adapt to dynamic changes in bandwidth and latency, resulting in low data transmission efficiency and poor stability. Secondly, surveillance data contains a large amount of private information such as facial features and vehicle trajectories. Existing privacy protection technologies often employ simple encryption or static desensitization, which not only easily degrades video visual quality but also suffers from low accuracy in extracting privacy information and imprecise authorization access control, failing to meet the differentiated protection needs of privacy data with different sensitivity levels. Simultaneously, the storage and retrieval efficiency of massive surveillance data is low. Traditional index structures do not combine video spatiotemporal characteristics and privacy levels, leading to long response times and high result redundancy for authorized users, making it impossible to quickly locate target data. Furthermore, the system lacks real-time monitoring and adaptive optimization mechanisms for hardware load, network quality, and privacy security during operation. When data volume increases or access demands change, hardware failures, network congestion, or privacy leaks are likely to occur, severely restricting the practicality and security of the IoT video surveillance system.
[0003] Therefore, this invention proposes an artificial intelligence-based method and system for big data privacy protection and efficient retrieval in IoT video surveillance to solve the above problems. Summary of the Invention
[0004] The purpose of this section is to outline some aspects of embodiments of the present invention and to briefly describe some preferred embodiments. Simplifications or omissions may be made in this section, as well as in the abstract and title of this application, to avoid obscuring the purpose of these documents; however, such simplifications or omissions should not be construed as limiting the scope of the invention.
[0005] In view of the problems existing in the above and / or prior art, the present invention is proposed.
[0006] Therefore, the purpose of this invention is to overcome the shortcomings of the prior art and provide an artificial intelligence-based method for privacy protection and efficient retrieval of big data in IoT video surveillance, characterized by the following steps:
[0007] Step S1: Build an adaptive acquisition architecture for multi-source heterogeneous video data. Based on the acquisition scenario priority and network status parameters, construct a protocol selection weight model to achieve adaptive switching of GMSL2, 4G, and WIFI protocols. Perform frame segmentation and DCT transformation on the acquired video data. Combine the spatiotemporal correlation of the monitoring video to construct a spatiotemporal joint DCT coefficient screening model. Select DCT blocks with a high proportion of zero coefficients as subsequent processing carriers. At the same time, introduce a rate-distortion optimization model in the preprocessing stage to balance video visual quality and data volume. Output standardized encoded video stream and supporting feature data.
[0008] Step S2: An improved YOLO target detection algorithm combined with inter-frame optical flow prediction is used to identify targets in video frames. A target stability coefficient is defined to filter objects for privacy information extraction. A spatiotemporal joint energy function is constructed by introducing DCT coefficient distribution features. An improved Graph-Cut technique is used to achieve accurate segmentation of privacy targets. The segmented privacy information is encoded and encrypted to form a privacy information bitstream. Based on high-zero coefficient DCT blocks, a dynamic bit allocation strategy and visual sensitivity coefficient are introduced to construct a dynamic offset formula. An improved rate-distortion optimization overhead function is combined with the Lagrange multiplier method to determine the optimal embedding scheme, completing the reversible watermark adaptive embedding and generating a video stream with hidden privacy and a corresponding key and mapping table.
[0009] Step S3: Construct a multi-dimensional privacy level assessment model based on the target stability coefficient and DCT coefficient distribution characteristics, divide privacy data into three levels (high, medium, and low) and adopt differentiated encryption strategies; construct a two-level storage architecture of "DDR4 cache + SATA hard disk archiving", optimize storage scheduling through storage utilization model and bus load coefficient; extract video key frames and use an improved ResNet-50 network to extract deep feature vectors, combine the spatiotemporal correlation of monitoring video to construct a dynamic hash function to map high-dimensional features into binary hash codes, construct a three-level hierarchical index based on hash codes and privacy levels, and introduce a dynamic index update mechanism to ensure that index performance remains stable as data volume increases;
[0010] Step S4: Employ a dual-factor authentication system combining biometrics and a dynamic key, along with a fixed personnel feature database, to verify identity and permissions. Decrypt privacy data using corresponding algorithms based on permission levels. Introduce a privacy level weighted histogram translation inverse algorithm to extract watermarks. Call a video restoration algorithm to fill gaps in privacy regions to output a complete original video and privacy information summary. Simultaneously, use natural language processing to parse user search requests into structured feature vectors. Utilize a three-level index for rapid filtering. Combine the authorized privacy level to calculate the weighted similarity between the search vector and the video hash code to determine candidate videos. Then, use an improved YOLO algorithm for secondary recognition, calculate the intersection-union ratio (IUU) of the target region and search features, and introduce an adaptive strategy to adjust the index node degree and hash function parameters. Finally, filter by permissions and output compliant search results and reports.
[0011] Step S5: By collecting hardware parameters through current monitoring chips and temperature sensors, monitoring network link transmission parameters in real time, and auditing authorized access records, hardware health, network transmission quality, and privacy security indicators are constructed respectively. The three-dimensional indicators are integrated to form a comprehensive system health status to achieve real-time monitoring of the operating status and anomaly warning. At the same time, based on the retrieval response time and result accuracy, an index node degree dynamic adjustment algorithm and a key frame extraction interval adjustment strategy are designed. Combined with storage utilization, an adaptive compression strategy is designed to adjust the compression ratio of low-sensitivity data. Based on the privacy level, encryption resources are dynamically allocated and the access audit frequency is adjusted to achieve adaptive optimization of system parameters and balance retrieval efficiency, storage overhead, and privacy protection strength.
[0012] Preferably, in step S1, the protocol selection weight model satisfies: ,in For fixed scene priority coefficients, This is a priority coefficient for mobile scenarios. These represent network bandwidth and transmission latency in a fixed scenario. These represent network bandwidth and transmission latency in mobile scenarios, respectively; in the spatiotemporal joint DCT coefficient selection model, the Pearson correlation coefficient of the inter-frame DCT coefficients satisfies: , The first Frame and the DCT coefficients of the frame, , These are the mean values of the DCT coefficients for the corresponding frames. The number of DCT blocks per frame; the rate-distortion optimization model satisfies: , For visual distortion, This is the increase in bit rate. These are the weighting coefficients.
[0013] Preferably, in step S2, the target stability coefficient satisfies: , The number of video frames within the time window. For the first Frame and the Correlation of DCT coefficients in the target region of the frame. The first Frame and the The target region area of the frame; the spatiotemporal joint energy function satisfies: , For the privacy target area to be segmented, For regional items, For boundary terms, For boundary weights, Weights representing distribution characteristics For the first Cauchy distribution probability density of the DCT coefficients corresponding to the pixels. The mean of the DCT coefficient distribution in the target region; the dynamic offset formula satisfies: , These are the original DCT coefficients. These are the DCT coefficients after embedding the watermark. Decimal values for privacy data. The visual sensitivity coefficient is used; the improved rate-distortion optimization cost function satisfies: , It is a static control factor. For time-related weights, Since it is a constant, solve using the Lagrange multiplier method: , To embed the DCT coefficient set, For Lagrange multipliers, For the first Number of bits to be embedded in the block This represents the total number of bits in the watermark.
[0014] Preferably, in step S3, the multidimensional privacy level assessment model satisfies: , The average inter-frame correlation between the privacy target region and the background region. The target stability coefficient, The parameters are Cauchy distribution parameters; high-sensitivity data encryption uses the Lorentz chaotic mapping algorithm: , For chaos control parameters, For adjustment coefficients, For continuously iterated encrypted data values; the storage utilization model satisfies: , Used storage capacity Total storage capacity Average access time, For storage device read / write cycles; the bus load factor satisfies: , The number of I2C bus requests per unit time. For average transmission time, For bus cycles; the dynamic hash function satisfies: , For random vectors, This is the offset. The width of the hash bucket. The hash code length; the optimal node degree of a B+ tree satisfies: V represents the total amount of video data. Average reading time This represents the average write time.
[0015] Preferably, in step S41, the biometric verification uses cosine similarity matching: This is the current user's facial feature vector. The authorized user feature vectors are stored in the database; the Lorentz chaotic mapping inverse algorithm satisfies: , To encrypt data, The original data after decryption; the improved histogram translation inverse algorithm satisfies: , These are the DCT coefficients after embedding the watermark. To restore the original coefficients, For the number of embedding bits, Privacy level weighting.
[0016] Preferably, in step S42, the weighted similarity calculation satisfies: , These are the weighting coefficients. To retrieve the cosine similarity between the feature vector and the hash code, For video privacy levels, The maximum access level for users; the intersection-union ratio (IUU) calculation satisfies: A represents the target region to be retrieved, and B represents the target region detected in the video; the degree of the B+ tree node is adaptively adjusted to satisfy: , This represents the original optimal node degree. The current retrieval response time; the hash function bucket width adjustment satisfies: , The original barrel width, This refers to the redundancy rate of the search results.
[0017] Preferably, in step S51, the hardware real-time power calculation satisfies: , For temperature coefficient, The reference temperature is used; the hardware health assessment meets the following requirements: , The rated power of the module, This is the module's maximum withstand temperature. For the signal strength of the communication module, Full signal strength; network quality assessment meets: , This is the maximum bandwidth of the link. The time delay attenuation coefficient is... For transmission delay, The packet loss rate; the permission matching degree calculation satisfies: , Number of visits per unit of time For the first The authorization level for this access. The requested access level is [level]; privacy and security indicators are met. , For the frequency of access to private data, The encryption integrity verification result is valid; the overall system health meets the following requirements: .
[0018] Preferably, in step S52, the B+ tree node degree optimization satisfies: , This represents the original optimal node degree. This represents the current retrieval response time; the keyframe extraction interval adjustment satisfies: , The average inter-frame correlation is used; the compression ratio is adaptively adjusted to satisfy: , For storage utilization; rate distortion control factor optimization satisfies: , The original control factor, Distortion of the target The current actual distortion; the access audit frequency adjustment meets the following requirements: , For privacy levels, the frequency is measured in "times per minute".
[0019] This invention provides an artificial intelligence-based method for protecting the privacy and efficiently retrieving big data from IoT video surveillance, which has the following beneficial effects:
[0020] 1. A two-layer acquisition architecture of "edge-distributed acquisition + central aggregation" is constructed. Combined with a dynamic protocol adaptation algorithm, the GMSL2, 4G and WIFI protocols are adaptively switched according to the priority of fixed and mobile scenarios and network status to solve the problem of heterogeneous protocols in multiple scenarios. Based on the spatiotemporal joint feature preprocessing model, while achieving a data compression ratio of 1:1.8, the PSNR of the preprocessed video is guaranteed to be no less than 38dB, providing high-quality data support for subsequent privacy protection and retrieval.
[0021] 2. A three-level privacy extraction system of "detection-segmentation-verification" and reversible watermarking technology with dynamic rate distortion optimization are adopted. The system accurately selects privacy targets and embeds watermarks in a differentiated manner by combining the target stability coefficient and DCT coefficient distribution characteristics. Through Lorentz chaotic mapping, RSA-2048 and AES-128 hierarchical encryption strategies, the system achieves secure storage of privacy data with different sensitivity levels. Authorized users can reversibly extract privacy information with PSNR of not less than 36dB by using two-factor authentication and an improved histogram translation inverse algorithm, thus balancing privacy protection and video usability.
[0022] 3. A three-level hierarchical index is constructed by integrating spatiotemporal features, including a scene tag hash table, a timestamp skip table, and a privacy level and hash code B+ tree; high-dimensional video features are mapped to 64-bit hash codes based on an improved LSH algorithm, and the search range is accurately filtered by combining privacy levels; the B+ tree node degree and hash bucket width are dynamically adjusted through an adaptive search strategy to shorten the search response time, and the matching accuracy of search results is improved through secondary filtering with IoU ≥ 0.6, meeting the needs of authorized users to quickly obtain compliant data.
[0023] 4. Construct a three-in-one real-time monitoring system integrating hardware, network, and privacy. Through real-time assessment of hardware health, network transmission quality, and privacy security indicators, promptly trigger hardware backup switching, network link optimization, or privacy security alerts. Combined with a feedback mechanism, dynamically adjust the index node degree, key frame extraction interval, data compression ratio, and privacy audit frequency to achieve adaptive optimization of the system as data volume increases and access demands change, ensuring system stability and security. Attached Figure Description
[0024] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the structures shown in these drawings without creative effort.
[0025] Figure 1 This diagram illustrates a method for privacy protection and efficient retrieval of big data in IoT video surveillance based on artificial intelligence.
[0026] Figure 2 This is a flowchart of step S1, IoT video surveillance data acquisition and preprocessing, in this invention;
[0027] Figure 3 This is a flowchart of step S2, intelligent extraction and embedding of privacy information based on reversible watermarking, in the present invention.
[0028] Figure 4 This is a flowchart of step S3, privacy data hierarchical encryption storage and retrieval index construction, in this invention;
[0029] Figure 5 This is a flowchart of step S4, the reversible extraction and retrieval optimization of authorized user privacy data, in this invention;
[0030] Figure 6 This is a flowchart of step S5, system dynamic monitoring and adaptive optimization, in this invention;
[0031] The objectives, features, and advantages of this invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0032] The technical solutions of the present invention will now be clearly and completely described with reference to the accompanying drawings of the embodiments of the present invention. It should be understood that the specific embodiments described herein are for illustration and explanation only and are not intended to limit the present invention.
[0033] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.
[0034] Secondly, the term "an embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places throughout this specification does not necessarily refer to the same embodiment, nor is it a single embodiment or an embodiment selectively excluded from other embodiments.
[0035] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the examples in the specification.
[0036] This invention proposes an artificial intelligence-based method for privacy protection and efficient retrieval of big data from IoT video surveillance, characterized by the following steps:
[0037] Step S1: IoT video surveillance data acquisition and preprocessing;
[0038] Step S2: Intelligent extraction and embedding of privacy information based on reversible watermarks;
[0039] Step S3: Privacy-protected data hierarchical encryption storage and retrieval index construction;
[0040] Step S4: Authorize the reversible extraction and retrieval optimization of user privacy data;
[0041] Step S5: System dynamic monitoring and adaptive optimization;
[0042] In a preferred embodiment of the present invention, in step S1, IoT video surveillance data acquisition and preprocessing, step S11 is first performed to build an adaptive acquisition architecture for multi-source heterogeneous video data, and then step S12 is performed to preprocess and optimize video data based on spatiotemporal joint features.
[0043] In step S1, IoT video surveillance data acquisition and preprocessing, a distributed acquisition system with a heterogeneous processor as the core is constructed. The Hi3520D processor is used as the core of the edge acquisition node in scenarios such as vehicles and parks, and integrates a 4G communication module and a WIF| module to realize wireless video backhaul in mobile scenarios. At the same time, an MPSoC chip is deployed as a centralized acquisition node in fixed scenarios such as battery swapping stations, and 16 AHD high-definition cameras are connected through the GMSL2 protocol link to form a two-layer architecture of "edge distributed acquisition + central centralized aggregation".
[0044] To address the issue of protocol heterogeneity across multiple scenarios, a dynamic protocol adaptation algorithm is introduced, based on the priority of the data collection scenario (with a fixed scenario priority coefficient). Mobile scenario priority coefficient) and network status parameters (bandwidth) (Latency), construct a protocol selection weight model to achieve adaptive switching between GMSL2, 4G, and WIFI protocols, as shown in the formula: ,in, These represent network bandwidth and transmission latency in a fixed scenario. These represent network bandwidth and transmission latency in mobile scenarios, respectively; when The GMSL2 protocol should be selected first. When selecting the 4G protocol, The system switches to the WIFI protocol at any time. During the acquisition process, the NVP6114 video AD module completes the conversion from analog to digital signals, outputting YUV4228bit format data. At the same time, the chip's built-in 9-channel audio codec synchronously acquires audio information, forming an "audio-video synchronous encapsulation" acquisition data packet.
[0045] In step S12, video data preprocessing and optimization based on spatiotemporal joint features, the acquired audio and video data packets are processed by frame segmentation. A 4x4 block DCT transform is performed on each video frame. Utilizing the characteristic that the quantized DCT coefficients in H.264 / AVC encoding follow a Cauchy distribution, and combining this with the spatiotemporal correlation of the surveillance video, a spatiotemporal joint DCT coefficient selection model is constructed. ,in, The mean of DCT coefficients represents the background area of the surveillance video. , The scaling parameter is used when the quantization step size QP = 10. When QP=20, a=1.2. These are the quantized DCT coefficient values. By calculating the Pearson correlation coefficient between the inter-frame DCT coefficients, frames with correlation of more than three consecutive frames are selected. DCT blocks, these blocks mostly correspond to static background areas, and the formula is: ,in, The first Frame and the DCT coefficients of the frame, , These are the mean values of the DCT coefficients for the corresponding frames. This represents the number of DCT blocks per frame.
[0046] Meanwhile, to balance the visual quality and data volume of the preprocessed video, a rate-distortion optimization model is introduced in the preprocessing stage: the quantization accuracy of the selected DCT blocks is adjusted, using the following formula: ,in, Visual distortion caused by preprocessing This is the increase in bit rate after preprocessing. The weighting coefficients are minimized to achieve a balance between distortion and bit rate. After preprocessing, a normalized H.264 encoded video stream is output, along with a DCT coefficient distribution feature table and an inter-frame correlation matrix.
[0047] In a preferred embodiment of the present invention, in step S2, intelligent extraction and embedding of privacy information based on reversible watermark, step S21, accurate extraction and modeling of privacy information fused with spatiotemporal features is performed first, followed by step S22, adaptive embedding of reversible watermark with dynamic rate distortion optimization.
[0048] In step S21, the precise extraction and modeling of privacy information by fusing spatiotemporal features, a three-level privacy information extraction system of "detection-segmentation-verification" is constructed based on the standardized H.264 encoded video stream, DCT coefficient distribution feature table, and inter-frame correlation matrix output in step S1. First, an improved YOLO target detection algorithm is used to identify targets in the video frames. Addressing the need to distinguish between fixed and temporary personnel in monitoring scenarios, a target stability coefficient is defined based on the inter-frame correlation matrix from step S1. Filter out The goal is to extract privacy information from temporary personnel, thereby reducing the probability of accidental extraction. in, The number of video frames within the time window. For the first Frame and the Correlation of DCT coefficients in the target region of the frame. The first Frame and the The area of the target region of the frame.
[0049] For the selected privacy targets, an improved Graph-Cut technique is used to achieve accurate segmentation. This involves introducing a DCT coefficient distribution feature term based on the traditional energy function, and utilizing the Cauchy distribution parameters from step S1... A spatiotemporal joint energy function is constructed to improve segmentation accuracy in occluded scenes. Its formula is:
[0050]
[0051] in, For the privacy target area to be segmented, For regional items, For boundary terms, For boundary weights, Weights representing distribution characteristics For the first Cauchy distribution probability density of the DCT coefficients corresponding to the pixels. The mean of the DCT coefficient distribution in the target region is given. After segmentation, the minimum bounding rectangle region of the privacy target is extracted and encrypted using H.264 / AVC encoding and AES-256 to form a privacy information bitstream. At the same time, a database of fixed personnel characteristics will be established to provide a basis for subsequent permission verification.
[0052] In step S22, the reversible watermark adaptive embedding with dynamic rate-distortion optimization, the high-zero-coefficient DCT blocks selected in step S1 are used as carriers. A reversible watermark embedding mechanism is designed based on the histogram translation algorithm. Considering the real-time requirements of surveillance video, a dynamic bit allocation strategy and visual distortion control are introduced to achieve a balance between large-capacity watermark embedding and low distortion. First, for each 4x4 DCT block, a zero-coefficient utilization rate is defined. The formula is: ,in, For the first The number of zero coefficients in a DCT block The total number of coefficients for a 4x4 DCT block, according to Value dynamically adjusts the number of embedded bits , hour , hour , Watermark embedding is not performed at all times to ensure that the watermark embedding capacity is adapted to the privacy information bitstream. size.
[0053] Based on the histogram translation algorithm, the DCT block coefficients are adjusted to embed the watermark. This improves upon the fixed offset strategy of traditional algorithms by introducing a visual sensitivity coefficient. A dynamic offset formula is constructed to reduce distortion in areas sensitive to human vision. The formula is:
[0054]
[0055] in, These are the original DCT coefficients. These are the DCT coefficients after embedding the watermark. for The decimal value of bit privacy data. for Visual sensitivity coefficient of location.
[0056] To further balance visual distortion and bit rate increase, an improved rate-distortion optimization overhead function is constructed, introducing time-related weights. The watermark embedding position of consecutive frames is optimized to avoid distortion accumulation caused by overlapping embedding regions in adjacent frames. It is a static control factor. For time-related weights, It is a constant. Visual distortion is determined by the formula: calculate, This represents the increase in bit rate. The minimum cost function is found using the Lagrange multiplier method. The optimal set of embedded DCT blocks and bit allocation scheme are determined. To embed the DCT coefficient set, For Lagrange multipliers, For the first Number of bits to be embedded in the block This represents the total number of bits for the watermark. After embedding, the video stream is re-encoded using H.264 / AVC to generate a video stream with hidden privacy, S. Simultaneously, a digital signature for the video is calculated to achieve video integrity authentication; the final output is... Privacy information encryption keys and embedded location mapping tables.
[0057] In a preferred embodiment of the present invention, in step S3, privacy data hierarchical encryption storage and retrieval index construction, step S31, privacy data dynamic hierarchical encryption and intelligent storage based on target features is performed first, followed by step S32, video hash index fused with spatiotemporal features and hierarchical retrieval architecture construction.
[0058] In step S31, dynamic hierarchical encryption and intelligent storage of privacy data based on target features, the hidden privacy video stream output in step S2 is received. A privacy-protected storage system is constructed based on the privacy information encryption key and the embedding location mapping table, which is "feature-driven, dynamically hierarchical, and adaptable storage". First, based on the target stability coefficient in step S2... Based on the distribution characteristics of DCT coefficients, a multidimensional privacy level assessment model is constructed, classifying privacy data into three levels: highly sensitive (…). Such as facial information of fixed personnel), medium sensitive information ( Such as vehicle location trajectory), low sensitivity ( (e.g., information about public area environments) to achieve dynamic classification of privacy levels. The multi-dimensional privacy level assessment model is as follows: ,in, The average inter-frame correlation between the privacy target region and the background region. The target stability coefficient, The Cauchy distribution parameter is used to fuse the target dynamic characteristics and coefficient distribution characteristics through this formula, thereby improving the accuracy of classification.
[0059] Different encryption strategies are adopted for data with different levels of privacy: highly sensitive data uses an encryption algorithm based on Lorentz chaotic mapping. Security is enhanced by leveraging its sensitivity to initial values; moderately sensitive data is encrypted using RSA-2048; and low-sensitivity data is encrypted using AES-128. For chaos control parameters, For adjustment coefficients, The encrypted data value is generated through continuous iteration, with the initial value generated from the hash value of the privacy information encryption key.
[0060] At the storage level, a two-tier storage architecture consisting of "DDR4 cache + SATA hard drive archiving" is constructed based on the MPSoC chip's storage controller. ,in, Used storage capacity Total storage capacity Average access time, For storage device read / write cycles, when Low-sensitivity data compression is triggered at certain times. DDR4 is used for caching real-time data. For video streams and frequently accessed low-sensitivity data, a dynamic cache replacement strategy is implemented using a storage utilization model. SATA hard drives are used for long-term storage of encrypted high / medium-sensitivity data and historical videos. Simultaneously, intelligent management of storage devices is achieved using the I2C bus, with bus load factor... Optimize data transmission timing to avoid conflicts from concurrent access by multiple devices. The number of I2C bus requests per unit time. For average transmission time, For bus cycle, when The bus arbitration mechanism is activated in time to prioritize the transmission of highly sensitive data.
[0061] In step S32, the construction of the video hash index and hierarchical retrieval architecture that integrates spatiotemporal features, based on the storage in S31... Video streams and encrypted privacy data are used to construct an efficient retrieval system based on "feature extraction - hash mapping - dynamic indexing," providing support for subsequent retrieval by authorized users. First, the gst-launch-1.0 command of the GStreamer framework is used to extract... For keyframes in the video, a ResNet-50 network is used to extract deep features from the keyframes, outputting 2048-dimensional feature vectors. The embedding position mapping table information from step S2 is incorporated into the feature vector.
[0062] To achieve efficient retrieval of high-dimensional features, an improved locality-sensitive hashing algorithm is introduced. A dynamic hash function is constructed by combining the spatiotemporal correlation of surveillance videos, mapping the 2048-dimensional feature vector into a 64-bit binary hash code. The number of hash buckets depends on the amount of video data. Dynamic adjustments are made to ensure uniform hash distribution. The improved locality-sensitive hashing algorithm is as follows: ,in, For random vectors, The offset is (). The hash bucket width is determined by the zero coefficient utilization rate of the DCT block in step S1. Dynamic adjustment , The hash code has a bit length.
[0063] Based on hash code Privacy level assessment results A three-tiered index is constructed: the first-level index uses monitoring scene tags as keys and is stored using a hash table; the second-level index uses timestamps as keys and uses a skip list structure to optimize range queries; the third-level index is based on privacy levels. The hash code prefix serves as the combined key, and a B+ tree structure is used to optimize retrieval efficiency. Simultaneously, a dynamic index update mechanism is introduced, with the formula: When the amount of newly added video data exceeds 20% of the current index capacity, the degree of the B+ tree nodes will be automatically adjusted. This ensures that index performance remains stable as data volume increases. Among other things, To achieve the optimal node degree in a B+ tree, Total video data Average reading time To average write time, this formula is used to balance index read and write performance. Finally, a mapping relationship of "hash code - index table - storage address" is constructed to provide efficient index support for authorized retrieval in step S4. Simultaneously, privacy level markers are incorporated into the index to ensure that search results are only accessible to authorized users with the corresponding level of privacy information.
[0064] In a preferred embodiment of the present invention, in step S4, reversible extraction and retrieval optimization of authorized user privacy data, step S41, multi-factor authorization authentication and reversible extraction of privacy data are performed first, followed by step S42, adaptive retrieval optimization and result filtering based on privacy level fusion.
[0065] In step S41, multi-factor authorization authentication and reversible extraction of privacy data, the hierarchical encrypted storage system and privacy level assessment model constructed in step S3 are integrated to design a full-process privacy data access mechanism of "identity authentication - permission matching - reversible extraction". First, for authorized users, a two-factor authentication method of "biometric feature + dynamic key" is adopted: the user's facial image is captured by a camera and matched with the fixed personnel feature database established in step S2 using cosine similarity, with the formula: ,in, This is the current user's facial feature vector. For the authorized user feature vectors stored in the database, when Once biometric verification is successful, the validity of the authorization is determined by the correctness of the dynamic key. In addition, the dynamic key bound to the hardware is also used for dual verification. Only after passing this dual verification can the user obtain access to the corresponding privacy level.
[0066] Based on the authenticated permission level, initiate the differentiated privacy data extraction process: for highly sensitive data, first use the Lorentz chaotic mapping inverse algorithm. Decryption, among which To encrypt data, The original data after decryption. Consistent with the encryption phase, based on the embedding position mapping table from step S2, the data is directly extracted after decryption with an RSA-2048 private key using an improved histogram translation inverse algorithm; low-sensitivity data is output after AES-128 decryption. The improved histogram translation inverse algorithm is shown in the formula:
[0067] in, These are the DCT coefficients after embedding the watermark. To restore the original coefficients, For the number of embedding bits, To establish privacy level weights, privacy level weights are introduced. Prioritize the recovery of the original DCT coefficients in areas sensitive to human vision to ensure that the PSNR of the extracted video is not less than 36dB.
[0068] After extracting privacy information, a video restoration algorithm is called to fill the privacy region holes in the original video. Combined with the inter-frame correlation matrix in step S1, a restoration region consistent with the background texture is generated, and finally, the complete original video and privacy information summary are output.
[0069] In step S42, the adaptive retrieval optimization and result filtering based on privacy levels, a retrieval optimization mechanism of "semantic parsing - fast filtering - precise matching" is designed based on the three-level hash index and feature mapping relationship constructed in step S3 to ensure that authorized users can efficiently obtain compliant monitoring data. First, the user inputs a retrieval request through the embedded GUI, and the system parses the request into a structured feature vector through natural language processing. The target type feature refers to the target stability coefficient in step S2. .
[0070] The three-level index in step S3 enables rapid filtering: the first-level index locates the hash bucket set of the target monitoring area based on scene tags; the second-level index filters corresponding skip list nodes by time range; and the third-level index combines the user-authorized privacy level. Only call B+ tree branches, using the formula Calculate the retrieval feature vector With video hash code The weighted similarity was used to select the top-20 candidate videos based on their similarity. These are the weighting coefficients. for and cosine similarity, For privacy level matching degree, when When included in the candidate set.
[0071] To improve retrieval accuracy, an improved YOLO object detection algorithm is used for secondary identification of candidate videos. Regions matching the retrieval target are extracted from the video frames, and the intersection-union ratio (IUU) of the target region and the retrieval features is calculated. Video segments with a loU ≥ 0.6 are retained. Simultaneously, an adaptive retrieval strategy is introduced: when the retrieval response time is [time value missing], where [missing information] To retrieve the target region, This refers to the target area detected in the video. At the same time, dynamically adjust the optimal node degree of the B+ tree in step S3. The formula is: , This represents the original optimal node degree. This refers to the current retrieval response time. The index hierarchy is reduced by increasing node degree; when the retrieval result redundancy rate... At the same time, optimize the bucket width of the hash function. This improves the uniformity of hash distribution.
[0072] Finally, the search results are filtered according to the user's permission level: users with high permissions can view the complete video and extracted privacy information, while users with low permissions only see the anonymized video, and a search report is output.
[0073] In a preferred embodiment of the present invention, in step S5, system dynamic monitoring and adaptive optimization, step S51, multi-dimensional system operation status real-time monitoring and anomaly warning is performed first, followed by step S52, system parameter adaptive optimization based on feedback mechanism.
[0074] In step S51, real-time monitoring and anomaly warning of multi-dimensional system operation status, the authorized retrieval and privacy extraction process in step S4 is connected to construct a dynamic monitoring system integrating "hardware-network-privacy" to achieve real-time perception and risk warning of the entire system's operation status. At the hardware level, a current monitoring chip is used to collect the voltage of the core module in real time. With current Through power calculation model Assess hardware load, among which, For temperature coefficient, Using the reference temperature, power calculation accuracy is improved through temperature correction, combined with chip temperature data collected by a temperature sensor. Build hardware health indicators ,in, The rated power of the module, This is the module's maximum withstand temperature. For the signal strength of the communication module, The full-scale signal strength ranges from [0,1]. The system is identified as a hardware malfunction, triggering a switchover to the backup module.
[0075] At the network level, based on the protocol adaptation algorithm in step S1, the bandwidth of the 4G / IFI / GMSL2 links is monitored in real time. Transmission delay Combined with packet loss rate, construct network transmission quality indicators , The rated power of the module, This is the module's maximum withstand temperature. For the signal strength of the communication module, The full-scale signal strength ranges from [0,1]. Initiate link optimization at the appropriate time.
[0076] In terms of privacy protection, the authorized access records in step S4 are audited in real time to count the frequency of access to privacy data. Permission matching degree With encryption integrity ,in, Number of visits per unit of time For the first The authorization level for this access. For the requested access level, when Time to take Build privacy and security metrics ,in, This indicates that the hash verification passed. This indicates a validation failure; the value range is [0,1]. Security alerts are triggered in a timely manner. Ultimately, by integrating indicators from three dimensions—hardware, network, and privacy—a comprehensive system health assessment is constructed. Real-time display via embedded GUI, when It triggers an audible and visual alarm and pushes anomaly logs to the management terminal.
[0077] In step S52, the system parameter adaptive optimization based on the feedback mechanism, a "performance-resource-privacy" collaborative optimization model is constructed based on the monitoring data from step S51 and the retrieval feedback from step S4. This model achieves a balance between efficiency and security by dynamically adjusting the core system parameters. Regarding retrieval optimization, to address the issue of excessively long retrieval response time in step S4, a dynamic adjustment algorithm for index node degree is designed based on the B+ tree index structure from S302. , This represents the original optimal node degree. To avoid memory overflow, the node degree limit is set to 1.5 times the original value to accommodate the current retrieval response time. Increasing the node degree reduces the index hierarchy, while also considering the inter-frame correlation from step S1. Adjust keyframe extraction interval ,in, The average correlation between frames is used to reduce the extraction frequency when the correlation is high, thereby reducing storage pressure.
[0078] In terms of storage optimization, considering the storage utilization of the S301... Design an adaptive compression strategy: when For low-sensitivity data, H.265 / HEVC encoding is used and the compression ratio is dynamically adjusted. , For storage utilization; when At the same time, the compression ratio is reduced to improve video restoration quality. Simultaneously, based on the reversible watermark rate-distortion model from step S2, the privacy information embedding parameters are optimized by adjusting control factors. ,in, The original control factor, Distortion of the target This represents the current actual distortion. Balancing visual distortion. With bit rate increase Ensure that the PSNR of the embedded video is ≥35dB.
[0079] Regarding privacy protection optimization, based on the access permission records in step S4, encrypted resources are prioritized for highly sensitive data, while the frequency of access auditing is dynamically adjusted. This enables focused monitoring of high-risk data. The frequency of access auditing is adjusted, where... For privacy levels, the frequency is measured in "times / minute", with the highest frequency for auditing highly sensitive data reaching 5 times / minute.
[0080] Through the above multi-dimensional adaptive optimization, the system can dynamically balance retrieval efficiency, storage overhead, and privacy protection strength under different operating conditions, ensuring stable and reliable operation even when the amount of monitoring data increases and access requirements change.
[0081] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
Claims
1. A method for privacy protection and efficient retrieval of big data in IoT video surveillance based on artificial intelligence, characterized in that, Includes the following steps: Step S1: Build an adaptive acquisition architecture for multi-source heterogeneous video data. Based on the acquisition scenario priority and network status parameters, construct a protocol selection weight model to achieve adaptive switching of GMSL2, 4G, and WIFI protocols. Perform frame segmentation and DCT transformation on the acquired video data. Combine the spatiotemporal correlation of the monitoring video to construct a spatiotemporal joint DCT coefficient screening model. Select DCT blocks with a high proportion of zero coefficients as subsequent processing carriers. At the same time, introduce a rate-distortion optimization model in the preprocessing stage to balance video visual quality and data volume. Output standardized encoded video stream and supporting feature data. Step S2: The improved YOLO target detection algorithm is combined with inter-frame optical flow prediction to identify video frame targets. The target stability coefficient is defined to screen the objects for privacy information extraction. The DCT coefficient distribution feature term is introduced to construct the spatiotemporal joint energy function. The improved Graph-Cut graph cut technology is used to achieve accurate segmentation of privacy targets. The segmented privacy information is encoded and encrypted to form a privacy information bitstream. Based on high zero-coefficient DCT blocks, a dynamic bit allocation strategy and visual sensitivity coefficient are introduced to construct a dynamic offset formula. Combined with an improved rate distortion optimization overhead function, the optimal embedding scheme is determined by the Lagrange multiplier method to complete the reversible watermark adaptive embedding and generate a video stream with hidden privacy and a matching key and mapping table. Step S3: Construct a multi-dimensional privacy level assessment model based on the target stability coefficient and DCT coefficient distribution characteristics, divide privacy data into three levels (high, medium, and low) and adopt differentiated encryption strategies; construct a two-level storage architecture of "DDR4 cache + SATA hard disk archiving", optimize storage scheduling through storage utilization model and bus load coefficient; extract video key frames and use an improved ResNet-50 network to extract deep feature vectors, combine the spatiotemporal correlation of monitoring video to construct a dynamic hash function to map high-dimensional features into binary hash codes, construct a three-level hierarchical index based on hash codes and privacy levels, and introduce a dynamic index update mechanism to ensure that index performance remains stable as data volume increases; Step S4: Employ a dual-factor authentication system combining biometrics and a dynamic key, along with a fixed personnel feature database, to verify identity and permissions. Decrypt privacy data using corresponding algorithms based on permission levels. Introduce a privacy level weighted, improved histogram translation inverse algorithm to extract watermarks. Call a video restoration algorithm to fill gaps in privacy regions to output a complete original video and privacy information summary. Simultaneously, use natural language processing to parse user search requests into structured feature vectors. Utilize a three-level index for rapid filtering. Combine authorized privacy levels to calculate the weighted similarity between the search vector and the video hash code to determine candidate videos. Then, use an improved YOLO algorithm for secondary recognition, calculate the intersection-union ratio (IUU) of the target region and search features, and introduce an adaptive strategy to adjust the index node degree and hash function parameters. Finally, filter by permissions and output compliant search results and reports. Step S5: By collecting hardware parameters through current monitoring chips and temperature sensors, monitoring network link transmission parameters in real time, and auditing authorized access records, hardware health, network transmission quality, and privacy security indicators are constructed respectively. The three-dimensional indicators are integrated to form a comprehensive system health status to achieve real-time monitoring of the operating status and anomaly warning. At the same time, based on the retrieval response time and result accuracy, an index node degree dynamic adjustment algorithm and a key frame extraction interval adjustment strategy are designed. Combined with storage utilization, an adaptive compression strategy is designed to adjust the compression ratio of low-sensitivity data. Based on the privacy level, encryption resources are dynamically allocated and the access audit frequency is adjusted to achieve adaptive optimization of system parameters and balance retrieval efficiency, storage overhead, and privacy protection strength.
2. The IoT video surveillance big data privacy protection and efficient retrieval method according to claim 1, characterized in that, In step S1, the protocol selection weight model satisfies: ,in For fixed scene priority coefficients, This is a priority coefficient for mobile scenarios. These represent network bandwidth and transmission latency in a fixed scenario. These represent network bandwidth and transmission latency in mobile scenarios, respectively; in the spatiotemporal joint DCT coefficient selection model, the Pearson correlation coefficient of the inter-frame DCT coefficients satisfies: , The first Frame and the DCT coefficients of the frame, , These are the mean values of the DCT coefficients for the corresponding frames. This refers to the number of DCT blocks per frame. The rate-distortion optimization model satisfies: , For visual distortion, This is the increase in bit rate. These are the weighting coefficients.
3. The IoT video surveillance big data privacy protection and efficient retrieval method according to claim 1, characterized in that, In step S2, the target stability coefficient satisfies: , The number of video frames within the time window. For the first Frame and the Correlation of DCT coefficients in the target region of the frame. The first Frame and the The target region area of the frame; the spatiotemporal joint energy function satisfies: , For the privacy target area to be segmented, For regional items, For boundary terms, For boundary weights, Weights representing distribution characteristics For the first Cauchy distribution probability density of the DCT coefficients corresponding to the pixels. The mean of the DCT coefficient distribution in the target region; the dynamic offset formula satisfies: , These are the original DCT coefficients. These are the DCT coefficients after embedding the watermark. Decimal values for privacy data. Visual sensitivity coefficient; The improved rate-distortion optimization cost function satisfies: , It is a static control factor. For time-related weights, Since it is a constant, solve using the Lagrange multiplier method: , To embed the DCT coefficient set, For Lagrange multipliers, For the first Number of bits to be embedded in the block This represents the total number of bits in the watermark.
4. The IoT video surveillance big data privacy protection and efficient retrieval method according to claim 1, characterized in that, In step S3, the multidimensional privacy level assessment model satisfies: , The average inter-frame correlation between the privacy target region and the background region. The target stability coefficient, The parameters are Cauchy distribution parameters; high-sensitivity data encryption uses the Lorentz chaotic mapping algorithm: , For chaos control parameters, For adjustment coefficients, The encrypted data value is obtained through continuous iteration; The storage utilization model satisfies: , Used storage capacity Total storage capacity Average access time, For storage device read / write cycles; the bus load factor satisfies: , The number of I2C bus requests per unit time. For average transmission time, For bus cycles; the dynamic hash function satisfies: , It is a random vector. This is the offset. The width of the hash bucket. The hash code length; the optimal node degree of a B+ tree satisfies: V represents the total amount of video data. Average reading time This represents the average write time.
5. The IoT video surveillance big data privacy protection and efficient retrieval method according to claim 1, characterized in that, In step S41, biometric verification uses cosine similarity matching: This is the current user's facial feature vector. The authorized user feature vectors are stored in the database; the Lorentz chaotic mapping inverse algorithm satisfies: , To encrypt data, The original data after decryption; the improved histogram translation inverse algorithm satisfies: , These are the DCT coefficients after embedding the watermark. To restore the original coefficients, For the number of embedding bits, Privacy level weighting.
6. The IoT video surveillance big data privacy protection and efficient retrieval method according to claim 1, characterized in that, In step S42, the weighted similarity calculation satisfies: , These are the weighting coefficients. To retrieve the cosine similarity between the feature vector and the hash code, For video privacy levels, The maximum access level for users; the intersection-union ratio (IUU) calculation satisfies: A represents the target region to be retrieved, and B represents the target region detected in the video; the degree of the B+ tree node is adaptively adjusted to satisfy: , This represents the original optimal node degree. The current retrieval response time; the hash function bucket width adjustment satisfies: , The original barrel width, This refers to the redundancy rate of the search results.
7. The IoT video surveillance big data privacy protection and efficient retrieval method according to claim 1, characterized in that, In step S51, the hardware real-time power calculation satisfies: , For temperature coefficient, The reference temperature is used; the hardware health assessment meets the following requirements: , The rated power of the module, This is the module's maximum withstand temperature. For the signal strength of the communication module, Full signal strength; network quality assessment meets: , This is the maximum bandwidth of the link. The time delay attenuation coefficient is... For transmission delay, The packet loss rate; the permission matching degree calculation satisfies: , Number of visits per unit of time For the first The authorization level for this access. The requested access level is [level]; privacy and security indicators are met. , For the frequency of access to private data, The encryption integrity verification result is valid; the overall system health meets the following requirements: .
8. The IoT video surveillance big data privacy protection and efficient retrieval method according to claim 1, characterized in that, In step S52, the B+ tree node degree optimization satisfies: , This represents the original optimal node degree. This represents the current retrieval response time; the keyframe extraction interval adjustment satisfies: , Inter-frame average correlation; The compression ratio adaptive adjustment satisfies: , For storage utilization; Rate distortion control factor optimization satisfies: , The original control factor, Distortion of the target The current actual distortion; the access audit frequency adjustment meets the following requirements: , For privacy levels, the frequency is measured in "times per minute".
Citation Information
Cited By
A hierarchical isolation retrieval method and system for implicit privacy data
CN122263175A