A secure video image acquisition system and method thereof

By performing one-time physical pairing and three-channel synchronous acquisition in the video acquisition system, combined with layered hybrid coding and end-side encryption, the problem of insufficient coupling between privacy protection and reliable traceability in video acquisition is solved, and strong binding protection and reliable traceability at the regional and time-time granularity are achieved.

CN121567844BActive Publication Date: 2026-07-21XIAN XINCHEN ELECTRONIC TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610094012.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-01-23
Publication Date
2026-07-21
Estimated Expiration
2046-01-23

AI Technical Summary

Technical Problem

Existing video capture technologies struggle to achieve synchronized protection at the regional and time-period granularity at the capture end, resulting in insufficient coupling between privacy protection and reliable traceability. Furthermore, the session key lifecycle cannot implement verifiable forgetting operations for specific time periods.

Method used

An encrypted channel is established between the camera and the user terminal through a one-time physical pairing. The device root key, privacy partition template and time capsule table are generated. Three-channel synchronous acquisition is performed and the capsule and partition numbers are marked. Layered hybrid encoding and end-side encryption are performed to ensure that the traceability mark is written to the public base layer and the partition privacy layer at the same time. The terminal side verifies the consistency between the base layer watermark and the explicit traceability and then decrypts according to the time capsule and partition permissions.

Benefits of technology

It achieves strong binding protection of region and time period granularity during video acquisition, ensuring the uniqueness of privacy partitions and temporal semantics, the continuity of traceability information, and supports selective decryption and reverse fragmentation, thus realizing reliable traceability and privacy protection of video content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121567844B_ABST
    Figure CN121567844B_ABST
Patent Text Reader

Abstract

The application discloses a secret-keeping video image acquisition system and method, and relates to the technical field of video monitoring.The system comprises the following steps: an encrypted channel of a camera and a user terminal is established through one-time physical pairing, a device root key, a privacy partition template and a time capsule table are generated, and are written into a secure storage to obtain an initialization configuration set; layered hybrid coding and end-side encryption are performed on a three-channel synchronous frame set, and a traceable mark is simultaneously written into a public base layer and a partition private layer, and a three-layer hybrid code stream is output; the three-layer hybrid code stream is reviewed for base layer watermark and explicit traceability consistency, and the partition private layer is decrypted and inversely scattered according to the time capsule and the partition authority, and a reconstructed video stream is obtained.The application realizes strong binding of public-privacy-traceability in layered hybrid coding and end-side encryption, wherein the public base layer carries a basic monitoring picture, the partition private layer performs session key encryption according to the partition and the capsule, and the traceable mark is simultaneously written into invisible watermark and an explicit field of a packet header.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of video surveillance technology, and in particular to a secure video image acquisition system and method. Background Technology

[0002] With the digitization of video acquisition and transmission links, surveillance equipment has gradually evolved from early local recording and full-stream compression to a collaborative form of cloud hosting and edge preprocessing. Regarding content security, on the one hand, session-level encryption and transmission encryption are used to ensure full-stream confidentiality; on the other hand, scalable video encoding, watermarking, and log recording are used to achieve content management and source identification. In terms of privacy handling, common practices include post-processing occlusion based on targets such as faces, static blurring in the rectification section, and access control and auditing policies executed in the cloud.

[0003] However, the aforementioned conventional approaches have the problem of difficulty in achieving simultaneous protection at the regional and time-period granularities at the acquisition end: full-stream encryption presents a binary characteristic of being fully visible or fully invisible, and post-processing occlusion relies on the recall and accuracy of the recognition algorithm, which is prone to omissions or false occlusions; structural and texture information are not separated, which may lead to the leakage of sensitive details even under low bitrate conditions; source tracing is mostly loosely bound, making it difficult to form a chain of consistency between the public and private layers at the frame level; the session key lifecycle is usually based on the entire stream, making it impossible to implement verifiable forgetting operations for specific time periods. Summary of the Invention

[0004] In view of the aforementioned existing problems, the present invention is proposed.

[0005] Therefore, this invention provides a secure video image acquisition method that solves the problem of insufficient coupling between privacy protection and reliable traceability at the regional and time-segment granularity during video acquisition.

[0006] To solve the above-mentioned technical problems, the present invention provides the following technical solution:

[0007] In a first aspect, the present invention provides a secure video image acquisition method, comprising:

[0008] An encrypted channel between the camera and the user terminal is established through a one-time physical pairing, generating the device root key, privacy partition template, and time capsule table and writing them into secure storage to obtain the initial configuration set;

[0009] Based on the initial configuration set, each frame of the image is simultaneously acquired in three channels: structure, texture, and source, and labeled with capsule and partition numbers to form a set of three-channel synchronous frames.

[0010] Layered hybrid coding and end-side encryption are performed on the three-channel synchronous frame set, and the traceability mark is written into the public base layer and the partition private layer at the same time, and the three-layer hybrid bitstream is output.

[0011] The consistency between the base layer watermark and explicit source tracing is verified for the three-layer hybrid bitstream. The video stream is decrypted according to time capsule and partition permissions, and the partition privacy layer is reversed to obtain the reconstructed video stream.

[0012] In a preferred embodiment of the secure video image acquisition method of the present invention, the specific steps for obtaining the initialization configuration set are as follows:

[0013] After the camera is powered on, it outputs a low-resolution prompt screen and displays a QR code containing a unique hardware identifier and the factory public key. The user terminal scans the code to complete a one-time physical pairing and establish an encrypted channel to obtain a secure handshake result.

[0014] The camera uses the secure handshake result to generate a device root key and embeds the derivation rules and device identifier into the security chip to form the basic key data.

[0015] The user terminal uses the key-based data to display the scene preview and delineates privacy partitions and privacy levels on the preview screen. At the same time, it organizes continuous time slices into time capsules and assigns access policies, generating partition templates and time capsule tables.

[0016] The partition template and time capsule table are written to secure storage, while the user terminal saves the access policy to form an initialization configuration set.

[0017] As a preferred embodiment of the secure video image acquisition method of the present invention, the specific steps for forming a three-channel synchronous frame set are as follows:

[0018] Read the initialization configuration set, locate the normal area and privacy area in each frame according to the partition template, and attach the capsule number of the current time period to obtain the frame-level annotation map;

[0019] Based on the frame-level annotation map, the ordinary areas and privacy partitions defined in the current frame are processed using preset edge extraction rules and brightness contrast, and the partition number and capsule number are recorded to obtain the structure channel frame;

[0020] Cut out the complete texture blocks of the privacy partition from the original image corresponding to the structure channel frame and label the position numbers to obtain a list of texture patches;

[0021] Attach the device identifier and previous frame summary to the texture list and write the partition number and capsule number to generate the source map;

[0022] The source graph is used to perform time synchronization and number verification on the structural channel frames and texture patch list, forming a three-channel synchronized frame set.

[0023] As a preferred embodiment of the secure video image acquisition method of the present invention, the specific steps of performing layered hybrid encoding and end-side encryption are as follows:

[0024] The camera receives a set of three synchronized frames, retains only the outline and brightness information in the structure channel, and fills the static area with a background reference image to obtain the basic monitoring frame;

[0025] The basic monitoring frames are sent to the video encoder for compression to generate a common base layer while retaining the partition number and capsule number, thus obtaining the base layer encoding result.

[0026] For texture pieces in the three-channel synchronous frame set, the texture pieces are rearranged according to the shuffling order in the partition template and the session key is selected according to the capsule number. The texture pieces are then encrypted and packaged at the partition level to form a partition privacy layer.

[0027] The source graph content corresponding to the three-channel synchronization frame set is retained and compressed to obtain the source graph set to be written with the source marker.

[0028] As a preferred embodiment of the secure video image acquisition method of the present invention, the specific steps for simultaneously writing the traceability marker into both the public base layer and the partitioned private layer are as follows.

[0029] The base layer encoding results are used to generate a base layer summary, which is then concatenated with the device identifier and the previous frame tag to form a chain-like traceability tag.

[0030] By embedding chain-based traceability tags into low-to-medium frequency locations in the public base layer and recording their mapping relationship with partition numbers and capsule numbers, a watermarked public base layer is obtained.

[0031] Based on the set of traceability graphs to be written with traceability tags, traceability tags are established, and chained traceability tags and explicit traceability information are written into the header of the partition privacy layer to obtain a traceability-consistent partition privacy layer.

[0032] The watermarked public base layer, the traceable private partition layer, and the traceability marker block are combined and encoded to output a three-layer hybrid bitstream.

[0033] As a preferred embodiment of the secure video image acquisition method described in this invention, the specific steps for verifying the consistency between the base layer watermark and the explicit source tracing of the three-layer hybrid bitstream are as follows:

[0034] After receiving the three-layer mixed bitstream, the terminal decodes the common base layer and extracts the invisible watermark to obtain the base layer verification information.

[0035] Explicit traceability information is parsed from the private layer header of the partition and chain traceability tags are extracted to form private layer verification information;

[0036] The chain traceability tags in the basic layer verification information are compared with the mapping relationship between the chain traceability tags, partition numbers and capsule numbers in the private layer verification information to obtain the consistency verification result;

[0037] Based on the consistency verification results, mark the decryptable partition set and the non-decryptable partition set, and generate the decryption preparation result.

[0038] As a preferred embodiment of the secure video image acquisition method described in this invention, the specific steps for decryption by time-segment capsules and partition permissions are as follows:

[0039] Based on the decryption preparation results, load the session key set corresponding to the current capsule number and match it with the partition number in the decryptable partition set to form a decryption authorization list;

[0040] According to the decryption authorization list, perform partition-level decryption on the authorized partitions in the private layer of the partition and retain the partition number and position number of the decrypted texture pieces to obtain a list of decrypted texture pieces.

[0041] As a preferred embodiment of the secure video image acquisition method of the present invention, the specific steps of the reverse fragmentation partitioning privacy layer are as follows:

[0042] Based on the shuffling order of the decrypted texture list and partition template record, the decrypted textures are reverse-shuffling and restored to their original arrangement according to the partition number and position number to obtain the restored texture set;

[0043] The restored texture patch set is overlaid onto the corresponding partition position in the common base layer, while keeping the partition number and capsule number consistent, to form a fused reconstruction frame;

[0044] The merged and reconstructed frames are arranged into a frame sequence according to time order, and the reconstructed video stream is output.

[0045] In a preferred embodiment of the secure video image acquisition method of the present invention, the specific steps for performing inverse shuffling processing on the decryption texture patch are as follows:

[0046] A position mapping table is generated based on the shuffling order of the partition template records and aligned item by item with the list of decrypted texture pieces to obtain the mapping alignment result;

[0047] Based on the mapping alignment results, a blank splicing canvas is created for each partition, and the corresponding decrypted texture pieces are placed sequentially according to the position number to obtain a partition-level splicing image;

[0048] Boundary merging and number verification are performed on the partition-level mosaic map to obtain the restored texture patch set.

[0049] Secondly, the present invention provides a secure video image acquisition system, comprising:

[0050] The pairing key module establishes an encrypted channel between the camera and the user terminal through a one-time physical pairing, generates the device root key, privacy partition template and time capsule table and writes them into secure storage to obtain the initial configuration set;

[0051] The synchronous acquisition module performs three-channel synchronous acquisition of each frame of the image based on the initial configuration set, namely structure, texture and source, and labels the capsule and partition number to form a set of three-channel synchronous frames.

[0052] The layered encryption module performs layered hybrid encoding and end-side encryption on the three-channel synchronous frame set, and writes the traceability mark into the public base layer and the partitioned private layer at the same time, and outputs a three-layer hybrid bitstream.

[0053] The verification and reconstruction module verifies the consistency between the base layer watermark and explicit source tracing of the three-layer hybrid bitstream, decrypts the data according to time capsules and partition permissions, and reverse-disassembles the partition privacy layer to obtain the reconstructed video stream.

[0054] The beneficial effects of this invention are as follows: By using one-time physical pairing and device root key derivation, combined with privacy partition templates and time capsule tables, an initial configuration set uniquely bound to the device is formed, ensuring that each subsequent frame has a definite partition and time semantics; In three-channel synchronous acquisition, the structure channel only retains contour and brightness contrast for public visibility, the texture channel strictly corresponds to the partition and capsule number, and presets the granular boundary for authorized reconstruction; the traceability channel writes the device identifier, partition number and previous frame summary at the frame level, making the source chain continuous; In layered hybrid encoding and end-side encryption, the public base layer carries the basic monitoring image, the partitioned private layer performs session key encryption according to the partition and capsule, and the traceability mark is written simultaneously as an invisible watermark and explicit field in the packet header, realizing a strong binding of public-private-traceability; On the terminal side, the consistency between the base layer watermark and explicit traceability is first verified, and then selective decryption and reverse disassembly are performed according to the capsule and partition permissions, and texture is superimposed only on the authorized area, while the unauthorized area maintains the public layer display. Attached Figure Description

[0055] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0056] Figure 1 A flowchart for a confidential video image acquisition method.

[0057] Figure 2 Generate a flowchart for initializing the configuration set.

[0058] Figure 3 A flowchart for forming a set of three-channel synchronization frames.

[0059] Figure 4 This is a flowchart of layered hybrid encoding and end-side encryption. Detailed Implementation

[0060] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0061] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.

[0062] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it a single or selective embodiment that is mutually exclusive with other embodiments.

[0063] Reference Figures 1-4 As an embodiment of the present invention, this embodiment provides a secure video image acquisition method, comprising the following steps:

[0064] S1. Establish an encrypted channel between the camera and the user terminal through a one-time physical pairing, generate the device root key, privacy partition template and time capsule table and write them to secure storage to obtain the initial configuration set.

[0065] S1.1. After the camera is powered on, it remains in an uninitialized state, only outputting a low-resolution prompt screen on the display interface. The prompt screen displays a QR code containing the camera's unique hardware identifier and the camera's factory public key. The user scans the QR code with a trusted user terminal near the camera, and obtains the camera's unique hardware identifier and the camera's factory public key for subsequent encrypted handshakes in the user terminal, triggering a one-time physical pairing request.

[0066] The user terminal uses the camera's unique hardware identifier and the camera's factory public key to send a one-time physical pairing request message to the camera via a short-range wireless link. After receiving the one-time physical pairing request message, the camera performs a handshake process based on the camera's factory public key. During the handshake process, encryption parameters are negotiated and the camera's unique hardware identifier and the user terminal's identity are confirmed. After the one-time physical pairing is completed, an encrypted channel is established between the camera and the user terminal, and a secure handshake result is generated at the camera end.

[0067] S1.2. The camera uses the secure handshake result to generate a set of device root keys that belong only to the camera in the security chip inside the camera, and writes the key derivation rules of the device root keys and the camera's unique hardware identifier into the camera's local secure storage to form the key base data.

[0068] The user terminal initiates a scene preview request to the camera through an established encrypted channel. The camera encodes the current shooting scene based on the key data and sends a low-resolution scene preview image in real time through the encrypted channel. The user terminal decrypts and displays the scene preview image. On the scene preview image, the user can directly define multiple privacy zones through touch operation, including areas where faces are likely to appear, workbench areas, doorway areas, etc., and set a corresponding privacy level for each privacy zone, generating a partition template with boundary information and privacy level labels. At the same time, on the same interface of the user terminal, the continuous time slice is divided into multiple time segments covering a predetermined time length, each time segment is defined as a time capsule, and an access policy is set for each time capsule, generating a time capsule table.

[0069] It should be noted that the scheduled time length refers to the configuration parameter explicitly set in the initialization interface, such as 5 minutes.

[0070] The partition template, time capsule table, and key basic data are combined and stored. The device root key, key derivation rules, camera unique hardware identifier, partition template, and time capsule table are written into the camera's local secure storage. At the same time, the user terminal saves the partition diagram corresponding to the partition template and the access policy description of each time capsule in its local storage. The combined storage result constitutes the initialization configuration set on the camera side.

[0071] S2. Based on the initial configuration set, each frame is simultaneously acquired through three channels: structure, texture, and source, and the capsule and partition numbers are labeled to form a set of three-channel synchronized frames.

[0072] S2.1. After initialization, the camera performs real-time acquisition, obtaining one original image frame at each sampling time. The original image of the current frame is recorded as the current frame image. The camera reads the partition template, time capsule table, camera's unique hardware identifier, and background reference image acquired during the initialization phase from the initialization configuration set. Combined with the current timestamp, the camera queries the time capsule table to find the time capsule number to which the current time point belongs. The retrieved time capsule number and the partition template are superimposed on the current frame image. In the image coordinate plane, the pixels are divided into normal areas and multiple privacy partitions according to the boundary area defined by the partition template. The corresponding partition number and the time capsule number to which the current frame belongs are recorded at each pixel position, generating a frame-level annotation map.

[0073] S2.2. Based on the frame-level annotation diagram, record the current frame as... The background reference image is denoted as The brightness difference is calculated at each pixel location, expressed as:

[0074] ;

[0075] in, In time At that time, the current frame's image is at coordinates The change in brightness relative to a static background is used to identify whether there is a moving target or a significant change in brightness at the current location; The coordinates of the current frame The pixel brightness value at that location is used to reflect the actual brightness of the scene at the current moment; As a background reference image in coordinates The pixel brightness value at that location is used to represent the static background brightness baseline recorded during the initialization phase; This is the time index of the current frame.

[0076] The brightness difference is compared with a preset brightness change threshold to obtain the motion region mask. Within the motion region, according to the normal region and privacy partition division results given by the frame-level annotation map, a preset edge extraction rule is applied. The neighborhood gradient is calculated by subtracting the edge strength from the horizontal and vertical gradients. The expression is as follows:

[0077] ;

[0078] in, In time Time, coordinates Edge strength values ​​at certain points highlight structural contour information. , , , In time At that time, the current frame is in The pixel brightness values ​​of neighboring locations represent the brightness gradient of the current location in the horizontal and vertical directions.

[0079] By combining the brightness difference and edge intensity values, the contour, moving contour block and brightness contrast results are marked at each pixel position, and the corresponding partition number and time capsule number are written to generate the structure channel frame.

[0080] It should be noted that the preset brightness change threshold is a configuration parameter explicitly set in the initialization interface, in units of 8-bit brightness levels (0–255), with a value of 16. The value is based on the fact that under common indoor illumination and camera light-sensing noise conditions, the background noise is usually less than 8 brightness levels. To avoid false detection, the threshold is fixed at twice the upper limit of the background noise to distinguish between background changes and effective motion. When the brightness difference is greater than or equal to the brightness change threshold, the current pixel position is marked as the foreground of the moving area; otherwise, it is marked as the static background.

[0081] S2.3. Based on the geometric boundaries of the partitions recorded in the structural channel frame and frame-level annotation diagram, perform texture clipping operation for each privacy partition in the original image of the current frame. Extract image blocks containing complete texture information from the original image according to the geometric boundaries of the privacy partitions. Represent the positional relationship of each image block in the original image as a position number. The position number includes the starting coordinates and size information of the image block in the current frame. Add a partition number and a time capsule number to each image block to form a texture patch list.

[0082] Read the camera's unique hardware identifier and the summary information generated in the previous frame. Concatenate the camera's unique hardware identifier, the current frame number, the current time capsule number, each partition number, and the summary information of the previous frame according to the preset field order to form a traceability tag for the current frame. At the same time, write the partition number and the time capsule number in the partition position of the traceability map to form the traceability map.

[0083] It should be noted that the preset field order is a fixed standard, namely, camera unique hardware identifier, frame number, time capsule number, partition number list length, partition number list, and previous frame summary information. This information is fixed as the basis for parsing in the initial configuration set to ensure consistency between writing and parsing.

[0084] S2.4. Using the current frame number as an index, perform a synchronization check on the structure channel frame, texture list, and source map. Use the partition number and time capsule number fields in the source map to check whether the partition number and time capsule number in the structure channel frame are consistent. Use the position mark in the source map and the partition boundary in the frame-level annotation map to check whether the position number in the texture list corresponds to the contour area in the structure channel frame. If the check passes, combine and store the structure channel frame, texture list, and source map of the current frame according to the unified frame number and time capsule number to form a three-channel synchronization frame set.

[0085] S3. Perform layered hybrid coding and end-side encryption on the three-channel synchronous frame set, and write the traceability mark into the public base layer and the partition private layer at the same time, and output the three-layer hybrid bitstream.

[0086] S3.1. Read the structure channel frame, frame-level annotation map, and motion region mask obtained from the three-channel synchronous frame set according to the frame sequence number. At the same time, read the background reference map from the initialization configuration set. Combined with the structure channel frame, generate a basic monitoring frame in the pixel coordinate plane of the current frame. The expression is:

[0087] ;

[0088] in, Based on the monitoring frame in coordinates The pixel brightness value at that location serves as the image content of the common base layer. For the motion region mask in time Time, coordinates The value at this location is 1, indicating that the current position belongs to the moving area, and 0, indicating that the current position belongs to the static background area. For the structure channel frame in coordinates The pixel brightness value at that location is used to represent contour information and brightness contrast information. As a background reference image in coordinates The pixel brightness value at that location is used to restore the static background image recorded during the initialization phase in a static area.

[0089] S3.2. The basic monitoring frame is used as the input to the video encoder to compress the basic monitoring frame and generate a common base layer bitstream containing only outline and background information. The partition number and time capsule number retained in S3.1 are written into the frame header of the common base layer bitstream to generate the common base layer corresponding to the current frame. The common base layer bitstream is input into the hash algorithm in byte order from frame header to frame tail to calculate the hash value of the entire common base layer bitstream. The obtained hash value is used as the base layer digest and recorded together with the base layer bitstream as the base layer encoding result.

[0090] The texture patch list and frame-level annotation map are read from the three-channel synchronous frame set according to the same frame sequence number. The partition template and time capsule table are read from the initialization configuration set. The camera searches for the shuffling order pre-configured for each privacy partition in the partition template, rearranges the image blocks belonging to the same privacy partition in the texture patch list according to the shuffling order, and obtains the rearranged texture patch sequence. The corresponding session key is selected according to the time capsule number to which the current frame belongs recorded in the time capsule table.

[0091] For each privacy partition, the rearranged texture sequence is encrypted using a session key. The encrypted texture, along with the corresponding partition number, time capsule number, and location number, is encapsulated into partition-level data. The partition-level data of all privacy partitions are combined to form the partition privacy layer bitstream and partition privacy layer digest of the current frame.

[0092] S3.3. While generating the partitioned privacy layer bitstream, read the source map from the three-channel synchronous frame set according to the current frame sequence number. Based on the camera's unique hardware identifier, current frame sequence number, current time capsule number, partition number, and previous frame summary information recorded in the source map, compress the source map into source marker image data, extract the source tag field from the source map, and combine the source tag field with the current frame's base layer summary and partitioned privacy layer summary to form the source marker content for the current frame. Organize the compressed source marker image data and the source marker content into a source marker block.

[0093] A concatenation operation is performed using the base layer summary, the camera's unique hardware identifier, and the previous frame's chained tracing tag field to form the current frame's chained tracing tag. The tracing tag content is extracted from the tracing tag block, and the encoding positions corresponding to low and medium frequencies are selected in the compressed data of the public base layer bitstream and written into the chained tracing tag according to the invisible watermark embedding rules. The correspondence between the chained tracing tag, partition number, and time capsule number is recorded in the description information of the public base layer, resulting in a watermarked public base layer bitstream. The chained tracing tag and explicit tracing information in the tracing tag content are written into the packet header field of the partitioned private layer bitstream. The chained tracing tag, partition number, and time capsule number are recorded in the header of each partition-level data, resulting in a tracing-consistent partitioned private layer bitstream.

[0094] Compressed data with traceability marker blocks is added to the bypass of the watermarked public base layer bitstream and the traceable private partition bitstream. The data is then frame-level multiplexed and encapsulated in time order to form a three-layer hybrid bitstream.

[0095] S4. Verify the consistency between the base layer watermark and explicit source tracing for the three-layer hybrid bitstream, decrypt by time capsule and partition permissions, and reverse break down the partition private layer to obtain the reconstructed video stream.

[0096] S4.1. The remote client receives the three-layer hybrid stream output by the camera from the cloud in chronological order. The cloud only forwards the three-layer hybrid stream without decryption. The client separates the common base layer stream, the partition private layer stream, and the traceability tag block from the three-layer hybrid stream. It decodes the common base layer stream to recover the basic monitoring frame containing only outline and background information. During the decoding process, it calls the invisible watermark extraction algorithm to extract the chain traceability tag and the mapping relationship between the partition number and the time capsule number recorded in the common base layer from the basic monitoring frame, and generates the base layer verification information.

[0097] The header of the private layer bitstream is parsed to read explicit tracing information and chained tracing tags. The explicit tracing information includes the camera's unique hardware identifier, the current frame number, the current time capsule number, and a list of partition numbers. The parsing result serves as the private layer verification information. The client compares the chained tracing tags in the basic layer verification information with those in the private layer verification information, and verifies whether the partition numbers and time capsule numbers recorded in both correspond one-to-one. If the chained tracing tags match and the mapping relationship between the partition numbers and time capsule numbers is consistent, a consistency verification result is generated indicating that the public basic layer and the private layer have the same source and have not been tampered with. From this result, a set of candidate decryptable partitions covering all partition numbers and time capsule numbers of the current frame is compiled.

[0098] S4.2. After obtaining the candidate decryptable partition set, the client confirms the traceability credibility of the three-layer hybrid bitstream based on the consistency verification result. According to the current logged-in user's identity and the access policy configured in the initialization phase, the client retrieves the session key set corresponding to the capsule number of the current frame's time period from the local security environment or the user terminal security environment, and reads the saved partition diagram and access policy description. Through the partition permission information recorded in the access policy description, the client matches the partition numbers in the candidate decryptable partition set with the allowed access partition numbers to form a decryption authorization list.

[0099] The client iterates through the private layer bitstream of the partitions according to the decryption authorization list. For the partition-level data in the private layer bitstream whose partition number belongs to the decryption authorization list, the client performs partition-level decryption operation. The client decrypts the ciphertext texture sequence in the corresponding partition-level data using the session key specified in the decryption authorization list and the same symmetric encryption algorithm as the encryption stage, to obtain the decrypted texture sequence of each authorized partition. The client also retains the partition number, time capsule number and location number carried in the partition-level data for each decrypted texture sequence, forming a decrypted texture list.

[0100] S4.3. After obtaining the list of decrypted texture pieces, the client retrieves the partition template and the shuffling order information recorded in the partition template from the initialization configuration set. Using the partition number and position number in the list of decrypted texture pieces as indexes, the client generates a position mapping table corresponding to the shuffling order for each partition. The client performs reverse rearrangement of the decrypted texture piece sequence of each partition according to the position mapping table, puts each decrypted texture piece back into the original position marked in the position mapping table, and arranges them in order from left to right and from top to bottom according to the position number, restoring the texture layout of each privacy partition in the current frame.

[0101] After reversing all partitions, a partition-level mosaic is generated for each partition. The partition-level mosaic corresponds to the mosaic result of all restored texture pieces in a privacy partition, and retains the partition number and time capsule number identifier. The client performs boundary alignment and number verification on all partition-level mosaics to confirm that the partition number and time capsule number in the partition-level mosaic are consistent with the records in the decrypted texture piece list and partition template, thus obtaining the set of restored texture pieces for the current frame.

[0102] S4.4. After the client completes the restoration of the texture patch set, it calls the basic monitoring frame. Based on the partition boundaries recorded in the partition template and the partition number and position coordinates recorded in the restored texture patch set, it performs a partition-by-partition overwrite operation on the basic monitoring frame. The restored texture patch area corresponding to each partition-level mosaic is written into the partition area at the same position in the common base layer. The outline and background display in the basic monitoring frame remain unchanged in the partition positions that have not been authorized for decryption. A fusion reconstruction frame for the current user is generated. The fusion reconstruction frame retains both the structural outline of the undecrypted partition and the texture details after decryption and inverse disassembly.

[0103] The client arranges each fused and reconstructed frame in ascending order of frame number, and encapsulates the continuous fused and reconstructed frames in chronological order to output a reconstructed video stream. During playback or real-time display, the reconstructed video stream can be overlaid with the camera's unique hardware identifier, chain-like traceability tags, and time capsule number sequence recorded in the traceability marker. This ensures controlled access to privacy areas while presenting users with customized visual video content based on zones and time capsules.

[0104] This embodiment also provides a secure video image acquisition system, including:

[0105] The pairing key module establishes an encrypted channel between the camera and the user terminal through a one-time physical pairing, generates the device root key, privacy partition template and time capsule table and writes them into secure storage to obtain the initial configuration set;

[0106] The synchronous acquisition module performs three-channel synchronous acquisition of each frame of the image based on the initial configuration set, namely structure, texture and source, and labels the capsule and partition number to form a set of three-channel synchronous frames.

[0107] The layered encryption module performs layered hybrid encoding and end-side encryption on the three-channel synchronous frame set, and writes the traceability mark into the public base layer and the partitioned private layer at the same time, and outputs a three-layer hybrid bitstream.

[0108] The verification and reconstruction module verifies the consistency between the base layer watermark and explicit source tracing of the three-layer hybrid bitstream, decrypts the data according to time capsules and partition permissions, and reverse-disassembles the partition privacy layer to obtain the reconstructed video stream.

[0109] This embodiment also provides a computer device applicable to a confidential video image acquisition method, comprising: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement the confidential video image acquisition method proposed in the above embodiment.

[0110] The computer device can be a terminal, comprising a processor, memory, communication interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, carrier networks, NFC (Near Field Communication), or other technologies. The display screen can be an LCD screen or an e-ink screen. The input devices can be a touch layer covering the display screen, buttons, a trackball, or a touchpad on the computer device's casing, or an external keyboard, touchpad, or mouse.

[0111] This embodiment also provides a storage medium storing a computer program, which, when executed by a processor, implements the secure video image acquisition method proposed in the above embodiments. The storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read Only Memory (EPROM), Programmable Red-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.

[0112] In summary, this invention forms an initial configuration set uniquely bound to the device through one-time physical pairing and device root key derivation, combined with privacy partition templates and time capsule tables, ensuring that each subsequent frame has a defined partition and time semantics. In three-channel synchronous acquisition, the structure channel retains only contour and brightness contrast for public visibility, the texture channel strictly corresponds to the partition and capsule number, and presets granular boundaries for authorized reconstruction. The tracing channel writes the device identifier, partition number, and previous frame summary at the frame level, ensuring continuity of the source chain. In layered hybrid encoding and end-side encryption, the public base layer carries the basic monitoring image, the partitioned private layer performs session key encryption according to the partition and capsule, and the tracing mark is written simultaneously as an invisible watermark and explicit field in the packet header, achieving a strong binding of public-private-tracing. On the terminal side, the consistency between the base layer watermark and explicit tracing is first verified, and then selective decryption and reverse disassembly are performed according to capsule and partition permissions. Texture is superimposed only on authorized areas, while unauthorized areas maintain the public layer display.

[0113] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.

Claims

1. A secure video image acquisition method, characterized in that: include: An encrypted channel between the camera and the user terminal is established through a one-time physical pairing, generating the device root key, privacy partition template, and time capsule table and writing them into secure storage to obtain the initial configuration set; The specific steps to obtain the initialization configuration set are as follows: After the camera is powered on, it outputs a low-resolution prompt screen and displays a QR code containing a unique hardware identifier and the factory public key. The user terminal scans the code to complete a one-time physical pairing and establish an encrypted channel to obtain a secure handshake result. The camera uses the secure handshake result to generate a device root key and embeds the derivation rules and device identifier into the security chip to form the basic key data. The user terminal uses the key-based data to display the scene preview and delineate privacy partitions and privacy levels on the preview screen. In the same interface of the user terminal, the continuous time slice is divided into multiple time segments covering a predetermined time length. Each time segment is defined as a time capsule and an access policy is set for each time capsule, generating a time capsule table and a partition template. Write the partition template and time capsule table to secure storage, and at the same time save the access policy on the user terminal to form an initialization configuration set; Based on the initial configuration set, each frame of the image is simultaneously acquired in three channels: structure, texture, and source, and labeled with capsule and partition numbers to form a set of three-channel synchronous frames. Layered hybrid coding and end-side encryption are performed on the three-channel synchronous frame set, and the traceability mark is written into the public base layer and the partition private layer at the same time, and the three-layer hybrid bitstream is output. The consistency between the base layer watermark and explicit source tracing is verified for the three-layer hybrid bitstream. The private layers of the partitions are decrypted according to time capsules and partition permissions and the partition privacy layers are reversed to obtain the reconstructed video stream. The specific steps for simultaneously writing the traceability marker into both the public base layer and the partitioned private layer are as follows: The base layer encoding results are used to generate a base layer summary, which is then concatenated with the device identifier and the previous frame tag to form a chain-like traceability tag. By embedding chain-based traceability tags into low-to-medium frequency locations in the public base layer and recording their mapping relationship with partition numbers and capsule numbers, a watermarked public base layer is obtained. Based on the set of traceability graphs to be written with traceability tags, traceability tags are established, and chained traceability tags and explicit traceability information are written into the header of the partition privacy layer to obtain a traceability-consistent partition privacy layer. The watermarked public base layer, the traceable private partition layer, and the traceability marker block are combined and encoded to output a three-layer hybrid bitstream.

2. The secure video image acquisition method as described in claim 1, characterized in that: The specific steps for forming the three-channel synchronization frame set are as follows: Read the initialization configuration set, locate the normal area and privacy area in each frame according to the partition template, and attach the capsule number of the current time period to obtain the frame-level annotation map; Based on the frame-level annotation map, the ordinary areas and privacy partitions defined in the current frame are processed using preset edge extraction rules and brightness contrast, and the partition number and capsule number are recorded to obtain the structure channel frame; Cut out the complete texture blocks of the privacy partition from the original image corresponding to the structure channel frame and label the position numbers to obtain a list of texture patches; Attach the device identifier and previous frame summary to the texture list and write the partition number and capsule number to generate the source map; The source graph is used to perform time synchronization and number verification on the structural channel frames and texture patch list, forming a three-channel synchronized frame set.

3. The secure video image acquisition method as described in claim 1, characterized in that: The specific steps for performing layered hybrid encoding and end-side encryption are as follows: The camera receives a set of three synchronized frames, retains only the outline and brightness information in the structure channel, and fills the static area with a background reference image to obtain the basic monitoring frame; The basic monitoring frames are sent to the video encoder for compression to generate a common base layer while retaining the partition number and capsule number, thus obtaining the base layer encoding result. For texture pieces in the three-channel synchronous frame set, the texture pieces are rearranged according to the shuffling order in the partition template and the session key is selected according to the capsule number. The texture pieces are then encrypted and packaged at the partition level to form a partition privacy layer. The source graph content corresponding to the three-channel synchronization frame set is retained and compressed to obtain the source graph set to be written with the source marker.

4. The secure video image acquisition method as described in claim 1, characterized in that: The specific steps for verifying the consistency between the base layer watermark and explicit source tracing of the three-layer hybrid bitstream are as follows: After receiving the three-layer mixed bitstream, the terminal decodes the common base layer and extracts the invisible watermark to obtain the base layer verification information. Explicit traceability information is parsed from the private layer header of the partition and chain traceability tags are extracted to form private layer verification information; The chain traceability tags in the basic layer verification information are compared with the mapping relationship between the chain traceability tags, partition numbers and capsule numbers in the private layer verification information to obtain the consistency verification result; Based on the consistency verification results, mark the decryptable partition set and the non-decryptable partition set, and generate the decryption preparation result.

5. The secure video image acquisition method as described in claim 1, characterized in that: The specific steps for decrypting capsules and partition permissions by time period are as follows. Based on the decryption preparation results, load the session key set corresponding to the current capsule number and match it with the partition number in the decryptable partition set to form a decryption authorization list; According to the decryption authorization list, perform partition-level decryption on the authorized partitions in the private layer of the partition and retain the partition number and position number of the decrypted texture pieces to obtain a list of decrypted texture pieces.

6. The secure video image acquisition method as described in claim 1, characterized in that: The specific steps for reversing the partitioning of the private layer are as follows. Based on the shuffling order of the decrypted texture list and partition template record, the decrypted textures are reverse-shuffling and restored to their original arrangement according to the partition number and position number to obtain the restored texture set; The restored texture patch set is overlaid onto the corresponding partition position in the common base layer, while keeping the partition number and capsule number consistent, to form a fused reconstruction frame; The merged and reconstructed frames are arranged into a frame sequence according to time order, and the reconstructed video stream is output.

7. The secure video image acquisition method as described in claim 6, characterized in that: The specific steps for performing inverse shredding on the decrypted texture fragment are as follows. A position mapping table is generated based on the shuffling order of the partition template records and aligned item by item with the list of decrypted texture pieces to obtain the mapping alignment result; Based on the mapping alignment results, a blank splicing canvas is created for each partition, and the corresponding decrypted texture pieces are placed sequentially according to the position number to obtain a partition-level splicing image; Boundary merging and number verification are performed on the partition-level mosaic map to obtain the restored texture patch set.

8. A secure video image acquisition system, based on the secure video image acquisition method according to any one of claims 1 to 7, characterized in that: include: The pairing key module establishes an encrypted channel between the camera and the user terminal through a one-time physical pairing, generates the device root key, privacy partition template and time capsule table and writes them into secure storage to obtain the initial configuration set; The synchronous acquisition module performs three-channel synchronous acquisition of each frame of the image based on the initial configuration set, namely structure, texture and source, and labels the capsule and partition number to form a set of three-channel synchronous frames. The layered encryption module performs layered hybrid encoding and end-side encryption on the three-channel synchronous frame set, and writes the traceability mark into the public base layer and the partitioned private layer at the same time, and outputs a three-layer hybrid bitstream. The verification and reconstruction module verifies the consistency between the base layer watermark and explicit source tracing of the three-layer hybrid bitstream, decrypts the data according to time capsules and partition permissions, and reverse-disassembles the partition privacy layer to obtain the reconstructed video stream.

Citation Information

Patent Citations

  • Safe traceable video coding and decoding system

    CN117278762A

  • Electronic voting using secure electronic identity device

    EP3145114A1