Video image processing method, apparatus, device, and storage medium

By encrypting the target area and implementing access control in the video surveillance system, the balance between privacy protection and target recognition is resolved, achieving the effect of accurately identifying targets while satisfying privacy protection.

CN120850325BActive Publication Date: 2025-12-30GUANGDONG INTELL VISION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511353998.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-22
Publication Date
2025-12-30
Estimated Expiration
2045-09-22

AI Technical Summary

Technical Problem

Existing technologies struggle to balance the needs of privacy protection and target recognition in video surveillance systems, failing to accurately identify targets while protecting the privacy of video data.

Method used

By identifying target regions in video images, encrypting the target regions, and storing the encrypted video data on a first server, the target pixel values ​​are recorded, encrypted, and stored on a second server. Access control is supported, allowing high-privilege users to decrypt and restore the original video data.

Benefits of technology

It achieves accurate reconstruction of the target area while satisfying privacy protection, meeting the requirements of privacy processing and target recognition, improving data security and processing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120850325B_ABST
    Figure CN120850325B_ABST
Patent Text Reader

Abstract

The application provides a video image processing method and device, equipment and a storage medium, relates to the technical field of image processing, and solves the problem that the video processed by privacy protection in the related art is difficult to meet the target recognition requirement. In the scheme, the target region in the image is encrypted, thereby obtaining the video data after the target region is encrypted and storing the video data in a first server. The target pixel value of the pixel point associated with the target region in the original video image and the encrypted pixel point is encrypted and stored in a second server. The scheme can further encrypt the corresponding data after the target region is encrypted, thereby improving the data security and accurately restoring the target region.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and in particular to a video image processing method, apparatus, device and storage medium. Background Technology

[0002] In video acquisition systems such as video surveillance systems and video recording systems, video data is acquired through video acquisition devices and stored on a server. When there is a need to query the video data, it is necessary to read the video data from the server and replay the original video footage. For example, in a video surveillance system monitoring a warehouse, if it is necessary to query the personnel entering and leaving the warehouse, it is necessary to display the faces in all video frames corresponding to the video data to identify the personnel entering and leaving the warehouse.

[0003] However, as people become more aware of privacy, video footage containing faces, tattoos, or other sensitive areas requires privacy protection measures. For example, faces can be blurred in video frames to prevent facial information leakage. While this method satisfies privacy requirements, it struggles in scenarios requiring raw video footage for target recognition, as it cannot provide clear images of faces. Current technology fails to balance the conflicting needs of privacy protection and target recognition. Summary of the Invention

[0004] This application provides a video image processing method, apparatus, device, and storage medium, which solves the problem that videos processed with privacy protection in related technologies are difficult to meet the requirements of target recognition. This solution can encrypt the target area and then further encrypt the corresponding data, thereby improving data security while accurately restoring the target area.

[0005] In a first aspect, this application provides a video image processing method applied to a video acquisition device in a video acquisition system. The video acquisition system includes a video acquisition device, a first server, and a second server. The video acquisition device is communicatively connected to both the first server and the second server. The method includes:

[0006] Extract video images from the first acquired video data and determine the image data corresponding to the target region in the video image. The target region is the image region set in the video image corresponding to the target object.

[0007] The image data is encrypted, the first encrypted data is determined, and the encrypted video image of the target area is obtained.

[0008] Based on all the video images encrypted in the target area, the second video data is re-synthesized and stored on the first server. The second video data is used for playback according to the playback request.

[0009] Based on a preset encryption algorithm, the target pixel values ​​associated with the image data and the first data are encrypted to obtain encrypted second data, which is then stored on a second server. The second data is used to decrypt the second video data according to a processing request to obtain the first video data.

[0010] Secondly, this application also provides a video image processing apparatus, which is applied to a video acquisition device in a video acquisition system. The video acquisition system includes a video acquisition device, a first server, and a second server. The video acquisition device is communicatively connected to the first server and the second server, respectively. The video image processing apparatus includes:

[0011] The data extraction module is configured to extract video images from the acquired first video data and determine the image data corresponding to the target area in the video image, wherein the target area is the image area set in the video image corresponding to the target object;

[0012] The first encryption module is configured to encrypt image data, determine the encrypted first data, and obtain the encrypted video image of the target area.

[0013] The image synthesis module is configured to re-synthesize second video data based on all encrypted video images of the target area and store the second video data on the first server. The second video data is used for playback according to playback requests.

[0014] The second encryption module is configured to encrypt the target pixel values ​​associated with the image data and the first data based on a preset encryption algorithm to obtain encrypted second data, and store the second data on a second server. The second data is used to decrypt the second video data to obtain the first video data according to a processing request.

[0015] Thirdly, this application also provides an electronic device comprising:

[0016] One or more processors;

[0017] A storage device is provided for storing one or more programs, which, when executed by one or more processors, enable the one or more processors to implement the video image processing method of this application.

[0018] Fourthly, this application also provides a storage medium for storing computer-executable instructions, which, when executed by a processor, are used to perform the video image processing method of this application.

[0019] This application's solution defines target regions for objects requiring privacy protection in video images and encrypts these target regions to obtain encrypted video data that meets privacy requirements. This encrypted data is stored on a first server. Furthermore, it records the target pixel values ​​of pixels in the original video image associated with the target region and the encrypted pixels, and then encrypts these target pixel values ​​and stores them on a second server. This facilitates subsequent reconstruction of the original video image for target recognition. Therefore, this solution improves data security while accurately reconstructing the target region, satisfying both privacy requirements and target recognition needs. Attached Figure Description

[0020] Figure 1 A flowchart illustrating the steps of a video image processing method provided in an embodiment of this application.

[0021] Figure 2 This is a schematic diagram illustrating the steps of encrypting image data according to an embodiment of this application.

[0022] Figure 3 This is a flowchart illustrating the processing of video images according to an embodiment of this application.

[0023] Figure 4 This is a schematic diagram of the structure of a video image processing apparatus provided in an embodiment of this application.

[0024] Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0025] The embodiments of this application will be further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are merely illustrative of the embodiments of this application and are not intended to limit the scope of this application. Furthermore, it should be noted that, for ease of description, the accompanying drawings only show the parts related to the embodiments of this application, not all structures. Those skilled in the art, after reading this specification, should be able to conceive that any combination of technical features can constitute an optional implementation method, provided that the technical features do not contradict each other.

[0026] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects, not to describe a specific order or sequence. It should be understood that such use of data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class, not limited in number; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship. In the description of this application, "multiple" means two or more, and "several" means one or more.

[0027] In video capture scenarios, such as monitoring specific locations or filming people, video capture systems acquire video data through video capture devices and store it on a server. In some scenarios, there is a need to query this video data, such as retrieving video data and playing back the original video footage. For example, in a video surveillance system monitoring a warehouse, to query who entered and exited the warehouse, it is necessary to display the faces in all video frames corresponding to the video data to identify the individuals who entered and exited the warehouse.

[0028] However, as people become more aware of privacy, they demand privacy protection measures for sensitive areas in video footage, such as faces, tattoos, and objects that need to be hidden. Current technologies typically use mosaic effects to address this, but when the original video footage needs to be played back, it's difficult to extract the original footage from the privacy-protected video data. For example, unmasked faces cannot be clearly obtained, making facial recognition difficult.

[0029] To address this issue, this application provides a video image processing method. This method is applied to a video acquisition device within a video acquisition system. The video acquisition device is an electronic device capable of recording video and processing its data, such as a video recorder or personal computer. The video acquisition system includes a video acquisition device, a first server, and a second server. The video acquisition device is communicatively connected to both the first and second servers to store relevant data.

[0030] Figure 1 This is a flowchart illustrating the steps of a video image processing method provided in an embodiment of this application. The method is applied to a video acquisition device, processes the acquired video data to meet user privacy requirements, and can also quickly restore the original video frame, facilitating the identification of target objects from the restored video frame. The specific steps are as follows:

[0031] Step S110: Extract video images from the acquired first video data and determine the image data corresponding to the target area in the video image. The target area is the image area set in the video image corresponding to the target object.

[0032] The first video data is unprocessed video data. All video images are extracted from the acquired first video data, and the location of a target object is identified in different frames of the video images. The target object is a pre-defined object to be identified, such as a face or other parts or objects requiring privacy protection. A corresponding image region is defined in the video image corresponding to the target object, and this region serves as the target area, which covers the location of the target object. Optionally, the target area is a pre-defined region on the video image corresponding to the location of the target object, and its shape can be rectangular. Furthermore, image data corresponding to the target area is determined in the video image. Optionally, the image data corresponds to the target area where the target object is located, and the image data includes the pixel values ​​of the pixels within that area, the frame number of the corresponding video image, etc. In one embodiment, the location of the target object is determined during the identification process to determine the position of the target area on the video image, and the corresponding image data is extracted from the video image. Optionally, the identification of targets can employ corresponding target detection algorithms, such as the Haar Cascades algorithm, R-CNN (Region-based Convolutional Neural Network), etc., to enable the identification of targets such as faces and specific objects.

[0033] Step S120: Encrypt the image data, determine the first encrypted data, and obtain the encrypted video image of the target area.

[0034] For each frame of video image, the image data of the corresponding target area is encrypted to obtain encrypted first data. This first data is then used to further process the target area in the video image, thereby combining it with the first data to encrypt the target area and obtain the encrypted video image. The encryption process is based on image data and can be performed by changing the position or value of pixels, thus converting identifiable content containing valid information into indistinguishable content. From the user's perspective, the target area in the video image changes from a clearly identifiable object to unidentifiable content.

[0035] In some embodiments, image data includes pixel values ​​of all pixels within the target area. For example, for black and white video, the pixel values ​​of the pixels in the image data extracted from each frame of video are grayscale values. For color video, the pixel values ​​of the pixels in the image data extracted from each frame of video are RGB (Red, Green, Blue) values, which are the intensity values ​​of the three different color channels, red, green, and blue, respectively, in either the RGB color mode or the RGBA (Red, Green, Blue, Alpha (corresponding to opacity)) color mode.

[0036] Step S130: Based on all the video images encrypted in the target area, re-synthesize the second video data and store the second video data on the first server. The second video data is used for playback according to the playback request.

[0037] After encrypting all video images, a new video data set is synthesized, resulting in the second video data. It's conceivable that the target areas in the video frame of this second video data are encrypted during playback. This second video data is stored on a first server, which can respond to playback requests from other devices or its own device by sending the second video data or playing it locally. For example, in a video surveillance scenario, using faces as the target object, the target areas corresponding to faces in each frame of the acquired video data (i.e., the first video data) are encrypted to obtain the second video data. Playing this second video data can be used to display the video frame of the currently monitored area to the user, and all faces in the video frame are encrypted to protect the privacy of others. Based on this, in practical applications, appropriate permissions can be set, such as a low-priority first permission, so that video playback devices logged into user accounts with first-priority permissions can only view the second video data, i.e., obtain the second video data from the first server, and can only provide users with the video frame containing face encryption, thereby protecting the privacy of others.

[0038] Step S140: Based on a preset encryption algorithm, the target pixel value associated with the image data and the first data is encrypted to obtain the encrypted second data, and the second data is stored on the second server. The second data is used to decrypt the second video data to obtain the first video data according to the processing request.

[0039] Furthermore, the target pixel values ​​associated with the image data and the first data are encrypted. For example, in one embodiment, the target pixel values ​​are obtained by performing matrix operations on a first matrix constructed from the image data and a second matrix constructed from the first data. These matrix operations include matrix addition, matrix subtraction, matrix multiplication, and matrix inversion. For instance, a cell can be represented in matrix form. In the first matrix corresponding to the image data, each element corresponds to a pixel, and the value of the element is the value of the corresponding pixel. Similarly, in the second matrix corresponding to the first data, each element corresponds to a pixel. It is understood that for the same cell, there exists a first matrix and a second matrix, and the pixel values ​​corresponding to the elements in the first matrix are the same as those corresponding to the elements in the second matrix. That is, the element in the first row and first column of the first matrix corresponds to the pixel in the first row and first column of the cell, and the element in the first row and first column of the second matrix also corresponds to the pixel in the first row and first column of the cell. However, it is worth noting that since the first data is encrypted, the values ​​of the elements in the second matrix differ from the corresponding values ​​of the elements in the first matrix. Therefore, matrix operations such as matrix subtraction and matrix addition can be used to process the first and second matrices to determine the target pixel values ​​of each element, thereby obtaining a matrix containing the target pixel values ​​of each element, and then encrypting it to obtain the second data.

[0040] Furthermore, the second data is stored in a second server. This second server can respond to processing requests from other devices or its own devices, thereby sending the second data to other devices so that they can use the second data to decrypt the second video data and obtain the first video data. Similarly, the second server can also use the second data to decrypt the second video data and obtain the first video data. That is, after parsing the second data to obtain the target pixel values ​​associated with the image data and the first data in each frame of encrypted video image, the encrypted video image of each frame is decrypted according to the corresponding target pixel values. For example, the target pixel value of each pixel is restored by comparing it with the corresponding pixel value in the first data, thereby restoring the pixel value of the same pixel in the video image extracted from the first video data, that is, obtaining the original pixel value, and then restoring the original pixel values ​​of all pixels in the target area, thereby obtaining each frame of video image in the first video data.

[0041] Optionally, in a video surveillance scenario, a face is used as the target object. After privacy processing, second video data is obtained and stored on the first server, while the second data is stored on the second server. In practical applications, appropriate permissions can be set, such as high-priority second permissions, so that video playback devices logged into with user accounts holding second permissions can decrypt and obtain the first video data. That is, they can obtain the second video data from the first server and the second data from the second server. By parsing the second data, the target pixel values ​​of each pixel in the target area are obtained, which are associated with the image data and the first data. Based on these values, the original pixel values ​​of each pixel in the target area in each frame of the second video data are restored, thus obtaining the pixel values ​​of the corresponding pixels in the video image of the first video data. This allows the face to be restored, resulting in the first video data captured by the video acquisition device, and the face in the video frame is restored for the purpose of identifying the target person. The target person is the person to be identified in this scenario. To address this, by further encrypting the second data and storing it on a second server, this solution better protects privacy. Even if other devices obtain the second video data, they cannot reconstruct the original video footage from a single data source. Thus, while meeting privacy protection requirements, it also allows for the reconstruction of the original video footage for target identification. This process is simple and yields good reconstruction results. Moreover, the encrypted second data, being a text file, occupies less memory compared to the original image, improving processing efficiency.

[0042] As described above, this solution sets target regions for objects requiring privacy protection in video images and encrypts these target regions to obtain encrypted video data that meets privacy requirements. This encrypted data is then stored on a first server. Furthermore, it records the target pixel values ​​of the pixels in the original video image associated with the target region and the encrypted pixels, and encrypts these target pixel values ​​before storing them on a second server. This facilitates subsequent reconstruction of the original video image for target recognition. Therefore, this solution improves data security while accurately reconstructing the target region, satisfying both privacy requirements and target recognition needs.

[0043] Figure 2 This is a schematic diagram illustrating the steps of encrypting image data according to an embodiment of this application, as shown below. Figure 2 As shown, offset operations are performed on each pixel within the target area to obtain the corresponding offset value. The offset operation is used to recalculate the pixel value for each pixel. The specific steps are as follows:

[0044] Step S210: Based on the pixel values ​​of all pixels in the target area, determine the offset value of each pixel after offset operation, and use the offset values ​​of all pixels as the first data.

[0045] Step S220: Change the pixel values ​​of all pixels in the target area according to the first data to obtain the encrypted target area, and then obtain the encrypted video image of the target area.

[0046] Understandably, after determining the pixel values ​​of all pixels within the target area, offset operations are performed on each pixel. During these offset operations, a new pixel value is determined for each pixel. For example, calculations are performed using multiple pixels in the pixel's neighborhood to obtain a new pixel value, which is then used as the offset value for that pixel. This process is repeated to obtain the offset values ​​for all pixels within the target area. These offset values ​​are then used as the first set of data, and the pixel values ​​of all pixels within the target area are changed accordingly to obtain the encrypted target area, thereby acquiring the encrypted video image of the target area.

[0047] In one embodiment, during the encryption process of pixels within a target area, the target area corresponding to the image data can be divided into multiple cells, with each cell processed separately to encrypt the pixel value of each pixel within each cell. Specifically, the target area is divided into multiple cells, and each cell contains multiple pixels; that is, each cell serves as a sub-region of the target area. Optionally, the multiple cells can be regions with the same shape and area, or the regions corresponding to each cell may have different shapes and contain different numbers of pixels.

[0048] For example, if the target area is a square region of 480×480 pixels, and the divided cells are square regions of 16×16 pixels, then 900 identical cells can be divided within the target area. It is conceivable that the more cells divided, the more refined the division of the entire target area, and the better the encryption effect. Therefore, the number of cells divided is set according to the actual encryption requirements.

[0049] Then, the pixel values ​​of all pixels within each cell are determined, and the corresponding average is calculated, which serves as the offset value to determine the first data. For example, in a video image extracted from a black and white video, the pixel values ​​of each pixel in the acquired image data are grayscale values. Assuming each cell is a 16×16 pixel square area, the average grayscale value of the 256 pixels within the cell is calculated. This average is used as the offset value for all pixels within that cell. The first data for the target area is then obtained by calculating the offset values ​​of the corresponding pixels in all cells. It's conceivable that if the target area is divided into 900 cells, the first data would include the offset values ​​of the pixels corresponding to all 900 cells.

[0050] Optionally, if the pixel value is an RGB value (i.e., the video image is a color image), the target area is divided into multiple cells, and the average intensity value of each color channel corresponding to all pixels in each cell is determined. For example, taking a cell containing N pixels as an example, pixel P1 corresponds to pixel values ​​(120, 154, 220), pixel P2 corresponds to pixel values ​​(152, 164, 150), ..., pixel P... n The corresponding pixel values ​​are (120, 154, 220); for the corresponding red color channel, pixel P1 has a value of 120, pixel P2 has a value of 152, ..., pixel P... n The corresponding value is 120; for the corresponding green color channel, the value of pixel P1 is 154, the value of pixel P2 is 164, ..., the value of pixel P... n The value is 154; for the corresponding blue color channel, the value of pixel P1 is 220, the value of pixel P2 is 150, ..., the value of pixel P... n The value is 220.

[0051] The mean intensity value of the corresponding red color channel is:

[0052]

[0053] The mean intensity value of the corresponding green color channel is:

[0054]

[0055] The mean intensity value of the corresponding blue color channel is:

[0056]

[0057] Then, the average value determined by the three color channels is used as the new RGB value, and this new RGB value is used as the offset value for each pixel within the current cell. Referring to the example above, the new RGB value is:

[0058]

[0059] Optionally, during encryption, the pixel values ​​of each pixel in the cell are changed according to the offset value of the cell, that is, after encryption, the values ​​of all pixels in the cell are the new RGB values ​​mentioned above.

[0060] Given the offset values ​​of all pixels within the target area, these offset values ​​are used as the first data within that area. This first data records the offset values ​​of each cell according to a preset data arrangement order. For example, the offset values ​​are arranged from top to bottom and left to right, meaning cells in the previous row precede cells in the next row, and within the same row, cells are arranged from left to right. By arranging the offset values ​​of each cell in this corresponding order, the data can be decrypted in this order, allowing for rapid extraction of the original pixel values.

[0061] Furthermore, the target area is encrypted according to the first data. This involves changing the pixel values ​​of all pixels within the corresponding cells of the target area according to the offset values ​​recorded in the first data, thus obtaining an encrypted video image of the target area. Specifically, the offset value of each cell determines the value of all pixels within that cell. For example, using the grayscale scheme described above, if the offset value of a pixel in a cell is 130, and the original grayscale values ​​of two pixels in that cell are 125 and 165 respectively, after encryption, the grayscale values ​​of both pixels will change to 130. Based on this, after updating the pixel values ​​of all cells within the target area, an encrypted video image of the target area is obtained. Therefore, by encrypting the target area, privacy protection processing of the target object can be achieved in the video image, meeting the user's privacy protection requirements.

[0062] In one embodiment, the encryption of the image data of the target area is still performed using Gaussian blur. For example, based on the Gaussian blur algorithm, all pixels within the target area are traversed, and the pixel value of each pixel is redefined as its offset value. Then, the offset value corresponding to each pixel is determined as the first data. Specifically, each pixel within the target area is processed sequentially, such as by performing a weighted average calculation based on the pixel values ​​of the pixels surrounding that pixel. For example, the pixel can be used as the center point, and a region with a radius of 10 pixels can be selected to perform a weighted average calculation on the pixels within that region, thereby determining the new pixel value of that pixel as the offset value. Optionally, when setting the weights, the weights can be set according to the distance between other pixels and the pixel used as the center point; for example, the closer the distance, the greater the weight, and the farther the distance, the smaller the weight. It is worth noting that the calculation for each pixel is based on the original pixel values ​​of other pixels; that is, during the Gaussian blur processing, the pixel values ​​of the pixels within the target area in the original video image are used for calculation.

[0063] It should be noted that in some embodiments, other encryption methods can be used for image data, such as randomly swapping the pixel values ​​corresponding to pixels in each cell, that is, the pixel values ​​after the pixel swap are the pixel values ​​of other pixels in the same cell.

[0064] By using Gaussian blurring, the pixels in the target area of ​​the encrypted video image are not the original pixel values, thus encrypting the original content of the area and effectively occluding the target object to protect its privacy.

[0065] Optionally, when multiple targets exist in the same video image and the distance between them is small, the target area corresponding to one target will overlap with the target areas of the other targets. Therefore, during the image data encryption process, the overlapping portion between the current target area and other target areas is determined and set as a single cell. Thus, for two target areas with an overlapping portion, the overlapping portion serves as a single cell dividing the two target areas. Consequently, when encrypting the pixel values ​​of pixels within the target areas, the encrypted values ​​of all pixels in the overlapping portion are identical in both target areas associated with the overlapping portion. This ensures that the overlapping portion is not encrypted multiple times during the encryption process for different target areas, thereby ensuring that the decrypted overlapping portion remains consistent with the original video image.

[0066] Optionally, during the encryption process of image data, if the overlapping portion of the current target region with other target regions is determined, the overlapping portion can be set as a cell. For two target regions with overlapping portions, the first percentage of the cell corresponding to the overlapping portion within each target region is determined. This first percentage represents the proportion of the overlapping portion on the target object within the target region. For example, if the target object is a face, for one target region, the proportion of the overlapping portion on the face within that region can be determined as the ratio of the number of pixels corresponding to the face in the overlapping portion to the total number of pixels in the cell of the overlapping portion. The other target region determines its corresponding first percentage in the same way. Then, by comparing the first percentages of different target regions, the target region with the larger first percentage is designated as the region to which the cell corresponding to the overlapping portion belongs. That is, during encryption, the overlapping portion is assigned to the target region with the larger first percentage, while other target regions do not process the overlapping portion to avoid duplicate encryption.

[0067] In one embodiment, the preset encryption algorithm is an asymmetric encryption algorithm, which has two different keys to form a key pair, including a public key and a private key. The method uses different encryption and decryption keys, such as using the public key as the first key for encrypting data and the private key as the second key for decrypting data. That is, the asymmetric encryption algorithm generates a key pair including a first key for encrypting data and a second key for decrypting data. First data and image data corresponding to each frame of video image are obtained, wherein the image data includes the pixel values ​​of all pixels within the target area, and the first data includes the encrypted values ​​of all pixels within the target area. Optionally, the asymmetric encryption algorithm includes algorithms such as RSA, DSA (Digital Signature Algorithm), and ECC (Elliptic Curve Cryptography).

[0068] Furthermore, a first key for encrypting data is determined in the key pair generated by the asymmetric encryption algorithm. Optionally, the first key is stored in the video capture device so that the encrypted data is uploaded to the second server after being encrypted using the first key. For pixels within the target area in the video image, the target pixel value for each pixel within the target area is determined based on the pixel value in the image data and the encrypted value in the first data. Optionally, the pixel value in the image data is subtracted from the offset value of the cell corresponding to the pixel in the first data to obtain the difference value for the corresponding pixel. That is, the difference value for the pixel is obtained by subtracting the encrypted value from the original pixel value. This difference value is then used as the target pixel value for all pixels within the target area. Optionally, the pixel value in the image data can also be added to the offset value of the cell corresponding to the pixel in the first data to obtain the target pixel value for the corresponding pixel.

[0069] Furthermore, the target pixel values ​​corresponding to all pixels within the target area are encrypted using the first key in the asymmetric encryption algorithm, and the encrypted data is used as the second data. It is conceivable that corresponding second data can be determined for each frame of video image. After encrypting the target pixel values ​​corresponding to all pixels within the target area of ​​the current video image using the first key, the obtained second data is stored through a second server, thus storing all the second data obtained from encrypting all video images in the second server.

[0070] For example, a video dataset contains N video frames, each containing a target object, and each video frame has a target region. After determining the image data and first data for the current frame, where the first data includes the offset values ​​corresponding to each cell within the target region of the video image, the original pixel value of each pixel within the target region is subtracted from the offset value of its corresponding cell in the first data to obtain the difference value of that pixel (which serves as the corresponding target pixel value). This yields the second data for the video frame, which is then uploaded to a second server. Therefore, the second server stores N sets of second data for this video data.

[0071] To address this, this solution encrypts the determined target pixel value using an asymmetric encryption algorithm. Data can be encrypted using the public key in asymmetric encryption, and can only be decrypted using the corresponding private key, thereby improving data security.

[0072] Optionally, when multiple target regions exist in the same video image, encrypting different target regions can yield multiple second data corresponding to different target regions. Then, according to a preset arrangement order of the multiple target regions, the second data corresponding to each target region is stored sequentially in the second server. It is conceivable that the second data belonging to different frames of video images can be identified by corresponding tags in the second server, while multiple second data belonging to the same frame of video images are stored sequentially according to a preset arrangement order. The preset arrangement order is a pre-defined data layout order, such as arranging the second data corresponding to each target region from top to bottom and from left to right. Taking the pixel at the center point of the video image as the origin of the coordinate system, taking two target regions in the same video image as an example, if the coordinate value (corresponding to the Y-axis coordinate value) of one pixel in one target region S1 is greater than the coordinate values ​​(corresponding to the Y-axis coordinate values) of all pixels in another target region S2, then target region S1 is determined to be above target region S2, and the second data corresponding to target region S1 is stored sequentially before the second data corresponding to target region S2. If the largest coordinate value among all pixels in the two target regions along the Y-axis is the same, and there exists a pixel coordinate value (corresponding to the X-axis coordinate value) in target region S1 and the coordinate values ​​of all pixels in target region S2 (corresponding to the X-axis coordinate values) along the X-axis, then it can be determined that target region S1 is to the left of target region S2, and the second data corresponding to target region S1 is stored before the second data corresponding to target region S2 in order.

[0073] Optionally, when multiple target regions exist in the same video image, second data corresponding to different target regions can be distinguished by coordinate indexes. For example, a coordinate index can be added to the second data, which includes the coordinate values ​​of the current target region in the video image. For instance, in one embodiment, the added coordinate index is the coordinate value of the top-left vertex of the target region. It is understood that the coordinate index is used to determine the target region associated with the second data in the video image when decrypting the resynthesized second video data. After obtaining and decrypting the second data, the target pixel values ​​of each pixel within the corresponding target region in the second data can be determined. Furthermore, the position of each target region in the video image can be determined according to the coordinate index to distinguish different target regions. Then, when restoring the pixels within the target region in the second video data, the second data corresponding to the target region can be determined according to the coordinate index to determine the corresponding target pixel values ​​and restore the original pixel values.

[0074] In some embodiments, the second key in the first key pair generated by the asymmetric encryption algorithm is aligned with the encryption time. Optionally, the first key pair generated by the asymmetric encryption algorithm is different each time it is run, wherein the key generated by the asymmetric encryption algorithm is determined based on two randomly selected prime numbers, and therefore the key pair used is different when encrypting different video images in the same video data. Furthermore, a timestamp corresponding to the time when the image data in each frame of video image is encrypted according to the first key is determined, that is, the encryption time of the image data is identified by the timestamp. For example, the timestamp is added to the encrypted data and stored in the first server, so as to identify the encryption time of the video image using the current first key by the timestamp.

[0075] Furthermore, a timestamp is associated with the second key corresponding to the first key and stored on the second server; that is, the second key and its associated timestamp are stored on the second server. The first and second keys in each key pair are corresponding. After determining the timestamp corresponding to the encryption time of the first key, the second key is associated with the timestamp, which can be used to determine the second key used when decrypting each frame of video image. It can be understood that the timestamp is used to align the second key with the image data encrypted by the first key in the same key pair when decrypting the second video data. That is, the encrypted image data corresponds to a timestamp, and the second key also has a timestamp. By comparing the two timestamps, the second key corresponding to the same timestamp is aligned with the encrypted image data so that the image data can be decrypted using the second key during the decryption process. Therefore, by setting timestamps to align the second key used for decrypting data with the image data in the second video data to be decrypted, it helps to better decrypt the data and recover the original video image.

[0076] For example, we will illustrate this using an application in a video surveillance scenario. Figure 3 The flowchart for processing video images provided in one embodiment of this application is shown in the figure. After the video acquisition device acquires first video data, it parses the first video data to obtain multiple frames of video images. In the figure, a human face is used as the target object. Face recognition is performed on each frame of video images to determine the image region where the target object is located, i.e., the target region, thereby determining the image data corresponding to the target region in each frame of video images. As shown in the figure, in a scenario of monitoring a warehouse, the acquired first video data includes multiple frames of video images. Video image I in the figure... s It is an unencrypted image containing a human face.

[0077] The image data is then encrypted. For example, the target area is divided into multiple cells, and the average pixel value of all pixels within each cell is calculated. If the pixel value is an RGB value, the average value is calculated for each color channel to determine the average pixel value for that cell, which is then used as the corresponding offset value. After determining the average pixel value for all cells within the target area, this is used as the first encrypted data. Then, based on the average pixel value for each cell in the first data, the pixel values ​​of all pixels within each cell are changed, resulting in the encrypted video image of the target area, as shown in Image I in the figure. g In image I g The target region in image I is represented by a dashed box. g The location within the target area is determined. The encrypted video images of the target area are re-synthesized and stored as a second video data on the first server.

[0078] Furthermore, the first data records the offset values ​​corresponding to each cell within the target area in the video image, which represent the encrypted pixel values ​​in the corresponding cells within the target area in the second video image; while the image data records the original pixel values ​​of the pixels in each cell within the target area in the video image. Based on this, the target pixel value corresponding to the same pixel point between the image data and the first data is calculated as the corresponding difference data. For example, when the difference is used as the target pixel value, the pixel value of the pixel point is an RGB value. The original pixel value recorded for a pixel point in the image data is (120, 154, 220), while the offset value corresponding to the cell where the pixel point is located in the first data is (152, 164, 150), and the corresponding difference value is (-32, -10, -70). After determining the difference values ​​corresponding to all pixels in the target area of ​​the video image frame, the difference data corresponding to the video image can be determined, and then the difference data is encrypted according to the encryption algorithm to obtain the second data. It is conceivable that the second data could be obtained by encrypting the difference data corresponding to a single video frame, or by encrypting the difference data of all video frames uniformly. Furthermore, the encrypted second data is stored on a second server. In the diagram, the second data is designated as D1, ..., D... n This represents the target pixel value calculated for each frame of the video image. One cell represents a pixel within the target area of ​​the corresponding video image and its corresponding target pixel value.

[0079] Device E1, acting as a video playback device for a user account with primary access, can only retrieve and play secondary video data from the primary server. However, the faces of people in the displayed video are encrypted; that is, the displayed video image is as shown in image I. gAs shown in Figure I, device E2, acting as a video playback device logged into by a user account with second-level privileges, can obtain second video data from the first server and second data from the second server. Furthermore, device E2 can decrypt the second data using the obtained key and then use the second data to decrypt the second video data, thereby playing the decrypted video. The faces of the people displayed in the video are the decrypted images, as shown in Figure I. s As shown, this displays the original video feed to the user. Based on this, during routine warehouse monitoring patrols, logging in with a low-priority, first-level account to access the video feed played on device E1 as described above is sufficient to meet the patrol requirements. However, for identifying targets such as employees entering the warehouse or those organizing shelves, logging in with a high-priority, second-level account to access the video feed played on device E2 as described above is necessary to obtain the original video feed for facial recognition. It should be noted that decrypting the video image is the reverse of the encryption process described above; after determining the encryption method, the process is reversed to complete the decryption of the video image.

[0080] In response, the video data processed by the video image processing method of this solution allows devices that only obtain the encrypted video data to play the video after the target object is encrypted, while devices that obtain both the encrypted video data and the second data can play the original video. This not only satisfies the user's privacy requirements for the target object, but also restores the image to facilitate the identification of the target object.

[0081] Figure 4 This is a schematic diagram of a video image processing apparatus according to an embodiment of this application. The apparatus is used to execute the video image processing method provided in the above embodiment and has corresponding functional modules and beneficial effects for executing the method. The video image processing apparatus is applied to a video acquisition device in a video acquisition system. The video acquisition system includes a video acquisition device, a first server, and a second server. The video acquisition device is communicatively connected to both the first server and the second server. As shown in the figure, the apparatus includes: a data extraction module 301, a first encryption module 302, an image synthesis module 303, and a second encryption module 304.

[0082] The data extraction module 301 is configured to extract video images from the acquired first video data and determine the image data corresponding to the target area in the video image, wherein the target area is the image area set in the video image corresponding to the target object;

[0083] The first encryption module 302 is configured to encrypt image data, determine the encrypted first data, and obtain the encrypted video image of the target area.

[0084] The image synthesis module 303 is configured to resynthesize second video data based on all encrypted video images of the target area and store the second video data on the first server. The second video data is used for playback according to playback requests.

[0085] The second encryption module 304 is configured to encrypt the target pixel values ​​associated with the image data and the first data based on a preset encryption algorithm, thereby obtaining encrypted second data. The second data is stored on the second server and is used to decrypt the second video data to obtain the first video data according to the processing request.

[0086] Based on the above embodiments, the image data includes the pixel values ​​of all pixels within the target area, and the first encryption module 302 is specifically configured as follows:

[0087] Based on the pixel values ​​of all pixels in the target area, determine the offset value of each pixel after offset operation, and use the offset values ​​of all pixels as the first data. The offset operation is used to recalculate the pixel value of each pixel.

[0088] By changing the pixel values ​​of all pixels within the target area according to the first data, the encrypted target area is obtained, and a video image of the encrypted target area is obtained.

[0089] Based on the above embodiments, the pixel value is an RGB value, where the RGB value is the intensity value of the three preset color channels corresponding to the pixel point. The first encryption module 302 is further configured as follows:

[0090] Determine the multiple cells within the target area, and determine the average intensity value of each color channel for all pixels within each cell;

[0091] The mean value of the three color channels is determined as the new RGB value, and the new RGB value is used as the offset value of each pixel in the current cell.

[0092] Given the offset values ​​of all pixels within the target area, the offset values ​​of all pixels are used as the first data within the target area, and the first data is recorded according to the preset arrangement order of the cells, which corresponds to the offset values ​​of the cells.

[0093] Based on the above embodiments, the first encryption module 302 is further configured as follows:

[0094] Based on the Gaussian blur algorithm, all pixels in the target area are traversed and the pixel value of each pixel is redefined as the offset value corresponding to each pixel.

[0095] The offset value corresponding to each pixel is determined as the first data.

[0096] Based on the above embodiments, the image data includes the pixel values ​​of all pixels within the target area, the first data includes the encrypted values ​​of all pixels within the target area, and the second encryption module 304 is specifically configured as follows:

[0097] The asymmetric encryption algorithm is used as the preset encryption algorithm, and the first key in the key pair generated by the asymmetric encryption algorithm is determined to be used to encrypt the data.

[0098] Based on the pixel values ​​in the image data and the encrypted values ​​in the first data, the target pixel values ​​of each corresponding pixel point in the target area are determined.

[0099] The target pixel values ​​corresponding to all pixels in the target area are encrypted using the first key in the asymmetric encryption algorithm, and the encrypted data is used as the second data.

[0100] The second data is stored on a second server.

[0101] Based on the above embodiments, the target pixel value is obtained by performing matrix operations on a first matrix constructed from the image data and a second matrix constructed from the first data.

[0102] Based on the above embodiments, when multiple target regions exist in the video image, the second encryption module 304 is further configured as follows:

[0103] According to the preset arrangement order corresponding to multiple target areas, the second data corresponding to each target area is stored in the second server in sequence.

[0104] Alternatively, a coordinate index can be added to the second data. The coordinate index includes the coordinates of the current target region in the video image. The coordinate index is used to determine the target region associated with the second data in the video image when decrypting the resynthesized second video data.

[0105] Based on the above embodiments, the asymmetric encryption algorithm generates a key pair including a first key and a second key for decrypting data, and the second encryption module 304 is further configured as follows:

[0106] Determine the timestamp corresponding to the moment when the image data in each frame of video is encrypted according to the first key;

[0107] A timestamp is associated with the second key corresponding to the first key and stored in the second server. The timestamp is used to align the second key with the image data encrypted by the first key in the same key pair when decrypting the second video data.

[0108] It is worth noting that in the embodiments of the above-mentioned device, the modules included are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be achieved; in addition, the specific names of each module are only for easy differentiation and are not used to limit the protection scope of the embodiments of this application.

[0109] Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. The device is used to execute the video image processing method provided in the above embodiment and has corresponding functional modules and beneficial effects for executing the method. As shown in the figure, the device includes a processor 401, a memory 402, an input device 403, and an output device 404. The number of processors 401 can be one or more; one processor 401 is shown as an example in the figure. The processor 401, memory 402, input device 403, and output device 404 can be connected via a bus or other means; a bus connection is shown as an example in the figure. The memory 402, as a computer-readable storage medium, can be used to store software programs, computer-executable programs, and modules, such as the program instructions / modules corresponding to the video image processing method in the embodiments of this application. The processor 401 executes various corresponding functional applications and data processing by running the software programs, instructions, and modules stored in the memory 402, thereby implementing the above-mentioned video image processing method.

[0110] The memory 402 may primarily include a program storage area and a data storage area. The program storage area may store the operating system and at least one application program required for a given function; the data storage area may store data recorded or created during use. Furthermore, the memory 402 may include high-speed random access memory and non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state storage device. In some embodiments, the memory 402 may further include memory remotely located relative to the processor 401, which can be connected to the device via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0111] The input device 403 can be used to input corresponding digital or character information to the processor 401, and to generate key signal inputs related to the user settings and function control of the device; the output device 404 can be used to send or display key signal outputs related to the user settings and function control of the device.

[0112] This application also provides a storage medium storing computer-executable instructions, which, when executed by a processor, are used to perform related operations in the video image processing method provided in any embodiment of this application.

[0113] Computer-readable storage media include both permanent and non-permanent, removable and non-removable media, and information storage can be achieved by any method or technology. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, disk storage or other magnetic storage devices, or any other non-transfer medium that can be used to store information accessible by a computing device.

[0114] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0115] Note that the above description is merely a preferred embodiment and the technical principles employed in this application. Those skilled in the art will understand that this application is not limited to the specific embodiments described herein, and various obvious changes, readjustments, and substitutions can be made without departing from the scope of protection of this application. Therefore, although this application has been described in detail through the above embodiments, this application is not limited to the above embodiments. Many other equivalent embodiments may be included without departing from the concept of this application, and the scope of this application is determined by the scope of the appended claims.

Claims

1. A method of processing a video image, characterized by, A video acquisition device applied to a video acquisition system, the video acquisition system comprising the video acquisition device, a first server and a second server, the video acquisition device being in communication connection with the first server and the second server respectively, the method comprising: extracting a video image from the obtained first video data and determining image data corresponding to a target region in the video image, the target region being an image region set for a target object in the video image; encrypting the image data to determine encrypted first data and obtain a video image of the target region after encryption; based on all the video images of the target region after encryption, re-composing second video data and storing the second video data in the first server, the second video data being used for playing according to a playing request; based on a preset encryption algorithm, encrypting target pixel values associated with the image data and the first data to obtain encrypted second data, and storing the second data in the second server, the second data being used for decrypting the second video data to obtain the first video data according to a processing request; wherein the image data comprises pixel values of all pixel points in the target region, and the encryption processing of the image data to determine the encrypted first data and obtain the video image of the target region after encryption comprises: based on the pixel values of all pixel points in the target region, determining offset values of each pixel after offset operation processing, and taking the offset values of all pixel points as the first data, the offset operation processing being used for re-computing pixel values of each pixel point; changing the pixel values of all pixel points in the target region according to the first data to obtain an encrypted target region, so as to obtain the video image of the target region after encryption.

2. The video image processing method of claim 1, wherein, The pixel value is an RGB value, and the RGB value is an intensity value of a pixel point corresponding to a preset three color channels; the determination of the offset values of each pixel after the offset operation processing based on the pixel values of all pixel points in the target region to take the offset values of all pixel points as the first data comprises: determining a plurality of cells divided in the target region, and respectively determining mean values of intensity values of each color channel corresponding to all pixel points in a cell; determining the determined mean values corresponding to the three color channels as new RGB values, and taking the new RGB values as the offset values corresponding to each pixel point in the current cell; in the case of determining the offset values of all pixel points in the target region, taking the offset values of all pixel points as the first data in the target region, and the first data recording the offset values corresponding to the cells according to a preset arrangement order of the cells.

3. The video image processing method of claim 1, wherein, the determination of the offset values of each pixel after the offset operation processing based on the pixel values of all pixel points in the target region to take the offset values of all pixel points as the first data comprises: based on a Gaussian blur algorithm, traversing all pixel points in the target region and re-determining pixel values of each pixel point as the offset values corresponding to each pixel point; The offset value corresponding to each pixel point is determined as the first data.

4. The method of any of claims 1-3, wherein, The image data includes pixel values of all pixel points in the target region, and the first data includes values after encryption of all pixel points in the target region. The target pixel value associated with the image data and the first data is encrypted based on a preset encryption algorithm to obtain encrypted second data, and the second data is stored in a second server, including: An asymmetric encryption algorithm is used as the preset encryption algorithm, and a first key for encrypting data in a key pair generated by the asymmetric encryption algorithm is determined. Based on the pixel values in the image data and the encrypted values in the first data, the target pixel values corresponding to each pixel point in the target region are determined. The target pixel values corresponding to all pixel points in the target region are encrypted according to the first key in the asymmetric encryption algorithm, and the encrypted data is used as the second data. The second data is stored in the second server.

5. The video image processing method of claim 4, wherein, The target pixel value is obtained by performing matrix operation processing on a first matrix constructed based on the image data and a second matrix constructed based on the first data.

6. The video image processing method of claim 4, wherein, In the case where multiple target regions exist in the video image, the method further includes: According to a preset arrangement order corresponding to the multiple target regions, the second data corresponding to each target region is sequentially stored in the second server. Or, a coordinate index is added to the second data, the coordinate index including a coordinate value of the current target region in the video image, and the coordinate index is used to determine the target region associated with the second data in the video image when decrypting the second video data obtained by re-composition.

7. The video image processing method of claim 4, wherein, The asymmetric encryption algorithm generates a key pair including the first key and a second key for decrypting data, and the method further includes: A timestamp corresponding to a time when the image data in each frame of video image is encrypted according to the first key is determined; The second key corresponding to the first key is associated with the timestamp, and is stored in the second server, and the timestamp is used to align the second key with the image data encrypted by the first key in the same key pair when decrypting the second video data.

8. A video image processing apparatus, characterized by comprising: A video acquisition device applied to a video acquisition system, the video acquisition system including the video acquisition device, a first server and a second server, the video acquisition device being in communication connection with the first server and the second server respectively, and the video image processing apparatus including: A data extraction module configured to extract a video image from the obtained first video data and determine image data corresponding to a target region in the video image, the target region being an image region set corresponding to a target object in the video image; A first encryption module configured to encrypt the image data, determine encrypted first data, and obtain a video image after encryption of the target region; An image synthesis module is configured to synthesize second video data based on all the video images of the target region after encryption and store the second video data in the first server, the second video data being used for playing according to a playing request; A second encryption module is configured to encrypt target pixel values associated with the image data and the first data based on a preset encryption algorithm to obtain second data after encryption, and store the second data in the second server, the second data being used for decrypting the second video data to obtain the first video data according to a processing request; The image data includes pixel values of all pixel points in the target region, and the first encryption module is specifically configured to: determine offset values of all pixel points in the target region after offset operation processing based on the pixel values of all pixel points in the target region, and use the offset values of all pixel points as the first data, the offset operation processing being used for recalculating pixel values of all pixel points; change the pixel values of all pixel points in the target region according to the first data to obtain the target region after encryption, so as to obtain the video image of the target region after encryption.

9. An electronic device, comprising: comprise: one or more processors; a storage device configured to store one or more programs, when the one or more programs are executed by the one or more processors, the one or more processors are caused to implement the video image processing method according to any one of claims 1-7.

10. A storage medium storing computer-executable instructions, wherein: The computer executable instructions, when executed by the processor, are used to perform the video image processing method according to any one of claims 1-7.

Citation Information

Patent Citations

  • Privacy-protected outsourcing image feature extraction and classification method

    CN115797653A

  • Data transmission method, device and system, computer equipment and storage medium

    CN116962589A