Encrypted image information acquisition apparatus and method
By using a mask for optical encryption in the camera lens and combining convolutional networks and neural networks to process image information, the problems of complex structure and software encryption risks in traditional cameras are solved, achieving miniaturization and efficient encryption.
Patent Information
- Application Number
- CN202210072981.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-21
- Publication Date
- 2025-11-11
- Estimated Expiration
- 2042-01-21
AI Technical Summary
Existing technologies struggle to achieve optical encryption during the image acquisition stage, and traditional cameras have complex lens structures that cannot be miniaturized. Furthermore, software encryption methods carry the risk of decryption.
A mask is used to replace the lens group in the camera lens. The image is encrypted through optical encryption, and the image information is processed and restored by combining convolutional networks and back-end neural networks.
This enables the miniaturization of the optical system, protects user privacy, reduces computational load, improves encryption speed, and makes real-time identification possible.
Smart Images

Figure CN114491592B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of optical instruments, and more specifically, to an encrypted image information acquisition device and method. Background Technology
[0002] With the improvement of camera pixels and image quality, as well as image analysis capabilities, sensitive biometric information can be increasingly extracted from images. For example, it is easy to capture a user's facial image from a photo, and there is even a high probability that a user's fingerprints, palm prints, iris scans, and other private information used for identity verification can be captured in the photo, posing a threat to the user's privacy.
[0003] Currently, image encryption mainly focuses on two levels: software and hardware. At the software level, after obtaining a high-fidelity visual image, pixelation blurring or replacement of sensitive areas is usually performed. For example, to protect facial privacy, facial regions in the image are replaced or altered to prevent facial recognition methods from effectively recognizing the modified face image. However, such methods require verification of facial information and also have a high risk of being decrypted. At the hardware level, cameras with complex lens groups are usually used in the image acquisition stage to obtain secure images without sensitive data. However, due to the complexity of the lens group arrangement and the limitation of the focusing distance, cameras with complex lens groups cannot achieve miniaturization of the optical system structure, limiting the application scenarios.
[0004] Therefore, there is a need for an encrypted image information acquisition device that can optically encrypt images during the image acquisition stage, and the optical device used for encryption has a simple structure. Summary of the Invention
[0005] To address the above problems, this disclosure utilizes a mask to replace the lens group in a camera lens and optically encrypts the acquired image, thereby completing the optical encryption of the image, and the optical device used for encryption has a simple structure.
[0006] The embodiments of this disclosure provide an encrypted image information acquisition device and method.
[0007] This disclosure provides an encrypted image information acquisition device, comprising: at least one mask layer and an optical sensing component, wherein the at least one mask layer is used to receive object light from a target object and generate an encrypted optical image signal based on the received object light, wherein the object light carries image information of the target object, and the encrypted optical image signal carries encrypted image information, wherein each mask layer includes a mask pattern; the optical sensing component is used to receive the encrypted optical image signal, convert the encrypted optical image signal into an electrical signal, and output the electrical signal, wherein the electrical signal includes the encrypted image information of the target object.
[0008] According to an embodiment of this disclosure, receiving object light from a target object and generating an encrypted optical image signal based on the received object light includes: performing convolution processing on the image information based on the received object light through the mask pattern, wherein the convolution network performing the convolution processing includes at least one convolutional layer.
[0009] According to embodiments of this disclosure, the at least one mask layer corresponds one-to-one with the at least one convolutional layer, wherein the mask pattern on each mask layer carries parameter information for performing convolution processing on the corresponding convolutional layer, wherein the parameter information includes at least the parameters of the convolution kernel.
[0010] According to embodiments of this disclosure, the mask pattern of each mask is composed of multiple mask holes, and the light transmittance of each mask hole is the same or different, wherein the light transmittance of each mask hole is determined by the parameters of its corresponding convolution kernel.
[0011] According to embodiments of this disclosure, the mask pattern is determined as follows: a convolutional model of the mask pattern is established, and based on the convolutional model of the mask pattern, training images used for training are encrypted to obtain encrypted training images; features are extracted from the encrypted training images using a back-end neural network to obtain feature-extracted images; a loss function for signal processing of the back-end neural network is determined based on the visual task to be completed; and the convolutional model of the mask pattern is trained based on the loss function to obtain parameter information of the trained convolutional model.
[0012] According to embodiments of this disclosure, the target object is the object to be photographed.
[0013] According to embodiments of this disclosure, the target object is an optical image signal carrying image information of the target object, and the encrypted image information acquisition device further includes a light generating component for generating an optical image signal carrying image information, wherein the optical image signal carrying image information is an incoherent light signal; wherein the light generating component is a plurality of point light sources, and the optical image signal carrying image information is a light signal generated by the plurality of point light sources; or the light generating component is a display, and the optical image signal carrying image information is a multi-pixel image generated by the display.
[0014] According to embodiments of this disclosure, the at least one mask layer is a single mask layer, and the convolutional network includes a convolutional layer, wherein the distance d between two adjacent point light sources or two adjacent pixels is... L The distance d between the light generating component and the photomask LM The distance d between the photomask and the optical sensing component MS The dimensions Δ of individual pixels on the optical sensing component have the following relationship: In this context, a single pixel on the optical sensing component is equivalent to a single pixel generated after the convolution calculation of the convolutional layer.
[0015] According to embodiments of this disclosure, the mask pattern is obtained by coating or etching on each layer of the mask.
[0016] According to an embodiment of this disclosure, the encrypted image information acquisition device further includes: an information extraction component, configured to receive the electrical signal output by the optical sensing component, and to extract features from the encrypted image information based on the electrical signal.
[0017] According to an embodiment of this disclosure, the information extraction component includes a deconvolutional network and a feature extraction network, wherein the deconvolutional network is used to deconvolve the encrypted image information to obtain recovered image information, and the feature extraction network is used to extract image features from the recovered image information.
[0018] According to embodiments of this disclosure, the deconvolutional network performs deconvolution on the encrypted image information using the following equation: X = F -1 (F(W)⊙F(Y)); where X is the recovered image information, W is the regularized pseudo-inverse of the point spread function, and Y is the encrypted image information, and
[0019] Embodiments of this disclosure provide a method for acquiring encrypted image information, comprising: receiving object light of a target object through at least one layer of a mask, and generating an encrypted optical image signal based on the received object light, wherein the object light carries image information of the target object, and the encrypted optical image signal carries encrypted image information, wherein each layer of the mask includes a mask pattern; converting the encrypted optical image signal into an electrical signal, and outputting the electrical signal, wherein the electrical signal includes the encrypted image information of the target object.
[0020] According to embodiments of this disclosure, the mask pattern is determined as follows: a convolutional model of the mask pattern is established, and based on the convolutional model of the mask pattern, training images used for training are encrypted to obtain encrypted training images; features are extracted from the encrypted training images using a back-end neural network to obtain feature-extracted images; a loss function for signal processing of the mask pattern and the back-end neural network is determined based on the visual task to be completed; and the convolutional model of the mask pattern is trained based on the loss function to obtain parameter information of the trained convolutional model.
[0021] The encrypted image information acquisition device disclosed herein replaces the lens group in the camera lens with a mask and performs optical encryption on the acquired image, protecting user privacy. It also omits modules such as de-mosaicing and noise reduction designed for perception enhancement in the image signal processing of traditional cameras, simplifying the hardware structure and realizing the miniaturization of the optical system structure. Compared with other digital encryption methods, it saves computation and improves encryption speed, making real-time recognition possible. Attached Figure Description
[0022] To more clearly illustrate the technical solutions of the embodiments of this disclosure, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are merely some exemplary embodiments of this disclosure. Those skilled in the art can obtain other drawings based on these drawings without any creative effort. The following drawings are not intentionally drawn to scale with actual dimensions; their focus is on illustrating the main points of the invention.
[0023] Figure 1 A schematic diagram of the structure of the encrypted image information acquisition device 100 according to an embodiment of the present disclosure is shown;
[0024] Figure 2 Another schematic diagram of the structure of the encrypted image information acquisition device 100 according to an embodiment of the present disclosure is shown;
[0025] Figure 3Another schematic diagram of the structure of the encrypted image information acquisition device 100 according to an embodiment of the present disclosure is shown;
[0026] Figure 4 A schematic diagram showing the relationship between the distances between the target object, the mask 1001, and the optical sensing component 1002 according to an embodiment of the present disclosure is shown.
[0027] Figure 5 A schematic diagram of a mask 1001 pattern according to an embodiment of the present disclosure is shown;
[0028] Figure 6 A flowchart illustrating a method for determining a mask pattern 600 according to an embodiment of the present disclosure is shown;
[0029] Figure 7 A schematic diagram of a U-Net network structure according to an embodiment of the present disclosure is shown;
[0030] Figure 8 A flowchart of an encrypted image information acquisition method 800 according to an embodiment of the present disclosure is shown. Detailed Implementation
[0031] To make the objectives, technical solutions, and advantages of this disclosure more apparent, exemplary embodiments according to this disclosure will now be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of this disclosure, and not all embodiments of this disclosure. It should be understood that this disclosure is not limited to the exemplary embodiments described herein.
[0032] As indicated in this application and claims, unless the context clearly indicates otherwise, the words "a," "an," "an," and / or "the" are not specifically singular and may include plural forms. Generally speaking, the terms "comprising" and "including" only indicate the inclusion of explicitly identified steps and elements, which do not constitute an exclusive list, and the method or apparatus may also include other steps or elements.
[0033] In this specification and accompanying drawings, steps and elements that are substantially the same or similar are indicated by the same or similar reference numerals, and repeated descriptions of these steps and elements are omitted. Furthermore, in the description of this disclosure, the terms "first," "second," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance or order.
[0034] Furthermore, the terms "upper," "lower," "vertical," and "horizontal," which relate to orientation or positional relationships, are used in this specification and accompanying drawings only for the convenience of describing embodiments according to this disclosure and are not intended to limit this disclosure. Therefore, they should not be construed as limiting this disclosure.
[0035] Furthermore, in this specification and accompanying drawings, if flowcharts are used to illustrate the steps of a method according to embodiments of this disclosure, it should be understood that the preceding or following steps are not necessarily performed in precise order. Instead, various steps may be processed in reverse order or simultaneously, unless expressly defined in the embodiments of this disclosure. Additionally, other operations may be added to these processes, or one or more steps may be removed from them.
[0036] Furthermore, in this specification and accompanying drawings, unless otherwise expressly stated, terms such as "connection" or "linked" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect.
[0037] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs. The terminology used herein is for the purpose of describing embodiments of the invention only and is not intended to limit the invention.
[0038] To facilitate the description of this disclosure, the following concepts related to this disclosure are introduced.
[0039] Optical encryption is a typical and efficient image encryption method. Its essence is to scramble and encode the inherent information of an image through optical transformation processes, such as interference, diffraction, and imaging, thereby achieving good encryption results.
[0040] Artificial intelligence (AI) is the theory, methods, technology, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that attempts to understand the essence of intelligence and produce a new kind of intelligent machine that can react in a way similar to human intelligence. AI studies the design principles and implementation methods of various intelligent machines, enabling them to possess the functions of perception, reasoning, and decision-making.
[0041] Artificial intelligence (AI) is a comprehensive discipline encompassing a wide range of fields, including both hardware and software technologies. Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies primarily include computer vision, speech processing, natural language processing, and machine learning / deep learning.
[0042] Computer vision is a science that studies how to enable machines to "see." More specifically, it refers to machine vision, which uses cameras and computers to replace human eyes in tasks such as target recognition, tracking, and measurement, and further performs image processing to create images more suitable for human observation or transmission to instruments for detection. As a scientific discipline, computer vision studies related theories and technologies, attempting to build artificial intelligence systems capable of extracting information from images or multidimensional data. Computer vision technologies typically include image processing, image recognition, image semantic understanding, image retrieval, OCR, video processing, video semantic understanding, video content / behavior recognition, 3D object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping (SLAM), and common biometric recognition technologies such as facial recognition and fingerprint recognition. In the embodiments provided in this application, by using computer vision technology, more information from images (e.g., medical images) can be acquired and provided to the user through computer processing.
[0043] Machine learning is a multidisciplinary field involving probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory, among others. It specifically studies how computers can simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to endow computers with intelligence; its applications span all areas of artificial intelligence. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and learn-by-doing.
[0044] In summary, the solutions provided by the embodiments of this disclosure involve optical encryption technology and artificial intelligence technology. The embodiments of this disclosure will be further described below with reference to the accompanying drawings.
[0045] Figure 1 A schematic diagram of the structure of an encrypted image information acquisition device 100 according to an embodiment of the present disclosure is shown.
[0046] like Figure 1As shown, an embodiment of this disclosure provides an encrypted image information acquisition device 100, including: at least one mask 1001 and an optical sensing component 1002. The at least one mask 1001 is used to receive object light 1004 of a target object 1003 and generate an encrypted optical image signal 1005 based on the received object light 1004. The object light 1004 carries image information of the target object, and the encrypted optical image signal 1005 carries encrypted image information. Each mask 1001 includes a mask pattern 1006. The optical sensing component 1002 is used to receive the encrypted optical image signal 1005, convert the encrypted optical image signal 1005 into an electrical signal, and output the electrical signal, wherein the electrical signal includes the encrypted image information of the target object.
[0047] Optionally, the target object 1003 is the object to be photographed, which can be a real person, object, animal, etc., which reflects light so that the target light carrying the image information of the object to be photographed enters the mask 1001. Any target object 1003 that can emit light and / or reflect light to the mask 1001 can be optically encrypted using the encrypted image information acquisition device 100 described in this disclosure. It should be understood that the light described in this disclosure can refer to various forms of visible or invisible light, such as diffuse light, laser light, monochromatic light, and polychromatic light.
[0048] Optionally, photomask 1001 typically includes a transparent substrate and a light-shielding film. Common materials for the transparent substrate include transparent glass (such as quartz glass, soda glass, low-expansion glass, etc.) and transparent resin. The light-shielding film typically includes rigid light-shielding films (e.g., chromium oxide film, iron oxide, molybdenum silicide, etc.) and latex. The light-shielding or light-transmitting patterns on photomask 1001 are currently achieved through coating or etching processes. It should be understood that these are merely examples of photomask materials and processing techniques; all transparent or light-shielding materials that can be processed into photomasks, and all processes that enable the processing of photomasks to form desired light-transmitting and light-shielding areas, can be applied to the photomasks described in this disclosure.
[0049] Optionally, the number of photomasks 1001 can be one or multiple. When the photomasks 1001 are multiple, they are placed parallel to each other between the target object and the photosensitive component. The mask pattern on each layer is used to encrypt the received optical signal and generate an encrypted optical image signal. Therefore, the processing result of one layer of photomasks can be the input optical signal of the next layer.
[0050] Optionally, the optical sensing component 1002 can be a photosensitive chip, a photoelectric conversion device, etc. For example, common optical sensing components include CCD (charge-coupled device) type optical sensing components and CMOS (metal-oxide-semiconductor) type optical sensing components. After the photosensitive image point of a CCD receives light, the photosensitive element generates a corresponding current, the magnitude of which corresponds to the light intensity. The photosensitive element directly outputs an analog electrical signal. In a CMOS, each photosensitive element directly integrates an amplifier and analog-to-digital conversion logic. When the photodiode receives light and generates an analog electrical signal, the electrical signal is first amplified by the amplifier in the photosensitive element and then directly converted into a corresponding digital signal. The main purpose of both CCD-type and CMOS-type optical sensing components is to convert the acquired light signal into an electrical signal that can be processed by subsequent circuits or computers. All devices that can convert light signals into signals that can be used by processing equipment can be considered optical sensing components as described in this disclosure.
[0051] For example, when a person stands on the ground and reflects sunlight, the reflected sunlight carries facial image information. This reflected sunlight passes through a mask and enters a CMOS sensor. The mask and CMOS sensor are placed coaxially, as close as possible to the CMOS sensor, along the direction of sunlight propagation. The facial image information in the reflected sunlight is encrypted by the mask, carrying the encrypted facial image information. The CMOS sensor receives the encrypted sunlight and generates an encrypted facial image, which cannot be directly recognized by the human eye. Based on the above, this disclosure protects user privacy by replacing the lens group in a camera lens with a mask and optically encrypting the acquired image. It also omits modules such as de-mosaicing and noise reduction designed for perception enhancement in traditional camera image signal processing, simplifying the hardware structure and achieving miniaturization of the optical system. Figure 2 Another schematic diagram of the structure of the encrypted image information acquisition device 100 according to an embodiment of the present disclosure is shown.
[0052] like Figure 2 As shown, according to an embodiment of this disclosure, the encrypted image information acquisition device 100 further includes a light generating component 1006 for generating an optical image signal 1007 carrying image information, wherein the optical image signal 1007 carrying image information is an incoherent light signal; wherein the light generating component 1006 is a plurality of point light sources, and the optical image signal 1007 carrying image information is a light signal generated by the plurality of point light sources; or the light generating component 1006 is a display, and the optical image signal 1007 carrying image information is a multi-pixel image generated by the display.
[0053] Alternatively, when using the light generating component 1006 to generate an optical image signal 1007 carrying image information, the mask 1001 receives the optical image signal 1007 carrying image information and encrypts the image information using the same method described above for the target object 1003.
[0054] Figure 3 Another schematic diagram of the structure of the encrypted image information acquisition device 100 according to an embodiment of the present disclosure is shown.
[0055] like Figure 3 As shown, if the electrical signal output by the optical sensing component 1002 needs to be further processed, the encrypted image information acquisition device 100 may further include an information extraction component 1008, which is used to receive the electrical signal output by the optical sensing component 1002 and extract features from the encrypted image information based on the electrical signal.
[0056] Optionally, the information extraction component 1008 described in this disclosure may be any component capable of calculating or processing signals, such as a computer, processor, server, integrated circuit, or a combination of one or more of the above components.
[0057] Optionally, the information extraction component includes a deconvolutional network and a feature extraction network, wherein the deconvolutional network is used to deconvolve the encrypted image information to obtain the recovered image information, and the feature extraction network is used to extract image features from the recovered image information.
[0058] Optionally, the deconvolutional network performs deconvolution on the encrypted image information using the following equation: X = F -1 (F(W)⊙F(Y)); where X is the recovered image information, W is the regularized pseudo-inverse of the point spread function (PSF) of the encrypted image information acquisition device 100, and Y is the encrypted image information, and
[0059] Optionally, the point extension function describes the response of the encrypted image information acquisition device 100 to a point source or point object.
[0060] Optionally, the extracted image features can be further applied to tasks such as feature classification or recognition.
[0061] Figure 4 A schematic diagram showing the relationship between the distances between the target object 1003, the mask 1001, and the optical sensing component 1002 according to an embodiment of the present disclosure is shown.
[0062] like Figure 4As shown, according to an embodiment of this disclosure, at least one mask 1001 is a mask, and the convolutional network includes a convolutional layer, wherein the distance d between two adjacent point light sources or two adjacent pixels is... L The distance d between the light generating component 1006 and the photomask 1001 LM The distance d between the photomask 1001 and the optical sensing component 1002 MS Since light travels in a straight line, according to the law of similar triangles (i.e., ...), The dimensions Δ of individual pixels on the optical sensing component 1002 have the following relationship: In this context, the size Δ of a single pixel on the optical sensing component 1002 is equivalent to the size of a single pixel generated after the convolution calculation of the convolutional layer.
[0063] The distance d between two adjacent point light sources or two adjacent pixels L This represents the resolution of the image information that the encrypted image information acquisition device 100 can acquire, according to the formula... It can be seen that the distance d between the mask 1001 and the optical sensing component 1002 is... MS The smaller the distance d between the light-generating component 1006 and the mask 1001, the better. LM The larger the resolution, the higher the resolution of the encrypted image information acquisition device 100. Therefore, the encrypted image information acquisition device 100 can be placed at a distance from the light generating component 1006. The mask 1001 should be placed as close as possible to the optical sensing component 1002, which can improve the resolution of the encrypted image information acquisition device 100 and reduce the size of the encrypted image information acquisition device 100.
[0064] After arranging the mask 1001 and the optical sensing component 1002 as close as possible, the distance between them can be obtained by measurement. Optionally, the distance measurement can be achieved using various high-precision ranging methods such as high-precision rulers, electrical ranging, and optical ranging.
[0065] Because the encrypted image information acquisition device 100 may have installation errors during installation, Δ and d MS The calculation and measurement may also introduce errors. In order to ensure the accuracy of the encrypted image information acquisition device 100, the encrypted image information acquisition device 100 can be fine-tuned according to the encryption result.
[0066] It should be understood that the fine-tuning process here is to verify whether the processing results of the device during initial installation deviate from the theoretical values. If errors exist, fine-tuning is used to reduce them. For devices that have already been debugged and verified, this fine-tuning process is not necessary.
[0067] Once the placement position of the encrypted image information acquisition device 100 is determined, it receives the object light 1004 of the target object 1003 and generates an encrypted optical image signal 1005. By comparing the resolution calculated by the encrypted image information acquisition device 100 with the resolution of the encrypted image information carried in the encrypted optical image signal 1005, the information loss during the image encryption process can be measured.
[0068] Based on the above, in this disclosure, the resolution of the encrypted image information acquisition device can be adjusted by changing the distance between the target object, the mask, and the optical sensing component.
[0069] Figure 5 A schematic diagram of a mask 1001 pattern according to an embodiment of the present disclosure is shown.
[0070] Optionally, the image information is convolved based on the received object light 104 using a mask pattern. At least one mask 1001 corresponds to at least one convolutional layer, and the mask pattern on each mask 1001 carries the parameters of the convolution kernel of its corresponding convolutional layer for convolution processing.
[0071] Optionally, the mask pattern of each mask 1001 is composed of multiple mask holes 1009, and the light transmittance of each mask hole 1009 is the same or different, wherein the light transmittance of each mask hole 1009 is determined by the parameters of its corresponding convolution kernel.
[0072] For example, the convolution kernel can be a binary distribution of 0 / 1, where 0 represents opaque and 1 represents fully transparent. The convolution kernel can also be an analog quantity, for example, 0 represents opaque, 1 represents fully transparent, and 0.5 represents half of the light being transmitted.
[0073] like Figure 5 The mask apertures 1009 on the mask 1001 shown have a binary distribution of 0 / 1, where 0 represents opaque and 1 represents fully transparent. Therefore, the convolution kernel corresponding to the mask apertures 1009 on the mask 1001 is:
[0074] Figure 6 A flowchart illustrating a method 600 for determining a mask pattern according to an embodiment of the present disclosure is shown. Figure 6 As shown, in step S601, a convolution model of the mask pattern is established, and based on the convolution model of the mask pattern, the training image used for training is encrypted to obtain an encrypted training image.
[0075] Optionally, the training image can be an image carrying information such as people, objects, animals, or text. The mask pattern carries the parameters of the convolution kernel. A convolution model of the mask is established, and the training image is convolved by the convolution model to obtain the encrypted training image.
[0076] In step S602, feature extraction is performed on the encrypted training image using a backend neural network to obtain a feature-extracted image.
[0077] Optionally, the backend neural network includes a deconvolutional network and a feature extraction network, wherein the deconvolutional network is used to deconvolve the encrypted image information to obtain the recovered image information, and the feature extraction network is used to extract image features from the recovered image information.
[0078] Alternatively, the deconvolutional network can be an inverse filter or a Wiener filter.
[0079] Optionally, the feature extraction network can be a network structure such as U-Net, LeNet-5, AlexNet, VGGNet, GoogLeNet, ResNet, DenseNet, etc., and the appropriate feature extraction network can be determined for different feature extraction needs.
[0080] Figure 7 A schematic diagram of a U-Net network structure according to an embodiment of the present disclosure is shown.
[0081] like Figure 7 As shown, the U-Net network is a type of convolutional neural network, initially applied in biological image processing, such as cell segmentation. The overall structure of the U-Net network consists of an encoding stage on the left and a decoding stage on the right. The encoding stage on the left is also called the contraction path, and its structure is a basic convolutional neural network structure, including basic convolutional layers, downsampling layers, and activation function layers. Because the feature images of each layer need to be contracted, no additional padding is needed during convolution operations. The decoding stage on the right is also called the dilation path. The network structure of the dilation path is usually a mirror image of the contraction path. Unlike the contraction path, after each convolutional part, instead of downsampling, it uses deconvolution to upsampling, thus enlarging the feature map for high-resolution image segmentation. Simultaneously, after each upsampling, the corresponding downsampling feature map is copied and concatenated with the current feature map, maximizing the preservation of effective information in the data and enabling high-precision segmentation tasks.
[0082] For example, the U-Net network takes a 572*572 encrypted training image as input, uses two repeated 3×3 convolutional kernels, employs a 2×2 pooling layer with a stride of 2 for downsampling, selects the ReLU linear activation unit as the activation function, and uses a 1×1 convolutional kernel for classification in the last layer of the network structure to obtain each final classification prediction result.
[0083] U-net can process input images of any size and can be trained with very few images. When extracting features from encrypted training images, the specific network structure can be adjusted according to the needs of the actual task.
[0084] Optionally, a classification network structure, such as the Inception-v2 or Inception-v3 network structure, can be connected after the feature extraction network. It should be understood that these are merely examples of network structures for classification, and all network structures implementing feature classification can be applied to this disclosure.
[0085] In step S603, based on the visual task to be completed, the loss function of the mask pattern and the signal processing of the back-end neural network is determined.
[0086] Optionally, different loss functions can be designed for different task objectives. For example, in a face recognition task, images containing faces are used for training, and encrypted face images are restored and features are extracted. The loss function can be a weighted sum of the cosine distance between the encrypted face image and the restored face image, and the triplet distance of the extracted face feature images, used to evaluate the degree of deviation between the predicted value and the true value.
[0087] In step S604, the convolutional model of the mask pattern is trained based on the loss function to obtain the parameter information of the trained convolutional model.
[0088] The mask pattern is determined based on the kernel parameters of the trained convolutional model, and then obtained by coating or etching on the mask plate 1001.
[0089] Based on the above, this disclosure combines optical encryption with a backend neural network to achieve joint optimization of the entire link through an end-to-end network, thereby optimizing the mask and improving the accuracy of recovering and recognizing encrypted images.
[0090] Figure 8 A flowchart of an encrypted image information acquisition method 800 according to an embodiment of the present disclosure is shown.
[0091] like Figure 8As shown, in step S801, the object light of the target object is received through at least one mask, and an encrypted optical image signal is generated based on the received object light. The object light carries the image information of the target object, and the encrypted optical image signal carries the encrypted image information. Each mask includes a mask pattern.
[0092] Optionally, the image information is convolved based on the received object light 104 using a mask pattern. At least one mask 1001 corresponds to at least one convolutional layer, and the mask pattern on each mask 1001 carries the parameters of the convolution kernel of its corresponding convolutional layer for convolution processing.
[0093] In step S802, the encrypted optical image signal is converted into an electrical signal and the electrical signal is output, wherein the electrical signal includes the encrypted image information of the target object.
[0094] The encryption method for acquiring encrypted image information disclosed herein enables the encryption process to be implemented in the optical domain through convolution. The result acquired by the optical sensing component is no longer the actual optical image, but the convolutional information of the actual image information after convolution processing. Therefore, this method can be used to protect and encrypt the actual image information.
[0095] Optionally, if subsequent recovery and feature extraction of the encrypted image information are required, the information extraction component 1008 can receive the electrical signal output by the optical sensing component 1002 and continue to process the image information based on the electrical signal.
[0096] The encrypted image information acquisition device and method disclosed herein replace the lens group in the camera lens with a mask and optically encrypt the acquired image, protecting user privacy. It also omits modules such as de-mosaicing and noise reduction designed for perception enhancement in the image signal processing of traditional cameras, simplifying the hardware structure and realizing the miniaturization of the optical system structure. Compared with digital encryption methods, it saves computation and improves encryption speed, making real-time recognition possible.
[0097] It should be noted that the flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing at least one executable instruction for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0098] In general, the various exemplary embodiments of this disclosure can be implemented in hardware or dedicated circuitry, software, firmware, logic, or any combination thereof. Some aspects can be implemented in hardware, while others can be implemented in firmware or software that can be executed by a controller, microprocessor, or other computing device. When aspects of embodiments of this disclosure are illustrated or described as block diagrams, flowcharts, or using some other graphical representation, it will be understood that the blocks, apparatuses, systems, techniques, or methods described herein can be implemented as non-limiting examples in hardware, software, firmware, dedicated circuitry or logic, general-purpose hardware or controllers or other computing devices, or some combination thereof.
[0099] The foregoing description is illustrative of the invention and should not be construed as limiting it. Although several exemplary embodiments of the invention have been described, those skilled in the art will readily understand that many modifications can be made to the exemplary embodiments without departing from the novel teachings and advantages of the invention. Therefore, all such modifications are intended to be included within the scope of the invention as defined in the claims. It should be understood that the foregoing description is illustrative of the invention and should not be construed as limiting it to the specific embodiments disclosed, and modifications to the disclosed embodiments and other embodiments are intended to be included within the scope of the appended claims. The invention is defined by the claims and their equivalents.
Claims
1. An encrypted image information acquisition device, comprising: At least one photomask and optical sensing components, wherein, The at least one mask is used to receive object light from a target object and generate an encrypted optical image signal based on the received object light, wherein the object light carries image information of the target object, and the encrypted optical image signal carries encrypted image information, wherein each mask includes a mask pattern. The optical sensing component is used to receive the encrypted optical image signal, convert the encrypted optical image signal into an electrical signal, and output the electrical signal, wherein the electrical signal includes the encrypted image information of the target object; The step of receiving the object light of the target object and generating an encrypted optical image signal based on the received object light includes: Using the mask pattern, the image information is convolved based on the received object light, wherein the convolutional network performing the convolutional processing includes at least one convolutional layer. The at least one mask layer corresponds one-to-one with the at least one convolutional layer. Each mask pattern carries parameter information for convolution processing of its corresponding convolutional layer, wherein the parameter information includes at least the parameters of the convolution kernel.
2. The encrypted image information acquisition device as described in claim 1, wherein, Each mask pattern consists of multiple mask holes, and the light transmittance of each mask hole may be the same or different. The light transmittance of each mask aperture is determined by the parameters of its corresponding convolution kernel.
3. The encrypted image information acquisition device as described in claim 2, wherein, The mask pattern is determined in the following manner: A convolutional model of the mask pattern is established, and the training image used for training is encrypted based on the convolutional model of the mask pattern to obtain the encrypted training image. Feature-extracted images are obtained by using a backend neural network to extract features from the encrypted training images. Based on the visual task to be completed, determine the loss function for signal processing of the back-end neural network; as well as Based on the loss function, the convolutional model of the mask pattern is trained to obtain the parameter information of the trained convolutional model.
4. The encrypted image information acquisition device as described in claim 1, wherein, The target object is the object to be photographed.
5. The encrypted image information acquisition device as described in claim 1, wherein, The target object is an optical image signal carrying image information of the target object. The encrypted image information acquisition device further includes a light generating component for generating an optical image signal carrying image information, wherein the optical image signal carrying image information is an incoherent light signal. Wherein, the light generating component comprises multiple point light sources, and the optical image signal carrying image information is a light signal generated by the multiple point light sources; or The light-generating component is a display, and the optical image signal carrying image information is a multi-pixel image generated by the display.
6. The encrypted image information acquisition device as described in claim 5, wherein, The at least one mask layer is a single mask layer, and the convolutional network includes a single convolutional layer, wherein the distance d between two adjacent point light sources or two adjacent pixels is... L The distance d between the light generating component and the photomask LM The distance d between the photomask and the optical sensing component MS The dimensions Δ of individual pixels on the optical sensing component have the following relationship: In this context, a single pixel on the optical sensing component is equivalent to a single pixel generated after the convolution calculation of the convolutional layer.
7. The encrypted image information acquisition device as described in claim 1, wherein, The mask pattern is obtained by coating or etching on each layer of the mask.
8. The encrypted image information acquisition device as described in claim 1, further comprising: An information extraction component is used to receive the electrical signal output by the optical sensing component and to extract features from the encrypted image information based on the electrical signal.
9. The encrypted image information acquisition device as described in claim 8, in, The information extraction component includes a deconvolutional network and a feature extraction network. The deconvolutional network is used to deconvolve the encrypted image information to obtain the recovered image information, and the feature extraction network is used to extract image features from the recovered image information.
10. The encrypted image information acquisition device as described in claim 9, wherein, The deconvolutional network performs deconvolution on the encrypted image information using the following equation: X=F -1 (F(W)⊙F(Y)); Where X is the recovered image information, W is the regularized pseudo-inverse of the point spread function of the encrypted image information acquisition device, and Y is the encrypted image information.
11. A method for acquiring encrypted image information, comprising: The object light of the target object is received through at least one mask, and an encrypted optical image signal is generated based on the received object light. The object light carries the image information of the target object, and the encrypted optical image signal carries the encrypted image information. Each mask includes a mask pattern. The encrypted optical image signal is converted into an electrical signal and the electrical signal is output, wherein the electrical signal includes the encrypted image information of the target object; The step of receiving object light from the target object through at least one mask and generating an encrypted optical image signal based on the received object light includes: The image information is convolved based on the received object light using the mask pattern, wherein the convolutional network performing the convolutional processing includes at least one convolutional layer. The at least one mask layer corresponds one-to-one with the at least one convolutional layer. Each mask layer carries the mask pattern carrying parameter information for convolution processing of its corresponding convolutional layer, wherein the parameter information includes at least the parameters of the convolution kernel. Each mask pattern consists of multiple mask holes, and the light transmittance of each mask hole may be the same or different. The light transmittance of each mask aperture is determined by the parameters of its corresponding convolution kernel.
12. The method for obtaining encrypted image information as described in claim 11, wherein, The target object is the object to be photographed, or the target object is an optical image signal carrying image information of the target object. The encrypted image information acquisition method further includes generating an optical image signal carrying image information, wherein the optical image signal carrying image information is an incoherent optical signal. The optical image signal carrying image information is a light signal generated by multiple point light sources; or The optical image signal carrying image information is a multi-pixel image generated by the display.
13. The method for obtaining encrypted image information as described in claim 12, wherein, The at least one mask layer is a single mask layer, and the convolutional network includes a single convolutional layer, wherein the distance d between two adjacent point light sources or two adjacent pixels is... L The distance d between the point light source or display and the photomask LM The distance d between the photomask and the optical sensing component MS The dimensions Δ of individual pixels on the optical sensing component have the following relationship: In this context, a single pixel on the optical sensing component is equivalent to a single pixel generated after the convolution calculation of the convolutional layer.
14. The method for obtaining encrypted image information as described in claim 12, further comprising: The system receives the electrical signal output from the optical sensing component and extracts features from the encrypted image information based on the electrical signal.
15. The method for obtaining encrypted image information as described in claim 14, wherein, The feature extraction of the encrypted image information based on the electrical signal includes: The encrypted image information is deconvolved to obtain the recovered image information; and Image features are extracted from the recovered image information.
16. The method for obtaining encrypted image information as described in claim 15, wherein, The encrypted image information is deconvolved using the following equation: X=F -1 (F(W)⊙F(Y)); Where X is the recovered image information, W is the regularized pseudo-inverse of the point spread function of the encrypted image information acquisition device, and Y is the encrypted image information.
Citation Information
Patent Citations
Image processing method and device, equipment and storage medium
CN113705307A
Camera and imaging system
WO2021075527A1