Optical task processing system and training method and device thereof
By encoding the incident light with optical features using an optical task processing system to generate a two-dimensional digital matrix, the target task processing result is directly generated. This solves the privacy leakage problem in the intelligent sensing architecture and achieves the effects of security and system miniaturization.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHPHOTONICS LTD
- Filing Date
- 2026-04-08
- Publication Date
- 2026-06-02
Smart Images

Figure CN122133733A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to fields such as metasurface optical coding, intelligent sensing, and deep learning, and particularly to optical task processing systems and their training methods and apparatus. Background Technology
[0002] Current intelligent sensing architectures typically employ an "image-then-computation" approach. This involves first capturing raw images of the scene using imaging devices, then processing these raw images (such as analysis and recognition) using computing units to ultimately obtain the task results. However, this approach carries a significant risk of privacy breaches, and security cannot be guaranteed. Summary of the Invention
[0003] This disclosure provides an optical task processing system and its training methods and apparatus.
[0004] An optical task processing system includes: an optical coding module and a task processing unit;
[0005] The optical coding module is used to encode the optical features of the incident light and generate a two-dimensional digital matrix based on the encoded light field. The incident light is the incident light that carries the information to be processed in the target scene. The task processing unit is used to generate the target task processing result corresponding to the target scene based on the two-dimensional digital matrix.
[0006] A training method for an optical task processing system includes: Obtain training samples, which include sample images and real labels, wherein the real labels are the real task processing results corresponding to the sample images; The optical task processing system described above is used to generate the prediction task processing result corresponding to the sample image, and the result is determined as the prediction label. The target loss is determined based on the predicted label and the true label, and the parameters of the optical coding module and / or task processing unit in the optical task processing system are updated based on the target loss.
[0007] A training device for an optical task processing system includes: a sample acquisition module, a label generation module, and a parameter update module; The sample acquisition module is used to acquire training samples, which include sample images and real labels, wherein the real labels are the real task processing results corresponding to the sample images. The label generation module is used to generate the prediction task processing result corresponding to the sample image using the optical task processing system described above, and determine it as the prediction label. The parameter update module is used to determine the target loss based on the predicted label and the real label, and update the parameters of the optical coding module and / or task processing unit in the optical task processing system based on the target loss.
[0008] An electronic device, comprising: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method described above.
[0009] A non-transitory computer-readable storage medium storing computer instructions for causing a computer to perform the methods described above.
[0010] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0011] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein: Figure 1 This is a schematic diagram of the composition structure of Embodiment 100 of the optical task processing system described in this disclosure; Figure 2 This is a schematic diagram of the first component structure of the optical coding module 11 described in this disclosure; Figure 3 This is a schematic diagram of the second component structure of the optical coding module 11 described in this disclosure; Figure 4 This is a schematic diagram of the third component structure of the optical coding module 11 described in this disclosure; Figure 5 This is a schematic diagram of the fourth component structure of the optical coding module 11 described in this disclosure; Figure 6 This is a schematic diagram of the fifth component structure of the optical coding module 11 described in this disclosure; Figure 7 This is a schematic diagram of the sixth component structure of the optical coding module 11 described in this disclosure; Figure 8 This is a schematic diagram of the seventh component structure of the optical coding module 11 described in this disclosure; Figure 9 This is a flowchart illustrating an embodiment of the training method for the optical task processing system described in this disclosure; Figure 10This is a schematic diagram of the composition structure of an embodiment 1000 of the optical task processing system described in this disclosure; Figure 11 A schematic block diagram of an electronic device 1100 that can be used to implement embodiments of the present disclosure is shown. Detailed Implementation
[0012] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0013] Furthermore, it should be understood that the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / " in this article generally indicates that the preceding and following related objects have an "or" relationship.
[0014] Figure 1 This is a schematic diagram of the structural composition of an embodiment 100 of the optical task processing system described in this disclosure. Figure 1 As shown, it includes: an optical coding module 11 and a task processing unit 12.
[0015] The optical encoding module 11 is used to encode the optical features of the incident light and generate a two-dimensional digital matrix based on the encoded light field. The incident light is the incident light carrying the information to be processed in the target scene.
[0016] The task processing unit 12 is used to generate the target task processing result corresponding to the target scene based on the two-dimensional digital matrix.
[0017] As can be seen, by adopting the scheme described in the above system embodiment, the optical coding module 11 can be used to encode the optical features of the incident light, and a two-dimensional digital matrix can be generated based on the encoded light field. Then, the task processing unit 12 can generate the target task processing result based on the two-dimensional digital matrix. The entire process does not involve the generation of images, that is, the original scene information is encoded into a matrix that cannot be intuitively understood. The task processing unit 12 only generates the target task processing result based on this matrix, thereby reducing the risk of privacy leakage and improving the security of the system.
[0018] The incident light can carry information to be processed in the target scene, such as faces, fingerprints, and information about specific objects.
[0019] In some specific implementations, the two-dimensional digital matrix constitutes an optical feature tensor related to the target task. This optical feature tensor is the raw data representation obtained after optically encoding the incident light field. At the data structure level, this optical feature tensor is directly configured as the input to the task processing unit (such as a decision circuit, a corresponding decoding algorithm, or a neural network model) without any intermediate image transformation. Specifically, the decision circuit can parse the final recognition, detection, or classification result from this tensor based on a preset mapping rule.
[0020] In other embodiments, the optical feature tensor can also be used as source data for image generation. Specifically, by processing the optical feature tensor through a task processing unit (such as a corresponding decoding algorithm or neural network model), images visible to the human eye or images with specific physical meaning (such as depth maps or spectral maps) can be generated to meet the needs of visual monitoring or manual interpretation. In other words, the technical solution of this disclosure supports both direct machine decision-making based on optical feature tensors and image generation paths based on the same tensor, and can be flexibly configured according to the application scenario.
[0021] Furthermore, the two-dimensional digital matrix corresponding to this optical feature tensor can be divided into multiple functional sub-regions, each sub-region's numerical set corresponding to a specific feature dimension. The spatial distribution of local maxima in the matrix can be used to directly characterize the target task results, including but not limited to the target's category, spatial location, or confidence level.
[0022] Figure 2 This is a schematic diagram of the first component structure of the optical coding module 11 described in this disclosure. Figure 2 As shown, it includes a metasurface unit 21 and an image sensor 22. The metasurface unit 21 is used to encode and modulate the incident light and output the encoded light field. The image sensor 22 is used to generate a two-dimensional digital matrix based on the encoded light field.
[0023] Specifically, the metasurface unit 21 comprises multiple metaatomic arrays with preset geometric parameters (such as diameter, height, and rotation angle). These metaatomic arrays are configured to perform at least one target-specific optical computational operation on the incident light field. This optical computational operation includes, but is not limited to, linear transformations (such as inner product operations, convolution, spatial filtering, and spectral weighted linear combinations), feature space mapping (such as projecting the input scene onto a low-dimensional feature space), or matrix operations in optical neural networks. These operations physically control the incident light field, allowing the encoded light field to directly carry target-task-related feature information.
[0024] Image sensor 22 may include a plurality of photosensitive units arranged in a two-dimensional manner, used to receive the light field encoded by metasurface unit 21, and spatially sample at least one optical physical quantity of the light field (e.g., light intensity distribution, phase distribution, polarization state, spectral intensity or a combination thereof), outputting the corresponding digital measurement value, thereby forming the two-dimensional digital matrix. This matrix can be directly used as input to task processing unit 12 without reconstructing it into a visual image.
[0025] The image sensor 22 can be any photoelectric detection array, such as a complementary metal-oxide-semiconductor (CMOS) sensor or a charge-coupled device (CCD), capable of converting optical signals into electrical signals and reading them out as a digital matrix.
[0026] In practical applications, the metasurface unit 21 can have various different compositions and structures, which will be introduced below.
[0027] Figure 3 This is a schematic diagram of the second component structure of the optical coding module 11 described in this disclosure. Figure 4 This is a schematic diagram of the third component structure of the optical coding module 11 described in this disclosure.
[0028] like Figure 3 and Figure 4 As shown, each of them can include: a metasurface unit 21 and an image sensor 22, and each metasurface unit 21 can include: a first substrate 211, a first micro / nano structure 212 (several, the number is unlimited) and a first protective layer 213.
[0029] The first micro / nano structure 212 is disposed on the first substrate 211. The first micro / nano structure 212 is used to encode the optical features of the incident light to obtain the encoded light field. The first protective layer 213 is used to fill and cover the first micro / nano structure 212, providing protection for the first micro / nano structure 212. The image sensor 22 is in close contact with the metasurface unit 21, receives the encoded light field, and outputs the corresponding two-dimensional digital matrix.
[0030] In some embodiments of this disclosure, the vertical spacing (d1) between the first micro / nano structure 212 and the image sensor 22 may be less than a first threshold.
[0031] The specific value of the first threshold can be determined according to actual needs, such as 1.5 mm. In extreme cases, d1 can approach zero. Here, "approaching zero" means that due to unavoidable microscopic surface roughness, manufacturing tolerances, or adsorption of interfacial atoms / molecules, d1 is not strictly a mathematical zero value in physics, but a very small distance at the nanometer or sub-nanometer level (such as less than 5 nm). This distance is small enough to achieve the expected optocoupler, signal transmission, or thermal management functions between the first micro / nano structure 212 and the image sensor 22, which is equivalent to direct contact in effect. When d1 approaches zero, it can be regarded as the first micro / nano structure 212 and the image sensor 22 achieving substantial contact or functional integration.
[0032] The material of the first substrate 211 may be at least one of the following: quartz glass, crystalline and amorphous silicon, alumina, silicon nitride, calcium fluoride, polymers with a refractive index of 1 to 1.5, composite polymers, etc.
[0033] The first micro / nano structure 212 can be a nanopillar, grating, metaatom, etc. After incident light irradiates the metasurface unit 21, the first micro / nano structure 212 can modulate the phase, amplitude, or polarization of the incident light to achieve specific optical feature encoding. Furthermore, the material of the first micro / nano structure 212 can be at least one of the following: titanium oxide, tantalum oxide, hafnium oxide, silicon nitride, photoresist, quartz glass, aluminum oxide, crystalline and amorphous silicon, gallium nitride, crystalline germanium, selenium sulfide, chalcogenide glass, etc.
[0034] The first protective layer 213 is used to fill and cover the first micro / nano structure 212, providing protection for the first micro / nano structure 212, i.e., preventing the first micro / nano structure 212 from being damaged due to exposure. The material of the first protective layer 213 may be at least one of the following: silicon dioxide, silicon nitride, and adhesive.
[0035] Figure 5 This is a schematic diagram of the fourth component structure of the optical coding module 11 described in this disclosure. Figure 6 This is a schematic diagram of the fifth component structure of the optical coding module 11 described in this disclosure. Figure 7 This is a schematic diagram of the sixth component structure of the optical coding module 11 described in this disclosure.
[0036] like Figure 5 , Figure 6 and Figure 7 As shown, each of them can include: a metasurface unit 21 and an image sensor 22, and each metasurface unit 21 can include: a first substrate 211, a first micro / nano structure 212 (several, unlimited in number), a second micro / nano structure 214 (several, unlimited in number), a first protective layer 213 and a second protective layer 215.
[0037] in, Figure 5In the first micro-nano structure 212, the first micro-nano structure 212 is arranged on the first substrate 211, and the second micro-nano structure 214 is arranged on the first protective layer 213. Figure 6 In the first substrate 211, the second micro-nano structure 214 is disposed on the first substrate 211, and the first micro-nano structure 212 is disposed on the second protective layer 215. Figure 7 In the first micro-nano structure 212 and the second micro-nano structure 214 are respectively arranged on both sides of the first substrate 211.
[0038] The second micro / nano structure 214 is used for pre-optimization of the incident light, and the second protective layer 215 is used to fill and cover the second micro / nano structure 214, providing protection for the second micro / nano structure 214. The first micro / nano structure 212 is used to encode the optical features of the pre-optimized light to obtain the encoded light field, and the first protective layer 213 is used to fill and cover the first micro / nano structure 212, providing protection for the first micro / nano structure 212.
[0039] In some embodiments of this disclosure, the vertical distance (d1) between the first micro / nano structure 212 and the image sensor 22 may be less than a first threshold, and the vertical distance (d2) between the second micro / nano structure 214 and the first micro / nano structure 212 may be greater than a second threshold, wherein the first threshold is greater than the second threshold.
[0040] The specific values of the first and second thresholds can be determined according to actual needs. For example, the first threshold can be 2mm, 1.5mm, or 1mm, and the second threshold can be 15nm, 10nm, 5nm, or 2nm.
[0041] By using the first threshold, the first micro / nano structure 212 and the image sensor 22 can be compactly integrated, thereby controlling the thickness of the system within the millimeter range and adapting to space-constrained scenarios. By using the second threshold, optical interference between the second micro / nano structure 214 and the first micro / nano structure 212 can be avoided, ensuring the accuracy of light modulation and encoding. In short, by using the two thresholds, both system miniaturization and optical performance can be balanced, which helps to facilitate efficient and accurate task processing.
[0042] After incident light illuminates the second micro / nano structure 214, the second micro / nano structure 214 can perform pre-optimization processing on the incident light, such as focusing and collimation. Then, the first micro / nano structure 212 can perform optical feature encoding on the pre-optimized light to obtain the encoded light field. The image sensor 22 is in close contact with the metasurface unit 21, receives the encoded light field, and outputs the corresponding two-dimensional digital matrix.
[0043] The type / morphology of the second micro / nano structure 214 can be nanopillars, gratings, superatoms, etc. The material of the second micro / nano structure 214 can be at least one of the following: titanium oxide, tantalum oxide, hafnium oxide, silicon nitride, photoresist, quartz glass, aluminum oxide, crystalline and amorphous silicon, gallium nitride, crystalline germanium, selenium sulfide, chalcogenide glass, etc.
[0044] The second protective layer 215 is used to fill and cover the second micro / nano structure 214, providing protection for the second micro / nano structure 214, i.e., preventing the second micro / nano structure 214 from being damaged due to exposure. The material of the second protective layer 215 may be at least one of the following: silicon dioxide, silicon nitride, or adhesive.
[0045] Figure 8 This is a schematic diagram of the seventh component structure of the optical coding module 11 described in this disclosure. Figure 8 As shown, it may include: a metasurface unit 21 and an image sensor 22, and the metasurface unit 21 may include: a first substrate 211, a first micro-nano structure 212 (a number of which are unlimited), a second micro-nano structure 214 (a number of which are unlimited), a first protective layer 213, a second protective layer 215, and a second substrate 216.
[0046] The first micro / nano structure 212 is disposed on the first substrate 211, the second micro / nano structure 214 is disposed on the second substrate 216, and the second substrate 216 is disposed on the first protective layer 213.
[0047] The second micro / nano structure 214 is used for pre-optimization of the incident light, and the second protective layer 215 is used to fill and cover the second micro / nano structure 214, providing protection for the second micro / nano structure 214. The first micro / nano structure 212 is used to encode the optical features of the pre-optimized light to obtain the encoded light field, and the first protective layer 213 is used to fill and cover the first micro / nano structure 212, providing protection for the first micro / nano structure 212.
[0048] The material of the second substrate 216 may be at least one of the following: quartz glass, crystalline and amorphous silicon, alumina, silicon nitride, calcium fluoride, polymers with a refractive index of 1 to 1.5, composite polymers, etc.
[0049] In some embodiments of this disclosure, the target task processing result may include: the target task processing result determined by comparing a two-dimensional digital matrix with a predetermined template, or the target task processing result output by a neural network model after inputting a two-dimensional digital matrix into a neural network model. In addition, the number of target task processing results may be one or more.
[0050] In other words, the task processing unit 12 can determine the target task processing result by comparing the two-dimensional numerical matrix with a predetermined template, or it can input the two-dimensional numerical matrix into a neural network model to obtain the target task processing result output by the neural network model, which is very flexible and convenient. The specific type of neural network model is not limited; it can be a convolutional neural network model or other deep learning models.
[0051] In addition, the target task processing result output by the task processing unit 12 can be one or multiple. For example, in a multi-task parallel recognition scenario, multiple target task processing results can be output.
[0052] In some embodiments of this disclosure, the optical coding module 11 may include: an optical coding module fabricated as a single chip using a wafer-level integration process.
[0053] Through the above processing, the system thickness can be significantly reduced, such as achieving an ultra-thin design with a system thickness of less than 3mm, to adapt to mobile / IoT devices, etc., to achieve ultra-thin and low-power deployment, and to enhance environmental robustness and adapt to reliable operation in multiple scenarios.
[0054] Taking the image sensor 22 as a CMOS sensor as an example, the integration of the metasurface unit 21 and the image sensor 22 can be carried out in the following two ways: 1) Direct monolithic integration: On the passivation layer of the sensor wafer after CMOS process, the metasurface unit 21 is directly prepared by compatible back-end nanofabrication (such as low temperature atomic layer deposition, nanoimprinting). The two interact directly through optical near-field coupling without the need for additional electrical connection; 2) Heterogeneous integration and stacked bonding: The separately prepared metasurface wafer and sensor wafer are vertically integrated by hybrid bonding technology. The bonding interface needs to form a transparent optical window to ensure unobstructed optical path, and electrical connection is achieved through copper interconnect (if the metasurface needs to be electrically controlled).
[0055] In addition, if the integration of multiple metasurfaces is involved, the following methods can be used: 1) Vertical stacking: Metasurface layers with different functions are bonded through nanoscale alignment and dielectric filling (such as spin-coated glass) to form a compact multilayer optical system; 2) In-plane coupling: Metasurfaces serve as input / output interfaces for on-chip optical networks, and optical signals are transmitted and processed in the chip plane through silicon waveguides or free space design; 3) Electrical interconnection: For active metasurface arrays, independent electrode addressing is achieved through interlayer vias and redistribution layers to realize dynamic and programmable optical coding.
[0056] The present disclosure also discloses the training method for the optical task processing system.
[0057] Accordingly, Figure 9 This is a flowchart illustrating an embodiment of the training method for the optical task processing system described in this disclosure. Figure 9As shown, the specific implementation methods are as follows.
[0058] In step 901, training samples are obtained, which include sample images and real labels. The real labels are the actual task processing results corresponding to the sample images.
[0059] In step 902, the optical task processing system is used to generate the prediction task processing result corresponding to the sample image and determine it as the prediction label.
[0060] In step 903, the target loss is determined based on the predicted label and the true label, and the parameters of the optical coding module and / or task processing unit in the optical task processing system are updated based on the target loss.
[0061] In other words, an end-to-end training method can be used to train the optical task processing system, that is, to jointly and synchronously optimize the optical coding module and the task processing unit. This allows the optical coding module and the task processing unit to be deeply adapted, reducing intermediate losses, improving processing accuracy, and enabling the system to maintain high recognition rate and stability in complex scenarios.
[0062] The optical coding module can perform optical feature encoding and other processing on the sample image to obtain a two-dimensional digital matrix. The task processing unit can generate a prediction task processing result based on the two-dimensional digital matrix, and then determine the target loss based on the prediction task processing result. The optical coding module and / or task processing unit can update the parameters according to the target loss until the predetermined convergence condition is met.
[0063] In the joint training phase, the optical feature encoding process can be implemented through computer simulation. For example, wave optics simulation tools (such as rigorous coupled-wave analysis, finite-difference time-domain method, etc.) can be used to simulate the light field distribution of the incident light corresponding to the sample image after being modulated by the metasurface, and generate the corresponding two-dimensional digital matrix based on the simulated light field distribution. This simulation process allows the gradient of the loss function to propagate back to the designable parameters of the metasurface (such as the geometric dimensions and rotation angles of the superatoms), thereby achieving indirect optimization of the optical encoding module.
[0064] Updating the parameters of an optical coding module refers to updating its configuration parameters, which may include the geometry, size, and spatial arrangement of the micro / nano structures within the metasurface unit. Updating the parameters of a task processing unit can refer to updating the model parameters of a neural network model.
[0065] In some embodiments of this disclosure, the number of predicted labels and the number of real labels are the same, and each predicted label corresponds to one real label. Accordingly, the method for determining the target loss based on the predicted labels and the real labels may include: in response to determining that the number of predicted labels is 1, determining a loss value based on the predicted labels and the real labels, and determining the loss value as the target loss; in response to determining that the number of predicted labels is greater than 1, determining the loss value corresponding to each predicted label based on the corresponding real label, and weighting and summing the loss values to obtain the target loss.
[0066] In other words, if there is only one predicted label, the loss value of the predicted label and the real label can be obtained directly, and this loss value can be determined as the target loss. There are no restrictions on how to obtain the loss value. For example, the loss value can be calculated using the mean squared error loss function. If there are multiple predicted labels, say three, the loss value corresponding to the three predicted labels can be determined according to the corresponding real labels. Then, the three loss values can be weighted and summed, and the weighted sum can be determined as the target loss.
[0067] By adopting the above processing method, the obtained target loss can accurately reflect the overall task performance, thereby improving training efficiency and training effect.
[0068] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this disclosure is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this disclosure. Secondly, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this disclosure. Furthermore, for parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.
[0069] The above is an introduction to the method embodiments. The following describes the solution described in this disclosure further through device embodiments.
[0070] Figure 10 This is a schematic diagram of the structural composition of an embodiment 1000 of the optical task processing system described in this disclosure. Figure 10 As shown, it includes: a sample acquisition module 1001, a label generation module 1002, and a parameter update module 1003.
[0071] The sample acquisition module 1001 is used to acquire training samples, which include sample images and real labels. The real labels are the real task processing results corresponding to the sample images.
[0072] The label generation module 1002 is used to generate the prediction task processing result corresponding to the sample image using the optical task processing system, and determine it as the prediction label.
[0073] The parameter update module 1003 is used to determine the target loss based on the predicted label and the real label, and to update the parameters of the optical coding module and / or task processing unit in the optical task processing system based on the target loss.
[0074] Optical task processing systems can provide Figure 1 The optical task processing system shown.
[0075] In some embodiments of this disclosure, the number of predicted labels and the number of real labels are the same, and each predicted label corresponds to one real label. Accordingly, the parameter update module 1003 may determine the target loss based on the predicted labels and the real labels in the following ways: in response to determining that the number of predicted labels is 1, a loss value is determined based on the predicted labels and the real labels, and the loss value is determined as the target loss; in response to determining that the number of predicted labels is greater than 1, the loss value corresponding to each predicted label is determined based on the corresponding real label, and the loss values are weighted and summed to obtain the target loss.
[0076] The specific workflow of the above-described device embodiments can be found in the relevant descriptions in the foregoing method and system embodiments, and will not be repeated here.
[0077] In summary, the solution described in this disclosure can reduce the risk of privacy leakage, improve system security, and enable ultra-thin, low-power deployment. Moreover, it has strong environmental robustness and can work stably in low-light or complex lighting environments. In addition, it can improve the accuracy of the acquired task processing results and reduce response latency, thus providing a new generation of contactless interaction solutions for privacy-sensitive and space-constrained scenarios.
[0078] Furthermore, the solutions described in this disclosure can be applied to various scenarios, demonstrating broad applicability. Examples include: consumer electronics products such as smartphones, tablets, smartwatches, smart bracelets, laptops, augmented reality (AR) glasses, and virtual reality (VR) glasses; smart homes and the Internet of Things (IoT) such as smart switch panels, smart kitchen devices, and bathroom fixtures; automobiles and vehicles such as in-vehicle infotainment systems; medical and monitoring applications such as aseptic control in operating rooms and intelligent monitoring in wards; and industrial and professional applications such as hazardous environments, clean rooms, and laboratories.
[0079] Furthermore, the collection, storage, use, processing, transmission, provision, and disclosure of user personal information involved in the technical solution disclosed herein all comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0080] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0081] Figure 11 A schematic block diagram of an electronic device 1100 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workbenches, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0082] like Figure 11 As shown, the electronic device 1100 includes a computing unit 1101, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 1102 or a computer program loaded from a storage unit 1108 into a random access memory (RAM) 1103. The RAM 1103 may also store various programs and data required for the operation of the electronic device 1100. The computing unit 1101, ROM 1102, and RAM 1103 are interconnected via a bus 1104. An input / output (I / O) interface 1105 is also connected to the bus 1104.
[0083] Multiple components in electronic device 1100 are connected to I / O interface 1105, including: input unit 1106, such as keyboard, mouse, etc.; output unit 1107, such as various types of displays, speakers, etc.; storage unit 1108, such as disk, optical disk, etc.; and communication unit 1109, such as network card, modem, wireless transceiver, etc. Communication unit 1109 allows electronic device 1100 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0084] The computing unit 1101 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1101 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSPs), and any suitable processor, controller, microcontroller, etc. The computing unit 1101 performs the various methods and processes described above, such as those described in this disclosure. For example, in some embodiments, the methods described in this disclosure can be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 1108. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 1100 via ROM 1102 and / or communication unit 1109. When the computer program is loaded into RAM 1103 and executed by the computing unit 1101, one or more steps of the methods described in this disclosure can be performed. Alternatively, in other embodiments, the computing unit 1101 may be configured to perform the methods described herein by any other suitable means (e.g., by means of firmware).
[0085] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard parts (ASSPs), systems-on-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0086] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0087] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory, read-only memory, erasable programmable read-only memory (EPROM), flash memory, optical fiber, compact disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0088] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a cathode ray tube (CRT) or liquid crystal display (LCD) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0089] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0090] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.
[0091] It should be understood that the various forms of processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.
[0092] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. An optical task processing system, characterized in that, include: Optical coding module and task processing unit; The optical coding module is used to encode the optical features of the incident light and generate a two-dimensional digital matrix based on the encoded light field. The incident light is the incident light that carries the information to be processed in the target scene. The task processing unit is used to generate the target task processing result corresponding to the target scene based on the two-dimensional digital matrix.
2. The system according to claim 1, characterized in that, The optical coding module includes: a metasurface unit and an image sensor; The metasurface unit is used to generate the encoded light field, and the image sensor is used to generate the two-dimensional digital matrix based on the encoded light field.
3. The system according to claim 2, characterized in that, The metasurface unit includes: a first substrate, a first micro / nano structure, and a first protective layer, wherein the first micro / nano structure is disposed on the first substrate; The first micro / nano structure is used to encode the optical features of the incident light to obtain the encoded light field, and the first protective layer is used to fill and cover the first micro / nano structure to provide protection for the first micro / nano structure.
4. The system according to claim 2, characterized in that, The metasurface unit includes: a first substrate, a first micro / nano structure, a second micro / nano structure, a first protective layer, and a second protective layer; The first micro / nano structure is disposed on the first substrate, and the second micro / nano structure is disposed on the first protective layer; or, the second micro / nano structure is disposed on the first substrate, and the first micro / nano structure is disposed on the second protective layer; or, the first micro / nano structure and the second micro / nano structure are respectively disposed on both sides of the first substrate. The second micro / nano structure is used to perform pre-optimization processing on the incident light, and the second protective layer is used to fill and cover the second micro / nano structure to provide protection for the second micro / nano structure. The first micro / nano structure is used to encode the optical features of the light after pre-optimization to obtain the encoded light field. The first protective layer is used to fill and cover the first micro / nano structure to provide protection for the first micro / nano structure.
5. The system according to claim 2, characterized in that, The metasurface unit includes: a first substrate, a second substrate, a first micro / nano structure, a second micro / nano structure, a first protective layer, and a second protective layer; The first micro / nano structure is disposed on the first substrate, the second micro / nano structure is disposed on the second substrate, and the second substrate is disposed on the first protective layer; The second micro / nano structure is used to perform pre-optimization processing on the incident light, and the second protective layer is used to fill and cover the second micro / nano structure to provide protection for the second micro / nano structure. The first micro / nano structure is used to encode the optical features of the light after pre-optimization to obtain the encoded light field. The first protective layer is used to fill and cover the first micro / nano structure to provide protection for the first micro / nano structure.
6. The system according to claim 4 or 5, characterized in that, The vertical distance between the first micro / nano structure and the image sensor is less than a first threshold. The vertical spacing between the second micro / nano structure and the first micro / nano structure is greater than a second threshold, and the first threshold is greater than the second threshold.
7. The system according to claim 1, characterized in that, The target task processing result includes: the target task processing result determined by comparing the two-dimensional digital matrix with a predetermined template, or the target task processing result output by the neural network model after inputting the two-dimensional digital matrix into the neural network model; The number of target task processing results is one or more.
8. The system according to claim 1, characterized in that, The optical coding module includes an optical coding module fabricated as a single chip using wafer-level integration technology.
9. A training method for an optical task processing system, characterized in that, include: Obtain training samples, which include sample images and real labels, wherein the real labels are the real task processing results corresponding to the sample images; The optical task processing system according to any one of claims 1-8 is used to generate the prediction task processing result corresponding to the sample image and determine it as the prediction label; The target loss is determined based on the predicted label and the true label, and the parameters of the optical coding module and / or task processing unit in the optical task processing system are updated based on the target loss.
10. The method according to claim 9, characterized in that, The number of predicted labels and the number of real labels are the same, and each predicted label corresponds to one real label; The step of determining the target loss based on the predicted label and the true label includes: In response to determining that the number of predicted labels is 1, a loss value is determined based on the predicted labels and the true labels, and the loss value is determined as the target loss; In response to determining that the number of predicted labels is greater than 1, the loss value corresponding to each predicted label is determined according to the corresponding real label, and the loss values are weighted and summed to obtain the target loss.
11. A training device for an optical task processing system, characterized in that, include: Sample acquisition module, label generation module, and parameter update module; The sample acquisition module is used to acquire training samples, which include sample images and real labels, wherein the real labels are the real task processing results corresponding to the sample images. The label generation module is used to generate the prediction task processing result corresponding to the sample image using the optical task processing system of any one of claims 1-8, and determine it as the prediction label; The parameter update module is used to determine the target loss based on the predicted label and the real label, and update the parameters of the optical coding module and / or task processing unit in the optical task processing system based on the target loss.
12. An electronic device, characterized in that, include: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 9-10.
13. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, The computer instructions are used to cause the computer to perform the method according to any one of claims 9-10.