Convolution Processing Method and Apparatus
Through mask processing and optical sensing components on the optical domain, convolutional calculation is realized, solving the problems of large size and large calculation volume of machine vision system, improving image recognition accuracy and processing efficiency, and reducing system costs.
Patent Information
- Application Number
- CN202210072390.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-21
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2042-01-21
AI Technical Summary
Existing machine vision systems require separate lens imaging equipment and back-end data processing equipment when processing images, resulting in a large system size and difficulty in adapting to application needs in micro-environments. The calculation of convolutional neural networks is large and requires high-performance post-processing components.
The convolution processing device consisting of a mask plate and an optical sensing component is used to realize convolution calculation through mask processing in the optical domain, reducing the calculation amount and power consumption of the post-processing component, and optical mask processing is performed using the mask pattern on the mask plate to convert it into an electrical signal for convolution processing of image information.
The size and weight of the convolution processing device are reduced, the performance requirements for post-processing components are reduced, the recognition accuracy and processing efficiency of image information are improved, and the protection and encryption of image information are realized.
Smart Images

Figure CN114418070B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the fields of image recognition and image processing, and more specifically, to a convolution processing method and apparatus in the optical domain. Background Art
[0002] Machine vision, as the "wise eye" in today's intelligent era, has a wide range of applications in many fields such as industrial inspection, smart home, and next-generation terminals. Images, as a very important visual information, are used as input information in many fields to assist the machine vision system to complete tasks.
[0003] However, there are many differences between machine vision and the human eye. Images in nature that can be understood by the human eye often require further processing by machine vision to complete tasks such as classification, detection, and recognition. The current processing flow of the machine vision system is that information such as images projects light onto a photosensitive chip through a lens imaging device. After the photosensitive chip collects the optical signal, it is converted into an electrical signal that can be processed by a computer. The computer calculates and processes the input electrical signal to implement subsequent image processing or image recognition tasks. In this processing method, a separate lens imaging device and a backend data processing device are required, and the entire machine vision system is relatively large in size, making it difficult to meet the usage requirements in a micro environment. In the context of the development of the machine vision system towards low power consumption, small size, and high computing efficiency, it is difficult to meet the application requirements in multiple scenarios.
[0004] The Convolutional Neural Network (CNN) is a feedforward neural network and is a popular neural network technology today. It is often applied to various image processing or image recognition tasks. Compared with the fully connected neural network, it has characteristics such as local area connection and weight sharing, with less computational complexity and stronger data processing capabilities. Currently, the processing of the convolutional neural network is basically implemented in backend data processing devices such as computers and processors, which requires a relatively high performance of the backend data processing device. Therefore, it is very necessary to study and improve the convolution processing process to make it have less computational complexity and be easy to implement. Summary of the Invention
[0005] To solve the above problems, the present disclosure provides a convolution processing method and apparatus.
[0006] Embodiments of the present disclosure provide a convolution processing device, the device comprising: a mask and an optical sensing component, wherein the mask is located between the target object and the optical sensing component, and is configured to receive the object light of the target object, the object light carrying the image information of the target object, the mask includes a mask pattern thereon, the mask pattern is determined based on a convolution layer for performing convolution processing on an image, the received object light is masked by the mask pattern to output a masked optical signal, and the masked optical signal corresponds to the convolution information obtained by performing convolution processing on the image information; the optical sensing component is configured to receive the masked optical signal, convert the optical signal into an electrical signal, and output the electrical signal.
[0007] According to an embodiment of the present disclosure, the convolution processing device further comprises: a light generating component, configured to generate an optical image signal carrying the image information of the target object, wherein the optical image signal carrying the image information of the target object is an incoherent optical signal, wherein the light generating component is a plurality of point light sources, and the optical image signal carrying the image information of the target object is a light signal generated by the plurality of point light sources; or the light generating component is a display, and the optical image signal carrying the image information of the target object is a multi-pixel image signal generated by the display.
[0008] Embodiments of the present disclosure provide a convolution processing device, wherein the distance d between two adjacent point light sources or two adjacent pixels L , the distance d between the light generating component and the mask LM , the distance d between the mask and the optical sensing component MS , and the size Δ of a single pixel on the optical sensing component have the following relationship: wherein, a single pixel on the optical sensing component is equivalent to a single pixel generated after the convolution calculation of the convolution layer.
[0009] Embodiments of the present disclosure provide a convolution processing device, wherein the spatial size of a single pixel on the optical sensing component is determined according to the relationship between the geometric blur during the transmission of the optical image signal and the diffraction blur generated by the optical image signal passing through the mask.
[0010] Embodiments of the present disclosure provide a convolution processing device, wherein the geometric blur d1 = Δ, the diffraction blur d2 = 2.44λd MS / Δ; the spatial size of a single pixel on the optical sensing component
[0011] Embodiments of the present disclosure provide a convolution processing device. The mask includes one or more mask regions, each mask region is independent of each other, and the mask patterns on each mask region are the same or different. The mask pattern is obtained by coating or etching on the mask.
[0012] Embodiments of the present disclosure provide a convolution processing device. The mask pattern of each mask region is composed of a plurality of mask holes, and the light transmission degrees of each mask hole are the same or different. Among them, the mask pattern of each mask region corresponds to a convolution layer, and the light transmission degree of each mask hole is determined by the parameters of its corresponding convolution layer.
[0013] According to an embodiment of the present disclosure, the convolution processing device further includes: a post-processing component, configured to receive the electrical signal output by the optical sensing component and generate a processing result of the image information based on the electrical signal.
[0014] Embodiments of the present disclosure provide a convolution processing method. The method includes: receiving object light of a target object, where the object light carries image information of the target object; performing optical masking processing on the image information through a mask to output a masked optical signal, and the masked optical signal corresponds to convolution information after performing convolution processing on the image information. Among them, the mask includes a mask pattern, and the mask pattern is determined based on the convolution layer for which convolution processing is to be performed on the image; converting the masked optical signal into an electrical signal; and generating a processing result of the image information based on the electrical signal.
[0015] According to an embodiment of the present disclosure, the convolution processing method further includes: generating an optical image signal carrying the image information of the target object through a light generating component, where the optical image signal carrying the image information of the target object is an incoherent optical signal, and the optical image signal carrying the image information of the target object is a light signal generated by a plurality of point light sources; or the optical image signal carrying the image information of the target object is a multi-pixel image signal generated by a display.
[0016] Embodiments of the present disclosure provide a convolution processing method: where the distance d between two adjacent point light sources or two adjacent pixels L , the distance d between the light generating component and the mask LM , the distance d between the mask and the optical sensing component MS , and the size Δ of a single pixel on the optical sensing component have the following relationship: Among them, a single pixel on the optical sensing component is equivalent to a single pixel generated after convolution calculation of the convolution layer.
[0017] Embodiments of the present disclosure provide a convolution processing method: Among them, the mask includes one or more mask regions, each mask region is independent of each other, and the mask patterns on each mask region are the same or different, and the mask pattern is obtained by coating or etching on the mask.
[0018] Embodiments of the present disclosure provide a convolution processing method: The mask pattern of each mask region is composed of a plurality of mask holes, and the light transmission degrees of each mask hole are the same or different. Among them, the mask pattern of each mask region corresponds to a convolution layer, and the light transmission degree of each mask hole is determined by the parameters of its corresponding convolution layer.
[0019] Embodiments of the present disclosure provide a convolution processing method and apparatus. According to the embodiments of the present disclosure, the convolution processing apparatus of the present disclosure includes: a mask and an optical sensing component. Among them, the mask is located between the target object and the optical sensing component and is used to receive the object light of the target object. The object light carries the image information of the target object. The mask includes a mask pattern on it. The mask pattern is determined based on the convolution layer for performing convolution processing on the image. The received object light is masked through the mask pattern to output a masked light signal. The masked light signal corresponds to the convolution information after performing convolution processing on the image information; the optical sensing component is used to receive the masked light signal, convert the light signal into an electrical signal, and output the electrical signal.
[0020] Through the convolution processing apparatus of the present disclosure, convolution calculation in the optical domain can be realized, thereby reducing the calculation amount of the post-processing component, reducing the performance requirements of the convolution neural network for the post-processing component, and reducing the power consumption of the post-processing component. Moreover, the present disclosure adopts a lensless convolution processing apparatus, effectively reducing the size, weight and cost of the convolution processing apparatus. Also, since the convolution process is realized in the optical domain, the result collected by the optical sensing component is no longer the actual optical image, but the convolution information after the actual image information is convolved. Therefore, the protection and encryption of the actual image information can be realized in this way.
[0021] In addition, the mask in the convolution processing apparatus provided by the present disclosure may include a plurality of mask regions, so as to complete parallel convolution operations optically. This way of mask regions can reduce the number of masks, save materials, effectively compress the volume of the parallel convolution processing system, and the parallel processing process of this convolution does not require more resources of the post-processing component, and can effectively improve the efficiency of convolution processing. At the same time, when light passes through the mask with different mask patterns, different convolution processing results may be obtained, and different features of the image information may be extracted. This multi-convolution processing method can improve the recognition accuracy of the image information and enhance the reliability of image processing or image recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] To more clearly illustrate the technical solutions of the embodiments of the present disclosure, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some exemplary embodiments of the present disclosure. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0023] Figure 1 is a schematic diagram showing a convolution processing device according to an embodiment of the present disclosure;
[0024] Figure 2 is a schematic flowchart showing a convolution processing method according to an embodiment of the present disclosure;
[0025] Figure 3 is a schematic diagram showing the propagation of light through a convolution processing device according to an embodiment of the present disclosure;
[0026] Figure 4A is a schematic diagram showing image information according to an embodiment of the present disclosure;
[0027] Figure 4B is a schematic diagram showing the result obtained after processing image information through a mask according to an embodiment of the present disclosure;
[0028] Figure 5 is a schematic diagram showing the relationship between the distances among a target object, a mask, and an optical sensing component according to an embodiment of the present disclosure; and
[0029] Figure 6 is a schematic diagram showing a mask with multiple mask regions according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0030] In order to make the objectives, technical solutions, and advantages of the present disclosure more obvious, the following will describe in detail exemplary embodiments according to the present disclosure with reference to the drawings. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all the embodiments of the present disclosure. It should be understood that the present disclosure is not limited by the exemplary embodiments described herein.
[0031] In addition, in this specification and the drawings, steps and elements having substantially the same or similar steps and elements are denoted by the same or similar reference numerals, and the repeated description of these steps and elements will be omitted.
[0032] In addition, in this specification and the accompanying drawings, elements are described in singular or plural forms according to embodiments. However, the singular and plural forms are appropriately selected for the proposed cases merely for convenience of explanation and are not intended to limit the present disclosure thereto. Therefore, the singular form may include the plural form, and the plural form may also include the singular form, unless the context clearly indicates otherwise.
[0033] In addition, in this specification and the accompanying drawings, the terms "first / second" are merely used to distinguish similar objects and do not represent a specific order for the objects. Understandably, "first / second" can be interchanged in a specific order or sequence when permitted, so that the embodiments of the present invention described herein can be implemented in an order other than that illustrated or described herein.
[0034] Image processing technology is a technology for processing image information by a computer. It mainly includes: image digitization, image enhancement and restoration, image data encoding, image segmentation, and image recognition, etc. Currently, image processing technology is mainly implemented based on a computer. An artificial neural network (ANN) is an algorithmic mathematical model that simulates the structure and behavior of a biological nervous system and performs distributed parallel information processing. The ANN adjusts the weight relationship between internal neurons to achieve the purpose of processing information. When processing an image, the computer takes the original image or an appropriately preprocessed image as the input signal of the neural network, and obtains the processed image signal or classification result at the output end of the neural network.
[0035] A convolutional neural network is a feedforward neural network that performs very well in image processing. The basic structure of a CNN consists of an input layer, a convolutional layer, a pooling layer (also called a sampling layer), a fully connected layer, and an output layer. Generally, several convolutional layers and pooling layers are taken, and the convolutional layer and the pooling layer are alternately arranged, that is, a convolutional layer is connected to a pooling layer, and then a convolutional layer is connected after the pooling layer, and so on. Each neuron in the output feature map of the convolutional layer is locally connected to its input, and the weighted sum of the local input is obtained by multiplying with the corresponding connection weights and then adding a bias value to obtain the input value of the neuron. This process is similar to the calculation process of convolution.
[0036] In summary, the solutions provided by the embodiments of the present disclosure involve fields such as image processing and neural networks. The embodiments of the present disclosure will be further described below with reference to the accompanying drawings.
[0037] Figure 1 It is a schematic diagram showing a convolutional processing device according to an embodiment of the present disclosure.
[0038] AsFigure 1 As shown in Figure 1 , the convolution processing device includes: a target object, an optical mask, and an optical sensing component.
[0039] Among them, the mask is located between the target object and the optical sensing component, and is used to receive the object light of the target object. The object light carries the image information of the target object. The mask includes a mask pattern, and the mask pattern is determined based on the convolutional layer for performing convolutional processing on the image. The received object light is masked through the mask pattern to output a masked optical signal, and the masked optical signal corresponds to the convolutional information after convolutional processing of the image information; the optical sensing component is used to receive the masked optical signal, convert the optical signal into an electrical signal, and output the electrical signal.
[0040] It should be understood that the image information described in this disclosure can directly come from the target object. For example, the target object can be a person or an object in reality, and it reflects light so that the target light carrying the image information of the target object enters the mask. The image information described in this disclosure can also come from a light generating component. For example: the light generating component can be a display for presenting an image, or a combination of a point light source or multiple point light sources. In any application scenario where light can be emitted and / or reflected to the mask, the convolution processing device described in this disclosure can be used to perform convolution processing operations. It should be understood that the light described in this disclosure can refer to various forms of visible or invisible light such as diffuse reflection light, monochromatic light, and composite light.
[0041] Optionally, in the case of using a light generating component to generate an optical image signal carrying the image information of the target object, the optical image signal carrying the image information of the target object is an incoherent optical signal. Among them, the light generating component is multiple point light sources, and the optical image signal carrying the image information of the target object is an optical signal generated by the multiple point light sources; or the light generating component is a display, and the optical image signal carrying the image information of the target object is a multi-pixel image signal generated by the display. The mask receives the optical image signal carrying the image information of the target object and processes the image information using the same method as described above for the target object.
[0042] A photomask generally includes a transparent substrate and a light-shielding film. Among them, common materials for the transparent substrate include: transparent glass (such as quartz glass, soda glass, low-expansion glass, etc.), transparent resin, etc. The light-shielding film usually includes a hard light-shielding film (such as chromium film, iron oxide, molybdenum silicide, etc.), latex, etc. The patterns for light shielding or light transmission on the photomask can be achieved through coating or etching processes. It should be understood that these are only some examples of the materials and processing techniques of the photomask, and all transparent or light-shielding materials that can be processed into a photomask, as well as all processes that can be used to process the photomask to form the required light-transmitting areas and light-shielding areas, can be applied to the photomask described in the present disclosure.
[0043] The optical sensing component can be a photosensitive chip, a photoelectric conversion device, etc. For example, common optical sensing components include CCD (Charge-Coupled Device) type optical sensing components and CMOS (Complementary Metal-Oxide-Semiconductor) type optical sensing components, etc. After the photosensitive pixel of the CCD receives light, the photosensitive element generates a corresponding current, and the magnitude of the current corresponds to the light intensity. The photosensitive element directly outputs an electrical signal in analog form. Each photosensitive element in the CMOS directly integrates an amplifier and analog-to-digital conversion logic. When the photosensitive diode receives light and generates an analog electrical signal, the electrical signal is first amplified by the amplifier in the photosensitive element, and then directly converted into a corresponding digital signal. Whether it is a CCD type optical sensing component or a CMOS type optical sensing component, the main purpose is to convert the collected optical signal into an electrical signal that can be processed by subsequent circuits or a computer. All devices that can convert optical signals into signals that can be utilized by processing devices can belong to the optical sensing components described in the present disclosure.
[0044] Optionally, if the processing result of the convolution needs to be further processed, the convolution processing device may further include a post-processing component. The post-processing component is used to receive the electrical signal output by the optical sensing component and generate a processing result for the image information based on the electrical signal. Optionally, the post-processing component described in the present disclosure can be any component that can perform calculations or processing on signals, such as a computer, a processor, a server, an integrated circuit, etc., or a combination of one or more of the above components.
[0045] Through Figure 1 It can be seen that the present disclosure adopts a lensless optical system, effectively reducing the size, weight, and cost of the system. Usually, the lens system requires high precision and has high requirements for usage conditions and the environment. The convolution processing device of the present disclosure can directly collect light without a lens. Compared with the traditional system that uses a lens to collect light, it is more convenient to operate and use and has a broader application prospect. Moreover, through the convolution processing device of the present disclosure, convolution calculation in the optical domain can be achieved, thereby reducing the calculation amount of the post-processing component, lowering the performance requirements of the convolution neural network for the post-processing component, and reducing the power consumption of the post-processing component.
[0046] Figure 2 is a schematic flowchart showing a convolution processing method according to an embodiment of the present disclosure.
[0047] As Figure 2 shown, in step 201, object light of a target object is received, and the object light carries image information of the target object.
[0048] Optionally, an optical image signal carrying the image information of the target object may be generated by an optical generation component, wherein the optical image signal carrying the image information of the target object is an incoherent light signal, and the optical image signal carrying the image information of the target object is a light signal generated by a plurality of point light sources; or the optical image signal carrying the image information of the target object is a multi-pixel image signal generated by a display.
[0049] In step 202, optical masking processing is performed on the image information through a mask plate to output a masked light signal, and the masked light signal corresponds to convolution information after performing convolution processing on the image information, wherein the mask plate includes a mask pattern, and the mask pattern is determined based on a convolution layer for which convolution processing is to be performed on the image.
[0050] Optionally, the distance d between two adjacent point light sources or two adjacent pixels L , the distance d between the optical generation component and the mask plate LM , the distance d between the mask plate and the optical sensing component MS , and the size Δ of a single pixel on the optical sensing component have the following relationship: wherein a single pixel on the optical sensing component is equivalent to a single pixel generated after convolution calculation of the convolution layer.
[0051] Optionally, the mask plate includes one or more mask regions, each mask region is independent of each other, and the mask patterns on each mask region are the same or different, and the mask pattern is obtained by coating or etching on the mask plate.
[0052] Optionally, the mask pattern of each mask region is composed of a plurality of mask holes, and the light transmission degrees of each mask hole are the same or different, wherein the mask pattern of each mask region corresponds to a convolution layer, and the light transmission degree of each mask hole is determined by the parameters of its corresponding convolution layer.
[0053] In step 203, the masked light signal is converted into an electrical signal.
[0054] Optionally, the process in step 203 can be implemented in various ways such as a photosensitive chip, a photoelectric conversion device, etc. The electrical signal formed after the conversion of the optical signal can be either an analog signal or a digital signal.
[0055] Through the convolution processing method of the present disclosure, the convolution process can be implemented in the optical domain. The result collected by the optical sensing component is no longer the actual optical image, but the convolution information after the actual image information is convolved. Therefore, the protection and encryption of the actual image information can be achieved through this method.
[0056] In step 204, based on the electrical signal, a processing result of the image information is generated.
[0057] It should be understood that the processing result of the image information generated here is the result of convolving the optical image signal. Optionally, if other calculations and processing are required subsequently, the post-processing component can be used to receive the electrical signal output by the optical sensing component and continue to process the image information based on the electrical signal.
[0058] Figure 3 FIG. is a schematic diagram showing the propagation of light through a convolution processing device according to an embodiment of the present disclosure.
[0059] According to an embodiment of the present disclosure, the image information can be propagated in the form of multiple point light sources. Figure 3 FIG. exemplarily shows five point light sources, and the light intensities corresponding to each point are represented by A, B, C, D, and E respectively. The matrix of the light source distribution is denoted as X. Assume that there is a mask area on the mask plate, and this mask area has 3 mask holes. Here, M1, M2, and M3 are used to represent the light transmission degrees of the mask holes on the mask plate, and its matrix is denoted as For the optical sensing component in the figure, the light intensity distributions received by the middle three points are S1, S2, and S3, and its matrix is denoted as Y. Then it can be expressed by the following formula:
[0060] Y = (S1 S2 S3) T
[0061]
[0062] X = (A B C D E) T
[0063] The light emitted by a point light source passes through a mask and is received and detected at an optical sensing component, and is converted into an electrical signal that is easy to be processed by a computer through photoelectric conversion. The light intensity at each point on the optical sensing component is the superposition of all the light intensities received after being processed by the mask. The mask pattern of each mask area is composed of multiple mask holes, and the light transmission degree of each mask hole is the same or different. For the mask holes on the mask, the light transmission degree corresponds to the parameters of the corresponding convolutional layer, that is, the convolution kernel. For example, M1, M2, and M3 can be a binary distribution of 0 and 1, where 0 represents non-light transmission and 1 represents full light transmission. Similarly, M1, M2, and M3 can be a binary distribution of 0 and 1, where 1 represents non-light transmission and 0 represents full light transmission. Optionally, in addition to digital values, M1, M2, and M3 can also be analog values. For example, 0 represents non-light transmission, 1 represents full light transmission, and 0.5 represents half light transmission...
[0064] Assume that except for the light transmission degrees of points 1, 2, and 3 on the mask corresponding to M1, M2, and M3 respectively, the rest are non-light transmissive (i.e., represented by 0), and without considering noise and attenuation of light during transmission, the received light S1, S2, and S3 on the optical sensing component can be calculated by the following formula:
[0065]
[0066] It can be seen that the above formula is in the form of multiplication and addition, and its corresponding matrix form is:
[0067]
[0068] From the above formula, it can be seen that M1, M2, and M3 have an effect similar to sliding forward in turn during convolution calculation. Rearranging the above formula into convolution form, we have:
[0069]
[0070] That is, the light signal (Y) detected on the optical sensing component after mask processing is the result of the light signal (X) from the point light source being processed by the mask pattern corresponding to the convolution kernel on the mask.
[0071] It should be understood that here multiple point light sources are taken as examples rather than limitations. Similarly, a multi-pixel image generated by a display can also be regarded as multiple point light sources for the same processing, where each pixel can be equivalent to a point light source; or the target object can be equivalent to an object composed of multiple point light sources, and the object light carrying the image information of the target object is equivalent to the light signals generated by multiple point light sources. In this case, the convolution processing process is similar to the above, so it will not be elaborated here.
[0072] Figure 4A is a schematic diagram showing the image information according to an embodiment of the present disclosure.
[0073] In Figure 4A there are two light sources, namely the first light source located at (1, 1) in the plane coordinate system and the second light source located at (4, 4) in the plane coordinate system. The horizontal and vertical spacings between the centers of the first light source and the second light source are both 3 unit distances.
[0074] For Figure 4A the image information in, the result obtained after being processed by the mask is as shown in Figure 4B shown.
[0075] Assume that the mask holes on the mask are distributed in a binary 01 pattern, where 0 represents non - light - transmissive and 1 represents fully light - transmissive. Then the convolution kernel corresponding to the mask holes on the mask is:
[0076] Since the horizontal and vertical spacings between the centers of the first light source and the second light source are both 3, in an ideal case, the horizontal and vertical spacings between the center of the first convolution result after convolution of the first light source and the center of the second convolution result after convolution of the second light source are both 3Δ, where the size of a single pixel on the optical sensing component is Δ, also known as the feature size.
[0077] Figure 5 is a schematic diagram showing the relationship between the distances among the target object, the mask, and the optical sensing component according to an embodiment of the present disclosure.
[0078] Assume that the distance between two adjacent point light sources or two adjacent pixels is d L (unit distance), the distance between the light - generating component and the mask is d LM , the distance between the mask and the optical sensing component is d MS , and the spatial size of the convolution kernel equivalent to a single pixel is Δ (i.e., the feature size). Since light travels in a straight line, according to the triangle similarity theory (i.e., ΔABC~ΔFEC), the corresponding light source spacing can be calculated.
[0079] For an incoherent light system, the optical system is affected by diffraction blur and geometric blur.
[0080] Among them, the geometric blur d1 during the light transmission process and the diffraction blur d2 generated by the light passing through the mask can be calculated by the following formulas:
[0081] d1 = Δ
[0082] d2 = 2.44λd MS / Δ
[0083] where λ is the wavelength of light, d MSis the distance between the mask and the optical sensing component.
[0084] For incoherent optical systems, the diffraction phenomenon is not obvious. The influence of geometric blur in the transmission process of optical image signals is greater than or equal to the diffraction blur caused by the optical image signals passing through the mask, that is, d1≥d2. After sorting, the characteristic size can be obtained. That is, the spatial size Δ of a single pixel on the optical sensing component can be determined based on the relationship between the geometric blur d1 during the transmission of the optical image signal and the diffraction blur d2 generated by the optical image signal passing through the mask.
[0085] At the same time, according to the formula It can be seen that the larger the characteristic size, the larger the volume of the entire optical domain processing system. Considering that the entire system should be as compact as possible, the value of Δ should be as small as feasible.
[0086] For example, for d MS =1.02mm, and the wavelength of visible light is λ=550nm, according to the formula Calculation shows that Δ≥37μm. In practical applications, in order to make calculation and processing easier, the characteristic size Δ can also be rounded up, that is, the characteristic size is taken as 40μm.
[0087] Considering that the entire system should be as small as possible, the distance d between the mask and the optical sensing component MS The distance between the mask and the optical sensing component can be measured after the mask and the optical sensing component are placed as close as possible. Alternatively, the distance measurement can be achieved using various high-precision distance measurement methods such as high-precision rulers, electrical distance measurement, and optical distance measurement.
[0088] Therefore, under a fixed optical domain convolution system (Δ, d MS has been determined or measured according to the above formula), according to the size of the actual target object or light generating component and the specific application, the formula Determine d LM and the specific location distribution of the target object or light generating components.
[0089] Since the convolution processing device may have installation errors during installation, Δ and d MS The calculation and measurement may also introduce errors. In order to ensure the accuracy of the convolution processing device, the convolution processing device can be fine-tuned according to the convolution result.
[0090] For example, for Figure 4A and Figure 4BIn the example, for the first light source located at (1, 1) in the plane coordinate system and the second light source located at (4, 4) in the plane coordinate system, theoretically, the horizontal spacing and the vertical spacing between the center of the first convolution result after the first light source passes through convolution and the center of the second convolution result after the second light source passes through convolution should both be 3Δ. If due to system errors, it is found that the horizontal spacing and the vertical spacing between the center of the first convolution result and the center of the second convolution result do not conform to 3Δ after convolution processing, then the system is finely adjusted so that the final actual result conforms to the theoretical result, thereby verifying that the processing accuracy of the convolution processing device can meet the usage requirements.
[0091] It should be understood that the system fine-tuning process here is to verify whether there is an error between the processing result and the theoretical value when the system is initially installed. If there is an error, the error is reduced through fine-tuning. For a system that has been verified through debugging, the fine-tuning process here is not necessary.
[0092] Figure 6 is a schematic diagram showing a mask plate having a plurality of mask regions according to an embodiment of the present disclosure.
[0093] In Figure 6 the shown mask plate includes one or more mask regions, each mask region is independent of each other, and the mask patterns on each mask region are the same or different. Therefore Figure 6 the shown mask plate can implement parallel calculation of multiple convolutions. Among them, the parameters of multiple convolutions can be the same or different. They can either simultaneously implement convolution processing on the same optical signal or simultaneously implement convolution processing on different optical signals. This mask pattern arrangement method can reduce the number of mask plates, save materials, and effectively compress the volume of the parallel convolution processing system.
[0094] Since multiple convolutions can be simultaneously processed in parallel in the optical domain, the parallel processing process of this convolution does not need to occupy more resources of post-processing components. Therefore, it can effectively improve the efficiency of parallel convolution processing. Moreover, when light passes through the mask plate engraved with different mask patterns, different convolution processing results may be obtained, and different features of image information may be extracted. This multi-convolution processing method can improve the recognition accuracy of image information and enhance the reliability of image processing or image recognition.
[0095] In summary, the embodiments of the present disclosure provide a convolution processing method and apparatus. According to the embodiments of the present disclosure, the convolution processing apparatus of the present disclosure includes: a mask and an optical sensing component. Among them, the mask is located between the target object and the optical sensing component, and is used to receive the object light of the target object. The object light carries the image information of the target object. The mask includes a mask pattern, and the mask pattern is determined based on the convolutional layer for performing convolution processing on the image. The received object light is masked through the mask pattern to output a masked optical signal, and the masked optical signal corresponds to the convolution information after performing convolution processing on the image information; the optical sensing component is used to receive the masked optical signal, convert the optical signal into an electrical signal, and output the electrical signal.
[0096] Through the convolution processing apparatus of the present disclosure, convolution calculation in the optical domain can be realized, thereby reducing the calculation amount of the post-processing component, reducing the performance requirements of the convolution neural network for the post-processing component, and reducing the power consumption of the post-processing component. Moreover, the present disclosure adopts a lensless convolution processing apparatus, effectively reducing the size, weight, and cost of the convolution processing apparatus. Furthermore, since the convolution process is realized in the optical domain, the result collected by the optical sensing component is no longer the actual optical image, but the convolution information after the actual image information is subjected to convolution processing. Therefore, in this way, the protection and encryption of the actual image information can be realized.
[0097] The present disclosure uses specific terms to describe the embodiments of the present disclosure. Such as "the first / second embodiment", "an embodiment", and / or "some embodiments" mean a certain feature, structure, or characteristic related to at least one embodiment of the present disclosure. Therefore, it should be emphasized and noted that the "an embodiment" or "one embodiment" or "an alternative embodiment" mentioned twice or more at different positions in this specification does not necessarily refer to the same embodiment. In addition, certain features, structures, or characteristics in one or more embodiments of the present disclosure can be appropriately combined.
[0098] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by those of ordinary skill in the art to which this invention belongs. It should also be understood that terms such as those defined in a commonly used dictionary should be interpreted as having a meaning consistent with their meaning in the context of the relevant art, and should not be interpreted in an idealized or overly formal sense unless expressly so defined herein.
[0099] The foregoing is a description of the invention and should not be construed as limiting thereof. Although several exemplary embodiments of the invention have been described, those skilled in the art will readily appreciate that many modifications can be made to the exemplary embodiments without departing from the novel teachings and advantages of the invention. Accordingly, all such modifications are intended to be included within the scope of the invention as defined by the claims. It should be understood that the foregoing is a description of the invention and should not be considered limited to the specific embodiments disclosed, and modifications to the disclosed embodiments as well as other embodiments are intended to be included within the scope of the appended claims. The invention is defined by the claims and their equivalents.
Claims
1. A convolution processing device, comprising: A mask plate and an optical sensing component, wherein, the mask plate is located between the target object and the optical sensing component, and is used to receive the object light of the target object. The object light carries the image information of the target object. The mask plate includes a mask pattern, which is determined based on the convolutional layer for performing convolutional processing on the image. The received object light is masked through the mask pattern to output a masked optical signal, and the masked optical signal corresponds to the convolutional information after performing convolutional processing on the image information; the optical sensing component is used to receive the masked optical signal, convert the optical signal into an electrical signal, and output the electrical signal; wherein, the mask plate includes one or more mask regions, each mask region is independent of each other, and the mask patterns on each mask region are the same or different. The mask pattern of each mask region is composed of a plurality of mask holes, and the light transmission degrees of each mask hole are the same or different, wherein, the mask pattern of each mask region corresponds to a convolutional layer, and the light transmission degree of each mask hole is determined by the parameters of the corresponding convolutional layer.
2. The convolutional processing device according to claim 1, further comprising: a light generating component, configured to generate an optical image signal carrying the image information of the target object, wherein the optical image signal carrying the image information of the target object is an incoherent optical signal, wherein, the light generating component is a plurality of point light sources, and the optical image signal carrying the image information of the target object is a light signal generated by the plurality of point light sources; or the light generating component is a display, and the optical image signal carrying the image information of the target object is a multi-pixel image signal generated by the display.
3. The convolutional processing device according to claim 2, wherein, The distance d between two adjacent point light sources or between two adjacent pixels L , the distance d between the light generating component and the mask LM , the distance d between the mask and the optical sensing component MS , and the size Δ of a single pixel on the optical sensing component have the following relationship: a single pixel on the optical sensing component is equivalent to a single pixel generated after the convolutional calculation of the convolutional layer.
4. The convolution processing device according to claim 3, wherein, The spatial dimension of a single pixel on the optical sensing component is determined according to the relationship between the geometric blur during the transmission of the optical image signal and the diffraction blur generated by the optical image signal passing through the mask plate.
5. The convolutional processing device according to claim 4, wherein, The geometric blur d1 = Δ, and the diffraction blur d2 = 2.44λd MS / Δ; The spatial dimension of a single pixel on the optical sensing component 6. The convolutional processing device according to claim 1, wherein, the mask pattern is obtained by coating or etching on the mask plate.
7. The convolutional processing device according to claim 1, further comprising: a post-processing component, configured to receive the electrical signal output by the optical sensing component and generate a processing result of the image information based on the electrical signal.
8. A convolutional processing method, comprising: receiving the object light of the target object, where the object light carries the image information of the target object; performing optical masking processing on the image information through a mask plate to output a masked optical signal, where the masked optical signal corresponds to the convolutional information after performing convolutional processing on the image information. The mask plate includes a mask pattern, and the mask pattern is determined based on the convolutional layer for performing convolutional processing on the image; Convert the optical signal after mask processing into an electrical signal; and Generate a processing result of the image information based on the electrical signal; wherein, the mask plate includes one or more mask regions, each mask region is independent of each other, and the mask patterns on each mask region are the same or different, the mask pattern of each mask region is composed of a plurality of mask holes, and the light transmission degrees of each mask hole are the same or different, wherein, the mask pattern of each mask region corresponds to a convolutional layer, and the light transmission degree of each mask hole is determined by the parameters of the corresponding convolutional layer.
9. The convolutional processing method according to claim 8, further comprising: Generate an optical image signal carrying the image information of the target object through a light generating component, wherein the optical image signal carrying the image information of the target object is an incoherent optical signal, wherein, the optical image signal carrying the image information of the target object is a light signal generated by a plurality of point light sources; or The optical image signal carrying the image information of the target object is a multi-pixel image signal generated by a display.
10. The convolutional processing method according to claim 9, wherein, The distance d between two adjacent point light sources or between two adjacent pixels L , the distance d between the light generating component and the mask LM , the distance d between the mask and the optical sensing component MS , and the size Δ of a single pixel on the optical sensing component have the following relationship: wherein, a single pixel on the optical sensing component is equivalent to a single pixel generated after the convolutional calculation of the convolutional layer.
11. The convolutional processing method according to claim 8, wherein, The mask pattern is obtained by coating or etching on the mask plate.
Citation Information
Patent Citations
Image processing method and device, computer device and storage medium
CN110675385A
Image processing method and device, equipment and storage medium
CN111754439A