LG light beam identification method based on physical feature information fusion neural network
Through the method of fusion neural network based on physical feature information, the high error and low efficiency problems of traditional LG beam recognition technology are solved, and high-precision and high-speed beam recognition are achieved, information capacity and model stability are improved, and the requirements for spatial light modulators are reduced.
Patent Information
- Application Number
- CN202510740587.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-05
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-06-05
AI Technical Summary
Traditional LG beam feature information extraction technology has high artificial error and low efficiency, and deep neural networks are limited by resolution and difficult to identify information between angular momentum modes of adjacent orbits with high accuracy due to their resolution.
Using a method of fusion neural network based on physical feature information, by observing the special characteristics of fractional-order LG beams in diffraction images, using improved neural network structure and loss function, we optimize the design of neural network training with less data to identify the beam's waist size and topological load changes.
High-precision and high-speed LG beam recognition is realized, reducing the performance requirements for spatial light modulators, improving the recognition difficulty and information capacity, and enhancing the stability and calculation speed of the model.
Smart Images

Figure CN120298809A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of cross - application of artificial intelligence and physical models, and specifically to an LG beam recognition method based on physical feature information fusion neural network. Background Technique
[0002] Traditional LG beam feature information extraction technologies have high human errors and low efficiency. The recognition accuracy largely depends on the features with obvious diffraction spot intensity. This method has great defects whether in experimental operations or instrument measurements. The rise of deep neural networks has brought hope for overcoming these deficiencies. Deep neural networks can not only accurately recognize integer - type vortex beams, but also recognize fractional - type vortex beams with very small intervals; however, although deep neural networks show unique advantages in vortex beam mode recognition, due to the limitation of the resolution of optical communication systems, the interval between adjacent orbital angular momentum modes cannot be infinitely small. How to accurately recognize a large amount of information carried by light at the receiving end has become an important key to improving optical communication capacity technology. Summary of the Invention
[0003] The purpose of the present invention is to provide an LG beam recognition method based on physical feature information fusion neural network for the deficiencies of the existing technology. Different from the traditional one - to - one recognition forward training and recognizing subtle topological charge changes, the present invention discovers through observation that fractional - order LG beams have special distinguishable characteristics in diffraction images, thus abandoning the previous dependence on topological charge recognition and turning to recognize the waist size. By using physical fusion deep neural networks, a robust neural network architecture with few - data training and a certain degree of interpretability is optimized and designed, so as to achieve high - precision and high - speed intelligent recognition of vortex beams.
[0004] To achieve the above purpose, the present invention adopts the following technical solutions.
[0005] An LG beam recognition method based on physical feature information fusion neural network includes the following steps: Step S1: Modulate and expand the incident beam through a polarization filter and then irradiate it on a spatial light modulator to generate Laguerre - Gaussian (LG) beam; Step S2: Use a lens to focus the LG beam generated in step S1 to produce a Fourier transform, and collect the Fraunhofer diffraction image of the LG beam through a CCD camera; Step S3: Input the collected Fraunhofer diffraction image of the LG beam into an improved neural network structure for processing, and through the optimized loss function, achieve accurate recognition of multi - physical feature information in the LG beam.
[0006] Specifically, the polarization filter described in step S1 includes a combination of a rotatable linear polarizer and a half-wave plate, which is used to control the light intensity and polarization direction of the LG beam; the spatial light modulator is built-in with a composite modulation function of a computer-generated hologram and a spiral phase plate, which is used to synchronously regulate the wavefront phase and intensity distribution of the LG beam.
[0007] Specifically, the lens in step S2 includes a pair of double convex lenses with a focal length ratio of 1:1 and an additional double convex lens. The pair of double convex lenses form a 4f system, that is, the initial beam radius of the LG beam is retained. Moreover, the first-order diffraction spot is filtered out at the focal plane of the first double convex lens to extract a pure LG beam. The subsequent two double convex lenses are focused at the focal plane to generate a Fourier transform to form a Fraunhofer diffraction image containing topological charge number information; the pixel size of the CCD camera is less than 1 / 5 of the diffraction ring spacing of the Fraunhofer diffraction image.
[0008] Specifically, in step S2, the lens is used to focus the LG beam generated in step S1 to generate a Fourier transform. The equation of the LG beam at Z = 0 is expressed as: ; In the above formula, represents the LG beam at Z = 0, is the polar coordinate on the incident surface, z is the propagation distance, r is the radial coordinate in the polar coordinate; is the generalized Laguerre polynomial, is the radial exponent, is the topological charge number of the LG beam; is the LG beam radius; is the radial intensity distribution of the LG beam; is the imaginary unit; is the azimuth angle in the polar coordinate; After the LG beam passes through the lens and forms a Fraunhofer diffraction spot at the focal plane, the expression of the Fraunhofer diffraction spot is abbreviated as the Fourier transform form of the incident light field as: ; In the above formula, is the Fourier transform of the incident light field; is the Fourier transform representation; respectively correspond to the frequency domain coordinates.
[0009] Specifically, the improved neural network structure in step S3 includes a convolutional layer, a pooling layer, a feature information fusion layer, and a fully connected layer; The convolutional layer is used to extract local features through a convolutional kernel, expand the receptive field by inserting holes in the convolutional kernel, and enhance the ability to extract global and local image information; The pooling layer includes a max-pooling layer and an adaptive pooling layer. The max-pooling layer is used to reduce the spatial size of the features while retaining the important information of the features, and the adaptive pooling layer is used to uniformly pool the input of any size into a fixed size to adapt to different input sizes; The physical feature information fusion layer is used to fuse the input matrix information. The input feature matrix is divided into three branches, and each branch undergoes a non-linear transformation through matrix dot product and Sigmoid function. Then, the features of each branch are fused at different nodes, and the features of multiple branches are combined through matrix addition operation. The fused features then pass through the ReLU activation function to output a more tightly fused and more stable feature matrix, enhancing the model's ability to recognize and express input information; The fully connected layer is used to map the extracted features to the output space to obtain the final prediction result.
[0010] Further, the convolutional layer includes a dilated convolutional layer and a standard convolutional layer. The dilated convolutional layer expands the receptive field of the convolutional kernel through a preset dilation rate, enabling the dilated convolutional layer to capture periodic features spanning a large physical scale in the input data while avoiding introducing additional parameters; the standard convolutional layer finely extracts the phase gradient between adjacent pixels and the detailed features of the singularity of the vortex center using the receptive field of a dense convolutional kernel; the dilated convolutional layer and the standard convolutional layer independently process the input through a parallel branch structure, and finally fuse the two feature maps through channel concatenation, thereby enhancing the ability to extract global and local image information.
[0011] Specifically, the improved neural network structure in step S3 enhances the robustness of the neural network in complex multi-task scenarios by optimizing the loss function. The design process of the loss function is as follows: First, calculate the cross-entropy loss k in the -th task scenario: ; In the above formula, C is the number of classes; represents the probability of the true label of sample x on the n -th class; is the predicted probability of sample x on the n -th class; Then, introduce the variance parameter to dynamically adjust the loss weights of each task, and divide the cross-entropy loss in each task scenario by the corresponding variance parameter , and then introduce the logarithmic term of the variance parameter to obtain the final loss function : ; In the above formula, is the cross-entropy loss in the k -th task scenario; is the learnable uncertainty parameter of the k -th task; The said logarithmic term is used to prevent gradient explosion. If the variance parameter is too large, will be very small. Introducing the logarithmic term of the variance parameter can avoid this kind of gradient explosion. Specifically: Usually, the program initially sets = 0, that is, = 1. If is too large (the task is difficult) → the gradient is negative → increase → exponentially increases → reduce the weight of this task ; if is too small (the task is easy) → the gradient is positive → decrease → exponentially decreases → increase the weight of this task . Compared with the prior art, the present invention has the following beneficial effects: 1. Based on the special distinguishable characteristics of the fractional-order LG beam in the diffraction image, the method of the present invention makes three outputs for the recognition of three physical characteristic parameters, outputs at different stages of the neural network according to the display difficulty of the spot characteristics, and combines the fusion technology of the deep and shallow layers in the neural network structure to achieve the fusion of physical information characteristics, making the prediction result of the neural network more stable and accurate; 2. Through the innovative binary phase encoding technology, the method of the present invention can efficiently generate LG beams with specific orbital angular momentum only using two phase states of 0 and π. Compared with the traditional multi-order phase modulation scheme, the performance requirements for the spatial light modulator SLM are greatly reduced, enabling ordinary SLM devices to also achieve high-quality LG beam output; 3. In the calculation process of the method of the present invention, the hologram is directly generated by an analytical expression, avoiding complex iterative optimization algorithms, and the calculation speed is increased by more than several times, which is especially suitable for experimental scenarios that require real-time dynamic regulation; at the same time, the binary phase structure has extremely strong anti-interference ability and higher tolerance to the calibration error of the spatial light modulator SLM; 4. The method of the present invention is different from the traditional method for identifying subtle topological charge changes. By simultaneously identifying three parameters, not only does it reduce the difficulty of task identification, but it also exponentially expands the information capacity, reducing the training volume of the neural network compared to the traditional single topological charge identification method. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Figure 1 is a flowchart of a method for identifying LG beams based on a neural network for fusing physical feature information according to the present invention; Figure 2 is a schematic diagram of the principle for collecting the Fraunhofer diffraction image of an LG beam according to the present invention; Figure 3 is a schematic diagram of the neural network structure in an embodiment of the present invention Figure 4 is a schematic diagram of the structure of the physical feature information fusion layer in an embodiment of the present invention; Figure 5 is in an embodiment of the present invention schematic diagram of the confusion matrix of the test set; Figure 6 is in an embodiment of the present invention schematic diagram of the confusion matrix of the test set; Figure 7 is in an embodiment of the present invention schematic diagram of the confusion matrix of the test set; Figure 8 is a spatial scatter plot of the recognition accuracy of the Fraunhofer diffraction image of an LG beam in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0013] To facilitate understanding and implementation of the present invention by those of ordinary skill in the art, the following details each step of the method proposed by the present invention. It should be understood that these embodiments are only for illustrating the present invention and not for limiting the scope of the present invention. In addition, it should be understood that after reading the content taught by the present invention, those skilled in the art can make various changes or modifications to the present invention, and these equivalent forms also fall within the scope defined by the appended claims of this application.
[0014] Embodiment As Figure 1 shown, the present invention discloses a method for identifying LG beams based on a neural network for fusing physical feature information, including the following steps: Step S1: Modulate and expand the incident beam through a polarization filter and then irradiate it on a spatial light modulator to generate a Laguerre-Gaussian (LG) beam; Step S2: Use a lens to focus the LG beam generated in step S1 to produce a Fourier transform, and collect the Fraunhofer diffraction image of the LG beam through a CCD camera; Step S3: Input the collected Fraunhofer diffraction images of the LG beam into the improved neural network structure for processing, and accurately identify the multi-physical feature information in the LG beam through the optimized loss function.
[0015] Specifically, as Figure 2 shown, the polarization filter described in step S1 includes a combination of a rotatable linear polarizer and a half-wave plate, which is used to control the light intensity and polarization direction of the LG beam; the spatial light modulator is built-in with a composite modulation function of a computer-generated hologram and a spiral phase plate, which is used to synchronously control the wavefront phase and intensity distribution of the LG beam.
[0016] Furthermore, the lens in step S2 includes a pair of double convex lenses with a focal length ratio of 1:1 and an additional double convex lens. The pair of double convex lenses form a 4f system, that is, the initial beam radius of the LG beam is retained , and at the focal plane of the first double convex lens, the first-order diffraction spot is filtered out to extract a pure LG beam. The subsequent two double convex lenses are focused at the focal plane to generate a Fourier transform to form a Fraunhofer diffraction image containing topological charge number information; the pixel size of the CCD camera is less than 1 / 5 of the diffraction ring spacing of the Fraunhofer diffraction image.
[0017] Specifically, in step S2, the lens is used to focus the LG beam generated in step S1 to generate a Fourier transform. The equation of the LG beam at Z = 0 is expressed as: ; In the above formula, represents the LG beam at Z = 0, is the polar coordinate on the incident surface, z is the propagation distance, r is the radial coordinate in the polar coordinate; is the generalized Laguerre polynomial, is the radial exponent, is the topological charge number of the LG beam; is the LG beam radius; is the radial intensity distribution of the LG beam; is the imaginary unit; is the azimuth angle in the polar coordinate; After the LG beam passes through the lens and forms a Fraunhofer diffraction spot at the focal plane, the expression of the Fraunhofer diffraction spot is abbreviated as the Fourier transform form of the incident light field as: ; In the above formula, is the Fourier transform of the incident light field; is the Fourier transform representation; respectively correspond to Frequency domain coordinates.
[0018] Specifically, the improved neural network structure described in step S3, as Figure 3 shown, the improved neural network structure includes a convolutional layer, a pooling layer, a physical feature information fusion layer, and a fully connected layer; The convolutional layer is used to extract local features through a convolutional kernel, expand the receptive field by inserting holes in the convolutional kernel, and enhance the ability to extract global and local information of the image; The pooling layer includes a max pooling layer and an adaptive pooling layer. The max pooling layer is used to reduce the spatial dimension of the features while retaining important feature information, and the adaptive pooling layer is used to uniformly pool inputs of any size into a fixed size to adapt to different input sizes; As Figure 4 shown, the physical feature information fusion layer is used to fuse the input matrix information. The input feature matrix is divided into three branches. Each branch undergoes a non-linear transformation through matrix dot product and the Sigmoid function, and then the features of each branch are fused at different nodes. The features of multiple branches are combined through matrix addition operation. The fused features are then passed through the ReLU activation function to output a more tightly fused and more stable feature matrix, enhancing the model's ability to recognize and express input information; The fully connected layer is used to map the extracted features to the output space to obtain the final prediction result.
[0019] Furthermore, the convolutional layer includes a dilated convolutional layer and a standard convolutional layer. The dilated convolutional layer expands the receptive field of the convolutional kernel through a preset dilation rate, enabling the dilated convolutional layer to capture periodic features spanning a large physical scale in the input data while avoiding introducing additional parameters; the standard convolutional layer finely extracts the phase gradient between adjacent pixels and the detailed features of the singularity of the vortex center using the receptive field of a dense convolutional kernel; the dilated convolutional layer and the standard convolutional layer independently process the input through a parallel branch structure, and finally fuse the two feature maps through channel concatenation, thereby enhancing the ability to extract global and local information of the image.
[0020] Specifically, the improved neural network structure in step S3 improves the robustness of the neural network in complex multi-task scenarios by optimizing the loss function. The design process of the loss function is as follows: First, calculate the cross-entropy loss k in the th task scenario: ; In the above formula, C is the number of classes; represents the true label of sample x in the nProbability on each category; is the sample x At the n predicted probability on each category; Then introduce the variance parameter Dynamically adjust the loss weights of each task, divide the cross-entropy loss in each task scenario by the corresponding variance parameter and then introduce the logarithmic term of the variance parameter to obtain the final loss function : ; In the above formula, is the cross-entropy loss in the k th task scenario; is the learnable uncertainty parameter of the k th task.
[0021] The technical effects of the method of the present invention will be further described below through a specific example.
[0022] The dataset used in this example is the Matlab simulation dataset; the deep neural network training environment is CPU 12thGen Intel(R) Core(TM) i7-12700KF 3.60 GHz, 32G RAM, GPU NVIDIA Geforce RTX3080, and the program is implemented using Pytorch in the python3.9 and window11 22000.1696 operating systems.
[0023] Refer to the neural network structure and the physical feature information fusion layer structure shown in Figure 3 and Figure 4 . Input the Fraunhofer diffraction image of the LG beam into the neural network for training. The neural network is divided into a deep layer and a shallow layer. The output channels of the shallow layer are and parameters with obvious feature information. The input of the deep layer is the combination of the output of the and channels of the shallow layer and the output of the original network. Finally, the three-channel information is used for the recognition of three parameters through the fully connected layer.
[0024] As shown in Figure 5 , Figure 6 , Figure 7 are the schematic diagrams of the confusion matrices of the three parameters P、 、L test sets respectively. From the analysis of the confusion matrices, it can be concluded that the neural network model proposed by the present invention is in P label, label and LIt performs excellently in the classification task of labels; the high values on the diagonal of the confusion matrix indicate that the model can accurately identify most samples, while the low values on the off-diagonal indicate fewer misclassification cases, showing that the method of the present invention has high accuracy and excellent recognition ability in the image classification task.
[0025] As Figure 8 shown is the spatial scatter plot of the recognition accuracy of the Fraunhofer diffraction image of the LG beam. Figure 8 In it, the red points represent the misidentified data, and the green points represent the data with all three parameters correctly identified. The spatial scatter plot further verifies that the predicted results and the actual values between the labels are highly consistent, and the number of misclassified points is extremely small. And Figure 8 it can be clearly seen the uniformity of the test set selected by the method of the present invention, indicating that the method of the present invention has good stability and accuracy in recognizing the physical feature information in the LG beam.
[0026] The above are only the preferred embodiments of the present invention, and are not limitations on the present invention in other forms. Any person skilled in the art may use the technical content disclosed above to make changes or modifications into equivalent embodiments with equivalent changes. However, any simple modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of the present invention without departing from the technical solution content of the present invention still fall within the protection scope of the technical solution of the present invention.
Claims
1. An LG beam recognition method based on the fusion of physical feature information and neural network, characterized in that It includes the following steps: Step S1: After modulating and expanding the incident light beam through a polarization filter, irradiate it on a spatial light modulator to generate a Laguerre-Gaussian (LG) beam; Step S2: Use a lens to focus the LG beam generated in Step S1 to produce a Fourier transform, and collect the Fraunhofer diffraction image of the LG beam through a CCD camera; Step S3: Input the collected Fraunhofer diffraction image of the LG beam into an improved neural network structure for processing, and through an optimized loss function, achieve accurate identification of multi-physical feature information in the LG beam.
2. The LG beam recognition method based on the physical feature information fusion neural network according to claim 1, wherein In Step S1, the polarization filter includes a combination of a rotatable linear polarizer and a half-wave plate, which is used to control the light intensity and polarization direction of the LG beam; the spatial light modulator is built-in with a composite modulation function of a computer-generated hologram and a spiral phase plate, which is used to synchronously regulate the wavefront phase and intensity distribution of the LG beam.
3. The LG beam recognition method based on the physical feature information fusion neural network according to claim 1, characterized in that, The lens in step S2 includes a pair of double convex lenses with a focal length ratio of 1:1 and another double convex lens. The pair of double convex lenses forms a 4f system, that is, the initial beam radius of the LG beam is retained , and the first-order diffraction spot is filtered out at the focal plane of the first double convex lens to extract a pure LG beam. The subsequent two double convex lenses are focused at the focal plane to generate a Fourier transform to form a Fraunhofer diffraction image containing topological charge number information; the pixel size of the CCD camera is less than 1 / 5 of the diffraction ring spacing of the Fraunhofer diffraction image.
4. A method for identifying LG beams based on a neural network that fuses physical feature information according to claim 1, characterized in that, In Step S2, using the lens to focus the LG beam generated in Step S1 to produce a Fourier transform, the equation of the LG beam at Z = 0 is expressed as: ; In the above formula, represents the LG beam at Z = 0, is the polar coordinate on the incident surface, z is the propagation distance, r is the radial coordinate in the polar coordinate; is the generalized Laguerre polynomial, is the radial exponent, is the topological charge number of the LG beam; is the radius of the LG beam; is the radial intensity distribution of the LG beam; is the imaginary unit; is the azimuth angle in the polar coordinate; The LG beam passes through the lens and forms a Fraunhofer diffraction spot at the focal plane. The expression of the Fraunhofer diffraction spot is abbreviated as the Fourier transform form of the incident light field as: ; In the above formula, is the Fourier transform of the incident light field; is the Fourier transform representation; respectively correspond to the frequency domain coordinates of.
5. The LG beam recognition method based on physical feature information fusion neural network according to claim 4, characterized in that, In Step S3, the improved neural network structure includes a convolutional layer, a pooling layer, a physical feature information fusion layer, and a fully connected layer; The convolutional layer is used to extract local features through a convolution kernel, expand the receptive field by inserting holes in the convolution kernel, and enhance the ability to extract global and local information of the image; The pooling layer includes a max pooling layer and an adaptive pooling layer. The max pooling layer is used to reduce the spatial size of the features while retaining the important information of the features, and the adaptive pooling layer is used to uniformly pool the input of any size into a fixed size to adapt to different input sizes; The physical feature information fusion layer is used to fuse the input matrix information. The input feature matrix is divided into three branches. Each branch undergoes a non-linear transformation through matrix dot product and Sigmoid function, and then the features of each branch are fused at different nodes. The features of multiple branches are combined through matrix addition operation. The fused features are then passed through a ReLU activation function to output a more tightly fused and more stable feature matrix, enhancing the model's ability to identify and express the input information; The fully connected layer is used to map the extracted features to the output space to obtain the final prediction result.
6. The LG beam recognition method based on the physical feature information fusion neural network according to claim 5, characterized in that, The convolutional layer includes an atrous convolutional layer and a standard convolutional layer. The atrous convolutional layer expands the receptive field of the convolution kernel through a preset atrous rate, enabling the atrous convolutional layer to capture periodic features spanning a large physical scale in the input data while avoiding introducing additional parameters; the standard convolutional layer finely extracts the phase gradient between adjacent pixels and the detailed features of the singularity of the vortex center using the receptive field of a dense convolution kernel; the atrous convolutional layer and the standard convolutional layer independently process the input through a parallel branch structure, and finally fuse the two feature maps through channel concatenation, thereby enhancing the ability to extract global and local information of the image.
7. A method for identifying LG beams based on a neural network for fusing physical feature information according to claim 5, characterized in that, In step S3, the improved neural network structure enhances the robustness of the neural network in complex multi-task scenarios by optimizing the loss function. The design process of the loss function is as follows: First, calculate the k cross-entropy loss for the th task scenario: ; In the above formula, C is the number of categories; represents the probability that the true label of the sample x is on the n th category; is the predicted probability of the sample x on the n th category; Then introduce the variance parameter Dynamically adjust the loss weights of each task, and divide the cross-entropy loss in each task scenario by the corresponding variance parameter Then introduce the logarithmic term of the variance parameter to obtain the final loss function : ; In the above formula, is the cross-entropy loss in the k th task scenario; is the learnable uncertainty parameter of the k th task.
Citation Information
Patent Citations
Coherent diffraction imaging method and its processing equipment
CN101482503A
Multimode vortex beam modal identification system based on feedforward neural network
CN111985320A
Interference and convolutional neural network mixed scheme measurement method
CN112836422A
Measuring device and measuring method for Laguerre Gaussian beam high-order topological charge
CN114720001A
Method for identifying high-order superposition state vortex beam based on convolutional neural network
CN116956137A