Three-dimensional space-oriented all-optical diffraction neural network gesture recognition method
By constructing a physical imaging model and a dynamic light field input mechanism in an all-optical diffraction neural network and optimizing the phase parameters of the diffraction layer, the problem of recognition failure caused by imaging scale scaling and defocusing effects in three-dimensional space is solved, achieving high accuracy and robust gesture recognition.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHANGCHUN UNIV OF SCI & TECH
- Filing Date
- 2026-01-15
- Publication Date
- 2026-04-21
AI Technical Summary
Existing all-optical diffraction neural networks cannot adapt to the scaling and defocusing effects caused by target movement in three-dimensional space, resulting in insufficient generalization ability and recognition failure.
By constructing a physical imaging model, establishing the mapping relationship between object distance and image distance, introducing a dynamic light field input mechanism, optimizing the phase parameters of the diffraction layer, and training the network using a random object distance sampling strategy, the network can be adapted to changes in three-dimensional space.
Within a wide object distance range of 0.1m to 1.0m, the recognition accuracy reaches over 84%, significantly improving the generalization ability and robustness of the all-optical diffraction neural network in three-dimensional space.
Smart Images

Figure CN121904841A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of interdisciplinary technology of optical computing and artificial intelligence, and specifically relates to a gesture recognition method using an all-optical diffraction neural network for three-dimensional space. Background Technology
[0002] With the development of technologies such as human-computer interaction and virtual reality (VR), dynamic gesture recognition has become a key technology. Traditional recognition methods are mostly based on electronic neural networks, which suffer from problems such as photoelectric conversion delay and high power consumption. Moreover, with the slowdown of Moore's Law, electronic computing is facing physical limits. All-optical diffraction neural networks (D2NNs), as an emerging architecture, utilize the parallelism and modulation capabilities of light waves, offering advantages such as low latency and zero-power computation. However, existing D2NN research mainly focuses on static targets in a two-dimensional plane. In actual gesture recognition, the target is in three-dimensional space and dynamically changes with distance. Changes in object distance lead to changes in imaging size (scale change) and changes in the light field propagation path. Traditional D2NN models trained on a fixed plane are highly sensitive to these changes in input distance and target size. When the target deviates from the reference distance, the recognition accuracy drops significantly, and the generalization ability is insufficient.
[0003] See the article "All-optical machine learning using diffractive deep neural networks" published by Lin et al. in the journal Science. This article discloses an all-optical deep neural network composed of multiple cascaded diffractive layers, which optimizes phase parameters through deep learning to achieve light speed classification of fixed planar images. However, this method is only trained for a single fixed object distance and ignores the imaging size scaling and defocusing effects caused by the movement of the target in three-dimensional space. The all-optical diffractive neural network has the problem of insufficient generalization ability and recognition failure due to its inability to adapt to the imaging scale scaling and defocusing effects caused by the movement of the target in three-dimensional space. Summary of the Invention
[0004] This invention addresses the shortcomings of existing all-optical diffraction neural networks, which suffer from insufficient generalization and recognition failures due to their inability to adapt to imaging scale scaling and defocusing effects caused by target movement in three-dimensional space. It provides a gesture recognition method based on an all-optical diffraction neural network for three-dimensional space. This method introduces physical imaging mechanisms during the network construction phase. First, a mapping model between object distance and image distance, and between object distance and magnification, is established using the thin lens imaging formula. Second, a dynamic light field input mechanism is constructed. When calculating the light field of the input layer, the image is scaled not only according to the magnification but also dynamically adjusted to its propagation distance to the first diffraction layer based on changes in the image plane position. During the training phase, a random object distance sampling strategy is employed, forcing the network to learn robust features that adapt to spatial changes, thereby optimizing the phase parameters of the diffraction layer.
[0005] To achieve the above objectives, the technical solution adopted by the present invention is as follows:
[0006] A gesture recognition method using an all-optical diffraction neural network for three-dimensional space, comprising the following steps:
[0007] Step 1: Construct the physical model and set its parameters;
[0008] Step 2: Establish the geometric mapping relationship for stereoscopic spatial imaging;
[0009] Step 3: Dynamic modeling of the input layer light field;
[0010] Step 4: Calculate the light field propagation from the dynamic imaging plane to the first diffraction layer of the D2NN diffraction layer;
[0011] Step 5: Network structure constraints and forward propagation;
[0012] Step 6: Detector layer signal extraction and category determination;
[0013] Step 7: D2NN diffraction layer training and parameter optimization.
[0014] The beneficial effects of this invention are as follows: By establishing a nonlinear mapping model of the object image and a dynamic light field propagation model, this invention accurately corrects the imaging scale scaling and defocus phase mismatch caused by changes in object distance in three-dimensional space, effectively solving the recognition failure problem caused by these factors. Experiments show that this method enables the model to achieve an average recognition accuracy of over 84% in a wide range of object distances from 0.1m to 1.0m, significantly improving the generalization ability and robustness of the all-optical diffraction neural network for dynamic targets in three-dimensional space. Attached Figure Description
[0015] Figure 1 This is a flowchart of a three-dimensional space-oriented all-optical diffraction neural network gesture recognition method according to the present invention;
[0016] Figure 2 This is a schematic diagram illustrating the construction of the physical model as described in this invention;
[0017] Figure 3 This is a schematic diagram of the optical path for imaging the gesture described in this invention at different object distances;
[0018] Figure 4 This is a schematic diagram of the D2NN diffraction layer structure described in this invention;
[0019] Figure 5 This is a schematic diagram of the interlayer connection structure of the D2NN diffraction layer described in this invention;
[0020] Figure 6 This is a schematic diagram of the detector structure described in this invention;
[0021] Figure 7 A comparison chart showing the recognition accuracy of existing methods and the method of the present invention at different object distances;
[0022] Figure 8 A comparison diagram of the phase visualization of the D2NN diffraction layer between existing methods and the method of this invention. Detailed Implementation
[0023] The technical solution of the present invention will now be clearly and completely described in conjunction with the accompanying drawings.
[0024] like Figure 1 As shown, a three-dimensional spatial all-optical diffraction neural network gesture recognition method includes the following steps:
[0025] Step 1: Construct the physical model and set its parameters;
[0026] Building physical models in computers, such as Figure 2 As shown, the model includes, along the optical axis, a gesture target 1, an imaging lens 2, a D2NN diffraction layer 3, and a detector 4.
[0027] The operating wavelength of this model The wavelength is set to 1550 nm, which is in the near-infrared band. The focal length of imaging lens 2 is f. The range of motion of the gesture target 1 is set to [0.1, 1] m. The median value of the imaging range is selected as the reference, that is, the reference object distance u = 0.168 m is set, and the corresponding reference image distance v is calculated according to the lens imaging formula.
[0028] Step 2: Establish the geometric mapping relationship for stereoscopic spatial imaging;
[0029] Based on the principles of geometric optics, a nonlinear mapping between the real-time position of the gesture target and the input parameters is established. For example... Figure 3 The diagram shows the optical path for imaging the gesture target 1 at different object distances. When the object distance of the gesture target 1 is u, according to the lens imaging formula... The corresponding real-time image distance v is calculated. Since changes in object distance lead to changes in imaging magnification, this invention defines an image size scaling factor M. As the object distance changes, the image size of the target also changes, and this change is determined by the optical magnification. Reference object distance Image distance Below, its reference magnification is: At this reference object distance, the image size is Therefore, the formula for image size change at any object distance is:
[0030]
[0031] This formula establishes how the object distance u determines the geometric size changes of the input image, providing a geometric basis for subsequent operations.
[0032] Step 3: Dynamic modeling of the input layer light field;
[0033] Based on the scaling factor M calculated in step two, the light field of the gesture target 1 in the input network is reconstructed. According to the Rayleigh–Sommerfeld diffraction theory, each diffraction layer of the D2NN diffraction layer 3 can be regarded as a secondary source of light waves. Its output function expression is:
[0034]
[0035] In the formula: and These represent the diffraction weight and transmission coefficient of the i-th neuron in layer l, respectively. The complex amplitude representing the input to the neuron is expressed as follows:
[0036]
[0037] In the formula: The wavelength representing the incident light wave. j is the imaginary unit. The amplitude representing the modulation of the light wave, This represents the phase of the modulation. Ultimately, the equation can be written as:
[0038]
[0039] In the formula: Represents the relative amplitude of the secondary wave. This refers to the additional phase delay of the secondary wave caused by the input wave of the neuron and its transmission coefficient.
[0040] Step 4: Calculate the light field propagation from the dynamic imaging plane to the first diffraction layer of the D2NN diffraction layer;
[0041] Based on the dynamic propagation distance from the input layer to the first diffraction layer of D2NN diffraction layer 3, angular spectrum diffraction theory is used to model and describe the propagation process of light waves between different diffraction layers. Let the third... Layer output complex amplitude is Its propagation distance in free space The result is:
[0042] ;
[0043] In the formula: and These represent the Fourier transform and its inverse transform, respectively. For spatial frequency; the angular spectral transfer function is defined as:
[0044]
[0045] In practical gesture recognition, the target moves freely in space, causing a spatial shift in the light field of the input layer, and the image size dynamically scales with changes in object distance. This non-static input effect is incorporated into the modeling of the D2NN diffraction layer 3, and the complex amplitude of the input layer is set as follows:
[0046]
[0047] in Indicates object distance The image size is relative to the reference size. scaling factor, and These represent the amplitude and phase of the incident light, respectively. Combining this with the dynamic propagation distance from the input layer to the first diffraction layer, the improved angular spectrum propagation formula is:
[0048]
[0049] This formula enables the first diffraction layer to accurately receive the light field distribution of the target's spatial motion characteristics, while the second to fifth diffraction layers at the rear still propagate the angular spectrum according to this formula.
[0050] like Figure 4 As shown, the D2NN diffraction layer 3 includes an input layer, a diffraction layer, and an output layer. The diffraction layer comprises a first diffraction layer, a second diffraction layer, a third diffraction layer, a fourth diffraction layer, and a fifth diffraction layer. Each diffraction layer contains N×N neurons, which adjust the phase of light to ultimately produce light intensity outputs at different locations in the output layer.
[0051] Coherent light carrying input information illuminates the input layer and propagates forward through diffraction. The light field passes through cascaded diffraction layers in sequence. The neurons in each layer modulate the phase of the transmitted light field, and the light waves between the layers achieve fully connected superposition and interference through diffraction. After multiple layers of optical evolution, the light field is finally focused and mapped onto the output layer plane.
[0052] Step 5: Network structure constraints and forward propagation;
[0053] In the structural design of D2NN diffraction layer 3, it is also necessary to satisfy the fully connected structure between diffraction layers, that is, each neuron in the previous layer can cover all neurons in the next layer. The expression for the maximum half-cone diffraction angle that satisfies this requirement can be written as:
[0054]
[0055] In the formula: This represents the size of the neuron. When the neuron size satisfies... At this point, the diffraction layers can achieve a maximum semi-cone diffraction angle of 90°, at which point the light wave can propagate uniformly throughout the entire space, and the distance between adjacent layers is no longer a constraint; while when the selected and The result At this time, geometric constraints need to be applied to the distance between layers to ensure that each neuron is within the range of... The diffracted light within the layer covers all neurons in the next layer. To constrain the distance along the diagonal line with the greatest distance in the diffracted layer, the distance constraint expression can be written as:
[0056]
[0057] like Figure 5 As shown, a single neuron on the left side. The emitted light waves diffract at the maximum semi-cone angle For forward diffusion, to ensure that information from neurons can be transmitted to neurons in the next diffraction layer, the range of the light wave must encompass the maximum geometric dimension of the next diffraction layer. That is, the length of the diagonal of the diffraction layer.
[0058] Step 6: Detector layer signal extraction and category determination;
[0059] After the light field undergoes phase modulation through the last D2NN diffraction layer 3, it propagates to the detector 4 at a predetermined distance. For example... Figure 6 As shown, detector 4 sets up 10 non-overlapping rectangular detection areas for gesture recognition tasks from 0 to 9. The number of areas K corresponds to the total number of gesture categories to be recognized. The system obtains the total light intensity energy of each detection area by calculating the square integral of the complex amplitude mode within that area. Finally, a maximum light intensity determination strategy is adopted, which compares the light intensity values of these 10 areas and determines the digital label corresponding to the area with the highest light intensity energy as the predicted category of the current input gesture.
[0060] Step 7: Training and parameter optimization of D2NN diffraction layer 3;
[0061] A dynamic training dataset containing multi-object distance samples is constructed. In each training iteration, the object distance u is first randomly sampled within the spatial range [0.1, 1.0] m according to the methods in steps two and three, and the corresponding dynamic input light field is generated. Then, the forward propagation light field is calculated according to steps four and five. Subsequently, using the detector 4 definition in step six, the cross-entropy loss function between the output light intensity distribution and the real gesture label is calculated. Using stochastic gradient descent (SGD) or Adam optimizer, the gradient of the loss function with respect to the phase of each diffraction layer neuron is calculated, and the phase parameters of all D2NN diffraction layers 3 are updated inversely. By repeating the above iterative process until the loss function converges, a set of phase parameters of an all-optical diffraction neural network that can adapt to changes in object distance in three-dimensional space and has high robustness is finally obtained.
[0062] Example:
[0063] To investigate the imaging patterns of gestures under different object distances and their impact on recognition performance, the American Sign Language Digit Dataset (ASL) was selected as the experimental dataset. This dataset, which is the main research object, includes ten categories of gesture images, with 500 images for each category, totaling 5000 images. 80% of these images were used as the training dataset, and the remaining 20% were used as the test dataset.
[0064] D2NN diffraction layer 3 consists of 5 diffraction layers. The input wavelength is selected as 1550nm eye-safe laser light. The neuron size is set to 3µm, and the number of neurons in each layer is set to 512×512. The maximum diffraction angle between network layers is 15°. The network construction needs to consider the distance constraint between D2NN diffraction layers, with the distance between adjacent D2NN diffraction layers being approximately 4mm. To effectively obtain the spatial distribution of diffracted light intensity at the output of the designed network model and achieve target recognition, detector 4 is introduced into the model output layer to simulate the actual detector target surface. Its size is consistent with the filled input image, both being 512×512 pixels. The detection area consists of square detection units with a side length of 35 pixels, arranged in three rows vertically. Each row has 3, 4, and 3 detection units respectively, for a total of 10 discrete detection areas. The spacing between detection units is adjusted through parameter settings to ensure that the entire array is evenly distributed at the center of detector 4, avoiding area overlap and detection blind spots.
[0065] By incorporating the propagation distance from the input layer to the first diffraction layer and the target size scaling factor into the model, the network can learn the imaging variation characteristics under different object distances during training. Based on an extended dataset, the improved model was trained and its recognition performance was evaluated on images with different sampling points ranging from 0.1 m to 1.0 m. The method of this invention can maintain stable performance over a wide range of object distances: Figure 7 As shown, the recognition rate of the method of the present invention remains stable at over 80%. In contrast, existing methods exhibit extremely strong distance sensitivity: they can only achieve a peak recognition rate of nearly 98% at the reference object distance set during training; however, once the test object distance deviates from this reference position, the recognition accuracy drops precipitously. This indicates that the method of the present invention solves the performance degradation problem caused by changes in input distance and target size in existing methods, not only maintaining high recognition accuracy at the reference input surface, but also exhibiting stronger generalization ability and robustness over extended distances.
[0066] After training, a visualization analysis was performed on the phase distribution of the D2NN diffraction layer using existing methods and the method of this invention. The results are as follows: Figure 8 As shown in the diagram, the two models are quite similar in overall phase, making it difficult to distinguish their differences with the naked eye. To further compare the local structural features of the two models, an 8×8 region at the center of the D2NN diffraction layer was selected for three-dimensional visualization, and the phase values were uniformly normalized to 0 to 1. It can be observed that the phase distribution of the two methods in this region exhibits distinctly different spatial structural characteristics: the phase height distribution of the existing method is relatively chaotic and lacks obvious regularity; while the method of this invention exhibits a topological morphology with specific undulating gradients in the central region. This significant structural difference strongly demonstrates that the training strategy proposed in this invention has successfully guided the diffraction layer parameters to undergo directional optimization, enabling it to learn a light field modulation mechanism at the physical level that can actively compensate for changes in object distance (such as defocusing and scale scaling). This is the fundamental reason why this invention possesses excellent robustness in three-dimensional space.
Claims
1. A gesture recognition method based on an all-optical diffraction neural network for three-dimensional space, characterized in that, The method includes the following steps: Step 1: Construct the physical model and set its parameters; Step 2: Establish the geometric mapping relationship for stereoscopic spatial imaging; Step 3: Dynamic modeling of the input layer light field; Step 4: Calculate the propagation of the light field from the dynamic imaging surface to the first diffraction layer; Step 5: Network structure constraints and forward propagation; Step 6: Detector layer signal extraction and category determination; Step 7: D2NN diffraction layer training and parameter optimization.
2. The all-optical diffraction neural network gesture recognition method for three-dimensional space according to claim 1, characterized in that, The specific steps of constructing the physical model and setting the physical model parameters are as follows: A physical model is constructed in the computer. The model includes, along the optical axis, a gesture target, an imaging lens, a D2NN diffraction layer, and a detector. The D2NN diffraction layer includes an input layer, a diffraction layer, and an output layer. The diffraction layer includes a first diffraction layer, a second diffraction layer, a third diffraction layer, a fourth diffraction layer, and a fifth diffraction layer cascaded together. Each diffraction layer includes N×N neurons. Coherent light carrying input information illuminates the input layer and propagates forward through diffraction. The light field passes through cascaded diffraction layers in sequence. The neurons in each layer modulate the phase of the transmitted light field, and the light waves between the layers achieve fully connected superposition and interference through diffraction. After multiple layers of optical evolution, the light field is finally focused and mapped onto the output layer plane. The operating wavelength of this model The wavelength is set to 1550 nm, which is in the near-infrared band; the focal length of the imaging lens is f; the range of motion of the gesture target is set to [0.1, 1] m; the median value of the imaging range is selected as the reference, that is, the reference object distance u = 0.168 m is set, and the corresponding reference image distance v is calculated according to the lens imaging formula.
3. The all-optical diffraction neural network gesture recognition method for three-dimensional space according to claim 2, characterized in that, The second step of establishing the geometric mapping relationship for stereoscopic spatial imaging specifically involves: Based on the principles of geometric optics, a nonlinear mapping between the real-time position of the gesture target and the input parameters is established; when the object distance of the gesture target is u, the lens imaging formula is applied. Calculate the corresponding real-time image distance v; define the image size scaling factor M, which changes the target's imaging size as the object distance changes, and this change is determined by the optical magnification. Reference object distance Image distance Below, its reference magnification is: At this reference object distance, the image size is Therefore, the formula for image size change at any object distance is: This formula establishes how the object distance u determines the geometric size changes of the input image, providing a geometric basis for subsequent operations.
4. The all-optical diffraction neural network gesture recognition method for three-dimensional space according to claim 3, characterized in that, The dynamic modeling of the input layer light field in step three specifically involves: Based on the scaling factor M calculated in step two, the light field of the gesture target in the input network is reconstructed. According to the Rayleigh–Sommerfeld diffraction theory, each neuron in the D2NN diffraction layer can be regarded as a secondary source of light waves; its output function expression is: In the formula: and These represent the diffraction weight and transmission coefficient of the i-th neuron in layer l, respectively. The complex amplitude representing the input to the neuron is expressed as follows: In the formula: The wavelength representing the incident light wave. j is the imaginary unit. The amplitude representing the modulation of the light wave, Represents the phase of modulation; ultimately, the equation can be written as: In the formula: Represents the relative amplitude of the secondary wave. This refers to the additional phase delay of the secondary wave caused by the input wave of the neuron and its transmission coefficient.
5. The all-optical diffraction neural network gesture recognition method for three-dimensional space according to claim 4, characterized in that, Step four, calculating the light field propagation from the dynamic imaging surface to the first diffraction layer, specifically involves: Combining the dynamic propagation distance from the input layer to the first diffraction layer of the D2NN, angular spectral diffraction theory is used to model and describe the propagation process of light waves between different diffraction layers; let the first... Layer output complex amplitude is Its propagation distance in free space The result is: ; In the formula: and These represent the Fourier transform and its inverse transform, respectively. For spatial frequency; the angular spectral transfer function is defined as: In practical gesture recognition, the target moves freely in space, causing a spatial shift in the light field of the input layer, and the image size dynamically scales with changes in object distance. Incorporating this non-static input effect into the modeling of the D2NN diffraction layer, the complex amplitude of the input layer is set as follows: in Indicates object distance The image size is relative to the reference size. scaling factor, and Let A and B be the amplitude and phase of the incident light, respectively. Combining this with the dynamic propagation distance from the input layer to the first diffraction layer, the improved angular spectrum propagation formula is: This formula enables the first diffraction layer to accurately receive the light field distribution of the target's spatial motion characteristics, while the second to fifth diffraction layers at the rear still propagate the angular spectrum according to this formula.
6. The all-optical diffraction neural network gesture recognition method for three-dimensional space according to claim 5, characterized in that, Step five, network structure constraints and forward propagation, specifically involves: In a D2NN diffraction layer, each neuron in the previous layer can cover all neurons in the next layer. The maximum half-cone diffraction angle that satisfies this requirement can be expressed as: In the formula: Represents the size of the neuron; when the neuron size satisfies At this point, the diffraction layers can achieve a maximum semi-cone diffraction angle of 90°, at which point the light wave can propagate uniformly throughout the entire space, and the distance between adjacent layers is no longer a constraint; while when the selected and The result At this time, geometric constraints need to be applied to the distance between layers to ensure that each neuron is within the range of... The diffracted light within the layer covers all neurons in the next layer. To constrain the distance along the diagonal line with the greatest distance in the diffracted layer, the distance constraint expression can be written as: single neuron The emitted light waves diffract at the maximum semi-cone angle Spread forward; among them This indicates the length of the diagonal of the diffraction layer.
7. The all-optical diffraction neural network gesture recognition method for three-dimensional space according to claim 6, characterized in that, The sixth step, detector layer signal extraction and category determination, specifically involves: After the light field passes through the phase modulation of the last D2NN diffraction layer, it propagates to the detector at a preset distance. The detector sets 10 non-overlapping rectangular detection areas for the gesture recognition task from 0 to 9. The number of areas K corresponds to the total number of gesture categories to be recognized. The system obtains the total light intensity energy of each detection area by calculating the square integral of the complex amplitude mode within that area. Finally, the maximum light intensity determination strategy is adopted, that is, the light intensity values of these 10 areas are compared, and the digital label corresponding to the area with the maximum light intensity energy is determined as the predicted category of the current input gesture.
8. The all-optical diffraction neural network gesture recognition method for three-dimensional space according to claim 7, characterized in that, The specific steps of step seven, training and parameter optimization of the D2NN diffraction layer 3, are as follows: A dynamic training dataset containing multi-object distance samples is constructed. In each training iteration, the object distance u is first randomly sampled within the spatial range of [0.1, 1.0] m according to the methods in steps two and three, and the corresponding dynamic input light field is generated. Then, the forward propagation light field is calculated according to steps four and five. Subsequently, the cross-entropy loss function between the output light intensity distribution and the real gesture label is calculated using the detector definition in step six. Using stochastic gradient descent or Adam optimizer, the gradient of the loss function with respect to the phase of each neuron in the diffraction layer is calculated, and the phase parameters of all D2NN diffraction layers are updated in reverse. By repeating the above iterative process until the loss function converges, a set of all-optical diffraction neural network phase parameters that can adapt to changes in object distance in three-dimensional space and have high robustness is finally obtained.