Auto-collimation large-range light spot center positioning method based on attention residual network

The local perception ability of self-collimated large-scale spot positioning is enhanced by the attention residual network, which solves the problem of high-precision coordinate positioning and realizes fast and high-precision spot center prediction.

CN120651098APending Publication Date: 2025-09-16HARBIN INST OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510593218.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-09
Publication Date
2025-09-16

AI Technical Summary

Technical Problem

It is impossible to directly and quickly obtain high-precision coordinate positioning results in self-collimation large-scale spot positioning, and deep networks may lead to problems such as gradient disappearance, gradient explosion, and fewer features.

Method used

A method based on attention residual network is adopted. By introducing residual blocks and attention mechanism, the local perception ability of the network is enhanced, the weight of the spot area is increased, and high-precision prediction of the spot center coordinates is achieved.

Benefits of technology

It achieves fast and convenient high-precision prediction of the center coordinates of the light spot with an error of less than ±0.1 pixel, and improves the light spot positioning accuracy by 88.34% and 75.34%, which is applicable to full-size resolution light spot images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120651098A_ABST
    Figure CN120651098A_ABST
Patent Text Reader

Abstract

The invention relates to an auto-collimation large-range light spot center positioning method, in particular to an auto-collimation large-range light spot center positioning method based on an attention residual network. In order to solve the problem that a high-precision coordinate positioning result cannot be directly and quickly obtained in auto-collimation large-range light spot positioning, the invention introduces a residual block and an attention mechanism to respectively solve the problems of gradient disappearance, gradient explosion and degradation possibly caused by a deep network in a large-scale image and few features in a small-scale light spot. According to the method, the local sensing capability of the network is enhanced, the light spot area weight is improved, and the purpose of reinforcement learning is achieved, so that a high-precision light spot center coordinate prediction result is accurately and rapidly obtained, online angle measurement can be more rapidly and conveniently achieved, and the method is suitable for light spot images with various full-size resolutions. The invention belongs to the technical field of auto-collimation angle measurement.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to a method for positioning the center of a large-scale self-collimating light spot, and belongs to the technical field of self-collimating angle measurement. Background Art

[0002] An autocollimator is a metrological instrument that uses the principle of light self-collimation to convert angular measurements into linear measurements. It is widely used for small angle measurements, flatness measurements of flat panels, and the straightness and parallelism of guide rails. However, currently, it is impossible to directly and quickly obtain high-precision coordinate positioning results using autocollimation for large-scale spot positioning. Furthermore, deep networks can cause gradient vanishing, gradient explosion, and degradation in large-scale images, while small-scale spots suffer from a lack of features. Summary of the Invention

[0003] This paper addresses the problem of being unable to directly and quickly obtain high-precision coordinate positioning results in self-collimating large-scale light spot positioning. A method for self-collimating large-scale light spot center positioning based on an attention residual network is proposed. This method introduces residual blocks and an attention mechanism to address the gradient vanishing, gradient explosion, and degradation phenomena that may be caused by deep networks in large-scale images, as well as the problem of fewer features in small-scale light spots. This method enhances the network's local perception capability, improves the light spot area weight, and achieves the goal of reinforcement learning. This allows for accurate and rapid prediction of high-precision light spot center coordinates, enabling faster and more convenient online angle measurement. The method is applicable to light spot images of various full-scale resolutions.

[0004] The technical solution adopted by the present invention to solve the above problems is: the steps of the present invention include:

[0005] Step 1: Prepare the experimental setup;

[0006] Step 2: Set the collection location for the sample;

[0007] Step 3: Input the self-collimated large-scale spot image collected by the autocollimator in the experiment into the attention residual network.

[0008] Furthermore, the process of preparing the experimental device in step 1 is as follows:

[0009] Step 101: Turn on the autocollimator and the standard angle measuring instrument to preheat;

[0010] Step 102: Keep the optical axes of the autocollimator, the reflective target to be measured, and the standard angle measuring instrument on the same horizontal line, and adjust the angle of the autocollimator so that the autocollimator reading is as close to zero as possible;

[0011] Step 103: Adjust the yaw angle knob and the pitch angle knob of the standard angle measuring instrument so that the reading of the standard angle measuring instrument is close to zero;

[0012] Step 104: After the autocollimator and the standard angle measuring instrument are preheated for 48 hours, the above-mentioned operation of fine-tuning the instrument angle is repeated once, so that the readings of both instruments are close to zero.

[0013] Furthermore, in step 2, the autocollimator spot image and standard instrument angle measurement data are collected every 10 inches. According to the formula X = f·tan2α, where f is the focal length of the collimating objective lens, it is equivalent to collecting samples and labels every 14 pixels of spot displacement. First, the reflective target starts collecting when the center pitch angle and yaw angle are both zero, and S-shaped collection begins in the positive X direction. Each time it steps to a collection point, it is held for 5 seconds, and then image and data collection is performed. The collection is continuous for 5 seconds, and an image and average angle measurement data are randomly selected as a set of training samples and labels. After collecting the autocollimator spot image of the upper half, the reflective target is returned to the zero position, and S-shaped collection begins in the negative X direction. The collection method is consistent with the above.

[0014] Furthermore, in step 3, the angle measurement data α and β collected by the standard instrument are converted into spot displacements X and Y according to the formulas X = f·tan2α and Y = f·tan2β, where f is the focal length of the collimating objective lens. The center coordinates of the spot are used as network labels to perform attention residual network modeling and training, thereby obtaining the correspondence between the self-collimating large-scale spot image and the spot center of mass.

[0015] The beneficial effects of the present invention are:

[0016] 1. The present invention directly inputs the full-scale, ultra-high-resolution, self-collimated, large-scale spot image of the sensor into the network. By using multiple serially connected residual blocks to extract features from the input image, it can avoid the gradient vanishing, gradient explosion, and degradation phenomena caused by the large number of parameters and deep networks, thereby enabling the network structure to converge to the target accuracy. Experimental results show that the present invention can enable the network structure to converge, ultimately achieving an error of less than ±0.1 pixel after the image data is directly input into the network, solving the problem that existing spot positioning algorithms cannot directly learn high-resolution, large-scale, self-collimated spot images.

[0017] 2. Based on the focusing principle and characteristics of human visual attention, this invention uses a maximum pooling layer and an average pooling layer to respectively perform feature enhancement extraction on the input image. After dimensional stacking, an attention factor is formed, which is multiplied by the output of the residual layer to enhance the local perception ability of the network, improve the weight of the spot area, and achieve the goal of reinforcement learning. Experimental results show that this method can enable the network to converge rapidly when the input image has very little effective information, and can improve the X and Y coordinate positioning accuracy of the spot by 88.34% and 75.34% respectively. This solves the problem that existing spot positioning algorithms cannot directly perform high-precision and fast positioning of spot images with very little effective information.

[0018] 3. The present invention obtains a local spot image by cropping to reduce the amount of calculation but loses some position information of the spot center relative to the original image center, and proposes a method of directly using the full-size spot image as the input of the neural network positioning model; this method can directly retain all the original information of the image, which is conducive to the network learning the absolute position of the spot, making the prediction result more accurate; and this method compresses the image preprocessing process, and can also realize rapid and efficient real-time measurement data display, which is convenient for the deployment of embedded measurement systems. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 This is a flow chart of the self-collimating large-scale spot center positioning algorithm based on the attention residual network;

[0020] Figure 2 It is a flowchart of data acquisition and neural network training;

[0021] Figure 3 is a schematic diagram of the angle measurement experimental setup;

[0022] Figure 3 In the middle, 1-autocollimator, 2-reflection target, 3-standard angle measuring instrument, 4-yaw angle knob, 5-pitch angle knob;

[0023] Figure 4 Schematic diagram of the attention residual network structure in Example 1;

[0024] Figure 4 In the figure, 1-first attention residual block, 2-second attention residual block, 3-third attention residual block, 4-fourth attention residual block, 5-fifth attention residual block, 6-sixth attention residual block;

[0025] Figure 5 Schematic diagram of the attention residual network structure in Example 2;

[0026] Figure 5In the figure, 1-first attention residual block, 2-second attention residual block, 3-third attention residual block, 4-fourth attention residual block, 5-fifth attention residual block, 6-sixth attention residual block. DETAILED DESCRIPTION

[0027] Specific implementation method 1: Figure 1 As shown in FIG, a self-collimating large-scale spot center positioning method based on the attention residual network includes the following steps:

[0028] Step 1: Prepare the experimental device. The process of preparing the experimental device is as follows:

[0029] Step 101: Turn on the autocollimator (1) and the standard angle measuring instrument (3) to preheat;

[0030] Step 102: Keep the optical axes of the autocollimator (1), the reflective target (2), and the standard angle measuring instrument (3) on the same horizontal line, and adjust the angle of the autocollimator (1) so that the reading of the autocollimator (1) is as close to zero as possible;

[0031] Step 103: Adjust the yaw angle knob (4) and the pitch angle knob (5) of the standard angle measuring instrument so that the reading of the standard angle measuring instrument is close to zero;

[0032] Step 104: After the autocollimator (1) and the standard angle measuring instrument (3) are preheated for 48 hours, the above-mentioned operation of fine-tuning the instrument angle is repeated once, so that the readings of both instruments are close to zero;

[0033] Step 2: Set the acquisition position for the sample; collect the autocollimator spot image and standard instrument angle measurement data every 10 inches. According to the formula X = f·tan2α, where f is the focal length of the collimator objective lens, it is equivalent to collecting a sample and label every 14 pixels of spot displacement. First, start the acquisition when the center pitch and yaw angles of the reflective target are both zero, and start S-shaped acquisition in the positive X direction. After each step to a collection point, hold it for 5 seconds, and then perform image and data acquisition. The acquisition is continuous for 5 seconds, and a random image and average angle measurement data are taken as a set of training samples and labels. After collecting the autocollimator spot image of the upper half, return the reflective target to zero position and start S-shaped acquisition in the negative X direction. The acquisition method is the same as above.

[0034] Step 3. Input the self-collimated large-scale light spot image collected by the autocollimator in the experiment into the attention residual network; convert the angle measurement data α and β collected by the standard instrument into the light spot displacement X and Y according to the formula X = f·tan2α and Y = f·tan2β, where f is the focal length of the collimating objective lens, and use the center coordinates of the light spot as the network label to model and train the attention residual network, so as to obtain the correspondence between the self-collimated large-scale light spot image and the light spot center of mass.

[0035] In this embodiment, the full-size ultra-high-resolution self-collimated large-scale light spot image of the sensor is used as the input of the attention residual network, and the actual X and Y coordinates of the center of the light spot are used as the output of the attention residual network. The correspondence between the full-size light spot image and the actual X and Y coordinates of the center of the light spot is established through the attention residual network, and the angle value corresponding to the coordinate is obtained through the self-collimation angle measurement formula.

[0036] Among them, such as Figure 2 As shown in Figure 2, the attention residual network training process is as follows:

[0037] After setting up the experimental setup and obtaining experimental data, the experimental data is fed into the constructed attention residual network for training. First, the network parameters are initialized to prepare for the training process.

[0038] The training set uses a random training method to reduce data set bias, improve model robustness, and avoid local optimal situations.

[0039] The network then uses forward propagation to calculate the network's predictions for a sample of self-collimated large-scale spot images and calculate the error between the predictions and the actual labels. If the error is greater than expected, the predictions are backpropagated using an adaptive learning rate optimization algorithm, continuously adjusting the network weights to reduce the error. If the error meets the required standard, the error distribution on the trained validation set is used to ensure the model's generalization ability. If the model performance is suboptimal, further adjustments are made to the network structure, number of layers, number of neurons, and other parameters.

[0040] like Figure 4 As shown, the attention residual network structure is as follows:

[0041] The attention residual network is mainly composed of a spot area attention module, an attention residual module, a dual spatial attention module and a multi-classification fully connected module.

[0042] Light Spot Area Attention Module: This module consists of a convolutional layer with a convolution kernel size of 512×512 and an output channel size of 512, a batch normalization (BN) layer, and a channel-based spatial attention (CBAM) layer. The resulting image is transformed using the Relu activation function. This module directly extracts features from the input sample using a large convolution kernel and uses the CBAM layer to enhance training of the results. The module output size is 6×6×512, which represents the area in the 6×6 grid where the light spot is more likely to fall. This is the light spot area attention factor.

[0043] Attention Residual Module: After an image sample is input, it undergoes feature extraction via a convolutional layer with an 8×8 kernel and 8 output channels, followed by a batch normalization layer. The feature extraction is then fed into six consecutive attention residual blocks. The first, second, and third attention residual blocks include a 3×3 residual block and a 7×7 CBAM block. The fifth, sixth, and seventh attention residual blocks include a 3×3 residual block and a 3×3 CBAM block. This module extracts multiple features from the input sample using multiple residual blocks and CBAM layers, thereby obtaining useful information about the sample.

[0044] Dual spatial attention module: This module consists of an average pooling layer and a maximum pooling layer to extract features from the input samples, and after stacking in the spatial dimension, it becomes a two-channel spatial attention factor, and is reduced in dimension through a 1×1 convolution layer to form a single-channel spatial attention coefficient. The first dual spatial attention module is downsampled by a pooling layer of size 96×96, and the spatial attention coefficient obtained is multiplied by the output matrix of the third attention residual block in the attention residual module. The second dual spatial attention module is downsampled by a pooling layer of size 48×48, and the spatial attention coefficient obtained is multiplied by the output matrix of the fourth attention residual block in the attention residual module. The dual spatial attention module is equivalent to a "short connection" that can retain more original light spot weight information, thereby achieving the purpose of strengthening the position information of the light spot area.

[0045] Multi-classification fully connected module: This module is mainly composed of two fully connected branches. The first branch is the output of the light spot area attention module. After a 1×1 average pooling layer for feature downsampling and feature flattening, it serves as the input of the first branch's fully connected layer. After a fully connected layer, it outputs 4 neurons, representing the possibility that the light spot is distributed in the upper left, upper right, lower left, and lower right areas of the image; the other branch is the output of the attention residual module multiplied by the output of the light spot area attention module, and then passes through a 1×1 average pooling layer for feature downsampling. After feature flattening, it serves as the input of the second branch's fully connected layer. After passing through two fully connected layers, it is merged with the first branch in the spatial dimension and then passes through another fully connected layer to obtain the two outputs of the network, representing the X and Y coordinates of the center of the large-scale light spot image.

[0046] Example

[0047] The full-size ultra-high-resolution autocollimation large-scale spot image of the sensor is used as the input of the attention residual network, and the actual X, Y coordinates of the center of the spot are used as the output of the attention residual network. The correspondence between the full-size spot image and the actual X, Y coordinates of the center of the spot is established through the attention residual network, and the angle value corresponding to the coordinate is obtained through the autocollimation angle measurement formula.

[0048] like Figure 1 As shown in the figure, the self-collimated large-scale light spot center positioning method based on the attention residual network is used to directly input the full-size ultra-high-resolution self-collimated large-scale light spot image of the sensor to obtain its high-precision X and Y coordinates of the light spot center to improve the angle measurement accuracy. The specific steps include:

[0049] Step 1: Simulate and generate a full-size autocollimated large-scale spot image of the sensor with 1536×1536 pixels, using the two-dimensional Gaussian formula I(x,y)=I0·exp(-[(x-x0) 2 +(y-y0) 2 ] / 2σ 2 ) generates a two-dimensional spot, I0 is the maximum intensity at the center of the spot, which is 255; the value of σ determines the width of the spot, and the σ value should be selected so that the spot diameter is approximately 32×32 pixels; (x0, y0) are the randomly generated coordinates of the center of the spot, generated within the pixel range of (-730, 730);

[0050] Step 2: Input the simulated self-collimated large-scale light spot image into the attention residual network, use the ideal center of mass of the light spot as the network label, and perform attention residual network modeling and pre-training to obtain the relationship between the simulated self-collimated large-scale light spot image and the center of mass of the light spot. Pre-training the network with simulated data can quickly converge to obtain a network with better performance, thereby greatly shortening the training time for subsequent adjustment of network parameters using experimental data.

[0051] Step 3: Prepare the experimental device, such as Figure 3 As shown, turn on the autocollimator 1 and the standard angle measuring instrument 3 to preheat, keep the optical axes of the autocollimator 1, the reflective target 2 to be measured, and the standard angle measuring instrument 3 on the same horizontal line, adjust the angle of the autocollimator 1 so that the reading of the autocollimator 1 is as close to zero as possible; adjust the yaw angle knob 4 and the pitch angle knob 5 of the standard angle measuring instrument so that the reading of the standard instrument angle measuring instrument is close to zero; wait for the autocollimator 1 and the standard angle measuring instrument 3 to preheat for 48 hours, and then repeat the above-mentioned fine-tuning operation of the instrument angles so that the readings of both instruments are close to zero;

[0052] Step 4: Set the acquisition position for the sample collection. Collect the autocollimator spot image and standard instrument angle measurement data every 10 inches. According to the formula X = f·tan2α, where f is the focal length of the collimator objective lens, it is equivalent to collecting a sample and label every 14 pixels of spot displacement. First, start the collection when the center pitch and yaw angles of the reflective target are both zero, and start S-shaped collection in the positive X direction. Each time it steps into a collection point, it is held for 5 seconds, and then image and data collection is performed. The collection is continuous for 5 seconds, and an image and average angle measurement data are randomly selected as a set of training samples and labels. After collecting the autocollimator spot image of the upper half, return the reflective target to zero position and start S-shaped collection in the negative X direction. The collection method is the same as above.

[0053] Step 5: Input the self-collimated large-scale light spot image collected by the autocollimator in the experiment into the attention residual network, and convert the angle measurement data α and β collected by the standard instrument into light spot displacements X and Y according to the formulas X=f·tan2α and Y=f·tan2β (where f is the focal length of the collimating objective lens). The center coordinates of the light spot are used as network labels to perform attention residual network modeling and training, thereby obtaining the correspondence between the self-collimated large-scale light spot image and the light spot centroid.

[0054] like Figure 2 As shown in Figure 2, the attention residual network training process is as follows:

[0055] First, initialize the network parameters to prepare for the training process.

[0056] The training set uses a random training method to reduce data set bias, improve model robustness, and avoid local optimal situations.

[0057] The network then uses forward propagation to calculate the network's predictions for a sample of self-collimated large-scale spot images and calculate the error between the predictions and the actual labels. If the error is greater than expected, the predictions are backpropagated using an adaptive learning rate optimization algorithm, continuously adjusting the network weights to reduce the error. If the error meets the required standard, the error distribution on the trained validation set is used to ensure the model's generalization ability. If the model performance is suboptimal, further adjustments are made to the network structure, number of layers, number of neurons, and other parameters.

[0058] like Figure 5 As shown, the attention residual network structure is as follows:

[0059] The attention residual network is mainly composed of a spot area attention module, an attention residual module, a dual spatial attention module and a multi-classification fully connected module.

[0060] Light Spot Area Attention Module: This module consists of a convolutional layer with a convolution kernel size of 512×512 and an output channel size of 512, a batch normalization (BN) layer, and a channel-based spatial attention (CBAM) layer. The resulting image is transformed using the Relu activation function. This module directly extracts features from the input sample using a large convolution kernel and uses the CBAM layer to enhance training of the results. The module output size is 6×6×512, which represents the area in the 6×6 grid where the light spot is more likely to fall. This is the light spot area attention factor.

[0061] Attention Residual Module: After an image sample is input, it undergoes feature extraction via a convolutional layer with an 8×8 kernel and 8 output channels, followed by a batch normalization layer. The feature extraction is then fed into six consecutive attention residual blocks. The first, second, and third attention residual blocks include a 3×3 residual block and a 7×7 CBAM block. The fifth, sixth, and seventh attention residual blocks include a 3×3 residual block and a 3×3 CBAM block. This module extracts multiple features from the input sample using multiple residual blocks and CBAM layers, thereby obtaining useful information about the sample.

[0062] Dual spatial attention module: This module consists of an average pooling layer and a maximum pooling layer to extract features from the input samples, and the outputs of the two are multiplied to form a single-channel spatial attention coefficient. The first dual spatial attention module is downsampled by a pooling layer of size 96×96, and the spatial attention coefficient obtained is multiplied by the output matrix of the third attention residual block in the attention residual module. The second dual spatial attention module is downsampled by a pooling layer of size 48×48, and the spatial attention coefficient obtained is multiplied by the output matrix of the fourth attention residual block in the attention residual module. The dual spatial attention module is equivalent to a "short connection" that can retain more original light spot weight information, thereby achieving the purpose of strengthening the position information of the light spot area.

[0063] Multi-classification fully connected module: This module is mainly composed of two fully connected branches. The first branch is the output of the light spot area attention module. After a 1×1 average pooling layer for feature downsampling and feature flattening, it serves as the input of the first branch's fully connected layer. After a fully connected layer, it outputs 4 neurons, representing the possibility that the light spot is distributed in the upper left, upper right, lower left, and lower right areas of the image; the other branch is the output of the attention residual module multiplied by the output of the light spot area attention module, and then passes through a 1×1 average pooling layer for feature downsampling. After feature flattening, it serves as the input of the second branch's fully connected layer. After passing through two fully connected layers, it is merged with the first branch in the spatial dimension and then passes through another fully connected layer to obtain the two outputs of the network, representing the X and Y coordinates of the center of the large-scale light spot image.

[0064] The above description is merely a preferred embodiment of the present invention and does not constitute any form of limitation to the present invention. Although the present invention has been disclosed as a preferred embodiment as above, it is not intended to limit the present invention. Any technician familiar with the present profession can make some changes or modifications to equivalent embodiments of equivalent changes using the technical content disclosed above without departing from the scope of the technical solution of the present invention. However, any simple modification, equivalent replacement and improvement of the above embodiments made according to the technical essence of the present invention, within the spirit and principles of the present invention, without departing from the content of the technical solution of the present invention, shall still fall within the scope of protection of the technical solution of the present invention.

Claims

1. A self-collimating large-scale spot center positioning method based on attention residual network, characterized by: The specific steps include: Step 1: Prepare the experimental setup; Step 2: Set the collection location for the sample; Step 3: Input the self-collimated large-scale spot image collected by the autocollimator in the experiment into the attention residual network.

2. The method for locating the center of a large-scale self-collimated light spot based on an attention residual network according to claim 1, characterized in that: The process of preparing the experimental device in step 1 is as follows: Step 101: Turn on the autocollimator (1) and the standard angle measuring instrument (3) to preheat; Step 102: Keep the optical axes of the autocollimator (1), the reflective target (2), and the standard angle measuring instrument (3) on the same horizontal line, and adjust the angle of the autocollimator (1) so that the reading of the autocollimator (1) is as close to zero as possible; Step 103: Adjust the yaw angle knob (4) and the pitch angle knob (5) of the standard angle measuring instrument so that the reading of the standard angle measuring instrument is close to zero; Step 104: After the autocollimator (1) and the standard angle measuring instrument (3) are preheated for 48 hours, the above-mentioned operation of fine-tuning the instrument angle is repeated once, so that the readings of the two instruments are close to zero.

3. The method for locating the center of a large-scale self-collimated light spot based on an attention residual network according to claim 1, characterized in that: In step 2, the autocollimator spot image and standard instrument angle measurement data are collected every 10 inches. According to the formula X = f·tan2α, where f is the focal length of the collimator objective lens, this is equivalent to collecting samples and labels approximately every 14 pixels of spot displacement. First, the reflective target is set to start collecting when the center pitch and yaw angles are both zero, and S-shaped collection is started in the positive X direction. Each time it steps to a collection point, it is held for 5 seconds before image and data collection is performed. The collection is continuous for 5 seconds, and an image and average angle measurement data are randomly selected as a set of training samples and labels. After collecting the autocollimator spot image of the upper half, the reflective target is returned to zero and S-shaped collection is started in the negative X direction. The collection method is the same as above.

4. The method for locating the center of a large-scale self-collimated light spot based on an attention residual network according to claim 1, characterized in that: In step 3, the angle measurement data α and β collected by the standard instrument are converted into spot displacements X and Y according to the formulas X = f·tan2α and Y = f·tan2β, where f is the focal length of the collimating objective lens. The center coordinates of the spot are used as network labels to perform attention residual network modeling and training, thereby obtaining the correspondence between the self-collimating large-scale spot image and the spot center of mass.

5. A self-collimating large-scale spot center positioning method based on attention residual network, characterized by: The specific structure includes: The attention residual network is mainly composed of the spot area attention module, the attention residual module, the dual spatial attention module and the multi-classification fully connected module; Light spot area attention module: This module consists of a convolution layer with a convolution kernel size of 512×512 and an output channel of 512, a batch normalization layer, and a channel-based and spatial attention layer. The results are transformed using the ReLu activation function. This module directly extracts features from input samples using a large convolution kernel and uses the CBAM layer to enhance the training of the results. The module output size is 6×6×512, which represents which area of ​​the grid is more likely to contain the light spot when the image is divided into a 6×6 grid. This is the light spot area attention factor. Attention residual module: After the image sample is input, a convolutional layer with a convolution kernel size of 8×8 and an output channel of 8 and a batch normalization layer performs feature extraction and inputs it into 6 consecutive attention residual blocks. The first, second, and third attention residual blocks include residual blocks with a convolution kernel size of 3×3 and CBAM blocks with a convolution kernel size of 7×7. The fifth, sixth, and seventh attention residual blocks include residual blocks with a convolution kernel size of 3×3 and CBAM blocks with a convolution kernel size of 3×3. This module extracts multiple features from the input sample through multiple residual blocks and CBAM layers to obtain useful information of the sample. Dual spatial attention module: This module extracts features from the input samples by an average pooling layer and a maximum pooling layer, and stacks them in the spatial dimension to form a two-channel spatial attention factor, which is then reduced to a single-channel spatial attention coefficient through a 1×1 convolution layer. The first dual spatial attention module is downsampled by a pooling layer of size 96×96, and the resulting spatial attention coefficient is multiplied by the output matrix of the third attention residual block in the attention residual module. The second dual spatial attention module is downsampled by a pooling layer of size 48×48, and the resulting spatial attention coefficient is multiplied by the output matrix of the fourth attention residual block in the attention residual module. The dual spatial attention module is equivalent to a "short connection" that can retain more original spot weight information, thereby enhancing the location information of the spot area. Multi-classification fully connected module: This module is mainly composed of two fully connected branches; the first branch is the output of the light spot area attention module, which is subjected to a 1×1 average pooling layer for feature downsampling and feature flattening, and then serves as the input of the first branch fully connected layer. After passing through a fully connected layer, 4 neurons are output, representing the possibility that the light spot is distributed in the upper left, upper right, lower left, and lower right areas of the image; the other branch is the output of the attention residual module multiplied by the output of the light spot area attention module, and then passes through a 1×1 average pooling layer for feature downsampling. After feature flattening, it serves as the input of the second branch fully connected layer. After passing through two fully connected layers, it is merged with the first branch in the spatial dimension, and then passes through a fully connected layer to obtain the two outputs of the network, representing the X and Y coordinates of the center of the large-scale light spot image.