Multimode optical fiber image transmission method, system, equipment and medium
Through the neural network model of sliding window self-attention mechanism, the anti-interference ability and dynamic environment adaptability of multimode fiber image transmission under disturbance conditions is solved, efficient image reconstruction and real-time monitoring are realized, and the generalization ability and computing efficiency of the model are improved.
Patent Information
- Application Number
- CN202510485308.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-07-25
AI Technical Summary
The existing multimode fiber image transmission technology has weak anti-interference ability, poor adaptability to dynamic environments, high computational complexity and insufficient model generalization ability under fiber disturbance conditions, resulting in a decrease in image reconstruction quality and unable to meet the high-precision requirements in actual applications.
A neural network model based on the sliding window self-attention mechanism is adopted, and image reconstruction is carried out through block partitioning module, linear embedding module, multi-layer sliding window self-attention block and image reconstruction module. Combining mixed training strategies and dynamic adjustment of calculation parameters, the model structure and training methods are optimized to enhance anti-interference ability and generalization ability.
It significantly improves the anti-interference ability and dynamic environment adaptability of the model, reduces the computational complexity, improves the SSIM and EME indicators of image reconstruction quality, and reduces the number of parameters, improving the generalization ability of the model under different perturbation conditions.
Smart Images

Figure CN120378584A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data transmission, and particularly to a multimode fiber optic image transmission method, system, device and medium. Background Art
[0002] Live working robots for transmission lines are important maintenance equipment in the power system, and are applied to live working, fault detection and repair tasks of transmission lines. These robots can safely and efficiently complete various operations in a high-voltage environment, ensuring the stable operation of the power system. The image transmission technology of live working robots is a key component, which can transmit high-resolution images in real time, helping operators remotely monitor and control the robots to ensure the safety and accuracy of operations. The current methods for realizing image transmission of transmission line operation robots include the following: multimode fiber optic image transmission based on convolutional neural network. Multimode fiber optic image transmission based on self-attention mechanism.
[0003] The existing methods for image transmission of transmission line operation robots have the following disadvantages:
[0004] Although the multimode fiber optic image transmission technology based on convolutional neural network (CNN) has achieved certain results in image reconstruction quality, it still faces several major problems. First, when the multimode fiber is disturbed (such as bending, vibration, temperature change, etc.), the CNN model has weak anti-interference ability, resulting in a significant decrease in image reconstruction quality, which makes the model need to be frequently retrained or calibrated in practical applications, increasing the operation complexity and cost. Second, the CNN model requires a large amount of training data to learn complex mapping relationships. The data acquisition process is complex and time-consuming, and it is difficult to ensure the diversity and representativeness of the data, resulting in limited generalization ability of the model. Finally, under new and unseen disturbance conditions, the image reconstruction quality of the model drops significantly, unable to meet the high-precision requirements in practical applications, limiting its applicability under different environmental conditions and affecting its wide promotion in practical scenarios.
[0005] Although the multimode fiber optic image transmission technology based on the self-attention mechanism has made significant progress in image reconstruction quality, it still faces some key challenges. First, although the self-attention mechanism can better learn long-range dependencies in the sequence, when the multimode fiber is disturbed, the robustness of the model is still limited, and the image reconstruction quality drops significantly when the fiber deforms or the environment changes, affecting the stability of the model in practical applications. Second, the model of the self-attention mechanism also requires a large amount of training data. The data acquisition process is complex and time-consuming, and it is difficult to ensure the diversity and representativeness of the data, resulting in limited generalization ability of the model and affecting its applicability under different environmental conditions. Finally, under new and unseen perturbation conditions, the image reconstruction quality drops significantly, unable to meet the high-precision requirements in practical applications, limiting its applicability under different environmental conditions and affecting its wide promotion in practical scenarios. Summary of the Invention
[0006] The purpose of the present invention is to provide a multimode fiber optic image transmission method, system, device and medium to solve the problems of weak anti-interference ability, poor adaptability to dynamic environment, high computational complexity and insufficient model generalization ability existing in the multimode fiber optic image transmission method in the prior art.
[0007] To achieve the above purpose, the following technical solutions are adopted.
[0008] A multimode fiber optic image transmission method includes the following steps.
[0009] Obtain the original image of the target area through the image acquisition device of the transmission line operation robot.
[0010] Convert the original image into an optical signal and transmit it through a multimode fiber to generate a speckle image with spatial perturbation characteristics.
[0011] Input the speckle image into a neural network model constructed based on the sliding window self-attention mechanism for image reconstruction.
[0012] Wherein the neural network model includes a block partition module, a linear embedding module, a multi-layer sliding window self-attention block and an image reconstruction module.
[0013] Output the reconstructed image to the transmission line status analysis system for real-time monitoring.
[0014] The sliding window self-attention block calculates local self-attention features within a fixed-size window through alternately set window multi-head self-attention and sliding window multi-head self-attention mechanisms, and realizes feature interaction between adjacent windows through the sliding window mechanism.
[0015] Optionally, the processing flow of the neural network model includes:
[0016] The input image is divided into multiple local blocks of 4×4 pixels by a block partitioning module;
[0017] The pixel values of each local block are mapped to high-dimensional feature vectors by a linear embedding module;
[0018] After being alternately processed by at least two groups of sliding window self-attention blocks, each group includes a combined structure of a window multi-head self-attention layer, a sliding window multi-head self-attention layer, and a multi-layer perceptron;
[0019] The deep features are compressed in the channel dimension by an average pooling layer;
[0020] Finally, the compressed feature vectors are reconstructed into an output image of the target size by a fully connected layer.
[0021] Optionally, the working process of the sliding window self-attention block includes:
[0022] In the window multi-head self-attention layer, the feature map is divided into fixed windows of 7×7, and multi-head self-attention calculation is performed inside each window;
[0023] In the sliding window multi-head self-attention layer, after performing a cyclic shift operation on the feature map, the windows are re-divided, and the attention calculation of non-adjacent regions is suppressed through a masking mechanism;
[0024] In the multi-layer perceptron, a two-layer fully connected structure is adopted for non-linear feature transformation;
[0025] After each calculation step, the original feature information is retained through a skip connection.
[0026] Optionally, the training method of the neural network model includes:
[0027] Construct a training data set containing various perturbation conditions, and simulate the perturbations of bending, vibration, and temperature changes of multimode optical fibers by adding Gaussian distribution noise to the real and imaginary parts of the transmission matrix respectively;
[0028] Adopt a hybrid training strategy to jointly train the speckle image pairs under different perturbation conditions by inputting them into the model;
[0029] Dynamically adjust the model parameters through an Adam optimizer, including the number of heads, window size, and network depth of the sliding window self-attention layer;
[0030] During the training process, synchronously verify the image reconstruction quality of the model under unknown perturbation configurations.
[0031] Optionally, the real-time monitoring includes:
[0032] Automatically identify the key components of the transmission line in the reconstructed image, including wire connection points, insulators, and tower structures;
[0033] Detect the abnormal state of the equipment by comparing with the historical image database, including wire wear, insulator discharge traces and tower inclination;
[0034] Generate a three-dimensional line state model by combining multi-angle reconstructed images;
[0035] Generate line maintenance decision suggestions based on real-time monitoring results.
[0036] Optionally, during the multimode fiber transmission:
[0037] When it is detected that the optical fiber transmission environment parameters exceed the preset threshold, dynamically adjust the calculation parameters of the sliding window self-attention block, including:
[0038] Increase the window division size proportionally according to the temperature change amplitude to compensate for the mode coupling interference caused by the fluctuation of the optical fiber refractive index. When the vibration frequency exceeds the preset frequency, automatically reduce the displacement step of the sliding window to enhance the local feature retention ability
[0039] In a strong electromagnetic interference environment, activate the noise suppression coefficient of the shielding mechanism to the highest level to filter the abnormal attention weights in non-adjacent regions
[0040] Based on the real-time collected optical fiber bending curvature data, adaptively adjust the feature fusion ratio between the sliding window self-attention layer and the multi-layer perceptron.
[0041] Optionally, perform the following data preprocessing before image reconstruction:
[0042] For the scene characteristics of the transmission line, apply directional data augmentation to the original speckle image: add strip noise patterns along the axis of the insulator to simulate wire discharge interference; superimpose metal reflection artifacts in the tower connection area to enhance the model's adaptability to high-reflection surfaces; generate a dynamic blur kernel according to the wire sag parameter to simulate image jitter caused by high-altitude wind vibration;
[0043] Perform feature domain alignment operations on the enhanced speckle image: extract the main frequency components of the speckle image through frequency domain analysis, establish a mapping relationship with the optical fiber perturbation type; automatically select the hierarchical processing path of the sliding window self-attention block according to the main frequency distribution characteristics; align the speckle distribution laws under different perturbation modes in the feature embedding space.
[0044] A multimode fiber image transmission system includes,
[0045] An image acquisition module, configured at the end of the transmission line operation robot, for acquiring high-resolution line state images;
[0046] An optical fiber transmission module, comprising a multimode optical fiber channel and its supporting optoelectronic conversion device, for converting an optical signal into a speckle image with spatial perturbation characteristics;
[0047] A deep learning reconstruction module, comprising a neural network model constructed based on a sliding window self-attention mechanism, for recovering an original line image from a speckle image;
[0048] A state analysis module, comprising an image recognition algorithm and a trend prediction model, for automatically detecting line equipment defects and generating maintenance suggestions.
[0049] An electronic device, comprising a processor and a memory, the processor being configured to execute a computer program stored in the memory to implement the described multimode optical fiber image transmission method.
[0050] A computer-readable storage medium storing at least one instruction, the at least one instruction, when executed by a processor, implementing the described multimode optical fiber image transmission method.
[0051] Compared with the prior art, the present invention has the following beneficial effects:
[0052] The present invention proposes a multimode optical fiber image transmission method. Through a neural network model constructed by a sliding window self-attention mechanism, it effectively solves the problem of the decline in image reconstruction quality in the prior art under the condition of optical fiber perturbation. The present invention improves the anti-interference ability, dynamic environment adaptability, reduces the computational complexity, and improves the model generalization ability. Compared with the prior art, this method calculates local self-attention features within a fixed-size window and realizes feature interaction between adjacent windows through a sliding window mechanism, significantly enhancing the anti-interference ability of the model. In addition, the hybrid training strategy further enhances the generalization ability of the model under different perturbation conditions, improving the SSIM and EME metrics of the image reconstruction quality by 0.04 and 0.79 respectively, while reducing the number of parameters by 25% and reducing the computational complexity.
[0053] The present invention further optimizes the structure and training method of the neural network model, improving the overall performance of the system. The processing flow of the neural network model, through the combination of a block partitioning module, a linear embedding module, and a multi-layer sliding window self-attention block, further reduces the dependence on a large amount of training data. The present invention respectively optimizes the working process of the sliding window self-attention block, the training method of the neural network model, and the real-time monitoring function, making the model more applicable in complex environments. The present invention further improves the adaptability of the system when the optical fiber transmission environment changes by dynamically adjusting calculation parameters and data preprocessing. The present invention applies this method to specific systems and devices, providing a complete solution for practical applications. Generally speaking, the present invention is superior to the prior art in terms of anti-interference ability, generalization ability, and calculation efficiency, and has significant innovation and practicality. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] Figure 1 FIG. 6 is a first schematic diagram of multimode fiber transmission in an embodiment of a multimode fiber image transmission method of the present invention;
[0055] Figure 2 FIG. 10 is a second schematic diagram of multimode fiber transmission in an embodiment of a multimode fiber image transmission method of the present invention;
[0056] Figure 3 FIG. 14 is a schematic diagram of the network structure in an embodiment of a multimode fiber image transmission method of the present invention;
[0057] Figure 4 FIG. 18 is a schematic diagram of the hybrid training architecture in an embodiment of a multimode fiber image transmission method of the present invention;
[0058] Figure 5 FIG. 22 is a block diagram of the structure of an electronic device in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0059] The present invention will be described in detail below with reference to the drawings and in conjunction with embodiments. It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments may be combined with each other.
[0060] The following detailed descriptions are all exemplary descriptions, aiming to provide further detailed descriptions of the present invention. Unless otherwise specified, all technical terms used in the present invention have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present application belongs. The terms used in the present invention are only for describing specific embodiments, and are not intended to limit the exemplary embodiments according to the present invention.
[0061] Embodiment 1
[0062] As Figures 1-4As shown in the figure, the present invention proposes a multimode fiber optic image transmission method, aiming to solve the problems of insufficient anti-interference ability, poor image reconstruction quality, and weak generalization ability in multimode fiber optic image transmission in complex environments through a neural network model based on a sliding window self-attention mechanism. The following are the specific implementation steps of this method:
[0063] Obtain the original image of the target area through the image acquisition device of the power transmission line operation robot. The image acquisition device is usually a high-resolution camera, installed at the end or key parts of the robot, and can take multi-angle pictures of the key areas of the power transmission line (such as wire connection points, insulators, towers, etc.). The collected original image contains rich line state information, providing basic data for subsequent image transmission and analysis.
[0064] Convert the collected original image into an optical signal and transmit it through a multimode fiber optic. Due to the characteristics of multimode fiber optic, the optical signal will undergo mode coupling and dispersion during transmission, resulting in the disruption of the original image information and forming a speckle image with spatial perturbation characteristics. A speckle image is a random light intensity distribution pattern, which contains the encoded information of the original image, but the content of the original image cannot be recognized when directly observed.
[0065] Before inputting the speckle image into the neural network model, a series of data preprocessing operations can be performed to enhance the model's adaptability to different perturbation modes. The specific operations include:
[0066] Directional data augmentation: For the characteristics of the power transmission line scene, apply directional data augmentation to the original speckle image. For example, add strip noise patterns along the axis of the insulator to simulate wire discharge interference; superimpose metal reflection artifacts in the tower connection area to enhance the model's adaptability to highly reflective surfaces; generate a dynamic blur kernel according to the wire sag parameters to simulate image jitter caused by high-altitude wind vibration.
[0067] Feature domain alignment: Perform feature domain alignment operations on the enhanced speckle image. Extract the main frequency components of the speckle image through frequency domain analysis and establish a mapping relationship with the fiber optic perturbation types. According to the main frequency distribution characteristics, automatically select the hierarchical processing path of the sliding window self-attention block and align the speckle distribution rules under different perturbation modes in the feature embedding space.
[0068] Input the preprocessed speckle image into a neural network model constructed based on the sliding window self-attention mechanism for image reconstruction. This neural network model includes the following key modules:
[0069] Block partitioning module: Divide the input image into multiple small blocks, such as local blocks of 4×4 pixels. This partitioning method facilitates subsequent feature extraction and self-attention calculation.
[0070] Linear embedding module: Maps the pixel values of each local block to high-dimensional feature vectors, providing richer feature representations for the subsequent self-attention mechanism.
[0071] Multi-layer sliding window self-attention block: The core module, which calculates local self-attention features within a fixed-size window through the alternately set window multi-head self-attention (W-MSA) and sliding window multi-head self-attention (SW-MSA) mechanisms, and realizes feature interaction between adjacent windows through the sliding window mechanism. The specific working process is as follows:
[0072] In the window multi-head self-attention layer, the feature map is divided into 7×7 fixed windows, and multi-head self-attention calculations are performed inside each window to capture local features.
[0073] In the sliding window multi-head self-attention layer, after performing cyclic shift operations on the feature map, the windows are re-divided, and a masking mechanism is used to suppress attention calculations in non-adjacent regions to avoid information interference.
[0074] In the multi-layer perceptron (MLP), a two-layer fully connected structure is adopted for non-linear feature transformation to further optimize the feature representation.
[0075] After each calculation step, the original feature information is retained through skip connections to ensure that the model can fully utilize the valid information in the input data.
[0076] Average pooling layer: Compresses the deep features in the channel dimension to reduce the number of parameters.
[0077] Fully connected layer: Reconstructs the compressed feature vector into an output image of the target size to complete image reconstruction.
[0078] The reconstructed image is output to the transmission line status analysis system for real-time monitoring. The status analysis system can automatically analyze the reconstructed image, identify key components of the transmission line (such as wire connection points, insulators, towers, etc.), and detect abnormal states of the equipment (such as wire wear, insulator discharge marks, tower tilt, etc.) by comparing with the historical image database. Combining multi-angle reconstructed images, the system can also generate a 3D line status model to provide a more intuitive reference for the maintenance and management of the transmission line. According to the real-time monitoring results, the system can generate line maintenance decision suggestions to help maintenance personnel timely discover and handle potential problems, ensuring the safe and stable operation of the transmission line.
[0079] During the multi-mode fiber transmission process, when it is detected that the fiber transmission environment parameters (such as temperature, vibration frequency, electromagnetic interference intensity, fiber bending curvature, etc.) exceed the preset thresholds, the system can dynamically adjust the calculation parameters of the sliding window self-attention block to optimize the image reconstruction effect. The specific adjustment methods include:
[0080] Increase the window division size proportionally according to the temperature change range to compensate for the mode coupling interference caused by the refractive index fluctuation of the optical fiber.
[0081] When the vibration frequency exceeds the preset frequency, automatically reduce the displacement step size of the sliding window to enhance the local feature retention ability.
[0082] In a strong electromagnetic interference environment, activate the noise suppression coefficient of the shielding mechanism to the highest level to filter the abnormal attention weights in non-adjacent regions.
[0083] Based on the real-time collected optical fiber bending curvature data, adaptively adjust the feature fusion ratio between the sliding window self-attention layer and the multi-layer perceptron.
[0084] To improve the generalization ability and anti-interference ability of the model, a hybrid training strategy is used to train the neural network model. The specific steps are as follows:
[0085] Construct a training data set containing various perturbation conditions, and simulate the bending, vibration and temperature change perturbations of multimode optical fibers by adding Gaussian noise to the real and imaginary parts of the transmission matrix respectively.
[0086] Input the speckle image pairs under different perturbation conditions into the model for joint training, so that the model can learn the common transmission characteristics under different perturbation modes.
[0087] Use the Adam optimizer to dynamically adjust the model parameters, including the number of heads, window size and network depth of the sliding window self-attention layer, etc.
[0088] During the training process, synchronously verify the image reconstruction quality of the model under unknown perturbation configurations to ensure that the model not only performs well on the training data, but also has strong generalization ability in practical applications.
[0089] Embodiment 2
[0090] The present invention also proposes a multimode optical fiber image transmission system, which is implemented based on the above method and can efficiently complete the full-process tasks of image acquisition, transmission, reconstruction and real-time monitoring. The following are the specific components and functions of the system:
[0091] The image acquisition module is configured at the end of the transmission line operation robot for acquiring high-resolution line state images. This module usually includes a high-resolution camera, an image sensor and other auxiliary devices, and can take multi-angle pictures of the key parts of the transmission line to obtain the original images containing rich line state information. The collected image data is the basic input of the whole system, and its quality directly affects the subsequent image transmission and reconstruction effects.
[0092] The fiber optic transmission module includes a multimode fiber optic channel and its supporting optoelectronic conversion device. The multimode fiber is used to transmit the image signal in the form of an optical signal. Although speckle phenomenon will occur during the transmission process, the original image information can be restored through the subsequent image reconstruction module. The supporting optoelectronic conversion device is used to convert the collected image signal into an optical signal and convert the optical signal into an electrical signal at the receiving end for subsequent processing. This module is the core transmission part of the system, and its performance determines the efficiency and quality of image transmission.
[0093] The deep learning reconstruction module includes a neural network model constructed based on the sliding window self-attention mechanism, which is used to restore the original line image from the speckle image. This module is the intelligent core of the system. Through the optimized neural network architecture and training strategy, it can efficiently reconstruct high-quality original images from complex speckle images. The specific technical details are the same as the image reconstruction method in Embodiment 1, including key modules such as a block partitioning module, a linear embedding module, a multi-layer sliding window self-attention block, an average pooling layer, and a fully connected layer, as well as a dynamic adjustment mechanism and data preprocessing operations.
[0094] The state analysis module includes an image recognition algorithm and a trend prediction model, which are used to automatically detect line equipment defects and generate maintenance suggestions. This module can automatically analyze the reconstructed image, identify key components of the transmission line (such as wire connection points, insulators, towers, etc.), and detect abnormal states of the equipment (such as wire wear, insulator discharge traces, tower tilt, etc.) by comparing with the historical image database. Combining multi-angle reconstructed images, the system can also generate a three-dimensional line state model, providing a more intuitive reference for the maintenance and management of the transmission line. According to the real-time monitoring results, the system can generate line maintenance decision suggestions to help maintenance personnel discover and handle potential problems in a timely manner, ensuring the safe and stable operation of the transmission line.
[0095] The system also has a dynamic adjustment mechanism, which can monitor the fiber optic transmission environment parameters (such as temperature, vibration frequency, electromagnetic interference intensity, fiber bending curvature, etc.) in real time during the multimode fiber optic transmission process. When it detects that the parameter exceeds the preset threshold, the system can automatically adjust the calculation parameters of the sliding window self-attention block to optimize the image reconstruction effect. The specific adjustment method is the same as the dynamic adjustment mechanism in Embodiment 1, including adjusting the window division size according to the temperature change, adjusting the sliding window displacement step according to the vibration frequency, and activating the noise suppression coefficient in a strong electromagnetic interference environment. This dynamic adjustment mechanism enables the system to always maintain high-efficiency image reconstruction ability in a complex transmission environment, further improving the practicability and reliability of the system.
[0096] To improve the model's adaptability to different perturbation patterns, the system can also be configured with a data preprocessing and enhancement module. This module can perform directional data enhancement and feature domain alignment operations on the original speckle image before image reconstruction. Specific operations include adding strip noise patterns along the axial direction of the insulator, superimposing metal reflection artifacts in the tower connection area, generating dynamic blur kernels according to the sag parameters of the wire, etc., to simulate various interference situations in actual applications. In addition, the main frequency components of the speckle image are extracted through frequency domain analysis, and a mapping relationship with the fiber perturbation type is established to further optimize the input data of the model and improve the accuracy and efficiency of image reconstruction.
[0097] The entire multimode fiber image transmission system is efficiently integrated through modular design, and data interaction and collaborative work are carried out between modules through standardized interfaces. The system can be flexibly configured and optimized according to actual application requirements, such as adjusting the parameter settings of the modules, adding or reducing certain functional modules, etc. In addition, the system also has good scalability and can be seamlessly docked with the existing power system monitoring platform to provide comprehensive support for the intelligent operation and maintenance of the power system.
[0098] Embodiment 3
[0099] As Figure 1 and Figure 2 shown, the laser beam-loaded image is scrambled into speckles after being transmitted through the multimode fiber, and the original image is reconstructed through the trained model to achieve multimode fiber image transmission.
[0100] As Figure 3 shown, the model in the method involved in the present invention will go through the following modules: PatchPartition, Linear Embeding, Swin Transformer Blocks, Avgpooling, and Fully Connected Layer. PatchPartition changes the dimension of the input image from 224×224 to 56×56×16. Specifically, the original image is decomposed into 16 dimensions on 4×4 small blocks. There are 56×56 small blocks on the 224×224 image, so the image is decomposed into 56×56×16.
[0101] In the sliding window self-attention block, window multi-head self-attention (W-MSA) calculates multi-head self-attention within a divided window of size 7×7. The sliding window self-attention block includes the following parts: layer normalization (LN), window multi-head self-attention (W-MSA), shifted window multi-head self-attention (SW-WSA), multi-layer perceptron (MLP), and skip connections are used after each module. The advantage of dividing the window is that the computational complexity does not grow quadratically with the number of image pixels, but linearly with the number of image pixels. Therefore, its complexity is much lower than that of ViT. When performing window multi-head self-attention calculation, self-attention is only calculated within each window, and there is no information interaction between the tokens in different windows. Therefore, shifted window multi-head self-attention is introduced to enable information interaction between adjacent windows to improve the modeling ability. The windows and shifted windows cooperate with each other, so the sum of the number of windows and shifted windows is always even. In the specific implementation of the sliding window, a cyclic shift operation is used, but it may cause non-adjacent tokens, such as tokens on the left and right sides of the image, to be divided into the same window. In this case, the tokens are not in adjacent positions and should not perform self-attention operations, resulting in information interference. To solve this problem, a mask calculation is introduced to make the tokens that should not calculate self-attention output close to zero after passing through the SoftMax activation function, which can isolate the calculation of non-adjacent tokens. The calculation expression of the sliding window self-attention block is shown in the equation. In each layer, the output features of (S)W-MSA (window multi-head self-attention calculation) and MLP (multi-layer perceptron calculation) are as follows
[0102]
[0103] The sliding window self-attention model provides a general backbone network based on self-attention through delicate design and efficient engineering implementation. The method proposed in this patent replaces the combination of convolutional layers and activation layers with sliding window self-attention blocks on the basis of the Real-Valued ANN network. The feature maps after passing through the sliding window self-attention blocks are averaged and pooled to merge channels, reducing the number of parameters. Finally, a fully connected layer is used to transform the features into the size of the target image.
[0104] As Figure 4As shown in the figure, in the hybrid training, speckle image pairs from multiple groups of configurations (these groups are called known configuration groups) jointly train the method proposed in this patent. The groups that do not participate in the training are called unknown configuration groups. The speckles from unknown configurations are used to reconstruct images and estimate the generalization ability of the training model, that is, the ability of the algorithm to resist different perturbation states of optical fibers. On the unknown configuration groups, images can be reconstructed, which proves that this training method, which we call hybrid training, can help neural networks have the ability to resist MMF perturbations in natural scene image transmission.
[0105] Specific image transmission process:
[0106] Image acquisition: The live working robot for transmission lines is equipped with a high-resolution image acquisition device, such as a high-definition camera. During operation, key parts of the transmission line, such as wire connection points, insulators, and towers, are photographed from multiple angles to obtain original images containing rich line state information, providing basic data for subsequent analysis.
[0107] Multi-mode fiber transmission: The collected original images are transmitted through multi-mode fibers in the form of optical signals. There are problems such as dispersion and mode coupling in multi-mode fibers, which will cause the optical signals of the original images to interfere with each other during transmission, resulting in a random speckle pattern formed by the outgoing light at the far end and the original image information being disrupted. Moreover, complex environmental factors around the transmission line, such as temperature changes, electromagnetic interference, and mechanical vibrations, will affect the transmission characteristics of multi-mode fibers, causing the speckle pattern to change with time and the environment, increasing the difficulty of image transmission and restoration.
[0108] Image reconstruction based on deep learning
[0109] Model construction: A neural network model based on self-attention is used to reconstruct the speckle images. The model contains multiple key modules. The input speckle image first enters the block partitioning module, which decomposes the 224×224 image into 16 dimensions on 4×4 small blocks to obtain a feature map of 56×56×16, facilitating subsequent processing. Then, the linear embedding module adjusts the feature map dimension to make it more suitable for self-attention calculation. The core sliding window self-attention block contains layer normalization, window multi-head self-attention (W-MSA), sliding window multi-head self-attention (SW-WSA), and multi-layer perceptron (MLP). Each part works together through skip connections. The window multi-head self-attention calculates the multi-head self-attention within a 7×7 window to capture local features; the sliding window multi-head self-attention enables information interaction between adjacent windows, enhancing the modeling ability and effectively learning long-range dependencies in the sequence. Finally, after the average pooling layer merges channels and reduces the number of parameters, the full connection layer transforms the features into the target image size and outputs the reconstructed image.
[0110] Model Training: The hybrid training method is used to improve the model performance. Based on the analysis of the experimentally collected data, the transmission matrix is randomly generated, and random noise with zero-mean normal distribution is added to the real and imaginary parts of the transmission matrix to simulate the changes in the transmission characteristics of multimode optical fibers under different perturbations, and a simulation data set is produced. The speckle image pairs (i.e., the known configuration groups) under various different perturbations are mixed to train the model, enabling the model to learn the common transmission characteristics of multimode optical fibers under different perturbations. The training process is based on the Python and PyTorch frameworks to build the environment. The Adam optimizer is used to continuously adjust the model parameters, such as the number of attention heads, the number of window self-attention layers and sliding window self-attention layers, etc., to minimize the loss function and improve the anti-perturbation ability and image reconstruction quality of the model.
[0111] Image Application and Analysis: The reconstructed image is used for the condition monitoring and fault diagnosis of transmission lines. Through the image recognition algorithm, it can automatically detect whether there are defects in line equipment, such as wire wear, broken strands, insulator breakage, discharge traces, and abnormal conditions such as tower inclination. Combining historical image data and real-time monitoring results, it is also possible to conduct trend analysis on the line condition, predict potential faults, provide a scientific basis for the maintenance and management of transmission lines, and ensure the safe and stable operation of transmission lines.
[0112] Example 4
[0113] As Figure 5 shown, an electronic device includes a processor and a memory. The processor is used to execute the computer program stored in the memory to implement a multimode optical fiber image transmission method.
[0114] The present invention also provides an electronic device 100 for implementing the multi-mode fiber optic image transmission method in the above embodiment; the electronic device 100 includes a memory 101, at least one processor 102, a computer program 103 stored in the memory 101 and executable on the at least one processor 102, and at least one communication bus 104. The memory 101 can be used to store the computer program 103. The processor 102 realizes the steps of the multi-mode fiber optic image transmission method in Embodiment 1 by running or executing the computer program stored in the memory 101 and calling the data stored in the memory 101. The memory 101 may mainly include a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.); the data storage area can store data created according to the use of the electronic device 100 (such as audio data, etc.). In addition, the memory 101 may include non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. The at least one processor 102 may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The processor 102 may be a microprocessor or the processor 102 may also be any conventional processor, etc. The processor 102 is the control center of the electronic device 100 and connects various parts of the entire electronic device 100 through various interfaces and lines. The memory 101 in the electronic device 100 stores multiple instructions to implement a multi-mode fiber optic image transmission method.
[0115] Embodiment 5
[0116] A computer-readable storage medium stores at least one instruction, and when the at least one instruction is executed by a processor, a multi-mode fiber optic image transmission method is implemented.
[0117] If the integrated module / unit of the electronic device 100 is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above-mentioned embodiment methods of the present invention, it can also be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-mentioned various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory and read-only memory (ROM, Read-Only Memory). Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code. The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the processes and / or blocks in the flowchart and / or block diagram can also be implemented. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the specified functions in the process Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks Figure 1 one process or multiple processes and / or blocksFigure 1 Steps of functions specified in one or more boxes.
[0118] As is known by common technical knowledge, the present invention can be implemented by other embodiments that do not depart from its spirit or essential features. Therefore, the above-disclosed embodiments are illustrative in all aspects and not exclusive. All changes within the scope of the present invention or within the scope equivalent to the present invention are encompassed by the present invention.
Claims
1. A multimode fiber optic image transmission method, characterized in that, Including the following steps, Obtain the original image of the target area through the image acquisition device of the transmission line operation robot; Convert the original image into an optical signal and transmit it through a multimode optical fiber to generate a speckle image with spatial perturbation characteristics; Input the speckle image into a neural network model constructed based on the sliding window self-attention mechanism for image reconstruction, where the neural network model includes a block partitioning module, a linear embedding module, multiple-layer sliding window self-attention blocks, and an image reconstruction module; Output the reconstructed image to the transmission line status analysis system for real-time monitoring; The sliding window self-attention block calculates local self-attention features within a fixed-size window through alternately set window multi-head self-attention and sliding window multi-head self-attention mechanisms, and realizes feature interaction between adjacent windows through the sliding window mechanism.
2. The multimode fiber optic image transmission method according to claim 1, characterized in that, The processing flow of the neural network model includes: Divide the input image into multiple local blocks of 4×4 pixels through the block partitioning module; Use the linear embedding module to map the pixel values of each local block into high-dimensional feature vectors; After being alternately processed by at least two groups of sliding window self-attention blocks, each group includes a combination structure of a window multi-head self-attention layer, a sliding window multi-head self-attention layer, and a multi-layer perceptron; Compress the deep features in the channel dimension through an average pooling layer; Finally, the fully connected layer reconstructs the compressed feature vectors into an output image of the target size.
3. A multimode fiber optic image transmission method according to claim 1, characterized in that, The working process of the sliding window self-attention block includes: Divide the feature map into 7×7 fixed windows in the window multi-head self-attention layer, and perform multi-head self-attention calculations within each window; Perform a cyclic shift operation on the feature map in the sliding window multi-head self-attention layer and then re-divide the windows, and suppress the attention calculations in non-adjacent regions through a masking mechanism; Adopt a two-layer fully connected structure in the multi-layer perceptron for non-linear feature transformation; Keep the original feature information through skip connections after each calculation step.
4. A multimode fiber optic image transmission method according to claim 1, characterized in that The training method of the neural network model includes: Construct a training data set containing various perturbation conditions, and simulate the bending, vibration, and temperature change perturbations of the multimode optical fiber by adding normal distribution noise to the real and imaginary parts of the transmission matrix respectively; Adopt a hybrid training strategy to input pairs of speckle images under different perturbation conditions into the model for joint training; Dynamically adjust the model parameters through the Adam optimizer, including the number of heads, window size, and network depth of the sliding window self-attention layer; Synchronously verify the image reconstruction quality of the model under unknown perturbation configurations during the training process.
5. A multimode fiber optic image transmission method according to claim 1, characterized in that, The real-time monitoring includes: Automatically identify key components of the transmission line in the reconstructed image, including wire connection points, insulators, and tower structures; Detect abnormal device states by comparing with the historical image database, including wire wear, insulator discharge traces, and tower tilt; Generate a three-dimensional line status model by combining multi-angle reconstructed images; Generate line maintenance decision suggestions according to the real-time monitoring results.
6. A multimode fiber optic image transmission method according to claim 1, characterized in that, During the multimode optical fiber transmission process: When it is detected that the optical fiber transmission environment parameters exceed the preset threshold, dynamically adjust the calculation parameters of the sliding window self-attention block, including: Increase the window division size proportionally according to the temperature change range to compensate for the mode coupling interference caused by the refractive index fluctuation of the optical fiber. When the vibration frequency exceeds the preset frequency, automatically reduce the displacement step size of the sliding window to enhance the local feature retention ability. In a strong electromagnetic interference environment, activate the noise suppression coefficient of the shielding mechanism to the highest level to filter the abnormal attention weights in non-adjacent regions. Based on the real-time collected optical fiber bending curvature data, adaptively adjust the feature fusion ratio between the sliding window self-attention layer and the multi-layer perceptron.
7. A multimode fiber optic image transmission method according to claim 1, characterized in that, Perform the following data preprocessing before image reconstruction: For the scene characteristics of the transmission line, apply directional data augmentation to the original speckle image: add strip noise patterns along the axial direction of the insulator to simulate the interference of wire discharge; superimpose metal reflection artifacts in the tower connection area to enhance the model's adaptability to high-reflection surfaces; generate a dynamic blur kernel according to the wire sag parameter to simulate the image jitter caused by high-altitude wind vibration. Perform a feature domain alignment operation on the enhanced speckle image: extract the main frequency components of the speckle image through frequency domain analysis and establish a mapping relationship with the optical fiber perturbation type; automatically select the hierarchical processing path of the sliding window self-attention block according to the main frequency distribution characteristics. Align the speckle distribution laws under different perturbation modes in the feature embedding space.
8. A multimode fiber optic image transmission system, based on the multimode fiber optic image transmission method according to any one of claims 1-7, characterized in that, Including, An image acquisition module, configured at the end of the transmission line operation robot, for acquiring high-resolution line state images. An optical fiber transmission module, including a multi-mode optical fiber channel and its supporting optoelectronic conversion device, for converting optical signals into speckle images with spatial perturbation characteristics. A deep learning reconstruction module, including a neural network model constructed based on the sliding window self-attention mechanism, for recovering the original line image from the speckle image. A state analysis module, including an image recognition algorithm and a trend prediction model, for automatically detecting line equipment defects and generating maintenance suggestions.
9. An electronic device, characterized in that, Including a processor and a memory, the processor is used to execute the computer program stored in the memory to implement a multi-mode optical fiber image transmission method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores at least one instruction, and when the at least one instruction is executed by the processor, it implements a multi-mode optical fiber image transmission method as described in any one of claims 1 to 7.
Citation Information
Cited By
Visitor identity automatic verification method and system based on intelligent access control
CN120877417A
Industrial equipment fault detection system and method based on computer vision
CN121095185A
A computer vision-based industrial equipment failure detection system and method
CN121095185B