A visible light communication decoding method and system based on deep learning

By constructing a CAE-LSTM-Net neural network based on deep learning, and combining channel attention and long short-term memory mechanisms, inter-symbol interference in visible light communication is mitigated, communication rate is improved and bit error rate is reduced, thus solving the performance degradation problem caused by inter-symbol interference in existing technologies.

CN116667926BActive Publication Date: 2026-06-02SOUTH CHINA UNIV OF TECH

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SOUTH CHINA UNIV OF TECH
Filing Date
2023-05-31
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Existing visible light communication technologies suffer from severe inter-symbol interference in high-speed communication, resulting in low communication rates and high bit error rates. Existing solutions are unable to effectively mitigate the performance degradation caused by such interference.

Method used

We employ a deep learning-based CAE-LSTM-Net neural network, combined with channel attention and long short-term memory mechanisms, to construct an encoder and equalizer. Through image preprocessing and sub-image segmentation, we decode the RGB channel binary data of each sub-image, mitigating inter-symbol interference and improving communication speed.

Benefits of technology

It achieves high data rate visible light communication, reduces bit error rate, fully extracts high-dimensional features of received images, and effectively mitigates performance degradation caused by inter-symbol interference.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116667926B_ABST
    Figure CN116667926B_ABST
Patent Text Reader

Abstract

The application discloses a visible light communication decoding method and system based on deep learning, and relates to the field of visible light communication.The CAE-LSTM-Net neural network is constructed;each frame of the collected video is subjected to image preprocessing and is segmented into a series of subimages;each subimage is decoded in one time step by using the constructed CAE-LSTM-Net neural network, and specifically, the subimage is input into the trained CAE-LSTM-Net neural network in time sequence, the CAE-LSTM-Net neural network outputs the category corresponding to the subimage, and the binary data of the RGB channel contained by the symbol in the subimage is obtained, so that the communication rate can be improved, the high-dimensional features of the received image can be fully extracted, the performance deterioration caused by the inter-symbol interference in space and time is relieved, and the cost is low.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of visible light communication technology, and in particular to a visible light communication decoding method and system based on deep learning. Background Technology

[0002] In recent years, visible light communication (VLC), as a short-range communication technology, has attracted widespread attention due to its high bandwidth and low latency; it is even considered one of the next wireless communication technologies (6G). VLC technology offers opportunities for applications such as indoor positioning, mobile payment, and navigation. In most research, VLC requires photodiodes as dedicated receivers, which has delayed its commercialization. To reduce its reliance on dedicated devices, optical camera communication (OCC), a special type of VLC, has emerged, primarily using charge-coupled devices (CCDs) or complementary metal-oxide-semiconductor (CMOS) cameras as receivers. Currently, CMOS-based commercial optical cameras are widely used in smartphones, providing favorable conditions for the development of OCC technology.

[0003] To achieve faster communication rates, previous researchers have improved the black-and-white LEDs at the transmitting end by replacing them with colored RGB-LEDs. This inherently gives the communication framework the potential to achieve faster speeds at the modulation end. Currently, the CSK modulation scheme recommended by the IEEE standard [IEEE Standard 802.15.7-2011; IEEE Standard for Local and Metropolitan Area Networks-Part 15.7: Short-Range Wireless Optical Communication Using Visible Light. IEEE: New York, NY, USA, 2011; pp.1-309.] is widely used in optical camera communication based on RGB-LEDs. In previous studies, such as [Hu, P.; Pathak, P.H.; Feng, X.; Fu, H.; Mohapatra, P. Colorbars: Increasing data rate of led-to-camera communication using color shift keying. In Proceedings of the 11th ACM Conference on Emerging Networking Experiments and Technologies, Heidelberg, Germany, 1-4 December 2015; pp.1-13.], the intensity of each channel of the RGB channel of the transmitting LED was modulated to mix the colors of the three channels to generate eight colors. At the decoding end, the data corresponding to these eight colors were decoded to recover the binary data of each channel in the RGB channel. In [47 - kbit / s RGB - LED light - based optical camera communication based on 2D-CNN and XOR-based data loss compensation.], an optical camera communication decoding framework based on RGB - LED light and CSK modulation scheme is proposed, which utilizes a two-dimensional convolutional neural network. At the same time, the scheme designs a special data packet construction form, which divides a single data packet into several parts and uses the proposed XOR mechanism to perform data recovery.Meanwhile, a patent titled "A Rolling Shutter Image Processing Method Based on Visible Light Communication" (publication number: CN115719359A) uses a black and white LED light to modulate the signal at the transmitting end, and extracts individual bright and dark symbol stripes by detecting the boundary, and then uses a threshold method to fit the symbol gray value for decoding.

[0004] However, the aforementioned scheme [Hu, P.; Pathak, PH; Feng, X.; Fu, H.; Mohapatra, P. Colorbars: Increasing data rate of led-to-camera communication using colorshift keying. In Proceedings of the 11th ACM Conference on Emerging Networking Experiments and Technologies, Heidelberg, Germany, 1-4 December 2015; pp. 1-13] has a relatively low implemented data rate (maximum only 5.2 Kbps), and its ability to cope with the strong inter-symbol interference that occurs in high-speed visible light communication is poor. In high-speed visible light communication, noise from the rolling shutter effect of CMOS sensors and overlapping exposure times will cause interference between color symbols. This inter-symbol interference will be very serious in high-speed visible light communication.

[0005] The Reed-Solomon coding used in this scheme cannot effectively cope with the performance degradation caused by such interference. The aforementioned scheme, "[47-kbit / s RGB-LED lamp-based optical camera communication based on 2D-CNN and XOR-based data loss compensation]", utilizes the XOR mechanism, which does not conform to the real communication environment because it introduces a large amount of redundant data into the data packets to mitigate data loss. In reality, most of the data transmitted by users is random; furthermore, this scheme uses a two-dimensional convolutional neural network to extract features from the sampled signal, which is essentially just an equalization method and can be replaced by a one-dimensional fully connected network, failing to fully utilize the potential of neural networks in high-dimensional feature extraction. The aforementioned scheme (publication number "CN115719359A", titled "A Rolling Shutter Image Processing Method Based on Visible Light Communication") uses a black and white LED lamp as the transmitter. Under the same symbol width at the receiver, a single symbol contains only one bit of data, which means that its achievable communication rate is significantly lower than that of the optical camera communication scheme using an RGB-LED lamp as the transmitter. Summary of the Invention

[0006] The purpose of this invention is to provide a visible light communication decoding method and system based on deep learning, which can improve the communication rate, fully extract the high-dimensional features of the received image, and learn in space and time to mitigate the performance degradation caused by inter-symbol interference.

[0007] To achieve the above objectives, the present invention provides the following solution:

[0008] A deep learning-based visible light communication decoding method includes the following steps:

[0009] Construct a CAE-LSTM-Net neural network;

[0010] Each frame of the captured video is preprocessed and segmented into a series of sub-images;

[0011] The constructed CAE-LSTM-Net neural network is used to decode each sub-image within one time step. Specifically, the sub-images are input into the trained CAE-LSTM-Net neural network in chronological order. The CAE-LSTM-Net neural network outputs the category corresponding to the sub-image, and obtains the binary data of the RGB channels contained in the symbol in the sub-image.

[0012] Furthermore, the constructed CAE-LSTM-Net neural network includes: an encoder neural network based on the channel attention mechanism and an equalizer neural network based on the long short-term memory mechanism.

[0013] Furthermore, the encoder neural network based on the channel attention mechanism includes a convolutional layer, a multi-channel high-level feature extraction unit, and a recalibration unit. The convolutional layer and the multi-channel high-level feature extraction unit are used to extract multi-channel high-level features of the input image to obtain a feature map. The recalibration unit is used to determine the weights between different channels of the feature map so as to recalibrate the channels of the feature map before channel fusion.

[0014] Furthermore, the equalizer neural network based on the long short-term memory mechanism includes a long short-term memory unit, and the input-output relationship in the long short-term memory unit is as follows:

[0015]

[0016] in, This is the weight matrix to be learned, where σ(.) represents the sigmoid activation function, φ(.) represents the tanh activation function, ⊙ represents the element-wise dot product, and the forget gate is used. Input gate and output gate The bias is the parameter to be learned. This is the final output corresponding to this input. Update the cell state value. This is the unit state.

[0017] Furthermore, before the step of performing image preprocessing on each frame of the acquired video and dividing it into a series of sub-images, the method further includes video acquisition, wherein the video acquisition method is to record the video in the form of 60 frames per second and a resolution of 1080×1920.

[0018] Furthermore, the process of preprocessing each frame of the acquired video and segmenting it into a series of sub-images specifically includes:

[0019] Frame extraction, which is to extract each frame of the video as an image frame to be decoded;

[0020] Column selection involves choosing a suitable pixel column based on the average gray level of the columns, and then expanding the signal of that pixel column into an image by copying it.

[0021] Amplitude normalization is the process of normalizing the pixel intensity of the RGB channels of an image.

[0022] Gamma correction is the process of compensating for the signal intensity of the RGB channels in an image using gamma correction.

[0023] Perform column selection again, that is, select the center column in the compensated image;

[0024] Data packet header location, which is to obtain the header position by fitting the header data using a threshold method;

[0025] Image segmentation is the process of dividing an image into multiple sub-images, starting from the obtained head position.

[0026] Furthermore, the method also includes:

[0027] Time-domain recovery involves restoring the time sequence of the obtained binary data.

[0028] This invention also provides a deep learning-based visible light communication decoding system, comprising:

[0029] Modules for building CAE-LSTM-Net neural networks;

[0030] The image preprocessing module is used to preprocess each frame of the captured video and divide it into a series of sub-images;

[0031] The offline decoding module is used to decode each sub-image within one time step using the constructed CAE-LSTM-Net neural network. Specifically, it includes: inputting the sub-images into the trained CAE-LSTM-Net neural network in chronological order; the CAE-LSTM-Net neural network outputting the category corresponding to the sub-image; and obtaining the binary data of the RGB channels contained in the symbol in the sub-image.

[0032] Furthermore, the system also includes a video acquisition device, which comprises:

[0033] The transmitting end includes a programmable gate array and RGB-LEDs for emitting visible light;

[0034] The receiving end includes a mobile phone with an optical camera based on a CMOS sensor for video acquisition. The video acquisition method is to record the video at 60 frames per second with a resolution of 1080×1920.

[0035] Furthermore, the image preprocessing module includes:

[0036] The frame extraction module is used to extract each frame of the video as an image frame to be decoded.

[0037] The column selection module is used to select a suitable pixel column by means of the average gray level of the column, and to expand the signal of the column into an image by means of copying;

[0038] The amplitude normalization module is used to normalize the pixel intensity of the RGB channels of the image;

[0039] The gamma correction module is used to compensate for the signal intensity of the RGB channels of an image using gamma correction.

[0040] The center column selection module is used to select the center column in the compensated image;

[0041] The data packet header localization module is used to fit the header data using a threshold method to obtain the header position.

[0042] The image segmentation module is used to segment the image into multiple sub-images, starting from the obtained head position.

[0043] According to specific embodiments provided by the present invention, the following technical effects are disclosed: The deep learning-based visible light communication decoding method and system provided by the present invention preprocesses each received frame, then sequentially segments it into a series of sub-images, and decodes one sub-image within one time step using a CAE-LSTM-Net neural network. Unlike existing OCC systems based on the rolling shutter effect of CMOS sensors, the present invention does not require traditional sampling, and is therefore less susceptible to sampling offset caused by inter-symbol interference. Furthermore, the present invention, for the first time, integrates spatial equalization and temporal equalization functions in the equalizer of the CAE-LSTM-Net neural network to further mitigate the performance degradation of the optical communication system caused by inter-symbol interference. The present invention can achieve high data rates, fully extract high-dimensional features of the received image, and significantly reduce the bit error rate. Attached Figure Description

[0044] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0045] Figure 1 This is a flowchart of the visible light communication decoding method based on deep learning according to the present invention;

[0046] Figure 2 This is a schematic diagram of the data packet construction method in an embodiment of the present invention;

[0047] Figure 3 This is a schematic diagram of the CAE-LSTM-Net neural network in an embodiment of the present invention;

[0048] Figure 4 This is a schematic diagram of the encoder neural network based on the channel attention mechanism in an embodiment of the present invention;

[0049] Figure 5 This is a schematic diagram of the structure of the Residual Unit and Recalibration Unit in an embodiment of the present invention;

[0050] Figure 6 This is a schematic diagram of a long short-term memory unit in an embodiment of the present invention;

[0051] Figure 7 This is a structural diagram of the visible light communication decoding system in an embodiment of the present invention. Detailed Implementation

[0052] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0053] The purpose of this invention is to provide a visible light communication decoding method and system based on deep learning, which can improve the communication rate, fully extract the high-dimensional features of the received image, and learn in space and time to mitigate the performance degradation caused by inter-symbol interference.

[0054] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0055] Example 1

[0056] Figure 1-7 As shown in the figure, an embodiment of the present invention provides a visible light communication decoding method based on deep learning, which includes the following steps:

[0057] Construct a CAE-LSTM-Net neural network;

[0058] Each frame of the captured video is preprocessed and segmented into a series of sub-images;

[0059] The constructed CAE-LSTM-Net neural network decodes each sub-image within one time step (Trained CAE-LSTM-Net for Decoding). Specifically, this includes: inputting the sub-images into the trained CAE-LSTM-Net neural network in chronological order; the CAE-LSTM-Net neural network outputting the category corresponding to the sub-image; and obtaining the binary data of the RGB channels contained in the symbol in the sub-image.

[0060] Figure 1 The CAE-LSTM-Net neural network training step (Initial CAE-LSTM-Net for Training) is required before the system starts working; this step is used for training. Figure 1 The neural network used in (g) is also required before the system starts working. This step is used to assess whether the system has been built well and meets the bit error rate requirements of the application.

[0061] The training process for the CAE-LSTM-Net neural network is as follows: The number of training epochs is set, for example, to 25, and an adaptive momentum estimation (Adam) optimizer is used after calculating the Softmax loss. To effectively search for the global optimum, some hyperparameters in the LSTM, such as the regularization term, the number of network blocks and convolutional kernels, and the number of hidden neurons, will vary under different symbol rates.

[0062] like Figure 2 As shown, in this OCC system, each data packet is transmitted three times within (1 / 60) second to ensure that it is fully recorded at least once in the captured image frame. Each data packet transmission has a 16-bit header and a payload.

[0063] In a further embodiment, the constructed CAE-LSTM-Net neural network includes: an encoder neural network based on a channel attention mechanism and an equalizer neural network based on a long short-term memory mechanism.

[0064] In a further embodiment, the encoder neural network based on the channel attention mechanism includes a convolutional layer, a multi-channel high-level feature extraction unit (Residual Unit), and a recalibration unit. The convolutional layer and the multi-channel high-level feature extraction unit are used to extract multi-channel high-level features of the input image to obtain a feature map. The recalibration unit is used to determine the weights between different channels of the feature map so as to recalibrate the channels of the feature map before channel fusion.

[0065] In a further embodiment, such as Figure 6As shown, the equalizer neural network based on the long short-term memory mechanism includes a long short-term memory unit, and the input-output relationship in the long short-term memory unit is as follows:

[0066]

[0067] in, This is the weight matrix to be learned, where σ(.) represents the sigmoid activation function, φ(.) represents the tanh activation function, ⊙ represents the element-wise dot product, and the forget gate is used. Input gate and output gate The bias is the parameter to be learned. This is the final output corresponding to this input. Update the cell state value. This refers to the unit state. This means that by utilizing Long Short-Term Memory (LSTM) units, the relationship between previous inputs and the current input and output can be constructed, thereby learning the temporal permutations and combinations of symbols to mitigate performance degradation caused by inter-symbol interference. The specific patterns learned by this equalizer help reduce the bit error rate and also help increase the data rate. This is because short-term memory can mitigate the impact of inter-symbol interference to reduce symbol error, while long-term memory allows the network to remember the packet header sequence, thus recognizing more packets to reduce packet loss and increase the data rate.

[0068] In a further embodiment, before the step of performing image preprocessing on each frame of the acquired video and dividing it into a series of sub-images, the method further includes video acquisition, wherein the video acquisition method is: recording in video format at 60 frames per second and a resolution of 1080×1920.

[0069] In a further embodiment, the step of performing image preprocessing on each frame of the acquired video and segmenting it into a series of sub-images specifically includes:

[0070] S001, Frame Extraction, which is to extract each frame of the video as the image frame to be decoded;

[0071] S002, Column Selection, which involves selecting a suitable pixel column by means of the average gray level of the column, and then expanding the signal of that column into an image by means of copying;

[0072] S003, Amplitude Normalization, which normalizes the pixel intensity of the RGB channels of the image;

[0073] S004, gamma correction, which is to compensate for the signal intensity of the RGB channels of the image using gamma correction;

[0074] S005. Perform column selection again, that is, select the center column in the compensated image;

[0075] S006. Header Location: This involves fitting the header data using a threshold method to obtain the header position. Since the symbols that make up the header are "white" and "black", the threshold method can be used to fit the data effectively.

[0076] S007. Image segmentation, which is to segment the image into multiple sub-images, starting from the obtained head position.

[0077] In a further embodiment, the method further includes: time-domain recovery, that is, restoring the time sequence of the obtained binary data.

[0078] Example 2

[0079] As shown in Figures 1-6, this embodiment provides a visible light communication decoding system based on deep learning, including: a building module for building a CAE-LSTM-Net neural network;

[0080] The image preprocessing module is used to preprocess each frame of the captured video and divide it into a series of sub-images;

[0081] The offline decoding module is used to decode each sub-image within one time step using the constructed CAE-LSTM-Net neural network. Specifically, it includes: inputting the sub-images into the trained CAE-LSTM-Net neural network in chronological order; the CAE-LSTM-Net neural network outputting the category corresponding to the sub-image; and obtaining the binary data of the RGB channels contained in the symbol in the sub-image.

[0082] In a further embodiment, such as Figure 7 As shown, the system also includes a video acquisition device, which comprises:

[0083] The transmitter includes a programmable gate array (FPGA, Xilinx Spalding6, XC6SLX16) and RGB-LEDs for emitting visible light;

[0084] The receiving end includes a mobile phone with an optical camera based on a CMOS sensor for video capture. The video capture method is to record video at 60 frames per second with a resolution of 1080×1920. Other parameters are automatically set by software.

[0085] In this embodiment, data packets constructed from random data are input into a programmable gate array (FPGA, Xilinx Spalding 6, XC6SLX16), which drives an RGB-LED light through its I / O ports. The data packets are transmitted at a baud rate of 32Kbauds to 46Kbauds. After 40cm of free-space signal transmission, a mobile phone (HUAWEI P20 PRO) with an optical camera based on a CMOS sensor is placed at the data receiving end. A convex-convex lens and a diffuser are arranged in front of the phone's lens.

[0086] The image preprocessing module includes:

[0087] The frame extraction module is used to extract each frame of the video as an image frame to be decoded.

[0088] The column selection module is used to select a suitable pixel column by means of the average gray level of the column, and to expand the signal of the column into an image by means of copying;

[0089] The amplitude normalization module is used to normalize the pixel intensity of the RGB channels of the image frame;

[0090] The gamma correction module is used to compensate for the signal strength of the RGB channels of the image frame using gamma correction.

[0091] The center column selection module is used to select the center column in the compensated image frame;

[0092] The data packet header localization module is used to fit the header data using a threshold method to obtain the header position.

[0093] The image segmentation module is used to segment the image frame into multiple sub-images, starting from the obtained head position.

[0094] In summary, the deep learning-based visible light communication decoding method and system provided by this invention uses RGB-LED lights for signal modulation, theoretically achieving a data rate three times higher than that based on a black-and-white LED light framework. A neural network framework based on channel attention and long short-term memory mechanisms is constructed, with the encoder based on the channel attention mechanism fully extracting the high-dimensional features of the received image. Addressing the problem of strong inter-symbol interference in high-speed optical camera communication, this invention fully considers mitigating the performance degradation caused by this problem in both space and time. This invention designs an encoder based on the channel attention mechanism and an equalizer based on the long short-term memory mechanism to learn patterns that mitigate the performance degradation caused by inter-symbol interference in both space and time.

[0095] The remaining technical features in this embodiment can be flexibly selected by those skilled in the art to meet different specific practical needs. However, it is obvious to those skilled in the art that these specific details are not necessary to implement this invention. In other instances, to avoid obscuring this invention, well-known components and structures are not specifically described, and all are within the scope of technical protection defined by the claims of this invention.

[0096] Modifications and variations made by those skilled in the art that do not depart from the spirit and scope of this invention should be within the protection scope of the appended claims. In the above description, numerous specific details have been set forth to provide a thorough understanding of the invention. However, it will be apparent to those skilled in the art that these specific details are not necessary to practice the invention. In other instances, to avoid obscuring the invention, well-known techniques, such as specific construction details, operating conditions, and other technical conditions, have not been specifically described.

[0097] This document uses specific examples to illustrate the principles and implementation methods of the present invention. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of the present invention. Furthermore, those skilled in the art will recognize that, based on the ideas of the present invention, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of the present invention.

Claims

1. A deep learning-based visible light communication decoding method, characterized in that, Includes the following steps: Construct a CAE-LSTM-Net neural network; Each frame of the captured video is preprocessed and segmented into a series of sub-images; The constructed CAE-LSTM-Net neural network is used to decode each sub-image in one time step. Specifically, the sub-images are input into the trained CAE-LSTM-Net neural network in chronological order. The CAE-LSTM-Net neural network outputs the category corresponding to the sub-image and obtains the binary data of the RGB channels contained in the symbol in the sub-image. The constructed CAE-LSTM-Net neural network includes: an encoder neural network based on the channel attention mechanism and an equalizer neural network based on the long short-term memory mechanism; The encoder neural network based on the channel attention mechanism includes a convolutional layer, a multi-channel high-level feature extraction unit, and a recalibration unit. The convolutional layer and the multi-channel high-level feature extraction unit are used to extract multi-channel high-level features of the input image to obtain a feature map. The recalibration unit is used to determine the weights between different channels of the feature map so as to recalibrate the channels of the feature map before channel fusion. The equalizer neural network based on the long short-term memory mechanism includes a long short-term memory unit, and the input-output relationship of the long short-term memory unit is as follows: ; in, This is the weight matrix to be learned, where σ(.) represents the sigmoid activation function, φ(.) represents the tanh activation function, ⊙ represents the element-wise dot product, and the gates are the forget gate (ft), input gate (it), and output gate. bf, bi, bg, and bo are the biases, which are the parameters to be learned, and ht is the final output corresponding to this input. This is the cell state update value, where Ct is the cell state; The process of preprocessing each frame of the acquired video and segmenting it into a series of sub-images specifically includes: Frame extraction, which is to extract each frame of the video as an image frame to be decoded; Column selection involves choosing a suitable pixel column based on the average gray level of the columns, and then expanding the signal of that column into the image by copying it. Amplitude normalization is the process of normalizing the pixel intensity of the RGB channels in an image frame. Gamma correction is used to compensate for the signal strength of the RGB channels in an image frame. Perform column selection again, that is, select the center column in the compensated image frame; Data packet header localization, which is to obtain the header position by fitting the header data using a threshold method; Image segmentation is the process of dividing an image frame into multiple sub-images, starting from the obtained head position.

2. The visible light communication decoding method based on deep learning according to claim 1, characterized in that, Before the step of preprocessing each frame of the acquired video and dividing it into a series of sub-images, the method of video acquisition is as follows: recording in video format at 60 frames per second and a resolution of 1080×1920.

3. The visible light communication decoding method based on deep learning according to claim 1, characterized in that, The method further includes: Time-domain recovery involves restoring the time sequence of the obtained binary data.

4. A visible light communication decoding system based on deep learning, characterized in that, include: Modules for building CAE-LSTM-Net neural networks; The image preprocessing module is used to preprocess each frame of the captured video and divide it into a series of sub-images; The offline decoding module is used to decode each sub-image within one time step using the constructed CAE-LSTM-Net neural network. This includes: inputting the sub-images into the trained CAE-LSTM-Net neural network in chronological order; the CAE-LSTM-Net neural network outputting the category corresponding to the sub-image; and obtaining the binary data of the RGB channels contained in the symbol in the sub-image. Specifically, the process of preprocessing each frame of the acquired video and dividing it into a series of sub-images includes: Frame extraction, which is to extract each frame of the video as an image frame to be decoded; Column selection involves choosing a suitable pixel column based on the average gray level of the columns, and then expanding the signal of that column into the image by copying it. Amplitude normalization is the process of normalizing the pixel intensity of the RGB channels in an image frame. Gamma correction is used to compensate for the signal strength of the RGB channels in an image frame. Perform column selection again, that is, select the center column in the compensated image frame; Data packet header localization, which is to obtain the header position by fitting the header data using a threshold method; Image segmentation is the process of dividing an image frame into multiple sub-images, starting from the obtained head position. The constructed CAE-LSTM-Net neural network includes: an encoder neural network based on the channel attention mechanism and an equalizer neural network based on the long short-term memory mechanism; The specific construction method of the CAE-LSTM-Net neural network is as follows: the encoder neural network based on the channel attention mechanism includes a convolutional layer, a multi-channel high-level feature extraction unit and a recalibration unit. The convolutional layer and the multi-channel high-level feature extraction unit are used to extract multi-channel high-level features of the input image to obtain a feature map. The recalibration unit is used to determine the weights between different channels of the feature map so as to recalibrate the channels of the feature map before channel fusion. The equalizer neural network based on the long short-term memory mechanism includes a long short-term memory unit, and the input-output relationship of the long short-term memory unit is as follows: ; in, This is the weight matrix to be learned, where σ(.) represents the sigmoid activation function, φ(.) represents the tanh activation function, ⊙ represents the element-wise dot product, and the gates are the forget gate (ft), input gate (it), and output gate. bf, bi, bg, and bo are the biases, which are the parameters to be learned, and ht is the final output corresponding to this input. Ct represents the cell state update value.

5. The visible light communication decoding system based on deep learning according to claim 4, characterized in that, The system also includes a video acquisition device, which comprises: The transmitting end includes a programmable gate array and RGB-LEDs for emitting visible light; The receiving end includes a mobile phone with an optical camera based on a CMOS sensor for video acquisition. The video acquisition method is to record the video at 60 frames per second with a resolution of 1080×1920.

6. The visible light communication decoding system based on deep learning according to claim 4, characterized in that, The image preprocessing module includes: The frame extraction module is used to extract each frame of the video as an image frame to be decoded. The column selection module is used to select a suitable pixel column by means of the average gray level of the column, and to expand the signal of the column into an image by means of copying; The amplitude normalization module is used to normalize the pixel intensity of the RGB channels of the image frame; The gamma correction module is used to compensate for the signal strength of the RGB channels of the image frame using gamma correction. The center column selection module is used to select the center column in the compensated image frame; The data packet header localization module is used to fit the header data using a threshold method to obtain the header position. The image segmentation module is used to segment the image frame into multiple sub-images, starting from the obtained head position.