A Dynamic Multi-Beacon Visible Light Communication Method, Device and Electronic Equipment
Through PPM-like encoding, sparse self-attention convolutional long short-term memory network and Gaussian mask processing combined with stream transformer decoding, the tracking and identity recognition problems of multiple dynamic beacons in visible light communication are solved, and the simultaneous communication and noise interference suppression of multiple beacons are achieved.
Patent Information
- Application Number
- CN202510660300.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-22
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-05-22
AI Technical Summary
The existing visible light communication methods have limitations in their scope of application and communication effects. Especially when factors such as background noise, motion noise and mutual interference between light sources are coupled, it is impossible to track, identify and communicate multiple dynamic beacons.
The beacon signal is encoded and classified by PPM-like encoding, identity distinction and dynamic tracking are performed through sparse self-attention convolutional long short-term memory network, channel separation is performed using Gaussian mask, and stream transformer based on the state reuse mechanism is used for decoding.
The simultaneous communication of multiple dynamic beacons is realized, which reduces the adverse effects of background noise, motion noise and mutual interference between light sources, expands the scope of application of visible light communication and improves the communication effect.
Smart Images

Figure CN120185713B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of visible light communication, and particularly to a dynamic multi-beacon visible light communication method, apparatus, and electronic device. Background Art
[0002] Visible light communication is a wireless optical transmission technology that uses visible light with wavelengths in the range of 380 nm to 790 nm for data communication. An event camera is a new generation of bionic asynchronous vision sensor with imaging characteristics of high dynamic range (>120 dB) and high temporal resolution (<10 μs). Different from the imaging method of traditional cameras that record the absolute value of brightness by framing, an event camera generates outputs only when brightness changes occur. This method effectively preserves the precious output bandwidth. At a farther distance, this characteristic provides higher system throughput while retaining the advantages of a camera-based system - spatial filtering ability. In addition, the fixed image background will be filtered out, greatly reducing the computational amount required to process the camera output. All these characteristics reveal the great potential of event cameras as visible light communication receivers.
[0003] Visible light communication based on event cameras is a continuation and development of visible light communication research using conventional RGB cameras as receivers. In the field of visible light communication with event cameras, early work combined the fine spatio-temporal resolution of event cameras with LEDs flashing at different frequencies. Censi et al. proposed an algorithm for attitude tracking using four infrared LED markers with different flashing frequencies. This technology uses flashing LED markers to trigger events and processes the input events using a particle filter algorithm to track the attitude of an unmanned aerial vehicle with an LED marker. Recently, the work of Wang et al. utilized an event camera to implement a static intelligent beacon and transmitted messages using the UART (Universal Asynchronous Receiver / Transmitter) protocol. The work of Von et al. used optical flow estimation based on pulsed neural networks and a Kalman filter algorithm to achieve identity discrimination and tracking of LED beacons with different flashing patterns in a scenario where the transmitter is fixed and the receiver is movable, but cannot transmit other information except the identity identifier. Due to factors such as background noise, motion noise, and mutual interference between light sources, many studies have further restricted the application scenarios of visible light communication: the technology proposed by Wang et al. mentioned above only realizes the communication function of a single static beacon and cannot achieve the tracking and identity discrimination of multiple targets or moving targets; the technologies proposed by Censi et al. and Von et al. mentioned above can track multiple movable targets and distinguish their identities, but cannot transmit information other than the identity identifier.
[0004] Regarding the problem that existing visible light communication methods have limitations in terms of applicable scope and communication effect, no effective solution has been proposed yet. Summary of the Invention
[0005] The present invention provides a dynamic multi-beacon visible light communication method, device, and electronic device to solve the defects of existing visible light communication methods having limitations in terms of applicable scope and communication effect, and to achieve the simultaneous tracking, identity recognition, and communication of multiple dynamic beacons when various challenges, such as background noise, motion noise, and mutual interference between light sources, are coupled together.
[0006] In a first aspect, the present invention provides a dynamic multi-beacon visible light communication method, including:
[0007] Obtain a beacon signal and perform encoding classification on the beacon signal;
[0008] Perform identity discrimination and dynamic tracking on the beacon signal through a sparse self-attention convolutional long short-term memory network;
[0009] Perform channel separation processing and signal filtering processing on the beacon signal using a Gaussian mask;
[0010] Decode the beacon signal using a streaming transformer based on a state reuse mechanism to generate a decoding result.
[0011] According to the dynamic multi-beacon visible light communication method provided by the present invention, performing encoding classification on the beacon signal includes:
[0012] Encode the beacon signal using a PPM-like encoding.
[0013] According to the dynamic multi-beacon visible light communication method provided by the present invention, encoding the beacon signal using a PPM-like encoding includes:
[0014] Perform waveform modulation on the data part of the beacon signal using a legal encoding, and perform waveform modulation on the identity identifier of the beacon signal using an illegal encoding.
[0015] According to the dynamic multi-beacon visible light communication method provided by the present invention, before performing identity discrimination and dynamic tracking on the beacon signal through a sparse self-attention convolutional long short-term memory network, it includes:
[0016] Determine a time interval according to the communication rate and compress the asynchronous event stream into an event frame.
[0017] According to the dynamic multi-beacon visible light communication method provided by the present invention, performing identity discrimination and dynamic tracking on the beacon signal through a sparse self-attention convolutional long short-term memory network includes:
[0018] Obtain a plurality of consecutive ones of the event frames;
[0019] Extract features of the beacon signal through a sparse self-attention convolutional long short-term memory network, and obtain the position, radius, and identity category of the beacon signal in the event frame through a classification output head and a tracking output head.
[0020] According to a dynamic multi-beacon visible light communication method provided by the present invention, a Gaussian mask is used to perform channel separation processing and signal filtering processing on the beacon signal, including:
[0021] Perform masked weighted multi-channel separation on the beacon signal, compress and reduce the dimension of the event frame, and generate a one-dimensional signal.
[0022] According to a dynamic multi-beacon visible light communication method provided by the present invention, a streaming transformer based on a state reuse mechanism is used to decode the beacon signal to generate a decoding result, including:
[0023] Obtain the one-dimensional signal of each beacon signal, perform feature extraction and sequence splicing on the one-dimensional signal to obtain intermediate data;
[0024] Perform streaming transformer decoding on the intermediate data based on the state reuse mechanism to obtain the decoding result.
[0025] According to a dynamic multi-beacon visible light communication method provided by the present invention, perform feature extraction and sequence splicing on the one-dimensional signal to obtain intermediate data, including:
[0026] Set the window size of the one-dimensional signal, perform a fast Fourier transform on the one-dimensional signal with windowing to generate a spectrogram;
[0027] Perform feature extraction on the spectrogram to obtain feature data;
[0028] Flatten the feature data into a one-dimensional vector, perform dimensionality reduction processing on the one-dimensional vector, and splice it with the sequence of the one-dimensional signal of the current time window to obtain the intermediate data.
[0029] In a second aspect, the present invention also provides a dynamic multi-beacon visible light communication device, including a transmitting end and a receiving end;
[0030] The transmitting end is used to obtain a beacon signal and perform encoding classification on the beacon signal; the receiving end includes:
[0031] A target recognition and tracking module, configured to perform identity discrimination and dynamic tracking on the beacon signal through a sparse self-attention convolutional long short-term memory network;
[0032] A Gaussian mask event processing module for performing channel separation processing and signal filtering processing on the beacon signal by using a Gaussian mask;
[0033] A decoding module for decoding the beacon signal by using a streaming transformer based on a state reuse mechanism to generate a decoding result.
[0034] In a third aspect, the present invention further provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the dynamic multi-beacon visible light communication method described in the first aspect above is implemented.
[0035] In a fourth aspect, the present invention further provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the dynamic multi-beacon visible light communication method described in the first aspect above is implemented.
[0036] In a fifth aspect, the present invention further provides a computer program product, including a computer program. When the computer program is executed by a processor, the dynamic multi-beacon visible light communication method described in the first aspect above is implemented.
[0037] Compared with the prior art, the present invention has the following beneficial effects:
[0038] The dynamic multi-beacon visible light communication method provided by the present invention differentiates different beacon signals through data frames with identity identifiers and a class PPM coding method, thereby making it possible for multiple beacon signals to communicate simultaneously; and realizes the tracking of dynamic multi-beacons through a sparse self-attention convolutional long short-term memory network, which can exclude other moving objects in the background and realize the identity differentiation and tracking positioning of multiple moving beacon signals. Then, the channel separation and signal filtering processing of the beacon signal are carried out by using the Gaussian mask method, which provides the possibility for multiple beacon signals to communicate simultaneously and also removes the influence of mutual interference between some light sources. Finally, a streaming transformer based on a state reuse mechanism is used to decode the communication data of the beacon signal to obtain a decoding result. Through the above process, different beacon signals can communicate simultaneously, and the adverse effects brought by factors such as background noise, motion noise, and mutual interference between light sources can be reduced, and the problems of limitations in the applicable range and communication effect of the existing visible light communication methods are solved. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] To more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.
[0040] Figure 1 is a flowchart of the dynamic multi-beacon visible light communication method provided by the present invention;
[0041] Figure 2 is a schematic diagram of the PPM-like coded symbol waveform in an embodiment of the present invention;
[0042] Figure 3 is a schematic diagram of the data frame format in an embodiment of the present invention;
[0043] Figure 4 is a schematic diagram of the network structure for identifying and tracking beacon signals in an embodiment of the present invention;
[0044] Figure 5 is a structure diagram of the sparse self-attention convolutional long short-term memory network in an embodiment of the present invention;
[0045] Figure 6 is a schematic diagram of the structure of the sparse self-attention module in an embodiment of the present invention;
[0046] Figure 7 is a schematic diagram of the structure of the streaming transformer in an embodiment of the present invention;
[0047] Figure 8 is a schematic diagram of the structure of the state-reused multi-head attention encoder in an embodiment of the present invention;
[0048] Figure 9 is a structural block diagram of the dynamic multi-beacon visible light communication device provided by the present invention;
[0049] Figure 10 is a structural block diagram of the receiving end in an embodiment of the present invention;
[0050] Figure 11 is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed implementation manners
[0051] To make the objectives, technical solutions, and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions in the present invention in conjunction with the accompanying drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Based on the embodiments in the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.
[0052] The present invention provides a dynamic multi-beacon visible light communication method, Figure 1 which is a flowchart of the dynamic multi-beacon visible light communication method provided by the present invention. As Figure 1 shown, the method includes the following steps:
[0053] Step S101, obtain beacon signals and perform encoding classification on the beacon signals;
[0054] Step S102, perform identity discrimination and dynamic tracking on the beacon signals through a sparse self-attention convolutional long short-term memory network;
[0055] Step S103, perform channel separation processing and signal filtering processing on the beacon signals by using a Gaussian mask;
[0056] Step S104, decode the beacon signals by using a streaming transformer based on a state reuse mechanism to generate a decoding result.
[0057] In this method, first, different beacon signals are distinguished by using a data frame with an identity identifier and a class PPM (Pulse Position Modulation) encoding method, thus making it possible for multiple beacon signals to communicate simultaneously. Then, the dynamic tracking of multiple beacons is realized through a sparse self-attention convolutional long short-term memory network. Since the convolutional long short-term memory network can extract the spatio-temporal features of an image stream and realize dynamic target recognition and tracking, the identity discrimination and tracking of multiple moving beacon signals can exclude other moving objects in the background and realize the identity discrimination and tracking positioning of multiple moving beacon signals. Furthermore, the channel separation and signal filtering processing of the beacon signals are performed by using a Gaussian mask method, which provides the possibility for multiple beacon signals to communicate simultaneously and also removes the influence of mutual interference between some light sources. Finally, the communication data of the beacon signals are decoded by using a streaming transformer based on a state reuse mechanism to obtain a decoding result. Through the above process, it is possible to realize the simultaneous communication of different beacon signals, and moreover, it is possible to reduce the adverse effects brought by factors such as background noise, motion noise, and mutual interference between light sources, and solve the problem that the existing visible light communication methods have limitations in terms of application range and communication effect.
[0058] In some of the embodiments, in step S101, performing encoding classification on the beacon signals includes: encoding the beacon signals by using class PPM encoding. Specifically, waveform modulation is performed on the data part of the beacon signals by using a legal encoding, and waveform modulation is performed on the identity identifier of the beacon signals by using an illegal encoding. Among them, the legal encoding is a standard PPM encoding, and the illegal encoding is not a standard PPM encoding. The specific encoding waveforms are as Figure 2 shown, Figure 2It is a schematic diagram of the waveform of the PPM-like encoded symbol in the embodiment of the present invention.
[0059] The standard PPM encoding, namely optical pulse position modulation, transmits information by changing the timing position of optical pulses. This makes PPM have the advantage of anti-noise interference, thus better combating the motion noise of the light source during communication. In addition, since there is only a high level in one time slot within any symbol time interval of PPM encoding, when applied in this system, the brightness of the beacon can be fixed within a small variation range. Illegal encoding has no such restriction of a single high-level time slot within one symbol time interval.
[0060] Assume that it takes time T to transmit one symbol, and T is evenly divided into three time slots 0, 1, and 2. In Figure 2 , the symbol "0" is represented by a high level in time slot 1 and low levels in other time slot positions; the symbol "1" is represented by a high level in time slot 0 and low levels in other time slot positions; other ways of occupying time slots are called "illegal encoding".
[0061] Figure 3 It is a schematic diagram of the data frame format in the embodiment of the present invention. As Figure 3 shown, in this embodiment, the data frame consists of a total of 9 symbols. The first 2 symbol headers are identity identifiers, and the last 7 symbols are data encoded in ASCII (American Standard Code for Information Interchange). Selecting 2 or 1 from the defined 5 kinds of illegal encodings can form a data frame header for identifying the type of beacon. A total of 25 beacon types can be identified by the 5 kinds of illegal encodings. This design of the identity identifier header makes beacon signals with exactly the same appearance also have unique identifiers, realizing the discrimination of different beacons. In order to avoid the confusion between the header and the semantic segment while not reducing the communication efficiency, a new symbol waveform modulation method - PPM-like encoding is proposed in this embodiment, that is, symbols are divided into legal encoding and illegal encoding. The data part is modulated with legal encoding waveforms, and the identity identifier is modulated with illegal encoding. PPM encoding has the advantage of anti-noise interference. Adopting the PPM-like encoding method helps to improve the signal-to-noise ratio and combat the motion noise of the light source during communication.
[0062] This embodiment adopts a beacon identification method in the communication frame that maps the identification sequence to unique illegal code positions to replace the start / stop bits of the data frame. On the one hand, it avoids the confusion between the header and the semantic segment while not reducing the communication efficiency. On the other hand, in target recognition and tracking, even if the beacons have exactly the same appearance, the algorithm can distinguish each target, and finally the decoding module can parse the communication information of different beacons.
[0063] In some of these embodiments, before step S102, where the beacon signals are identified and dynamically tracked through a sparse self-attention convolutional long short-term memory network, it includes: determining a time interval according to the communication rate and compressing the asynchronous event stream into event frames.
[0064] Exemplarily, let the event with index in the event stream be , denoted as a quadruple:
[0065]
[0066] where represents the pixel position where event occurs; is the timestamp, indicating the time when the event occurs; represents the polarity, that is, the direction of the brightness change. Usually, the event camera records the change of brightness increase or brightness decrease, represented by +1 or -1 respectively. The frame compression algorithm is as follows: Select the time interval , and let the event frame at time t be , and its calculation formula is:
[0067]
[0068] where, when , , at this time, equation is equal to 1; represents the polarity of event , when the polarity is positive , when the polarity is negative .
[0069] After passing through the frame compression algorithm, the frame compression result at this moment is output and temporarily stored in the memory. The specific number of frames stored depends on the input data window size and step size set for target recognition and tracking, and it is necessary to ensure that the input data of the target tracking module is not lost first.
[0070] In this embodiment, the selection of the time interval depends on the frequency of the communication transmission signal, and it needs to meet the Nyquist frequency requirement of the transmission signal: when the symbol transmission duration is T and the encoding method of the present invention is adopted, it is necessary to meet .
[0071] On this basis, in step S102, the beacon signal is identified and dynamically tracked through a sparse self-attention convolutional long short-term memory network, including: obtaining a number of consecutive event frames; extracting features of the beacon signal through the sparse self-attention convolutional long short-term memory network, and obtaining the position, radius, and identity category of the beacon signal in the event frame through the classification output head and the tracking output head.
[0072] Figure 4 It is a schematic diagram of the network structure for identifying and tracking the beacon signal in an embodiment of the present invention. As Figure 4 shown, the feature extraction part consists of 3 layers of sparse self-attention convolutional long short-term memory networks, and each layer is connected through a BN layer, a ReLU activation layer, and a max pooling layer for downsampling. Among them, these 3 convolutional long short-term memory networks respectively contain 8, 16, and 32 hidden nodes; then, the extracted feature data is flattened into a one-dimensional vector through a Flatten Layer; the output head respectively includes a classification network and a localization network. The classification network consists of 2 fully connected layers. The first fully connected layer contains 128 hidden nodes, and the second fully connected layer outputs the probabilities of n categories; the localization network also consists of 2 fully connected layers. The first fully connected layer is the same as the classification network, and the second fully connected layer outputs a vector with a length of 3n , where represents the coordinate position of the nth category, represents the radius size of the target.
[0073] In order to increase the sparsity of the network and improve the ability of the network to extract features, in this embodiment, a new sparse self-attention convolutional long short-term memory network is designed. The structure of this network block is as Figure 5 shown, Figure 5 It is a structural diagram of the sparse self-attention convolutional long short-term memory network in an embodiment of the present invention. The sparse self-attention convolutional long short-term memory network is composed of a convolutional long short-term memory network with a sparse self-attention module added. The structure of the sparse self-attention module is as Figure 6 shown, Figure 6 It is a schematic diagram of the structure of the sparse self-attention module in an embodiment of the present invention. The sparse self-attention module modifies the update equation of the convolutional long short-term memory network, and its update equation is as follows:
[0074]
[0075]
[0076]
[0077]
[0078]
[0079]
[0080]
[0081] Among them, represents the hidden state at time t-1, represents the hidden state after the self-attention mechanism and concatenation operation, represents the concatenation operation, represents the attention mechanism. Performing self-attention operation on the hidden state is beneficial to extracting more important features in the input signal. , respectively represent the outputs of the input gate i, forget gate f, candidate memory cell g, memory cell c, and output gate o at time t. , tanh represents the inverse tangent function, W represents the trainable weight, X represents the input; b represents the bias.
[0082] According to the above equations, it can be seen that in the convolutional long short-term memory network with a sparse self-attention module, the previous moment hidden state in the expressions of the input gate i, forget gate f, candidate memory cell g, memory cell c, and output gate o is all changed to The calculation formula of
[0083]
[0084] By the formula, the hidden state is sparsified, reducing the computational complexity of the network. Among them, is a threshold that can be set. By adjusting the value, the degree of being sparsified can be adjusted.
[0085] When performing inference, first set the input data window size and step size. The selection of the window size and step size needs to ensure that the network has sufficient labeled data input. Then, input the event frame stream into the network for inference. Finally, output the target category and tracking vector detected by the network. If the probability of a certain output category does not exceed the threshold and is not activated, then the data at the corresponding position in the tracking vector will be meaningless.
[0086] Exemplarily, each time p consecutive event frames are input, after passing through the feature extraction network, classification output head, and tracking output head, the position of the output beacon signal in is , the radius r, and the category of the beacon signal. In the prior art, for the tracking of beacon signals, some traditional image processing methods such as the blob (Binary large object) algorithm are used, and this method can only achieve the tracking of a single beacon. There are also some that use ordinary target tracking neural networks, but this method cannot distinguish beacons with the same appearance. In order to better extract the spatio-temporal features in the event frame stream, so as to distinguish and track multiple dynamic beacons, this embodiment designs a sparse self-attention convolutional long short-term memory network block for feature extraction. The sparse self-attention convolutional long short-term memory network block adds a sparse self-attention module on the basis of the convolutional long short-term memory network. The sparse self-attention module modifies the update formula of the convolutional long short-term memory network, that is, changing the hidden state at the previous moment of the update formula to , introducing sparsity through operation, and introducing the attention mechanism through and operation.
[0087] Since each beacon signal in communication has a defined identity identification frame header, its LED blinking pattern carries the identity identification, and the convolutional long short-term memory network can extract the spatio-temporal features of the image stream to complete the dynamic target recognition and tracking. Therefore, this embodiment can distinguish the identities and track multiple moving beacon signals, and can exclude other moving objects in the background.
[0088] This embodiment applies the sparse self-attention module to target recognition and tracking. Compared with the ordinary method of using the convolutional long short-term memory network to extract spatio-temporal features, it adds an attention mechanism, so that important features are more concerned. In addition, it also introduces sparsity to the hidden state of the network, saving computing resources.
[0089] In some of these embodiments, in step S103, a Gaussian mask is used to perform channel separation processing and signal filtering processing on the beacon signal, including: performing masked weighted multi-channel separation on the beacon signal, compressing and reducing the dimension of the event frame to generate a one-dimensional signal.
[0090] This embodiment uses the results of target recognition and tracking and the two-dimensional Gaussian function to compress and reduce the dimension of the event frame into multiple one-dimensional signals that can be used as the input of the decoding module. Assuming that the target recognition and tracking module detects n beacons, the position and size of each beacon signal are represented by the vector . The following operations are performed on the beacon signal i:
[0091] First step, calculate its Gaussian mask function:
[0092]
[0093] Among them, represents a Gaussian mask function, , the size of which depends on , satisfying .
[0094] In the second step, use to perform mask weighting on , and the calculation formula is:
[0095]
[0096] where W×H is the size of , is the finally obtained one-dimensional signal, q and k represent the pixel point coordinates. For n beacon signals, after the above operations, one-dimensional signals of n channels will be output, and each channel corresponds to a one-dimensional signal. For the i-th beacon signal, its corresponding one-dimensional signal is denoted as . Since there are more motion events at the beacon edge compared to the center position, through the weighting of the Gaussian mask function, the motion noise of the beacon itself can be removed to a certain extent, and a larger proportion of flicker events can be retained, facilitating decoding. After this processing, the beacon signals that are communicating are separated from each other, and respective decodable one-dimensional signals are generated.
[0097] Exemplarily, using a two-dimensional Gaussian function for masking, the two-dimensional event frame stream is further compressed and dimension-reduced into a one-dimensional signal that can be used as the input of the decoding module , realizing the mask weighting and multi-channel separation of beacon signals. On the one hand, it provides the possibility for multi-beacon signals to communicate simultaneously. On the other hand, the processing method of Gaussian function mask weighting reduces the influence of motion events at the light source edge and also removes the influence of mutual interference between some light sources.
[0098] Based on the above embodiments, in step S104, a streaming transformer based on a state reuse mechanism is used to decode the beacon signals to generate a decoding result, including: obtaining the one-dimensional signal of each beacon signal, performing feature extraction and sequence splicing on the one-dimensional signal to obtain intermediate data; performing streaming transformer decoding based on the state reuse mechanism on the intermediate data to obtain the decoding result.
[0099] In this embodiment, performing feature extraction and sequence splicing on the one-dimensional signal to obtain intermediate data includes: setting the window size of the one-dimensional signal, performing a fast Fourier transform on the one-dimensional signal by windowing to generate a spectrogram; performing feature extraction on the spectrogram to obtain feature data; flattening the feature data into a one-dimensional vector, performing dimensionality reduction processing on the one-dimensional vector, and splicing it with the sequence of the one-dimensional signal in the current time window to obtain intermediate data.
[0100] To fuse the deep spectral features and shallow temporal features of the sequence, a feature extraction module is designed in this embodiment. The feature extraction module includes a feature extraction network and a fusion algorithm. The feature extraction network extracts the deep spectral features of the one-dimensional signal and the fusion algorithm flattens and reduces the dimension of the extracted features and then concatenates them with the original time series as the input to the transformer encoder. First, take the time length of transmitting one frame of data, i.e., 9T, as a time window, and perform a short-time Fourier transform on the current time window and concatenate it with the results of the FFT of the previous p frames of data to form a spectrogram as the input to the feature extraction network. The feature extraction network includes a 2D convolutional block followed by 5 groups of Bottleneck blocks. The repetition times, strides, and output channels are shown in Table 1. The use of Bottleneck blocks can reduce the computational amount and the number of parameters on the one hand, thereby improving the network efficiency, and on the other hand, it is beneficial to multi-scale feature fusion. Then, the extracted features are flattened into a one-dimensional vector and reduced in dimension through a linear fully connected layer. Finally, the dimension-reduced features are concatenated with the current time window sequence as the input to the transformer encoder.
[0101] Table 1 Feature Extraction Network Structure Parameter Table
[0102]
[0103] As Figure 7 shown, Figure 7 is a schematic structural diagram of the streaming transformer in the embodiment of the present invention. Since the self-attention encoder-decoder of the classical transformer needs to calculate the attention weights on the entire sequence, it is difficult to implement streaming decoding; and if the transformer network is forced to input in frames, the decoding performance will be poor. For beacon communication, if the frame division position is not aligned with the sending end, the decoding effect will be even worse. The streaming transformer used in this embodiment can establish longer sequence dependencies through the state reuse mechanism of the encoder, thereby improving the decoding performance while reducing the latency.
[0104] As Figure 8 shown, Figure 8 is a schematic structural diagram of the state reuse multi-head attention encoder in the embodiment of the present invention. Compared with the multi-head attention encoder of the classical transformer, the present invention makes changes to the input of the multi-head attention module of the transformer encoder, thereby introducing the state reuse mechanism. The following is the modified calculation formula:
[0105]
[0106]
[0107]
[0108]
[0109] wherein, refers to the output after the -th input frame passes through the -th encoder layer, represents the Query of the τ -th input frame in the l -th encoder layer, represents the Key of the τ -th input frame in the l -th encoder layer, represents the Value of the τ -th input frame in the l -th encoder layer, represents multiple attention heads, represents the -th attention head, represents the attention mechanism, are all trainable matrix parameters, and the function represents stopping the gradient. According to the above formula, it can be seen that during the processing of the current frame, the state reuse encoder utilizes the state after encoding the previous frame, which enables the network to still establish longer sequence dependencies even when the frame length is short, and also solves the problem of poor decoding effect caused by misalignment of frame positions, thus achieving low-latency streaming decoding with shorter frames.
[0110] Exemplarily, set the window size of the one-dimensional signal, window the signal input over a long time period and perform a fast Fourier transform to generate a spectrogram. Then input the spectrogram into a feature extraction network stacked by one layer of conv2d and five layers of Bottleneck to extract features from the spectrogram of the signal over a long time interval to obtain feature data. Next, flatten the extracted feature data into a one-dimensional vector, and then perform dimensionality reduction through one layer of linear fully connected layer, and finally combine it with the current time window The sequences are spliced to obtain intermediate data. For the decoding of beacon signals, previous work often adopted the traditional communication decoding scheme of filtering - sampling - decoding. However, this method requires strict alignment of the signal's frame positions. Once misaligned, the communication information of the current frame will be lost. Therefore, the present invention designs a streaming Transformer to solve this problem. In this embodiment, based on the classical Transformer, it is also composed of an encoder and a decoder. However, in order to learn sequence - dependence features over a longer time to achieve streaming decoding, a state - reuse mechanism is introduced into the multi - head attention module of the encoder.
[0111] Compared with the traditional communication decoding scheme of filtering - sampling - decoding, in the decoding process of this embodiment, the long - time frequency deep features and short - time time - domain signals are first fused and then input into the Transformer for decoding. It has stronger ability to exclude other motion events and irrelevant events, improving the system robustness. The use of the state - reuse mechanism in the Transformer encoder also makes it unnecessary to align the signal's frame positions during the decoding process, solving the problem of difficult alignment.
[0112] In summary, the present invention proposes a dynamic multi - beacon visible - light communication method, which can simultaneously complete the tracking, identification, and communication of multiple dynamic beacons. First, the present invention proposes an encoding module that encodes communication frames using PPM - like encoding, solving the problem of differentiating beacons with the same appearance. Then, the present invention designs an object recognition and tracking module to solve the discrimination and tracking of multiple beacons. This module introduces a new sparse self - attention convolutional long - short - term memory network block, thereby increasing the sparsity of network parameters and improving the network's feature extraction ability. Next, in order to exclude other irrelevant event noises and simultaneously enable the decoding of data of multiple beacons, the present invention designs an event - frame dimensionality reduction method based on a Gaussian - function mask to compress the event - frame stream into a multi - channel one - dimensional time series. Finally, the present invention designs a decoding module that fuses multi - scale time - frequency features for analysis and decoding, and improves the decoding performance while reducing latency through the state - reuse mechanism.
[0113] The present invention also provides a dynamic multi - beacon visible - light communication device. The dynamic multi - beacon visible - light communication device provided by the present invention will be described below. The dynamic multi - beacon visible - light communication device described below can be mutually corresponded and referred to the dynamic multi - beacon visible - light communication method described above. Figure 9 is the structural block diagram of the dynamic multi - beacon visible - light communication device provided by the present invention. As Figure 9 shown, the device includes a transmitting end and a receiving end;
[0114] The transmitter is used to obtain beacon signals and perform encoding and classification on the beacon signals. The transmitter mainly includes a beacon with an LED light and a beacon main control to complete the waveform modulation module algorithm and LED light drive. Figure 10 is the structural block diagram of the receiver in the embodiment of the present invention, as Figure 10 shown, the receiver includes:
[0115] A target recognition and tracking module, which is used to distinguish the identity of beacon signals and perform dynamic tracking on the beacon signals through a sparse self-attention convolutional long short-term memory network;
[0116] A Gaussian mask event processing module, which is used to perform channel separation processing and signal filtering processing on the beacon signals by using a Gaussian mask;
[0117] A decoding module, which is used to decode the beacon signals by using a streaming transformer based on a state reuse mechanism to generate a decoding result.
[0118] In this device, an event camera is used as the receiver. The characteristics of the event camera enable it to filter out static backgrounds well and reduce background noise. When this device is in use, first, the transmitter distinguishes different beacon signals through data frames with identity identifiers and a class PPM (Pulse Position Modulation) encoding method, thus making it possible for multiple beacon signals to communicate simultaneously. Then, the target recognition and tracking module realizes the dynamic tracking of multiple beacons through a sparse self-attention convolutional long short-term memory network. Since the convolutional long short-term memory network can extract the spatio-temporal features of the image stream and realize dynamic target recognition and tracking, the identity of multiple moving beacon signals can be distinguished and tracked, and other moving objects in the background can be excluded, realizing the identity distinction and tracking positioning of multiple moving beacon signals. The Gaussian mask event processing module then performs channel separation and signal filtering processing on the beacon signals by using the Gaussian mask method, which provides the possibility for multiple beacon signals to communicate simultaneously and also removes the influence of mutual interference between some light sources. Finally, the decoding module decodes the communication data of the beacon signals by using a streaming transformer based on a state reuse mechanism to obtain a decoding result. Through the above process, it is possible to realize the simultaneous communication of different beacon signals, and moreover, it is possible to reduce the adverse effects brought by factors such as background noise, motion noise, and mutual interference between light sources, and solve the problem that the existing visible light communication methods have limitations in the applicable range and communication effect.
[0119] In addition, as shown in the figure, this device also includes an event stream compression frame module, which is used to determine the time interval according to the communication rate and compress the asynchronous event stream into an event frame.
[0120] For the target recognition and tracking module, since the event changes caused by data transmission are much greater than those caused by movement, the target recognition and tracking module does not need to be as fast as the decoding module. Therefore, the present invention decouples it from other modules to achieve asynchronous tracking and decoding, saving computing resources.
[0121] For the decoding module, the number of instances of the decoding module is dynamically adjusted according to the results of the target recognition and tracking module, that is, each detected beacon signal corresponds to a decoding module, and each decoding module only receives the one-dimensional signal of the corresponding beacon signal , so as to achieve simultaneous decoding of communication signals of multiple beacon signals. Specifically, the decoding module includes a feature extraction module and a streaming Transformer, and the feature extraction module and the streaming Transformer module need to be jointly trained to achieve high-quality streaming decoding.
[0122] Figure 11 Illustrates a schematic physical structure diagram of an electronic device, as Figure 11 shown, the electronic device may include: a processor 1101, a communication interface 1102, a memory 1103, and a communication bus 1104. Among them, the processor 1101, the communication interface 1102, and the memory 1103 complete mutual communication through the communication bus 1104. The processor 1101 can call the logical instructions in the memory 1103 to execute the dynamic multi-beacon visible light communication method, and the method includes:
[0123] Obtain beacon signals and perform encoding classification on the beacon signals;
[0124] Use a sparse self-attention convolutional long short-term memory network to distinguish the identities and dynamically track the beacon signals;
[0125] Use a Gaussian mask to perform channel separation processing and signal filtering processing on the beacon signals;
[0126] Use a streaming Transformer based on a state reuse mechanism to decode the beacon signals and generate decoding results.
[0127] In addition, when the logical instructions in the above-mentioned memory 1103 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0128] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the dynamic multi-beacon visible light communication method provided by the above-mentioned various methods. The method includes:
[0129] Obtain a beacon signal and perform encoding classification on the beacon signal;
[0130] Perform identity discrimination and dynamic tracking on the beacon signal through a sparse self-attention convolutional long short-term memory network;
[0131] Perform channel separation processing and signal filtering processing on the beacon signal using a Gaussian mask;
[0132] Perform decoding on the beacon signal using a streaming transformer based on a state reuse mechanism to generate a decoding result.
[0133] On another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is implemented to execute the dynamic multi-beacon visible light communication method provided by the above-mentioned various methods. The method includes:
[0134] Obtain a beacon signal and perform encoding classification on the beacon signal;
[0135] Perform identity discrimination and dynamic tracking on the beacon signal through a sparse self-attention convolutional long short-term memory network;
[0136] Perform channel separation processing and signal filtering processing on the beacon signal using a Gaussian mask;
[0137] Perform decoding on the beacon signal using a streaming transformer based on a state reuse mechanism to generate a decoding result.
[0138] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative effort.
[0139] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0140] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or equivalently replace some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A dynamic multi-beacon visible light communication method, characterized in that Including: The sending end acquires a beacon signal and encodes and classifies the beacon signal; The receiving end compresses an asynchronous event stream into event frames, inputs a plurality of consecutive event frames into a sparse self-attention convolutional long short-term memory network to distinguish the identity and dynamically track the beacon signal, and obtains the position, radius, and identity category of the beacon signal in the event frames; Performing channel separation processing and signal filtering processing on the beacon signal by using a Gaussian mask, including: using the results of target recognition and tracking and a two-dimensional Gaussian function to compress and reduce the dimension of the event frames into a plurality of one-dimensional signals as inputs to the decoding module; Decoding the beacon signal by using a streaming transformer based on a state reuse mechanism to generate a decoding result.
2. The dynamic multi-beacon visible light communication method according to claim 1, characterized in that, Encoding and classifying the beacon signal, including: Encoding the beacon signal by using a PPM-like encoding.
3. The dynamic multi-beacon visible light communication method according to claim 2, wherein Encoding the beacon signal by using a PPM-like encoding, including: Performing waveform modulation on the data part of the beacon signal by using a legal encoding, and performing waveform modulation on the identity identifier of the beacon signal by using an illegal encoding; The legal encoding is a standard PPM encoding, and the illegal encoding is a non-standard PPM encoding.
4. The dynamic multi-beacon visible light communication method according to claim 1, wherein Before distinguishing the identity and dynamically tracking the beacon signal by using a sparse self-attention convolutional long short-term memory network, including: Determining a time interval according to the communication rate and compressing the asynchronous event stream into event frames.
5. The dynamic multi-beacon visible light communication method according to claim 4, characterized in that, Distinguishing the identity and dynamically tracking the beacon signal by using a sparse self-attention convolutional long short-term memory network, including: Acquiring a plurality of consecutive event frames; Extracting features of the beacon signal by using a sparse self-attention convolutional long short-term memory network, and obtaining the position, radius, and identity category of the beacon signal in the event frames through a classification output head and a tracking output head.
6. The dynamic multi-beacon visible light communication method according to claim 5, wherein Performing channel separation processing and signal filtering processing on the beacon signal by using a Gaussian mask, including: Performing mask weighted multi-channel separation on the beacon signal, compressing and reducing the dimension of the event frames, and generating one-dimensional signals.
7. The dynamic multi-beacon visible light communication method according to claim 6, wherein Decoding the beacon signal by using a streaming transformer based on a state reuse mechanism to generate a decoding result, including: Acquiring the one-dimensional signal of each beacon signal, extracting features and performing sequence splicing on the one-dimensional signal to obtain intermediate data; Performing streaming transformer decoding on the intermediate data based on a state reuse mechanism to obtain the decoding result.
8. The dynamic multi-beacon visible light communication method according to claim 7, wherein Extracting features and performing sequence splicing on the one-dimensional signal to obtain intermediate data, including: Setting the window size of the one-dimensional signal, performing a fast Fourier transform on the one-dimensional signal with windowing to generate a spectrogram; Extracting features from the spectrogram to obtain feature data; Flattening the feature data into a one-dimensional vector, performing dimensionality reduction processing on the one-dimensional vector, and splicing it with the sequence of the one-dimensional signal of the current time window to obtain the intermediate data.
9. A dynamic multi-beacon visible light communication device, characterized in that, Including a sending end and a receiving end; The sending end is used to acquire a beacon signal and encode and classify the beacon signal; the receiving end includes: The target recognition and tracking module is used to compress the asynchronous event stream into event frames, input several consecutive event frames into a sparse self-attention convolutional long short-term memory network to distinguish the identity and dynamically track the beacon signal, and obtain the position, radius, and identity category of the beacon signal in the event frames; The Gaussian mask event processing module is used to perform channel separation processing and signal filtering processing on the beacon signal by using a Gaussian mask, including: using the results of target recognition and tracking and a two-dimensional Gaussian function to compress and reduce the dimension of the event frame into multiple one-dimensional signals as the input of the decoding module; The decoding module is used to decode the beacon signal by using a streaming transformer based on a state reuse mechanism to generate a decoding result.
10. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the dynamic multi-beacon visible light communication method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Indoor positioning method, equipment and medium
CN119573737A
Absolute pointing device and method and interaction device
CN119729073A