Visual information recognition method based on pulse transformer network

By introducing a prior knowledge embedding layer of Fourier basis functions or wavelet basis functions into the pulse Transformer network, the problem of low accuracy in visual information recognition is solved, achieving higher recognition accuracy and reduced energy consumption.

CN119339282BActive Publication Date: 2025-11-18INST OF AUTOMATION CHINESE ACAD OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411213211.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-30
Publication Date
2025-11-18
Estimated Expiration
2044-08-30

AI Technical Summary

Technical Problem

Existing pulse Transformer networks have low accuracy in visual information recognition due to their discrete pulse coding method.

Method used

A prior knowledge embedding layer is constructed using Fourier basis functions or wavelet basis functions to replace the attention module in the pulse Transformer network, thereby building a visual information recognition model.

Benefits of technology

It improves the accuracy of visual information recognition while reducing energy consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119339282B_ABST
    Figure CN119339282B_ABST
Patent Text Reader

Abstract

The application provides a visual information recognition method based on a pulse Transformer network, and the method comprises the following steps: obtaining a pulse sequence of visual information to be recognized; constructing a visual information recognition model based on a pulse Transformer network; wherein one or more basis functions in a Fourier basis function or a wavelet basis function are used to construct a priori knowledge embedding layer; the priori knowledge embedding layer is used to replace an attention module in the pulse Transformer network; the pulse sequence is input into the visual information recognition model to obtain a recognition result of the visual information to be recognized. By using the above method, the application solves the problem that the accuracy of visual information recognition using the pulse Transformer network is low due to the discrete pulse coding mode of the real value of the pulse neural network in the related art.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of image detection, and in particular to a visual information recognition method based on a pulse Transformer network. BACKGROUND

[0002] The pulse Transformer network combines the advantages of SNN in processing pulse data and the ability of Transformer in capturing long-distance dependencies, and is suitable for processing pulse data with spatiotemporal characteristics. Through the self-attention mechanism, the pulse Transformer network can capture long-distance dependencies in pulse data and improve the modeling capability of the model.

[0003] The pulse Transformer network can realize the recognition function of visual information, but due to the discrete pulse coding mode of the pulse neural network for real values, the accuracy of using the pulse Transformer network for visual information recognition is low. SUMMARY

[0004] The present application provides a visual information recognition method based on a pulse Transformer network, which solves the problem of low accuracy in using the pulse Transformer network for visual information recognition due to the discrete pulse coding mode of the pulse neural network for real values in the prior art.

[0005] The present application provides a visual information recognition method based on a pulse Transformer network, which solves the problem of low accuracy in using the pulse Transformer network for visual information recognition due to the discrete pulse coding mode of the pulse neural network for real values in the prior art.

[0006] Obtaining a pulse sequence of visual information to be recognized;

[0007] Based on the pulse Transformer network, a visual information recognition model is constructed; wherein one or more basis functions of Fourier basis functions or wavelet basis functions are used to construct a priori knowledge embedding layer; and the priori knowledge embedding layer is used to replace the attention module in the pulse Transformer network;

[0008] The pulse sequence is input into the visual information recognition model to obtain the recognition result of the visual information to be recognized.

[0009] According to the visual information recognition method based on the pulse Transformer network provided by the present application, the pulse sequence is input into the visual information recognition model to obtain the recognition result of the visual information to be recognized, which comprises:

[0010] The pulse sequence is input into the convolutional encoding layer of the visual information recognition model for preliminary encoding;

[0011] input the preliminary coded pulse sequence into a deep coding layer in the visual information recognition model for deep coding;

[0012] input the deep coded pulse sequence into a classification layer in the visual information recognition model to obtain the recognition result of the to-be-recognized visual information.

[0013] According to the visual information recognition method based on the pulse Transformer network provided by the application, the prior knowledge embedding layer is located in the deep coding layer.

[0014] According to the visual information recognition method based on the pulse Transformer network provided by the application, the pulse sequence of the to-be-recognized visual information comprises:

[0015] When the to-be-recognized visual information is a static RGB image, the static RGB image is input into an RGB coding module to obtain the pulse sequence of the to-be-recognized visual information.

[0016] According to the visual information recognition method based on the pulse Transformer network provided by the application, the pulse sequence of the to-be-recognized visual information comprises:

[0017] When the to-be-recognized visual information is DVS data, the DVS data is directly taken as the pulse sequence of the to-be-recognized visual information.

[0018] According to the visual information recognition method based on the pulse Transformer network provided by the application, the prior knowledge embedding layer is constructed by using one or more basis functions in a Fourier basis function or a wavelet basis function, and the prior knowledge embedding layer comprises:

[0019] The prior knowledge embedding matrix is constructed by using different frequency sine functions and cosine functions in the Fourier basis function to obtain the prior knowledge embedding layer; or

[0020] The prior knowledge embedding matrix is constructed by using different frequency cosine functions in the Fourier basis function to obtain the prior knowledge embedding layer.

[0021] According to the visual information recognition method based on the pulse Transformer network provided by the application, the prior knowledge embedding layer is constructed by using one or more basis functions in a Fourier basis function or a wavelet basis function, and the prior knowledge embedding layer comprises:

[0022] The prior knowledge embedding matrix is constructed by using one or more discrete wavelet basis functions to obtain the prior knowledge embedding layer.

[0023] The application further provides a visual information recognition device based on a pulse Transformer network, comprising:

[0024] A visual information processing unit is used to acquire pulse sequences of visual information to be recognized.

[0025] The model building unit is used to build a visual information recognition model based on the pulse Transformer network; wherein, one or more basis functions, such as Fourier basis functions or wavelet basis functions, are used to build a prior knowledge embedding layer; the prior knowledge embedding layer replaces the attention module in the pulse Transformer network;

[0026] A visual information recognition unit is used to input the pulse sequence into the visual information recognition model to obtain the recognition result of the visual information to be recognized.

[0027] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the visual information recognition method based on the pulse Transformer network as described above.

[0028] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the visual information recognition method based on pulse Transformer networks as described above.

[0029] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the visual information recognition method based on a pulse Transformer network as described above.

[0030] The visual information recognition method based on pulsed Transformer networks provided by this invention constructs a visual information recognition model using a pulsed Transformer network and introduces Fourier basis functions or wavelet basis functions to construct a prior knowledge embedding layer to replace the attention module of the pulsed Transformer network. This improves the accuracy of visual information recognition while reducing the energy consumption of the visual information recognition process. Attached Figure Description

[0031] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0032] Figure 1This is one of the flowcharts of the visual information recognition method based on pulse Transformer network provided by the present invention.

[0033] Figure 2 This is a schematic diagram illustrating the principle of the visual information recognition method based on pulse Transformer network provided by the present invention.

[0034] Figure 3 This is the second flowchart of the visual information recognition method based on pulse Transformer network provided by the present invention.

[0035] Figure 4 This is a schematic diagram illustrating the principle of the visual information recognition model provided by the present invention.

[0036] Figure 5 This is a schematic diagram of the structure of the visual information recognition device based on the pulse Transformer network provided by the present invention.

[0037] Figure 6 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation

[0038] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0039] The terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such descriptions can be used interchangeably where appropriate to allow embodiments to be implemented in a sequence other than that illustrated or described in this application. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or device that includes a series of steps or modules is not necessarily limited to those explicitly listed, but may include other steps or modules not explicitly listed or inherent to such processes, methods, products, or devices. The naming or numbering of steps appearing in this application does not imply that the steps in the method flow must be performed in the chronological / logical order indicated by the naming or numbering. The execution order of named or numbered process steps can be changed according to the desired technical purpose, as long as the same or similar technical effect is achieved. The module division described in this application is a logical division. In practical applications, there may be other division methods. For example, multiple modules may be combined or integrated into another system, or some features may be ignored or not executed. In addition, the shown or discussed mutual coupling, direct coupling, or communication connection may be through some interface, and the indirect coupling or communication connection between units may be electrical or other similar forms, none of which are limited in this application. Furthermore, the modules or sub-modules described as separate components may or may not be physically separated, may or may not be physical units, or may be distributed in multiple circuit units. Some or all of the units can be selected to achieve the purpose of the solution in this application according to actual needs.

[0040] In this invention, the spiking Transformer network can achieve the function of visual information recognition. However, due to the discrete pulse encoding method of the spiking neural network for real values, the accuracy of the model using the spiking Transformer network for visual information recognition is low. To solve the problems existing in related technologies, this invention constructs a visual information recognition model based on an improved spiking Transformer network, thereby improving the accuracy of the visual information recognition model in recognizing visual information.

[0041] In this invention, the visual information recognition results can be applied to various fields, such as dynamic scenes, including moving object detection, tracking, and behavior analysis. The visual information recognition model captures and analyzes changes in the scene by recognizing real-time visual information, providing accurate dynamic information. The specific applications of the visual information recognition results in this embodiment will not be elaborated upon further.

[0042] The relevant concepts in the embodiments of the present invention will be explained below.

[0043] The Fourier transform is a mathematical method for converting signals from the time domain to the frequency domain. It is based on the concept of Fourier series, which shows that any periodic function can be represented as the sum of sine and cosine functions. The Discrete Fourier Transform (DFT) is a discrete version of the continuous Fourier Transform, used to process discrete signals. The Fast Fourier Transform (FFT) is an efficient algorithm for calculating the DFT, with a time complexity far lower than directly calculating the DFT.

[0044] Wavelet transform is a mathematical tool for signal analysis that provides a time-frequency representation of a signal. Unlike Fourier transform, wavelet transform allows for the analysis of local characteristics of a signal, making it very useful when dealing with non-stationary signals (signals whose frequency varies over time). Discrete wavelet transform (DWT) is a discrete version of continuous wavelet transform.

[0045] The following is combined with Figures 1-5 The specific contents of this invention are described below.

[0046] Figure 1 This is one of the flowcharts of the visual information recognition method based on pulse Transformer network provided by the present invention, including steps S101-S103.

[0047] Step S101: Obtain the pulse sequence of the visual information to be identified.

[0048] Step S102: Construct a visual information recognition model based on the pulse Transformer network; wherein, one or more basis functions, either Fourier basis functions or wavelet basis functions, are used to construct a prior knowledge embedding layer; the prior knowledge embedding layer is used to replace the attention module in the pulse Transformer network.

[0049] Specifically, the prior knowledge embedding layer can be a Fourier transform embedding layer or a wavelet transform embedding layer. The dimension of the prior knowledge embedding matrix is ​​consistent with the QK matrix in the self-attention module, and the value of each basis function remains fixed during training.

[0050] Step S103: Input the pulse sequence into the visual information recognition model to obtain the recognition result of the visual information to be recognized.

[0051] In one possible implementation, in step S101, obtaining the pulse sequence of the visual information to be identified includes: when the visual information to be identified is a static RGB image, inputting the static RGB image into the RGB encoding module to obtain the pulse sequence of the visual information to be identified.

[0052] Specifically, such as Figure 2As shown, a static RGB image is input to an RGB data encoding module to obtain an encoded pulse sequence, which is then used as input to a visual information recognition model. In an embodiment of the present invention, for a static RGB image, the time step of the pulse sequence is... The time step can be determined based on the size of the static RGB image. The time step is proportional to the size of the static RGB image. In embodiments of the present invention, the time step... It is generally set between 1 and 4.

[0053] In one possible implementation, in step S101, obtaining the pulse sequence of the visual information to be identified includes: when the visual information to be identified is DVS data, directly using the DVS data as the pulse sequence of the visual information to be identified.

[0054] In embodiments of the present invention, the DVS data (Dynamic Vision Sensor Dataset) is generated by an event camera, which generates events based on pixel-level brightness changes. Each event includes the pixel's position, a timestamp, and the polarity of the brightness change (brightening or darkening). Because DVS data records discrete events that change over time, it is essentially a pulse sequence, and therefore can be directly used as a pulse sequence of visual information to be identified.

[0055] Figure 3 This is the second flowchart illustrating the visual information recognition method based on pulse Transformer networks provided by the present invention.

[0056] In one possible implementation, in step S103, the pulse sequence is input into the visual information recognition model to obtain the recognition result of the visual information to be recognized, including steps S301-S303.

[0057] Step S301: Input the pulse sequence into the convolutional coding layer in the visual information recognition model for preliminary encoding.

[0058] Step S302: Input the pre-encoded pulse sequence into the deep coding layer of the visual information recognition model for deep coding.

[0059] Step S303: Input the deep-encoded pulse sequence into the classification layer of the visual information recognition model to obtain the recognition result of the visual information to be recognized.

[0060] In one possible implementation, the prior knowledge embedding layer is located within the deep coding layer.

[0061] Figure 4 This is a schematic diagram illustrating the principle of a visual information recognition model provided by the present invention.

[0062] The following is combined with Figure 4 The following provides a detailed explanation of steps S301-S303.

[0063] Specifically, in step S301, the convolutional coding layer consists of a two-dimensional convolutional layer, a normalization layer, a LIF spiking neuron layer, and a max pooling layer. First, the two-dimensional convolutional layer encodes the input pulse sequence; then, the normalization layer, the LIF spiking neuron layer, and the max pooling layer pulse the initially encoded pulse sequence.

[0064] Specifically, in step S302, the deep coding layer consists of a prior knowledge embedding layer, a normalization layer, a LIF spiking neuron layer, a residual layer, and a sparse connection layer.

[0065] Specifically, in step S303, the classification layer classifies the pulse sequence information obtained after depth encoding to obtain the recognition result of visual information.

[0066] In one possible implementation, for each module of the pulsatile Transformer network, a LIF neuron layer is added at the output of each module to simulate... Neuronal membrane potentials accumulate and pulses are fired, and the resulting pulse sequence is used as input for the next module.

[0067] In an embodiment of the present invention, the membrane potential of any LIF neuron in the visual information recognition model is updated based on the following formula:

[0068] (X) (V) ))

[0069]

[0070]

[0071] in, The membrane potential time constant, for Input current at any given time for The initial neuronal membrane potential at time t. This is the resting potential. For the updated membrane potential, as The initial membrane potential at time t. The membrane potential after the last update, as The initial membrane potential at time t. The function is used to confirm the flag bit. The value of .

[0072] when Greater than the issuance threshold At that time, the flag bit =1, confirming the output pulse delivery.

[0073] Dangdang Less than or equal to the issuance threshold At that time, the flag bit =0, no output pulse is emitted.

[0074] In embodiments of the present invention, the prior knowledge embedding layer replaces the self-attention module in a conventional pulse Transformer network.

[0075] In one possible implementation, a prior knowledge embedding layer is constructed using one or more basis functions, either Fourier basis functions or wavelet basis functions, including:

[0076] A prior knowledge embedding matrix can be constructed using sine and cosine functions of different frequencies from the Fourier basis functions to obtain the prior knowledge embedding layer; or a prior knowledge embedding matrix can be constructed using cosine functions of different frequencies from the Fourier basis functions to obtain the prior knowledge embedding layer.

[0077] In the embodiments of the present invention, the Fourier basis functions have globality and orthogonality. The advantage of the prior knowledge embedding layer constructed based on the above Fourier basis functions lies mainly in the extraction of global information.

[0078] In one possible implementation, a prior knowledge embedding layer is constructed using one or more basis functions, either Fourier basis functions or wavelet basis functions, including:

[0079] A prior knowledge embedding matrix is ​​constructed using one or more discrete wavelet basis functions to obtain the prior knowledge embedding layer.

[0080] In embodiments of the present invention, one or more discrete wavelet basis functions from Biorthogonal 1.1 Wavelet, HaarWavelet, and Daubechies 1 Wavelet are used as a combination to construct a prior knowledge embedding matrix to obtain the prior knowledge embedding layer. All three discrete wavelet basis functions possess orthogonality and compact support, and the advantage of wavelet basis functions lies in the extraction of local information, making them more suitable for pulse signal processing.

[0081] In one possible implementation, a prior knowledge embedding layer is constructed using one or more basis functions, either Fourier basis functions or wavelet basis functions, including:

[0082] When using multiple Fourier basis functions to construct a prior knowledge embedding layer, initial weighting parameters are configured for each Fourier basis function, and these initial weighting parameters are continuously adjusted as the visual information recognition model is trained.

[0083] When using multiple wavelet basis functions to construct a prior knowledge embedding layer, initial weighting parameters are configured for each Fourier basis function, and these initial weighting parameters are continuously adjusted as the visual information recognition model is trained.

[0084] This invention, by employing the above method, achieves the following beneficial effects: It constructs a visual information recognition model using a Transformer network based on a pulse framework, and introduces Fourier basis functions or wavelet basis functions to construct a prior knowledge embedding layer to replace the attention module of the pulse Transformer network. This improves the accuracy of visual information recognition while reducing the energy consumption of the visual information recognition process.

[0085] The visual information recognition device based on the pulse Transformer network provided by the present invention will be described below. The visual information recognition device based on the pulse Transformer network described below can be referred to in correspondence with the visual information recognition method based on the pulse Transformer network described above.

[0086] Figure 5 A schematic diagram of the structure of the visual information recognition device based on the pulse Transformer network provided in the embodiments of this application includes:

[0087] The visual information processing unit 510 is used to acquire the pulse sequence of visual information to be recognized.

[0088] Model building unit 520 is used to build a visual information recognition model based on the pulse Transformer network; wherein, one or more basis functions, either Fourier basis functions or wavelet basis functions, are used to build a prior knowledge embedding layer; the prior knowledge embedding layer replaces the attention module in the pulse Transformer network.

[0089] The visual information recognition unit 530 is used to input the pulse sequence into the visual information recognition model to obtain the recognition result of the visual information to be recognized.

[0090] In one possible implementation, the visual information recognition unit 530 is specifically used to input the pulse sequence into the convolutional coding layer in the visual information recognition model for preliminary encoding; input the pre-encoded pulse sequence into the deep coding layer in the visual information recognition model for deep encoding; and input the deep-encoded pulse sequence into the classification layer in the visual information recognition model to obtain the recognition result of the visual information to be recognized.

[0091] In one possible implementation, the prior knowledge embedding layer is located within the deep coding layer.

[0092] In one possible implementation, the visual information processing unit 510 is specifically used to input the static RGB image into the RGB encoding module to obtain the pulse sequence of the visual information to be recognized when the visual information to be recognized is a static RGB image.

[0093] In one possible implementation, the visual information processing unit 510 is specifically used to directly use the DVS data as the pulse sequence of the visual information to be identified when the visual information to be identified is DVS data.

[0094] In one possible implementation, the model building unit 520 is specifically used to construct a prior knowledge embedding matrix using sine and cosine functions of different frequencies in the Fourier basis functions to obtain a prior knowledge embedding layer; or to construct a prior knowledge embedding matrix using cosine functions of different frequencies in the Fourier basis functions to obtain a prior knowledge embedding layer.

[0095] In one possible implementation, the model building unit 520 is specifically used to construct a prior knowledge embedding matrix using one or more discrete wavelet basis functions to obtain a prior knowledge embedding layer.

[0096] Figure 6 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 6 As shown, the electronic device may include a processor 610, a communications interface 620, a memory 630, and a communication bus 640. The processor 610, communications interface 620, and memory 630 communicate with each other via the communication bus 640. The processor 610 can call logical instructions from the memory 630 to execute a visual information recognition method based on a pulse Transformer network, including the following steps:

[0097] Acquire the pulse sequence of the visual information to be identified;

[0098] A visual information recognition model is constructed based on the pulsating Transformer network. A prior knowledge embedding layer is constructed using one or more basis functions, either Fourier or wavelet basis functions. This prior knowledge embedding layer replaces the attention module in the pulsating Transformer network.

[0099] The pulse sequence is input into the visual information recognition model to obtain the recognition result of the visual information to be recognized.

[0100] Furthermore, the logical instructions in the aforementioned memory 630 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, essentially, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0101] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can perform a visual information recognition method based on a pulse Transformer network provided by the above methods, including the following steps:

[0102] Acquire the pulse sequence of the visual information to be identified;

[0103] A visual information recognition model is constructed based on the pulsating Transformer network. A prior knowledge embedding layer is constructed using one or more basis functions, either Fourier or wavelet basis functions. This prior knowledge embedding layer replaces the attention module in the pulsating Transformer network.

[0104] The pulse sequence is input into the visual information recognition model to obtain the recognition result of the visual information to be recognized.

[0105] In another aspect, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements a visual information recognition method based on a pulse Transformer network provided by the methods described above, comprising the following steps:

[0106] Acquire the pulse sequence of the visual information to be identified;

[0107] A visual information recognition model is constructed based on the pulsating Transformer network. A prior knowledge embedding layer is constructed using one or more basis functions, either Fourier or wavelet basis functions. This prior knowledge embedding layer replaces the attention module in the pulsating Transformer network.

[0108] The pulse sequence is input into the visual information recognition model to obtain the recognition result of the visual information to be recognized.

[0109] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0110] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0111] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A visual information recognition method based on a pulse Transformer network, characterized in that, include: Acquire the pulse sequence of the visual information to be identified; A visual information recognition model is constructed based on a pulse Transformer network. A prior knowledge embedding layer is constructed using one or more Fourier basis functions. The prior knowledge embedding layer replaces the attention module in the pulse Transformer network. Constructing the prior knowledge embedding layer using one or more Fourier basis functions includes: constructing a prior knowledge embedding matrix using sine and cosine functions of different frequencies from the Fourier basis functions to obtain the prior knowledge embedding layer; or constructing a prior knowledge embedding matrix using cosine functions of different frequencies from the Fourier basis functions to obtain the prior knowledge embedding layer. The pulse sequence is input into the visual information recognition model to obtain the recognition result of the visual information to be recognized.

2. The visual information recognition method based on pulse Transformer network according to claim 1, characterized in that, The step of inputting the pulse sequence into the visual information recognition model to obtain the recognition result of the visual information to be recognized includes: The pulse sequence is input into the convolutional coding layer of the visual information recognition model for preliminary encoding. The initially encoded pulse sequence is input into the deep coding layer of the visual information recognition model for deep encoding. The deep-encoded pulse sequence is input into the classification layer of the visual information recognition model to obtain the recognition result of the visual information to be recognized.

3. The visual information recognition method based on pulse Transformer network according to claim 2, characterized in that, The prior knowledge embedding layer is located in the deep coding layer.

4. The visual information recognition method based on pulse Transformer network according to claim 1, characterized in that, The pulse sequence for acquiring the visual information to be identified includes: When the visual information to be identified is a static RGB image, the static RGB image is input into the RGB encoding module to obtain the pulse sequence of the visual information to be identified.

5. The visual information recognition method based on pulse Transformer network according to claim 1, characterized in that, The pulse sequence for acquiring the visual information to be identified includes: When the visual information to be identified is DVS data, the DVS data is directly used as the pulse sequence of the visual information to be identified.

6. A visual information recognition device based on a pulse Transformer network, characterized in that, include: A visual information processing unit is used to acquire pulse sequences of visual information to be recognized. A model building unit is used to construct a visual information recognition model based on a pulse Transformer network. Specifically, it employs one or more Fourier basis functions to construct a prior knowledge embedding layer; the prior knowledge embedding layer replaces the attention module in the pulse Transformer network. Constructing the prior knowledge embedding layer using one or more Fourier basis functions includes: constructing a prior knowledge embedding matrix using sine and cosine functions of different frequencies from the Fourier basis functions to obtain the prior knowledge embedding layer; or constructing a prior knowledge embedding matrix using cosine functions of different frequencies from the Fourier basis functions to obtain the prior knowledge embedding layer. A visual information recognition unit is used to input the pulse sequence into the visual information recognition model to obtain the recognition result of the visual information to be recognized.

7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the visual information recognition method based on the pulse Transformer network as described in any one of claims 1 to 5.

8. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the visual information recognition method based on the pulse Transformer network as described in any one of claims 1 to 5.