A Memristor-Based Text Sentiment Detection System and Method
Through the text sentiment detection system based on memristor, the memory wall problem of text sentiment detection on embedded devices is solved, the number of network parameters is reduced and the computing architecture is optimized, realizing low-power text sentiment detection and efficient feature extraction.
Patent Information
- Application Number
- CN202310091466.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-06
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2043-02-06
AI Technical Summary
The existing text sentiment detection system faces memory wall bottlenecks when deployed on embedded devices, making it difficult to effectively pay attention to the internal connection between the local and overall conversations, and the huge number of network parameters, resulting in excessive computing resource overhead.
A text emotion detection system based on memristors is designed, including a preprocessing module, an attention calculation module and an emotion classification module. The attention calculation module and linear mapping operation are constructed using memristor cross-arrays. Combined with CMOS technology, the number of network parameters is reduced and the computing architecture is optimized.
It realizes low-power text sentiment detection on embedded devices, reducing the number of network parameters by three times, and improving the global and local feature extraction capabilities of dialogue text information, reducing additional circuit area overhead.
Smart Images

Figure CN115994221B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of text sentiment recognition, and particularly to a text sentiment detection system and method based on a memristor. Background Art
[0002] With the rapid development of Web 2.0 and 5G networks, a large number of social media on the Internet have affected all aspects of people's lives. Due to the increase in the number of users, a large amount of data in different formats has been uploaded to the network for sharing, including photos, voices, and texts. Therefore, a large amount of available text data has recently received a lot of attention.
[0003] Due to the expansion of NLP (sentiment analysis, machine reading comprehension) tasks, deep neural networks (DNNs) have been widely used to process text data, including recurrent neural networks (RNNs) and long short-term memory (LSTM) networks. Although both RNN and LSTM networks have solved some problems in text information processing, they still cannot fully understand the context relationship. Therefore, attention-based networks are used to implement NLP tasks and perform better than RNN and LSTM networks in processing context information. The novel network model Transformer, which uses a multi-head attention mechanism as an encoder and a decoder, is becoming a computing paradigm in NLP. Recently, several works have applied the pre-trained BERT model, which is a variant of Transformer commonly used in NLP, to detect emotions in dialogue tasks and obtained a good score. However, due to the rapid increase in the parameters of attention-based network models, more memory and computing resources are required when deploying attention-based networks to embedded devices. On the other hand, the classic von Neumann architecture separates the computing and memory units, which is called the "memory wall", resulting in an increase in energy overhead and computing latency when performing a large number of data transmissions. This indicates that implementing a text sentiment detection system in practical applications will be severely hindered by the memory wall. In particular, social media and other applications can already be applied to embedded devices such as mobile phones, tablets, and even smart watches. Therefore, deploying a text sentiment detection system on embedded devices is a very promising direction for embedded applications. However, in order to overcome the bottleneck of the memory wall, a new computing architecture needs to be established for the embedded sentiment detection system. For this purpose, the present embodiment provides a feasible solution based on a memristive neuromorphic computing system.
[0004] The memristor is an ideal in-memory computing device because it is a two-terminal device with characteristics such as dynamic resistance, low power consumption, and easy integration. In addition, memristive crossbars have been proven to be useful for implementing the hardware deployment of convolutional neural networks (CNNs) and are also widely used in the development of various intelligent applications. However, only a few hardware solutions have been proposed for implementing attention-based architectures. Currently, a character recognition task has been successfully completed using a memristor circuit, demonstrating that memristors can be used to deploy attention-based networks on hardware. But there is still much work to be done. First, the memristor hardware implementation solution for multi-layer attention networks is still vacant. In addition, due to the large number of network parameters and the overhead of computing resources, it is still a huge challenge to implement text sentiment detection on embedded devices.
[0005] Briefly, sentiment detection in text conversations has a wide range of applications in the field of human-computer interaction. Different from detecting a single sentence, detecting the potential emotions in dialogue text requires modeling the dependence between context and local information. However, traditional sentiment detection networks do not pay attention to the internal connection between the local and the whole of the dialogue. At the same time, due to the huge number of parameters in the sentiment detection network and the huge scale of text information, it is still challenging to deploy the text sentiment detection system on embedded devices currently. Summary of the Invention
[0006] The present invention provides a text sentiment detection system and method based on memristors, and the technical problems to be solved are: how to pay attention to the internal connection between the local and the whole of the dialogue, and how to reduce the number of network parameters and the system area.
[0007] To solve the above technical problems, the present invention provides a text sentiment detection system based on memristors, including a preprocessing module, an attention calculation module, and a sentiment classification module;
[0008] The preprocessing module is used to input the dialogue text data and perform positional encoding and token embedding on the dialogue text data to obtain the embedded text data;
[0009] The attention calculation module includes a multi-layer attention circuit and a forward propagation circuit; the attention circuit of the first layer is used to extract the multi-head attention of the embedded text data, and the forward propagation circuit is used to propagate the output of the attention circuit of the current layer to the attention circuit of the next layer as input until the last layer of the attention circuit, and the final attention output is sent to the sentiment classification module;
[0010] The emotion classification module includes a max pooling circuit, a third layer regularization circuit, a first pointwise convolution circuit, a depth convolution circuit, a Swish activation function circuit, a second pointwise convolution circuit, a fourth layer regularization circuit, a fully connected layer circuit, and a second SoftMax activation function circuit connected in sequence, which are used to perform corresponding max pooling, regularization, pointwise convolution, depth convolution, Swish activation function calculation, pointwise convolution, regularization, fully connected, and SoftMax activation function calculation operations on their respective inputs, and finally output the emotion detection result;
[0011] The attention extraction operation, weight mapping operation in the attention calculation module, the first pointwise convolution circuit, the depth convolution circuit, the second pointwise convolution circuit, and the fully connected layer circuit are all constructed based on a memristive crossbar array. Every two columns of the memristive crossbar array correspond to positive and negative weights in the neural network, and the output of every two columns corresponds to an output voltage.
[0012] Specifically, the attention circuit includes a multi-head attention circuit constructed based on a memristive crossbar array, a first memory cell module circuit for storing voltage signals, a first multiply-accumulate circuit and a second multiply-accumulate circuit connected to the first memory cell module circuit, and also includes a first SoftMax activation function circuit. The multi-head attention circuit is used to perform multi-head attention extraction on the embedded text data and output multi-head Q, K, and V signals. The first multiply-accumulate circuit is used to perform multiply-accumulation on the multi-head Q and K signals to obtain the total Q and K signals. The second multiply-accumulate circuit is used to perform multiply-accumulation on the multi-head V signals to obtain the total V signal. The first SoftMax activation function circuit is used to calculate the total Q and K signals using the softmax function, and the calculation result is multiplied by the total V signal with weights and then output to the forward propagation circuit;
[0013] The output of the attention circuit is expressed as:
[0014]
[0015] Among them, Concat{} represents concatenation, softmax() represents the softmax function, d h = d / h represents the dimension of the head, h represents the number of heads, d is the dimension of the input Q, K, V, Q h 、K h 、V h respectively represent the Q, K, V values of the single-head output, and W Z represents the mapping matrix.
[0016] Specifically, the forward propagation circuit includes a first weight mapping circuit, a first layer regularization circuit, a second weight mapping circuit, a third weight mapping circuit, and a second layer regularization circuit connected in sequence, and further includes a second memory unit module circuit; wherein the second layer regularization circuit is used to connect the attention circuit of the next layer, and the second memory unit module circuit is used to store the attention obtained in each calculation during the forward propagation process.
[0017] The first weight mapping circuit, the second weight mapping circuit, and the third weight mapping circuit are all constructed based on the memristive crossbar array.
[0018] The forward propagation process of the forward propagation circuit is expressed as:
[0019] Forward = LN{W3[W2(LN(W1·Attention + b1)) + b2] + b3},
[0020] where Forward represents the output of this forward propagation, Attention represents the input of this forward propagation, W1 and b1 respectively represent the weight and bias of the first weight mapping circuit, W2 and b2 respectively represent the weight and bias of the second weight mapping circuit, W3 and b3 respectively represent the weight and bias of the third weight mapping circuit, and LN() represents the regularization function.
[0021] The output of a single channel of the depth convolution circuit is expressed as:
[0022]
[0023] where k is the kernel width of the depth convolution, represents the kernel weight, represents the input of the depth convolution circuit, and U represents the number of channels;
[0024] The output of the depth convolution circuit is expressed as:
[0025]
[0026] The output of the second pointwise convolution circuit is expressed as:
[0027]
[0028] W pw1 、W pw2 and W dw respectively represent the weights of the convolution kernels of the first pointwise convolution circuit, the second pointwise convolution circuit, and the depth convolution circuit, W P1 represents the weight of the third layer regularization circuit, Max(Vatten m,n) represents the output result of the max pooling circuit, and Swish() represents the Swish function.
[0029] Specifically, the sentiment classification module further includes a third memory unit module circuit connected between the max pooling circuit and the third layer regularization circuit, a fourth memory unit module circuit connected between the first pointwise convolution circuit and the depth convolution circuit, and a fifth memory unit module circuit connected between the first pointwise convolution circuit and the fourth layer regularization circuit, which is used to store the output signals of the upper-level circuit and provide them to the lower-level circuit.
[0030] Specifically, the first memory unit module circuit, the second memory unit module circuit, the third memory unit module circuit, the fourth memory unit module circuit, and the fifth memory unit module circuit are all constructed based on the memory unit module circuit. The memory unit module circuit includes a plurality of memory unit circuits connected in an array, and the number of memory unit circuits is the same as the number of voltage signals to be stored;
[0031] For the memory unit circuits in each row of the memory unit module circuit, all the voltage signal input terminals are connected together, all the voltage signal output terminals are connected together, all the control terminals of the first MOS switches are connected together, and all the control terminals of the second MOS switches are connected together; the memory unit module circuit inputs and outputs voltages in sequence according to the clock beats.
[0032] Specifically, the memory unit circuit includes a first MOS switch, a second MOS switch, a first operational amplifier, and a capacitor. The control terminal of the first MOS switch and the control terminal of the second MOS switch are connected to a clock control signal. The inverting input terminal of the first operational amplifier is connected to the output terminal of the first MOS switch. The non-inverting input terminal of the first operational amplifier is connected to the capacitor and then grounded. The output terminal of the first operational amplifier is connected to the input terminal of the second MOS switch. The input terminal of the first MOS switch is used as the voltage signal input terminal of the memory unit circuit, and the output terminal of the second MOS switch is used as the voltage signal output terminal of the memory unit circuit;
[0033] Specifically, the memristive crossbar array adopts a memristor model with a 2M structure.
[0034] Specifically, both the Friends and EmotionPush datasets are used to train the network to train this text sentiment detection system, and weighted cross-entropy is used as the training loss.
[0035] The present invention also provides a text sentiment detection method based on memristors, including: constructing a text sentiment detection system, training and testing the constructed text sentiment detection system, inputting the text to be detected into the trained text sentiment detection system, and the text sentiment detection system outputs a sentiment detection result.
[0036] The main contributions of a text sentiment detection system and method based on memristors provided by the present invention are as follows:
[0037] 1. A text sentiment detection system (called MTEDS) with co - design of software and hardware is proposed, realizing the hardware deployment of the network model MLA - DSCN. By constructing the linear mapping and matrix multiplication operations in the system using a memristive cross - array, and the remaining circuit modules are also formed using CMOS technology. Compared with the traditional computing architecture (GPU), the hardware solution designed by the present invention has lower power consumption and reduces the additional circuit area overhead caused by the separation of memory and computing units;
[0038] 2. A lightweight network model called MLA - DSCN is proposed to implement the text sentiment detection task. By integrating the depth - separable convolution module into the network, the improved network model MLA - DSCN can not only focus on the global information of the input text information but also improve its information extraction ability at the local level. Without sacrificing accuracy, the number of network parameters is reduced by 3 times. Description of the Drawings
[0039] Figure 1 It is a diagram showing the change of memristance values during the programming process provided by an embodiment of the present invention;
[0040] Figure 2 It is a co - design idea diagram of software and hardware of MTEDS provided by an embodiment of the present invention;
[0041] Figure 3 It is a system framework diagram and application diagram of MTEDS provided by an embodiment of the present invention;
[0042] Figure 4 It is a structural diagram of the text sentiment detection module provided by an embodiment of the present invention;
[0043] Figure 5 It is a circuit diagram of MMLAN provided by an embodiment of the present invention;
[0044] Figure 6 It is a schematic diagram of a 2M - structure memristive cross - array provided by an embodiment of the present invention;
[0045] Figure 7 It is a circuit diagram of an analog MUM provided by an embodiment of the present invention;
[0046] Figure 8It is the circuit diagram of MEC provided by the embodiment of the present invention;
[0047] Figure 9 It is the display diagram of the input and output of each layer of the network in the training stage provided by the embodiment of the present invention;
[0048] Figure 10 It is the analysis diagram of the activation function circuit provided by the embodiment of the present invention;
[0049] Figure 11 It is the performance display diagram of the MUM circuit provided by the embodiment of the present invention;
[0050] Figure 12 It is the simulation result diagram of the memristor under different LHRs provided by the embodiment of the present invention;
[0051] Figure 13 It is the trade-off analysis diagram of different memory module sizes provided by the embodiment of the present invention. Detailed implementation manners
[0052] The following specifically clarifies the implementation manners of the present invention in conjunction with the attached drawings. The given embodiments are only for illustrative purposes and should not be construed as limitations on the present invention. The attached drawings are only for reference and illustration and do not constitute limitations on the protection scope of the patent of the present invention, because many changes can be made to the present invention without departing from the spirit and scope of the present invention.
[0053] With the development of artificial intelligence, the attention mechanism has gradually become an important part of neural networks. The traditional single-head attention can be calculated in the following way:
[0054]
[0055] Assume that the input feature information X is transferred to the query matrix Q, key matrix K, and value matrix V after word embedding, where d is the dimension of the input Q, K, and V, and SoftMax() represents the SoftMax function.
[0056] The multi-head attention is concatenated (Concat) by multiple single-head attentions and then mapped to the final output by the W Z matrix. In addition, the calculation process is shown in formula (2), where h indicates the number of heads, and the dimension of each head d h = d / h.
[0057]
[0058] In addition, the multi-head attention mechanism is just a dimensional extension of the single-head attention mechanism. Therefore, the hardware implementation strategy of the single-head attention mechanism can be used to prove the effectiveness of the multi-head attention mechanism.
[0059] Before the emergence of the attention mechanism, one of the benchmarks for handling sentiment detection in conversations was the autoencoder sentiment classifier based on CNN-DCN. As attention mechanisms gradually achieved remarkable results in different learning fields, they began to be used in sentiment detection tasks. In addition, it was found that pre-trained language models are more suitable for handling NLP tasks (sentiment analysis, machine reading comprehension) because they are pre-trained on a substantial unlabeled corpus to provide deep context embeddings, which makes the methods developed on top of them perform better. In addition, the embedding techniques for sentences and conversations, as well as the choice of classifier, have a substantial impact on the final sentiment detection results under the same pre-trained model architecture. Traditional embedding methods in NLP include Bag-of-Word (BOW), term frequency-inverse document frequency (TFIDF), and WordPiece. In addition, classification methods include Random Forest (RF), TextCNN, whose initial word embeddings are GloVe, and a classifier implemented through the SELU activation function. In addition, CNN has recently been widely used in the field of attention-based deep learning and has achieved SOTA results in most sub-field tasks.
[0060] Therefore, in this embodiment, for all input corpora in different channels, pointwise convolution with a kernel size of 1×1 is adopted. During the depth convolution process, all features in the same channel have to go through the convolution program. The output of the depthwise separable convolution module is denoted as DS U , where the kernel weights U is the output dimension of the depth convolution, k is the kernel width of the depth convolution, and X is the input feature.
[0061]
[0062] In many works, the swish activation function shown in Equation (4) below has achieved better results than the ReLu activation function. The hyperparameter β determines the output form of the function. When β = 0.1, the swish function is close to a linear function, and when β = ∞, it is similar to the ReLu activation function. x represents the input feature, and sigmoid() represents the sigmoid function.
[0063] Swish(x) = x · sigmoid(βx) (4)
[0064] Many memristor models have been proposed. To simulate artificial neurons, it is expected that memristors can process pulse signals. Literature 1 (Y. Zhang, X. Wang, Y. Li, and E. G. Friedman, “Memristive model for synaptic circuits,” IEEE Trans. Circuits Syst. II, Exp. Briefs, vol. 64, no. 7, pp. 767–771, Jul. 2017.) presents a voltage-controlled threshold model where the memristance R T- ,V T+ changes only when the applied voltage exceeds the voltage threshold [V m . In addition, the expression of the threshold characteristic is related to time t, where W[R m (t)] represents the window function, and the memristance is expressed as the minimum and maximum memristances [R on ,R off .
[0065] The memristor model is shown as follows:
[0066]
[0067]
[0068] where μ v represents the average ion mobility, i0 is a constant, i on and i off represent the currents in the corresponding on and off states respectively, and D represents the length of the memristor model. It is worth mentioning that these parameters jointly determine the characteristics of the memristor, where the magnitude of the switching ratio R on / R off directly affects the mapping relationship between the neural network and the memristor, as well as the energy consumption of a single memristor. In addition, the parameter configuration of the peripheral circuit is thus determined. To explore the influence of different switching ratios on the memristor, Table 1 shows three different switching ratio-based AIST memristors, where LHR represents the switching ratio R on / R off . Figure 1 Further illustrates the change of the memristor over time, showing the change of the memristance during the programming process. Pulses with amplitudes of 1V, 1.5V, and 2V are applied to both sides of the memristor respectively within a 0.5s period. In an ideal case, it can be configured in a graded manner. However, it is worth noting that it is difficult for actual memristor devices to achieve any level of memristance between the on and off states, so this limitation and influence are considered in the performance analysis.
[0069] Table 1 Simulation configuration of the memristor model
[0070]
[0071] For embedded devices, the design of hardware and software is inseparable. Figure 2 Illustrates the co - design strategy of the present invention. On the software side, an improved attention - based network model is designed and proposed using the PyTorch platform to achieve a higher F1 score and more accurate emotion detection, and then the trained model is mapped to the corresponding hardware to provide a practical solution for embedded applications. However, due to the large number of parameters of the attention network and its specific large number of multiply - accumulate operations, a new type of computing architecture: the memristive neuromorphic computing system, is considered an ideal framework for hardware deployment due to its ideal in - memory computing characteristics.
[0072] In addition, considering the impact of limited physical hardware resources, this means that additional energy and time may be required for signal processing. Finally, linear mapping realizes the uploading of network model weights. It should be noted that the bit precision of the memristor model and the power consumption of a single neuron affect the entire system in the actual hardware implementation. Therefore, the trained network needs to be further quantized as the input for implementing MTEDS deployment to embedded devices.
[0073] Figure 3 Shows the entire framework of MTEDS and illustrates the process of performing emotion computing and human - machine interaction. This example shows that MTEDS implanted in an embedded device can help the machine understand the emotions in human conversations and respond to the corresponding emotional expressions. The improved model MLA - DSCN is mainly implemented in two parts: the attention calculation module (MMLAN) and the memristive emotion classification module (MEC). Figure 4 Illustrates the corresponding network model structure of MTEDS. The model is correspondingly segmented at the network layer to facilitate the design of the hardware circuit.
[0074] From a macroscopic perspective, convolution has been proven to effectively solve the problem that the attention mechanism cannot focus on local features. In addition, using a multi-layer attention network can better simulate the context information of text in NLP. However, to facilitate the deployment of embedded devices, lighter convolution operations are considered a promising option combined with multi-layer attention networks and are implemented for the first time in this embodiment. The proposed network MLA-DSCN allows the model to extract context feature information more comprehensively and performs particularly well when processing dialogue-based text rather than single sentences. From a microscopic perspective, a deep convolutional layer and two pointwise convolutional layers are added to the MEC. The present invention does not use the GLU layer because the addition of GLU reduces the amount of feature information extracted, and the difference comes from the max pooling layer added to the top of the MEC. This means that the global information calculated by the attention has been further extracted without additional operations to filter out redundant information. In addition, a corresponding hardware implementation scheme is designed based on the improved network. In the sentiment detection program, the dialogue is first embedded and encoded by a computer during the preprocessing process. Further, the embedded text information is input into the MMLAN and then sent to the MEC to identify the final detection results of emotions, including natural, happy, sad, and angry. Therefore, MTEDS can be widely deployed on embedded devices and applied as an agent platform in fields such as social media, opinion analysis robots, and emotion stewards.
[0075] As Figure 3 shown, an memristor-based text sentiment detection system provided by an embodiment of the present invention generally includes a preprocessing module (for data preprocessing, position embedding, and word embedding), an attention calculation module (for attention calculation), and a sentiment classification module (for sentiment classification).
[0076] As Figure 4 shown, the attention calculation module includes a multi-layer attention circuit and a forward propagation circuit; the attention circuit of the first layer is used to extract the multi-head attention of the embedded text data, and the forward propagation circuit is used to propagate the output of the attention circuit of the current layer to the attention circuit of the next layer as input until the attention circuit of the last layer, and the final attention output is sent to the sentiment classification module.
[0077] In this embodiment, the MMLAN is implemented by fusing memristor and CMOS technologies. Figure 5 Shows the detailed hardware scheme implementation strategy of the MMLAN. As Figure 5As shown in the figure, the attention circuit includes a multi-head attention circuit constructed based on a memristive crossbar array, a first memory cell module circuit for storing voltage signals, a first multiply-accumulate circuit and a second multiply-accumulate circuit connected to the first memory cell module circuit, and further includes a first SoftMax activation function circuit. The multi-head attention circuit is used to perform multi-head attention extraction on the embedded text data to output multi-head Q, K, and V signals. The first multiply-accumulate circuit is used to perform multiply-accumulation on the multi-head Q and K signals to obtain total Q and K signals. The second multiply-accumulate circuit is used to perform multiply-accumulation on the multi-head V signals to obtain total V signals. The first SoftMax activation function circuit is used to calculate the total Q and K signals using the softmax function, and the calculation result is multiplied by the total V signal with weights and then output to the forward propagation circuit.
[0078] As Figure 4 , Figure 5 shown in the figure, the forward propagation circuit includes a first weight mapping circuit, a first layer regularization circuit, a second weight mapping circuit, a third weight mapping circuit and a second layer regularization circuit connected in sequence, and further includes a second memory cell module circuit. The second layer regularization circuit is used to connect to the attention circuit of the next layer, and the second memory cell module circuit is used to store the attention calculated each time during the forward propagation process. Among them, the first weight mapping circuit, the second weight mapping circuit, and the third weight mapping circuit are all constructed based on a memristive crossbar array.
[0079] It is worth mentioning that in order to reduce the energy consumption during the digital-to-analog conversion (DAC) and analog-to-digital conversion (ADC) processes, the MUM (memory cell module circuit) temporarily stores the input voltage and sequentially completes the subsequent two matrix multiplication operations. Three different memristive crossbar arrays are used to represent Q, K, and V, where K and V usually have the same weight distribution in the attention calculation. In previous work, K and V were deployed on one memristive crossbar array to reduce power consumption and circuit area. However, in this embodiment, three different memristive crossbar arrays and one MUM are used to ensure that the stored intermediate voltage signals are not affected by the transmission between layers, and reduce the signal distortion effects caused by the parallel and series connections of other circuit modules, further improving the speed of the hardware for processing large-scale feature information and reducing the circuit complexity. In addition, the linear mapping is also achieved by loading the weights onto the memristive crossbar array. However, due to the integration limitations of the memristive crossbar array in practice, the designed hardware circuit cannot be deployed in the ideal size d of the software design. In other words, to complete the feature extraction of a conversation, multiple clock cycles are required. Therefore, it is necessary to have an additional MUM at the end of the MMLAN circuit to save the intermediate information and ensure that a complete conversation information has been processed before being sent to the MEC circuit.
[0080] The output of the memristive crossbar array depends on the type of the memristor model and the structure of the memristive crossbar array. The memristive crossbar array frameworks widely used in neural networks mainly include 1M, 2M, and 1T1M. In this embodiment, the 2M structure (see Document 2: M. Prezioso et al., “Training and operation of an integrated neuromorphic network based on metal-oxide memristors,” Nature, vol. 521, no. 7550, pp. 61–64, 2015) is used to implement linear mapping and matrix multiplication operations.
[0081] Figure 6 illustrates the specific process of the input voltage signal where i represents the length of the input signal, h is the number of heads, and each complete attention layer corresponds to one head. Every two columns of the memristive crossbar array correspond to the positive and negative weights in the neural network, and the outputs of two columns of memristors correspond to an output voltage where j represents the dimension of the output signal. In MMLAN, the parameters of the attention layer are Q, K, and In addition, since memristors cannot represent negative values, each pair of memristors is used to form a neuron G i,j , whose conductance value represents a weight. In addition, since the final outputs of every two columns of the memristive crossbar array are currents and Then, the currents are converted into voltages through a volt-ampere converter, and the intermediate voltages are stored in the MUM. Specifically, Formulas (7) and (8) show the entire process of voltage conversion. Where n and m represent the nth row and the mth th column of the crossbar array respectively. Rf represents the resistance value of the resistor in the circuit.
[0082]
[0083]
[0084] Figure 7The circuit diagram of the memory cell module circuit MUM is shown. Specifically, the memory cell module circuit includes a plurality of memory cell circuits connected in an array, and the number of memory cell circuits is the same as the number of voltage signals to be stored. The memory cell circuit includes a first MOS switch, a second MOS switch, a first operational amplifier, and a capacitor. The control terminal of the first MOS switch is connected to the control terminal of the second MOS switch to receive a clock control signal. The inverting input terminal of the first operational amplifier is connected to the output terminal of the first MOS switch. The non-inverting input terminal of the first operational amplifier is connected to the capacitor and then grounded. The output terminal of the first operational amplifier is connected to the input terminal of the second MOS switch. The input terminal of the first MOS switch is used as the voltage signal input terminal of the memory cell circuit, and the output terminal of the second MOS switch is used as the voltage signal output terminal of the memory cell circuit. For each row of memory cell circuits in the memory cell module circuit, all the voltage signal input terminals are connected together, all the voltage signal output terminals are connected together, all the control terminals of the first MOS switches are connected together, and all the control terminals of the second MOS switches are connected together. The memory cell module circuit sequentially inputs and outputs voltages according to the rhythm of the clock.
[0085] Due to the adoption of MUM, the energy consumption of the ADC and DAC can be effectively avoided, and during the calculation process, the intermediate voltage signal can be temporarily stored in a stable manner. Four wires are connected in each memory cell, where CL c (=CL) and CR c (=CR) are clock signals used to control two NMOS switches on each side of the operational amplifier. In addition, each memory cell sequentially inputs and outputs voltages according to the rhythm of the clock, and then transmits them through wires RI r and RO r respectively. It is worth mentioning that the capacitor C plays an important role in MUM. It not only stores the intermediate voltage but also converts the size of the input signal. In addition, the operational amplifier maintains the temporarily stored intermediate voltage in a stable state to avoid signal distortion.
[0086] The neural network realizes a large number of forward propagation processes by introducing functions in different layers of the MMLAN, including linear mapping (corresponding to the first weight mapping circuit, the second weight mapping circuit, and the third weight mapping circuit), layer regularization (corresponding to the first layer regularization circuit and the second layer regularization circuit), and the SoftMax activation function (corresponding to the first SoftMax activation function circuit). Therefore, the corresponding functional circuits will be described sequentially below.
[0087] The SoftMax activation function is mainly used for the classification of multi-class networks, converting the outputs of all classes into a probability distribution within the range of [0,1], and the sum of which is 1. As Figure 5As shown, the first SoftMax activation function circuit is composed of multiple exponential function circuits. The SoftMax activation function circuit consists of an exponential function circuit (Exp), an inverter circuit, an operational amplifier, and a divider. The exponential function circuit is composed of two symmetric transistors, successfully overcoming the temperature drift problem. Its output voltage V exp is as shown in the following formula.
[0088]
[0089] V in represents the input voltage.
[0090] In addition, the output voltage of the SoftMax activation function circuit is expressed as follows:
[0091]
[0092] N represents the number of input voltages of the SoftMax activation function circuit.
[0093] The forward propagation process of the forward propagation circuit is expressed as:
[0094] Forward = LN{W3[W2(LN(W1·Attention + b1)) + b2] + b3} (11)
[0095] Among them, Forward represents the output of this forward propagation, Attention represents the input of this forward propagation, W1 and b1 respectively represent the weight and bias of the first weight mapping circuit, W2 and b2 respectively represent the weight and bias of the second weight mapping circuit, W3 and b3 respectively represent the weight and bias of the third weight mapping circuit, and LN() represents the regularization function. and However, different from the memristive crossbar array described before, in this example, a row of memristor crossbar array is added to represent the bias weights b1, b2, and b3 corresponding to the three weight circuits.
[0096] In the circuit, Attention represents the multi - head attention extracted by the multi - head attention unit, and LN() represents the function for linear mapping. Then the output of the linear mapping circuit can also be expressed by formula (7) and formula (8).
[0097] The LN operation is a common data normalization operation, alleviating the "gradient vanishing" problem in the network training process. As shown in the following formula. α and β are the trainable parameter and the shift parameter respectively. x i represents the i th channel of the input feature signal, μ l and σ lare the mean and variance of the samples under different input channels, and ε is a constant whose initial value is almost close to 0.
[0098]
[0099] The LN layer retains the constant mean and variance values during the training phase for use during the inference phase. Therefore, to simplify the complexity of the hardware implementation of LN, the traditional formula is rewritten as a linear expression.
[0100]
[0101] Input voltage V x corresponds to the elements of the input sample, and V LN represents the regularized output voltage, Vb is the control signal for performing the LN operation, and M1 represents the resistance value of the memristor.
[0102]
[0103] Figure 8 Illustrates the hardware implementation scheme of MEC. It can be seen that the emotion classification module also includes a third memory unit module circuit connected between the max pooling circuit and the third layer regularization circuit, a fourth memory unit module circuit connected between the first pointwise convolutional circuit and the depth convolutional circuit, and a fifth memory unit module circuit connected between the first pointwise convolutional circuit and the fourth layer regularization circuit, which is used to store the output signals of the upper-level circuit and provide them to the lower-level circuit. The third memory unit module circuit, the fourth memory unit module circuit, and the fifth memory unit module circuit are all constructed based on the memory unit module circuit.
[0104] The emotion classifier is similar to a decoder. It receives the global information processed by the attention layer and then performs a depth-separated convolutional operation to obtain more local information. In addition, the final emotion detection is completed by a fully connected layer and a SoftMax activation function. In addition, MEC contains a max pooling circuit that performs a max embedding operation on the voltage from MMLAN. In other words, the most valuable information from MMLAN is filtered out before the convolutional operation. The max pooling circuit consists of paired NMOS devices, and its size is adjusted according to the number of input features. Its output is described by formula (15), where M represents the size of the max pooling.
[0105] V max =Max(V1,V2...V M ) (15)
[0106] The characteristic signals of each of the K scenario dialogues are fed into a max-pooling circuit, which includes all the corpora in the dataset. After passing through max-embedding, each dialogue outputs a signal of size U×d, and In addition, the embedded features are fed row by row into the third layer normalization circuit (LN circuit), followed by a pointwise convolution, a depth convolution, and another pointwise convolution. Since the information processed by the pointwise convolution and the depth convolution comes from different dimensions, a memory unit module circuit is inserted between them to temporarily store the intermediate voltage and perform dimensional transformation of the features. In addition, the Swish circuit is applied before the next pointwise convolution, aiming to normalize the output of the depth convolution while solving the problem of vanishing gradients caused by the saturation of the sigmoid function. Equation (16) shows the expression of the Swish circuit, with the hyperparameter β = 1.
[0107]
[0108] V bias represents the reference voltage, v sh represents its input voltage, V sh represents its output voltage, and exp() represents the exponential function.
[0109] After two pointwise convolutions and one depth convolution, the circuit output can be expressed as:
[0110]
[0111]
[0112] W pw1 、W pw2 and W dw represent the weights of the two pointwise convolution kernels and the depth convolution kernel respectively, W P1 represents the weight of the third layer normalization circuit, Max(Vatten m,n ) represents the output result of the max-pooling circuit, and Swish() represents the Swish function. It should be noted that the size of the convolution kernel of the depth convolution is kd 2 , so a total of d - kd 2 +1 depth convolution operations are required, which are completed by parallel operations. In addition, the generated intermediate voltage is saved in the MUM. Then the output voltage of the regularization operation is fed into the fully connected layer circuit (with weights and biases W4 and b4 respectively) and the second SoftMax activation function circuit (also constructed based on the SoftMax activation function circuit), and the result is shown as:
[0113] result = Max{SoftMax[W4·LN(V pw2 ) + b4]} (19)
[0114] W4 and b4 represent the weights and biases of the fully connected layer circuit respectively, and Max() represents taking the maximum probability as the output.
[0115] Correspondingly, the embodiment of the present invention further provides a text sentiment detection method based on memristors, including: constructing the above text sentiment detection system, training and testing the constructed text sentiment detection system, inputting the text to be detected into the trained text sentiment detection system, and the text sentiment detection system outputs the sentiment detection result.
[0116] To analyze the performance and effect of MTEDS, a series of experiments were conducted in this embodiment by comparing the proposed system with other related dialogue text sentiment detection works.
[0117] EmotionLines is a dialogue dataset consisting of two subsets, Friends and EmotionPush. The Friends dataset is a speech-based dataset that compiles multi-party conversations from a famous comedy series. The text data in the EmotionPush dataset is collected through social network messengers (such as Facebook). 4000 conversations in each dataset include 1000 original English conversations and 3000 enhanced conversations back-translated into French, German, and Italian respectively. Each conversation contains a large amount of corpus, and each corpus is labeled with seven emotions. In this embodiment, four emotions including happy, sad, angry, and natural are selected as the label candidates of the present invention and are regarded as benchmarks in the performance evaluation shown in Table 8. In addition, approximately 3000 corpus from 240 conversations are provided to the evaluation process; the distribution of the test data is the same as that of the training set.
[0118] This embodiment preprocesses the EmotionLines dataset. First, all utterances are tokenized and lowercased. By introducing special tokens between utterances, all tokens in the same conversation are concatenated. It is worth mentioning that all tokens are embedded through the WordPiece method. After positional encoding and token embedding, the text information is input into the corresponding multi-layer attention network and trained using the PyTorch platform.
[0119] Figure 9Represent the inputs and outputs of each layer of the network during the training phase, which are related to the size of the memristive crossbar array and the circuit in the MTEDS inference phase. Additionally, during the entire training phase, dropout is utilized to avoid overfitting of the data, and its value affects the final network accuracy. Furthermore, to improve the generalization ability of the network, in this embodiment, both the Friends and EmotionPush datasets are used to train the network. However, due to the unbalanced data distributions of the two datasets, it is difficult to train the two datasets simultaneously. Therefore, weighted cross-entropy (WCE) is used as the training loss to weight the minority human category data.
[0120] Moreover, by incorporating weighted balanced warming into the WCE loss, the model can learn small-scale sentiment data faster, enabling the training process to run more effectively. Additionally, the Adam optimizer is adopted and works with a predetermined learning rate. Table 2 shows some key parameters used during the training process, which are determined after a large number of experiments.
[0121] Table 2 Network parameters in the training phase
[0122]
[0123] To evaluate the prediction performance, mainly the F1 score is evaluated. If each data is only labeled as one category, then it is equivalent to the weighted accuracy (WA) and is calculated by the following formula. Additionally, to compare the performance of the proposed model with other network topologies, the last 20% of the English conversation data in the Friends dataset is used as the validation set to evaluate the model.
[0124] In this embodiment, all circuits are implemented based on memristors and CMOS technology. Additionally, the parameters of the remaining electronic components of the circuit are shown in Table 3. According to the trained network structure and the corresponding weights, the specific circuit structure and area size of the MTEDS are determined. Moreover, the weights of each network layer are uploaded to the memristive crossbar array to complete the hardware deployment of MMLAN and MEC for the inference task of emotion detection. Additionally, according to the given parameters, circuit simulation is performed using SPICE to simulate the hardware implementation of emotion detection. It is worth mentioning that the adopted memristor model has a high switching ratio R off / R on , which meets the application of the memristive neural computing system in practical scenarios. Additionally, the clock of the circuit is set to 10 μs, which is sufficient for almost all circuits in this embodiment. However, due to the non-linearity of the activation function circuit, the simulation time is set to 10 times the clock to obtain stable and accurate outputs.
[0125] Table 3 Circuit parameters and their configurations
[0126]
[0127] After network training, MTEDS has achieved the best text sentiment detection results so far. First, the methods used in the previous work were respectively experimented on the validation dataset, and their F1 scores in four emotion categories (natural, happy, sad, and angry) were calculated. The specific results are shown in Table 4. BOW and TFIDF are traditional baselines, and their F1 scores reached 0.81, but the scores for angry and sad were relatively low. TextCNN and the causal model TextCNN (C-TextCNN) used the weighted loss method and produced a rather good balance. Based on the original TextCNN, using the causal corpus model, both the previous embeddings and the target embeddings were mapped into the model.
[0128] Table 4 Results of different neural networks on the Friends validation dataset
[0129]
[0130] In addition, pre-trained base and large-scale tested BERT models were used to predict the results. As shown in Table 5, better performance than the early works was achieved by fusing the depth-separable convolutional module and the multi-layer attention network. The proposed model MLA-DSCN has a good balance, includes data with large and small sample sizes, and obtained superior results. Notably, the F1 scores for the small sample size data (sad and angry) were increased by 24% and 14% respectively. In addition, since the attention-based mechanism method achieved better results than other networks, other similar works were also listed for comparison.
[0131] Table 5 Performance on the EmotionPush and Friends test datasets
[0132]
[0133] As shown in Table 6, the accuracy of the proposed model on the EmotionLines dataset is 72.09%, while the F1 score is found to be 0.79. Based on the inference results of the test dataset, it is obvious that the proposed MTEDS achieves SOTA performance in terms of the final F1 score. In addition, in this embodiment, no additional dataset is required to pre-train the model to outperform previous models. In other words, the proposed structure is proven to have excellent feature extraction ability and greater generalization ability, making MTEDS more universal in practical application scenarios. Moreover, the designed hardware deployment scheme provides a feasible solution for the edge deployment of MLA-DSCN. The parameters of the whole network model are only 110.6M, while the parameters of the Friends-BERT-large model are 340M. Therefore, the proposed system further improves the feasibility and scalability of the network in actual embedded applications.
[0134] Table 6 Comparison of existing methods based on the EmotionLines dataset
[0135]
[0136] After completing the deployment process of MTEDS, when performing inference on the hardware device, the stability and correctness of the proposed circuit need to be further considered. Since the designed circuit includes MOS transistors, their non-linear characteristics become the main factor affecting the circuit stability. Therefore, within the specified clock cycle, transient analysis and DC analysis are performed on different circuits respectively to test the stability of the circuit under different input modes.
[0137] Different from previous activation function designs, the proposed activation function circuit fully simulates the entire mathematical model of the activation function and performs well within the rated input range [-2V, 2V]. This is attributed to the fine-tuning of specific component parameters in the circuit and the reasonable scheduling of the clock cycle. Due to the presence of transistors in the activation function circuit, the circuit has high non-linear characteristics during transient analysis. Therefore, to ensure the stability of the circuit, a clock signal with a length 10 times that of each transient signal is applied to the output of each transient signal.
[0138] For transient analysis, all voltages are limited within the range of [-1V, 1V]. It is worth mentioning that during the actual operation of the circuit, the intermediate voltage values in the circuit are usually at the mV level or even lower, and this operation is usually achieved by scaling the input voltage with an operational amplifier. The simulated output of the circuit is represented as a smooth curve and compared with the ideal output voltage curve for DC analysis.
[0139] The relevant analysis of the activation function is shown in Figure 10Among them: (a) input pulse, (b) comparison between DC analysis and inference output and SoftMax ideal output, (c) transient analysis result of SoftMax circuit, (d) comparison between DC analysis and Swish circuit, (e) transient analysis result of Swish circuit. In Figure 10 Among them, the sum of ideal DC voltages V total_ideal and the inferred circuit output V total are almost exactly the same, both equal to 1V, meeting the actual usage requirements. In addition, although the input of the pulse signal causes slight fluctuations in the output signal of the SoftMax circuit, the circuit still remains almost stable after operation, and the sum of all its outputs is equal to 1V. Therefore, it can be concluded that the presence of transistors and capacitors in the exponential function circuit is the main cause of this unstable signal, and embedding a 1nF capacitor in the exponential function circuit can effectively alleviate this instability. Like the SoftMax circuit, the same test also applies to the Swish circuit, as shown in Figure 10 (d). It is worth mentioning that in the specific experiment, the bias voltage of the multiplication and division operations is fine-tuned to the mV level to make the Swish circuit operate under the most ideal conditions.
[0140] The capacitance of the memory cell is the most important factor affecting the intermediate voltage. To prevent the mutual access between the memory cell and the computing circuit and avoid the frequent conversion between analog and digital signals, the MUM circuit is adopted in operations involving dimensional transformation and inter-layer transmission. To ensure the stability of each circuit branch, the memory cell needs to maintain the accuracy of the intermediate voltage signal within different clock cycles, that is, no matter how many clock cycles pass, it is necessary to ensure that the voltage at the input end is as equal as possible to the voltage at the output end. Therefore, an operational amplifier is used to maintain the voltage drop on the capacitor. On this basis, this embodiment explores the performance of MUM under different capacitances. Figure 11 shows the relationship between the input and output signals under five capacitance values (1nF, 10pF, 1pF, 0.1pF, and 1fF). In addition, the solid line represents the input and the dashed line represents the output. It is found that when the capacitance is equal to 1nF, the input and output can be stable and consistent. Therefore, the final capacitance size is determined to be 1nF, so that the temporary storage of the intermediate voltage can be accurately achieved.
[0141] After the co - design of the model and the circuit is completed, errors occur when loading weights on each neuron, which is the first place where non - linear errors appear. When using the AIST - based model for linear weight mapping, there is still a 2% error. In addition, due to the non - linear characteristics of memristors themselves, the produced memristors are still not perfectly linearly graded in practical applications. This remains a problem in actual implementation. In this embodiment, the weights of the memristor model are assumed to be linearly graded. In addition, the training weights based on the PyTorch platform are stored as 32 - bit floating - point numbers. However, for actual embedded device deployment, it is often necessary to quantize the weights to 8 - bit integers to meet the requirements of hardware deployment such as ARM and FPGA. Therefore, in this embodiment, the trained MLA - DSCN is also quantized to 8 - bit to match the interface with the designed hardware. It is worth mentioning that the benefits of quantization are not only to implement the interface between hardware and software, but also to greatly reduce the memory overhead, from 110.67MB to 25.06MB, while the F1 score only drops by 2.19%. This is not only beneficial to the deployment of emerging memory architectures, but also provides a feasible solution for traditional embedded devices. Table 7 shows the performance of the final MTEDS.
[0142] Table 7 Hardware deployments with accuracy loss at different bit - precision levels
[0143]
[0144] Each neuron consists of two memristors. The voltage input to the circuit is at the mV level, so the magnitude of the voltage is limited between [-3.3mV, 3.3mV] during calculation. The time period is set to 10us, and 50 pulsed voltages with an amplitude of 3.3mV are input into the memristor to simulate the stability of the memristor model within the threshold voltage range and its corresponding maximum power consumption. As Figure 12 shown, after the weight mapping is completed, the memristor is in a stable state. The energy consumption of the memristive neuron is:
[0145]
[0146] According to the simulation test results, the switching ratio of the memristor is not the factor determining the power consumption of a single neuron, while the magnitude of R on directly affects the maximum power consumption of the neuron. In addition, according to Figure 8For the specific structure of the MTEDS, the area overhead of the entire system is calculated in Table 8. The energy consumption depends on the power overhead of each circuit module and the calculation time. It should be noted that the calculation time is related to the quantity and size of the input information, as well as the size of the memristive crossbar array. In this embodiment, the energy consumption of the MTEDS is evaluated based on a single conversation because the corpus of each conversation is embedded with 768-dimensional feature information through max pooling. Therefore, an ideal design is a memristive crossbar array with a size of 768*768. In this case, the information processing of a single conversation can be completed with one input and one clock cycle. Table 8 shows the power consumption and area of the MTEDS based on the entire conversation, where the abbreviations of each circuit module are represented by their representative parameters. In integrated circuits, MOSEFETs, transistors, resistors, and other basic components are all at the nanoscale. Therefore, when calculating the area overhead of the overall circuit module, the size of each component is set to 100 nm 2 , where one operational amplifier consists of 8 MOSEFETs. In addition, since the electrical components in the proposed circuit can be reused by scheduling clock signals, the memristive crossbar array and its peripheral circuits are the sources of most of the energy and area overhead in the MTEDS. However, the size of the peripheral circuits is limited within the actually produced memristive crossbar array.
[0147] Table 8 Main power and area consumption of the MTEDS
[0148]
[0149] Therefore, the available device size must be considered in actual hardware deployment. This embodiment considers two different strategies. The first is to use only one memristive crossbar array and then store the results of multiple operations through MUM. The other is to use multiple memristive crossbar arrays for parallel processing. Both of these schemes need to consider the power consumption of the memristive crossbar array and the power consumption of the additional peripheral auxiliary circuits due to the change in the size of the memristive crossbar array. However, the ideal state still cannot be achieved in actual applications. The main size of the memristive crossbar array used is 1K. Therefore, multiple memristive crossbar arrays are used for parallel processing in the simulation experiment. In addition, in order to minimize the power consumption and area overhead in the hardware design as much as possible, the peripheral circuits are reused in this paper, and more clock cycles are required to complete one inference. With the further integration of the memristive crossbar array, larger-sized memristive crossbar arrays are produced, such as 8K memristor arrays (128 rows and 64 columns). In addition, as Figure 13As shown, the performance of MTEDS under different cross - array sizes is reported. Thus, the energy consumption for completing one dialogue inference is 1.9 mJ, and the total time cost is 17.15 ms. Finally, ten emotion detection inferences were performed on the GPU platform (Geforce RTX3070) for comparison with the proposed MTEDS. The average time for these 10 inferences is 6.42 ms. Therefore, the average time for completing one inference on the GPU platform is determined to be 6.42 ms / 240 dialogues = 26.75 μs. The overall energy consumption is 26.75 μs × 220 W = 5.885 mJ. The energy consumption of MTEDS is only about 32.3% of the GPU overhead, but the running time is 640 times that of the GPU.
[0150] The experimental results confirm that the MTEDS proposed in the embodiments of the present invention has lower energy and area overheads than the standard neuromorphic computing architecture. In addition, the memory function of the memristor enables it to simulate neurons in the memristive in - memory computing architecture. This is due to the threshold model of the memristor, which allows the memristor to remain stable when the operating voltage is lower than the threshold voltage.
[0151] This embodiment proposes a hardware implementation scheme for text emotion detection, namely MTEDS. First, by integrating the depth - separable convolution module into the multi - layer attention network, the improved network model MLA - DSCN can not only focus on the global information of the input text information but also improve its information extraction ability at the local level. Without sacrificing accuracy, the number of network parameters is reduced by 3 times. Then, the proposed network achieves the current best F1 - score for text - based emotion detection tasks on the EmotionLines dataset. Further, the corresponding hardware parts such as size and weight are also determined. The stability and power consumption of the proposed circuit are also analyzed through SPICE simulation, verifying the feasibility of practical applications in the real world. In particular, under the same benchmark conditions, the proposed MTEDS can effectively reduce energy consumption and area overhead compared with the GPU, while the full - analog configuration of the circuit increases the computing time of MTEDS. The MTEDS proposed in this embodiment not only has better text emotion detection ability but also has a lighter structure and lower energy consumption, which provides a good solution for the application of deploying text emotion detection systems and methods on embedded devices.
[0152] The above - mentioned embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited by the above - mentioned embodiments. Any other changes, modifications, substitutions, combinations, and simplifications made without departing from the spirit and principle of the present invention shall be equivalent replacement methods and are all included in the protection scope of the present invention.
Claims
1. A memristor-based text sentiment detection system, characterized in that: It includes a preprocessing module, an attention calculation module, and an emotion classification module; The preprocessing module is used to input the discourse text data and perform positional encoding and token embedding on the discourse text data to obtain the embedded text data; The attention calculation module includes multiple layers of attention circuits and a forward propagation circuit; The attention circuit of the first layer is used to extract the multi-head attention of the embedded text data, and the forward propagation circuit is used to propagate the output of the attention circuit of the current layer to the attention circuit of the next layer as input until the last layer of the attention circuit, and the final attention output is sent to the emotion classification module; The emotion classification module includes a max pooling circuit, a third-layer regularization circuit, a first pointwise convolutional circuit, a depth convolutional circuit, a Swish activation function circuit, a second pointwise convolutional circuit, a fourth-layer regularization circuit, a fully connected layer circuit, and a second SoftMax activation function circuit connected in sequence, which are used to perform corresponding max pooling, regularization, pointwise convolution, depth convolution, Swish activation function calculation, pointwise convolution, regularization, fully connected, and SoftMax activation function calculation operations on their respective inputs, and finally output the emotion detection result; The attention extraction operation, weight mapping operation in the attention calculation module, the first pointwise convolutional circuit, the depth convolutional circuit, the second pointwise convolutional circuit, and the fully connected layer circuit are all constructed based on the memristive crossbar array. Each two columns of the memristive crossbar array correspond to positive and negative weights in the neural network, and the output of each two columns corresponds to an output voltage; The attention circuit includes a multi-head attention circuit constructed based on the memristive crossbar array, a first memory unit module circuit for storing voltage signals, a first multiply-accumulate circuit and a second multiply-accumulate circuit connected to the first memory unit module circuit, and also includes a first SoftMax activation function circuit; the multi-head attention circuit is used to extract the multi-head attention of the embedded text data and output the Q, K, and V signals of multiple heads. The first multiply-accumulate circuit is used to perform multiply-accumulation on the Q and K signals of multiple heads to obtain the total Q and K signals. The second multiply-accumulate circuit is used to perform multiply-accumulation on the V signals of multiple heads to obtain the total V signal. The first SoftMax activation function circuit is used to calculate the total Q and K signals using the softmax function, and the calculation result is multiplied by the total V signal with weights and then output to the forward propagation circuit; The forward propagation circuit includes a first weight mapping circuit, a first-layer regularization circuit, a second weight mapping circuit, a third weight mapping circuit, and a second-layer regularization circuit connected in sequence, and also includes a second memory unit module circuit; wherein the second-layer regularization circuit is used to connect the attention circuit of the next layer, and the second memory unit module circuit is used to store the attention calculated each time during the forward propagation process; The first weight mapping circuit, the second weight mapping circuit, and the third weight mapping circuit are all constructed based on the memristive crossbar array.
2. The text sentiment detection system based on a memristor according to claim 1, characterized in that, The output of the attention circuit is expressed as: Among them, Concat{} represents concatenation, softmax() represents the softmax function, d h = d / h represents the dimension of the head, h represents the number of heads, d is the dimension of the input Q, K, V, and Q h , K h , V h respectively represent the Q, K, V values of the single-head output, and W Z represents the mapping matrix.
3. The text sentiment detection system based on a memristor according to claim 2, wherein The forward propagation process of the forward propagation circuit is expressed as: Forward = LN{W3[W2(LN(W1·Attention + b1)) + b2] + b3}, where Forward represents the output of the forward propagation, Attention represents the input of the forward propagation, W1 and b1 respectively represent the weight and bias of the first weight mapping circuit, W2 and b2 respectively represent the weight and bias of the second weight mapping circuit, W3 and b3 respectively represent the weight and bias of the third weight mapping circuit, and LN() represents the regularization function.
4. The text sentiment detection system based on a memristor according to claim 3, characterized in that The output of a single channel of the deep convolutional circuit is expressed as: where k is the kernel width of the depth convolution, represents the kernel weights, represents the input of the depth convolution circuit, and U represents the number of channels; The output of the deep convolutional circuit is expressed as: The output of the second pointwise convolutional circuit is expressed as: W pw1 , W pw2 and W dw respectively represent the weights of the convolution kernels of the first point-to-convolution circuit, the second point-to-convolution circuit, and the depth convolution circuit. W P1 represents the weight of the third-layer regularization circuit, Max(Vatten m,n ) represents the output result of the max pooling circuit, and Swish() represents the Swish function.
5. The text sentiment detection system based on a memristor according to claim 4, characterized in that: The sentiment classification module further includes a third memory unit module circuit connected between the max pooling circuit and the third layer regularization circuit, a fourth memory unit module circuit connected between the first pointwise convolutional circuit and the deep convolutional circuit, and a fifth memory unit module circuit connected between the first pointwise convolutional circuit and the fourth layer regularization circuit, which are used to store the output signals of the upper-level circuits and provide them to the lower-level circuits.
6. The text sentiment detection system based on a memristor according to claim 5, wherein: The first memory unit module circuit, the second memory unit module circuit, the third memory unit module circuit, the fourth memory unit module circuit, and the fifth memory unit module circuit are all constructed based on the memory unit module circuit. The memory unit module circuit includes a plurality of memory unit circuits connected in an array, and the number of memory unit circuits is the same as the number of voltage signals to be stored; For the memory unit circuits in each row of the memory unit module circuit, all the voltage signal input terminals are connected together, all the voltage signal output terminals are connected together, all the control terminals of the first MOS switches are connected together, and all the control terminals of the second MOS switches are connected together; the memory unit module circuit sequentially inputs and outputs voltages according to the clock beats.
7. The text sentiment detection system based on a memristor according to claim 6, characterized in that: The memory unit circuit includes a first MOS switch, a second MOS switch, a first operational amplifier, and a capacitor. The control terminal of the first MOS switch is connected to the control terminal of the second MOS switch with a clock control signal. The inverting input terminal of the first operational amplifier is connected to the output terminal of the first MOS switch. The non-inverting input terminal of the first operational amplifier is connected to the capacitor and then grounded. The output terminal of the first operational amplifier is connected to the input terminal of the second MOS switch. The input terminal of the first MOS switch is used as the voltage signal input terminal of the memory unit circuit, and the output terminal of the second MOS switch is used as the voltage signal output terminal of the memory unit circuit.
8. The text sentiment detection system based on a memristor according to claim 7, wherein: The memristive crossbar array adopts a memristor model with a 2M structure.
9. The text sentiment detection system based on a memristor according to claim 1, wherein: Both the Friends and EmotionPush datasets are used to train the network to train the text sentiment detection system, and weighted cross-entropy is used as the training loss.
10. A memristor-based text sentiment detection method, characterized in that, Including: Construct the text sentiment detection system according to any one of claims 1 to 9, train and test the constructed text sentiment detection system, input the text to be detected into the trained text sentiment detection system, and the text sentiment detection system outputs the sentiment detection result.
Citation Information
Patent Citations
Multi-head attention memory network for short text sentiment classification
CN112784532A
Semantic sentiment analysis method fusing in-depth features and time sequence models
US11194972B1