Natural language processing method and device based on hybrid neural network model, equipment and medium
By constructing a hybrid neural network model that combines spiking neural networks and BERT neural networks, feature extraction and information transmission are optimized, solving the problems of high computational overhead and energy consumption in traditional models, and achieving efficient and low-energy natural language processing.
Patent Information
- Application Number
- CN202510305249.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-11-04
- Estimated Expiration
- 2045-03-14
AI Technical Summary
Traditional artificial neural network models have high computational overhead and energy consumption in natural language processing, becoming a bottleneck for computing resources.
A hybrid neural network model is constructed, combining a spiking neural network model and a BERT neural network model. Feature extraction and information transmission are optimized through periodic integral and firing neurons, periodic spiking attention modules, and hybrid unit-driven attention modulation modules.
It improves the efficiency and accuracy of natural language processing tasks, reduces energy consumption, is suitable for low-resource environments, maintains accuracy comparable to traditional models, and significantly reduces energy consumption.
Smart Images

Figure CN120409454B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of computer, and particularly relates to a natural language processing method and device based on a hybrid neural network model, equipment and a medium. BACKGROUND
[0002] At present, with the rapid development of artificial intelligence technology, artificial intelligence models have been successfully applied in the field of natural language processing (NLP), and have shown good application prospects.
[0003] In the related art, traditional artificial neural network (ANN) models such as BERT and GPT are usually used in the field of NLP. However, in actual application, it is found that the traditional ANN model has high computational overhead and energy consumption, and the consumption of computing resources becomes a major bottleneck.
[0004] Therefore, the technical problems in the related art need to be improved. SUMMARY
[0005] The embodiments of the present application provide a natural language processing method and device based on a hybrid neural network model, which can effectively combine the high-precision advantages of artificial neural network models and the low-power computing advantages of spiking neural network models, improve the efficiency and accuracy of natural language processing tasks, and provide strong support for the development of large-scale artificial intelligence model industry.
[0006] In one aspect, the embodiments of the present application provide a natural language processing method based on a hybrid neural network model, which comprises the following steps:
[0007] acquiring a natural language text to be processed;
[0008] inputting the natural language text to be processed into a hybrid neural network model to obtain a natural language processing result output by the hybrid neural network model;
[0009] The hybrid neural network model is constructed based on a spiking neural network model and a BERT neural network model.
[0010] Optionally, the hybrid neural network model comprises an embedding layer, a plurality of encoding feature extraction modules and a prediction head connected in sequence.
[0011] The encoding feature extraction module comprises a periodic integral and firing neuron, a periodic spiking attention module and a hybrid unit driven attention modulation module.
[0012] Optionally, the periodic integral and firing neuron is designed with a periodic modulation input current mechanism and a periodic reset voltage mechanism.
[0013] The periodic modulation input current mechanism includes periodically modulating the input current of the periodic integral and firing neuron, and the periodic reset voltage mechanism includes periodically resetting the output voltage of the periodic integral and firing neuron.
[0014] Optionally, the periodic pulse attention module is designed with a pulsing processing mechanism, an adaptive convolution mechanism, and a periodic feature fusion mechanism.
[0015] The pulsing processing mechanism includes calling the periodic integral and firing neuron to convert a feature sequence into a pulse sequence, the adaptive convolution mechanism includes adjusting the convolution kernel size according to the number of heads and the sequence length in the multi-head attention mechanism to complete one-dimensional convolution, and the periodic feature fusion mechanism includes periodically fusing the original input features and the features output by the adaptive convolution.
[0016] Optionally, the hybrid unit driven attention modulation module is designed with a dual-path hybrid encoding mechanism and a joint loss function mechanism.
[0017] The dual-path hybrid encoding mechanism includes a BERT neural network model path and a pulse neural network model path, and the key matrix in the multi-head attention mechanism is extracted and weighted fused through the two paths, and the joint loss function mechanism includes introducing a joint loss function, and the joint loss function includes a task loss and a regularization term.
[0018] Optionally, the encoding feature extraction module has a residual connection.
[0019] Optionally, the hybrid neural network model is trained based on the following steps:
[0020] A plurality of historical natural language texts are obtained, and a plurality of natural language processing results corresponding to the historical natural language texts are obtained.
[0021] Each of the historical natural language texts is used as a sample, and the natural language processing result corresponding to each of the historical natural language texts is used as a sample label corresponding to the sample, and a training data set is constructed.
[0022] The hybrid neural network model is pre-trained using the training data set.
[0023] On the other hand, the embodiment of the present application provides a natural language processing device based on a hybrid neural network model, and the device comprises:
[0024] The text acquisition module is configured to acquire a natural language text to be processed.
[0025] The text processing module is configured to input the natural language text to be processed into a hybrid neural network model to acquire a natural language processing result output by the hybrid neural network model.
[0026] The hybrid neural network model is constructed based on a spiking neural network model and a BERT neural network model.
[0027] In another aspect, the embodiments of the present application provide an electronic device, which comprises a memory and a processor. The memory stores a computer program, and the processor implements the natural language processing method based on the hybrid neural network model when executing the computer program.
[0028] In another aspect, the embodiments of the present application provide a computer readable storage medium, which stores a computer program. The computer program is executed by a processor to implement the natural language processing method based on the hybrid neural network model.
[0029] The embodiments of the present application construct a hybrid neural network model based on a spiking neural network model and a BERT neural network model, and then complete a natural language processing task through the hybrid neural network model. The embodiments effectively combine the high-precision advantage of an artificial neural network model and the low-power computing advantage of a spiking neural network model, improve the efficiency and accuracy of the natural language processing task, and provide strong support for the development of large-scale artificial intelligence model industry. BRIEF DESCRIPTION OF DRAWINGS
[0030] Figure 1 is an implementation environment schematic diagram of a natural language processing method based on a hybrid neural network model provided by the embodiments of the present application;
[0031] Figure 2 is a flow schematic diagram of a natural language processing method based on a hybrid neural network model provided by the embodiments of the present application;
[0032] Figure 3 is an architecture schematic diagram of a hybrid neural network model provided by the embodiments of the present application;
[0033] Figure 4 is a flow schematic diagram of data processing of a periodic spiking attention module provided by the embodiments of the present application;
[0034] Figure 5 is a flow schematic diagram of data processing of a hybrid cell driven attention modulation module provided by the embodiments of the present application;
[0035] Figure 6is a structural schematic diagram of a natural language processing device based on a hybrid neural network model provided by an embodiment of the present application.
[0036] Figure 7 is a hardware structural schematic diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION
[0037] In order to make the purposes, technical solutions and advantages of the present application clearer, the present application is further described in detail below in combination with the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and do not limit the present application. When the following description refers to the drawings, the same numbers in different drawings represent the same or similar elements unless otherwise indicated. The implementations described in the following exemplary embodiments do not represent all the implementations consistent with the embodiments of the present application. They are only examples of devices and methods consistent with some aspects of the embodiments of the present application as detailed in the appended claims.
[0038] It can be understood that the terms "first", "second", and the like used in the present application can be used herein to describe various concepts, but unless specifically stated, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of the embodiments of the present application, the first information can also be referred to as the second information, and similarly, the second information can also be referred to as the first information. Depending on the context, the word "if" as used herein can be interpreted as "when" or "when" or "in response to determining".
[0039] The terms "at least one", "multiple", "each", "any" and the like used in the present application include one, two or more than two, multiple includes two or more than two, each refers to each of the corresponding multiple, and any refers to any one of the multiple.
[0040] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as understood by those skilled in the art to which the present application belongs. The terms used herein are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.
[0041] Currently, with the rapid development of artificial intelligence technology, artificial intelligence models have been successfully applied in the field of natural language processing, showing good application prospects.
[0042] In the related art, traditional artificial neural network models such as BERT, GPT, etc. are usually used in the field of NLP. However, it is found in actual application that the traditional ANN model has high computational overhead and energy consumption, and the consumption of computing resources becomes a major bottleneck.
[0043] Therefore, the embodiment of the present application provides a natural language processing method, device and equipment based on a hybrid neural network model and a medium. The hybrid neural network model is constructed based on a spiking neural network model and a BERT neural network model, and the natural language processing task is completed through the hybrid neural network model. The high-precision advantage of the artificial neural network model and the low-power computing advantage of the spiking neural network model are effectively combined to improve the efficiency and accuracy of the natural language processing task, and provide strong support for the development of large-scale artificial intelligence model industry.
[0044] The specific implementation of the embodiment of the present application will be described in detail below with reference to the accompanying drawings. First, a natural language processing method based on a hybrid neural network model provided in the embodiment of the present application is described with reference to the accompanying drawings.
[0045] Please refer to Figure 1 , Figure 1 is an implementation environment schematic diagram of a natural language processing method based on a hybrid neural network model provided in the embodiment of the present application. In the implementation environment, the main hardware and software subjects involved include a terminal processor 110 and a server 120.
[0046] Specifically, the terminal processor 110 can be installed with a related control program of the natural language processing method based on a hybrid neural network model, and the server 120 is a background server of the control program. The terminal processor 110 and the background server 120 are communicatively connected. The natural language processing method based on a hybrid neural network model provided in the embodiment of the present application can be executed on the terminal processor 110 side.
[0047] The server 120 can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and basic cloud computing services such as big data and artificial intelligence platforms.
[0048] In addition, the server 120 can also be a node server in a blockchain network.
[0049] The terminal processor 110 and the server 120 can establish a communication connection through a wireless network. The wireless network uses standard communication technology and / or protocols, and the network can be set up as the Internet or any other network, such as but not limited to a local area network (LAN), a metropolitan area network (MAN), a wide area network (WAN), a mobile or wireless network, a private network, or any combination of virtual private networks. In addition, the same communication connection method or different communication connection methods can be used between the above-mentioned software and hardware subjects, and the application does not make specific limitations.
[0050] Of course, it can be understood that Figure 1 The implementation environment in the above is only some optional application scenarios of the natural language processing method based on the hybrid neural network model provided in the embodiments of the application, and the actual application is not fixed to the software and hardware environment shown in the above. Figure 1 The application does not make specific limitations.
[0051] As Figure 2 shown, Figure 2 is a flowchart of a natural language processing method based on a hybrid neural network model provided in an embodiment of the application, and specifically includes but is not limited to steps 100 to 200.
[0052] Step 100, obtaining a natural language text to be processed.
[0053] In the embodiments of the application, the natural language text to be processed can be obtained by responding to the natural language text input by the user, which can come from various data sources such as documents, chat conversations, social media, etc.
[0054] Step 200, inputting the natural language text to be processed into a hybrid neural network model to obtain a natural language processing result output by the hybrid neural network model.
[0055] The hybrid neural network model is constructed based on a spiking neural network model and a BERT neural network model.
[0056] In the embodiments of the application, the obtained natural language text data to be processed is input into the hybrid neural network model, and the hybrid neural network model performs feature extraction and understanding, and finally outputs the processed natural language processing result, such as text classification, sentiment analysis, question and answer, etc.
[0057] In practical applications, the hybrid neural network model is constructed based on the pulse neural network model and the BERT neural network model, which can effectively utilize the powerful language understanding ability of the BERT model to capture semantic information in the text, and utilize the event-driven sparsity and binary activation characteristics of the pulse neural network (SNN) model inspired by biology to reduce energy consumption while ensuring computing efficiency.
[0058] Therefore, the hybrid neural network model provided in the present application combines the pulse neural network model and the BERT neural network model, utilizes the high-precision advantages of the artificial neural network model and the low-power computing advantages of the pulse neural network model, improves the performance and efficiency of the hybrid neural network model in the natural language processing field, and provides strong support for the development of large-scale artificial intelligence model industry, for example, has significant advantages in deployment in low-resource environments (such as edge computing devices).
[0059] Specifically, as an optional implementation, the hybrid neural network model includes an embedding layer, a plurality of encoding feature extraction modules, and a prediction head connected in sequence.
[0060] The encoding feature extraction module includes a periodic integral and firing neuron, a periodic pulse attention module, and a hybrid unit driven attention modulation module.
[0061] In the embodiments of the present application, please refer to Figure 3 , Figure 3 is a schematic diagram of a hybrid neural network model provided in an embodiment of the present application, which mainly includes an embedding layer, a plurality of encoding feature extraction modules, and a prediction head.
[0062] It can be understood that the encoding feature extraction module, as a basic structure for encoding and extracting feature information in the hybrid neural network model, can stack a plurality of encoding feature extraction modules between the embedding layer and the prediction head, thereby forming the hybrid neural network model provided in the present application. In addition, the structure of the Nth encoding feature extraction module shown in Figure 3 is not intended to limit the hybrid neural network model provided in the present application to only one encoding feature extraction module.
[0063] As Figure 3As shown, the encoding feature extraction module includes a periodic integrate-and-fire neuron (PIF), a periodic spiking attention module (PSAM), and a hybrid unit-driven attention modulation module (HUAM).
[0064] The natural language text (for example, a sentence) to be processed input into the hybrid neural network model is first converted into a matrix form by an embedding layer, which can be composed of a word vector dimension, a position coding dimension, and a sentence dimension.
[0065] Further, the converted feature matrix is input into the encoding feature extraction module as initial data, and three paths are used for feature extraction. The first path is to use the periodic integrate-and-fire neuron PIF provided by the present application for separate feature extraction, the second path is to use the hybrid unit-driven attention modulation module HUAM provided by the present application for feature extraction, and the third path is to use the linear transformation layer Linear for feature extraction. After the feature extraction is completed, data conversion is needed to facilitate subsequent data calculation.
[0066] Further, the output result after feature extraction by PIF and HUAM and data conversion is input into the QK T matrix multiplication operation, and then input into the scaling module SCALE and the periodic spiking attention module PSAM provided by the present application for feature extraction and data conversion, and then the result is input into the normalization module softmax for score calculation to generate attention weights. Finally, the features extracted by the linear transformation layer Linear (as shown in Figure 3 ) are combined, and input into the multilayer perceptron (MLP) for further feature extraction.
[0067] Finally, the above feature extraction process is repeatedly performed by the plurality of encoding feature extraction modules set, and finally a weight is output and input into the prediction head for task processing to obtain the natural language processing result output by the hybrid neural network model.
[0068] As an optional implementation, the encoding feature extraction module has a residual connection.
[0069] In the embodiments of the present application, a residual connection can be provided in each encoding feature extraction module. By adding the output attention weight and the input feature, the structure of "Attention+Input" in the Transformer is formed, as shown inFigure 3 As shown, there are two residual connections, such as Figure 3 The two points shown represent residual connections performed on the input features input to the encoding feature extraction module and the input features input to the MLP, respectively.
[0070] Specifically, as an optional implementation, the periodic integral and firing neurons are designed with a periodic modulation input current mechanism and a periodic voltage reset mechanism.
[0071] The periodic modulation input current mechanism includes periodically modulating the input current of the periodic integral and the firing neuron; the periodic voltage reset mechanism includes periodically resetting the output voltage of the periodic integral and the firing neuron.
[0072] In the embodiments of this application, the existing Integrate-and-Fire (IF) neuron model typically fires a pulse after the voltage accumulates to a threshold. This mechanism focuses too much on the instantaneous accumulation effect of voltage and ignores the periodic characteristics of neuron firing, making it difficult to effectively transmit continuous and fine-grained important information. Especially in tasks that require high-precision information transmission (such as Natural Language Processing NLP), this defect can lead to a large loss of information, thereby affecting the performance of the neural network model.
[0073] Taking the existing IF neuron as an example, its voltage accumulation process is shown in the following formula (1):
[0074] (1)
[0075] Where V(t) represents the membrane potential of the neuron at time step t; I(t) represents the input current at time step t; and ∆t represents the time step.
[0076] When V(t) ≥ Vth (where Vth is a preset voltage threshold), the neuron fires a pulse and resets the membrane potential to a specific value via a hard reset or soft reset. However, in some cases, if the input current I(t) is small, the voltage may fail to reach the threshold Vth within multiple time steps, thus preventing pulse firing and leading to information loss. This is particularly problematic in shallow layers of deep networks, severely impacting the feature extraction capabilities of subsequent networks and ultimately harming the overall network performance.
[0077] Therefore, the application provides a periodic integral and firing neuron PIF, and a periodic modulation input current mechanism and a periodic reset voltage mechanism are designed in the PIF. The periodic modulation input current mechanism includes periodically modulating the input current of the PIF, and the periodic reset voltage mechanism includes periodically resetting the output voltage of the PIF.
[0078] Specifically, the application periodically modulates the input current by introducing a differentiable transform function (such as a tangent function) to simulate the periodic behavior of biological neurons and suppress the influence of extreme current values on voltage accumulation.
[0079] In practical applications, the voltage accumulation process of the PIF provided by the application is shown in the following formula (2):
[0080] (2)
[0081] wherein, is a scaling factor for adjusting the non-linear amplitude of the tangent function, which can be set according to the specific scene use requirements.
[0082] That is, the application can effectively weaken the influence of extreme inputs by scaling the input current to the interval [-1, 1] using the tangent function.
[0083] Further, the application also designs a periodic reset voltage mechanism to replace the existing hard reset and soft reset mode, so that the neuron adjusts the firing state of the pulse according to the periodic dynamics at each time step, improves the pulse firing frequency, and improves the stability of information transmission of the neural network.
[0084] In practical applications, the PIF voltage reset of the application can be shown in the following formula (3) and formula (4):
[0085] (3)
[0086] (4)
[0087] wherein, one time step corresponds to one period, A is the period amplitude, which can be set according to the specific scene use requirements; T is the period length.
[0088] That is, when the membrane potential is negative, it means that it is in the trough of the period, and the PIF will not trigger a firing pulse, and the voltage of the next time step will be guided to the peak position of the period, and the voltage The periodic modulation term can be added to the voltage of the neuron to adjust the voltage towards the peak of the periodic wave. In this way, even if the input is small, the periodic mechanism can be used to gradually accumulate to a higher potential, preparing for the next possible pulse emission; when the membrane potential is positive and exceeds the voltage threshold , the PIF will be triggered to emit a pulse, and the voltage at the next time step will be guided to the trough position of the periodic wave. By adding a periodic modulation term to the current voltage , the voltage can be quickly reduced to guide the voltage to the trough of the periodic wave.
[0089] Therefore, the present application simulates the periodic excitation characteristics of neurons by periodically resetting the voltage mechanism, raising the voltage by a sine function at the trough and lowering the voltage by a cosine function after firing, so that the voltage changes along the preset periodic trajectory at each time step, making the pulse distribution more uniform in the time dimension, thereby effectively avoiding information loss caused by small current that cannot emit a pulse.
[0090] It can be understood that by introducing the PIF and designing the periodic modulation input current mechanism and the periodic reset voltage mechanism, the present application not only dynamically suppresses the interference of extreme input to the network during the voltage accumulation stage, but also avoids information loss through the periodic mechanism during the voltage reset stage.
[0091] Therefore, the hybrid network model of the present application can improve the information processing capability of SNN in NLP tasks by introducing periodic integration and firing neurons, so that it can achieve more efficient information transmission in fewer time steps and reduce the problem of information loss caused by low emission rate.
[0092] Specifically, as an optional implementation, the periodic pulse attention module is designed with a pulsing processing mechanism, an adaptive convolution mechanism, and a periodic feature fusion mechanism.
[0093] The pulsing processing mechanism includes converting the feature sequence into a pulse sequence by calling the periodic integration and firing neuron; the adaptive convolution mechanism includes adjusting the convolution kernel size according to the number of heads and the sequence length in the multi-head attention mechanism, thereby completing one-dimensional convolution; and the periodic feature fusion mechanism includes periodically fusing the original input features and the features output by the adaptive convolution.
[0094] In the embodiments of the present application, the existing BERT model is based on the Transformer model as the model backbone, and the Transformer realizes the global information modeling capability in the natural language processing task by virtue of its multi-head self-attention mechanism. However, its information transmission process mainly depends on the following three encoding dimensions: word vector dimension: capturing the semantic features of the word itself; position encoding dimension: by adding absolute or relative position encoding information, the model has sequence perception ability; inter-sentence dimension: realizing the comprehensive modeling of inter-sentence information.
[0095] However, in the actual training process, the neural network model not only needs the initial global position information, but also needs to dynamically capture the changing relationship between contexts. Because of the discrete characteristics of the SNN network, the information captured in the pulsing of the Transformer structure will be damaged to some extent, and cannot be flexibly adjusted according to the changes of the weights in the training process, resulting in insufficient performance in capturing dynamic context information. This deficiency will reduce the modeling ability of the network for long-distance dependencies and affect the overall performance.
[0096] That is, when combining the existing SNN network model with the ANN network model, due to the deficiencies of the ANN network in long-distance dependency modeling, information transmission stability and precision optimization after pulsing, there will be a problem of insufficient precision in the natural language processing task,
[0097] Therefore, the present application introduces a periodic pulse attention module PSAM, and designs a pulsing processing mechanism, an adaptive convolution mechanism and a periodic feature fusion mechanism in the PSAM.
[0098] Among them, the pulsing processing mechanism includes calling the periodic integral and firing neuron PIF to convert the feature sequence into a pulse sequence, and the PIF neuron is used to convert the serialized intermediate result into a pulse in the PSAM module, which can effectively model dynamic signals by combining the sparse excitation characteristics of the PIF neuron similar to biological nerves.
[0099] Further, to realize dynamic modeling, the PSAM module is designed with an adaptive convolution mechanism, including adjusting the convolution kernel size according to the number of heads in the multi-head attention mechanism and the sequence length, so as to complete one-dimensional convolution. In the convolution process, the size of the convolution kernel can be dynamically adjusted according to the number of heads in the Transformer multi-head attention mechanism and the length of the input feature sequence, which can be specifically represented as wherein, a is a proportion factor for controlling the scaling of the convolution kernel; h is the number of heads in the multi-head attention mechanism; N is the length of the input feature sequence; and ⌊•⌋ represents the rounding operation.
[0100] Further, the application also introduces a periodic feature fusion mechanism for the PSAM module, including periodic feature fusion of the original input features and the features output by the adaptive convolution, which further improves the network's ability to perceive dynamic relationships by repeatedly exciting and decaying pulses using the periodic features of the SNN. The feature fusion can be as shown in the following formula (5):
[0101] (5)
[0102] wherein, represents the feature result of the output after weighted fusion; is the feature result output after one-dimensional convolution in the adaptive convolution mechanism; is the original input feature; Fusion(.) is the feature fusion operation, which can be weighted summation or splicing.
[0103] Therefore, through the pulsing processing mechanism, the application can capture the dynamic weight changes in the training process in real time and enhance the perception of context information. Through the adaptive convolution mechanism and the periodic feature fusion mechanism, the receptive field of the convolution kernel can be flexibly adjusted, effectively solving the problem of insufficient long-distance dependency modeling capability in the pulsing Transformer structure, and the sparse pulse mechanism of PIF reduces the computational complexity and enhances the robustness of the model to noise.
[0104] In practical applications, please refer to Figure 4 , Figure 4 is a flowchart of data processing of a periodic pulse attention module provided by an embodiment of the application.
[0105] First, the multi-head attention mechanism of the Transformer generates a series of intermediate feature representations, which are input into the PSAM module and subjected to serialization processing, i.e., converting a two-dimensional feature matrix into a one-dimensional sequence, so as to complete the serialization processing for subsequent operations. The two-dimensional feature matrix can be connected as a long vector by row units, thereby converting into a one-dimensional sequence.
[0106] Further, through the pulsing processing mechanism introduced by the application, the one-dimensional sequence is subjected to pulsing processing, and a sequence of pulse signal representations is obtained according to the pulse outputs of all time steps.
[0107] Further, after completing the pulsing processing, the pulse signal sequence output can be subjected to adaptive convolution through the adaptive convolution mechanism introduced by the application, thereby completing the adaptive convolution.
[0108] Exemplarily, the sequence data can be first converted in data form, split and stacked into a square shape, then split according to a behavior unit and according to a column unit, wherein the horizontal data operation is row merging and the vertical data operation is column merging.
[0109] The split data is subjected to convolution and an activation function is called, finally the convolution structure of the horizontal data operation and the vertical data operation is merged, and a periodic feature fusion mechanism is called to periodically fuse the original input features and the features output by the adaptive convolution, and finally the feature result is output.
[0110] Specifically, as an optional implementation, the hybrid unit driven attention modulation module is designed with a dual-path hybrid encoding mechanism and a joint loss function mechanism.
[0111] The dual-path hybrid encoding mechanism includes a BERT neural network model path and a spiking neural network model path, and the key matrix in the multi-head attention mechanism is extracted and weighted fused through the two paths respectively; the joint loss function mechanism includes introducing a joint loss function, and the joint loss function includes a task loss and a regularization term.
[0112] In the Transformer architecture, the attention mechanism realizes feature interaction by generating query (Q), key (K) and value (V) matrices. However, if only the Q matrix is spiking (sparse coding based on SNN), and the encoding mode of the K matrix is inconsistent, a series of problems will be caused. First, the spiking Q matrix presents a significant binary feature (0 or 1), while the continuous K matrix retains high precision but lacks sparsity, which leads to unbalanced feature matching. For example, two similar feature values (such as a difference of 0.1) in the Q matrix may be mapped to the same spiking state, and then in the calculation of similarity, the subtle difference is weakened, affecting the accurate allocation of attention weights. Second, if the K matrix is completely spiking, although the calculation efficiency can be improved, the semantic continuity as the "key" will be lost, resulting in a decrease in feature expression ability, especially in language tasks that require high-precision modeling of long-distance dependencies, the model performance will be significantly impaired. Therefore, the inconsistent encoding modes of the Q matrix and the K matrix will affect the effectiveness of the attention mechanism, and cannot balance precision and efficiency.
[0113] Therefore, the application introduces a hybrid unit driven attention modulation module HUAM, and designs a dual-path hybrid encoding mechanism and a joint loss function mechanism.
[0114] The dual-path hybrid encoding mechanism includes a BERT neural network model path and a spiking neural network model path, and the key matrix in the multi-head attention mechanism is extracted and weighted fused through the two paths respectively.
[0115] Specifically, the HUAM decomposes the generation of the K matrix into two parallel feature extraction paths. One path is a traditional ANN network (e.g., a BERT neural network model) path that captures fine-grained semantic information to ensure high-precision modeling of key attributes. The other path is an SNN network (e.g., a PIF provided by the present application) path that enhances feature discrimination and suppresses noise interference to improve computational efficiency through sparse pulse coding.
[0116] Further, the output results of the two paths are fused by weighting. The process of weighted fusion can be shown in the following formula (6):
[0117] (6)
[0118] wherein, represents the weighted fusion output result of the HUAM module; represents the output result of the BERT neural network model path; represents the output result of the pulse neural network model path; is the coefficient of the BERT neural network model path; is the coefficient of the pulse neural network model path.
[0119] In actual applications, the outputs of the two paths are fused by dynamic mixing coefficients and , wherein, , ∈[0, 1], which can be dynamically optimized by a loss function to achieve task-driven adaptive adjustment. For example, when high-precision semantic modeling is required (such as language tasks) tends to 1 to activate the enhanced BERT neural network model path; when efficient noise suppression is required (such as dynamic perception tasks), tends to 1 to activate the enhanced pulse neural network model path.
[0120] Therefore, in order to further balance the model precision and computational efficiency, the HUAM module is also designed with a joint loss function mechanism, which includes a task loss and a regularization term.
[0121] Specifically, the task loss can adopt a cross-entropy loss to ensure the accuracy of attention weight distribution and support end-to-end optimization of long-distance dependencies. Specifically, the cross-entropy loss maximizes the similarity between the prediction and the true label, prompting the model to learn how to distribute attention weights, especially for tasks that require long-distance dependencies, which helps to enhance the model's ability to transfer information over long distances.
[0122] The regularization term can be designed as a double constraint term, and a hybrid regularization term can be introduced to balance the contributions of the BERT neural network model path and the spiking neural network model path.
[0123] The hybrid regularization term can include two main constraint terms. The first is the weight coefficient conservation, which ensures that the sum of the two coefficients approaches 1 by constraining the coefficients and The second is the smoothing constraint, which controls the changes of the coefficients and by introducing a regularization term with a coefficient μ The HUAM module balances the contributions of the BERT neural network model path and the spiking neural network model path through the dynamically generated coefficients
[0124] The HUAM module balances the contributions of the BERT neural network model path and the spiking neural network model path through the dynamically generated coefficients The HUAM module balances the contributions of the BERT neural network model path and the spiking neural network model path through the dynamically generated coefficients The HUAM module balances the contributions of the BERT neural network model path and the spiking neural network model path through the dynamically generated coefficients The HUAM module balances the contributions of the BERT neural network model path and the spiking neural network model path through the dynamically generated coefficients
[0125] Therefore, the regularization term loss function of the present application can be represented by the following formula (7):
[0126] + (7)
[0127] wherein, is the regularization term loss function; μ is the adjustment coefficient factor, which can be set according to the specific use scenario requirements.
[0128] Further, the joint loss function of the present application can be represented by the following formula (8):
[0129] (8)
[0130] wherein, represents the joint loss function of the HUAM module; represents the task loss composed of the cross-entropy loss; is the regularization term loss function; is the adjustment coefficient factor of the regularization term loss function, which can be set according to the specific use scenario requirements.
[0131] Therefore, by introducing the joint loss function mechanism, the HUAM module can simultaneously optimize the accuracy (through the task loss) and the model stability (through the regularization term), achieving the Pareto Optimal between the accuracy and the performance, providing the adaptive feature modulation capability for the network model architecture, and being able to flexibly adjust the path contribution under different computing requirements.
[0132] In practical applications, please refer to Figure 5 , Figure 5 is a flowchart of data processing of a hybrid unit driven attention modulation module provided by an embodiment of the present application.
[0133] First, the text data input to the encoding feature extraction module will complete data embedding through the embedding layer, and generate query (Q), key (K) and value (V) matrices through the attention mechanism in the Transformer architecture.
[0134] Further, the K matrix is extracted, and feature extraction is performed through the double path, i.e., the BERT neural network model path and the spiking neural network model path, respectively, and then weighted fusion is performed through the dynamic mixing weight, and attention score calculation is performed, and then input to the feedforward network to complete task prediction.
[0135] Further, the hybrid loss function can also be calculated according to the prediction result, wherein the hybrid loss function mainly includes the task loss function and the regularization term loss function.
[0136] Finally, the task loss function and the regularization term loss function are combined to obtain the hybrid loss function, and then the model parameter optimization and dynamic adjustment process are completed through the hybrid loss function.
[0137] Therefore, by introducing the HUAM module in the hybrid neural network model, the contribution proportion of the BERT neural network model path and the spiking neural network model path is adaptively adjusted, cross-paradigm information fusion is realized in the self-attention mechanism, and thus the modeling capability of the model for dynamic context and fine-grained information is enhanced.
[0138] It can be understood that the hybrid neural network model provided in the application can realize more efficient information transmission in fewer time steps and reduce information loss by introducing periodic integral and firing neuron PIF; the precision problem of the BERT neural network model in processing long-distance dependence after pulsing is solved by introducing the PSAM module, and the capture ability of context information in the NLP task is effectively improved; by introducing the HUAM module, the contribution ratio of the BERT neural network model path and the pulse neural network model path is adaptively adjusted, cross-paradigm information fusion is realized in the self-attention mechanism, and the modeling ability of the model for dynamic context and fine-grained information is enhanced.
[0139] That is, the hybrid neural network model of the application not only solves the problem of insufficient precision of the pulse neural network model, but also significantly reduces energy consumption, and performs well in multiple benchmark tests in the natural language processing field, maintains the precision comparable to the traditional ANN model, and greatly reduces the model energy consumption, which is helpful for the development and wide application of large-scale artificial intelligence models, and provides strong support for the development of large-scale artificial intelligence model industry.
[0140] Exemplarily, the general natural language understanding task GLUE (General Language Understanding Evaluation) is one of the most common evaluation standards in the field of natural language processing. According to the existing data, through experiments and simulations on public data sets, the performance indicators of the hybrid neural network model provided in the application are shown in Table 1 below. Table 1 is a model performance indicator evaluation table.
[0141] Table 1 Model performance indicator evaluation table
[0142]
[0143] Therefore, it can be determined that the hybrid neural network model provided in the application not only achieves a precision similar to that of the ANN model, but also reduces the energy consumption by about 1 times compared to the ANN model with the same precision.
[0144] Specifically, as an optional implementation, the hybrid neural network model is trained based on the following steps:
[0145] A plurality of historical natural language texts are obtained, and a plurality of natural language processing results corresponding to the historical natural language texts are obtained;
[0146] Each of the historical natural language texts is taken as a sample, and the natural language processing result corresponding to each of the historical natural language texts is taken as a sample label corresponding to the sample, and a training data set is constructed;
[0147] Pre-training the hybrid neural network model using the training data set.
[0148] In the embodiment of the present application, a plurality of historical natural language texts and natural language processing results of the plurality of historical natural language texts in the historical NLP task processing process can be obtained, and any one historical natural language text is taken as a sample and the natural language processing result of the any one historical natural language text is taken as a sample corresponding sample label to form a set of training samples. A training data set is constructed by obtaining a plurality of training samples, and then the training data set is input to a hybrid neural network model. The model parameters in the hybrid neural network model are adjusted according to each output result of the hybrid neural network model, and finally the pre-training process of the hybrid neural network model is completed.
[0149] Among them, it can be considered that the pre-training process of the hybrid neural network model is completed after reaching the pre-set pre-training times; or it can be considered that the pre-training process of the hybrid neural network model is completed when the training output result of the hybrid neural network model converges.
[0150] In actual application, after constructing the training data set, it can also be divided into a training set, a validation set and a test set, and the training set accounts for 80% of the training data set, the validation set accounts for 10% of the training data set, and the test set accounts for 10% of the training data set. The training set is used for model training, and the validation set and the test set are used for model verification.
[0151] Please refer to Figure 6 , Figure 6 is a structural schematic diagram of a natural language processing device based on a hybrid neural network model provided by the embodiment of the present application. The embodiment of the present application also provides a natural language processing device based on a hybrid neural network model, which can realize the natural language processing method based on the hybrid neural network model. The device comprises:
[0152] The text acquisition module 610 is configured to acquire the natural language text to be processed.
[0153] The text processing module 620 is configured to input the natural language text to be processed into the hybrid neural network model to acquire the natural language processing result output by the hybrid neural network model.
[0154] The hybrid neural network model is constructed based on the pulse neural network model and the BERT neural network model.
[0155] It can be understood that the contents in the above method embodiments are applicable to the device embodiments. The device embodiments specifically realize the same functions as the above method embodiments, and achieve the same beneficial effects as the above method embodiments.
[0156] Please refer to Figure 7 , Figure 7 is a hardware structure schematic diagram of an electronic device provided by an embodiment of the present application, and the electronic device comprises:
[0157] The processor 701 can be implemented in the form of a general-purpose CPU (Central Processing Unit), a microprocessor, an ASIC (Application Specific Integrated Circuit), or one or more integrated circuits, and is used to execute related programs to implement the technical solutions provided by the embodiments of the present application.
[0158] The memory 702 can be implemented in the form of a ROM (Read Only Memory), a static storage device, a dynamic storage device, or a RAM (Random Access Memory). The memory 702 can store an operating system and other application programs. When the technical solutions provided by the embodiments of the present application are implemented by software or firmware, the related program codes are stored in the memory 702 and are called and executed by the processor 701 to implement the natural language processing method based on a hybrid neural network model according to the embodiments of the present application.
[0159] The input / output interface 703 is used to realize information input and output.
[0160] The communication interface 704 is used to realize the communication interaction between the device and other devices. The communication can be realized in a wired manner (for example, a USB, a network cable, etc.) or in a wireless manner (for example, a mobile network, WIFI, Bluetooth, etc.).
[0161] The bus 705 is used to transmit information between various components (for example, the processor 701, the memory 702, the input / output interface 703, and the communication interface 704) of the device.
[0162] The processor 701, the memory 702, the input / output interface 703, and the communication interface 704 are connected to each other in the device through the bus 705.
[0163] The embodiments of the present application further provide a computer readable storage medium, which stores a computer program. The computer program is executed by a processor to implement the natural language processing method based on a hybrid neural network model.
[0164] It can be understood that the contents in the above method embodiments are all applicable to the present storage medium embodiments, the present storage medium embodiments specifically implement the functions same as those of the above method embodiments, and achieve the same beneficial effects as those of the above method embodiments.
[0165] The memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory can include a high-speed random access memory and can also include a non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state memory device. In some embodiments, the memory can optionally include a memory disposed remotely with respect to the processor, and these remote memories can be connected to the processor through a network. Examples of the above network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and a combination thereof.
[0166] The natural language processing method, device, equipment and medium based on the hybrid neural network model provided by the embodiments of the present application, which constructs a hybrid neural network model based on a spiking neural network model and a BERT neural network model, and then completes a natural language processing task through the hybrid neural network model, effectively combines the high-precision advantages of an artificial neural network model and the low-power computing advantages of a spiking neural network model, improves the efficiency and accuracy of the natural language processing task, and provides strong support for the development of large-scale artificial intelligence model industry.
[0167] The embodiments described in the embodiments of the present application are used to more clearly illustrate the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those skilled in the art can know that, with the evolution of technology and the appearance of new application scenarios, the technical solutions provided by the embodiments of the present application are also applicable to similar technical problems.
[0168] Those skilled in the art can understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and can include more or fewer steps than the figures shown, or combine certain steps, or different steps.
[0169] The device embodiments described above are only schematic, and the units described as separate components can or can not be physically separate, that is, can be located in one place, or can be distributed on multiple network units. Part or all of the modules can be selected according to actual needs to achieve the purpose of the present embodiment.
[0170] Those skilled in the art can understand that all or some of the steps in the above disclosed method, the functional modules / units in the system and the device can be implemented as software, firmware, hardware and their appropriate combinations.
[0171] The terms "first", "second", "third", "fourth" and the like in the description of this application and in the claims, if any, are used for distinguishing between similar elements and not necessarily for describing a particular sequential or chronological order. It is to be understood that the use of these terms herein is to be construed to cover a changeable order, sequence or arrangement, if any. Further, the terms "comprising", "having", "including", and "containing" and any variations thereof, are intended to cover a non-exclusive inclusion, such that a process, method, system, product or apparatus that comprises, has, includes or contains a list of elements is not necessarily limited to those elements, but can include other elements not expressly listed or inherent to such process, method, system, product or apparatus.
[0172] It should be understood that, in the application, "at least one" means one or more, and "multiple" means two or more. "And / or" is used to describe the relationship between associated objects, which means that there can be three relationships, for example, "A and / or B" can mean that there are three cases: only A, only B, and A and B at the same time, where A and B can be singular or plural. The character " / " generally represents an "or" relationship between the associated objects. "At least one of the following" or similar expressions means any combination of these items, including any combination of single or multiple items. For example, at least one of a, b or c can mean a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.
[0173] In several embodiments provided in the application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only illustrative, for example, the division of the above units is only a logical function division, and actual implementation can have another division manner, for example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the displayed or discussed each other can be indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.
[0174] The units described above as separate components can or can not be physically separated, and the components shown as units can or can not be physical units, that is, they can be located in one place, or they can be distributed on multiple network units. Part or all of the units can be selected according to actual needs to achieve the purpose of the embodiment scheme.
[0175] In addition, each function unit in each embodiment of the present application can be integrated in one processing unit, or each unit can be physically present separately, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software function unit.
[0176] When the integrated unit is realized in the form of a software function unit and sold or used as an independent product, it can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application, essentially or in part, or all or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, and includes multiple instructions used to cause a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the methods in the embodiments of the present application. The foregoing storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various other media that can store programs.
[0177] The preferred embodiments of the embodiments of the present application are described above with reference to the accompanying drawings, and are not intended to limit the scope of the embodiments of the present application. Any modifications, equivalent replacements and improvements made by those skilled in the art without departing from the scope and essence of the embodiments of the present application shall be within the scope of the embodiments of the present application.
Claims
1. A method for natural language processing based on a hybrid neural network model, characterized in that, The method comprises the following steps: acquiring natural language text to be processed; inputting the natural language text to be processed into a hybrid neural network model to acquire a natural language processing result output by the hybrid neural network model; wherein the hybrid neural network model is constructed based on a pulse neural network model and a BERT neural network model; the hybrid neural network model comprises an embedding layer, a plurality of encoding feature extraction modules and a prediction head connected in sequence; the encoding feature extraction module comprises a periodic integral and firing neuron, a periodic pulse attention module and a hybrid unit driven attention modulation module; the periodic integral and firing neuron is designed with a periodic modulation input current mechanism and a periodic reset voltage mechanism; wherein the periodic modulation input current mechanism comprises periodically modulating the input current of the periodic integral and firing neuron; and the periodic reset voltage mechanism comprises periodically resetting the output voltage of the periodic integral and firing neuron; the periodic pulse attention module is designed with a pulsing processing mechanism, an adaptive convolution mechanism and a periodic feature fusion mechanism; wherein the pulsing processing mechanism comprises calling the periodic integral and firing neuron to convert a feature sequence into a pulse sequence; the adaptive convolution mechanism comprises adjusting the size of a convolution kernel according to the number of heads and the sequence length in a multi-head attention mechanism, thereby completing one-dimensional convolution; and the periodic feature fusion mechanism comprises periodically fusing original input features and features output by adaptive convolution; the hybrid unit driven attention modulation module is designed with a dual-path hybrid encoding mechanism and a joint loss function mechanism; wherein the dual-path hybrid encoding mechanism comprises a BERT neural network model path and a pulse neural network model path, and the key matrix in the multi-head attention mechanism is extracted and weightedly fused through the two paths; and the joint loss function mechanism comprises introducing a joint loss function, which comprises a task loss and a regularization term; the encoding feature extraction module has a residual connection.
2. The natural language processing method based on a hybrid neural network model according to claim 1, characterized in that, The hybrid neural network model is obtained based on the following steps: acquiring a plurality of historical natural language texts and acquiring natural language processing results corresponding to the plurality of historical natural language texts; using each historical natural language text as a sample and using the natural language processing result corresponding to each historical natural language text as a sample label corresponding to the sample to construct a training data set; pre-training the hybrid neural network model using the training data set.
3. A hybrid neural network model based natural language processing apparatus for implementing the hybrid neural network model based natural language processing method according to claim 1, characterized by, The device comprises: a text acquisition module configured to acquire natural language text to be processed; a text processing module configured to input the natural language text to be processed into a hybrid neural network model to acquire a natural language processing result output by the hybrid neural network model; wherein the hybrid neural network model is constructed based on a pulse neural network model and a BERT neural network model.
4. An electronic device, comprising: The electronic device comprises a memory and a processor, the memory stores a computer program, and the processor implements the natural language processing method based on the hybrid neural network model according to any one of claims 1-2 when executing the computer program.
5. A computer-readable storage medium storing a computer program, the computer program comprising instructions that, when executed by a computer, cause the computer to perform the method of any one of claims 1 to 4. The computer program is executed by the processor to implement the natural language processing method based on the hybrid neural network model according to any one of claims 1-2.