EVENT-DRIVEN ATTENTION MECHANISM AND METHOD THAT DOES NOT REQUIRE MULTIPLICATION FOR NATURAL LANGUAGE PROCESSING MODELS

TR202609793A2Pending Publication Date: 2026-06-22ISTANBUL GELISIM UNIVSI
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
TR · TR
Patent Type
Applications
Current Assignee / Owner
ISTANBUL GELISIM UNIVSI
Filing Date
2026-06-17
Publication Date
2026-06-22

Smart Images

  • Figure 00000019_0000
    Figure 00000019_0000
Patent Text Reader

Abstract

The invention is a method and architecture for the energy-efficient implementation of attention mechanisms used in natural language processing models with event-driven and non-multiplication computational methods, and its features include: at least one data input module (1) that receives raw data, at least one embedding module (2) that converts the raw data into numerical vector representations, at least one event-based encoding module (3) that converts the numerical vector representations into temporally sensitive event sequences, at least one dynamic thresholding and activation control module (4) that evaluates the generated event sequences according to at least one dynamic threshold value and selects only events exceeding the specified threshold value as active events, a query submodule (5.1) that generates query vectors from active events, and a key submodule (5.1) that generates key vectors.2) It should include an aggregation-based accumulator unit (5.3) that calculates similarity scores between query and key vectors using only aggregation operations, and a value submodule (5.4) that generates value vectors; at least one integration and firing module (6) that processes the values ​​obtained from the aggregation-based attention calculation module (5) according to the integration and firing principle that mimics biological neuron behavior; at least one output analysis module (7) that converts the generated impulse sequences into standard numerical representations; and at least one final output module (8) that presents the converted data to the user as meaningful text, classification result, or natural language processing output.
Need to check novelty before this filing date? Find Prior Art

Description

1 TARIFF MULTIPLICATION OPERATION FOR NATURAL LANGUAGE PROCESSING MODELS NON-NECESSARY EVENT-DRIVEN ATTENTION MECHANISM AND METHOD 5 Technological Field: The invention relates to attention mechanisms used in natural language processing models. Energy efficient 10 with event-driven and multiplication-free computation methods It is concerned with a method and architecture for achieving this in this way. State of the Art: Large language models, widely used in natural language processing in recent years, include 15 It is primarily based on the Transformers architecture. These architectures include... The attention field mechanism examines the relationships between elements in input sequences. one of the fundamental components that enables its identification is the query, key, and dense matrix multiplication operations performed between value vectors It operates based on... 20 In current Transformer-based systems, attention calculations involve processing information from the data. constant at every inference step, regardless of its intensity or semantic significance and is carried out in a way that creates a high computational load. This situation, Computational complexity can increase to 25, especially in applications where long input sequences are processed. increase, high memory bandwidth requirement and significant energy consumption This is the reason. Floating-point number-based calculation is used in the known technique. These methods require continuous and intensive processing of data. The systems, 30 even for data sections that lack meaningful informational content or remain unchanged 2 The process continues, but this approach leads to unnecessary energy consumption and processing. This leads to the inefficient use of resources. Especially mobile devices, wearable technologies, Internet of Things (IoT) systems, intelligent sensor networks, autonomous platforms, space systems, and edge artificial intelligence with limited battery capacity. In intelligence applications, these computationally intensive requirements significantly impact performance. This creates limitations. Existing systems struggle with the information density of input data. Because it independently performs a constant and high amount of calculations, the information even for data segments that do not contain or are of low importance, energy is consumed This continues. High energy consumption leads to reduced battery life and thermal overload in processors. This leads to an increase in load and creates a need for additional cooling. Data Similarly, high electricity consumption in the centers leads to high operating costs. This leads to an increase in pollution and negatively impacts sustainability goals. Reducing the computational load caused by attention mechanisms in the current technology 15 For this purpose, quantization, pruning, model compression, and knowledge distillation various optimization methods such as distillation and low-precision calculation methods These techniques are used. However, these approaches are fundamentally based on reducing the number of model parameters and simplifying data representations or focusing on performing calculations with lower bit precision 20 and is a matrix multiplication-based calculation that forms the basis of the attention mechanism. It does not change its structure. Therefore, it provides a certain level of improvement in processing load. Even if this is possible, the energy consumption resulting from multiplication operations and memory access costs And hardware requirements are still ongoing. On the other hand, various approaches exist in the literature to reduce the need for matrix multiplication. multiplication-free or reduced multiplication count architectures Although these approaches have been proposed, they are not widely used in natural language processing applications. There are significant technical challenges that prevent its widespread use. In particular, the accurate representation of contextual relationships in language models is crucial for semantic 30. matrix in terms of preserving integrity and maintaining model accuracy Completely eliminating multiplication operations presents significant difficulties. 3 Therefore, current collision-free architectures suffer from a loss of accuracy in natural language processing tasks. limited scalability, compatibility issues with existing Transformer architectures, and widely accepted due to reasons such as inadequate performance in real-time applications He hasn't seen it. 5 In conclusion, the current attention mechanisms used in natural language processing models high energy consumption, high matrix multiplication load, hardware dependency, increased thermal requirements, limited mobile applicability, and inadequate conditions in energy-constrained environments. Considering that it has various technical problems such as performance issues. when taken into account, it eliminates the intense matrix multiplication load in attention mechanisms. while removing it, the model can maintain accuracy, reduce energy consumption, and save resources. New methods are needed that can maintain their applicability on limited platforms. It is heard. Description of the invention: This discovery is relevant to natural language processing (NLP) models. especially Large Language Models (LLM) and Transformer-based Increasing the energy efficiency of attention mechanisms used in architecture, 20 to reduce processing latency and to use hardware resources more efficiently event-driven, impulse-based (spiking) and an attention mechanism, method and software architecture that does not require multiplication It is related. The invention combines artificial intelligence, natural language processing, neuromorphic computing, and edge AI. embedded systems, artificial intelligence accelerators, semiconductor technologies, and energy-efficient It is located in the technical fields of information processing systems. The invention, thanks to its event-driven and collision-free attention mechanism, 30 It offers significant technical advantages compared to existing Transformer-based architectures. The system only processes data flow where the information content meets specified importance thresholds. 4 It performs calculations when the data exceeds the limit, for data that is uninformative or irrelevant. It does not consume energy, thus providing ultra-low power consumption and extending battery life. It significantly extends the timeframe. The invention involves intensive matrix calculations of attention. Performing operations based solely on addition, instead of multiplication operations, This leads to high energy consumption and large silicon area usage on processors. 5 This reduces or eliminates the need for multiplication units. The situation is changing with lower-cost processors, embedded systems, neuromorphic hardware, and Enables working with high hardware efficiency on artificial intelligence accelerators. This reduces the thermal load on the processor. This reduces overheating, prevents devices from overheating, and uses fan-like active cooling. The need for these systems is being reduced. Furthermore, event-driven architecture aggregates data. Instead of a processing approach, it processes events only as they occur, providing real-time processing. It improves responsiveness and significantly reduces processing latency. The unique structure of the invention combines existing linear attention with sparse attention. Unlike quantization and pruning-based methods, attention is 15 by changing the mathematical basis of the mechanism, matrix-matrix multiplication operations It transforms into an aggregation-based computation approach. Also, continuous value text. hybrid that transforms data into impulse-based representations while preserving semantic integrity Energy-efficient processing of natural language data thanks to the encoding layer. This makes it possible. In addition, energy consumption is 20% of the input data. It scales dynamically depending on complexity and event intensity. Thus, the system uses resources only to the extent needed, maximizing energy efficiency. It elevates it to the highest level. The invention relates to 25 high-performance and energy-intensive graphics processing units (GPUs). by reducing dependency, lower cost processors, embedded systems and neuromorphic It can be implemented on hardware. Thanks to its event-driven structure, the processing load is reduced. It changes dynamically depending on the density of the input data, especially sparse data. Energy consumption and processing time are significantly reduced in these structures. This allows the battery to... Its lifespan is extended and resource utilization efficiency is increased. In addition, intensive 30 Reducing the computational load lowers the thermal load on the processor. This prevents devices from overheating and eliminates the need for active cooling systems. is being reduced. In current technology, attention should be paid to Transformer-based architectures that are widely used. calculations; 5 between query, key and value representations. It is based on dense matrix multiplication operations performed. This calculation This approach requires high processing power, especially for inference. These processes involve significant energy consumption and a high memory bandwidth requirement. This leads to increased processing latency, higher hardware costs, and thermal load generation. is. 10 The primary goal of this invention is to identify the issues that arise in existing Transformer-based attention mechanisms. High energy consumption, heavy processing load, thermal issues, high hardware requirements. The goal is to overcome technical problems such as requirements and limited scalability. Another objective of the invention is to enable large language models to be used only with high-performance data. not on central units and graphics processing units, but on energy-constrained mobile devices, embedded also sustainably on systems and edge AI platforms. The goal is to enable it to be operated. To this end, the invention aims to enable the energy-efficient operation of biological nervous systems. an impulse-based information processing approach inspired by the principles of natural language processing It adapts to its field. Within the scope of the invention, data is obtained from 25 continuously flowing, sliding structures, as in traditional architectures. It is not processed in the form of decimal numbers; instead, it is processed as event streams or The system is represented as sequences of impulses (spike trains). The system only contains information. the threshold values ​​of its content that can be predetermined or dynamically updated It performs a calculation if it exceeds the limit. 6 This event-driven operating principle ensures that the data remains unchanged and devoid of information. or no action is taken for data segments deemed unimportant, thus Unnecessary energy consumption is prevented. One of the key innovative elements of the invention is the 5 that forms the basis of the attention mechanism. The goal is to completely eliminate dense matrix multiplication operations. In this context... Query, relationships between key and value representations, traditional matrix-matrix multiplication Instead of operations based on addition, only addition-based operations are used to determine this. Thus This leads to high energy consumption and significant silicon area usage on processors. The need for multiplication units is reduced or completely eliminated. is being removed. The computational architecture used in the invention involves the integration of biological neurons and It utilizes the principles of integrate-and-fire ignition, and requires careful calculations. It is performed solely based on addition operations. This approach 15 Thanks to this, computational complexity is reduced, processing delays are lowered, and Hardware efficiency is being increased. In an application example of the invention, after the embedding process of the input data... The values ​​are passed to the event-based coding layer. In this layer, continuous values ​​are pulsed. 20 They are converted into impulse sequences. The generated impulse sequences are then used in a dynamic thresholding mechanism. Only events that meet the specified conditions are considered. It is included in the calculation. The query and key are used during the attention calculation. Relationships between representations are evaluated through summation-based processes, The results obtained are processed according to the integration and ignition principle, and the necessary 25 In these cases, the data is transferred to the next processing layer. The invention also includes the representation of text data in a continuous feature space. enabling the transformation into impulse-based representations while preserving semantic integrity. A hybrid coding layer can be used. Thanks to this hybrid structure, continuous and event-based 30-bit coding is possible. semantic-based information processing methods can be used together, analyzing natural language data. Energy efficiency can be increased while preserving the content. 7 Due to the event-driven nature of the invention, the processing load depends on the complexity of the input data and the information required. It changes dynamically depending on the intensity and frequency of the event. Thus Energy consumption does not remain at a constant value; it only changes when and where it is needed. Calculations are made to the extent required. This situation provides dynamic energy scaling capability 5 By providing this, significant energy savings can be achieved, especially in sparse data structures. It allows for this to be done. The current invention modifies the mathematical basis of the attention mechanism by introducing the matrix. It completely eliminates the calculation approach based on multiplication and attention 10 calculations in an event-driven, summation-based, and dynamic energy scaling manner It enables its realization. In this respect, the invention is not only an improvement of existing methods. not its optimization, but a radical change in the way the attention mechanism is implemented. It offers a groundbreaking and unique computing paradigm. Thanks to the invention;  reducing energy consumption,  reducing transaction delays,  reducing hardware costs,  reducing the memory access load, 20  Reducing the thermal load on the processor,  reducing the need for active cooling,  extending battery life,  the ability to run large language models on low-resource hardware,  such as reducing reliance on high-performance graphics processing units 25 Significant technical gains are being achieved. The invention includes mobile devices, smartphones, wearable technologies, and the Internet of Things. applications, smart sensor networks, edge artificial intelligence systems, semiconductor technologies, Artificial intelligence accelerators, neuromorphic processors, autonomous systems, space and defense 30 energy consumption, process delay, and especially in industrial applications 8 in numerous industrial areas where hardware costs are of critical importance It is applicable. The invention also enables AI assistants to run on the device, and offline natural language processing. applications, real-time translation systems, low-power sensor nodes, 5 security systems, smart cameras, autonomous platforms, and energy-constrained communication. It can be used in applications such as infrastructure. As a result, the invention has attracted attention in large language models and natural language processing systems. its mechanism is event-driven, impulse-based, and does not require multiplication. 10 enabling its implementation, it is energy efficient, hardware-friendly and low-cost. It introduces a new, albeit delayed, computing architecture. Explaining the Figures: The invention will be described by referring to the attached figures, so that the features of the invention can be explained. It will be understood and appreciated more clearly, but the purpose of this invention is this obvious It is not about limiting it with regulations. On the contrary, the invention is defined by the accompanying claims. all alternatives, modifications, and options that could be included within the defined area The aim is to cover their equivalences. The details shown are only for 20 of the present invention. It is shown to illustrate the preferred arrangements and both the methods shaping, as well as the most useful and conceptual features of the invention's rules and principles It should be understood that these drawings are presented to provide an easily understandable definition. In these drawings; Figure 1: Schematic of the modular architecture of the system subject to the invention. 25 It is the appearance. Illustrations that will help understand this invention are shown in the attached image. They are numbered and their names are given below. 9 Explanation of References: 1. Data input module 2. Embedded module 3. Event-based coding module 5 4. Dynamic thresholding and activation control module 5. Addition-based attention calculation module 5.1. Query submodule 5.2. Switch submodule 5.3. Accumulation-based accumulator unit 10 5.4. Value submodule 6. Integration and ignition module 7. Output analysis module 8. Final output module Description of the Invention: The invention includes at least one data input module (1) that receives raw data and converts that raw data into digital form. at least one embedding module (2) that converts to vector representations, digital vector at least one event that transforms their representations into event sequences sensitive to temporal changes 20 the coding module based on (3), the generated event sequences at least one dynamic threshold By evaluating events based on their value, only those exceeding the specified threshold are activated as active events. select at least one dynamic thresholding and activation control module (4), active a query submodule (5.1) that generates query vectors from events, key vectors a key submodule (5.2) that creates the similarity between query and key vectors 25 an addition-based system that calculates scores using only addition operations including an accumulator unit (5.3) and a value submodule (5.4) that generates value vectors. a summation-based attention calculation module (5), summation-based attention calculation values ​​obtained from module (5) mimicking biological neuron behavior at least one integration and ignition system operating according to the integration and ignition principle 30 module (6) converts the generated impulse sequences into standard numeric representations at least one The output analysis module (7) provides the converted data to the user as meaningful text. at least one final output presented as a classification result or natural language processing output. It includes module (8). Detailed Description of the Invention: The terminology used here is intended solely to describe specific applications. and does not limit the scope of the invention. Used here The term "and / or" refers to any of the items listed as related. It includes one and all combinations thereof. Also, the singular "one" used here, "one" The terms "number" and "specified" are used in their plural forms unless the context explicitly indicates otherwise. It is designed to include singular forms such as those mentioned above. Furthermore, the terms used in this specification... The terms "includes" and / or "contains" refer to the specified features, steps, processes, It indicates the presence of elements and / or components, but one or more other features, steps, processes, elements, components and / or groups thereof It will be understood that it does not exclude its existence or addition. 15 Unless otherwise noted, all terms used herein (including technical and scientific terms), in the sense that a person with general knowledge in the field to which this invention belongs would generally understand it They have the same meaning. Furthermore, as defined in commonly used dictionaries... 20 The terms will have a meaning consistent with the context of the relevant field and this explanation. it should be interpreted as idealized unless otherwise explicitly defined here. or it will be understood that it will not be interpreted in an overly formal sense. The description of the invention will reveal a series of techniques and steps involved. These are: Each of them offers benefits individually, and at the same time, one or more of them or 25 In some cases, all of the other techniques described may be used in combination. Accordingly, To ensure clarity, each step in the disclosure of the invention should be presented as accurately as possible. By doing this, unnecessary repetition of all combinations will be avoided. Together, the specification and claims, such combinations fully constitute an invention and claims. It should be read with the understanding that it falls within its scope. 30 11 The system architecture of the invention is basically a data input module (1), embedding module (2), event based coding module (3), dynamic thresholding and activation control module (4), collection-based attention calculation module (5), integration and ignition module (6), output It consists of the analysis module (7) and the final output module (8). The modules are functionally linked to each other in a way that allows them to communicate with each other using data. and sequentially and / or simultaneously within an integrated pipeline architecture. It is configured to work. In this invention, the addition-based attention calculation module (5) has a query sub-module within itself. module (5.1), key submodule (5.2), summation-based accumulator unit (5.3) and value 10 It includes sub-module (5.4). These sub-modules are the energy of the attention mechanism. to enable it to be carried out efficiently and with low computational complexity It is structured. In the first stage of the invention, raw text data is entered into the system via the data input module (1), 15 the script, document content, user query, or natural language processing task Any input is provided. The input received by the data input module (1) is embedded. It is transferred to module (2). Embedding module (2) is known in the field of natural language processing. word embeddings, sub-word embeddings, contextual embeddings, or pre-trained Using methods like language model embeddings, text data can be converted into numerical vectors. It transforms each token into a vector in a high-dimensional semantic space. This corresponds to the representation of semantic information found in natural language data. Syntactic and contextual relationships are represented numerically and then processed. A standardized data structure is being created for these steps. In the invention, the continuous values ​​generated by the embedded module (2) are encoded using event-based coding. The event-based coding module (3) transfers the continuous values ​​to the module. It transforms them into temporal sequences of events or impulse sequences. During this transformation... Data is measured not by its absolute magnitude, but by its change over specific time intervals. It is coded according to the amount. (If there is no change in the data, "0" (silence) is produced.) 30 Thus, the system processes only the changes that carry information, eliminating unnecessary calculations. 12 reducing its activities and energy in line with the event-driven working principle. It ensures efficiency. In an application example, the difference between consecutive time samples is predetermined. If a change threshold is exceeded, a pulse signal 5 is generated for the relevant data element. is being generated. Data that does not change or remains below the defined level. No events are generated for the elements. This allows the system to function semantically. It creates sparse event sequences that represent meaningful changes and eliminates unnecessary calculations. It reduces the load. Thanks to infrequent event representation, the number of memory accesses and data transfers is reduced. By reducing the load and processor usage intensity, the system's total energy consumption is reduced to 10. is being reduced. In the invention, the event-based coding module (3) converts continuous data into impulse sequences. a hybrid encoding structured to preserve semantic integrity during conversion It uses the following approach. Thus, the contextual relationships contained in natural language data 15 While preserving data, unnecessary data processing load is reduced. This is the hybrid coding system in question. this approach combines the advantages of both consistently valued representation and event-based processing. It combines the computational efficiency provided by its architecture. The generated event The sequences are passed to the dynamic thresholding and activation control module (4). In the invention, the dynamic thresholding and activation control module (4) controls each coded event. by evaluating its amplitude, intensity, frequency, temporal distribution, or significance level It decides whether the incident will be processed or not. The signal strength threshold value. The underlying events are considered low-information-content data, noise, or irrelevant data. They are being evaluated and removed from the processing line. Thus, the system has a low 25% An activation structure is being created, and only data with high informational value will be used. Events exceeding the threshold value are marked as active events and added to the collection. The attention-based calculation module (5) is transmitted. In an application example, the threshold value can be defined as a constant, or an alternative 30 In the application example, the threshold value depends on the density of the input data, system load, and hardware. dynamically, taking into account the current state and previous layer outputs. 13 It can be updated. Thanks to the dynamic thresholding mechanism, the system handles variable data. It can adapt to workflows and different working conditions, thus ensuring accuracy. An optimal balance can be struck between calculation costs. Active events that were found to exceed the threshold value in the invention are sum-based attention calculation 5. It is processed within module (5). Addition-based attention calculation module (5), query submodule (5.1), key submodule (5.2), collection-based accumulator unit It consists of the (5.3) and value submodule (5.4). The query submodule (5.1) is active Using the data obtained from the events, which information should the relevant event focus on? It creates query vectors that represent what is needed. In other words, query 10 vectors indicate which contextual information is important at the system's current processing step. It creates a representation structure for determination. The key submodule (5.2) is active. events that represent the characteristic features of the information available within the system It generates key vectors. Key vectors determine which information is associated with the query. It is used to determine this. In an application example, query 15 vectors and key vectors, temporal characteristics of active events, event density, event frequency, event amplitude, contextual embedding information, or previous layer outputs It can be created using the query submodule (5.1). key vectors generated by the key submodule (5.2) The similarity relationship between them is 20 within the collection-based accumulator unit (5.3). is calculated. Aggregation-based accumulator unit (5.3), query and key vectors It calculates the similarity scores between them using only addition operations, and It accumulates the results obtained over time. In this context, the accumulator unit (5.3), without the need for matrix multiplication or multiplication-accumulation (MAC) operations It enables the calculation of attention scores. 25 In an application example, the accumulator unit (5.3) is between the query and key vectors. sum of absolute differences, Manhattan distance, logarithmic addition operations, sign based addition methods or other similar methods that only require addition. It can use the following criteria. 30 14 The value submodule (5.4) represents the content information related to active events. It forms value vectors. Value vectors determine which attention mechanism results in which value vectors. It creates the data structure that determines where the information will be transferred to the next layer. Similarity scores calculated by the accumulator unit (5.3) based on the collection, value Attention 5 is associated with the value vectors generated by submodule (5.4). The output is obtained. Thus, the system obtains what it needs through query vectors. Identifying information, and the relationships between existing information through key vectors. It evaluates and selects the relevant content through value vectors for the next step. It transfers it to the layer. Thanks to this structure, it converts energy-dense matrix multiplication units. The attention mechanism can be activated without needing it. Furthermore, the aforementioned 10 architecture, low-power processors, edge computing devices, embedded systems, and It can be effectively applied to neuromorphic devices. The invention includes matrix multiplication operations, which form the basis of attention calculations. It is completely eliminated, and the relationship between the query and the key vectors is only 15. It is calculated using a unique algorithm that requires an "addition (ADD)" operation. In an application example, these calculations would involve logarithmic addition or Manhattan substitution. The sum of absolute differences, distance-based, requiring only the "Addition (ADD)" operation. 20 requiring logarithmic addition, sign-based addition, or simply addition operations. This can be done using other similarity criteria. Alternative approach In these examples, the similarity criteria in question are based on application requirements and hardware. depending on the limitations, delay tolerance, or target energy consumption values It can be selected and adapted. In the system of the invention, the following are obtained from the collection-based accumulator unit (5.3) The results are transferred to the integration and ignition module (6). Integration and firing module (6), membrane inspired by the working principle of biological neurons It uses a potential-based model. In this context, from the accumulator unit The incoming data are accumulated over specific time intervals, and the resulting membrane 30 Its potential is compared with a predefined ignition threshold. Membrane When its potential reaches or exceeds the firing threshold, the neuron "fires". (spike) occurs, meaning that an impulse signal is generated by the module and the output is sent to the next one. It transmits the output to the decoding module (7), which is the layer. The membrane potential If it remains below the firing threshold, no output is generated, and The system remains silent. This approach allows only meaningful information to be collected. Data is processed and energy consumption is reduced. This is how biological neurons work. 5 It is a digital copy of the energy saving principle. The integration and ignition in question... This mechanism enables efficient modeling of temporal dependencies and It improves the overall processing efficiency of the system by preventing unnecessary activations. The output analysis module (7) included in the invention, the spike from the last layer 10 It converts data in the form of strings into standard numerical representations. In the application example, the data obtained as a result of this transformation is the word. probabilities, classification outputs, semantic representations, or other natural language processing outputs It can be created in this way. The converted data is sent to the final output module (8) is transmitted and the user receives meaningful text, command output, classification result or 15 This is presented as a result of other natural language processing tasks. Thus, the event Sparse temporal representations obtained from the data-based processing architecture are provided by the user. They are converted into interpretable standard output formats. According to the invention's methodology, the process involves digitizing the raw text data received into the system. It starts with the data input and embedding step, where the data is converted into vectors, then the word The vectors in question are sensitive to temporal changes via an event-based coding module. They are converted into impulse sequences. The generated events are then subjected to dynamic thresholding and Only those exceeding the specified threshold value are evaluated in the activation control step. Active events are processed. Events exceeding the threshold value are added to the total. The query and key vectors are transferred to the attention-based calculation module, where they are processed. The relationship between them is based solely on addition operations, rather than matrix multiplication operations. It is calculated using [method]. The values ​​obtained as a result of the calculation are integrated. and accumulated in the ignition mechanism until predetermined ignition conditions are met. If this is provided, output is generated. In the final stage, the generated impulse sequences are 30. In the output analysis step, it is converted into standard numerical representations and presented to the user. 16 meaningful text, as a result of classification or other natural language processing outputs It is presented. In one application example of the invention, a smart device works offline on a mobile device. The personal assistant system is being discussed. This system allows the user to communicate via voice or text. 5 It is configured to process commands with low energy consumption. User The command "Schedule a dentist appointment for tomorrow at 09:00" given by the system, data It is received by the input module (1) and transferred to the embedding module (2). The embedding module (2), It converts the tokens that make up the command into numerical vectors. For example, “tomorrow”, “09.00”, The expressions "dentist" and "appointment" are represented as vectors representing their semantic features. The vectors are encoded. The generated vectors are sent to the event-based coding module (3). Event-based The encoding module (3) analyzes the semantic changes between successive markers and It only generates a spike signal for tokens with high information content. In the application example, words like "for" and "create" have low contextual significance. While no events are generated for the keywords “tomorrow”, “09.00”, “dentist” and “appointment”, 15 events are generated for these keywords. Impulse signals are generated. The generated impulse signals undergo dynamic thresholding and activation. The intensity and frequency of each event are transferred to the control module (4). This module transmits the intensity and frequency of each event. and assesses its contextual significance. For example, information containing date and time. While a lower threshold value can be applied for markers, for general purpose words... A higher threshold value can be applied. Thus, only the specified threshold value of 20 Excessive active events are being processed. Active event collection-based attention calculation. It is passed to the module (5). The query submodule (5.1) needs the current step of the system. It creates query vectors that represent the information it hears. The Key submodule (5.2) is active. It generates key vectors that represent characteristic features of events. (Aggregation) The accumulator unit based on (5.3) only relates the query and key vectors to 25 It calculates using aggregation-based operations. In an application example, the system uses "appointment". By determining the relationship between the indicators "tomorrow" and "09.00", and identifying that they are the same. It determines that it is located within a contextual structure. During this calculation, the matrix... Multiplication, dot product, or multiplication-accumulation operations are not used. Instead, The sum of the absolute differences between vector components or other summation-based calculations. 30 Similarity criteria are used. The Value submodule (5.4) contains content related to active events. It creates value vectors representing its information and is generated by the accumulator unit (5.3). 17 It correlates the obtained similarity scores with the resulting attention output, integration, and... The outputs are transferred to the ignition module (6). The integration and ignition module (6) transmits the attention outputs. It accumulates over a specific time window. The accumulated values ​​are predefined. When the firing threshold is exceeded, the module generates a pulse signal. In this application example, When the relationship between the markers “appointment”, “tomorrow” and “09.00” reaches a sufficient level, 5 The system detects the user's intention to create a calendar. The generated impulse signals are output. The output is transferred to the analysis module (7). The output analysis module (7) processes the impulse sequences. by converting the user's intent and relevant parameters into standard numerical representations. The final output module (8) determines the new information in the device calendar using the information obtained. It creates a record and tells the user, "Your dentist appointment is tomorrow at 9:00 AM." It gives feedback in the form of "created." In this application example, the system, the command It does not continuously process all the identifiers it contains, only their informational value. It evaluates high-profile events. Thus, traditional Transformer-based attention lower processing load, fewer memory accesses and lower performance compared to other mechanisms. The natural language processing task is performed using energy consumption. Thanks to this, the invention enables smart 15 phones, smartwatches, headphones, embedded systems and other energy-constrained end-of-line artificial intelligence It can be effectively implemented on intelligence platforms. 25

Claims

18 REQUESTS 1- The invention provides an event-driven approach for natural language processing models that does not require multiplication. It is an attention mechanism and method, and its characteristic is;  at least one data input module (1) that receives the raw data, 5  at least one that converts the raw data in question into numerical vector representations embedded module (2),  numerical vector representations of event sequences sensitive to time changes at least one event-based coding module that converts (3),  The generated event sequences are based on at least one dynamic threshold value. By evaluating the situation, only events exceeding the specified threshold are activated. select at least one dynamic thresholding and activation control module (4),  a query submodule that generates query vectors from active events (5.1), a key submodule (5.2) that generates key vectors, query and key 15 Similarity scores between vectors can only be obtained through summation operations. a summation-based accumulator unit (5.3) that calculates using and an aggregate containing a value submodule (5.4) that generates value vectors attention-based calculation module (5),  The values ​​obtained from the addition-based attention calculation module (5) are 20 integration and firing that mimics biological neuron behavior at least one integration and ignition module (6) which operates according to the principle,  at least one that converts the generated impulse sequences into standard numerical representations output analysis module (7),  Transformed data provides the user with meaningful text and classification results. 25 or at least one final output module that presents it as natural language processing output (8) is included.