Attention computation systems and methods in quantized transformer models

CN122837784APending Publication Date: 2026-09-29TESLA INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610376432.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2025-07-14
Filing Date
2026-03-25
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

实践时,该步骤可能包括:计算指数和执行浮点算术,这些浮点算术在缺乏专用浮点单元或未对超越函数进行优化支持的嵌入式硬件上的开销可能过高

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122837784A_ABST
    Figure CN122837784A_ABST
Patent Text Reader

Abstract

Disclosed herein are methods and systems for navigation paradigms. One method includes receiving a set of quantized values resulting from a matrix multiplication of a quantized query vector and a quantized key vector, the set of quantized values corresponding to sensor data associated with an ego; identifying a maximum value using an integer compare operation; for at least one quantized value in the set of quantized values, calculating a difference between the quantized value and the maximum value; performing a bit shift operation on the difference between the quantized value and the maximum value to generate a shifted value corresponding to a quantization scale that is a power of 2; performing a power of 2 operation on the shifted value to generate a power of 2 value; aggregating the power of 2 values to generate an aggregated value; and executing a model using the aggregated value to output a navigation instruction to be executed by the ego.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-references to related applications This application claims priority to U.S. Provisional Patent Application No. 63 / 780,060, filed March 28, 2025, the entire contents of which are incorporated herein by reference for all purposes. Technical Field

[0002] This application relates to performing efficient attention computations in quantization transformer models, including optimized softmax operations suitable for deployment on edge computing platforms for autonomous navigation and other real-time applications. Background Technology

[0003] Autonomous navigation techniques for autonomous vehicles and robots (sometimes referred to as "egos") have advanced with the development of computer technology. These advancements allow for safer and more reliable autonomous navigation of egos. Transformer models are increasingly being adopted in autonomous navigation systems because they can handle complex multimodal data and capture long-range dependencies within sequential inputs. Unlike traditional convolutional neural networks or recurrent neural networks, transformers can simultaneously focus on both spatial and temporal features across high-dimensional inputs. This capability enables more accurate perception, trajectory prediction, and decision-making in dynamic environments. Furthermore, the scalability and attention-based architecture of transformer models make them well-suited for integrating contextual information that can be incorporated into autonomous navigation decisions.

[0004] However, deploying the transformer model in real-time (or near-real-time) autonomous navigation systems presents significant computational challenges, especially when such systems operate on edge computing platforms with limited processing power and energy resources. The attention mechanism, a core component of the transformer model, requires intensive matrix operations followed by normalization using the softmax function. In practice, this step may involve calculating the exponent and performing floating-point arithmetic, which can be prohibitively expensive on embedded hardware lacking dedicated floating-point units or optimized support for transcendental functions. Summary of the Invention

[0005] Although quantization techniques have been introduced to improve runtime efficiency and reduce memory usage by converting model parameters and activations to a lower precision format, softmax computation remains a computational bottleneck that places excessive demands on computing resources. Traditional softmax implementations require dequantization of intermediate values ​​to perform floating-point exponentiation and normalization, thus offsetting many of the performance benefits gained through quantization. Additionally, to ensure numerical stability and maintain model accuracy, some solutions may introduce clamping functions, such as ReLU6 or bounded clamps (clamping layers), which constrain the dynamic range of activation values. However, these functions are not compatible with efficient hardware operation and often require additional multiplication or division during dequantization.

[0006] Existing and traditional approaches to mitigating these inefficiencies (such as replacing exponential functions with base-2 powers) offer limited hardware compatibility or require specialized instructions that are not universally applicable across all platforms. Consequently, current transformer implementations in autonomous navigation systems struggle to maintain real-time (or near-real-time) responsiveness without sacrificing model fidelity or significantly increasing power consumption.

[0007] Therefore, a more efficient method for performing attention normalization in quantization transformer models, tailored for deployment on edge computing platforms, is needed. This method should minimize reliance on floating-point operations, achieve efficient scaling using integer-friendly instructions, and maintain the accuracy of the attention mechanisms required for safe and reliable autonomous navigation.

[0008] The optimized base-2 softmax scheme disclosed in this paper continuously improves both the training and inference workflows of machine learning models (e.g., models using transformers) by replacing traditional exponentiation and division with hardware-friendly shifts, table lookups, and single normalized traversals, thereby reducing latency and power consumption while maintaining numerical fidelity. The methods and systems discussed in this paper reduce computational overhead and memory traffic, thereby achieving higher throughput and lower power consumption on edge processors.

[0009] In some embodiments, the methods and systems discussed herein provide an optimized "base-2 softmax" pipeline that replaces the floating-point exponentiation and division traditionally used to normalize attention scores in transformer models with a sequence of integer-only operations: (i) constraining the query-key accumulator to a power of 2 quantization scale; (ii) shifting each score at this scale instead of dequantizing via multiplication; (iii) converting the shifted integers to weights via a single base-2 exponentiation instruction (such as exp2 / fscale); (iv) summing these weights using narrow-word integer addition; and / or (v) normalizing by multiplying each weight by the integer reciprocal of the sum. Because each arithmetic step after matrix multiplication remains in the fixed-point domain, the methods and systems discussed herein reduce latency, memory traffic, and energy consumption while providing numerically similar probabilities, enabling faster (and sometimes real-time) transformer inference.

[0010] In some embodiments, the technology described herein relates to a method for navigating an individual by ingesting sensor data, the method comprising: one or more processors receiving a set of quantized values ​​generated by matrix multiplication of a quantization query vector and a quantization key vector in an attention layer of a transformer model, the set of quantized values ​​corresponding to sensor data associated with the individual; one or more processors identifying a maximum value from the set of quantized values ​​using an integer comparison operation; for at least one quantized value in the set of quantized values, one or more processors calculating a difference between the quantized value and the maximum value; one or more processors performing a shift operation on the difference between the quantized value and the maximum value corresponding to a quantization scale raised to a power of 2 to generate a shifted value; one or more processors performing a power-to-base 2 operation on the shifted value to generate a power-to-base 2 value; one or more processors aggregating the power-to-base 2 value to generate an aggregated value; and one or more processors using the aggregated value to execute a model to output navigation instructions to be executed by the individual.

[0011] In some embodiments, the techniques described herein relate to a method in which the execution model further includes: one or more processors estimating the localization time or contact time of at least one dynamic object detected in a camera image associated with the self; and when the estimated localization time or contact time meets a threshold, navigation instructions guide evasive steering or braking maneuvers.

[0012] In some embodiments, the techniques described herein relate to a method in which a quantization scale raised to the power of 2 is selected such that at least one shifted value falls within a defined range.

[0013] In some embodiments, the technology described herein relates to a method in which navigation instructions include self-control commands, which include at least one of: steering angle adjustment, lane change initiation, longitudinal acceleration or deceleration, and target speed setpoint.

[0014] In some embodiments, the techniques described herein relate to a method in which sensor data corresponds to two-dimensional image data captured by at least one camera associated with the self.

[0015] In some embodiments, the techniques described herein relate to a method that further includes: one or more processors using aggregated values ​​to normalize the set of quantized values; and one or more processors using the aggregated normalized values ​​to execute a model to output navigation instructions to be executed by the processor itself.

[0016] In some embodiments, the techniques described herein relate to a method in which normalization includes dividing each of the base-2 exponentially-operated values ​​by the aggregated value to produce a probability distribution over a set of quantized values.

[0017] In some embodiments, the techniques described herein relate to a computer system for navigating an individual by ingesting sensor data. The computer system includes a computer-readable medium having a non-transitory instruction set that, when executed by at least one processor, causes the at least one processor to: receive a set of quantized values ​​generated by matrix multiplication of a quantization query vector and a quantization key vector in an attention layer of a transformer model, the set of quantized values ​​corresponding to sensor data associated with the individual; identify a maximum value from the set of quantized values ​​using an integer comparison operation; for at least one quantized value in the set of quantized values, calculate the difference between the quantized value and the maximum value; perform a shift operation on the difference between the quantized value and the maximum value corresponding to a quantization scale raised to a power of 2 to generate a shifted value; perform a power-to-base operation on the shifted value to generate a power-to-base value; aggregate the power-to-base value to generate an aggregated value; and execute a model using the aggregated value to output navigation instructions to be executed by the individual.

[0018] In some embodiments, the techniques described herein relate to a computer system in which the execution model further includes: estimating the localization time of at least one dynamic object detected in a camera image associated with the self; and, when the estimated localization time meets a threshold, guiding evasive steering or braking maneuvers.

[0019] In some embodiments, the techniques described herein relate to a computer system in which a quantization scale raised to the power of 2 is selected such that at least one shifted value falls within a defined range.

[0020] In some embodiments, the technology described herein relates to a computer system in which navigation instructions include self-control commands, which include at least one of: steering angle adjustment, lane change initiation, longitudinal acceleration or deceleration, and target speed setpoint.

[0021] In some embodiments, the techniques described herein relate to a computer system in which sensor data corresponds to two-dimensional image data captured by at least one camera associated with the self.

[0022] In some embodiments, the techniques described herein relate to a computer system in which instructions further cause at least one processor to: normalize a set of aggregated values ​​using aggregated values; and execute a model using the normalized aggregated values ​​to output navigation instructions to be executed by the processor itself.

[0023] In some embodiments, the techniques described herein relate to a computer system in which normalization includes dividing each of the base-2 exponentiated values ​​by the aggregated value to produce a probability distribution over a set of quantized values.

[0024] In some embodiments, the techniques described herein relate to a computer system for navigating an individual by ingesting sensor data. The computer system includes at least one processor configured to: receive a set of quantized values ​​generated by matrix multiplication of a quantization query vector and a quantization key vector in an attention layer of a transformer model, the set of quantized values ​​corresponding to sensor data associated with the individual; identify a maximum value from the set of quantized values ​​using an integer comparison operation; for at least one quantized value in the set of quantized values, calculate the difference between the quantized value and the maximum value; perform a shift operation on the difference between the quantized value and the maximum value corresponding to a quantization scale raised to a power of 2 to generate a shifted value; perform a power-to-base operation on the shifted value to generate a power-to-base value; aggregate the power-to-base value to generate an aggregated value; and execute a model using the aggregated value to output navigation instructions to be executed by the individual.

[0025] In some embodiments, the techniques described herein relate to a computer system in which the execution model further includes: estimating the localization time of at least one dynamic object detected in a camera image associated with the self; and, when the estimated localization time meets a threshold, guiding evasive steering or braking maneuvers.

[0026] In some embodiments, the techniques described herein relate to a computer system in which a quantization scale raised to the power of 2 is selected such that at least one shifted value falls within a defined range.

[0027] In some embodiments, the technology described herein relates to a computer system in which navigation instructions include self-control commands, which include at least one of: steering angle adjustment, lane change initiation, longitudinal acceleration or deceleration, and target speed setpoint.

[0028] In some embodiments, the techniques described herein relate to a computer system in which sensor data corresponds to two-dimensional image data captured by at least one camera associated with the self.

[0029] In some embodiments, the techniques described herein relate to a computer system in which at least one processor is further configured to: normalize a set of aggregated values ​​using aggregated values; and execute a model using the normalized aggregated values ​​to output navigation instructions to be executed by the processor itself. Attached Figure Description

[0030] Non-limiting embodiments of this disclosure are described by way of example, including accompanying drawings, which are schematic and not intended to be drawn to scale. Unless indicated as background art, the drawings illustrate various aspects of this disclosure.

[0031] Figure 1A The illustration shows components of a self-quantification system according to an embodiment.

[0032] Figure 1B The illustration shows various sensors associated with a vehicle (or other type of self) according to an embodiment.

[0033] Figure 1C The diagram illustrates the components of the body according to an embodiment.

[0034] Figure 2 The illustration shows a flowchart of execution in a quantization system according to an embodiment.

[0035] Figure 3 The illustration shows a visualization of the quantification performed using the methods and systems discussed herein, according to an embodiment. Detailed Implementation

[0036] Now, reference is made to the illustrative embodiments depicted in the accompanying drawings, which will be described herein using specific language. However, it should be understood that this is not intended to limit the scope of the claims or this disclosure. Any changes and modifications to the inventive features illustrated herein, and additional applications of the principles of the subject matter illustrated herein, that can be conceived by those skilled in the art, should be considered within the scope of the subject matter disclosed herein. Other embodiments and / or other changes may be used without departing from the spirit or scope of this disclosure. The illustrative embodiments described in the detailed description are not intended to limit the presented subject matter.

[0037] Figure 1A These are non-restricted examples of system components that can implement the methods and systems discussed in this paper. Figure 1A The diagram illustrates the components of an AI-enabled visual data analysis system 100. System 100 may include an analysis server 110a, a system database 110b, an administrator computing device 120, entities 140a-b (collectively referred to as one or more entities 140), entity computing devices 141a-c (collectively referred to as entity computing devices 141), and a server 160. System 100 is not limited to the components described herein and may include additional or other components not shown for brevity, which are considered to be within the scope of the embodiments described herein.

[0038] The components described above can be connected via network 130. Examples of network 130 may include, but are not limited to, private or public local area networks (LANs), wireless local area networks (WLANs), metropolitan area networks (MANs), wide area networks (WANs), and the Internet. Network 130 may include wired and / or wireless communications according to one or more standards and / or via one or more transmission media.

[0039] Communication via network 130 can be performed according to various communication protocols, such as Transmission Control Protocol and Internet Protocol (TCP / IP), User Datagram Protocol (UDP), and the Institute of Electrical and Electronics Engineers (IEEE) communication protocol. In various examples, network 130 may include wireless communication according to the Bluetooth specification set or another standard or proprietary wireless communication protocol. In another example, network 130 may also include communication via cellular networks, including, for example, GSM (Global System for Mobile Communications), CDMA (Code Division Multiple Access), or EDGE (Enhanced Data for Global Evolution) networks.

[0040] System 100 illustrates an example of a system architecture and components that can be used to train and execute one or more AI models (such as one or more AI models 110c). Specifically, as Figure 1A As depicted and described herein, analysis server 110a can use data retrieved from self 140 to execute one or more AI models 110c to make navigation decisions (e.g., using data streams 172 and 176). When one or more AI models 110c are trained, each self 140 in self 140 can access and execute one or more trained AI models 110c. For example, a vehicle 140a with self computing device 141a can transmit its camera feeds to one or more trained AI models 110c and can determine the occupancy status of its surrounding environment (e.g., data stream 174). Furthermore, data ingested and / or predicted by one or more AI models 110c relative to self 140 (at inference time) can also be used to improve one or more AI models 110c. Therefore, system 100 depicts a continuous loop that can periodically improve the accuracy of one or more AI models 110c. Moreover, system 100 depicts a loop in which data received from self 140 can be used at the training phase in addition to the inference phase.

[0041] Analysis server 110a can be configured to collect, process, and analyze navigation data (e.g., images captured during navigation) and various sensor data collected from self 140. This collected data can then be processed and organized into a training dataset. The training dataset can then be used to train one or more AI models, such as AI model 110c. Analysis server 110a can also be configured to collect visual data from self 140. Using AI model 110c (trained using the methods and systems discussed herein), analysis server 110a can generate navigation decisions for self 140.

[0042] exist Figure 1A In the diagram, AI model 110c is illustrated as a component of system database 110b, but AI model 110c can be stored in different or separate components (such as cloud storage devices or any other data repository accessible by analytics server 110a).

[0043] The analytics server 110a can also be configured to display an electronic platform illustrating various training attributes used to train the AI ​​model 110c. This electronic platform can be displayed on the administrator computing device 120, allowing analysts to monitor the training of the AI ​​model 110c. An example of an electronic platform generated and hosted by the analytics server 110a could be a web-based application or website configured to display the training dataset and / or training status / metrics of the AI ​​model 110c collected from itself 140.

[0044] The analytics server 110a can be any computing device that includes a processor and a non-transitory machine-readable storage device capable of performing the various tasks and processes described herein. Non-limiting examples of such computing devices may include workstation computers, laptop computers, server computers, etc. While system 100 includes a single analytics server 110a, system 100 may include any number of computing devices operating in a distributed computing environment, such as a cloud environment.

[0045] Self 140 can represent various electronic data sources that transmit data associated with its previous or current navigation session to analysis server 110a. Self 140 can be any device configured for navigation, such as vehicle 140a and / or truck 140c. Self 140 is not limited to vehicles and can also include robotic devices. For example, self 140 can include robot 140b, which can represent a general-purpose bipedal autonomous humanoid robot capable of navigating various terrains. Robot 140b can be equipped with software to achieve balance, navigation, perception, or interaction with the physical world. Robot 140b can also include various cameras configured to transmit visual data to analysis server 110a.

[0046] Although referred to herein as "self," self 140 may or may not be an autonomous device configured for automated navigation. For example, in some embodiments, self 140 may be controlled by a human operator or a remote processor. Self 140 may include various sensors, such as Figure 1B The sensors depicted herein can be configured to collect data as the self 140 navigates various terrains (e.g., roads). Analysis server 110a can collect the data provided by the self 140. For example, analysis server 110a can obtain navigation session and / or road / terrain data (e.g., images of roads navigated by the self 140) from various sensors, such that the collected data is ultimately used by AI model 110c for training purposes.

[0047] As used herein, a navigation session corresponds to a journey along the route of the vehicle 140, whether the journey is autonomous or human-controlled. In some embodiments, the navigation session may be used for data collection and model training purposes. However, in some other embodiments, the vehicle 140 may refer to a vehicle purchased, rented, leased, etc., by a consumer, and the purpose of the journey may be categorized as everyday use. A navigation session may begin when the vehicle 140 moves from a non-moving location beyond a threshold distance (e.g., 0.1 miles, 100 feet) or exceeds a threshold speed (e.g., exceeding 0 mph, exceeding 1 mph, exceeding 5 mph). A navigation session may end when the vehicle 140 returns to a non-moving location and / or is turned off (e.g., when the driver leaves the vehicle).

[0048] Self 140 can represent a collection of self-entities monitored by analysis server 110a to train one or more AI models 110c. For example, the driver of vehicle 140a can authorize analysis server 110a to monitor data associated with their respective vehicles. As a result, analysis server 110a can utilize various methods discussed herein to collect sensor / camera data and generate training datasets to train one or more AI models 110c accordingly. Analysis server 110a can then apply one or more trained AI models 110c to analyze the data associated with self 140 and predict navigation decisions. Moreover, additional / continuous data associated with self 140 can be processed and added to the training dataset, causing analysis server 110a to recalibrate one or more AI models 110c accordingly. Thus, system 100 depicts a loop in which navigation data received from self 140 can be used to train one or more AI models 110c. Self 140 may include a processor that executes one or more trained AI models 110c for navigation purposes. During navigation, the self 140 can collect additional data about its navigation session, and this additional data can be used to calibrate one or more AI models 110c. That is, the self 140 represents a self that can be used to train, execute / use, and recalibrate one or more AI models 110c. In a non-limiting example, the self 140 represents vehicles purchased by customers that can use one or more AI models 110c for autonomous navigation while simultaneously improving upon one or more AI models 110c.

[0049] Self 140 can be equipped with various technologies that allow it to collect data from its surroundings and (potentially) navigate autonomously. For example, Self 140 can be equipped with an inference chip to run autonomous driving software.

[0050] Each of the various sensors in the self 140 can monitor the collected data associated with different navigation sessions and transmit it to the analysis server 110a. Figures 1B to 1C A block diagram of a sensor integrated within body 140 according to an embodiment is illustrated. Relative to Figures 1B to 1C The number and location of each sensor discussed can depend on Figure 1A The types of self discussed herein. For example, robot 140b may include sensors different from those of vehicle 140a or truck 140c. For example, robot 140b may not include airbag activation sensor 170q. Moreover, the locations of the sensors in vehicle 140a and truck 140c may differ from those in other vehicles. Figure 1C The positions shown in the diagrams are different.

[0051] As discussed herein, various sensors integrated within each self 140 can be configured to measure and associate various data with each navigation session. Analysis server 110a can periodically collect the data monitored and collected by these sensors, wherein the data is processed according to the methods described herein and used to train AI model 110c and / or execute AI model 110c to generate an occupancy map.

[0052] Self 140 may include user interface 170a. User interface 170a may refer to self computing device (e.g., Figure 1A The user interface 170a is a user interface for the self-computing device 141 in the vehicle. The user interface 170a can be implemented as a display screen, head-up display, touchscreen, etc., integrated with or coupled to the vehicle's interior. The user interface 170a may include input devices such as touchscreens, knobs, buttons, keyboards, mice, gesture sensors, steering wheels, etc. In various embodiments, the user interface 170a can be adapted to provide input to other devices or sensors of the self-computing device 140 (e.g., ...). Figure 1B The sensor (such as controller 170c) shown in the figure provides user input (e.g., as some kind of signal and / or sensor information).

[0053] User interface 170a may also be implemented using one or more logical devices that can be adapted to execute instructions, such as software instructions, to implement any of the various processes and / or methods described herein. For example, user interface 170a may be adapted to establish communication links, transmit and / or receive communications (e.g., sensor signals, control signals, sensor information, user input, and / or other information), or perform various other processes and / or methods. In another example, a driver may use user interface 170a to control the temperature of body 140 or activate its features (e.g., autonomous driving or steering system 170o). Therefore, user interface 170a can be combined with other sensors described herein to monitor and collect driving session data. User interface 170a may also be configured to display various data generated / predicted by analysis server 110a and / or AI model 110c.

[0054] Orientation sensor 170b can be implemented as one or more of a compass, buoy, accelerometer, and / or other digital or analog devices capable of measuring the orientation of body 140 (e.g., the magnitude and direction of roll, pitch, and / or yaw relative to one or more reference orientations such as gravity and / or magnetic north). Orientation sensor 170b can be adapted to provide heading measurements for body 140. In other embodiments, orientation sensor 170b can be adapted to use time series of orientation measurements to provide roll rate, pitch rate, and / or yaw rate for body 140. Orientation sensor 170b can be positioned and / or adapted to make orientation measurements relative to a specific coordinate system of body 140.

[0055] The controller 170c can be implemented as any suitable logic device (e.g., a processing device, microcontroller, processor, application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), memory storage device, memory reader, or other device or combination of devices), which can be adapted to execute, store, and / or receive appropriate instructions, such as software instructions that implement control loops for controlling various operations of self 140. Such software instructions can also implement methods for processing sensor signals, determining sensor information, providing user feedback (e.g., via user interface 170a), querying device operating parameters, selecting device operating parameters, or performing any of the various operations described herein.

[0056] The communication module 170e can be implemented as any wired and / or wireless interface configured to transmit sensor data, configuration data, parameters, and / or other data and / or signals to... Figure 1A Any of the features shown (e.g., analysis server 110a). As described herein, in some embodiments, the communication module 170e may be implemented in a distributed manner, such that the parts of the communication module 170e are distributed across... Figure 1B Implemented within one or more of the components and sensors shown. In some embodiments, the communication module 170e may delay the transmission of sensor data. For example, when self 140 lacks network connectivity, the communication module 170e may store the sensor data in a temporary data storage device, and transmit the sensor data when self 140 is identified as having correct network connectivity.

[0057] The speed sensor 170d can be implemented as an electronic pitot tube, a metering gear or wheel, a water velocity sensor, a wind speed sensor, a wind rate sensor (e.g., direction and amplitude), and / or a device capable of measuring or determining the linear velocity of itself 140 (e.g., in the surrounding medium and / or aligned with the longitudinal axis of itself 140) and providing such measurement as a sensor signal that can be transmitted to various devices.

[0058] The gyroscope / accelerometer 170f can be implemented as one or more electronic sextants, semiconductor devices, integrated chips, accelerometer sensors, or other systems or devices capable of measuring the angular velocity / acceleration and / or linear acceleration (e.g., direction and magnitude) of body 140 and providing such measurements as sensor signals that can be transmitted to other devices, such as analysis server 110a. The gyroscope / accelerometer 170f can be positioned and / or adapted to perform such measurements relative to a specific coordinate system of body 140. In various embodiments, the gyroscope / accelerometer 170f can be used with… Figure 1B The other elements depicted are implemented together in a common housing and / or module to ensure a common reference frame or known transformations between reference frames.

[0059] The Global Navigation Satellite System (GNSS) 170h can be implemented as a global positioning satellite receiver and / or another device capable of determining the absolute and / or relative position of itself 140, for example, based on radio signals received from space and / or ground sources, and capable of providing measurements such as sensor signals that can be transmitted to various devices. In some embodiments, the GNSS 170h can be adapted (e.g., using a time series of position measurements) to determine the rate, velocity, and / or yaw rate of itself 140, such as the yaw component of its absolute rate and / or angular rate.

[0060] Temperature sensor 170i can be implemented as a thermistor, an electrical sensor, an electrical thermometer, and / or other device capable of measuring the temperature associated with body 140 and providing such measurement as a sensor signal. Temperature sensor 170i can be configured to measure the ambient temperature associated with body 140, such as cabin or dashboard temperature, for example, which can be used to estimate the temperature of one or more components of body 140.

[0061] The humidity sensor 170j can be implemented as a relative humidity sensor, an electrical sensor, an electrical relative humidity sensor, and / or another device capable of measuring the relative humidity associated with itself 140 and providing such measurement as a sensor signal.

[0062] The steering sensor 170g can be adapted to physically adjust the heading of the self 140 based on one or more control signals and / or user input provided by a logic device such as a controller 170c. The steering sensor 170g may include one or more actuators and control surfaces (e.g., a rudder or other type of steering or adjustment mechanism) of the self 140, and can be adapted to physically adjust the control surfaces to various positive and / or negative steering angles / positions. The steering sensor 170g can also be adapted to sense the current steering angle / position of such steering mechanism and provide this measurement.

[0063] The propulsion system 170k can be implemented as a propeller, turbine, or other thrust-based propulsion system, a mechanical wheeled and / or tracked propulsion system, a wind / sail-based propulsion system, and / or other types of propulsion systems that can be used to power the self-body 140. The propulsion system 170k can also monitor the direction of the self-body 140's power and / or thrust relative to the self-body 140's reference coordinate system. In some embodiments, the propulsion system 170k can be coupled to and / or integrated with the steering sensor 170g.

[0064] Passenger restraint sensor 170l can monitor seatbelt detection and locking / unlocking assemblies, as well as other passenger restraint subsystems. Passenger restraint sensor 170l may include various environmental and / or state sensors, actuators, and / or other devices that facilitate the operation of safety mechanisms associated with the operation of self 140. For example, passenger restraint sensor 170l can be configured to... Figure 1B Other sensors depicted receive motion and / or status data. Passenger restraint sensor 170l can determine whether a safety measurement (e.g., seatbelt) is being used.

[0065] like Figure 1CAs described herein, camera 170m can refer to one or more cameras integrated within body 140, and may include multiple cameras integrated (or modified) into body 140. Camera 170m can be an internally or externally oriented camera of body 140. For example, as... Figure 1C As depicted, the vehicle 140 may include one or more inward-facing cameras 170m-1. These cameras can monitor and collect footage of the passengers of the vehicle 140. The vehicle 140 may also include a front-view camera 170m-2, a camera 170m-3 (e.g., integrated into the door frame), and a rear-view camera 170m-4.

[0066] In some embodiments, the methods and systems discussed herein can operate using only 2D sensors (e.g., 2D cameras), explicitly excluding depth cameras, time-of-flight (ToF) sensors, and other dedicated depth sensing technologies. The AI ​​models and processing pipelines discussed herein can be trained to extract spatial and environmental information solely from monocular or stereo 2D image input, without relying on depth estimation hardware. This ensures compatibility with 2D camera systems that only transmit captured images without any additional depth data, while maintaining robust performance in autonomous navigation and visual data analysis.

[0067] refer to Figure 1B The radar 170n and ultrasonic sensor 170p can be configured to monitor the distance of the self 140 to other objects, such as other vehicles or stationary objects (e.g., trees or garage doors). Figure 1C As depicted, radar 170n and ultrasonic sensor 170p can be integrated into body 140. Body 140 may also include autonomous driving or steering system 170o, which is configured to use data collected from the main navigation body 140 via various sensors (e.g., radar 170n, speed sensor 170d, and / or ultrasonic sensor 170p).

[0068] Therefore, the autonomous driving or steering system 170o can analyze various data collected by one or more sensors described herein to identify driving data. For example, the autonomous driving or steering system 170o can calculate the risk of forward contact based on its own speed 140 and its distance from another vehicle on the road. The autonomous driving or steering system 170o can also determine whether the driver is touching the steering wheel. The autonomous driving or steering system 170o can transmit the analyzed data to various features discussed herein, such as an analysis server.

[0069] The airbag activation sensor 170q can anticipate or detect contact and cause activation or deployment of one or more airbags. The airbag activation sensor 170q can transmit data regarding airbag deployment, including data associated with the event that caused the deployment.

[0070] Return to reference Figure 1A Administrator computing device 120 can represent a computing device operated by a system administrator. Administrator computing device 120 can be configured to display data retrieved or generated by analytics server 110a (e.g., various analytical metrics and risk scores), whereby the system administrator can monitor various models used by analytics server 110a, review feedback, and / or facilitate the training of one or more AI models 110c maintained by analytics server 110a.

[0071] One or more self-entities 140 can be any device configured to navigate various routes, such as vehicle 140a or robot 140b. (See also: [link to relevant information]) Figures 1B to 1C The self 140 discussed herein may include various telemetry sensors. Self 140 may also include a self-computing device 141. Specifically, each self may have its own self-computing device 141. For example, truck 140c may have a self-computing device 141c. For simplicity, self-computing devices are collectively referred to as one or more self-computing devices 141. Self-computing device 141 may control the presentation of content on the infotainment system of self 140, process commands associated with the infotainment system, aggregate sensor data, manage the delivery of data to electronic data sources, receive updates and / or transmit messages. In one configuration, self-computing device 141 communicates with an electronic control unit. In another configuration, self-computing device 141 is an electronic control unit. Self-computing device 141 may include a processor and a non-transitory machine-readable storage medium capable of performing the various tasks and processes described herein. For example, one or more AI models 110c described herein may be stored and executed (or directly accessed) by self-computing device 141. Non-limiting examples of self-computing device 141 may include a vehicle multimedia and / or display system.

[0072] In operation, as depicted in data flow 172, one or more entities 140 may collect image data from their cameras and transmit the image data to a processor (locally located on one or more entities 140) and / or an analysis server 110a. The processor may then execute one or more AI models 110c to predict navigation decisions for one or more entities 140.

[0073] In converter architectures deployed for edge inference, one of the key techniques for improving computational efficiency (in terms of both power consumption and performance) is quantization. Quantization of weights and activations reduces memory footprint on embedded devices and allows such models to utilize low-power integer computational cores, such as the 8-bit (int8) arithmetic units commonly found in embedded accelerometers. However, training a quantized model is typically more complex than training a floating-point model. To facilitate efficient quantization, activation functions are often constrained or clamped to a narrower range more suitable for quantization. For example, multi-head attention modules in quantized converters often contain clamping operations such as ReLU6 or other bounded activation functions to limit the dynamic range of the activation output.

[0074] The computational component of the transformer model is the attention mechanism, where a lookup key matrix multiplication is followed by the application of the softmax function. The results of the lookup key multiplication are typically accumulated in a high-bit-width accumulator (e.g., a 32-bit integer when multiplying by an int8 operand). Traditional methods for calculating the softmax function based on these accumulator outputs involve multiple traversals and heavily rely on floating-point arithmetic, as illustrated below: First pass (maximum extraction) Each input value x from the accumulator is converted from an integer to a floating-point (i2f) and dequantized by multiplying by the quantization scales (scale_q and scale_k) associated with the query projection and key projection. The maximum value across elements is then computed using floating-point comparisons.

[0075] max_x = -inf For element x: { x_fp = i2f(x) x_dequant = x_fp scale_q scale_k max_x = max(max_x, x_dequant) } Second traversal (exponentiation and accumulation) Each element is again converted to floating-point, dequantized, offset relative to the previously calculated maximum value, and exponentially calculated. Then, the exponentially calculated values ​​are summed to calculate the normalized denominator.

[0076] acc_x = 0 For element x: { x_fp = i2f(x) x_dequant = x_fp scale_q scale_k x_neg = x_dequant - max_x x_exp = exp(x_neg) acc_x = acc_x + x_exp } Third traversal (normalization and output) The reciprocal of the accumulated value is calculated and used to normalize each exponentiated value element. The final softmax output is then stored.

[0077] recip_acc_x = 1 / acc_x For element x: { x_fp = i2f(x) x_dequant = x_fp scale_q scale_k x_neg = x_dequant - max_x x_exp = exp(x_neg) res = x_exp recip_acc_x Storage (resources) } Because these traditional methods of calculating softmax rely on multiple floating-point conversions, multiplications, and exponentiation operations, they are inefficient for edge hardware. Even though the optimal approximation of the exponential function typically requires multiple floating-point operations, all of these operations are computationally expensive on embedded processors.

[0078] To address this inefficiency, modifications can be made to the underlying neural network architecture to better align computation with hardware acceleration instructions. Specifically, by forcing the quantization scale used for query projection and key projection to be powers of 2, clamping operations can also be tuned to bounded powers of 2, such as ReLU4, ReLU8, clamp (-4, 4), or clamp (-8, 8). This results in activation values ​​remaining within a quantization-friendly range and allows dequantization to be implemented as a simple shift operation in the integer domain, rather than expensive floating-point multiplication.

[0079] Furthermore, since the goal of softmax might be to convert the logit into a normalized probability score, the modified implementation can utilize power-law operations to the base 2 (e.g., calculating 2...). x Instead of e xThis is better suited for efficient hardware instructions (such as fscale). Since the dequantized inputs can still be represented as integers after scaling, these inputs can be processed directly using power-law techniques with base 2, resulting in a more concise and efficient implementation of the softmax operation for the quantization converter model.

[0080] This optimization results in a fused multi-head attention kernel that significantly reduces the number of instructions required and avoids expensive floating-point operations, making it particularly suitable for real-time inference on edge devices.

[0081] Based on the traditional attention mechanism implemented in the quantizer model, the softmax function after query key multiplication typically uses multiple traversals for computation, each involving floating-point arithmetic. These operations include floating-point conversions, multiplications, additions, and exponential evaluations, which are computationally intensive and energy-inefficient, especially on edge computing platforms where floating-point resources are limited or scarce.

[0082] Figure 2 A flowchart illustrating a method 200 performed in a quantization system according to an embodiment is shown. Method 200 may include steps 210-280. However, other embodiments may include additional or alternative steps or may omit one or more steps entirely. Method 200 is described as being performed by an analysis server (e.g., a computer similar to analysis server 110a). However, one or more steps of method 200 may be performed by a... Figure 1A and Figure 1B The distributed computing system described herein can operate any number of computing devices (e.g., processors of self-component 140 and / or self-component computing device 141). For example, one or more computing devices can execute locally. Figure 2 Some or all of the steps described in the document.

[0083] Now, reference will be made to illustrative embodiments of the base-2 softmax technique, which will be described using specific language. It should be understood that this is not intended to limit the scope of the claims or this disclosure. Any changes and modifications to the features illustrated herein, as well as additional applications of the principles of the subject matter, that will be apparent to those skilled in the art should be considered within the scope of the subject matter disclosed herein.

[0084] The optimized base-2 softmax operation described in this paper exhibits data-agnostic properties, allowing it to be universally applied across model layers without relying on data-specific properties. This allows the optimized softmax to deliver a uniform speedup across training and inference workloads, providing performance improvements regardless of the characteristics of the underlying data. By transforming traditional exponential and division operations into efficient shift, lookup, and one or more normalization steps, Method 200 can provide consistent numerical accuracy across different input distributions. In doing so, executing Method 200 reduces computational latency and power consumption, making the process particularly suitable for deployment across a variety of applications, including but not limited to edge computing platforms and real-time systems (e.g., autonomous navigation systems that need to make rapid decisions based on constantly changing conditions around them).

[0085] The methods and systems discussed in this paper (e.g., method 200) utilize scaling factors raised to powers of 2 to facilitate shift operations instead of floating-point multiplication. Therefore, these techniques not only maintain numerical stability and model fidelity but also support a wide range of model architectures and data types. This data-agnostic approach allows for integration into current processing pipelines, thereby improving computational efficiency and enabling broad application across different technological environments and systems employing converter models.

[0086] At step 210, the analysis server may receive a set of quantized values ​​generated by matrix multiplication of quantized query vectors and quantized key vectors in the attention layer of the transformer model, which corresponds to sensor data associated with the entity. In some embodiments, the analysis server may act as a host compute node that executes the transformer's attention kernel for real-time (or near-real-time) sensing. For example, the analysis server may be a processor that performs and / or communicates with the entity's sensors (e.g., a 2D camera) locally on the entity. In some embodiments, consecutive sensor data frames (e.g., camera images) may be ingested, where the data is first embedded and quantized into 8-bit integer representations. Inside the attention layer, the analysis server / model may perform integer matrix multiplication between the quantized query (Q) vector representing the model's current focus and the quantized key (K) vector encoding contextual information from the same or adjacent frames.

[0087] For example, in one illustrative embodiment, a front-facing camera mounted on the device captures image frames divided into multiple fixed-size regions (e.g., 256 tiles). A lightweight quantization stage converts each region into a short integer vector. An analytics server then computes an integer similarity score between each pair of regions, resulting in a matrix of integer values ​​reflecting the degree of relevance of each region to every other region. This matrix of integer similarity scores constitutes a “quantized value set,” which is then passed to the base-2 softmax process described herein without any intermediate conversion to floating-point format.

[0088] At step 220, the analysis server can identify the maximum value from the set of quantized values ​​using an integer comparison operation. In some embodiments, this occurs after the analysis server iteratively compares pairs of integer values ​​to find a single maximum value in the set. In some embodiments, because these values ​​may remain in fixed-point integer format throughout the operation, the comparison can be performed using a low-latency SIMD “max” instruction without any conversion to floating-point. The result of the reduction can establish a reference level relative to which each subsequent value can be offset, thereby limiting the dynamic range of downstream power-law operations to base 2.

[0089] At step 230, the analysis server can calculate the difference between a quantized value and the maximum value for at least one quantized value in the quantized value set. After obtaining the maximum value in the integer field, the analysis server can iterate over each remaining quantized attention value and perform integer subtraction to identify the difference between the quantized value and the maximum value identified in step 220.

[0090] At step 240, the analysis server can perform a shift operation corresponding to a quantization scale raised to a power of 2 on the difference between the quantized value and the maximum value to generate a shifted value. Next, the analysis server can convert each value to a reference scale (e.g., the difference calculated in step 230) by shifting integers instead of multiplying by a floating-point factor. The shift operation eliminates the expensive floating-point multiplication typically required for dequantization, while keeping the computation completely and intact within a fixed-width integer data path. In some embodiments, to allow for compatibility, the analysis server can choose a scale raised to a power of 2 such that at least one centered fraction, after being right-shifted, falls within a predefined safe range (e.g., –8 to 0) or any other defined range.

[0091] At step 250, the analysis server can perform a base-2 exponentiation operation on the shifted values ​​to generate base-2 exponentiated values. The analysis server can invoke a dedicated base-2 exponentiation protocol that converts each shifted value into its corresponding weight. For example, the base-2 exponentiation operation can use a single hardware native instruction (e.g., the exp2 intrinsic or fscale micro-operation) to convert each shifted integer domain fraction into its corresponding weight. In some embodiments, because the aforementioned quantization and centering steps limit each fraction to a small integer range aligned with a power-of-2 scale, this conversion produces numerically accurate softmax weights while incurring only a single-cycle delay per element.

[0092] At step 260, the analysis server can aggregate the base-2 exponentiation values ​​to generate an aggregated value. After converting each shifted value to its corresponding base-2 weight, the analysis server can invoke a low-latency accumulator that iteratively (or in parallel) aggregates all generated base-2 exponentiation values. In some embodiments, because each weight remains within the fixed-point integer or narrow floating-point range guaranteed by the aforementioned quantization and exponentiation stages, aggregation can be accomplished through a series of word-length-constrained addition operations.

[0093] At step 270, the analysis server can use the aggregated value to normalize the set of quantized values. Once the scalar sum is available, the analysis server can perform the normalization protocol. In a non-limiting example, a power-law value to the base 2 can be divided by the aggregated value to produce a probability that the sum across the set is 1. Because the numerator and denominator reside in the same fixed-point domain established by the aforementioned integer pipeline, the division can be implemented either by calculating the integer reciprocal of a single exp2 / shift pair or by hardware-supported short-word integer division.

[0094] At step 280, the analysis server can use the normalized values ​​to execute the model to output navigation commands to be executed by the vehicle itself. The analysis server can feed softmax-normalized attention weights into the rest of the transformer stack (first applying the weights to the corresponding value vectors to form a context embedding, and then propagating the embedding through subsequent layers of the model) to generate inference outputs specifying navigation commands for the vehicle (e.g., steering angle, lane change command, or acceleration setpoint). In some embodiments, attention calculations, including base-2 softmax normalization, can be performed using low-precision integer arithmetic, and downstream layers can consume the resulting probabilities without incurring format conversion, thereby maintaining the model's low latency and low-power profile while delivering time-critical control commands suitable for onboard execution by the vehicle's motion planning subsystem.

[0095] Although the embodiments discussed with respect to method 200 include navigation decisions, method 200 is not limited to application in autonomous navigation or specific types of navigation data. Therefore, the methods discussed herein can be applied to all types of data.

[0096] In some embodiments, normalized attention weights generated by optimized base-2 softmax are propagated downstream of the transformer to generate a fusion vector (e.g., from consecutive camera frames). The model can then calculate a Time-To-Location (TTL) metric for each dynamically tracked object (such as a vehicle, cyclist, or pedestrian) within its field of view based on this fusion vector. When the model determines that the TTL of any object is below a safe threshold, the analytics server generates navigation instructions that command avoidance steering adjustments, braking pulses, target rate adjustments, lane change initiation, longitudinal acceleration or deceleration setpoints, or coordinated combinations thereof, enabling the model to mitigate or avoid impending contact while maintaining a low-power edge-deployable profile provided by the disclosed attention computation techniques.

[0097] As used in this article, location time or contact time indicates the estimated time it takes for a self to arrive at a location (whether that location is a specified location or a defined location or whether that location indicates an object, such as a wall or another self).

[0098] As used herein, a model can include any computer model, including artificial intelligence models, whether or not a transformer is used. Additionally or alternatively, the model may not be categorized as an artificial intelligence or machine learning model.

[0099] In some embodiments, the analysis server can provide (e.g., method 200) a hardware-optimized softmax sequence that minimizes floating-point operations while leveraging power-of-2 scaling and efficient integer field arithmetic. The proposed sequence can be implemented as follows: First iteration (maximum integer value determined) The maximum value is determined from the quantization accumulator value using integer operations.

[0100] max_x_int = -int_min For each x in the input element: { max_x_int = max(max_x_int, x) } Second traversal (preparation and accumulation for exponentiation): Each input is shifted and exponentially operated on using hardware-efficient base-2 functions. These values ​​are scaled using displacements derived from the logarithmic scale of the query projection and key projection.

[0101] acc_x = 0 For each x in the input element: { x_neg = x – max_x_int x_neg = x_neg>>(log2(scale_q) + log2(scale_k)) x_exp = fscale(x_neg) acc_x = acc_x + x_exp } Third traversal (normalization and output calculation) Each exponentiation value is normalized by the reciprocal of the sum and stored as the final softmax output.

[0102] recip_acc_x = 1 / acc_x For each x in the input element: { x_neg = x – max_x_int x_neg = x_neg>>(log2(scale_q) + log2(scale_k)) x_exp = fscale(x_neg) res = x_exp recip_acc_x Storage (resources) } This optimized approach avoids dequantization and floating-point multiplication by using a base-2 quantization scale. The shift operations used to scale the values ​​are natively supported and hardware-accelerated on common processor architectures, including the Advanced Reduced Instruction Set Computer Machine (ARM) and x86 platforms.

[0103] A comparative analysis of instruction-level operand counts between the traditional (“native”) implementation and the proposed hardware-friendly implementation highlights the efficiency benefits. The following table illustrates... Figure 3 The operations described in the document for each softmax traversal.

[0104] As described, the proposed method reduces the number of floating-point operations by replacing multiplication and exponentiation with hardware-optimized fscale operations that compute integer addition, shifting, and powers of 2. This results in an efficient softmax implementation well-suited for low-power real-time inference on embedded and edge devices. Furthermore, because the technique uses a quantization scale constrained to powers of 2, it simplifies dequantization and improves compatibility with fixed-point processing pipelines. Thus, this fused attention kernel achieves the same probability normalization as traditional softmax while significantly reducing computational overhead—making it particularly advantageous for deployment in resource-constrained environments increasingly employing transformer models.

[0105] The methods and systems described in this paper allow for the practical execution of transformer models on edge devices embedded within themselves (e.g., autonomous vehicles and robots). Using the methods and systems discussed herein, performance bottlenecks in attention computation can be reduced (or sometimes eliminated), thereby aligning model execution with hardware capabilities and enabling real-time autonomous navigation in constrained environments.

[0106] In one embodiment, the methods and systems described herein (e.g., optimized base-2 softmax techniques) can be implemented within an autonomous navigation system operating on a self-driving vehicle. The autonomous navigation system may include a transformer-based neural network model configured to process real-time sensor inputs, such as multi-view camera images, Global Positioning System (GPS) data, and / or high-resolution maps. The model can be designed to perform multimodal perception tasks, including lane detection, obstacle identification, and dynamic object tracking. As part of its attention mechanism, the model can perform matrix multiplication between a quantized query tensor and a quantized key tensor, followed by a softmax operation to normalize the attention weights. The autonomous navigation system runs inference on an embedded processor within the vehicle, such as an ARM-based automotive-grade System-on-Chip (SoC) or a low-power Graphics Processing Unit (GPU) accelerator, where floating-point resources are limited and efficiency is critical for maintaining real-time responsiveness.

[0107] To optimize inference performance, the softmax operation is replaced with the disclosed hardware-efficient attention mechanism. Instead of using floating-point exponentiation and arbitrary scaling, the model employs power-of-2 quantization for both query and key projections, achieving inverse quantization through integer bit shifts. The exponentiation step is computed using hardware-accelerated base-2 functions (such as the fscale instruction), eliminating the need for floating-point exp(x) calls.

[0108] This implementation reduces computational load, memory bandwidth, and latency, enabling the converter model to process high-dimensional sensor data within real-time operational constraints. Therefore, even in complex and dynamic environments such as urban intersections, highway junctions, and pedestrian-intensive areas, autonomous vehicles can make timely and accurate decisions for safe navigation.

[0109] In a non-restrictive example, the analytics server can retrieve a vector of quantized values ​​consisting of five quantized integer values ​​and a quantization scale raised to the power of 2: x = [3, 7, 5, 2, 6]. The quantization scales for the query projection and the key projection can be set to 2, such that scale_q = 2 and scale_k = 2. Since the scaling factor is a power of 2, their base-2 logarithms can be efficiently computed as integers. In this case, log2(scale_q) + log2(scale_k) = log2(2) + log2(2) = 1 + 1 = 2, meaning that dequantization can be achieved by simply right-shifting by 2 bits (equivalent to dividing the value by 4 in hardware).

[0110] In the first iteration of the softmax algorithm, the analysis server uses integer arithmetic to determine the maximum value of the quantized input vector. This value serves as a reference to stabilize subsequent exponentiation steps and prevent numerical overflow or underflow. In the input values ​​[3, 7, 5, 2, 6], the maximum value is 7 (max_x_int = 7).

[0111] During the second traversal, the method uses power-to-base 2 to calculate the exponentiation value. First, each input value is subtracted from the maximum value to produce a centered value, which is then right-shifted by 2 bits. The calculation steps for each element are as follows: For x = 3, the difference from the maximum value is -4, which shifts to the right to -1, resulting in 2^(-1) = 0.5.

[0112] For x = 7, the difference is 0, the shift is still 0, therefore 2^0 = 1.0.

[0113] For x = 5, the difference is -2, which is shifted to -1, resulting in 0.5.

[0114] For x = 2, the difference is -5, which is shifted to -2, resulting in 2^(-2) = 0.25.

[0115] For x = 6, the difference is -1, which is shifted to -1, resulting in 0.5.

[0116] The sum of these exponentiation values ​​is 0.5 + 1.0 + 0.5 + 0.25 + 0.5 = 2.75. This value is used as the normalization factor in the final traversal.

[0117] In the third iteration, each exponentiation value is divided by the cumulative sum of 2.75 to obtain a normalized probability that sums to 1. The final normalized value is: For x = 3: 0.5 / 2.75 ≈ 0.1818 For x = 7: 1.0 / 2.75 ≈ 0.3636 For x = 5: 0.5 / 2.75 ≈ 0.1818 For x = 2: 0.25 / 2.75 ≈ 0.0909 For x = 6: 0.5 / 2.75 ≈ 0.1818 Therefore, the final output vector of the softmax class is approximately [0.1818, 0.3636, 0.1818, 0.0909, 0.1818].

[0118] This example illustrates how the proposed base-2 quant softMax produces normalized attention weights similar to the traditional softMax function, but by using hardware-friendly operations such as integer subtraction, bitwise shifting, and powers of 2. This approach improves computational efficiency by eliminating floating-point multiplication and expensive exponential functions.

[0119] In another non-limiting example, the methods and systems discussed herein can be used to autonomously navigate a self. In one example, the self approaches a complex four-way intersection where pedestrians, cyclists, and lateral traffic must be tracked in real time. A transformer-based perception and planning stack running on the self processor ingests eight synchronized camera streams, labels each frame as a spatial patch, and feeds the resulting int8 query tensors, key tensors, and value tensors to its multi-head attention module. Utilizing the optimized base-2 softmax disclosed herein, the attention layer performs Q·K integer matrix multiplication, identifies the maximum logarithm, applies a power-of-2 shift and a single power-of-2 operation to obtain normalized attention weights, and produces contextual embeddings without resorting to floating-point exponentiation or multiplication. Because the entire softmax computation is performed using the methods discussed herein, the model can ingest data faster (and use less computational power) than conventional implementations. Consequently, the downstream motion planning network receives the normalized attention output more quickly. The reduced computational load allows the processor to dedicate more energy margin to navigation decisions, enabling the system to accurately predict when a cyclist will enter its lane from the right and issue predictive deceleration and steering adjustments that allow for safe and comfortable lane changes when navigating intersections.

[0120] The various illustrative logic blocks, modules, circuits, and algorithm steps described in conjunction with the embodiments disclosed herein can be implemented as electronic hardware, computer software, or a combination of both. To clearly illustrate this interchangeability between hardware and software, various illustrative components, blocks, modules, circuits, and steps have been generally described above in terms of their functionality. Whether this functionality is implemented as hardware or software depends on the specific application and the design constraints imposed on the system as a whole. Those skilled in the art can implement the described functionality in different ways for each specific application, but such implementation decisions should not be construed as departing from the scope of the invention.

[0121] Implementations in computer software can be implemented as software, firmware, middleware, microcode, hardware description languages, or any combination thereof. Code segments or machine-executable instructions can represent any combination of procedures, functions, subroutines, programs, routines, subroutines, modules, software packages, classes, or instructions, data structures, or program statements. Code segments can be coupled to other code segments or hardware circuitry by passing or receiving information, data, arguments, attributes, or memory contents. Information, arguments, attributes, data, etc., can be passed, forwarded, or transmitted by any suitable means, including memory sharing, message passing, token passing, network transmission, etc.

[0122] The actual software code or dedicated control hardware used to implement these systems and methods does not limit the invention. Therefore, since the operation and behavior of the systems and methods are described without reference to specific software code, it should be understood that the software and control hardware can be designed to implement the systems and methods based on the descriptions herein.

[0123] When implemented in software, functionality can be stored as one or more instructions or code on a non-transitory computer-readable or processor-readable storage medium. The steps of the methods or algorithms disclosed herein can be embodied in a processor-executable software module that may reside on a computer-readable or processor-readable storage medium. A non-transitory computer-readable or processor-readable medium includes both computer storage media and tangible storage media that facilitate the transfer of a computer program from one place to another. A non-transitory processor-readable storage medium can be any available medium accessible to a computer. By way of example, and not limitation, such a non-transitory processor-readable medium may include random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), compact disc read-only memory (CD-ROM) or other optical disc storage devices, magnetic disk storage devices or other magnetic storage devices, or any other tangible storage medium that can be used to store desired program code in the form of instructions or data structures and is accessible to a computer or processor. As used herein, “disc” and “plate” include compact discs (CDs), laser discs, optical discs, digital versatile discs (DVDs), floppy disks, and Blu-ray discs, where discs typically reproduce data magnetically, while plates reproduce data optically using lasers. Combinations of the above should also be included within the scope of computer-readable media. Additionally, operations of methods or algorithms may reside as one or any combination or set of code and / or instructions on a non-transitory processor-readable medium and / or a computer-readable medium, which may be incorporated into a computer program product.

[0124] The foregoing description of the disclosed embodiments is intended to enable any person skilled in the art to make or use the invention. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be applied to other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not intended to be limited to the embodiments shown herein, but is to be given the widest scope consistent with the appended claims and the principles and novel features disclosed herein.

[0125] While various aspects and embodiments have been disclosed, other aspects and embodiments are contemplated. The disclosed aspects and embodiments are for illustrative purposes and are not intended to be limiting, wherein the true scope and spirit are indicated by the appended claims.

Claims

1. A method for navigating oneself by acquiring sensor data, the method comprising: One or more processors receive a set of quantized values, which is generated by matrix multiplication of the quantization query vector and the quantization key vector in the attention layer of the converter model, and the set of quantized values ​​corresponds to sensor data associated with the self. The one or more processors use integer comparison operations to identify the maximum value from the set of quantized values; For at least one quantized value in the set of quantized values, the one or more processors calculate the difference between the quantized value and the maximum value; The one or more processors perform a shift operation on the difference between the quantized value and the maximum value, corresponding to a quantization scale raised to a power of 2, to generate a shifted value; The one or more processors perform a power-to-2 operation on the shifted value to generate a power-to-2 value; The one or more processors aggregate the base-2 exponentiation values ​​to generate an aggregated value; as well as The one or more processors use the aggregated value execution model to output navigation instructions to be executed by the processor itself.

2. The method of claim 1, wherein executing the model further comprises: The one or more processors estimate the localization time of at least one dynamic object detected in a camera image associated with the self; And when the estimated positioning time meets a threshold, the navigation command guides evasive steering or braking maneuvers.

3. The method of claim 1, wherein the quantization scale raised to the power of 2 is selected such that at least one shifted value falls within a defined range.

4. The method according to claim 1, wherein the navigation instructions include self-control commands, the self-control commands including: Steering angle adjustment, lane change start, longitudinal acceleration or deceleration, and target speed set point are at least one of these.

5. The method of claim 1, wherein the sensor data corresponds to two-dimensional image data captured by at least one camera associated with the self.

6. The method according to claim 1, further comprising: The one or more processors use the aggregated value to normalize the quantized value set; as well as The one or more processors use an aggregated normalized value execution model to output navigation instructions to be executed by the processor itself.

7. The method of claim 1, wherein normalization comprises: Each of the base-2 exponentiation values ​​is divided by the aggregate value to generate a probability distribution over the set of quantized values.

8. A computer system for navigating oneself by ingesting sensor data, the computer system comprising a computer-readable medium having a non-transitory instruction set, which, when executed by at least one processor, causes the at least one processor to: Receive a set of quantized values, which is generated by matrix multiplication of the quantization query vector and the quantization key vector in the attention layer of the converter model, and the set of quantized values ​​corresponds to the sensor data associated with the self. The maximum value is identified from the set of quantized values ​​using integer comparison operations; For at least one quantized value in the set of quantized values, calculate the difference between the quantized value and the maximum value; Perform a shift operation corresponding to a quantization scale raised to a power of 2 on the difference between the quantized value and the maximum value to generate a shifted value; Perform a power-to-2 operation on the shifted value to generate a power-to-2 value; Aggregate the values ​​raised to the power of 2 to generate an aggregated value; as well as The aggregated value is used to execute the model to output navigation instructions to be executed by the entity.

9. The computer system of claim 8, wherein executing the model further comprises: Estimate the localization time of at least one dynamic object detected in a camera image associated with the self; And when the estimated positioning time meets a threshold, the navigation command guides evasive steering or braking maneuvers.

10. The computer system of claim 8, wherein the quantization scale raised to the power of 2 is selected such that at least one shifted value falls within a defined range.

11. The computer system of claim 8, wherein the navigation instructions include self-control commands, the self-control commands comprising: Steering angle adjustment, lane change start, longitudinal acceleration or deceleration, and target speed set point are at least one of these.

12. The computer system of claim 8, wherein the sensor data corresponds to two-dimensional image data captured by at least one camera associated with the self.

13. The computer system of claim 8, wherein the instructions further cause the at least one processor to: The aggregated values ​​are used to normalize the quantized value set; and The model is executed using aggregated normalized values ​​to output navigation instructions to be executed by the entity itself.

14. The computer system of claim 8, wherein normalization comprises: Each of the base-2 exponentiation values ​​is divided by the aggregate value to generate a probability distribution over the set of quantized values.

15. A computer system for navigating oneself by acquiring sensor data, the computer system comprising at least one processor configured to: Receive a set of quantized values, which is generated by matrix multiplication of the quantization query vector and the quantization key vector in the attention layer of the converter model, and the set of quantized values ​​corresponds to the sensor data associated with the self. The maximum value is identified from the set of quantized values ​​using integer comparison operations; For at least one quantized value in the set of quantized values, calculate the difference between the quantized value and the maximum value; Perform a shift operation corresponding to a quantization scale raised to a power of 2 on the difference between the quantized value and the maximum value to generate a shifted value; Perform a power-to-2 operation on the shifted value to generate a power-to-2 value; Aggregate the values ​​raised to the power of 2 to generate an aggregated value; as well as The aggregated value is used to execute the model to output navigation instructions to be executed by the entity.

16. The computer system of claim 15, wherein executing the model further comprises: Estimate the localization time of at least one dynamic object detected in a camera image associated with the self; And when the estimated positioning time meets a threshold, the navigation command guides evasive steering or braking maneuvers.

17. The computer system of claim 15, wherein the quantization scale raised to the power of 2 is selected such that at least one shifted value falls within a defined range.

18. The computer system of claim 15, wherein the navigation instructions include self-control commands, the self-control commands comprising: Steering angle adjustment, lane change start, longitudinal acceleration or deceleration, and target speed set point are at least one of these.

19. The computer system of claim 15, wherein the sensor data corresponds to two-dimensional image data captured by at least one camera associated with the self.

20. The computer system of claim 15, wherein the at least one processor is further configured to: The aggregated values ​​are used to normalize the quantized value set; and The model is executed using aggregated normalized values ​​to output navigation instructions to be executed by the entity itself.