Near memory computing chip
Through the hybrid bonding of memory chips and logic chips and Bank-level parallel memory access method, combined with on-chip data cache and vector processing unit optimization, the flexibility, scalability and energy efficiency problems in large-scale model applications are solved, and efficient data access and computing performance improvements are achieved.
Patent Information
- Application Number
- CN202510660323.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-21
- Publication Date
- 2025-07-22
AI Technical Summary
The existing integrated storage and near-access computing technologies have problems such as insufficient flexibility and scalability, limited bandwidth at the system level, limited computing power density, and still need to be improved in large-scale application scenarios, and it is difficult to fully consider the strict requirements of computing power, storage bandwidth and energy consumption.
The hybrid bonding method of memory chip and logic chip is adopted, combined with Bank-level parallel memory access and hybrid bonding integration, to achieve high bandwidth and high parallelism data access, and optimize computing efficiency through on-chip data cache and vector processing units, combined with advanced cache replacement algorithms and distributed near-memory computing architecture dynamic management cache, supporting flexible storage and computing mode switching.
It significantly improves data access efficiency, reduces the complexity and cost of hardware upgrades, improves overall computing performance and storage efficiency, meets the needs of large-scale model inference, and has a wide range of application prospects.
Smart Images

Figure CN120353755A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of chip technology, and in particular to a near-memory computing chip. Background Art
[0002] In the past few decades, driven by Moore's Law, processor performance has been significantly improved. However, traditional computers use the von Neumann architecture, and data processing and storage are physically separated, resulting in huge computing delays and power consumption during data transmission. In addition, although the traditional logic gate-based computing method is versatile and robust, it has low computational efficiency when performing high-dimensional operations such as tensor multiplication, addition, and nonlinear operations, and requires a lot of hardware resources and time. In the current field of computing technology, although storage-computing integration and near-storage computing technologies have alleviated the drawbacks of traditional computing architectures to a certain extent, the existing single technical solutions still have many shortcomings when dealing with large model application scenarios.
[0003] In-memory computing embeds computing units in storage units, showing significant computing density and energy efficiency advantages at the macro unit design level. However, in the actual application of in-memory computing technology, there are the following problems:
[0004] Insufficient flexibility and scalability: In-memory computing uses a design that tightly couples computing and storage units. This highly integrated architecture leads to insufficient flexibility and scalability of the system. In large-model applications, model size and computing requirements are constantly changing with the development of algorithms. Chips based on in-memory computing usually only support fixed computing operations and data paths, and it is difficult to flexibly adapt to different types of computing tasks. At the same time, existing in-memory computing implementation solutions lack good hardware scalability and are difficult to achieve large-scale distributed deployment, limiting their widespread application in large-model deployment scenarios.
[0005] Limited bandwidth at the system level: The scale of large models often far exceeds the size of hardware macro units, making it difficult to store weights inside the array. In order to meet the needs of model operations, the weights inside the array need to be updated frequently. When the in-memory computing chip needs to exchange data with external devices, the lack of efficient external data transmission and collaboration mechanisms leads to low data transmission efficiency. This not only increases the complexity of data transmission, but also greatly weakens the performance and energy efficiency of hardware at the system level. During the deployment of large models, this problem will lead to a significant decrease in computing efficiency and a significant increase in energy consumption.
[0006] Near-memory computing alleviates the problem of limited system performance and energy efficiency caused by data transmission by physically shortening the distance between storage and computing components. However, there are the following problems in the actual application of in-memory computing technology:
[0007] Limited computing power density: Current research on in-memory computing mainly focuses on improving the bandwidth provided by storage chips and reducing the power consumption of data transmission. Regarding how to fully utilize the provided bandwidth and how to efficiently complete the calculation of the transmitted data, the current research is not yet mature. Considering costs, the storage chip part needs to ensure sufficient storage density to meet the storage requirements of large models. Therefore, the high-computing-power-density computing circuit adapted to it has become a new design challenge. With the increase in data volume and computational complexity, how to improve the computing power density of in-memory computing chips has become a bottleneck faced by current in-memory computing chips.
[0008] The overall energy efficiency still needs to be improved: Although in-memory computing shortens the data transmission path and reduces part of the power consumption, in the actual task deployment process, due to a large amount of power consumption overhead in both the storage and computing parts, it faces severe power supply, heat dissipation, and maintenance cost problems. There is an urgent need for high-energy-efficiency computing solutions to solve the obstacle of high power consumption overhead to the application deployment of large models.
[0009] In summary, certain research results have been achieved in current memory-computation integration and in-memory computing technologies at home and abroad. However, existing research and applications mostly focus on single technologies, and there is still a blank in the systematic solution that organically combines the two and is specifically designed for large model applications. Large model applications have strict requirements for computing power, storage bandwidth, and energy consumption. Single technologies are difficult to comprehensively consider, unable to fully utilize the synergistic advantages of the two technologies, and difficult to effectively support the efficient operation of large model task deployment. Summary of the Invention
[0010] To address the above problems, the present invention provides an in-memory computing chip, including a storage chip and a logic chip. The storage chip and the logic chip are connected by a hybrid bonding method. The storage chip includes a plurality of storage module groups, and each storage module group includes a plurality of storage modules; the logic chip includes a plurality of memory controller groups corresponding to the storage module groups. Each memory controller group is provided with a plurality of memory controllers corresponding to the plurality of storage modules within the storage module group, and a plurality of NPU Banks corresponding to the plurality of memory controller groups.
[0011] Optionally, the plurality of memory controllers within each memory controller group are connected in one-to-one correspondence with the plurality of memory controllers within other memory controller groups.
[0012] For the near-memory computing chip provided by this application, considering the need to improve the data access bandwidth of the storage chip, a parallel memory access method based on the Bank hierarchy is adopted for the interconnection between the storage chip and the logic chip. Since each storage module corresponds to a Memory Controller located in the logic chip. During large model inference, these Memory Controllers can access the corresponding storage modules in parallel, greatly improving the data access efficiency. In addition, the integration method of Hybrid Bonding has significantly increased the wiring density and distance between the two chips compared to traditional two-dimensional packaging, thus achieving an order-of-magnitude increase in the memory access bandwidth of the logic chip.
[0013] The near-memory computing chip provided by this application has a storage mode and a computing mode.
[0014] In the storage mode, the Memory Controller is configured to be responsible for data refreshing and execution of Error Correction Code (ECC) in addition to performing data read and write operations, to ensure the integrity of the data and the correctness of the function.
[0015] In this architecture, the access behavior of the storage chip is exactly the same as that of a standard memory. This design enables the near-memory computing chip to directly replace the DRAM memory in the current near-memory computing chip without additional architecture modification, significantly reducing the complexity and cost of hardware upgrade.
[0016] In the computing mode, each NPU Bank of the logic chip can drive multiple memory controllers simultaneously to achieve parallel access to multiple storage modules; at the same time, since each NPU Bank can only locally access a fixed number of corresponding storage modules, to reduce the latency and overhead caused by global data access. This design combines high bandwidth, high parallelism, and flexible switching between storage and computing modes, not only meeting the requirements of large model inference, but also significantly improving the overall computing performance and storage efficiency through innovative architecture design, and has broad application prospects.
[0017] Optionally, the NPU Bank includes:
[0018] A storage unit for data interaction with the corresponding storage module;
[0019] An on-chip data cache for temporarily storing intermediate data generated during the execution of computing tasks;
[0020] A vector processing unit for assisting in completing vector computing tasks; and
[0021] An SRAM in-memory computing core for executing computing tasks.
[0022] The storage unit LSU is used to interact with the corresponding storage module. During the calculation process, the storage unit LSU accurately schedules the reading and writing of data to ensure that the data in the storage module is transmitted to the logic chip in a timely and accurate manner.
[0023] The on-chip data cache SPAD Mem. is used to temporarily store the intermediate data generated during the execution of the calculation task. Since the access speed of the on-chip data cache is much higher than that of the storage chip, temporarily storing the intermediate data through it can effectively reduce the number of accesses of the logic chip to the storage chip and reduce the data transmission power consumption.
[0024] Preferably, in this embodiment, by adopting an advanced cache replacement algorithm, the on-chip data cache SPAD Mem. is configured to be able to dynamically manage the cache space to ensure that the hot data is always stored in the cache, greatly improving the data access efficiency and further optimizing the overall performance of the system.
[0025] The vector processing unit Vector PU is used to assist in completing the vector calculation task. The vector processing unit Vector PU closely cooperates with the SRAM CIM Core based on the in-memory computing technology. According to the calculation requirements of the SRAM CIM Core, it provides the necessary vector data support to jointly complete complex calculation tasks.
[0026] Optionally, the near-memory computing chip further includes an on-chip network connected to the multiple NPU Banks to realize data interaction between the multiple NPU Banks.
[0027] Optionally, the on-chip network includes multiple computing cores. Each computing core is correspondingly provided with a router, and the routers are connected to form a mesh topology structure. Each NPU Bank is provided with a network interface for connecting to the mesh topology structure.
[0028] Optionally, each computing core includes a network interface unit for connecting to the corresponding router.
[0029] Optionally, the router is configured to be able to detect the traffic of its link. In response to the traffic of the link reaching a preset value, the router can automatically adjust the data sending rate.
[0030] Optionally, the computing core can feedback the available cache space to the computing core upstream of the data sending link, and the computing core upstream of the sending link can adjust the data sending amount in response to the feedback.
[0031] Optionally, the on-chip network follows the AXI4 communication protocol.
[0032] In summary, for the near-memory computing chip provided in this application, a parallel memory access method based on the Bank hierarchy is adopted for interconnection between the memory chip and the logic chip. Since each memory module corresponds to a Memory Controller located in the logic chip. During large model inference, these Memory Controllers can access the corresponding memory modules in parallel, greatly improving the data access efficiency. In addition, the integration method of Hybrid Bonding has significantly increased the wire density and distance between the two chips compared with traditional two-dimensional packaging, thus realizing an order-of-magnitude increase in the memory access bandwidth of the logic chip. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] Figure 1 FIG. is a schematic structural diagram of a near-memory computing chip provided by an embodiment of the present invention;
[0034] Figure 2 FIG. is a schematic structural diagram of a near-memory computing chip provided by an embodiment of the present invention;
[0035] Figure 3 FIG. is a schematic diagram of the on-chip network communication protocol of a near-memory computing chip provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0036] Next, the technical solutions in the embodiments of the present application will be clearly and completely described with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0037] This embodiment provides a near-memory computing chip, including a DRAM Die memory chip and a Logic Die logic chip. The DRAM Die memory chip and the Logic Die logic chip are connected in a hybrid bonding manner. Specifically, the DRAM Die memory chip includes a plurality of memory module groups, and each memory module group includes a plurality of DRAM Bank memory modules; the Logic Die logic chip includes a plurality of memory controller groups corresponding to the memory module groups. Each memory controller group is provided with a plurality of Memory Controllers corresponding to a plurality of memory module groups, and a plurality of NPU Banks corresponding to a plurality of memory controller groups; wherein, the plurality of Memory Controllers in each memory controller group are connected in one-to-one correspondence with the plurality of Memory Controllers in other memory controller groups.
[0038] In this architecture, in order to further improve the data access bandwidth of the DRAM Die of the storage chip, a parallel access method based on the Bank level is adopted for interconnection between the DRAM Die of the storage chip and the Logic Die of the logic chip. Since each memory module corresponds to a Memory Controller located in the Logic Die of the logic chip. During large model inference, these Memory Controllers can access the corresponding memory modules in parallel, greatly improving the data access efficiency. In addition, the integration method of Hybrid Bonding has significantly improved the wire density and distance between the two chips compared with traditional two-dimensional packaging, thus achieving an order-of-magnitude increase in the memory access bandwidth of the logic chip.
[0039] The in-memory computing chip provided in this embodiment has a storage mode and a computing mode.
[0040] In the storage mode, the Memory Controller is configured to be responsible for data refresh and execution of Error Correction Code (ECC) in addition to performing data read and write operations, so as to ensure the integrity of data and the correctness of functions.
[0041] In this architecture, the access behavior of the chip is completely consistent with that of a standard memory. This design enables the in-memory computing chip to directly replace the DRAM memory in the current in-memory computing chip without additional architectural modifications, significantly reducing the complexity and cost of hardware upgrades.
[0042] In the computing mode, each NPU Bank of the Logic Die can drive multiple memory controller groups simultaneously to achieve parallel access to multiple memory modules; at the same time, since each NPU Bank can only locally access a fixed number of corresponding memory modules to reduce the latency and overhead caused by global data access. This design combines high bandwidth, high parallelism, and flexible switching between storage and computing modes, not only meeting the requirements of large model inference, but also significantly improving the overall computing performance and storage efficiency through innovative architectural design, and has broad application prospects.
[0043] Optionally, the in-memory computing chip further includes a STANDARD MEMORY INTERFACE, which is connected to each Memory Controller and can communicate with an external system / device.
[0044] Optionally, each DRAM Bank of the storage chip DRAM Die includes a memory array Memory Array and peripheral circuits Peripheral Circuits; among them, the peripheral circuits Peripheral Circuits are auxiliary circuits used to support the normal operation of the memory array, and may include circuits such as a row decoder, a column decoder, a sense amplifier, a write driver, a timing control circuit, a precharge circuit, a redundancy and repair circuit, etc.
[0045] Optionally, each NPU Bank includes a load store unit LSU, a on-chip data cache SPAD Mem., a vector processing unit Vector PU, and a SRAM in-memory computing core SRAM CIM Core.
[0046] The load store unit LSU is connected to the corresponding memory controller group and is used to interact with the corresponding storage module. During the calculation process, the load store unit LSU accurately schedules the reading and writing of data to ensure that the data in the storage module is transmitted to the logic chip Logic Die in a timely and accurate manner.
[0047] The on-chip data cache SPAD Mem. is used to temporarily store the intermediate data generated during the execution of the calculation task. Since the access speed of the on-chip data cache is much higher than that of the storage chip DRAM Die, temporarily storing the intermediate data through it can effectively reduce the number of accesses to the storage chip DRAM Die and reduce the data transmission power consumption.
[0048] Preferably, in this embodiment, by adopting an advanced cache replacement algorithm, the on-chip data cache SPAD Mem. is configured to be able to dynamically manage the cache space to ensure that the hot data is always stored in the cache, greatly improving the data access efficiency and further optimizing the overall performance of the system.
[0049] Based on the pre-determined characteristics of the data stream during AI calculation, the cache replacement algorithm provided in this embodiment combines the weight stationary characteristic native to in-memory computing and the output stationary data stream design adapted to the distributed near-memory computing architecture to dynamically manage the cache.
[0050] Specifically, in terms of in-memory computing: Utilize the weight stationary characteristic and adopt a scheme to reduce weight replacement. In AI calculation, the weight data is usually relatively fixed. In this way, the data transfer overhead can be reduced. Because in traditional computing, the transfer of weight data between memory and computing units consumes a lot of time and energy, while in-memory computing integrates the computing function into the storage unit and directly calculates the weight data in the storage unit, avoiding frequent weight data transfer and improving the computing efficiency.
[0051] In terms of the distributed near-memory computing architecture: Considering the on-chip storage resource overhead and the on-chip network hardware overhead, an output stationary data stream design is adopted. In a distributed computing system, multiple computing nodes work together. If other data stream methods are used, intermediate results may need to be frequently transferred between the logic chip and the storage chip. However, the output stationary data stream design enables intermediate results not to be frequently transferred between different chips, but to be processed and stored on-chip as much as possible. In this way, hot data (i.e., data that is frequently accessed and processed, such as intermediate results, frequently used weights, etc.) can always be stored in the cache. When these data need to be accessed again, they can be directly obtained from the cache, greatly improving the data access efficiency and thus optimizing the overall performance of the system.
[0052] In this embodiment, an algorithm for dynamically managing the cache is combined with the weight stationary characteristic of in-memory computing and the output stationary data stream design adapted to the distributed near-memory computing architecture to adapt to the distributed near-memory computing architecture and dynamically manage the cache to improve AI computing performance.
[0053] The Vector Processing Unit (Vector PU) is used to assist in completing vector computing tasks. The Vector PU closely cooperates with the SRAM CIM Core based on in-memory computing technology, and provides the necessary vector data support according to the computing requirements of the SRAM CIM Core to jointly complete complex computing tasks.
[0054] The SRAM CIM Core is implemented based on in-memory computing technology and is configured for matrix-vector and matrix-matrix multiplication computing operations that are abundant in neural networks. In the SRAM CIM Core, the storage unit not only undertakes the data storage function, but also directly participates in the computing process, avoiding the frequent transmission of data between the storage and computing units in the traditional architecture, greatly reducing the data transmission delay and power consumption. By designing a dedicated computing circuit in the storage unit, the SRAM CIM Core can complete multiple multiplication and accumulation operations within a single clock cycle, achieving high computing power density computing. At the same time, the SRAM CIM Core optimizes the computing algorithm according to the computing characteristics of large models to improve the computing energy efficiency, providing a computing solution with high computing power density and high computing energy efficiency for large models.
[0055] In the SRAM CIM Core, the algorithm optimization is mainly reflected in two key aspects to give full play to the hardware advantages of in-memory computing and the near-memory computing architecture:
[0056] Support data quantization: During the calculation process of large models, data is usually stored and calculated in high-precision form, which consumes a large amount of storage resources and computing resources. By supporting data quantization technology, high-precision data is converted into low-precision data. For example, by converting 32-bit floating-point activation values into 16-bit floating-point numbers and quantizing weights into 4-bit fixed-point numbers, the computing efficiency can be improved while ensuring the quality of calculation results. Low-precision data occupies less storage space, can reduce the amount of data transfer during calculation, and in the in-memory computing and near-memory computing architectures, the computing operations required to process low-precision data are also correspondingly reduced, enabling more efficient use of hardware resources, enhancing the computing speed, and reducing energy consumption.
[0057] Optimize the computing data flow: Reasonably plan the flow path and method of data in the in-memory computing and near-memory computing architectures. For example, according to the storage structure of the hardware and the distribution of computing units, design the most suitable data reading, processing, and storage order to avoid redundant data transfer and ineffective waiting. By optimizing the computing data flow, data can flow more smoothly between storage units and computing units, making full use of the characteristics of in-memory computing that directly performs calculations within storage units and the advantages of near-memory computing that reduce long-distance data transmission, improving the parallelism and efficiency of computing, and thus achieving higher computing power density and computing energy efficiency.
[0058] In this embodiment, in the SRAM CIM Core, the two technologies of data quantization and optimization of computing data flow are combined, and for the characteristics of large model calculations, optimizations are made to adapt to the in-memory computing and near-memory computing architectures, forming a set of solutions specifically for large models to provide high computing power density and high computing energy efficiency.
[0059] Optionally, as Figure 1 and Figure 2 shown, the near-memory computing chip further includes a Network-on-chip (NoC). The NPU Bank of the NoC is configured with a network interface for connecting to the NoC. The NoC includes multiple computing cores CIM Core. The NPU Bank is connected to the computing core CIM Core through a router Router (R), and multiple routers Router form a mesh topology structure.
[0060] Specifically, the computing core CIM Core includes a Network Interface Unit (NIU) for connecting to the corresponding router Router.
[0061] The Network-on-Chip (NoC) provided by this embodiment, where each computing core CIM Core acts as a node and is connected to multiple adjacent nodes through corresponding routers, forming a distributed network connection including multiple information transmission links. This topology provides multiple paths for data transmission, effectively avoiding network congestion and improving the reliability of data transmission. At the same time, the Mesh topology facilitates the addition and deletion of nodes. When the system needs to be expanded, new nodes can be added to the Network-on-Chip (NoC) and connected to adjacent nodes to achieve seamless expansion of the system. In addition, each node is equipped with a standardized interface to support the access of different types of IP cores (Intellectual Property Cores), ensuring the compatibility and flexibility of the network architecture.
[0062] Preferably, the distributed network of the Network-on-Chip (NoC) can monitor the network traffic situation in real time. When the traffic on a certain link approaches or reaches the saturation state, the corresponding computing core CIM Core will automatically adjust the data sending rate to avoid network congestion.
[0063] At the same time, a credit-based flow control method is adopted. The computing core of the receiving node can feedback the available buffer space to the computing core of the sending node, and the sending node controls the data sending volume according to the feedback information to ensure the stability of data transmission. This adaptive flow control mechanism can dynamically adjust the flow according to the change of network load, improve the utilization rate of network resources, and ensure the stable operation of the system under different load conditions.
[0064] The Network-on-Chip (NoC) follows standardized communication protocols such as AXI4 to ensure the compatibility and interoperability between different IP cores. Based on the standard interface of bus semantics, each node can be accessed as a master node or a slave node to the interconnected network and communicate through transactions. During data transmission, the transaction requests sent by the master node are encapsulated and transmitted in the format specified by the protocol, and the slave node parses the requests according to the protocol and returns responses. The adoption of standardized communication protocols not only facilitates the integration of new IP cores, reduces the complexity of system design, but also provides a solid foundation for the expansion and upgrade of the Network-on-Chip, enabling the system to adapt to changing application requirements.
[0065] Such as Figure 3As shown, the Network-on-Chip (NoC) needs to follow a standardized protocol to support the rapid integration of IPs using the same communication protocol. Using the AXI4 bus protocol or other communication protocols, the standard interface is based on bus semantics. Each node accesses the interconnection network as a master or a slave node and communicates through transactions. For a write operation transaction issued by the master node, the network channel encapsulates and transmits the destination address of the slave node in the address channel, encapsulates it as a message header, and then encapsulates the data in the write data channel as the message body. The message is split into data packets and injected into the network through the routing ports of the node routers. After receiving the data packets, the slave node reassembles the message, extracts the address and control information from the header and feeds it into the AXI write address channel, and extracts the data from the message body and feeds it into the AXI write data channel. After all the messages are received, the network interface of the slave node creates a write response message and returns it to the master node. After the response message passes through the write response channel of the master node, the write transaction is completed.
[0066] So far, the technical solutions of this application have been described in conjunction with the accompanying drawings. However, it is easy for those skilled in the art to understand that the protection scope of this application is obviously not limited to the above specific implementation manners. Without departing from the principle of this application, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will all fall within the protection scope of this application.
Claims
1. A near-memory computing chip, characterized in that, It includes a storage chip and a logic chip, which are connected by a hybrid bonding method. The storage chip includes a plurality of storage module groups, and each storage module group includes a plurality of storage modules; the logic chip includes a plurality of memory controller groups corresponding to the storage module groups, and each memory controller group is provided with a plurality of memory controllers corresponding to the plurality of storage modules in the storage module group, and a plurality of NPU Banks are provided corresponding to the plurality of memory controller groups.
2. The near-memory computing chip according to claim 1, wherein A plurality of the memory controllers in each memory controller group are connected in one-to-one correspondence with a plurality of the memory controllers in other memory controller groups.
3. The near-memory computing chip according to claim 2, wherein The NPU Bank includes: A storage unit for data interaction with the corresponding storage module; An on-chip data cache for temporarily storing intermediate data generated during the execution of a computing task; A vector processing unit for assisting in completing vector computing tasks; and An SRAM in-memory computing core for executing computing tasks.
4. The near-memory computing chip according to any one of claims 1-3, characterized in that, It further includes an on-chip network connected to the plurality of NPU Banks to realize data interaction between the plurality of NPU Banks.
5. The near-memory computing chip according to claim 4, wherein The on-chip network includes a plurality of computing cores, each computing core is correspondingly provided with a router, and the routers are connected to form a mesh topology structure, and each NPU Bank is provided with a network interface for connecting to the mesh topology structure.
6. The near-memory computing chip according to claim 5, wherein Each computing core includes a network interface unit for connecting to the corresponding router.
7. The near-memory computing chip according to claim 5, wherein The router is configured to be able to detect the traffic of its link, and in response to the traffic of the link reaching a preset value, the router can automatically adjust the data transmission rate.
8. The near-memory computing chip according to claim 5, wherein The computing core can feedback the available cache space to the computing core upstream of the data transmission link, and the computing core upstream of the transmission link can adjust the data transmission amount in response to the feedback.
9. The near-memory computing chip according to claim 4, wherein The on-chip network follows the AXI4 communication protocol.