Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

479 results about "Local memories" patented technology

Tensor Memory Accelerator Enhancements

One embodiment provides a graphics processor comprising a memory interface and a graphics core cluster including a plurality of graphics cores and tensor processing circuitry. The tensor processing circuitry includes a local memory, a tensor accelerator coupled with the local memory, the tensor accelerator configured to perform a matrix multiply and accumulate operation, and a tensor data movement accelerator configured to asynchronously transfer tensor data between a global memory coupled to the memory interface and the local memory. The tensor data movement accelerator includes circuitry configured to translate the tensor data from a first tensor format to a second tensor format.
Owner:INTEL CORP

Hardware compression of sparse matrix content

The disclosure describes hardware compression of sparse matrix content. One embodiment provides a graphics processor, the graphics processor comprising: a base die, the base die comprising a plurality of chiplet slots; and a plurality of chiplets, the plurality of chiplets being coupled with the plurality of chiplet slots. At least one chiplet of the plurality of chiplets comprises: a graphical core cluster comprising a plurality of processing elements; a shared local memory coupled with the plurality of processing elements; a plurality of matrix engines coupled to the shared local memory; and codec circuitry coupled with the shared local memory and the plurality of matrix engines. Codec circuitry is configured to decode matrix data stored in a first format in a shared local memory into a second format for consumption by a plurality of matrix engines.
Owner:INTEL CORP

GPU asynchronous matrix multiply accumulate applications

One embodiment provides a graphics processor comprising a base die including a plurality of chiplet sockets and a plurality of chiplets coupled with the plurality of chiplet sockets. At least one of the plurality of chiplets including a plurality of processing elements, a distributed shared local memory coupled with the plurality of processing elements, a plurality of matrix engines coupled with the distributed shared local memory, and an asynchronous matrix multiply accumulate (MMA) controller configured to perform an asynchronous MMA operation via the plurality of matrix engines.
Owner:INTEL CORP

Tensor data moving accelerator

The invention relates to a tensor data movement accelerator. One embodiment provides a graphics processor comprising a memory interface and a graphics core cluster comprising a plurality of graphics cores and tensor processing circuitry. The tensor processing circuit includes: a local memory; a tensor accelerator coupled with the local memory, the tensor accelerator configured to perform a matrix multiply-accumulate operation; and a tensor data movement accelerator configured to asynchronously transfer tensor data between a global memory and a local memory coupled to the memory interface. The tensor data moving accelerator includes circuitry configured to convert tensor data from a first tensor format to a second tensor format.
Owner:INTEL CORP

Distributed energy intelligent matching method for heavy truck charging load scheduling

The invention relates to the technical field of distributed computing, and discloses a heavy truck charging load scheduling-oriented distributed energy intelligent matching method, which comprises the following steps of: establishing a charging service computing power mapping table at a scheduling node; maintaining a shadow counter in a local memory, extracting a pre-estimated computing power consumption value according to an event type and accumulating the pre-estimated computing power consumption value to the shadow counter, and executing linear numerical deduction on the shadow counter according to a reference logic subtraction rate so as to simulate a scheduling data throughput evolution process; adjusting a linear deduction rate parameter according to the deviation between the state feedback data and the numerical value of the shadow counter; and distributing the charging matching task to a charging station edge computing node of which the shadow counter value does not exceed a preset logic saturation threshold value. According to the invention, through an open-loop estimation and closed-loop calibration mechanism of a local logic state, an instantaneous congestion risk caused by physical feedback lag is eliminated; and logic state consistency and self-adaptive distribution of the whole network computing power resources are realized.
Owner:SOX (XIAMEN) TECH CO LTD

Data processing method, processor, chip, display card and electronic equipment

The invention relates to a data processing method, a processor, a chip, a display card and electronic equipment, and relates to the field of intelligent computing, and the method comprises the following steps: reading a part of a first matrix from a local memory or a register of a computing unit in response to M tensor computing engines in a tensor computing engine cluster, reading the second matrix from the tensor memory, and executing matrix operation by the M tensor calculation engines according to part of the first matrix and the second matrix respectively to obtain operation results corresponding to the M tensor calculation engines; and the M tensor calculation engines respectively write operation results corresponding to the M tensor calculation engines into registers of the M tensor calculation engines. According to the embodiment of the invention, the computing power can be increased under the condition that the sizes and the bandwidths of the local memory and the register in the computing unit are not increased.
Owner:MOORE THREADS TECH CO LTD

Consistency interconnection control device and method, electronic device, product and computing system

The invention discloses a consistency interconnection control device and method, an electronic device, a product and a computing system, and relates to the technical field of computers, the consistency interconnection control device is provided and comprises a consistency interconnection network, a first controller and a first connector, the consistency interconnection network comprises a plurality of interconnected consistency interconnection nodes, the consistency interconnection node is connected with the first card slot for installing the acceleration card through the first connector, and the first controller manages the state of the cache line of the acceleration card and controls the consistency interconnection network to execute the consistency interconnection read-write task between different acceleration cards, so that the consistency interconnection system for the acceleration computing cluster is realized. By implementing the cache consistency between the acceleration cards, when the acceleration computing cluster executes distributed computing, the cache can be hit firstly, so that access to a local memory and memories of other acceleration cards is reduced, the problem of high time consumption of communication between different accelerators of the acceleration computing cluster is solved, and the efficiency of distributed computing is improved.
Owner:SHANDONG HAILIANG INFORMATION TECH RES INST

Low latency scratch memory path

An apparatus and method for efficiently processing vector memory accesses on an integrated circuit. In various implementations, a computing system includes a processing circuit with multiple compute circuits for executing wavefronts of a parallel data application. Each compute circuit includes a local memory subsystem for accessing data not found in vector register files of the compute circuit. The local memory subsystem includes a first execution pipeline and a second execution pipeline. The second execution pipeline processes vector stack access instructions that access temporary data such as stack data of a function call used by each wavefront that is generated based on the function call. The first execution pipeline processes other types of vector memory access instructions and includes multiple complex pipeline stages not found in the second execution pipeline. Thus, the second execution pipeline has a latency less than the latency of the first execution pipeline.
Owner:ADVANCED MICRO DEVICES INC

Memory allocation method and device

The embodiment of the invention provides a memory allocation method and device.The method comprises the steps that under the condition that page table page allocation is determined to be conducted on a target page table of a target process, a resource binding mechanism corresponding to the target process is obtained; according to a memory node binding constraint in the resource binding mechanism, determining a target memory node corresponding to the target process, and allocating a page table page for the target page table at the target memory node; when it is determined that the local memory of the target memory node is insufficient, so that the page table page allocation fails, a target physical page in the local memory is migrated to a first remote memory node, and the step of allocating the page table page to the target page table at the target memory node continues to be executed until the page table page allocation succeeds. By forcibly distributing the pages of the page table at the target memory node and combining with an intelligent memory migration mechanism, the performance loss caused by remote access of the page table is fundamentally eliminated.
Owner:ALIBABA CLOUD COMPUTING CO LTD

Low latency logical unit for a memory system

Methods, systems, and devices for low latency logical (L3A) unit for a memory system are described. For example, a logical unit, such as an L3Alogical unit or L3A logical unit number (LUN), may include a storage area for storing information for system swap operations, including, for example, a logical-block-address (LBA) range for one or more single-level-cells (SLCs). The logical unit may include a logical-to-physical (L2P) mapping table stored in a local memory of a memory system controller, such as in static random access memory (SRAM). In some examples, a reserved storage area may be overprovisioned, and the logical unit may be associated with a higher priority and a larger granularity than one or more other logical units. Further, one or more read-only (RO) descriptors stored to one or more registers, one or more provisioning parameters, or both, may be defined for the logical unit.
Owner:MICRON TECHNOLOGY INC

Separated memory-oriented page migration method and system

ActiveCN120631262AInput/output to record carriersMemory systemsRemote memory accessTerm memory
The invention discloses a separated memory-oriented page migration method and system, and belongs to the field of data storage. The page migration method comprises the following steps that: a node dynamically samples the access condition of a page through NUMA Balancing, tracks and marks the access record of the page; static screening is carried out in combination with an LRU list, hot pages are obtained through identification, and other pages are marked as cold pages; storing the identified hot pages and cold pages in a hot to-be-migrated queue and a cold to-be-migrated queue respectively; the background kernel thread continuously obtains to-be-migrated hot pages and cold pages from the hot to-be-migrated queue and the cold to-be-migrated queue, asynchronous page migration is conducted, the hot pages are migrated to the node from the far-end memory node, and the cold pages of the node are migrated to the far-end memory node; when the page is migrated, the access of the user to the page is decoupled. Preferentially migrating a real high-frequency access page to a local memory node; the hit rate of local memory access is improved, and the delay of remote memory access is reduced.
Owner:HUAZHONG UNIV OF SCI & TECH

Structured data storage method and system based on natural language transformation

The invention discloses a structured data storage method based on natural language transformation, which comprises the following steps of: a system initialization configuration stage: deploying a protocol adapter in a local memory of a PC (Personal Computer) client, and loading natural language processing pipeline configuration parameters; a heterogeneous data acquisition stage: capturing a multi-source text data stream through the protocol adapter, uniformly converting the multi-source text data stream into a standardized data packet, and sending the standardized data packet to a message queue theme; a text cleaning stage: a named entity recognition stage: inputting the pure text data into an NER module deployed with a language model loader; in the conditional feature extraction stage, feature vectors are generated for texts meeting preset conditions on the basis of entity type tags in the entity recognition result; and a consistent storage stage: inserting the entity identification results into a relational database in batches, and updating the entity mapping relationship in the cache. According to the method, the intelligent level of cache management is remarkably improved, and the access fluency of the key data of the user is guaranteed.
Owner:TIANJIN AUTOHOME DATA INFORMATION TECH CO LTD

Memory access method and electronic equipment

The invention provides a memory access method and electronic equipment, and is applied to the technical field of accelerator systems. The memory access method comprises the following steps: determining a target memory access strategy from a preset memory access strategy set according to a local memory access feature of any computing core; under the condition that the target memory access strategy is a transmission optimization strategy, position information of the to-be-accessed data in a target cache line is determined according to a target cache line address in an initial memory access request initiated by any computing core; generating a bit mask for identifying the to-be-accessed data according to the position information; and adding the bit mask to the initial memory access request to form a memory access request, sending the memory access request to a target device to which the target cache line address belongs, and returning the to-be-accessed data by the target device according to the bit mask in the memory access request.
Owner:INSPUR SUZHOU INTELLIGENT TECH CO LTD

GPU asynchronous direct memory access application

The invention discloses a GPU asynchronous direct memory access application. One embodiment provides a graphics processor, comprising: a base die comprising a plurality of chiplet slots; and a plurality of chiplets, the plurality of chiplets being coupled with the plurality of chiplet slots. A chiplet of the plurality of chiplets includes: a graphics core cluster, the graphics core cluster including a plurality of graphics cores; a distributed shared local memory, the distributed shared local memory including a shared local memory within each of the plurality of graphics cores; and a direct memory access engine within each of the plurality of graphics cores, the direct memory access engine configured to asynchronously copy data from a memory device to the distributed shared local memory.
Owner:INTEL CORP

Memory allocation method and device and storage medium

The invention discloses a memory allocation method and device and a storage medium, and relates to the field of computers. Performing hierarchical division on the memory nodes based on access performance of the memory nodes, constructing an affinity group in combination with an access relationship between the computing nodes and the memory nodes, and incorporating local memory nodes and extended memory nodes into a memory allocation range; meanwhile, weight factors reflecting the relative access capability of the memory nodes are distributed to the memory nodes in the affinity group, so that the memory can be distributed in a weighted mode according to performance differences. The memory resource allocation does not simply depend on a physical topology and a local priority strategy, but comprehensively considers the bandwidth and delay difference of the heterogeneous memory, thereby avoiding the problems of excessive consumption of high-performance memory nodes and idle low-performance nodes; the problems of performance bottleneck and resource allocation imbalance caused by the fact that related technical strategies cannot deal with heterogeneous memory differences are solved, and the effects of balancing and efficiently utilizing the memory and improving the overall performance and resource utilization rate of the system are achieved.
Owner:SHANDONG YINGXIN COMP TECH CO LTD

Hardware compression for sparse matrix content

One embodiment provides a graphics processor comprising a base die including a plurality of chiplet sockets and a plurality of chiplets coupled with the plurality of chiplet sockets. At least one of the plurality of chiplets include a graphics core cluster including a plurality of processing elements, a shared local memory coupled with the plurality of processing elements, a plurality of matrix engines coupled with the shared local memory, and codec circuitry coupled with the shared local memory and the plurality of matrix engines. The codec circuitry is configured to decode matrix data stored in the shared local memory in a first format into a second format for consumption by the plurality of matrix engines.
Owner:INTEL CORP

Intelligent leakage detection method and device for underground pipeline

The invention discloses an underground pipeline intelligent leakage detection method and device, and relates to the field of intelligent leakage detection.The method comprises the steps that a high-frequency pressure sequence of a target pipe section is collected, and based on the Boltzmann superposition principle, the high-frequency pressure sequence is used for constructing a neural lag operator network by taking the environment temperature as a heat rheological regulation factor; taking the full historical stress-strain memory field tensor as input, carrying out non-local memory characteristic and hysteresis loop analysis, carrying out point-by-point differential processing on a high-frequency pressure sequence and the intrinsic nonlinear hysteresis response signal of the pipe, stripping non-stationary background fluctuation, avoiding modulation interference of material nonlinearity on a fluid signal, and obtaining a non-linear hysteresis response signal of the pipe. The method comprises the following steps: extracting the abnormal residual error of the fidelity fluid dynamics, constructing a spatio-temporal evolution Poincare section, inputting the spatio-temporal evolution Poincare section into a convolutional neural network, judging the real leakage negative pressure wave and the random drift of the sensor through the morphological difference of attractors on the spatio-temporal evolution Poincare section, and eliminating the periodic false alarm caused by the rheological property of the material.
Owner:陕西昌硕科技有限公司

High-performance processor, processor cluster and electronic equipment

ActiveCN120540708AMachine execution arrangementsComputer architectureScalar processor
The invention provides a high-performance processor, a processor cluster and electronic equipment. The high-performance processor comprises a scalar processor, a vector processor and a local memory, the scalar processor and the vector processor share a local memory; the vector processor only accesses the local memory and is only called and executed by the scalar processor; connection is established between the scalar processor and the vector processor; the scalar processor is connected with the global memory; the scalar processor is used for acquiring instructions and parameters from the global memory; and after the scalar processor determines that the execution condition is met, calling the vector processor to execute the parameter-based execution task. According to the high-performance processor provided by the invention, after the scalar processor obtains the instruction and the parameter from the global memory, when the execution condition is determined to be met, the vector processor is called to execute the task based on the parameter, so that a plurality of heterogeneous cores are cooperatively used to meet different calculation requirements.
Owner:SHANGHAI SMARTLOGIC TECHNOLOGY LTD

Industrial chain knowledge graph dynamic updating method and system based on multi-source data

The invention provides an industrial chain knowledge graph dynamic updating method and system based on multi-source data, and relates to the technical field of data processing.The method comprises the steps that information of industrial classification nodes is obtained through a database query interface, multi-level sorting processing is conducted on an industrial classification list composed of the nodes, and an analysis template is initialized; when a user side selects a node, triggering a sub-tree acquisition request, querying a direct sub-industry classification list based on an industry classification list, searching a sub-industry classification node in a recursive lazy loading mode, and constructing a classification sub-tree structure; constructing a multi-level cache architecture, storing the analysis template and the classification sub-tree structure of the high-frequency access to a local memory cache, and storing the industry classification list and the classification sub-tree structure to a distributed cache; and querying and calling the template in the multi-level cache architecture, and outputting industrial classification template data, so that closed-loop synchronization from data source change to cache update to front-end interface display is realized, data consistency is guaranteed, and smooth and non-perceptual user experience is provided.
Owner:SHUZU TECHNOLOGY (NANJING) CO LTD

Task processing method and computing device

The embodiment of the invention provides a task processing method and computing equipment, and the method comprises the steps: obtaining a processing task which comprises target application information; according to a mapping relation and the target application information, a target memory allocation proportion corresponding to the processing task is determined, the mapping relation comprises multiple pieces of application information and a memory allocation proportion corresponding to each piece of application information, and the memory allocation proportion is the proportion between a local memory used by the computing device and a computing fast connection CXL memory; and processing the processing task according to the target memory allocation proportion. In the method, the computing device can flexibly and accurately determine the adaptive target memory allocation proportion for each processing task according to the target application information in each processing task and the mapping relation, and accurately adjust the local memory and the CXL memory used by the computing device based on the target memory allocation proportion. Therefore, it is ensured that the computing device has high processing performance on the processing task.
Owner:XFUSION DIGITAL TECH CO LTD

Memory management method and related equipment

The invention provides a memory management method and related equipment, and relates to the technical field of cloud computers. The method is applied to an aggregation host, and the aggregation host comprises a plurality of computing nodes; the method comprises the following steps: in response to starting an aggregation host, reserving a local memory for each computing node; after the aggregation host is started, creating a kernel thread for each computing node; and in response to starting the load in the aggregation host, copying a memory page table corresponding to the load from a computing fast link (CXL) memory to a local memory reserved by each computing node through a kernel thread of each computing node. According to the method and the device, the local memory is reserved for the computing node, the memory page table is copied to the local, the dependence on a CXL memory is reduced, and the reserved local memory is closer to the computing node, so that the access speed is higher, the memory access speed and efficiency can be improved, and the problem that the performance is limited when a super aggregation host runs a service with a high requirement on the memory is solved.
Owner:CHINA TELECOM CORP LTD +1

Data transmission method and device, equipment and storage medium

The invention discloses a data transmission method and device, equipment and a storage medium, and the method comprises the steps: receiving different types of local memory semantic data sent locally, and obtaining a queue identifier corresponding to the local memory semantic data; merging the local memory semantic data with the same queue identifier and belonging to the same type to obtain a local merged data block; and adding a packet header to the local merged data block through a direct memory access (RDMA) component to obtain encapsulated data, and sending the encapsulated data to the remote equipment through the switch. Local memory semantic data with the same queue identifier and belonging to the same type are merged to obtain a local merged data block, and the local merged data block is packaged once and then sent to remote equipment, so that efficient transmission of memory semantic data of a small data packet is realized through queue merging; and after a plurality of small data packets are combined, the packet header only needs to be added once, so that the extra overhead of the packet header in the packaging process is reduced, and the transmission efficiency of the effective load is improved.
Owner:SHANGHAI SUIYUAN TECH CO LTD

Firmware management engine for integrated system memory management and firmware provisioning

Methods, systems, and devices for providing firmware management using a firmware management engine of a cloud computing system are described. The firmware management engine supports using a hot-plugged memory (e.g., Compute Express Link (CXL) pooled memory) to temporarily migrate system memory (e.g., Central Processing Unit (CPU) local memory or CXL attached memory) of a server node, while performing firmware management operations (e.g., a firmware update) on the server node. The hot-plugged memory can be CXL pooled memory that is a shared at a rack level, the CXL pooled memory can be used to store system memory data of system memory while updating and activating new memory initialization firmware code. A virtual machine associated with the server node stays operational temporarily using the hot-plugged memory. Firmware management is performed on a server node while virtual machines associated with the server node and the system memory of the server node remain operational.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Method and system for directly storing audio and video streams to cloud

The invention relates to a method and system for directly storing audio and video streams to a cloud end, and the method comprises the steps: enabling an audio and video collection terminal to carry out the coding and packaging of a video stream according to a GOP as a unit, and generating a transmission stream segment; establishing a data channel with a cloud server based on a link parameter provided by the signaling server, and uploading the transport stream fragment to the cloud server; after an upload success response is received, deleting the local transport stream segments, and continuously transmitting the remaining transport stream segments until a termination signal is detected; in response to the termination signal, sending a termination request to the cloud server; and based on the termination request, the cloud server performs merging operation on all the transport stream fragments which are accumulatively uploaded at this time, and a complete transport stream file is generated and stored. Through application of the method and the device, the problem that the memory pressure of the audio and video acquisition terminal and the cloud management efficiency cannot be considered at the same time due to the fixed slice duration is solved, and reasonable optimization of the local memory and cloud resources is realized.
Owner:E SURFING VISION TECHNOLOGY CO LTD

Inline materialization of database pages in postgresql-compatible systems

PendingUS20260133987A1Database distribution/replicationSpecial data processing applicationsRelational database management systemTerm memory
Techniques discussed herein relate to an object-relational database management system (ODMS) (e.g., a PostgreSQL ODMS) that provides in-line materialization of database pages corresponding to read requests. A replica node of the ODMS may be configured to share access to a shared block storage volume with a primary node. The replica node may receive log updates from the primary indicating changes to the object-relational database. The read replica, in response to receiving a read request, may update a previous version of a database page, store the database page in local memory, and replay the log updates to update the database page in local memory. Data may be provided from the updated database page in response to the read request.
Owner:ORACLE INT CORP

Telecommunications switch infrastructure for real-time audio interpretation and recording via artificial intelligence agents

A telecommunications switch-type platform positioned between a private-branch exchange and a data network dynamically hosts artificial-intelligence agents that interpret, transcribe, translate and transliterate live call traffic while the session is in progress. Machine-readable instructions stored in local memory instantiate, migrate and terminate the agents on demand across dedicated processing resources such as field-programmable gate arrays, application-specific integrated circuits or system-on-chip devices, thereby maintaining end-to-end latency below 250 milliseconds. An integrated call-audio recorder captures bidirectional media streams for secure archiving without interrupting service. The platform exposes a network interface that passes both signalling and media traffic, enabling seamless deployment inside call-centre or enterprise environments and compatibility with softswitch architectures conforming to class-4 Computational Interpretation of Communication using Artificial Intelligence standards. The same functional stack is deliverable as a computer-implemented method and as a non-transitory machine-readable medium storing the instructions executed by the processing resources.
Owner:CUNNINGHAM CHERYL

Mechanism for performing distributed power management of a multi-GPU system by powering down links based on previously detected idle conditions

Systems, apparatuses, and methods for efficient power management of a multi-node computing system are disclosed. A computing system includes multiple nodes that receive tasks to process. The nodes include a processor, local memory, a power controller, and multiple link interfaces for transferring messages with other nodes across links. Using a distributed approach for power management, negotiation for powering down components of the computing system occurs without performing a centralized system-wide power down. Each node is able to power down its links, its processor and other components regardless of whether other components of the computing system are still active or powered up. A link interface initiates power down of a link with delay or without delay based on a prediction of whether a link idle condition leads to the link interface remaining idle for at least a target idle threshold period of time.
Owner:ADVANCED MICRO DEVICES INC

Memory-efficient encoding and decoding of boundary conditions in the lattice boltzmann method

Methods, systems, and apparatus, including medium-encoded computer program products include: receiving an allocation of storage locations in a first memory, where the storage locations are used for storing representations of particle populations for respective lattice units of a discretized space, where the first memory is a random access memory of a processing system; receiving boundary conditions data; encoding the boundary condition data in unused portions of the storage locations of the first memory during at least one time step of a lattice Boltzmann modelling process; providing the boundary conditions data from the first memory to a second memory for applying, in the second memory, the boundary conditions at the one or more boundaries during the at least one time step, where the second memory is a local memory of the processing system; and providing a result of the at least one time step of the lattice Boltzmann modelling process.
Owner:AUTODESK INC

Test system, method and equipment based on memory extension equipment, medium and product

The invention discloses a test system, method, device, medium and product based on a memory extension device, and relates to the technical field of computers, a processor and a memory extension interconnection component are connected through a memory extension bus, and a simulation test of the memory extension device is added. The memory expansion interconnection assembly is connected with the memory expansion device through the serial bus, so that the processor is directly connected with the memory expansion device, and the memory expansion interconnection assembly is located at the position outside the processor. Delay caused by a local memory access path is avoided, and the transmission efficiency is improved. The technical problem that the simulation processing result obtained according to the program duration under a traditional memory access path is low in accuracy can be solved, and the technical effects that the simulation test of the memory extension device is achieved, meanwhile, the function of full-system simulation is achieved, and the simulation accuracy of memory access of the mounted CXL memory extension device is improved are achieved.
Owner:SHANDONG HAILIANG INFORMATION TECH RES INST