Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

598 results about "Local memories" patented technology

Process scheduling method and device based on NUMA architecture, electronic equipment and storage medium

The invention discloses a process scheduling method and device based on an NUMA framework, electronic equipment and a storage medium, and relates to the technical field of computer systems.The process scheduling method comprises the steps that in the process running process, comprehensive performance scores of all NUMA nodes are calculated by obtaining performance index data corresponding to all the NUMA nodes respectively, and the performance index data of all the NUMA nodes are obtained; the target NUMA node is selected based on the comprehensive performance score to be bound with the execution core of the current running process, and the required memory page is pre-allocated for the current running process in the local memory area of the target NUMA node, so that the current running process can access the local memory of the bound NUMA node, the probability that the process accesses the memory across the NUMA node is reduced, and the memory access efficiency is improved. Therefore, the technical problem that the memory access delay is remarkably increased due to the fact that the process frequently accesses the memory across the NUMA nodes can be solved, and the technical effects of shortening the memory access delay and improving the system performance are achieved.
Owner:INSPUR SUZHOU INTELLIGENT TECH CO LTD

Adaptive dynamic loading strategy for rendering large three-dimensional models

Methods, systems, and apparatus, including medium-encoded computer program products for loading and rendering include: obtaining a 3D spatial access tree data structure encoding location information for objects in a 3D model of an environment, wherein the 3D model is stored on a remote computer system; ranking a set of the objects in the 3D model to form an object hierarchy based at least on distances between each object of the set of objects and a specified viewpoint for a user within the environment, as determined using the three-dimensional spatial access tree data structure; selecting a proper subset of the set of objects to be rendered based on the object hierarchy and a current model load limit; downloading the proper subset to the local memory; and rendering the proper subset from the local memory to the display device based on the specified viewpoint within the environment for the user.
Owner:AUTODESK INC

Tensor Memory Accelerator Enhancements

One embodiment provides a graphics processor comprising a memory interface and a graphics core cluster including a plurality of graphics cores and tensor processing circuitry. The tensor processing circuitry includes a local memory, a tensor accelerator coupled with the local memory, the tensor accelerator configured to perform a matrix multiply and accumulate operation, and a tensor data movement accelerator configured to asynchronously transfer tensor data between a global memory coupled to the memory interface and the local memory. The tensor data movement accelerator includes circuitry configured to translate the tensor data from a first tensor format to a second tensor format.
Owner:INTEL CORP

Direct memory access (DMA) engine with network interface capabilities

Examples described herein include one or more processors; a network interface; and a direct memory access (DMA) engine communicatively coupled to the one or more processors. In some examples, the DMA engine is to receive a DMA data access request and based on an address in the DMA data access request corresponding to a remote memory device, the DMA engine is to cause the network interface to generate at least one packet for transmission to the remote memory device. In some examples, if the source address corresponds to a local memory device and the destination address corresponds to a remote memory device, the DMA engine is to cause the network interface to generate at least one packet for transmission to the remote memory device.
Owner:INTEL CORP

Data processing method and device and electronic equipment

The embodiment of the invention provides a data processing method and device and electronic equipment. In the embodiment, when a data writing request is received, whether a target cache region corresponding to a target file indicated by the request exists locally or not is checked firstly, if not, the target cache region is created in a local memory, and the size of the target cache region is the same as that of a storage page in an NAND Flash medium; and then, before to-be-written target data is written into the NAND Flash medium, the target data is written into the target cache region based on the size of the available storage space of the target cache region, and on this basis, when the size of the available storage space in the target cache region is a preset value, namely when the target cache region is full, the data in the target cache region is written into the NAND Flash medium. In this way, write amplification caused by the fact that the size of the target data to be written is smaller than that of the storage page can be effectively avoided, and therefore the PE service life of the NAND Flash medium is prolonged.
Owner:XINHUASAN INFORMATION TECH CO LTD

Method and system for managing coupon and providing coupon-based targeted advertisement by using artificial intelligence

Embodiments relate to utilizing artificial intelligence (AI) technology to enable comprehensive and systematic management of coupons that can be equivalent to marketing strategies or benefits. A first analysis on a raw data stored as text, image, video, or any combination thereof in the IT device is performed with an AI tool to automatically detect one or more coupons from the raw data. The result of the first analysis is treated as the detected coupons that will be stored in a coupon database or a local memory. A second analysis on the coupon database or local memory is performed with the AI tool to extract a detailed coupon information from each of the detected coupons. The detailed coupon information corresponding to each of the detected coupons is clustered based on one or more predetermined criteria and stored as the clustered result in the coupon database or local memory.
Owner:SK PLANET CO LTD

Hardware compression of sparse matrix content

The disclosure describes hardware compression of sparse matrix content. One embodiment provides a graphics processor, the graphics processor comprising: a base die, the base die comprising a plurality of chiplet slots; and a plurality of chiplets, the plurality of chiplets being coupled with the plurality of chiplet slots. At least one chiplet of the plurality of chiplets comprises: a graphical core cluster comprising a plurality of processing elements; a shared local memory coupled with the plurality of processing elements; a plurality of matrix engines coupled to the shared local memory; and codec circuitry coupled with the shared local memory and the plurality of matrix engines. Codec circuitry is configured to decode matrix data stored in a first format in a shared local memory into a second format for consumption by a plurality of matrix engines.
Owner:INTEL CORP

GPU asynchronous matrix multiply accumulate applications

One embodiment provides a graphics processor comprising a base die including a plurality of chiplet sockets and a plurality of chiplets coupled with the plurality of chiplet sockets. At least one of the plurality of chiplets including a plurality of processing elements, a distributed shared local memory coupled with the plurality of processing elements, a plurality of matrix engines coupled with the distributed shared local memory, and an asynchronous matrix multiply accumulate (MMA) controller configured to perform an asynchronous MMA operation via the plurality of matrix engines.
Owner:INTEL CORP

Tensor data moving accelerator

The invention relates to a tensor data movement accelerator. One embodiment provides a graphics processor comprising a memory interface and a graphics core cluster comprising a plurality of graphics cores and tensor processing circuitry. The tensor processing circuit includes: a local memory; a tensor accelerator coupled with the local memory, the tensor accelerator configured to perform a matrix multiply-accumulate operation; and a tensor data movement accelerator configured to asynchronously transfer tensor data between a global memory and a local memory coupled to the memory interface. The tensor data moving accelerator includes circuitry configured to convert tensor data from a first tensor format to a second tensor format.
Owner:INTEL CORP

Distributed energy intelligent matching method for heavy truck charging load scheduling

The invention relates to the technical field of distributed computing, and discloses a heavy truck charging load scheduling-oriented distributed energy intelligent matching method, which comprises the following steps of: establishing a charging service computing power mapping table at a scheduling node; maintaining a shadow counter in a local memory, extracting a pre-estimated computing power consumption value according to an event type and accumulating the pre-estimated computing power consumption value to the shadow counter, and executing linear numerical deduction on the shadow counter according to a reference logic subtraction rate so as to simulate a scheduling data throughput evolution process; adjusting a linear deduction rate parameter according to the deviation between the state feedback data and the numerical value of the shadow counter; and distributing the charging matching task to a charging station edge computing node of which the shadow counter value does not exceed a preset logic saturation threshold value. According to the invention, through an open-loop estimation and closed-loop calibration mechanism of a local logic state, an instantaneous congestion risk caused by physical feedback lag is eliminated; and logic state consistency and self-adaptive distribution of the whole network computing power resources are realized.
Owner:SOX (XIAMEN) TECH CO LTD

Data processing method, processor, chip, display card and electronic equipment

The invention relates to a data processing method, a processor, a chip, a display card and electronic equipment, and relates to the field of intelligent computing, and the method comprises the following steps: reading a part of a first matrix from a local memory or a register of a computing unit in response to M tensor computing engines in a tensor computing engine cluster, reading the second matrix from the tensor memory, and executing matrix operation by the M tensor calculation engines according to part of the first matrix and the second matrix respectively to obtain operation results corresponding to the M tensor calculation engines; and the M tensor calculation engines respectively write operation results corresponding to the M tensor calculation engines into registers of the M tensor calculation engines. According to the embodiment of the invention, the computing power can be increased under the condition that the sizes and the bandwidths of the local memory and the register in the computing unit are not increased.
Owner:MOORE THREADS TECH CO LTD

Consistency interconnection control device and method, electronic device, product and computing system

The invention discloses a consistency interconnection control device and method, an electronic device, a product and a computing system, and relates to the technical field of computers, the consistency interconnection control device is provided and comprises a consistency interconnection network, a first controller and a first connector, the consistency interconnection network comprises a plurality of interconnected consistency interconnection nodes, the consistency interconnection node is connected with the first card slot for installing the acceleration card through the first connector, and the first controller manages the state of the cache line of the acceleration card and controls the consistency interconnection network to execute the consistency interconnection read-write task between different acceleration cards, so that the consistency interconnection system for the acceleration computing cluster is realized. By implementing the cache consistency between the acceleration cards, when the acceleration computing cluster executes distributed computing, the cache can be hit firstly, so that access to a local memory and memories of other acceleration cards is reduced, the problem of high time consumption of communication between different accelerators of the acceleration computing cluster is solved, and the efficiency of distributed computing is improved.
Owner:SHANDONG HAILIANG INFORMATION TECH RES INST

Code compiling method, electronic equipment and storage medium

The embodiment of the invention provides a code compiling method, electronic equipment and a storage medium. The code compiling method is used for compiling a code comprising at least one thread local storage variable, and comprises the following steps: in response to existence of a target thread local storage variable, allocating the at least one target thread local storage variable to a thread local memory based on an attribute of the at least one target thread local storage variable, obtaining a thread local memory address corresponding to each target thread local storage variable, wherein the target thread local storage variable is the thread local storage variable used by the first function; updating an intermediate representation corresponding to the code based on a thread local memory address corresponding to each target thread local storage variable; and generating a machine code corresponding to the code based on the updated intermediate representation. According to the code compiling method, thread local storage model support under a parallel computing architecture is provided, and the problem that traditional thread local storage implementation cannot be directly applied to a GPU architecture is solved.
Owner:SHANGHAI BIREN TECH CO LTD

Low latency scratch memory path

An apparatus and method for efficiently processing vector memory accesses on an integrated circuit. In various implementations, a computing system includes a processing circuit with multiple compute circuits for executing wavefronts of a parallel data application. Each compute circuit includes a local memory subsystem for accessing data not found in vector register files of the compute circuit. The local memory subsystem includes a first execution pipeline and a second execution pipeline. The second execution pipeline processes vector stack access instructions that access temporary data such as stack data of a function call used by each wavefront that is generated based on the function call. The first execution pipeline processes other types of vector memory access instructions and includes multiple complex pipeline stages not found in the second execution pipeline. Thus, the second execution pipeline has a latency less than the latency of the first execution pipeline.
Owner:ADVANCED MICRO DEVICES INC

Memory allocation method and device

The embodiment of the invention provides a memory allocation method and device.The method comprises the steps that under the condition that page table page allocation is determined to be conducted on a target page table of a target process, a resource binding mechanism corresponding to the target process is obtained; according to a memory node binding constraint in the resource binding mechanism, determining a target memory node corresponding to the target process, and allocating a page table page for the target page table at the target memory node; when it is determined that the local memory of the target memory node is insufficient, so that the page table page allocation fails, a target physical page in the local memory is migrated to a first remote memory node, and the step of allocating the page table page to the target page table at the target memory node continues to be executed until the page table page allocation succeeds. By forcibly distributing the pages of the page table at the target memory node and combining with an intelligent memory migration mechanism, the performance loss caused by remote access of the page table is fundamentally eliminated.
Owner:ALIBABA CLOUD COMPUTING CO LTD

Software programmable flexible and dynamic optical transceivers

An optical transceiver includes an electro-optic front end; a digital-to-analog converter (DAC) and an analog-to-digital converter (ADC) connected to the electro-optic front end; and one or more Field Programmable Gate Arrays (FPGAs) connected to the DAC and the ADC, wherein the one or more FPGAs are connected to one or more of a local memory and a remote storage for loading FPGA bit files, and wherein the one or more FPGAs are loaded with a forward error correction (FEC) encoding app and a FEC decoding app. The FEC encoding app and the FEC decoding app can be selected based on any of an optical application and a standard compliance requirement.
Owner:CIENA CORP

Heterogeneous memory page migration device and method and readable storage medium

The invention discloses a heterogeneous memory page migration device and method and a readable storage medium, and relates to the technical field of data processing, the heterogeneous memory page migration device comprises a first computing unit, a migration unit and a memory unit, and the migration unit comprises a communicator and a processor; a communicator configured to receive an address translation request; the processor is configured to transmit page contents of a remote page to a local memory frame according to a remote physical frame number in a first page table item when determining that the target page in the address conversion request is the remote page, and replace the remote physical frame number in the first page table item with a local physical frame number; and the communicator is further configured to return the second page table item to the first computing unit, so that the problem of relatively long time consumption when the computing unit accesses a remote page and triggers a page error in the related technology is solved, and the time consumption for migrating the heterogeneous memory page is saved.
Owner:INSPUR SUZHOU INTELLIGENT TECH CO LTD

Low latency logical unit for a memory system

Methods, systems, and devices for low latency logical (L3A) unit for a memory system are described. For example, a logical unit, such as an L3Alogical unit or L3A logical unit number (LUN), may include a storage area for storing information for system swap operations, including, for example, a logical-block-address (LBA) range for one or more single-level-cells (SLCs). The logical unit may include a logical-to-physical (L2P) mapping table stored in a local memory of a memory system controller, such as in static random access memory (SRAM). In some examples, a reserved storage area may be overprovisioned, and the logical unit may be associated with a higher priority and a larger granularity than one or more other logical units. Further, one or more read-only (RO) descriptors stored to one or more registers, one or more provisioning parameters, or both, may be defined for the logical unit.
Owner:MICRON TECHNOLOGY INC

Separated memory-oriented page migration method and system

ActiveCN120631262AInput/output to record carriersMemory systemsRemote memory accessTerm memory
The invention discloses a separated memory-oriented page migration method and system, and belongs to the field of data storage. The page migration method comprises the following steps that: a node dynamically samples the access condition of a page through NUMA Balancing, tracks and marks the access record of the page; static screening is carried out in combination with an LRU list, hot pages are obtained through identification, and other pages are marked as cold pages; storing the identified hot pages and cold pages in a hot to-be-migrated queue and a cold to-be-migrated queue respectively; the background kernel thread continuously obtains to-be-migrated hot pages and cold pages from the hot to-be-migrated queue and the cold to-be-migrated queue, asynchronous page migration is conducted, the hot pages are migrated to the node from the far-end memory node, and the cold pages of the node are migrated to the far-end memory node; when the page is migrated, the access of the user to the page is decoupled. Preferentially migrating a real high-frequency access page to a local memory node; the hit rate of local memory access is improved, and the delay of remote memory access is reduced.
Owner:HUAZHONG UNIV OF SCI & TECH

Structured data storage method and system based on natural language transformation

The invention discloses a structured data storage method based on natural language transformation, which comprises the following steps of: a system initialization configuration stage: deploying a protocol adapter in a local memory of a PC (Personal Computer) client, and loading natural language processing pipeline configuration parameters; a heterogeneous data acquisition stage: capturing a multi-source text data stream through the protocol adapter, uniformly converting the multi-source text data stream into a standardized data packet, and sending the standardized data packet to a message queue theme; a text cleaning stage: a named entity recognition stage: inputting the pure text data into an NER module deployed with a language model loader; in the conditional feature extraction stage, feature vectors are generated for texts meeting preset conditions on the basis of entity type tags in the entity recognition result; and a consistent storage stage: inserting the entity identification results into a relational database in batches, and updating the entity mapping relationship in the cache. According to the method, the intelligent level of cache management is remarkably improved, and the access fluency of the key data of the user is guaranteed.
Owner:TIANJIN AUTOHOME DATA INFORMATION TECH CO LTD

Memory access method and electronic equipment

The invention provides a memory access method and electronic equipment, and is applied to the technical field of accelerator systems. The memory access method comprises the following steps: determining a target memory access strategy from a preset memory access strategy set according to a local memory access feature of any computing core; under the condition that the target memory access strategy is a transmission optimization strategy, position information of the to-be-accessed data in a target cache line is determined according to a target cache line address in an initial memory access request initiated by any computing core; generating a bit mask for identifying the to-be-accessed data according to the position information; and adding the bit mask to the initial memory access request to form a memory access request, sending the memory access request to a target device to which the target cache line address belongs, and returning the to-be-accessed data by the target device according to the bit mask in the memory access request.
Owner:INSPUR SUZHOU INTELLIGENT TECH CO LTD

GPU asynchronous direct memory access application

The invention discloses a GPU asynchronous direct memory access application. One embodiment provides a graphics processor, comprising: a base die comprising a plurality of chiplet slots; and a plurality of chiplets, the plurality of chiplets being coupled with the plurality of chiplet slots. A chiplet of the plurality of chiplets includes: a graphics core cluster, the graphics core cluster including a plurality of graphics cores; a distributed shared local memory, the distributed shared local memory including a shared local memory within each of the plurality of graphics cores; and a direct memory access engine within each of the plurality of graphics cores, the direct memory access engine configured to asynchronously copy data from a memory device to the distributed shared local memory.
Owner:INTEL CORP

Memory allocation method and device and storage medium

The invention discloses a memory allocation method and device and a storage medium, and relates to the field of computers. Performing hierarchical division on the memory nodes based on access performance of the memory nodes, constructing an affinity group in combination with an access relationship between the computing nodes and the memory nodes, and incorporating local memory nodes and extended memory nodes into a memory allocation range; meanwhile, weight factors reflecting the relative access capability of the memory nodes are distributed to the memory nodes in the affinity group, so that the memory can be distributed in a weighted mode according to performance differences. The memory resource allocation does not simply depend on a physical topology and a local priority strategy, but comprehensively considers the bandwidth and delay difference of the heterogeneous memory, thereby avoiding the problems of excessive consumption of high-performance memory nodes and idle low-performance nodes; the problems of performance bottleneck and resource allocation imbalance caused by the fact that related technical strategies cannot deal with heterogeneous memory differences are solved, and the effects of balancing and efficiently utilizing the memory and improving the overall performance and resource utilization rate of the system are achieved.
Owner:SHANDONG YINGXIN COMP TECH CO LTD

Save-restore engine for access control

Methods, systems, and apparatus, for implementing a save-restore engine in a computing device. One of the apparatus includes a power manager configured to control power provided to a plurality of power domains on the device, wherein each power domain has a respective client device, wherein each respective client device has a respective access control (AC) component that is configured to control which other components on the device can communicate with the respective client device; and a save-restore engine (SRE) configured to save, in an isolated local memory, configuration data for an AC component located in a power domain affected by the power manager initiating a power collapse operation, and wherein the SRE is configured to restore, from the isolated local memory, the configuration data of the AC component when the power manager restores the power to the power domain of the AC component.
Owner:GOOGLE LLC

Hardware compression for sparse matrix content

One embodiment provides a graphics processor comprising a base die including a plurality of chiplet sockets and a plurality of chiplets coupled with the plurality of chiplet sockets. At least one of the plurality of chiplets include a graphics core cluster including a plurality of processing elements, a shared local memory coupled with the plurality of processing elements, a plurality of matrix engines coupled with the shared local memory, and codec circuitry coupled with the shared local memory and the plurality of matrix engines. The codec circuitry is configured to decode matrix data stored in the shared local memory in a first format into a second format for consumption by the plurality of matrix engines.
Owner:INTEL CORP

Intelligent leakage detection method and device for underground pipeline

The invention discloses an underground pipeline intelligent leakage detection method and device, and relates to the field of intelligent leakage detection.The method comprises the steps that a high-frequency pressure sequence of a target pipe section is collected, and based on the Boltzmann superposition principle, the high-frequency pressure sequence is used for constructing a neural lag operator network by taking the environment temperature as a heat rheological regulation factor; taking the full historical stress-strain memory field tensor as input, carrying out non-local memory characteristic and hysteresis loop analysis, carrying out point-by-point differential processing on a high-frequency pressure sequence and the intrinsic nonlinear hysteresis response signal of the pipe, stripping non-stationary background fluctuation, avoiding modulation interference of material nonlinearity on a fluid signal, and obtaining a non-linear hysteresis response signal of the pipe. The method comprises the following steps: extracting the abnormal residual error of the fidelity fluid dynamics, constructing a spatio-temporal evolution Poincare section, inputting the spatio-temporal evolution Poincare section into a convolutional neural network, judging the real leakage negative pressure wave and the random drift of the sensor through the morphological difference of attractors on the spatio-temporal evolution Poincare section, and eliminating the periodic false alarm caused by the rheological property of the material.
Owner:陕西昌硕科技有限公司

High-performance processor, processor cluster and electronic equipment

ActiveCN120540708AMachine execution arrangementsComputer architectureScalar processor
The invention provides a high-performance processor, a processor cluster and electronic equipment. The high-performance processor comprises a scalar processor, a vector processor and a local memory, the scalar processor and the vector processor share a local memory; the vector processor only accesses the local memory and is only called and executed by the scalar processor; connection is established between the scalar processor and the vector processor; the scalar processor is connected with the global memory; the scalar processor is used for acquiring instructions and parameters from the global memory; and after the scalar processor determines that the execution condition is met, calling the vector processor to execute the parameter-based execution task. According to the high-performance processor provided by the invention, after the scalar processor obtains the instruction and the parameter from the global memory, when the execution condition is determined to be met, the vector processor is called to execute the task based on the parameter, so that a plurality of heterogeneous cores are cooperatively used to meet different calculation requirements.
Owner:SHANGHAI SMARTLOGIC TECHNOLOGY LTD

Industrial chain knowledge graph dynamic updating method and system based on multi-source data

The invention provides an industrial chain knowledge graph dynamic updating method and system based on multi-source data, and relates to the technical field of data processing.The method comprises the steps that information of industrial classification nodes is obtained through a database query interface, multi-level sorting processing is conducted on an industrial classification list composed of the nodes, and an analysis template is initialized; when a user side selects a node, triggering a sub-tree acquisition request, querying a direct sub-industry classification list based on an industry classification list, searching a sub-industry classification node in a recursive lazy loading mode, and constructing a classification sub-tree structure; constructing a multi-level cache architecture, storing the analysis template and the classification sub-tree structure of the high-frequency access to a local memory cache, and storing the industry classification list and the classification sub-tree structure to a distributed cache; and querying and calling the template in the multi-level cache architecture, and outputting industrial classification template data, so that closed-loop synchronization from data source change to cache update to front-end interface display is realized, data consistency is guaranteed, and smooth and non-perceptual user experience is provided.
Owner:SHUZU TECHNOLOGY (NANJING) CO LTD

Data reading method and computing device

The embodiment of the invention provides a data reading method and computing equipment, and the method comprises the steps: under the condition that a process reads a first memory object through a target memory virtual address, if it is recognized that an uncorrectable error fault exists in a first memory, replacing a current page table with a target page table; a first memory object is stored in the first memory; the first memory is located in a physical memory in a main memory of the computing device; the main memory is a local memory; the current page table is used for indicating a mapping relation between the target memory virtual address and the first memory; the target page table is used for indicating a mapping relation between the target memory virtual address and a second memory; the second memory is located in a backup memory; the backup memory is physically isolated from the main memory; reading a second memory object from the second memory based on the target page table; wherein the second memory object is a copy of the first memory object, and according to the embodiment of the invention, the effect of improving the business processing efficiency can be realized.
Owner:XFUSION DIGITAL TECH CO LTD

Task processing method and computing device

The embodiment of the invention provides a task processing method and computing equipment, and the method comprises the steps: obtaining a processing task which comprises target application information; according to a mapping relation and the target application information, a target memory allocation proportion corresponding to the processing task is determined, the mapping relation comprises multiple pieces of application information and a memory allocation proportion corresponding to each piece of application information, and the memory allocation proportion is the proportion between a local memory used by the computing device and a computing fast connection CXL memory; and processing the processing task according to the target memory allocation proportion. In the method, the computing device can flexibly and accurately determine the adaptive target memory allocation proportion for each processing task according to the target application information in each processing task and the mapping relation, and accurately adjust the local memory and the CXL memory used by the computing device based on the target memory allocation proportion. Therefore, it is ensured that the computing device has high processing performance on the processing task.
Owner:XFUSION DIGITAL TECH CO LTD