Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

12303 results about "Computer architecture" patented technology

In computer engineering, computer architecture is a set of rules and methods that describe the functionality, organization, and implementation of computer systems. Some definitions of architecture define it as describing the capabilities and programming model of a computer but not a particular implementation. In other definitions computer architecture involves instruction set architecture design, microarchitecture design, logic design, and implementation.

AI-Optimized Memory Fabric for Large Contexts and Multimodal Workloads

A coherent, intelligent, packet-switched memory fabric enables predictive, cache-coherent access across distributed compute, accelerator, and memory resources using a Memory-Fabric Transaction Layer Protocol (MF-TLP). MF-TLP defines routable packet formats for read, write, vectorized, atomic, reduction, collective, and predictive-prefetch transactions executed by memory-centric network interface controllers (MC-NICs). Each MC-NIC performs packet parsing, address translation, coherence management, and near-memory arithmetic or tensor operations while coordinating with MF-TLP-aware switches providing hierarchical directory control, multi-path routing, and in-network aggregation. Vectorized and multimodal packets encode multiple addresses or tensor offsets to reduce scatter / gather overhead, and programmable caching and quality-of-service modules manage tiered memory and tenant fairness. MF-TLP supports extension headers for predictive prefetch, collective coordination, and tenant governance, operating across hierarchical leaf-spine topologies using Ultra-Ethernet Transport, InfiniBand, or CXL fabrics. The system delivers scalable, low-latency, memory-centric orchestration for large-language-model training, multimodal AI, and data-intensive analytics.
Owner:QOMPLX INC

AI Serving Hardware and Software Frontier Enhancements

A computer system implements a unified framework integrating an adaptive elastic funnel (AEF) with a convergent intelligence fabric (CIF) for multi-agent AI collaboration. The system provides a universal multi-modal key-value subsystem for sharing partial computations, implements hybrid placement strategies for dynamic memory management, and incorporates quantum-resistant secure enclaves. The architecture integrates hardware acceleration through GPU-FPGA hybrid caching and neuromorphic processors, applies adaptive energy and thermal management across hardware generations, and implements autonomous flash resource orchestration with multi-dimensional wear management. The system orchestrates tensor workflows using hierarchical scheduling, enables cross-agent collaboration with privacy preservation, and supports continuous learning without catastrophic forgetting. This integration delivers unprecedented computational efficiency and security in high-dimensional decision-making environments while supporting incremental adoption through modular interfaces.
Owner:QOMPLX INC

Adaptive Real Time Image and Video Processing Using PCM-Enhanced Visual Strategy Caching and Multi-Stage Cognitive Routing

A system and method for adaptive image and video processing using a Persistent Cognitive Machine (PCM) architecture with visual strategy caching. The system receives degraded input media and extracts degradation fingerprints to query a PCM-based visual strategy cache containing previously successful processing strategies. When matching cached strategies are found above a relevance threshold, they are retrieved and applied directly. When no match exists, the input is processed through transform-domain networks to generate new strategies. A pattern synthesizer combines multiple strategies for complex degradation types. The system evaluates processing effectiveness using a feedback controller and stores successful strategies in the hierarchical cache. This cognitive approach enables real-time processing with continuously improving performance as the cache learns from successful patterns. The adaptive architecture eliminates redundant processing while maintaining high-quality output, making it suitable for diverse imaging and video applications requiring efficient enhancement capabilities with superior performance over traditional methods.
Owner:ATOMBEAM TECH INC

Methods for delta-QP signaling for decoder parallelization in hevc

ActiveUS20120183049A1Color television with pulse code modulationColor television with bandwidth reductionComputer architectureCoded block flag
By implementing a new bitstream for a Delta-Quantization Parameter (DQP), a decoder is able to implement parallel decoding of multiple coding units within a largest coding unit. In some embodiments, the DQP is placed immediately after the mode information of the first coding unit. In some embodiments, the DQP is placed after the mode information of the first non-skipped coding unit. In some embodiments, the DQP is placed after the first non-zero coded block flag.
Owner:SONY GROUP CORP

Method and apparatus for efficient access to multidimensional data structures and / or other large data blocks

A parallel processing unit comprises a plurality of processors each being coupled to a memory access hardware circuitry. Each memory access hardware circuitry is configured to receive, from the coupled processor, a memory access request specifying a coordinate of a multidimensional data structure, wherein the memory access hardware circuit is one of a plurality of memory access circuitry each coupled to a respective one of the processors; and, in response to the memory access request, translate the coordinate of the multidimensional data structure into plural memory addresses for the multidimensional data structure and using the plural memory addresses, asynchronously transfer at least a portion of the multidimensional data structure for processing by at least the coupled processor. The memory locations may be in the shared memory of the coupled processor and / or an external memory.
Owner:NVIDIA CORP

Multi-protocol dynamic switching method and system based on FPGA

The invention relates to the field of electronic communication, and provides a multi-protocol dynamic switching method and system based on an FPGA. The method comprises the following steps: carrying out regional division on an FPGA chip through a layout planning tool to obtain a static logic region and a plurality of reconfigurable logic partitions; performing hanging processing on the current protocol module based on the protocol switching trigger signal, and writing running state data of the current protocol module into a state cache unit to obtain protocol context snapshot data; according to the target protocol type, bit stream loading is carried out on the corresponding reconfigurable logic partition through configuration of an access port, and a reconfiguration protocol module is obtained; and reading protocol context snapshot data, and performing state recovery on the reconstruction protocol module to obtain a target protocol data path. According to the invention, rapid protocol switching is realized, the utilization rate of hardware resources is improved, and the system deployment cost is reduced.
Owner:TIANJIN JINHANG COMP TECH RES INST

In-memory computing circuit chip based on magnetic cache and computing device

The embodiment of the invention discloses an in-memory computing circuit based on a magnetic cache, and the circuit comprises at least one magnetic cache unit, at least one in-memory computing unit, and a timer. The magnetic cache unit in the at least one magnetic cache unit is used for caching data output by the corresponding in-memory computing unit as to-be-processed data within the corresponding data retention time; the timer is used for respectively setting data retention time for the at least one magnetic cache unit; and the in-memory computing unit in the at least one in-memory computing unit is used for extracting the data to be processed from the corresponding magnetic cache unit for calculation and outputting the computed data to other magnetic cache units. According to the embodiment of the invention, the invention achieves the flexible adjustment of the data retention time of the magnetic cache unit in various in-memory calculation scenes, and achieves the provision of a high-capacity cache for the data needed by in-memory computing under the lower power consumption.
Owner:NANJING HOUMO TECH CO LTD

Industrial personal computer and multi-graphics card collaborative parallel operation acceleration system

The invention discloses an industrial personal computer and multi-graphics card collaborative parallel computation acceleration system, which relates to the technical field of industrial resource allocation and parallel computation, and comprises a resource monitoring and predicting module, a resource management module and a prediction type resource preparation module, the task splitting and collaborative execution module comprises a task splitting module, a cross-node collaborative module and a collaborative operation engine; the intelligent scheduling and dynamic resource allocation module comprises an intelligent scheduler, a dynamic resource allocation module and a conflict avoidance module. According to the method, the GPU video memory utilization rate, the core utilization rate, the temperature, the video memory fragment rate, the available video memory total amount, the CPU core total utilization rate, the load condition and the idle core number index are collected in real time through a resource monitoring module, and the video memory capacity, the GPU core occupancy rate and the CPU load requirement are predicted in advance before a task is submitted in combination with a gradient boosting decision tree and a neural network prediction model; and resources are reserved, so that the scheduling delay is remarkably reduced, and the scheduling hit rate is improved.
Owner:ZHUHAI SHININGDA TECH CO LTD

Two-level context caching and eviction for scatter-gather DMA

One aspect of the instant disclosure may provide a system and method for processing scatter-gather direct memory access (S-G DMA) instructions. During operation, the system may receive an S-G DMA instruction associated with a message and gather instruction context for the S-G DMA instruction. An S-G DMA processor may process the S-G DMA instruction based on the gathered instruction context and determine whether there exists a pending S-G DMA instruction associated with the message. In response to the presence of the pending S-G DMA instruction, the system stores the instruction context in a hot context cache at an address corresponding to the pending S-G DMA instruction. In response to the absence of the pending S-G DMA instruction, the system stores the instruction context in a cold context cache.
Owner:HEWLETT PACKARD ENTERPRISE DEV LP

Memory processing method based on intelligent agent, storage medium and electronic device

The embodiment of the invention provides a memory processing method based on an intelligent agent, a storage medium and an electronic device, a memory architecture of the intelligent agent comprises a first memory layer, a second memory layer and a third memory layer, the first memory layer is used for storing historical dialogue information of interactive dialogue between at least one interactive object and the intelligent agent, and the second memory layer is used for storing historical dialogue information of interactive dialogue between at least one interactive object and the intelligent agent. The second memory layer is used for storing structured events extracted from historical dialogue information, and the third memory layer is used for storing pattern induction memory obtained by performing pattern induction on the structured dialogue events; the method comprises the steps that in response to current round dialogue input information of a target interaction object, target storage information associated with the current round dialogue input information is retrieved from at least one memory layer in a memory architecture, and the current round dialogue input information and the target storage information are assembled into a current cue word; and submitting the current prompt word to a specified interaction model through the intelligent agent, and outputting a response result of the specified interaction model to the interaction object.
Owner:ZTE CORP

Large language model accelerator architecture based on three-dimensional NAND flash memory

The invention discloses a large language model accelerator architecture based on a three-dimensional NAND flash memory, and belongs to the technical field of calculation, reckoning or counting. The architecture comprises a three-dimensional NAND flash memory used for executing feedforward neural network calculation; the auxiliary calculation unit is used for executing attention mechanism calculation; the DRAM chip is used for storing attention mechanism related weights and KV cache; and the interconnection resource is used for realizing data interaction among the components. Wherein the three-dimensional NAND flash memory comprises a logic chip and an NAND array chip, the logic chip is used for controlling and executing calculation, and the NAND array chip is used for storing weights and participating in calculation. The invention further provides a scheduling method based on KV cache awareness. The scheduling method comprises decomposition and dynamic allocation of calculation tasks. Through collaborative design of the hardware architecture and the scheduling method, the memory wall bottleneck in large language model calculation can be relieved, the energy consumption is reduced, and the overall calculation performance is improved.
Owner:SOUTHEAST UNIV

Methods and systems for prioritization of group computing tasks

A system for prioritization of group computing tasks is described. The system includes at least a processor and a memory communicatively connected to the at least a processor. The memory contains instructions configuring the at least a processor to detect a plurality of active nodes communicatively connected in a group computing environment and receive a plurality of computing tasks associated with the plurality of active nodes for execution in the group computing environment. The at least a processor is also configured to determine a computing demand associated with each of the plurality of computing tasks on the group computing environment and establish a priority for the plurality of computing tasks as a function of the computing demand associated with each of the plurality of computing tasks.
Owner:PARRY LABS LLC

Intelligent edge computing cooperative processing system based on integrated circuit

The invention relates to the technical field of edge computing, and discloses an intelligent edge computing co-processing system based on an integrated circuit, which comprises a heterogeneous computing cluster module, a hardware computing unit set, a multi-core control processor based on RISC-V, a programmable pulse tensor computing array and a reconfigurable engine oriented to streaming processing, according to the intelligent edge computing cooperative processing system based on the integrated circuit, zero-delay data exchange is realized through silicon intermediate layer integration of the heterogeneous computing cluster module and a snakelike data channel, and traditional bus arbitration delay is eliminated through real-time operation code analysis and optimal optical communication path mapping of the hardware task routing matrix (TRF) module; the priority path distribution of the task scheduling subsystem is synchronously coordinated with the time-sensitive task, so that the collaborative efficiency of the computing unit is improved in multiple dimensions, the non-blocking transmission of the high-priority task is ensured, and the effect of enhancing the overall processing capability and response speed of the intelligent edge computing is achieved.
Owner:QIQIHAR QISAN MACHINE TOOL

Parallel computing method and device, electronic equipment and storage medium

The invention provides a parallel computing method and device, electronic equipment and a storage medium, and relates to the technical field of parallel computing, and the method comprises the steps: carrying out the first protocol operation of a target tensor based on each computing core in each stream processor cluster, and generating a data block containing the computing result of each computing core; writing a data block generated by each stream processor cluster into a shared cache; under the condition that each stream processor cluster completes the first protocol operation, reading a data block written by each stream processor cluster from the shared cache; and executing a second protocol operation on the data block read from the shared cache to generate a calculation result of the target tensor. According to the method and device provided by the invention, the parallel architecture and memory access characteristics of the artificial intelligence chip can be better matched, the unnecessary calculation delay and synchronization overhead of the cross-flow processor cluster in the parallel calculation process are reduced, the bandwidth utilization rate of the shared cache is improved, and the overall performance and calculation efficiency of parallel calculation are remarkably improved.
Owner:SHANGHAI BIREN TECH CO LTD

Clock domain crossing synchronization circuit and method based on multi-phase clock

The invention relates to the technical field of integrated circuits, and provides a clock domain crossing synchronization circuit and method based on a multiphase clock, and the circuit comprises a selector, an asynchronous clock synchronization module and a digital signal processor. Wherein serial data to be transmitted and a low-frequency clock are input into the selector, and the selector is used for outputting corresponding RX parallel data to the asynchronous clock synchronization module according to the serial data and the low-frequency clock; the asynchronous clock synchronization module is used for outputting corresponding alignment data to the digital signal processor according to the input RX parallel data, the clock of the digital signal processor and the low-frequency clock; the asynchronous clock synchronization module is used for performing phase alignment on the RX parallel data and a clock of the digital signal processor; and the digital signal processor is used for carrying out data processing on the alignment data. According to the invention, stable transmission of low-delay cross-clock domain data can be realized.
Owner:SHANGHAI SHENGLIANKE SEMICONDUCTOR CO LTD

Parallel processor dynamic resource allocation system and method based on reconfigurable hardware

The invention discloses a parallel processor dynamic resource allocation system and method based on reconfigurable hardware, and relates to the technical field of computer chips, and the system comprises a workload monitoring module which is responsible for monitoring the workload type and the resource demand of a task processed by a parallel processor in real time, and determining the task type and the resource demand by identifying the task type and the resource demand; a monitoring result is fed back to the resource allocation control module; the resource allocation control module is responsible for generating a corresponding resource allocation control signal based on feedback information of the workload monitoring module and dynamically configuring the reconfigurable hardware module; the reconfigurable hardware module is responsible for realizing efficient adaptation to diversified tasks through a plurality of reconfigurable units formed by programmable logic devices according to different working load dynamic reconfiguration functions and connection modes; and the data caching and transmission module is responsible for cross-module data circulation and global data storage. Different task requirements can be accurately adapted, and the resource utilization rate is improved.
Owner:YUANQIXIN (SHANDONG) SEMICONDUCTOR TECHNOLOGY CO LTD

Clock optimization method, device and equipment and computer readable storage medium

The invention discloses a clock optimization method, device and equipment and a computer readable storage medium, which are applied to the field of integrated circuits, and comprise the following steps: reading a netlist of each fragment, and traversing each netlist to generate an instance tree corresponding to each fragment; performing reverse tracing and forward tracing on each clock signal based on each instance tree to obtain a clock network path corresponding to each clock signal; determining a redundant clock according to each clock network path, and performing corresponding processing on the redundant clock; the redundant clock comprises an isolated redundant clock and an equivalent redundant clock. According to the method, the instance tree is generated by traversing the netlist, and the complete clock network path is obtained by combining reverse tracing and forward tracing, so that isolated redundant clocks and equivalent redundant clocks can be comprehensively captured, and automatic tracking of the clock network path and accurate identification of the redundant clocks are realized; the redundant clock is correspondingly processed, so that unnecessary clock generation logic and wiring resource occupation can be reduced, the chip power consumption and the area overhead are reduced, and the resource utilization rate is improved.
Owner:SHANDONG BOSUAN ZHIXIN INFORMATION TECHNOLOGY CO LTD

Automated hardware-aware deployment of machine learning pipelines on chipsets

A method or system for implementing a machine learning pipeline on a chipset comprising a plurality of hardware compute elements. The system accesses a hardware-agnostic functional description of the machine learning pipeline, wherein the description specifies a plurality of functional modules, including at least one machine learning model. Hardware specifications of the chipset are accessed to identify the available hardware compute elements. Based on the hardware specifications, the functional modules are synthesized into a plurality of interconnected executable components configured to execute on at least two different hardware compute elements. An implementation package is generated, comprising the executable components and metadata describing interconnections between them. The implementation package is then deployed to the chipset, where the executable components are executed by the identified hardware compute elements.
Owner:SIMA TECHNOLOGIES INC

Memory architecture-oriented dual-precision general matrix multiplication optimization method and system

The invention belongs to the related technical field of high-performance computing, and provides a memory architecture-oriented dual-precision general matrix multiplication optimization method and system in order to solve the problems of limited computing power and access efficiency and the like in the prior art. Decomposing the matrix into a plurality of sub-matrix blocks according to the slave core array topology; the slave core receives the sub-matrix blocks issued by the master core, divides the sub-matrix blocks into small sub-matrix blocks based on a uniform blocking rule, loads the small sub-matrix blocks to an independent buffer area of a local data memory based on a DMA double-buffer protocol, divides the small sub-matrix blocks in the buffer area into SIMD vectors according to the SIMD unit characteristics of the slave core, and sends the SIMD vectors to the slave core; vectorization calculation and caching operation are alternately switched according to an iteration period through different independent buffer areas; and after all the slave cores finish calculation, the master core collects results written back to the master memory by the slave cores to obtain a final operation result, and double breakthrough of calculation power and memory access efficiency is realized.
Owner:QILU UNIVERSITY OF TECHNOLOGY (SHANDONG ACADEMY OF SCIENCES) +1

Sub-cell, MAC array and bit-width reconfigurable mixed-signal in-memory computing module

A mixed-signal in-memory computing sub-cell requires only 9 transistors for 1-bit multiplication. In one aspect, a computing cell is constructed from a plurality of such sub-cells that share a common computing capacitor and common transistors. As a result, the average number of transistors in each sub-cell is close to 6. Also proposed is a MAC array for performing MAC operations, which includes a plurality of the computing cells each activating the sub-cells therein in a time-multiplexed manner. Also proposed is an in-memory mixed-signal computing module for digitalizing parallel analog outputs of the MAC array and for performing other tasks in the digital domain.
Owner:REEXEN TECH CO LTD

Test instruction generation method and device, electronic equipment, medium and product

The embodiment of the invention discloses a test instruction generation method and device, electronic equipment, a medium and a product. The method comprises the steps that hardware constraint information of a specified processor architecture is obtained, and the hardware constraint information comprises a plurality of threads supported by the specified processor architecture, the number of operators to be executed by each thread and an address space corresponding to each thread; for each thread, according to the number of operators to be executed by the thread, calling different operators through each execution engine, and generating a target operator queue corresponding to the thread; determining a test instruction stream corresponding to a target operator queue of the thread according to the address space corresponding to the thread; wherein each target operator in the target operator queue corresponds to one test instruction in the test instruction stream, and an output address of at least one test instruction in the test instruction stream is an input address of a subsequent test instruction; and generating a test instruction file of the specified processor architecture according to the test instruction stream corresponding to each thread.
Owner:SHANGHAI ORIENTAL COMPUTER TECHNOLOGY CO LTD

Hardware accelerator facing triple sparse matrix multiplication, equipment and application method thereof

The invention discloses a hardware accelerator and equipment oriented to triple sparse matrix multiplication and an application method thereof.The hardware accelerator comprises a high-bandwidth memory HBM, a crossbar switch network and an on-chip processing unit which are connected in sequence, and the on-chip processing unit comprises a hierarchical cache module, a global controller and a plurality of computing chips; each calculation piece comprises an RA calculation array, a TP calculation array and a local controller, wherein the RA calculation array and the TP calculation array are respectively used for executing front-end operation T = R * A and rear-end operation C = T * P in triple sparse matrix multiplication. The method aims at solving the problem that when a traditional universal processor processes triple sparse matrix multiplication, due to irregular memory access, uneven calculation load and sharp increase of middle parts and results, huge off-chip data carrying is confronted with serious performance and energy efficiency bottlenecks, and the calculation performance and energy efficiency of triple sparse matrix multiplication are improved.
Owner:NAT UNIV OF DEFENSE TECH

Clock reset system level verification method and device based on general verification methodology

The invention relates to the technical field of chip testing, and discloses a clock reset system-level verification method and device based on general verification methodology, and the method comprises the steps: firstly analyzing a clock topological relation defined in a system control unit architecture document, and generating a clock reset information table; on the basis of the information table, an executable verification environment component and a virtual interface component are automatically generated through a script, the verification environment component comprises an excitation forwarding and processing component and can generate clock and reset signals matched with framework definition, and the virtual interface component can generate actual clock and reset waveforms; and finally, constructing a verification platform based on the components, calling the excitation sequence in the parameterized test sequence library, and carrying out system-level clock reset verification on the digital chip. And meanwhile, the parameterized test sequence library supports flexible adjustment of a verification scene and ensures that a test case fits a clock frequency defined by an architecture document, and a reset mechanism can improve the efficiency, accuracy and reliability of digital chip clock reset system-level verification.
Owner:JINAN MAIWEI INTELLIGENT TECHNOLOGY CO LTD

Method for quickly forwarding message from PCIE interface to WIA interface

The invention relates to the technical field of industrial communication networks, in particular to a method for quickly forwarding messages from a PCIE (Peripheral Component Interface Express) interface to a WIA (Wireless Interface Architecture) interface, which comprises the following steps of: simultaneously receiving PCIE messages and WIA-FA protocol messages through a hardware logic circuit; analyzing the message in a data link layer to obtain an analysis result comprising data load and address information; packaging the analysis result according to a pre-configuration rule to generate a data frame conforming to a target protocol; and constructing a sending descriptor for the generated data frame, and sending the data frame from a corresponding interface through a direct memory access mechanism. According to the method, parallel analysis and protocol conversion of messages are achieved through a hardware logic circuit, fast address conversion is achieved through a pre-configured address mapping table, and zero-copy data transmission is achieved through a direct memory access mechanism. According to the invention, software protocol stack processing is replaced by a hardware processing flow, so that communication delay and CPU resource consumption are effectively reduced.
Owner:BONCHREE (SHANGHAI) COMMUNICATION CO LTD

Storage and calculation integrated chip architecture, chip stacking and packaging structure and terminal equipment

According to the storage and calculation integrated chip architecture, the chip stacking packaging structure and the terminal equipment, physical tight coupling of storage and calculation is realized through the 3D stacking design of the first storage unit and the logic calculation unit, the problem of physical separation of traditional storage and calculation is solved, the data carrying distance and the cross-chip communication overhead are remarkably reduced, and the communication efficiency is improved. And the chip area and the packaging process complexity are reduced, and the application scenarios of end-side artificial intelligence reasoning and mass data reading and writing are supported. According to the storage and calculation integrated chip architecture, the first storage unit and the logic calculation unit are directly interconnected through the vertical interconnection technology, the data access delay is reduced, the requirement of end side equipment for low delay is met, and the energy efficiency ratio of the storage and calculation integrated chip architecture is remarkably reduced.
Owner:SHANGHAI GUANGYU XINCHEN TECHNOLOGY CO LTD

10T-SRAM in-memory multiplication unit, in-memory operation unit and circuit

The invention discloses a 10T-SRAM (10 T-Static Random Access Memory) in-memory multiplication unit, an in-memory operation unit and a circuit, and relates to the technical field of integrated circuit design. The invention provides a 10T-SRAM in-memory multiplication unit consisting of a 6T-SRAM part and an MAC (Media Access Control) calculation part consisting of four NMOS (N-channel Metal Oxide Semiconductor) tubes, on one hand, a 1bit weight is stored in the 6T-SRAM part, on the other hand, a 4bit unsigned number is divided into high 2bit input X [3: 2] and low 2bit input X [1: 0], and the high 2bit input X [3: 2] and the low 2bit input X [1: 0] are respectively input into the MAC calculation part, so that the multiplication of the 4bit unsigned number and the 1bit weight is realized in an in-memory calculation mode, and the calculation efficiency is improved. And the voltage drop is reflected through bit lines RBLN and RBLP. The in-memory multiplication circuit solves the problem that the existing in-memory multiplication circuit has few input bits and cannot meet the calculation requirement.
Owner:ANHUI UNIV

PCIe (Peripheral Component Interconnect Express) switching chip for realizing calculation in transmission, switch with PCIe switching chip and communication system

The invention discloses a PCIe (Peripheral Component Interconnect Express) switching chip for realizing calculation in transmission, a switch with the PCIe switching chip and a communication system with the PCIe switching chip, the switching chip comprises a physical layer, a data link layer, a transaction layer and a calculation module, and is used for responding to the condition that a specific prefix of a transaction layer data packet is analyzed by the transaction layer and carries a preset calculation instruction; if yes, the calculation module analyzes the calculation instruction analyzed by the transaction layer so as to extract a calculation parameter field in the calculation instruction; the calculation module generates and executes a task of calculating the data in transmission according to the calculation parameter field, and calculation does not need to be started after all the data to be calculated are in place; firstly, part of the currently received data to be calculated is calculated to obtain an intermediate calculation result; and aggregating the subsequent newly received data to be operated and the intermediate operation result. According to the invention, by inserting the programmable calculation path into the PCIe switch, an innovative mechanism of'transmission while calculation 'is realized, the system delay is reduced, and the bandwidth occupation is reduced.
Owner:SHANGHAI XINLIJI SEMICON CO LTD

Communication method between graphics processors, product, equipment and medium

The invention discloses a communication method between graphics processors, a product, equipment and a medium, relates to the technical field of high-performance calculation and artificial intelligence acceleration, and is applied to a stand-alone system comprising a plurality of graphics processors and a central photoelectric hybrid switching chip constructed based on interconnection of an electric switching matrix and an optical switching matrix. The graphics processor is connected with the central photoelectric hybrid switching chip through an optical link and an electric link; the method comprises the following steps: performing data classification on to-be-transmitted data of a source graphics processor; if the to-be-transmitted data is a control flow, the source graphics processor is controlled to send the to-be-transmitted data to an electric switching matrix through an electric link, and then the to-be-transmitted data is routed to a target graphics processor in the graphics processors; and if the to-be-transmitted data is a data stream, the source graphics processor is controlled to send the to-be-transmitted data to the optical switching matrix through the optical link, and then the to-be-transmitted data is routed to the target graphics processor. Interconnection between graphics processors is optimized to improve communication efficiency between graphics processors.
Owner:SHANDONG YUNHAI GUOCHUANG CLOUD COMPUTING EQUIP IND INNOVATION CENT CO LTD

CXL switching chip of multi-core-grain architecture and cross-core-grain routing switching method thereof

The invention belongs to the field of PCIe (Peripheral Component Interconnect Express) chips, and relates to a CXL switching chip of a multi-core-grain architecture and a cross-core-grain routing switching method thereof, and a plurality of CXL switching core grains with uniform specifications are packaged and interconnected to form the CXL switching chip; each CXL switching core particle comprises a plurality of ports, each port can be configured to be in a CXL mode and a D2D mode, and the port of each CXL switching core particle at least comprises a physical layer, a CXL link layer, a CXL transaction layer, a D2D routing layer, a bypass and a multipath selection module; the CXL switching chip further comprises a multi-VC on-chip CXL switching network, and the multi-VC on-chip CXL switching network comprises at least one order-preserving switching VC and one out-of-order switching VC. According to the method, CXL switching chips of various specifications can be flexibly constructed, the design complexity and verification difficulty of the large-scale or multi-port CXL switching chips are effectively reduced, and the production yield of the chips is improved.
Owner:BEIJING SHUDU INFORMATION TECH CO LTD