Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

3318results about "Digital computer details" patented technology

Image processing apparatus and method, and storage medium

A pattern which does not appear at a flat portion in normal binarization processing is set as a code pattern, and a code formed from this pattern is attached. At this time, code attachment with little degradation in image quality is implemented by selecting an unnoticeable pattern.
Owner:CANON KK

Efficient remote pointer sharing for enhanced access to key-value stores

A method to share remote DMA (RDMA) pointers to a key-value store among a plurality of clients. The method allocates a shared memory and accesses the key-value store with a key from a client and receives an information from the key-value store. The method further generates a RDMA pointer from the information, maps the key to a location in the shared memory, and generates a RDMA pointer record at the location. The method further stores the RDMA pointer and the key in the RDMA pointer record and shares the RDMA pointer record among the plurality of clients.
Owner:IBM CORP

Ai-based energy edge platforms, systems, and methods

In some embodiments, a configured artificial intelligence system includes a plurality of intelligence models; a scoring system configured to generate know-your-model scores that quantify suitability for specific tasks of each model; a model execution system configured to provide standardized execution environment for the plurality of intelligence models; a training and reinforcement system configured to monitor outcomes relating to decisions or predictions made by the plurality of intelligence models and use outcome data as feedback to reinforce model performance; and a governance and analysis system configured to ensure model operations comply with governance standards. The intelligence controller may be configured to receive task requests, analyze task complexity, decompose tasks into manageable subtasks, and dynamically select appropriate models from the plurality of intelligence models to execute each subtask based on model suitability and performance characteristics.
Owner:STRONG FORCE EE PORTFOLIO 2022 LLC

Lossless and lossy automatic hardware compression in graphics-to-graphics network links

An apparatus to facilitate lossless and lossy automatic hardware compression in graphics-to-graphics network links is disclosed. The apparatus includes compressor / decompressor circuitry (CDC) integrated with physical layer (PHY) intellectual property (IP) hardware circuitry for a graphics processor unit (GPU)-to-GPU communication link communicably coupling a first GPU to one or more other GPUs, the CDC to: receive a data message from the first GPU, wherein the data message is in an uncompressed format; determine that a compression process is to be applied to the data message; apply the compression process to the data message to generate a compressed data message; and cause a GPU link IP hardware circuitry that comprises the PHY IP hardware circuitry to transmit the compressed data message over the GPU-to-GPU communication link.
Owner:INTEL CORP

Remote memory access systems and methods

The present disclosure relates to systems and methods remote memory access between systems. In particular, some implementations relate to remote memory access using data processing units that can reduce loads on central processing units or other system components. Some implementations utilize scheduling algorithms to optimize memory transfers. Some implementations relate to data processing unit hardware that includes programmable logic, which can be configured for scheduling, data processing, and the like.
Owner:THE ALIGNED CO

Data processing chip based on heterogeneous computing architecture

The invention discloses a data processing chip based on a heterogeneous computing architecture, and relates to the technical field. The data processing chip based on the heterogeneous computing architecture comprises a data processing unit for receiving and preprocessing data to be transmitted; the clustering compression engine is used for carrying out clustering compression processing on the preprocessed data to be transmitted; the modulation and coding control unit is used for carrying out modulation and coding processing on the data to be transmitted after clustering compression processing; the redundant coding unit is used for carrying out redundant coding processing on the modulated and coded data to be transmitted; and outputting the to-be-transmitted data subjected to redundant coding processing to a protocol stack and interface unit of external equipment according to a set communication protocol. Through collaborative design of a clustering compression engine in a chip and a modulation coding control unit, clustering compression aiming at data structure characteristics and modulation strategy matching in a dynamic channel state are realized; the problems of fixed compression and modulation modes and lack of cooperation in the prior art are solved.
Owner:SHENZHEN JINCHAO CLOUD CONTROL TECH CO LTD

Intelligent edge computing cooperative processing system based on integrated circuit

The invention relates to the technical field of edge computing, and discloses an intelligent edge computing co-processing system based on an integrated circuit, which comprises a heterogeneous computing cluster module, a hardware computing unit set, a multi-core control processor based on RISC-V, a programmable pulse tensor computing array and a reconfigurable engine oriented to streaming processing, according to the intelligent edge computing cooperative processing system based on the integrated circuit, zero-delay data exchange is realized through silicon intermediate layer integration of the heterogeneous computing cluster module and a snakelike data channel, and traditional bus arbitration delay is eliminated through real-time operation code analysis and optimal optical communication path mapping of the hardware task routing matrix (TRF) module; the priority path distribution of the task scheduling subsystem is synchronously coordinated with the time-sensitive task, so that the collaborative efficiency of the computing unit is improved in multiple dimensions, the non-blocking transmission of the high-priority task is ensured, and the effect of enhancing the overall processing capability and response speed of the intelligent edge computing is achieved.
Owner:QIQIHAR QISAN MACHINE TOOL

Parallel computing method and device, electronic equipment and storage medium

The invention provides a parallel computing method and device, electronic equipment and a storage medium, and relates to the technical field of parallel computing, and the method comprises the steps: carrying out the first protocol operation of a target tensor based on each computing core in each stream processor cluster, and generating a data block containing the computing result of each computing core; writing a data block generated by each stream processor cluster into a shared cache; under the condition that each stream processor cluster completes the first protocol operation, reading a data block written by each stream processor cluster from the shared cache; and executing a second protocol operation on the data block read from the shared cache to generate a calculation result of the target tensor. According to the method and device provided by the invention, the parallel architecture and memory access characteristics of the artificial intelligence chip can be better matched, the unnecessary calculation delay and synchronization overhead of the cross-flow processor cluster in the parallel calculation process are reduced, the bandwidth utilization rate of the shared cache is improved, and the overall performance and calculation efficiency of parallel calculation are remarkably improved.
Owner:SHANGHAI BIREN TECH CO LTD

Data transmission method based on remote direct memory access, computer equipment and medium

The invention relates to the technical field of computers and provides a data transmission method based on remote direct memory access, computer equipment and a medium. The method comprises the following steps: when a sending side sends a plurality of request messages associated with a remote direct memory access transmission request, respectively writing a queue index value of a first queue at a receiving side into respective extension heads of the request messages; on the receiving side, based on the queue index value of the first queue in the extension header of the received request message, the message content is written into the position corresponding to the queue index value in the first queue, and the effective identifier of the corresponding position is set to be effective; and on the receiving side, when the effective identifier of the first position in the first queue is effective, reading contents in the positions with the effective identifiers being effective in the first queue one by one according to the sequence of the first queue, so as to generate a plurality of response messages. Therefore, the problem of out-of-order receiving is solved, and resources are saved.
Owner:ZHUHAI XINGYUN ZHILIAN TECH CO LTD

GPU BOX, server, interconnection system and high-speed interconnection and data interaction method

The embodiment of the invention discloses a GPU BOX, a server, an interconnection system and a high-speed interconnection and data interaction method. The GPU BOX high-speed interconnection method is characterized in that the GPU BOX comprises a GPU and a control mainboard, an internal PCIe switch, an internal storage device, a first DMA controller and a system controller are integrated in the control mainboard, and a second DMA controller is integrated in the GPU; the system controller receives the reading instruction and performs data read-write initialization on the first DMA controller and the second DMA controller based on the reading instruction; the first DMA controller after data read-write initialization reads cache data from the internal storage device, and transmits the cache data to the GPU via the internal PCIe switch through the established GDS connection; wherein the GDS connection between the GPU and the internal storage device is established through the system controller, the second DMA controller, the internal PCIe switch and the first DMA controller. According to the method, high-speed direct data transmission of the GPU BOX can be achieved, CPU data transfer is not needed, and the data transmission efficiency is greatly improved.
Owner:BEIJING RONGXIN ZHIYUAN TECHNOLOGY CO LTD

GPU cluster connection method and device, switch and storage medium

The invention provides a GPU cluster connection method and device, a switch and a storage medium, relates to the technical field of communication, and is used for solving the problems that a GPU cluster of a single-layer fat tree architecture is relatively small in scale and a GPU cluster of a two-layer fat tree architecture is relatively low in efficiency in related technologies. The GPU cluster connection method comprises the steps that Leaf switches with the same serial number of different server clusters are interconnected, and each server cluster comprises a plurality of Leaf switches and a plurality of servers; each server comprises a plurality of graphic processing units (GPUs); and each switch is in communication connection with the GPU in the server cluster. According to the method, a GPU cluster with a larger scale and higher efficiency can be constructed in a transverse expansion mode.
Owner:CHINA UNITED NETWORK COMM GRP CO LTD +1

Ensemble and averaging for deep neural network inference with non-volatile memory arrays

To reduce noise for compute in memory vector-matrix multiplications using non-volatile memory, such in large language models or other neural networks, the matrix of values is independently programmed in multiple arrays. The input vector is then independently applied to the multiple arrays to generate a corresponding set of multiple results. The independent results can then be combined by ensemble algorithms, as an averaging, to generate the output. The combined output can then be used as input to a subsequent layer of weights, which can again be independently written in to multiple arrays.
Owner:WESTERN DIGITAL TECHNOLOGIES INC

Extension of network control system into public cloud

Some embodiments provide a method for a first data compute node (DCN) operating in a public datacenter. The method receives an encryption rule from a centralized network controller. The method determines that the network encryption rule requires encryption of packets between second and third DCNs operating in the public datacenter. The method requests a first key from a secure key storage. Upon receipt of the first key, the method uses the first key and additional parameters to generate second and third keys. The method distributes the second key to the second DCN and the third key to the third DCN in the public datacenter.
Owner:VMWARE INC

Adaptive dynamic routing method and system for NoC (Network on Chip)

The invention discloses a self-adaptive dynamic routing method and system for an NoC (Network on Chip), and relates to the technical field of routing management. Comprising the steps that the local congestion degree of each routing node is collected through a communication link, the global congestion degree is generated through a weighted average mechanism, each routing node integrates delay counters in the east direction, the west direction, the south direction and the north direction according to a local router, the delay counters are used for monitoring the time interval when adjacent data packets leave a buffer area in real time, and the global congestion degree is generated; taking a time interval of leaving a buffer area of adjacent data packets as output delay, spreading the output delay to neighbor routing nodes, calculating local congestion degree and global congestion degree of the routing nodes, generating all possible shortest paths and a candidate direction set corresponding to output directions according to addresses of a source routing node and a destination routing node, and sending the candidate direction set to the destination routing node; screening an optimal path through the local congestion degree and the global congestion degree in sequence; and when a plurality of equivalent paths exist, the routing direction is randomly selected to balance the network load.
Owner:SHANDONG INSPUR SCI RES INST CO LTD

Enhanced Harvard Architecture Reduced Instruction Set Computer (RISC) with Debug Mode Access of Instruction Memory within a Unified Memory Space

A Harvard-architecture computer simultaneously reads instructions using an instruction bus and accesses data over a separate data bus. A debug module has an external interface to an external debugger that initiates a debugging session by setting a debug bit in a debug mode register. When the debug bit is set, a first mux to disconnect an instruction pointer and an instruction buffer from the instruction bus and instead connects the data bus to the instruction bus. A second mux disconnects a load / store unit in an execution core from the instruction bus and instead connects the debug module to the instruction bus. The external debugger sees a unified memory space and writes addresses to the debug module that are sent through the second mux to access the data memory, and instruction addresses are sent through the second mux and the first mux to the instruction memory to read or write instructions.
Owner:HIGH TECH TECH LTD

Coordinating processing tasks between one-dimensional processing engines and two-dimensional processing engines

In various examples, systems and methods are disclosed relating to coordinating and synchronizing the actions of different types of processors with low latency. Different types of processors may perform better at different types of tasks. By coordinating the processing of a one-dimensional processor such as a vector processing unit (VPU) and the processing of a two-dimensional processor such as a pixel processing engine (PPE), an overall speed of task completion can be improved.
Owner:NVIDIA CORP

Three-dimensional stacked memory controller supporting transverse expansion

The invention relates to the technical field of integrated circuit design, in particular to a three-dimensional stacked memory controller supporting lateral expansion, which comprises a memory control logic unit connected with an interface expansion unit and a mode configuration unit, the interface expansion unit and the mode configuration unit are connected with a near memory calculation unit, the near memory calculation unit is connected with the memory control logic unit, and the memory control logic unit is connected with the interface expansion unit. And the on-chip cache unit is connected with the memory control logic unit. The on-chip cache unit is used for storing high-frequency access data; the memory control logic unit is used for forming a hierarchical memory architecture; the interface expansion unit is used for transversely accessing the low-speed memory; the mode configuration unit enables on-chip high-speed cache, three-dimensional stacked memory and low-speed memory access to be independently Bypass; the near memory calculation unit is used for providing transposition, stepping, compression and direct memory calculation functions. Therefore, the problems that in the prior art, a three-dimensional stacked memory and a low-speed memory cannot be efficiently organized and matched, development is complex, and a hierarchical framework is missing are solved.
Owner:WUXI CORE FIELD MICROELECTRONICS CO LTD

Heterogeneous computing low-delay communication method and system

The invention relates to the technical field of computers, discloses a heterogeneous computing low-delay communication method and system, and aims to solve the problem of high delay caused by high communication protocol overhead, lack of dynamic scheduling collaboration, memory migration redundancy and non-uniform cross-node communication abstraction in existing heterogeneous computing. The method comprises the following steps: receiving a task scheduling request and analyzing a task dependency graph; tasks are dynamically allocated based on node loads and link states; rDMA, NVLink or PCIe straight-through protocols are adaptively selected according to node types to establish communication channels; zero-copy data exchange is realized through a shared memory mapping buffer area; hardware timestamps are utilized to synchronize feedback delays with PTP to optimize scheduling. The system comprises a heterogeneous computing node cluster, a unified communication scheduling controller, a low-delay communication protocol stack, a shared memory mapping buffer area and a communication delay sensing task distributor. According to the scheme, the communication delay is remarkably reduced, and the throughput and the task execution efficiency are improved.
Owner:BEIJING TOPMOO TECH

Tracking data center build health

Skills and skills metadata may be used to define a process for building a data center. Skills of one service may depend on skills corresponding to the same or different service. A dependency graph may be generated based on these dependencies. The graph may specify an order by which orchestration operations are to be performed to build the services, thereby building the data center. During execution of the process for building the data center, health states corresponding to the skills may be tracked (based at least in part on alarms and / or namespaces associated with the skills). When an unhealthy skill is identified, the system may traverse the dependency graph to identify a root cause (e.g., failed operations corresponding to a skill on which the unhealthy skill directly / indirectly depends). A notification and / or various options may be provided to address the unhealthy state of one or both skills.
Owner:ORACLE INT CORP

AI load-oriented network topology awareness and communication optimization method, device and related system

The invention relates to the technical field of network communication, and discloses an AI load-oriented network topology awareness and communication optimization method, an AI load-oriented network topology awareness and communication optimization device and a related system. The method comprises the following steps: acquiring gradient transmission delay data among computing nodes in a distributed AI training system, and constructing a gradient distribution table of gradient tensors and RDMA intelligent network card queue resources; performing aggregation operation on the gradient tensor in the gradient distribution table on a programmable data plane of the RDMA intelligent network card to obtain pre-aggregated gradient data; calculating a transmission efficiency index of the pre-polymerized gradient data and solving to obtain a bandwidth allocation coefficient; and adjusting an MTU frame parameter, a WQE queue depth and a credit window of the RDMA intelligent network card according to the bandwidth allocation coefficient, and generating a hybrid execution scheme. According to the method, intelligent distribution of the gradient tensor between hardware unloading and software processing is achieved, and the performance close to the optimal performance is obtained under the condition that resources are limited.
Owner:GUANGZHOU BAOYUN INFORMATION TECH CO LTD

LLM-based network troubleshooting using expert-curated recipes

In one implementation, a device receives an input request for a large language model-based network troubleshooting agent regarding an issue in a network. The large language model-based network troubleshooting agent performs a lookup of a recipe based on the input request, wherein the recipe comprises contextual information for the issue. The device generates, by the large language model-based network troubleshooting agent, a prompt for a large language model based on the input request and on the recipe. The device provides, by the large language model-based network troubleshooting agent, the prompt to the large language model to troubleshoot the issue in the network.
Owner:CISCO TECHNOLOGY INC

Pulse neural network hardware accelerator and data processing method

The invention discloses a pulse neural network hardware accelerator and a data processing method, and the accelerator is characterized in that a low-power-consumption three-stage pipeline CPU module is used for receiving input data, scheduling an SNN network acceleration instruction, and sending the input data to an asynchronous edge SNN hardware accelerator module through a coprocessor interface; the asynchronous edge SNN hardware accelerator module comprises a pulse data encoding and decoding module, L neuromorphic kernels and an on-chip network, the pulse data encoding and decoding module encodes input data into a pulse form and sends the pulse form into the neuromorphic kernels, and the neuromorphic kernels are used for performing calculation based on the data in the pulse form; the network-on-chip is used for communication between the neuromorphic kernels, and the connection between neurons before and after synapses in the neuromorphic kernels is realized by adopting a synaptic cross array. According to the invention, the data processing acceleration performance can be greatly improved.
Owner:WUHAN UNIV +1

Ensuring model context protocol server integrity for artificial intelligence agents

A system detects changes in model context protocol (“MCP”) processes, and performs a remedial action. A server-sent events (“SSE”) bridge sends a request to an MCP server. A first list of resource profiles, including tools and instructions, is received from the MCP server. This is stored and compared against a later list of tools and instructions. When a difference is detected, a user interface can display the difference to an administrator. Security rules cause the SSE bridge to be blocked from sending commands to the impacts tool or MCP server until approved by the administrator.
Owner:AIRIA LLC

Domain-specific verification tool for cache coherence protocol

The invention relates to a modeling and formalized verification method for cache consistency protocol verification, and the method comprises the steps: building a protocol model which comprises a plurality of processor nodes, directory nodes and a communication mechanism through a structured modeling mode, and constructing an asynchronous message mechanism to simulate disordered concurrent communication; protocol behavior rules such as processor requests, directory responses and network receiving are defined, protocol property invariants are set, semantic constraints such as data consistency, confirmation before writing and sharing state legality are covered, model verification is carried out in a state space traversal mode, breadth-first search and a symmetry recognition mechanism are supported, and the method is suitable for the protocol behavior rules. A protocol error or deadlock state is effectively found, a traceable error path is generated, a verifier can automatically generate source codes through a compiler and operate on a host platform, and rapid verification and result output of a protocol model are achieved. The method has good protocol adaptability and model reusability, and is suitable for formalized verification of various cache consistency protocols.
Owner:SHAOXIN LABORATORY

Memory system with processor in memory (PIM)

A memory system includes a memory interface including a first sub-channel interface associated with a first plurality of memory banks and a first plurality of processor in memory (PIM) blocks and further including a second sub-channel interface associated with a second plurality of memory banks and a second plurality of PIM blocks. During a particular mode of operation associated with the memory system, the first sub-channel interface is configured to communicate, with a host device, one or more memory access commands associated with the first plurality of memory banks. During the particular mode of operation, the second plurality of PIM blocks are configured to perform, concurrently with communication of the one or more memory access commands, one or more PIM operations associated with the second plurality of memory banks. The second sub-channel interface is configured to be disabled during the particular mode of operation.
Owner:QUALCOMM INC

Data transmission method and device, heterogeneous system, and coherent interconnect processing component

The present application relates to the technical field of data storage, and specifically discloses a data transmission method and device, a heterogeneous system, and a coherent interconnect processing component. The coherent interconnect processing component installed on a device enables corresponding coherent interconnect interfaces on the basis of the number of other devices to be interconnected in a heterogeneous system where said device is located, initializes physical communication links to obtain memory interconnect parameters, establishes memory coherent interconnect communication links between the devices on the basis of the memory interconnect parameters, allocates corresponding cache spaces from said device, and on the basis of the memory coherent interconnect communication links and the cache spaces, performs cache coherent transaction processing of said device and the other devices, so as to implement memory coherent interconnect transmission requests between said device and the other devices. Therefore, topologies can be dynamically sensed, and memory coherent interconnect communication links between devices in a heterogeneous system can be automatically established, thereby mitigating the problem that different device interconnect protocols need to be managed in the heterogeneous system, achieving inter-device cache coherence, and further improving the access performance.
Owner:LANGCHAO ELECTRONIC INFORMATION IND CO LTD

AI accelerator integrated circuit chip with integrated cell-based fabric adapter

An integrated circuit formed on (i) a single semiconductor die or (ii) a plurality semiconductor dies that are integrated into a single package. The integrated circuit may include a communication interface including a serializer / deserializer (SerDes) interface; a fabric adapter communicatively coupled to the communication interface; a plurality of inference engine clusters, each inference engine cluster including a respective memory element and / or memory interface; and a data interconnect communicatively coupling each respective memory element and / or memory interfaces of the plurality of inference engine clusters to the fabric adapter. The fabric adapter may be configured to facilitate remote direct memory access (RDMA) read and write services and / or datagram communication over a cell-based switch fabric to and from the respective memory elements and / or memory interfaces of the plurality of inference engine clusters via the data interconnect.
Owner:TENSORDYNE INC

Controlling remote electronic device with wearable electronic device

In one embodiment, a method includes displaying functions associated with a remote device at a wearable computing device, wherein the remote device is paired with the wearable computing device, receiving a user input from a user at the wearable computing device, determining intended operations of one or more of the functions by the user based on the user input, sending instructions for executing the intended operations of the one or more of the functions associated with the remote device to the remote device, and displaying execution results of the intended operations of the one or more of the functions associated with the remote device at the wearable computing device.
Owner:SAMSUNG ELECTRONICS CO LTD

High-performance CPU-GPU (Central Processing Unit-Graphics Processing Unit) coprocessing architecture of audio frequency integrated signal processor

The invention discloses a high-performance CPU-GPU (Central Processing Unit-Graphics Processing Unit) cooperative processing architecture of an audio frequency integrated signal processor, belonging to the technical field of audio frequency integrated signal processing, the architecture is based on a dynamic priority scheduling model and realizes efficient collaboration of a CPU and a GPU through hierarchical resource management and protocol level optimization, a hardware layer adopts a multi-GPU cluster and a distributed storage node, and the CPU-GPU cooperative processing architecture has the advantages that the multi-GPU cluster and the distributed storage node are integrated; high-concurrency task processing is supported; the transmission layer fuses RapidIO and an Ethernet protocol, and adapts to a short frame control signal and a long packet data stream respectively; and the application layer calculates task resource demands through dynamic priority weights, and ensures low time delay of key tasks in combination with a normalized allocation algorithm. According to the architecture, in an audio signal processing scene, the task scheduling efficiency is improved by 40%, the short frame transmission delay is as low as 0.5, the throughput of a long data stream reaches 100 Gbps, and the requirements for real-time performance and calculation precision in a complex acoustic environment can be met.
Owner:CHINA SHIP DEV & DESIGN CENT +1