Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

2175results about "Digital computer details" patented technology

Image processing apparatus and method, and storage medium

A pattern which does not appear at a flat portion in normal binarization processing is set as a code pattern, and a code formed from this pattern is attached. At this time, code attachment with little degradation in image quality is implemented by selecting an unnoticeable pattern.
Owner:CANON KK

Efficient remote pointer sharing for enhanced access to key-value stores

A method to share remote DMA (RDMA) pointers to a key-value store among a plurality of clients. The method allocates a shared memory and accesses the key-value store with a key from a client and receives an information from the key-value store. The method further generates a RDMA pointer from the information, maps the key to a location in the shared memory, and generates a RDMA pointer record at the location. The method further stores the RDMA pointer and the key in the RDMA pointer record and shares the RDMA pointer record among the plurality of clients.
Owner:IBM CORP

Ai-based energy edge platforms, systems, and methods

In some embodiments, a configured artificial intelligence system includes a plurality of intelligence models; a scoring system configured to generate know-your-model scores that quantify suitability for specific tasks of each model; a model execution system configured to provide standardized execution environment for the plurality of intelligence models; a training and reinforcement system configured to monitor outcomes relating to decisions or predictions made by the plurality of intelligence models and use outcome data as feedback to reinforce model performance; and a governance and analysis system configured to ensure model operations comply with governance standards. The intelligence controller may be configured to receive task requests, analyze task complexity, decompose tasks into manageable subtasks, and dynamically select appropriate models from the plurality of intelligence models to execute each subtask based on model suitability and performance characteristics.
Owner:STRONG FORCE EE PORTFOLIO 2022 LLC

Parallel computing method and device, electronic equipment and storage medium

The invention provides a parallel computing method and device, electronic equipment and a storage medium, and relates to the technical field of parallel computing, and the method comprises the steps: carrying out the first protocol operation of a target tensor based on each computing core in each stream processor cluster, and generating a data block containing the computing result of each computing core; writing a data block generated by each stream processor cluster into a shared cache; under the condition that each stream processor cluster completes the first protocol operation, reading a data block written by each stream processor cluster from the shared cache; and executing a second protocol operation on the data block read from the shared cache to generate a calculation result of the target tensor. According to the method and device provided by the invention, the parallel architecture and memory access characteristics of the artificial intelligence chip can be better matched, the unnecessary calculation delay and synchronization overhead of the cross-flow processor cluster in the parallel calculation process are reduced, the bandwidth utilization rate of the shared cache is improved, and the overall performance and calculation efficiency of parallel calculation are remarkably improved.
Owner:SHANGHAI BIREN TECH CO LTD

Extension of network control system into public cloud

Some embodiments provide a method for a first data compute node (DCN) operating in a public datacenter. The method receives an encryption rule from a centralized network controller. The method determines that the network encryption rule requires encryption of packets between second and third DCNs operating in the public datacenter. The method requests a first key from a secure key storage. Upon receipt of the first key, the method uses the first key and additional parameters to generate second and third keys. The method distributes the second key to the second DCN and the third key to the third DCN in the public datacenter.
Owner:VMWARE INC

Coordinating processing tasks between one-dimensional processing engines and two-dimensional processing engines

In various examples, systems and methods are disclosed relating to coordinating and synchronizing the actions of different types of processors with low latency. Different types of processors may perform better at different types of tasks. By coordinating the processing of a one-dimensional processor such as a vector processing unit (VPU) and the processing of a two-dimensional processor such as a pixel processing engine (PPE), an overall speed of task completion can be improved.
Owner:NVIDIA CORP

Heterogeneous computing low-delay communication method and system

The invention relates to the technical field of computers, discloses a heterogeneous computing low-delay communication method and system, and aims to solve the problem of high delay caused by high communication protocol overhead, lack of dynamic scheduling collaboration, memory migration redundancy and non-uniform cross-node communication abstraction in existing heterogeneous computing. The method comprises the following steps: receiving a task scheduling request and analyzing a task dependency graph; tasks are dynamically allocated based on node loads and link states; rDMA, NVLink or PCIe straight-through protocols are adaptively selected according to node types to establish communication channels; zero-copy data exchange is realized through a shared memory mapping buffer area; hardware timestamps are utilized to synchronize feedback delays with PTP to optimize scheduling. The system comprises a heterogeneous computing node cluster, a unified communication scheduling controller, a low-delay communication protocol stack, a shared memory mapping buffer area and a communication delay sensing task distributor. According to the scheme, the communication delay is remarkably reduced, and the throughput and the task execution efficiency are improved.
Owner:BEIJING TOPMOO TECH

Tracking data center build health

Skills and skills metadata may be used to define a process for building a data center. Skills of one service may depend on skills corresponding to the same or different service. A dependency graph may be generated based on these dependencies. The graph may specify an order by which orchestration operations are to be performed to build the services, thereby building the data center. During execution of the process for building the data center, health states corresponding to the skills may be tracked (based at least in part on alarms and / or namespaces associated with the skills). When an unhealthy skill is identified, the system may traverse the dependency graph to identify a root cause (e.g., failed operations corresponding to a skill on which the unhealthy skill directly / indirectly depends). A notification and / or various options may be provided to address the unhealthy state of one or both skills.
Owner:ORACLE INT CORP

AI load-oriented network topology awareness and communication optimization method, device and related system

The invention relates to the technical field of network communication, and discloses an AI load-oriented network topology awareness and communication optimization method, an AI load-oriented network topology awareness and communication optimization device and a related system. The method comprises the following steps: acquiring gradient transmission delay data among computing nodes in a distributed AI training system, and constructing a gradient distribution table of gradient tensors and RDMA intelligent network card queue resources; performing aggregation operation on the gradient tensor in the gradient distribution table on a programmable data plane of the RDMA intelligent network card to obtain pre-aggregated gradient data; calculating a transmission efficiency index of the pre-polymerized gradient data and solving to obtain a bandwidth allocation coefficient; and adjusting an MTU frame parameter, a WQE queue depth and a credit window of the RDMA intelligent network card according to the bandwidth allocation coefficient, and generating a hybrid execution scheme. According to the method, intelligent distribution of the gradient tensor between hardware unloading and software processing is achieved, and the performance close to the optimal performance is obtained under the condition that resources are limited.
Owner:GUANGZHOU BAOYUN INFORMATION TECH CO LTD

LLM-based network troubleshooting using expert-curated recipes

In one implementation, a device receives an input request for a large language model-based network troubleshooting agent regarding an issue in a network. The large language model-based network troubleshooting agent performs a lookup of a recipe based on the input request, wherein the recipe comprises contextual information for the issue. The device generates, by the large language model-based network troubleshooting agent, a prompt for a large language model based on the input request and on the recipe. The device provides, by the large language model-based network troubleshooting agent, the prompt to the large language model to troubleshoot the issue in the network.
Owner:CISCO TECHNOLOGY INC

Pulse neural network hardware accelerator and data processing method

The invention discloses a pulse neural network hardware accelerator and a data processing method, and the accelerator is characterized in that a low-power-consumption three-stage pipeline CPU module is used for receiving input data, scheduling an SNN network acceleration instruction, and sending the input data to an asynchronous edge SNN hardware accelerator module through a coprocessor interface; the asynchronous edge SNN hardware accelerator module comprises a pulse data encoding and decoding module, L neuromorphic kernels and an on-chip network, the pulse data encoding and decoding module encodes input data into a pulse form and sends the pulse form into the neuromorphic kernels, and the neuromorphic kernels are used for performing calculation based on the data in the pulse form; the network-on-chip is used for communication between the neuromorphic kernels, and the connection between neurons before and after synapses in the neuromorphic kernels is realized by adopting a synaptic cross array. According to the invention, the data processing acceleration performance can be greatly improved.
Owner:WUHAN UNIV +1

Ensuring model context protocol server integrity for artificial intelligence agents

A system detects changes in model context protocol (“MCP”) processes, and performs a remedial action. A server-sent events (“SSE”) bridge sends a request to an MCP server. A first list of resource profiles, including tools and instructions, is received from the MCP server. This is stored and compared against a later list of tools and instructions. When a difference is detected, a user interface can display the difference to an administrator. Security rules cause the SSE bridge to be blocked from sending commands to the impacts tool or MCP server until approved by the administrator.
Owner:AIRIA LLC

Memory system with processor in memory (PIM)

A memory system includes a memory interface including a first sub-channel interface associated with a first plurality of memory banks and a first plurality of processor in memory (PIM) blocks and further including a second sub-channel interface associated with a second plurality of memory banks and a second plurality of PIM blocks. During a particular mode of operation associated with the memory system, the first sub-channel interface is configured to communicate, with a host device, one or more memory access commands associated with the first plurality of memory banks. During the particular mode of operation, the second plurality of PIM blocks are configured to perform, concurrently with communication of the one or more memory access commands, one or more PIM operations associated with the second plurality of memory banks. The second sub-channel interface is configured to be disabled during the particular mode of operation.
Owner:QUALCOMM INC

Data transmission method and device, heterogeneous system, and coherent interconnect processing component

The present application relates to the technical field of data storage, and specifically discloses a data transmission method and device, a heterogeneous system, and a coherent interconnect processing component. The coherent interconnect processing component installed on a device enables corresponding coherent interconnect interfaces on the basis of the number of other devices to be interconnected in a heterogeneous system where said device is located, initializes physical communication links to obtain memory interconnect parameters, establishes memory coherent interconnect communication links between the devices on the basis of the memory interconnect parameters, allocates corresponding cache spaces from said device, and on the basis of the memory coherent interconnect communication links and the cache spaces, performs cache coherent transaction processing of said device and the other devices, so as to implement memory coherent interconnect transmission requests between said device and the other devices. Therefore, topologies can be dynamically sensed, and memory coherent interconnect communication links between devices in a heterogeneous system can be automatically established, thereby mitigating the problem that different device interconnect protocols need to be managed in the heterogeneous system, achieving inter-device cache coherence, and further improving the access performance.
Owner:LANGCHAO ELECTRONIC INFORMATION IND CO LTD

Data processing unit, distributed system and chip

The invention relates to a data processing unit, a distributed system and a chip, and the method comprises the steps: a first RDMA engine receives a first network message sent by a previous network node, analyzes the first network message to obtain first data, writes the first data into a first data buffer region, and reports a first receiving notification to a first processor in a reduction stage; the first processor reads the first data from the first data buffer area according to the first receiving notification, reads the second data from the local GPU memory, performs reduction calculation based on the first data and the second data to obtain third data, writes the third data into the first data buffer area, and issues a first sending notification to the first RDMA engine; the first RDMA engine reads third data from the first data buffer area according to the first sending notification, generates a second network message according to the third data, and sends the second network message to the next network node; according to the invention, the computing power of the network node and the utilization rate of the bandwidth can be improved.
Owner:SHENZHEN JAGUAR MICROSYSTEMS CO LTD

Satellite-borne high-performance calculation module and system based on SpaceVPX standard and control method

The invention discloses a spaceborne high-performance calculation module and system based on a SpaceVPX standard and a control method. The calculation module comprises a board card following the SpaceVPX standard, and a core processor module, a storage module, a communication interface module, a monitoring module, a power management module and a VPX connector which are integrated on the board card. The core processor module adopts an anti-radiation reinforced high-performance processor; the storage module comprises a DDR4, a NORFLASH and an SSD (Solid State Disk); the communication interface module comprises an Ethernet, an SRIO high-speed interface and a debugging interface; the monitoring module collects state information through an I2C bus; the power management module provides stable power supply; the VPX connector leads out an SRIO bus, an ETH bus and an I2C bus, and star interconnection with other units on a satellite is achieved. Through standardized and modular design, the problems that an existing satellite-borne computing architecture is fixed, poor in expansibility and insufficient in reliability are solved, and high-performance, high-reliability and flexibly-combined satellite-borne data processing capacity is achieved.
Owner:XIAN MICROELECTRONICS TECH INST

Embedded system based on NVME oF RDMA

The invention provides an embedded system based on NVME oF RDMA. The system comprises an initiating end, a target end and an NVME-oF Target unit, wherein the initiating end is used for initiating a data reading request or a data writing request to a storage target of the target end; the target end is used for executing corresponding data operation according to the request of the initiating end; the initiating end is connected with the target end through an optical fiber; the NVMe-oF Target unit is arranged in the target end, and the NVMe-oF Target unit is arranged in the target end; according to the system, a software and hardware collaborative storage access architecture is constructed, and a kernel protocol stack bypass mechanism is adopted, so that the processing delay in a data transmission path is effectively reduced, the real-time requirement in an embedded scene is met, and the high-delay bottleneck of data transmission on a standard Linux platform is solved.
Owner:NAT UNIV OF DEFENSE TECH

GPU cluster, redundancy optimization method of GPU cluster, electronic equipment, storage medium and computer program product

The invention relates to a GPU cluster, a redundancy optimization method of the GPU cluster, electronic equipment, a storage medium and a computer program product, the GPU cluster comprises a plurality of GPU cabinets, and each GPU cabinet comprises a plurality of GPU nodes, a GPU extension frame and an interconnection module; for any GPU cabinet, each GPU node in the GPU cabinet comprises a plurality of physical GPUs, and a GPU expansion frame in the GPU cabinet comprises a plurality of standby virtual GPUs; the interconnection module in the GPU cabinet is used for carrying out GPU interconnection on a plurality of GPU nodes and a plurality of GPU expansion frames in the GPU cabinet; and the plurality of standby virtual GPUs in the GPU extension frame in the GPU cabinet are used for providing redundant computing resources for the plurality of physical GPUs in each GPU node in the GPU cabinet. According to the embodiment of the invention, the GPU card-level redundancy capability can be effectively realized.
Owner:MOORE THREADS TECH CO LTD

Instruction synchronization method and artificial intelligence chip

The invention provides an instruction synchronization method and an artificial intelligence chip, and relates to the technical field of artificial intelligence chips, and the method comprises the steps: after a processing core sends a calculation instruction, sending a barrier instruction with the same memory address, enabling the calculation instruction and the barrier instruction to enter a cache in an order-preserving manner, and guaranteeing that the cache receives the calculation instruction and achieves calculation; therefore, a large amount of memory fence operation is reduced, and the processing time delay of the memory access instruction is reduced. And for the plurality of processing cores bound with the same target barrier identifier, recording the number of synchronized instructions of the plurality of processing cores in real time. When the number of the synchronized instructions is equal to the expected synchronization number, returning a synchronization success message to the plurality of processing cores so as to realize instruction synchronization of the plurality of processing cores; in the process, each processing core does not need to circularly read the atomic accumulation result, so that the occupation of bandwidth resources such as bus bandwidth and direct interconnection bandwidth is greatly reduced, the instruction synchronization overhead is reduced, and the efficiency of the whole system is improved.
Owner:SHANGHAI BIREN TECH CO LTD

Communication optimization method for topological table persistent storage in efficient parallel computing

The invention provides a communication optimization method for topological table persistent storage in efficient parallel computing, belongs to the technical field of storage communication, and aims to avoid pseudo sharing by aligning memory allocation through cache lines, expand local topological coverage through a topological entropy increment driven prefetching mechanism, and improve the reliability of the topological table persistent storage. Boundary processing is accelerated through a pre-calculation period mapping lookup table and a frequency domain transfer function vector, synchronization overhead is optimized through a concurrency control mechanism perceived by a read-write ratio, targeted cache preloading is achieved through stability and jitter degree two-dimensional evaluation, bandwidth consumption is reduced through an increment synchronization mechanism, and the stability of the cache is improved. The communication template is selected or the communication parameters are generated through adaptive conversion by matching the matching degree decision, and the technical problem that the parallel computing performance is reduced due to the fact that the communication overhead is too large in the topological table persistent storage process is solved.
Owner:青岛国实科技集团有限公司

Gesture-based input to a wearable device

In one embodiment, a method includes capturing, by an optical sensor on or in a body of a wearable device, an image in a field of view of the optical sensor and detecting, by the wearable device and from the image, at least a portion of a hand of a wearer of the wearable device. The method further includes detecting, by the wearable device and from the image, a gesture of at least a portion of the hand of the wearer, the gesture including one or more of (1) a motion of the at least a portion of the hand of the wearer (2) an orientation of the at least a portion of the hand of the wearer or (3) an identification of the detected at least a portion of the hand of the wearer; and determining, by the device, an input to the device based on the detected gesture.
Owner:SAMSUNG ELECTRONICS CO LTD

Request broadcast method, multi-chip device, storage medium and computer program product

The present application provides a method of broadcasting a DVM request for a multi-chip device, each chip in the multi-chip device comprising a plurality of dies, each die of the plurality of dies comprising a plurality of DN domains, each of the plurality of DN domains comprising a DN and a plurality of connection cores, the plurality of DNs of each die comprising a master DN, and the plurality of connection cores comprising a plurality of connection cores. The method comprises the steps that a first connection core initiates a first DVM request to a first DN of a first DN domain to which the first connection core belongs, a crystal grain to which the first DN belongs is a first crystal grain, and a chip to which the first crystal grain belongs is a first chip; the first DN broadcasts the first DVM request to other DNs of the first crystal grain except the first DN; the first master DN of the first die broadcasts the first DVM request to master DNs of other dies of the first chip than the first die. The invention further provides a multi-chip device, a computer readable storage medium and a computer program product.
Owner:SANECHIPS TECH CO LTD

Data Processing Method, Switching Board, Data Processing System and Data Processing Apparatus

A data processing method, a switching board, a data processing system and a data processing apparatus are provided. The method includes: receiving a first fault instruction sent by a controller; on the basis of the first fault instruction, acquiring first target data; and transmitting the first target data to a second host system, so as to instruct the second host system to control a second processor to continue to process first data according to the first target data. wherein the second host system is a host system in a normal operating state among a plurality of host systems connected to a CXL switching device, and the second processor is a processor allocated to the second host system by the CXL switching device.
Owner:INSPUR SUZHOU INTELLIGENT TECH CO LTD

Shared memory communication system and method between processing units, electronic equipment and medium

The invention provides a shared memory communication system and method between processing units, electronic equipment and a medium, and the system comprises the following steps: a first processing unit is used for configuring a memory address of a preset communication flag bit in a shared memory as a specified address, and initiating a memory access request for the specified address to a cache module to poll the communication flag bit; and the cache module is used for receiving the memory access request, acquiring the latest data of a specified address from the shared memory under the condition of determining that the access address in the memory access request is the specified address, and returning the latest data to the first processing unit. The first processing unit can obtain the latest state of the communication flag bit in time, so that communication synchronism and data consistency between the processing units are ensured, communication errors and data errors caused by cache inconsistency are avoided, the reliability and stability of shared memory communication are improved, and the service life of the shared memory is prolonged. Meanwhile, the problem that the efficiency is reduced due to the fact that effective cache lines are cleared in an existing solution is solved.
Owner:YUANQIXIN (SHANDONG) SEMICONDUCTOR TECHNOLOGY CO LTD

System and method for facilitating data request management in a network interface controller (NIC)

A network interface controller (NIC) capable of facilitating efficient data request management is provided. The NIC can be equipped with a command queue, a message chopping unit (MCU), and a traffic management logic block. During operation, the command queue can store a command issued via a host interface. The MCU can then determine a type of the command and a length of a response of the command. If the command is a data request, the traffic management logic block can determine whether the length of the response is within a threshold. If the length exceeds the threshold, the traffic management logic block can pace the command such that the response is within the threshold.
Owner:HEWLETT PACKARD ENTERPRISE DEV LP

Class-based queueing for scalable multi-tenant RDMA traffic

Techniques and apparatus for data networking are described. In one example, a method of queuing Remote Direct Memory Access (RDMA) packets includes receiving a first RDMA packet having a first quality-of-service (QoS) data field; based on a value of the first QoS data field, queueing the first RDMA packet in a first queue of a plurality of queues; receiving a second RDMA packet having a second QoS data field; and based on a value of the second QoS data field, queueing the second RDMA packet in a second queue of the plurality of the queues, the second queue being different than the first queue.
Owner:ORACLE INT CORP

Chip system, data transmission method and related equipment

According to the chip system, the data transmission method and the related equipment provided by the embodiment of the invention, the chip system comprises IO core particles and a plurality of calculation core particles, each calculation core particle is configured with a corresponding core particle center, and each calculation core particle is connected with at least two physical interconnection ports in the IO core particles through the core particle center; the core grain center comprises a splitting unit and a merging unit, and the splitting unit is used for distributing data streams sent by the computing core grains to at least one physical interconnection port according to service types; the merging unit is used for recombining the data stream received from the physical interconnection port or directionally forwarding the data stream to the computing core grain, and the data throughput rate and the overall transmission efficiency in a large-scale parallel computing scene are remarkably improved on the premise that the hardware specification of the external IO core grain does not need to be modified.
Owner:SUZHOU YIZHU INTELLIGENT TECH CO LTD

Zero-knowledge-proof-oriented CPU-GPU data processing and transmission optimization system and method

Disclosed in the present invention is a zero-knowledge-proof-oriented CPU-GPU data processing and transmission optimization system, which is used for executing multi-scalar multiplication operations on the basis of elliptic curve cryptography, so as to generate zero-knowledge proofs. The system comprises a CPU and a GPU, wherein a data transmission channel is provided between the CPU and the GPU; the CPU is used for acquiring and storing point sets and scalar sets that are involved in multi-scalar multiplication operations, and using a set parallelized data transmission mechanism to communicate with the GPU; and the GPU executes multi-scalar multiplication computation on the basis of the point sets and the scalar sets, and returns a computation result to the CPU. By means of the coordinated effect of a plurality of technical features, the present invention significantly improves the speed and efficiency of zero-knowledge proof generation and optimizes resource utilization efficiency.
Owner:SHENZHEN INST OF ADVANCED TECH CHINESE ACAD OF SCI

Data loading method and device, storage medium and equipment

The invention relates to a data loading method and device, a storage medium and equipment, the method is used in an intelligent calculation data processing system comprising a host, a CSD and a SmartNIC, and a PCIe P2P data path is established between the CSD and the SmartNIC; the method comprises the steps that a host generates collaborative descriptors and respectively issues the collaborative descriptors to a CSD and a SmartNIC; the CSD reads source data from a storage medium according to source address information in the collaborative descriptor, executes a processing operation corresponding to the first operator identifier to obtain intermediate data, and directly transmits the intermediate data to a storage area of the SmartNIC through a PCIe P2P data path; and the SmartNIC executes a processing operation corresponding to the second operator identifier on the intermediate data to obtain target data, and transmits the target data to the computing device according to the target address information in the collaborative descriptor.
Owner:GUANGDONG UCAP INTERNET INFORMATION TECH