Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

379 results about "Pipeline (computing)" patented technology

In computing, a pipeline, also known as a data pipeline, is a set of data processing elements connected in series, where the output of one element is the input of the next one. The elements of a pipeline are often executed in parallel or in time-sliced fashion. Some amount of buffer storage is often inserted between elements.

Federated distributed graph-based computing platform with hardware management

A federated distributed AI reasoning and action platform utilizing decentralized, partially observable hierarchical computing for neuro-symbolic reasoning. It features a federated Distributed Computational Graph (DCG) system integrating core components like pipeline orchestration, transformers, and marketplaces. The platform enables privacy-preserving dynamic resource allocation, intelligent task scheduling, and variable information sharing across diverse computing environments. By coordinating with an AI-based operating system and analyzing performance metrics, environmental conditions, and resource availability, the system optimizes efficiency across AI workloads and decision-making processes. This results in an adaptive, power-efficient, and scalable AI-enabled data processing system capable of handling complex tasks while maintaining peak performance under various operating conditions.
Owner:QOMPLX INC

Hybrid expert model reasoning method based on cooperation of CPU and GPU

The invention discloses a hybrid expert model reasoning method based on cooperation of a CPU (Central Processing Unit) and a GPU (Graphics Processing Unit), and belongs to the field of deep learning. According to the method, a CPU-GPU computing framework of a hybrid expert model is constructed, heterogeneous computing resource loads are effectively balanced, and the hardware utilization rate is remarkably increased; an intelligent cache management mechanism based on dynamic priority scores is provided, high-demand experts are reserved preferentially, and the transmission overhead caused by cache missing is reduced; through pipeline parallel design for separating calculation and transmission tasks, CPU calculation and PCIe transmission are overlapped in the GPU execution period, and delay is effectively hidden. In addition, in combination with a multi-layer expert activation prediction prospective prefetching mechanism, the expert cache hit rate is improved. The method is compatible with hybrid expert models of different scales and structures, and stable and efficient reasoning acceleration is realized on a resource-limited heterogeneous platform.
Owner:PEKING UNIV

Distributed training scheduling and communication optimization method and system of multi-modal large model on domestic computing power platform

The invention discloses a distributed training scheduling and communication optimization method and system of a multi-modal large model on a domestic computing power platform. The method comprises the following steps: virtualizing a heterogeneous computing unit of a preset platform into a virtual device pool, and fusing first-order gradient of a multi-modal sample and Hessian matrix information based on quantitative perception training to generate a sample sensitivity grading atlas; virtual device pool attributes and the sensitivity grading atlas are used as input, an optimal hybrid parallel configuration scheme is automatically generated through a configuration search algorithm, and a parallel combination mode, resource mapping and a high-sensitivity sample scheduling strategy are defined; a distributed training code of an integrated communication optimization strategy is automatically generated according to a configuration scheme, pipeline parallel communication and data parallel gradient synchronization constraint are executed in a topology adjacent equipment subset, and a hierarchical aggregation mechanism is adopted; and dynamically screening a core training set and scheduling a calculation task to complete distributed training. According to the method, efficient cooperative training of the multi-modal large model on the domestic computing power platform is realized.
Owner:GUANGXI POWER GRID CORP

Computing power resource fusion method based on distributed flow pipeline

The invention relates to the technical field of distributed computing and computing power scheduling, in particular to a computing power resource fusion method based on a distributed flow pipeline, which comprises the following steps of: disassembling a user computing power request into a flow pipeline unit for packaging computing logic, an input / output interface and a resource demand label; meanwhile, computing power types, real-time load rates, memory occupancy rates, network round-trip delays, geographic positioning and energy consumption data of cloud edge end nodes are collected, a hierarchical topology network framework is constructed based on the collected data, node computing power available values are calculated, an inter-node transmission cost matrix is generated, a fault probability prediction model is constructed, and a fault probability prediction model is constructed. And finally, analyzing a data dependency relationship of the pipeline unit through a four-dimensional joint decision engine, executing dynamic mapping, and preferentially mapping the high-computing-power demand unit to a GPU cluster node, so as to realize non-interruption reconstruction during pipeline topology operation. And the global computing power resource utilization rate, the task operation efficiency and the service stability are improved.
Owner:LANZHOU YUNFAN ZHILIAN TECH CO LTD +2

Cluster-oriented large model parallel method and device and electronic device

The invention relates to a cluster-oriented large model parallelization method and device and an electronic device.The method is applied to the field of large models.The method comprises the steps that operator information of a cluster-oriented large model in a preset microprocessing batch and a preset parallelization mode is obtained, the operator information comprises operator time information and operator memory information of operators in the cluster-oriented large model; the cluster comprises one or more types of accelerators; determining an initial operator parallel configuration strategy of a plurality of assembly lines in the large model based on the operator information, model memory information required by the large model and a memory extreme value of an accelerator; and performing recursion processing on the initial operator parallel configuration strategy according to a preset load balancing mode to obtain a target operator parallel configuration strategy, and running the cluster-oriented large model based on the target operator parallel configuration strategy. According to the method and the device, the utilization efficiency of chip calculation performance during large model parallel configuration is improved, and high efficiency and wide application range of large model training are realized.
Owner:ZHEJIANG LAB

Case and computing device

The embodiment of the invention provides a case and computing equipment, the case is applied to the computing equipment, the case comprises a shell, the computing equipment comprises a computing module, a power supply module and a pipeline group which are arranged in the shell, the computing module comprises a liquid cooling module and a computing force plate, the liquid cooling module is provided with a cooling flow channel for a cooling medium to flow, and the cooling flow channel is provided with a liquid outlet. The cooling flow channel is communicated with the pipeline group, a containing cavity and a pipeline cavity are defined in the case, the containing cavity is used for containing the calculation module and the power module, and the pipeline cavity is used for containing the pipeline group, so that the calculation module, the power module and the pipeline group are integrated, maintenance is easy, and occupied space of the case is reduced; and the appearance of the case is neater and more attractive.
Owner:CANAAN CREATIVE CO LTD

BLAS3 structured operator accelerated computing system based on Hopper architecture GPU

The invention provides a BLAS3 structured operator accelerated computing system based on a Hopper architecture GPU, and relates to the technical field of computers. The system comprises: a calculation unit discrimination module for determining a calculation unit used by a current operator during operation, and estimating the maximum row dimension upper bound of the current operator in a tensor core execution path; an instruction sensing block parameter determination module dynamically determines the optimal block size and number of the input matrix in real time; the block matrix loading and aligning module divides an input matrix and a matrix to be updated into sub-matrixes by taking the block size as a basic block and completes loading of the corresponding sub-matrixes; the operator kernel function execution module completes shared memory structured parallel loading and storage of a double-precision floating-point number array of a sub-matrix corresponding to the input matrix, and calls a tensor core to carry out multiply-add accumulation calculation; and the assembly line and concurrent scheduling module adds the block calculation tasks into corresponding task sets and performs multi-stream concurrent scheduling on the task sets.
Owner:NORTHEASTERN UNIV CHINA

Machine learning pipelines

Techniques described herein may be implemented in the context of a computing resource service provider. A machine learning (ML) service provides an interface to clients which can be used to create, read, update, and delete ML pipelines. ML pipelines can be converted to a human-readable format, which a server persists. Clients of a machine learning service can start, stop, and resume executions a ML pipeline based on the human-readable format of the ML pipeline.
Owner:AMAZON TECH INC

Multi-modal basic model convolution operation optimization framework method suitable for GPU / DCU

The invention provides a multi-modal basic model convolution operation optimization framework method suitable for a GPU / DCU. The method is used for solving the technical problems that the optimization process of an existing convolution operator optimization method is time-consuming and cross-layer operator fusion is difficult to capture. The method comprises the following steps: receiving a convolution operator through a parameterized input interface; collecting shape features of each convolution operator in the structured parameter set by using a shape analyzer to realize convolution feature analysis; judging whether batch normalization and nonlinear activation function fusion calculation is started or not; a heuristic optimizer dynamically selects a calculation graph optimization strategy based on convolution parameter characteristics; a calculation task is adapted to a physical architecture of the GPU / DCU, so that hardware resources are utilized to the maximum extent; a convolution calculation kernel function oriented to DCU / GPU architecture optimization is generated based on the convolution parameter features; kernel function assembly line execution is achieved through a DCU / GPU hardware task queue; gradient synchronization among multiple computing units is realized by adopting atomic operation at an equipment end. According to the method, the execution efficiency of the convolution operation in the multi-modal basic model can be remarkably improved.
Owner:HENAN POLYTECHNIC

Controlling execution of machine learning models

In an example, an apparatus is described. The apparatus comprises processing circuitry comprising a control module. The control module determines whether a computing device communicatively coupled to the control module is in a specified state for executing a machine learning model controlled by a third party entity. In response to determining that the computing device is in the specified state, the control module is to send, to an attestation module in a data processing pipeline associated with the computing device, an indication that the computing device is in the specified state.
Owner:HEWLETT PACKARD DEVELOPMENT COMPANY LP

Low latency scratch memory path

An apparatus and method for efficiently processing vector memory accesses on an integrated circuit. In various implementations, a computing system includes a processing circuit with multiple compute circuits for executing wavefronts of a parallel data application. Each compute circuit includes a local memory subsystem for accessing data not found in vector register files of the compute circuit. The local memory subsystem includes a first execution pipeline and a second execution pipeline. The second execution pipeline processes vector stack access instructions that access temporary data such as stack data of a function call used by each wavefront that is generated based on the function call. The first execution pipeline processes other types of vector memory access instructions and includes multiple complex pipeline stages not found in the second execution pipeline. Thus, the second execution pipeline has a latency less than the latency of the first execution pipeline.
Owner:ADVANCED MICRO DEVICES INC

Specific file detection baked into machine learning pipelines

A set of features including a first feature and a second feature is received at a server. A subset of the set of features is determined for use in generating a model usable by a device to locally make a malware classification decision. The device has reduced computing resources as compared to computing resources of the server. The subset of the set of features is used to generate the model. The generated model includes the first feature and does not include the second feature. A determination is made, at a time subsequent to the generation of the model, that an updated model should be deployed to the device. An updated model is generated.
Owner:PALO ALTO NETWORKS INC

Intelligent telescoping and tuning method for reinforcement learning driving in big data pipeline

The invention relates to a reinforcement learning driven intelligent telescoping and tuning method in a big data pipeline, and relates to the technical field of big data processing and computing, and the method comprises the steps: (1) data pipeline operation and state monitoring; (2) extracting resource indexes and features; (3) data preprocessing and model initialization; (4) a dynamic adjustment algorithm based on LinUCB is operated; (5) a periodic reset and quick response mechanism; (6) applying and evaluating a resource configuration result; (7) saving and continuously optimizing an adjusting and optimizing result; according to the method, on the basis of a dynamic adjustment strategy of a context tiger machine algorithm, system state characteristics of an Apache NiFi data pipeline are collected in real time, an online learning mechanism is combined, the thread count, memory allocation and CPU core number configuration are optimized step by step, and intelligent scaling and tuning of the data pipeline are achieved on the basis, so that the throughput rate, the resource utilization rate and the load balancing capacity are improved.
Owner:QILU UNIVERSITY OF TECHNOLOGY (SHANDONG ACADEMY OF SCIENCES) +1

Detection and correction device and method based on RISC-V architecture

The invention discloses a detection and correction device and method based on an RISC-V architecture, belongs to the technical field of processor micro-architectures, and aims to solve the technical problem of how to carry out error detection and correction on an instruction execution process, hardware resources, system reliability, instruction storage, a data path and related system performance in the RISC-V architecture. Comprising a computing core, a cross validation module, a dynamic voting module, a lock step execution module and a hardware multiplexing module, the processor is used for comparing and checking the instruction execution results output by the three calculation cores based on a triple modular redundancy technology and a dynamic voting fault-tolerant technology, and determining the instruction execution result with correct voting through a two-out-of-three voting mode; the lock step execution module is used for controlling lock step execution and assembly line rollback operation of the calculation core; and the hardware multiplexing module is used for multiplexing the same RISC-V pipeline to process normal execution, lock step and rollback operations.
Owner:SHANDONG INSPUR SCI RES INST CO LTD

Hashing method, hash calculation system, electronic equipment and storage medium

The invention provides a hash method, a hash calculation system, electronic equipment and a storage medium, and the method comprises the steps: filling data, and obtaining a to-be-compressed message; inputting the to-be-compressed message and the first group of message words into a first pipeline structure for compression to obtain a first compressed message; inputting the first compressed message into a second pipeline structure; and compressing the second compressed message output by the nth parallel computing unit based on the (n + 1) th message word output by the nth parallel computing unit, and generating an (n + 2) th message word based on the (n + 1) th message word for the (n + 2) th parallel computing unit to compress. Thus, when the compression function module carries out compression, the message word extension module carries out extension calculation of a next message word at the same time to form parallel calculation of the compression function and message word extension, the compression function module and / or the parallel calculation unit are / is connected through the register to form pipeline calculation, and the compression efficiency and the throughput rate of the compression process are improved.
Owner:SHANDONG YUNHAI GUOCHUANG CLOUD COMPUTING EQUIP IND INNOVATION CENT CO LTD

Real-time target detection method and system based on RTSP flow and NPU collaborative optimization

A real-time target detection method and system based on RTSP flow and NPU collaborative optimization belong to the technical field of computer vision and artificial intelligence, and are characterized by comprising the following steps: adopting dynamic memory optimization management of a hybrid pipeline architecture, and performing single-time continuous copying through a CPU (Central Processing Unit); through deep integration of innovative technologies such as DMA direct transmission NPU continuous memory pool management, dynamic batch processing scheduling, hybrid assembly line processing, intelligent equipment load balancing, parallel preprocessing optimization and vectorization post-processing, RTSP flow collaborative optimization and the like, the NPU utilization rate is improved to 85% or above, the overall average FPS is improved by 200% or above, the assembly line parallelism degree achieves three times of performance gain, and the production efficiency is greatly improved. The data transmission delay is reduced by 80%, the system stability is remarkably improved, performance degradation is avoided after long-time operation, and the method is suitable for various scenes such as edge calculation and cloud reasoning.
Owner:XIAN KEYWAY TECH

Industrial algorithm model scheduling method and system for industrial big data

The invention relates to the technical field of industrial big data processing and algorithm scheduling, and particularly discloses an industrial algorithm model scheduling method and system for industrial big data, and the method comprises the steps: constructing an algorithm dependency graph of an execution pipeline composed of a plurality of industrial algorithms, bottleneck nodes influencing the overall performance are identified through a critical path analysis algorithm; carrying out operation fusion on algorithm pairs with close dependency relationships; optimal data partitioning strategies are determined for different algorithms, and an elastic execution pipeline is realized; allocating the most matched heterogeneous computing resources for the algorithm operation on the critical path, and adjusting the execution priority of the algorithm operation in real time according to the system load and the operation scene; a fine-grained synchronization system is realized; the industrial algorithm execution efficiency is improved, and the high-performance requirement of an industrial big data processing scene is met.
Owner:ZHONGKE YUZHOU (GUANGDONG) TECHNOLOGY SERVICE CO LTD

Model training method and apparatus based on hybrid parallelism manner, and device

This application discloses a model training method and apparatus based on a hybrid parallelism manner. In this method, a neural network model is divided into a plurality of pipeline stages, and each pipeline stage includes a plurality of sub-stages of the neural network model. Computing nodes corresponding to the plurality of pipeline stages are invoked in a hybrid parallelism manner according to a sequence of sub-stages in the neural network model. When iterative training is performed on a network layer in a corresponding pipeline stage, because sub-stages at same locations in adjacent pipeline stages are consecutive in the neural network model, the computing node does not need to wait for completion of forward propagation of a previous pipeline stage, and can perform forward propagation on the corresponding pipeline stage only after forward propagation of the 1st sub-stage in the previous pipeline stage is completed.
Owner:HUAWEI TECH CO LTD

Graph spatial split

A method for reducing latency and increasing throughput in a reconfigurable computing system includes receiving a compute graph for execution on a reconfigurable dataflow processor comprising a grid of compute units and grid of memory units interconnected with a switching array. The compute graph includes a node specifying an operation on a tensor. The node may be split into multiple nodes that each specify the operation on a distinctive portion of the tensor to produce a first modified compute graph. The first modified compute graph may be executed. In addition, the multiple nodes may be within a single meta-pipeline stage and may be processed in parallel. Furthermore, the compute graph may further comprise a separate node for gathering the distinctive portions of the tensor into a complete tensor, to produce a second modified compute graph.
Owner:SAMBANOVA SYSTEMS INC

Layered self-adaptive full block pre-filling scheduling method and system for large language model reasoning

The invention discloses a hierarchical self-adaptive full block pre-filling scheduling method and system for large language model reasoning, and the method comprises the steps: carrying out the hierarchical portrait analysis of a to-be-served model, and dividing the to-be-served model into partitions with different calculation characteristics according to the calculation intensity and memory access characteristics of each layer; then, making a layering and partitioning strategy based on a partitioning result, allocating a larger partitioning size to a calculation-intensive partition, allocating a smaller partitioning size to a memory bandwidth-intensive partition, and generating a layering and partitioning mapping table; and finally, when the online scheduling is executed, querying the mapping table according to the request processing progress to determine the block target size, and jointly forming a batch processing unit by the decoding task and the pre-filled block with the heterogeneous size under the constraint of the iteration time budget to be executed. According to the method, accurate matching of calculation and bandwidth resources is achieved, the system throughput can be effectively improved, tail delay and fluctuation thereof can be remarkably reduced, bubbles under pipeline parallelism are reduced, and the method is suitable for various attention mechanisms and distributed reasoning architectures.
Owner:ZHEJIANG LAB

Abnormal data real-time filtering method and system based on edge calculation in dynamic environment system

The invention discloses an abnormal data real-time filtering method and system based on edge computing in a dynamic environment system, the method is executed by an edge computing gateway, sliding window weighted average preprocessing is equivalently realized by adopting integer shift operation, and the single computing overhead is controlled within 50 clock cycles; dynamically and adaptively adjusting dynamic reference model parameters based on the ratio of the network load to the sensor sampling frequency; performing multi-stage anomaly filtering of a hard threshold value, a mutation rate and a statistical interval on the data through a three-stage pipeline judgment structure executed by atomization; an FPGA hardware queue manager independent of a main processor bypasses a TCP stack to push abnormal data at the highest priority, and redundant data is stored in a zero-copy annular buffer area managed by DMA. According to the invention, the problems of intranet congestion, server I / O bottleneck and alarm delay under the centralized architecture of the traditional dynamic loop system are solved, and the fault, fault, fault, fault, fault and fault are realized on resource-limited platforms such as Cortex-M4 and the like; the average alarm delay is 8 milliseconds; and the data compression rate is more than 90%.
Owner:BEIJING ZHONGYI YUETAI SCI & TECH

Artificial intelligence chip, parallel method for vector and scalar execution pipeline, computing device, medium and program product

The invention relates to an artificial intelligence chip, a method for parallel vector and scalar execution assembly lines, a computing device, a medium and a program product. The artificial intelligence chip comprises an execution unit, the execution unit is configured with a vector execution assembly line and a scalar execution assembly line, and the scalar execution assembly line at least comprises a scalar instruction decoding unit which is configured to at least obtain an operand type, address information and scalar operation control information of a scalar instruction; a scalar instruction operand acquisition unit configured to acquire an operand source of a scalar instruction; and a scalar instruction operation unit configured to execute scalar calculation at least based on an operand type, an operand source and scalar operation control information of the scalar instruction, and write a calculation result to the scalar register group included in the execution unit. According to the method, the utilization rate and the actual computing power of hardware resources of the execution unit of the artificial intelligence chip can be remarkably improved.
Owner:SHANGHAI BIREN TECH CO LTD

Heterogeneous multi-source data dynamic synchronization method and system based on unified space-time framework

The invention discloses a heterogeneous multi-source data dynamic synchronization method and system based on a unified space-time framework, and the method comprises the following steps: collecting heterogeneous multi-source data through a real-time monitoring interface, carrying out the preprocessing, and eliminating the difference of the heterogeneous multi-source data in space-time reference and data format; identifying the change of the heterogeneous multi-source data by adopting an incremental detection algorithm, generating an incremental update packet only containing difference information of the heterogeneous multi-source data, and sending the incremental update packet to a message queue; the message queue performs hierarchical management on the incremental update package, and performs distributed parallel processing on difference information by using a distributed computing cluster; constructing a dynamic spatio-temporal index and positioning conflict data based on the spatio-temporal index; fusing the difference information, fusing conflict data in the difference information by adopting a conflict resolution mechanism, and finally storing the conflict data in a hierarchical storage architecture. Through increment detection, distributed pipeline processing and dynamic index optimization, the balance of quick response, zero-conflict fusion and efficient resource utilization is realized.
Owner:NAT UNIV OF DEFENSE TECH +1

Superconducting quantum measurement and control microcontroller based on extended instruction set and measurement and control system

The invention discloses a superconducting quantum measurement and control microcontroller and measurement and control system based on an extended instruction set, and relates to the technical field of quantum computing, and the superconducting quantum measurement and control microcontroller comprises a state machine which is used for controlling system operation and managing the execution sequence and state switching of instructions; the assembly line is of a four-step two-stage assembly line structure, a current instruction is sequentially subjected to instruction fetching, decoding, execution and write-back stages, instruction fetching of a next instruction is started to be executed in the write-back stage, and each instruction is completed in three clock periods; the code word transmitting instruction is used for generating a control pulse signal of the superconducting quantum bit; the time sequence control instruction is matched with the time delay module and is used for controlling a time interval in quantum bit operation; the feedback transmitting instruction is used for adjusting the state of the quantum bit in real time according to the feedback signal and supporting quantum error correction; according to the superconducting quantum measurement and control microcontroller and the measurement and control system, efficient and low-delay quantum bit control and real-time feedback are realized.
Owner:UNIV OF SCI & TECH OF CHINA +1

Data transform acceleration using metadata stored in accelerator memory

A method includes determining an address associated with a data transform command in a container data structure which is in the data transform accelerator. The data transform accelerator is in communication with a host computing unit. In response to a determination that the address is in the container data structure, the method includes accessing the data transform command based on the address. The data transform command is in the host computing unit. The method includes obtaining metadata based on information in the data transform command. The metadata is in the data transform accelerator or spread out in the host computing unit memory and in the memory of data transform accelerator. The method includes configuring a data transform pipeline based on the metadata. The metadata can be shared in its entirety or partially by multiple data transform commands grouped together.
Owner:MAXLINEAR INC

AI multi-agent and digital twinborn fusion scheduling process visualization method, medium and system

The invention provides an AI multi-agent and digital twinborn fusion production scheduling process visualization method, medium and system, and belongs to the technical field of AI multi-agent production scheduling. Large-scale parallel computing is achieved by constructing a GPU three-layer CUDA processing architecture, a first layer data preprocessing grid executes data cleaning in parallel, a second layer data preprocessing grid executes data cleaning in parallel, and a third layer data preprocessing grid executes data cleaning in parallel; the second layer of negotiation analysis grid operates an agent interaction recognition model based on Transform to perform parallel mode recognition, the third layer of visual calculation grid operates an LSTM-CNN fused production scheduling process mapping model to generate visual data, and calculation resource allocation is dynamically adjusted through a multi-head attention mechanism. The calculation performance is optimized by adopting a pipeline parallel and data parallel strategy, and the memory access delay is hidden by combining an asynchronous data transmission and calculation overlapping technology, so that the technical problem that the real-time parallel processing of the multi-agent high-frequency negotiation data cannot be realized is solved.
Owner:BEIJING NANCAL RUIYUAN DIGITAL TECH CO LTD

Data transform acceleration using metadata stored in accelerator memory

A method includes determining an address associated with a data transform command in a container data structure which is in the data transform accelerator. The data transform accelerator is in communication with a host computing unit. In response to a determination that the address is in the container data structure, the method includes accessing the data transform command based on the address. The data transform command is in the host computing unit. The method includes obtaining metadata based on information in the data transform command. The metadata is in the data transform accelerator or spread out in the host computing unit memory and in the memory of data transform accelerator. The method includes configuring a data transform pipeline based on the metadata. The metadata can be shared in its entirety or partially by multiple data transform commands grouped together.
Owner:MAXLINEAR INC

Accelerator architecture for near IO pipeline computing, and ai acceleration system

The present application relates to the field of accelerators, and discloses an accelerator architecture for near IO pipeline computing, and an AI acceleration system. A multi-channel direct access module comprises N DRAM controllers, and the N DRAM controllers are connected to DRAMs in a one-to-one correspondence; each DRAM controller is at least connected to k DMA controllers, and is connected to a corresponding core group cluster by means of the DMA controllers; a pipeline synchronization ring is connected to N core group clusters and comprises M cascaded forward transmission blocks and M cascaded backward transmission blocks, and the head-end forward transmission block and the tail-end forward transmission block are respectively connected to a data receiving module and a data transmitting module; the output of the tail-end forward transmission block is cascaded to the first backward transmission block; the ith forward transmission block and the (M-i)th backward transmission block correspond to each other in a front-rear direction, and are jointly connected to at least one computing core group. The accelerator of the architecture replaces a traditional multi-level cache structure, reduces the time delay caused multi-layer search, and accelerates the computing rate between the interior of the accelerator and the accelerator.
Owner:STORAGEX TECHNOLOGY INC

FFT processor, FFT computing method, system on chip, integrated circuit, and sensor

Disclosed herein are an FFT processor, an FFT computing method, a radar signal processing system-on-chip, an integrated circuit, and an electromagnetic wave sensor. The FFT processor comprises two cascaded FFT kernels, and the FFT processor has at least two operating modes among a large-point-number FFT mode, a pipeline mode and an independent parallel mode, wherein when the FFT processor is in the large-point-number FFT mode, the two FFT kernels are configured to decompose FFT of N points into two instances of FFT; when the FFT processor is in the pipeline mode, the former FFT kernel of the two cascaded FFT kernels is configured to perform distance FFT, and the latter FFT kernel of the two cascaded FFT kernels is configured to perform Doppler FFT; and when the FFT processor is in the independent parallel mode, the two FFT kernels are configured to independently process data of different channels in parallel.
Owner:CALTERAH SEMICON TECH (SHANGHAI) CO LTD

Design method and device of many-core processor architecture, equipment and storage medium

The invention discloses a design method, device and equipment of a many-core processor architecture and a storage medium, and relates to the technical field of distributed computing. Comprising the steps of 1, dividing all cores into groups, 2, integrating a CPU assembly line, an intra-core private memory, a hardware FIFO queue and a DMA engine in each core, 3, labeling the cores in each group according to coordinates, storing shared data of processors in the group by using an intra-group shared memory, and storing the shared data of the processors in the group. The method comprises the following steps: step 1, transmitting and receiving an NoC data packet through an intra-group main core control inter-group communication interface for intra-group communication, and step 4, planning an inter-group optimal transmission path for inter-group communication according to the transmitted core coordinates of a source processor group and a target processor group. According to the method, the shortest path between the groups is automatically calculated, and NoC transmission is dynamically optimized; hardware FIFO and a DMA engine are adopted to realize low-delay data transmission, and a main core efficiently exchanges data among groups through an NoC to realize response.
Owner:SHANDONG INSPUR SCI RES INST CO LTD