Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

146 results about "Pipeline (computing)" patented technology

In computing, a pipeline, also known as a data pipeline, is a set of data processing elements connected in series, where the output of one element is the input of the next one. The elements of a pipeline are often executed in parallel or in time-sliced fashion. Some amount of buffer storage is often inserted between elements.

Distributed training scheduling and communication optimization method and system of multi-modal large model on domestic computing power platform

The invention discloses a distributed training scheduling and communication optimization method and system of a multi-modal large model on a domestic computing power platform. The method comprises the following steps: virtualizing a heterogeneous computing unit of a preset platform into a virtual device pool, and fusing first-order gradient of a multi-modal sample and Hessian matrix information based on quantitative perception training to generate a sample sensitivity grading atlas; virtual device pool attributes and the sensitivity grading atlas are used as input, an optimal hybrid parallel configuration scheme is automatically generated through a configuration search algorithm, and a parallel combination mode, resource mapping and a high-sensitivity sample scheduling strategy are defined; a distributed training code of an integrated communication optimization strategy is automatically generated according to a configuration scheme, pipeline parallel communication and data parallel gradient synchronization constraint are executed in a topology adjacent equipment subset, and a hierarchical aggregation mechanism is adopted; and dynamically screening a core training set and scheduling a calculation task to complete distributed training. According to the method, efficient cooperative training of the multi-modal large model on the domestic computing power platform is realized.
Owner:GUANGXI POWER GRID CORP

BLAS3 structured operator accelerated computing system based on Hopper architecture GPU

The invention provides a BLAS3 structured operator accelerated computing system based on a Hopper architecture GPU, and relates to the technical field of computers. The system comprises: a calculation unit discrimination module for determining a calculation unit used by a current operator during operation, and estimating the maximum row dimension upper bound of the current operator in a tensor core execution path; an instruction sensing block parameter determination module dynamically determines the optimal block size and number of the input matrix in real time; the block matrix loading and aligning module divides an input matrix and a matrix to be updated into sub-matrixes by taking the block size as a basic block and completes loading of the corresponding sub-matrixes; the operator kernel function execution module completes shared memory structured parallel loading and storage of a double-precision floating-point number array of a sub-matrix corresponding to the input matrix, and calls a tensor core to carry out multiply-add accumulation calculation; and the assembly line and concurrent scheduling module adds the block calculation tasks into corresponding task sets and performs multi-stream concurrent scheduling on the task sets.
Owner:NORTHEASTERN UNIV CHINA

Real-time target detection method and system based on RTSP flow and NPU collaborative optimization

A real-time target detection method and system based on RTSP flow and NPU collaborative optimization belong to the technical field of computer vision and artificial intelligence, and are characterized by comprising the following steps: adopting dynamic memory optimization management of a hybrid pipeline architecture, and performing single-time continuous copying through a CPU (Central Processing Unit); through deep integration of innovative technologies such as DMA direct transmission NPU continuous memory pool management, dynamic batch processing scheduling, hybrid assembly line processing, intelligent equipment load balancing, parallel preprocessing optimization and vectorization post-processing, RTSP flow collaborative optimization and the like, the NPU utilization rate is improved to 85% or above, the overall average FPS is improved by 200% or above, the assembly line parallelism degree achieves three times of performance gain, and the production efficiency is greatly improved. The data transmission delay is reduced by 80%, the system stability is remarkably improved, performance degradation is avoided after long-time operation, and the method is suitable for various scenes such as edge calculation and cloud reasoning.
Owner:XIAN KEYWAY TECH

Layered self-adaptive full block pre-filling scheduling method and system for large language model reasoning

The invention discloses a hierarchical self-adaptive full block pre-filling scheduling method and system for large language model reasoning, and the method comprises the steps: carrying out the hierarchical portrait analysis of a to-be-served model, and dividing the to-be-served model into partitions with different calculation characteristics according to the calculation intensity and memory access characteristics of each layer; then, making a layering and partitioning strategy based on a partitioning result, allocating a larger partitioning size to a calculation-intensive partition, allocating a smaller partitioning size to a memory bandwidth-intensive partition, and generating a layering and partitioning mapping table; and finally, when the online scheduling is executed, querying the mapping table according to the request processing progress to determine the block target size, and jointly forming a batch processing unit by the decoding task and the pre-filled block with the heterogeneous size under the constraint of the iteration time budget to be executed. According to the method, accurate matching of calculation and bandwidth resources is achieved, the system throughput can be effectively improved, tail delay and fluctuation thereof can be remarkably reduced, bubbles under pipeline parallelism are reduced, and the method is suitable for various attention mechanisms and distributed reasoning architectures.
Owner:ZHEJIANG LAB

Abnormal data real-time filtering method and system based on edge calculation in dynamic environment system

The invention discloses an abnormal data real-time filtering method and system based on edge computing in a dynamic environment system, the method is executed by an edge computing gateway, sliding window weighted average preprocessing is equivalently realized by adopting integer shift operation, and the single computing overhead is controlled within 50 clock cycles; dynamically and adaptively adjusting dynamic reference model parameters based on the ratio of the network load to the sensor sampling frequency; performing multi-stage anomaly filtering of a hard threshold value, a mutation rate and a statistical interval on the data through a three-stage pipeline judgment structure executed by atomization; an FPGA hardware queue manager independent of a main processor bypasses a TCP stack to push abnormal data at the highest priority, and redundant data is stored in a zero-copy annular buffer area managed by DMA. According to the invention, the problems of intranet congestion, server I / O bottleneck and alarm delay under the centralized architecture of the traditional dynamic loop system are solved, and the fault, fault, fault, fault, fault and fault are realized on resource-limited platforms such as Cortex-M4 and the like; the average alarm delay is 8 milliseconds; and the data compression rate is more than 90%.
Owner:BEIJING ZHONGYI YUETAI SCI & TECH

Artificial intelligence chip, parallel method for vector and scalar execution pipeline, computing device, medium and program product

The invention relates to an artificial intelligence chip, a method for parallel vector and scalar execution assembly lines, a computing device, a medium and a program product. The artificial intelligence chip comprises an execution unit, the execution unit is configured with a vector execution assembly line and a scalar execution assembly line, and the scalar execution assembly line at least comprises a scalar instruction decoding unit which is configured to at least obtain an operand type, address information and scalar operation control information of a scalar instruction; a scalar instruction operand acquisition unit configured to acquire an operand source of a scalar instruction; and a scalar instruction operation unit configured to execute scalar calculation at least based on an operand type, an operand source and scalar operation control information of the scalar instruction, and write a calculation result to the scalar register group included in the execution unit. According to the method, the utilization rate and the actual computing power of hardware resources of the execution unit of the artificial intelligence chip can be remarkably improved.
Owner:SHANGHAI BIREN TECH CO LTD

FFT processor, FFT computing method, system on chip, integrated circuit, and sensor

Disclosed herein are an FFT processor, an FFT computing method, a radar signal processing system-on-chip, an integrated circuit, and an electromagnetic wave sensor. The FFT processor comprises two cascaded FFT kernels, and the FFT processor has at least two operating modes among a large-point-number FFT mode, a pipeline mode and an independent parallel mode, wherein when the FFT processor is in the large-point-number FFT mode, the two FFT kernels are configured to decompose FFT of N points into two instances of FFT; when the FFT processor is in the pipeline mode, the former FFT kernel of the two cascaded FFT kernels is configured to perform distance FFT, and the latter FFT kernel of the two cascaded FFT kernels is configured to perform Doppler FFT; and when the FFT processor is in the independent parallel mode, the two FFT kernels are configured to independently process data of different channels in parallel.
Owner:CALTERAH SEMICON TECH (SHANGHAI) CO LTD

IO Pipeline of a Database System and Applications Thereof

A store and compute sub-system of a database system, wherein the store and compute sub-system includes pluralities of computing nodes of a plurality of computing device clusters, wherein a set of computing nodes of the pluralities of computing nodes is operable to implement a first input / output (IO) pipeline for a first segment of a plurality of segments of a dataset to support execution of a query, wherein, the first IO pipeline functions to convert long-term storage (LTS) data of the first segment into first query ready raw data. The set of computing nodes is further operable to implement a second IO pipeline for a second segment of the plurality of segments of the dataset to support execution of the query, wherein, the second IO pipeline functions to convert LTS data of the second segment into second query ready raw data.
Owner:OCIENT HOLDINGS LLC

Intelligent task scheduling and collaborative execution method and heterogeneous computing system

The invention discloses an intelligent task scheduling and collaborative execution method and a heterogeneous computing system, and the method comprises the steps: receiving and analyzing a computing task, and extracting the computing, data and constraint features of the computing task; decomposing the task into sub-tasks with a dependency relationship; acquiring load, computing power and power consumption states of the NPU, the GPU and the CPU in real time; on the basis of task characteristics and real-time states, multiple allocation schemes are evaluated through a pre-constructed cost model, and an optimal scheduling decision for allocating the sub-tasks among the heterogeneous computing units is dynamically generated; a flow line collaborative execution plan across at least two computing units is arranged according to a decision, sub-task execution is scheduled, meanwhile, cross-unit data flow and synchronization are managed, and a corresponding system comprises an NPU, a GPU, a CPU, a system memory and an intelligent heterogeneous resource scheduler for executing the method. According to the method, fine scheduling, self-adaptive load balancing and multi-objective optimization of heterogeneous computing resources are realized, and the comprehensive efficiency and energy efficiency of the system are remarkably improved.
Owner:DONGGUAN HUAMING TENG TECH CO LTD

Large model reasoning optimization method for SLO perception in edge heterogeneous computing power network

The invention discloses a large model reasoning optimization method for SLO perception in an edge heterogeneous computing power network, aiming at an edge heterogeneous computing power computing cluster scene providing large model reasoning service, requests are strategically scheduled to make full use of heterogeneous computing power of a cluster, so that SLO heterogeneous requests are met to the maximum extent; the method comprises the following steps: firstly, separating a pre-filling stage and a decoding stage of a reasoning process to different nodes, and respectively adopting a data parallel deployment strategy and an assembly line parallel deployment strategy to ensure that the computing power of equipment is fully utilized; then, according to heterogeneous computing power characteristics of decoding nodes, a distributed assembly line non-uniform deployment strategy is realized to improve the model reasoning throughput; finally, round rewards are collected in offline exploration, a large model is used for assisting in encoding of potential rewards, a decoder is trained to decode agency rewards in each step, a reinforcement learning PPO algorithm is assisted in training a scheduler, effective scheduling requests are achieved, the resource utilization rate is increased, meanwhile, cluster energy consumption is reduced, and the throughput of large model reasoning service is increased.
Owner:JINAN UNIVERSITY

Data processing method and electronic equipment

The invention discloses a data processing method and electronic equipment, and relates to the technical field of computers, and the method comprises the steps: dividing a complete reasoning process of a request into at least one processing stage according to the number of layers of a current reasoning model, and determining a model layer of each stage; and determining a current processing stage, and executing pre-filling reasoning on the request by the corresponding model layer in the first processing node to obtain a key value cache. The second processing node carries out decoding reasoning on the key value cache to obtain the lexical element sequence of the current stage, meanwhile, a new current processing stage is determined, the pre-filling reasoning step and the decoding reasoning step are repeated until the lexical element sequences of all stages are obtained, and the lexical element sequences are spliced to obtain a final result. Therefore, the key value cache is stored and managed by taking the model layer as a unit, pipeline parallelism of three tasks of pre-filling reasoning, network transmission and decoding reasoning is realized, and the problems that computing resources are idle and TTFT is prolonged due to the fact that the key value cache is transmitted by taking the whole request as the unit are solved.
Owner:INSPUR SUZHOU INTELLIGENT TECH CO LTD

Hardware acceleration method and system of ZSTD data compression algorithm based on FPGA

The invention discloses a hardware acceleration method and system for a ZSTD data compression algorithm based on an FPGA. Unified hardware implementation and collaborative optimization are carried out on three key links of LZ77 character matching, Huffman coding and finite state entropy coding in the ZSTD data compression algorithm. According to the system, a multi-channel parallel LZ77 character matching hardware architecture is adopted, and character matching is executed on continuous positions in a plurality of parallel windows at the same time. A modular design and a deep pipeline structure are adopted, and a buffer and pipeline mechanism is introduced among modules such as an LZ77 matching module, a Huffman coding module and a finite state entropy coding module. The 375MB / s compression throughput can be achieved under the 100MHz clock frequency, the compression speed is remarkably increased while the compression ratio is guaranteed, and the method is suitable for scenes with high requirements for compression performance of data centers, edge computing and embedded systems and has high practicability and popularization value.
Owner:XIDIAN UNIV

Data migration method and device based on calculation pipeline, processor and related product

The invention provides a data migration method and device based on a computing pipeline, a processor and a related product. The data migration method based on the calculation pipeline comprises the following steps: receiving a data migration instruction of to-be-processed data; the data migration instruction is used for indicating a target data migration operation on the to-be-processed data; determining a target complexity classification corresponding to the target data migration operation based on a corresponding relationship between a preset data migration operation and the complexity classification; if the target complexity is classified as a preset complex operation type, converting the to-be-processed data and the data migration instruction into a general calculation task; the preset complex operation type comprises data conversion and data analysis in the data migration operation; and executing the general-purpose computing task through the general-purpose computing pipeline to obtain target data after the target data migration operation is completed. According to the method, the hardware and software complexity of the graphics processor for realizing data migration can be reduced.
Owner:VERISILICON MICROELECTRONICS (CHENGDU) CO LTD +1

Interleaved pipeline scheduling method, apparatus, electronic device, storage medium, and program

This disclosure provides interleaved pipeline scheduling methods, apparatus, devices, and storage media, particularly relating to the computer technology field, and more particularly to the artificial intelligence technology field, including deep learning, neural networks, and parallel computing. [Solution] A specific solution includes obtaining micro-batch scheduling parameters based on the pipeline splitting dimension, interleave dimension, and cumulative count; determining the transmission interval of a micro-batch requiring cache scheduling based on the micro-batch scheduling parameters; and performing cache scheduling on the computation results of the micro-batch requiring delayed transmission based on the transmission interval and cache units. According to this disclosure, a cache unit, for example, a cache queue, can be created, the computation results of some computation units can be cached in the cache unit, and the computation results can be transmitted with a delay. This allows for delayed transmission of the computation results of some micro-batches, improving the applicability of interleaved pipeline scheduling.
Owner:BEIJING BAIDU NETCOM SCI & TECH CO LTD

Pipelined architecture based computing engine acceleration apparatus, method, device and medium

The application provides a pipeline structure-based computing engine acceleration device, method, equipment and medium in the technical field of neural network hardware acceleration, the device comprises: a circular row buffer, comprising: a first multiplexer, used for determining the storage position of n feature map pixels of an input feature map, wherein n is a positive integer greater than 0; m random access memories, used for caching n feature map pixels according to the storage position, wherein m is a positive integer greater than 0; a second multiplexer, used for outputting the feature map pixels written into the m random access memories; a stream driving filler, used for performing boundary filling on the feature map pixels output by the second multiplexer according to the position of the feature map pixels and the boundary parameter of the convolution layer in the neural network; a computing unit array, used for performing convolution calculation on the filled feature map pixels; and a weight memory, used for caching the weight data of the neural network.
Owner:INST OF SEMICONDUCTORS - CHINESE ACAD OF SCI

Cross-data center large model training system architecture and resource allocation method and system

PendingCN121957895AImplement collaborative trainingEfficient collaborative utilizationResource allocationBiological modelsWide areaData center
The invention provides a system architecture for cross-data center large model training and a resource allocation method and system, and belongs to the technical field of cross-wide area distributed large model training. According to the method, large-scale model cooperative training across multiple data centers can be realized, and the bottleneck that the computing power of a single data center is limited is broken through. Through unified modeling and scheduling of calculation, memory and network resources, task loads can be intelligently allocated according to hardware performance and network bandwidth of different data centers, and efficient collaborative utilization of computing power resources is realized. The training task of the super-large-scale model can be rapidly completed in the heterogeneous computing power environment, and the training time is remarkably shortened. The provided flexible parallelism degree allocation method can be adaptive to different task and resource conditions, the proportion of data parallelism, model parallelism and pipeline parallelism is automatically adjusted, the parallelism efficiency is improved, and the communication overhead is reduced. A training time estimation function is integrated, the overall time delay and resource requirements can be predicted before task execution, and a basis is provided for scheduling decision making.
Owner:BEIJING JIAOTONG UNIV

Improved systolic array calculation device for improving assembly line efficiency

The invention relates to an improved systolic array computing device capable of improving pipeline efficiency, the device comprises an input characteristic value cache array, a multiply-accumulate array and an output characteristic value cache array, and the multiply-accumulate array is composed of N * N multiply-accumulate units. In order to improve the utilization rate of the array, data paths connected end to end are added in the horizontal direction and the vertical direction respectively, so that the array forms a Torus annular interconnection structure; input characteristic values are injected through diagonal positions, partial sum results are annularly accumulated in the vertical direction, so that continuous flow of input and partial sum is achieved, the structure remarkably reduces assembly line filling and emptying delay, hardware redundancy caused by input and output FIFO arrays in a traditional systolic array is avoided, calculation of a 512 * 512 matrix through a 256 * 256 array is taken as an example, and the structure has the advantages of being simple in structure, convenient to operate and low in cost. The calculation delay is reduced from 2560 clock cycles to 2176 clock cycles, and the array utilization rate is improved to 94%. The method has the advantages of being high in efficiency, low in cost, good in expansibility and suitable for neural network and matrix calculation acceleration.
Owner:SHAOXIN LABORATORY

A heterogeneous acceleration system and method for analog modulation recognition of large-scale multi-channel signals

PendingCN122285217ASolve the problem of computing congestionImprove data throughputVideo memoryComputer architecture
This invention discloses a heterogeneous acceleration system and method for analog modulation recognition of large-scale multi-channel signals, comprising: a CPU host processing end and a GPU device computing end; the CPU host processing end is used for asynchronous pipelined processing of data reading, task scheduling, and result distribution, realizing parallel execution of data reading, GPU computing, and result distribution on the time axis; the GPU device computing end completes the entire process of signal preprocessing and deep learning model inference within the GPU memory after the original signal is transmitted to the video memory via the PCIe bus, and does not exchange data with the CPU host processing end until the recognition result is generated. This invention eliminates data transport and I / O blocking through a three-stage asynchronous pipeline architecture and a closed-loop processing link in the entire video memory, and combines parallel preprocessing operators, inference acceleration engines, multi-stream scheduling mechanisms, and dynamic video memory pools to achieve high throughput and low latency recognition of large-scale multi-channel analog modulation signals.
Owner:SHANGHAI UNIV

Hierarchical data driving method, system and equipment for digital twin archives and medium

The invention relates to the technical field of data driving, and particularly provides a hierarchical data driving method, system and device for digital twin archives and a medium, and the method comprises the steps: dividing physical entity data into three processing levels according to real-time requirements, the first level being the highest in real-time performance, and the third level being the lowest in real-time performance; establishing a parallel and resource-isolated data processing pipeline for each level; first-level data is directly written into a memory database, and a digital twin engine actively queries through a high-frequency timer to drive a virtual model to realize millisecond-level synchronization; performing state aggregation on the second-level data through stream calculation, and pushing the second-level data to an engine to realize second-level state updating; and persistently storing the third-level data in various databases, and accessing the third-level data by an engine when receiving a query request to realize post analysis and backtracking. According to the method, the hierarchical parallel processing assembly line is established, so that accurate optimization of data driving is realized.
Owner:SHANDONG ZHENGCHEN TECH CO LTD

Python operator scheduling method and device based on streaming computing framework, and storage medium

The invention discloses a python operator scheduling method and device based on a streaming computing framework, a storage medium and computer equipment. The method comprises the steps that a task submitting end submits a task execution request to a streaming computing management service; the streaming computing management service responds to the task execution request, generates a target task containing a main program file, and sends the target task to a streaming computing cluster; the streaming computing cluster receives a target task, and executes a main program file through a Python execution environment integrated in a streaming computing framework, so as to call a user-defined data source component and a user-defined data output component running on a Java virtual machine through a communication channel between a Python virtual machine and the Java virtual machine, and constructing a data processing assembly line through the user-defined data source component, the user-defined data output component and the data processing logic defined in the main program file, and triggering the data processing assembly line to be executed in the streaming computing framework.
Owner:BEIJING GUODIAN ZHISHEN CONTROL TONGDY

Automatic selection of computer hardware configuration for data processing pipelines

A method for recommending a computer hardware configuration, including: receiving, by a processor, a machine-readable specification of a computing task; extracting, by the processor, a plurality of features from the machine-readable specification of the computing task; supplying, by the processor, the plurality of features to a reinforcement learning model to generate a proposed computer hardware configuration to execute the computing task; and providing, by the processor, the proposed computer hardware configuration to a user.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Federal learning acceleration method based on parallel sampling and training of in-batch real-time data of assembly line

PendingCN121684100AMachine learningKnowledge based modelsEvent synchronizationAlgorithm
The invention discloses a federated learning acceleration method and system based on pipeline in-batch real-time data sampling and training. According to the method, in-batch data sampling is provided on the algorithm level aiming at the problems that computing resources of edge equipment are limited and importance is outdated and gradient deviation is caused by existing static sampling: the importance of samples in a fixed mini-batch is evaluated in real time by utilizing a latest model in each round of iteration, and a dynamic micro-batch is constructed; and through a gradient correction coefficient based on a sampling probability reciprocal, distribution deviation is eliminated, and unbiased training is realized. On the system level, an assembly line parallel mechanism based on a CPU-GPU heterogeneous architecture is designed, overlapping execution of sampling and training is achieved through a double-thread-double-flow concurrent model, data competition is solved through an annular buffer area and an event synchronization mechanism, and overhead is reduced in cooperation with mixing precision reasoning. According to the method, the hardware utilization rate can be remarkably improved, the model convergence precision is improved while the training time is greatly shortened, and the method is suitable for various edge computing scenes.
Owner:NANJING UNIV OF AERONAUTICS & ASTRONAUTICS

Intelligent edge data transmission terminal based on AI Internet of Things

An intelligent edge data transmission terminal based on AI Internet of Things belongs to the cross technical field of edge computing, Internet of Things and artificial intelligence, and is a bionic concentric sphere type four-layer intelligent architecture. The architecture is composed of an intelligent core layer, a collaborative execution layer, a heterogeneous interface layer and an autonomous optimization layer from inside to outside, data and instruction streams radiate outwards with the intelligent core layer as an original point and drive the whole system to operate, meanwhile, state information of the outer layer is fed back inwards, and a closed-loop intelligent system evolving continuously is formed; the intelligent core layer is a highly collaborative organic whole, and the core of the intelligent core layer is an initiative lightweight hybrid AI engine; the collaborative execution layer customizes a storage, calculation and transmission integrated assembly line constructed by an AI SoC on a hardware bottom layer through an FPGA; the heterogeneous interface layer constructs a highly intelligent, adaptive and tough all-thing interconnection physical basis; and the autonomous optimization layer is a core enabling layer for ensuring the terminal to realize long-term stable operation and continuous autonomous evolution.
Owner:BEIJING ZHONGDIAN ZHICHENG TECH CO LTD

A learning-based lossless lightweight compression, decompression, random access and query processing method and system for GPU

This invention relates to a learning-based lossless lightweight compression, decompression, random access, and query processing method and system for GPUs. The method includes: processing input data using the GPU's native learning compression pipeline and directly generating learning-based lossless lightweight compressed SLAP layout data in GPU memory; decompressing and reading the target SLAP layout data using a warp collaborative learning decompression module to obtain decompressed data; selectively accessing the target SLAP layout data on the GPU side based on target point or range requests, and calling the warp collaborative learning decompression module to decompress a portion of the target SLAP layout data to obtain decompressed data; and outputting the decompressed data to a query operator for further processing or directly returning the query result, depending on actual needs. This invention can be widely applied in the fields of data compression, columnar storage, GPU parallel computing, and database acceleration.
Owner:RENMIN UNIVERSITY OF CHINA

Convolutional code parallel pipeline decoding acceleration system and method based on memory-computing integrated architecture

The application discloses a convolution code parallel pipeline decoding acceleration system and method based on a memory-compute integrated architecture, comprising: a global data scheduling module, which is used for slicing convolution code data from a magnetic tape storage device according to a set rule and scheduling the data through a multi-level cache mechanism; a memory-compute integrated unit array, which is used for storing data tiles and intermediate results output from the global data scheduling module and performing convolution operation, path metric calculation and surviving path selection through a reconfigurable computing unit; a parallel pipeline controller, which is used for dynamically allocating decoding tasks and controlling the pipeline beat of the memory-compute integrated unit array; an adaptive resource configuration module, which is used for monitoring the system load in real time and dynamically adjusting data distribution strategies and computing resource scheduling; and a check and error correction unit, which is used for checking and correcting the decoding results output from the memory-compute integrated unit array and then outputting the results; the decoding acceleration system and method realize efficient and low-delay convolution code decoding.
Owner:HANGZHOU INTERNATIONAL INNOVATION INSTITUTE OF BEIHANG UNIVERSITY

An AI multi-agent and digital twin fusion production process visualization method, medium and system

The application provides an AI multi-agent and digital twin fusion production scheduling process visualization method, medium and system, belonging to the technical field of AI multi-agent production scheduling. Large-scale parallel computing is realized by constructing a GPU three-layer CUDA processing architecture. The first layer of data preprocessing grid performs data cleaning in parallel. The second layer of negotiation analysis grid runs a Transformer-based agent interaction recognition model for parallel pattern recognition. The third layer of visualization calculation grid runs an LSTM-CNN fusion production scheduling process mapping model to generate visualization data. The multi-head attention mechanism dynamically adjusts the allocation of computing resources. The pipeline parallel and data parallel strategies optimize the computing performance. The asynchronous data transmission and computing overlap technology hides the memory access delay, solving the technical problem that multi-agent high-frequency negotiation data cannot be processed in real time.
Owner:BEIJING NANCAL RUIYUAN DIGITAL TECH CO LTD

Full-link traffic recording method, full-link traffic playback method and computing equipment

The invention discloses a full-link traffic recording method, which comprises the following steps of: recording a request of each transaction, request information of the transaction, an external dependence response result, abnormal information and execution time consumption through pipeline middleware of Asp.netCore; the external dependency of each transaction is recorded through a dynamic weaving structure provided by Harmony.lib, recording of the external dependency is achieved, and recording content of the external dependency comprises response parameters of the service, an external dependency response result and execution time consumption. Through configuration of recording and comparison rules, multi-dimensional verification and problem positioning are realized, the problems of high invasiveness, incomplete dependence coverage, poor playback matching and low verification efficiency of a traditional test tool are solved, and the method is adaptive to various requirements of interface regression testing, online problem reproduction and the like in a distributed scene.
Owner:SHANGHAI EHI CAR RENTAL CO LTD +2

Chip computing inference method and system of hybrid expert model

The application provides a chip computing reasoning method and system of a mixed expert model, relates to the technical field of artificial intelligence computing, and comprises the following steps: acquiring input data to calculate expert activation weights, performing time sequence tracking and clustering analysis to identify an expert combination, and preloading the weights to a shared cache; selecting a target expert based on the activation weights, loading the weights; dividing data and constructing a pipeline level, and mapping to different computing units; controlling each unit to perform calculation according to the pipeline, and performing weighted summation to obtain a result. The application improves the reasoning efficiency and reduces the memory access overhead through expert activation mode prediction and pipeline parallel processing.
Owner:XINQIAO (BEIJING) SEMICONDUCTOR CO LTD

A signal processing combination performance checking system based on DSP communication test

PendingCN122332233APathPingData stream
The application discloses a signal processing combination performance verification system based on DSP communication test, relates to the real-time test technical field of a digital signal processor, and specifically comprises: a homomorphism copying module, a slot injection module and a performance judgment module; the dynamic configuration parameters of a current DSP operation core are captured in real time through hardware sniffing of a bus, and the dynamic configuration parameters are mirrored to shadow registers; a peek unit is used to identify the bubble period of a pipeline through prediction logic; at the moment when the main computing unit is in a waiting period, the operation of the shadow register group is enabled by hardware, a preset characteristic test vector is input into a computing engine, and a shadow operation is completed; the real-time data flow of a main path and the reference data flow of a shadow path are compared in time domain and frequency domain by a hardware comparison array, the second derivative of a response time delay is calculated, a comprehensive performance factor is output, online verification of the combination performance of a DSP signal processing is realized, fault types can be accurately distinguished, and the reliability and maintainability of the system are improved.
Owner:BEIJING SHIJICHEN DATA TECH CO LTD