Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

95 results about "Data reuse" patented technology

Artificial intelligence model training resource adaptive distribution system

The invention belongs to the technical field of artificial intelligence, and discloses an artificial intelligence model training resource adaptive distribution system. The method comprises the following steps: acquiring and calculating graph structure data and hardware resource state data in real time, and calculating a data reuse rate and generating a candidate operator fusion scheme by constructing an operator execution time sequence constraint matrix and identifying an operator cluster of data locality characteristics; a resource competition hotspot prediction mechanism is introduced, memory bandwidth occupation fluctuation characteristics are analyzed, a resource conflict probability is calculated for a fusion scheme, and a dynamic balance optimization model of fusion income and resource conflicts is constructed. An optimal operator fusion decision sequence and a resource allocation strategy are generated through iterative solution, and accurate dynamic adjustment of computing resources in the training process is achieved. The training efficiency and the resource utilization rate are improved, the energy consumption is reduced, and the system stability is enhanced.
Owner:YANGZHOU HUAZHISHENG INFORMATION TECHNOLOGY CO LTD

Marketing mobile terminal dynamic behavior compliance monitoring method and system

The invention discloses a marketing mobile terminal dynamic behavior compliance monitoring method and system, belongs to the technical field of information transmission monitoring, and aims to solve the problems that existing monitoring is insufficient in non-text information processing, redundant in sensitive word judgment, incomplete in violation expression variant coverage and low in data reuse rate. The method comprises the steps that dynamic behavior information is captured, a composite retrieval code retrieval history library is constructed, and whether sensitive word recognition is conducted or not is judged; if not, generating a standardized word segmentation sequence; comparing the segmented words to determine matched segmented words and unmatched segmented words, calling a historical synonym set of the matched segmented words, positioning core semantics of the unmatched segmented words, screening evading expressions to generate a synonym set of the unmatched segmented words, and constructing a synonym data set; calling a matched word segmentation historical judgment result, establishing a synonym sequence calculation probability, and judging a sensitive word in combination with a significant proportion; executing a processing flow and updating the history library and the word segmentation library; according to the invention, efficient and accurate compliance monitoring is realized, the data reuse rate is improved, and the effectiveness of a monitoring closed loop is guaranteed.
Owner:JIANGSU ELECTRIC POWER INFORMATION TECH

Sweeping robot cleaning control system with self-adaptive mode adjustment

The invention discloses a sweeping robot cleaning control system with a self-adaptive mode adjustment function, and relates to the technical field of intelligent cleaning robot control systems, and the system comprises a multi-mode sensing module, a central processing module and a cleaning execution module. The multi-modal sensing module collects multi-dimensional data such as a hyperspectral image, acoustic response and gas concentration of ground stains, and the central processing module extracts features through the data fusion unit and deeply fuses the features to generate analysis results of stain types, adhesion strength and chemical components. The decision control unit dynamically generates a cleaning instruction in combination with a preset strategy and the real-time state of the robot, the cleaning execution module executes a targeted cleaning action, and meanwhile, the system constructs a dynamic stain atlas to achieve historical data reuse and prospective adjustment, so that the cleaning efficiency and the resource utilization efficiency are remarkably improved, and the cleaning effect is improved. The method is suitable for intelligent cleaning scenes of home and office environments.
Owner:DONG GUAN DA YIN SU JIAO ZHI PIN YOU XIAN GONG SI

Sonar intelligent interpretation method and system

The invention provides a sonar intelligent interpretation method and system, and the method comprises the steps: integrating end-side equipment (disposed on a ship-borne platform) and side-side equipment (disposed in a shore-based machine room) to form a sonar intelligent interpretation system; through the intelligent sonar interpretation system, full-process automatic processing of sonar interpretation from data acquisition of sonar images to interpretation is realized, the interpretation efficiency of the sonar images is effectively improved, and in addition, the intelligent sonar interpretation system also can be used for realizing automatic interpretation of the sonar images on the basis of mass sonar image data (namely sonar images stored in a data warehouse) obtained through actual measurement. Model training of the sonar interpretation model and training examination of interpretation personnel are synchronously achieved, a sustainable closed-loop optimization system is formed, and the problems that in traditional sonar interpretation, the interpretation accuracy is low, the data reuse rate is poor, and personnel training examination is difficult are solved.
Owner:BEIJING ZHONGKE STRON CLOUD INTELLIGENT TECH CO LTD

Model reasoning acceleration method and device, electronic equipment and nonvolatile storage medium

The invention discloses a model reasoning acceleration method and device, electronic equipment and a nonvolatile storage medium. The method comprises the steps that in the process that a natural language processing model carries out general matrix calculation, a first matrix and a second matrix are divided into a plurality of blocks respectively, general matrix calculation is executed through an image processing unit, and the first matrix and the second matrix are matrixes needing general matrix calculation; loading the blocks from a global memory of the image processing unit to a shared memory of the image processing unit, and performing matrix operation by adopting threads corresponding to the blocks in the shared memory; and writing a result obtained by calculation of each thread back to a corresponding position in the global memory to obtain a matrix calculation result corresponding to matrix calculation of the first matrix and the second matrix. The technical problem of low model reasoning calculation efficiency caused by high memory access delay and low data reuse rate of a global memory of an image processing unit during reasoning of a large model in related technologies is solved.
Owner:CHINA TELECOM CORP LTD

Lightweight reasoning acceleration architecture based on instruction set architecture

The invention relates to a lightweight reasoning acceleration architecture based on an instruction set architecture, and the architecture comprises the steps: initializing hardware level configuration, building a software level environment, and constructing a basic operation environment of the lightweight reasoning acceleration architecture; a model analysis mode is adopted to reduce the memory demand and carry out allocation optimization on the memory, and a reasoning execution basis is constructed; completing calculation parameter configuration by adopting modes of accelerator parameter configuration, weight data preloading and output data loading; based on the calculation task and monitoring information of the calculation process, calculation resources are dynamically scheduled, and efficient scheduling and execution of the calculation task are achieved. The energy efficiency is improved based on a data reuse strategy, memory access optimization and power consumption management; a comprehensive anomaly detection system is constructed, and anomaly conditions are detected and reported; the collaborative optimization of reducing the power consumption and improving the calculation efficiency is achieved, the real-time performance and the stability are both considered, and the efficient and reliable lightweight reasoning effect is achieved in the edge calculation scene.
Owner:CHANGZHOU UNIV

GPGPU thread block scheduling method and system based on data space locality

The invention belongs to the field of GPGPU chip design, and particularly relates to a GPGPU thread block scheduling method and system based on data space locality, and the method comprises the steps: carrying out the statistics of Bank feature information of access data of all thread blocks, and classifying the thread blocks accessing the same Bank into the same Bank group according to the Bank feature information; preferentially distributing the thread blocks of the same Bank group to the same programmable multiprocessor until the resources of the programmable multiprocessor are saturated; setting a private row cache space for each Bank in each programmable multiprocessor; when the data in the private line cache space is updated, synchronously updating the corresponding data in the L1 cache, the L2 cache and the DRAM; counting the hit rate of the private line cache in real time, and closing the private line cache function when the hit rate is lower than a preset threshold value. High-delay external storage access is reduced, the data reuse rate is improved, and the execution efficiency is remarkably improved.
Owner:SHANDONG INSPUR SCI RES INST CO LTD

Data customization production delivery method and system based on block chain and cloud

The invention discloses a data customization production delivery method and system based on a block chain and cloud, and relates to the technical field of data processing and information security, and the method comprises the steps that a data customization party and a production party register and bind a public chain address in the system; deploying an intelligent contract with an automatic fund transfer function on a public chain based on the transaction demand; the customization party prepares a production environment, including creating a production cloud container and a cloud database, and generating an asymmetric key pair; the producer completes data production in the isolated cloud container, encrypts the whole-process data by using the public key of the customizing party and stores the encrypted data into the private block chain; and the system then starts an independent data verification container, decrypts the data by using a built-in private key, carries out integrity comparison and quality verification, and automatically triggers an intelligent contract to carry out settlement according to a verification result. The problems that the data is easy to reuse, the ownership is unclear and the settlement efficiency is low in data customization production are effectively solved, and safe, credible and efficient data production and delivery are realized.
Owner:HEFEI DAZHIHUI CAIHUI DATA TECH CO LTD

Data reuse computing architecture

Disclosed is an improved computer architecture for generating an electronic form by having a user input a portion of form data and obtaining the remaining portion of the format data from a data storage that stores reusable data for various forms. A master data object is configured to store “request-agnostic data,” which is typically that portion of form data that does not differ, or is common, between various forms. The data that differs between various forms, such as the data that is specific to a form, may be considered as “request-specific data.” When a form generation request is received, the user may be prompted to input request-specific data, but not the request-agnostic data. The system automatically obtains the request-agnostic data from the master data object, and integrates the request-agnostic data with the request-specific data to generate the form.
Owner:THE BANK OF NEW YORK MELLON

A method and system for tactical competition data interaction based on offline physical battlefields

This invention discloses a tactical competitive data interaction method and system based on an offline physical battlefield, relating to the field of offline real-world entertainment and multi-platform client linkage technology. The method includes: players registering and binding their accounts via any mainstream mini-program or H5 client such as WeChat, Alipay, Douyin, Meituan, and Xiaohongshu; setting up minefields (including proprietary toy mines), wearable smart equipment (including smart helmets, smart vests, and smart toy guns), escape room puzzle areas, and withdrawal points on the offline physical battlefield; anti-theft tags embedded in items on the battlefield and in the minefield; and an alarm sounding when a player dies without leaving with all items; collecting player offline behavior data and synchronizing it to the server; unified settlement at the cashier, with data synchronized to various mini-programs and H5 clients; enabling data reuse, online virtual auctions for offline redemption, and lottery activities; and the system including an offline physical terminal, multi-platform clients, and a server, with the server handling data processing, algorithm calculations, and anti-theft management. This invention solves the problems of existing offline tactical competitive experiences being monotonous, inefficient in operation, prone to item loss, and insufficient online-offline linkage, enhancing player immersion and stickiness, forming a unique technological barrier, and facilitating lightweight promotion.
Owner:HEBEI YUAN CULTURE COMM CO LTD

An image FHOG feature extraction device

The application discloses an image FHOG feature extraction device and belongs to the technical field of image processing. (a,b) The histogram cache module comprises K bank cache units, which are denoted as KBank (m,n) The histogram feature vector of Cell(i,j) is cached; m=i%c; n=j%c; the cache mode is convenient for accumulating the contribution values of the same pixel generated in the same period to each associated Cell into the histogram feature vector of the corresponding associated Cell; all pixels only need to be processed once, data reusability is greatly improved, and the FHOG feature of the image can be extracted at a faster operation speed under the premise of meeting the accuracy requirement.
Owner:HUAZHONG UNIV OF SCI & TECH

A computing power multiplication pre-storage system, method and embedded device

This invention discloses a computing power multiplication pre-storage system, method, and embedded device, belonging to the field of computing architecture and data reuse technology. Addressing the problems of redundant repetitive computation, conflict risks and irreparable damage in data integrity verification, and inefficient request-response modes in existing technologies, this invention proposes a closed-loop "pre-storage-reuse-verification" system. The system includes: a hotspot identification and adaptive pre-storage module, used to identify high-frequency computing tasks, dynamically adjust the pre-storage range, and automatically eliminate low-frequency data, achieving self-evolution and convergence of the data set; a multi-level trusted caching module, adopting a layered storage architecture to achieve layered management of data value, and possessing automatic backfilling between layers, periodic consistency verification, and automatic repair mechanisms; a mathematically based integrity verification module, configuring corresponding differential or recursive verification logic for a wide range of computing tasks, from basic data types and various mathematical functions to operational operators, providing a zero-conflict, reversible, and repairable lightweight data verification and restoration method; and an asynchronous push processing engine, which immediately returns after receiving a request, submitting the computing task to a background lock-free queue for batch processing, achieving decoupling between computation and requests. This invention pioneers the deep integration of differential verification and computation acceleration, reconstructing the traditional "request-computation-response" logic into a "pre-store-query-retrieve" push-pull interaction logic. By reusing repeated calculations through module collaboration, it not only significantly reduces computation latency, but more importantly, it improves the effective utilization of computing power, greatly saves hardware resources and energy consumption, and ensures data credibility and integrity. It is applicable to fields such as artificial intelligence, scientific computing, engineering simulation, financial technology, and video processing.
Owner:马光雷

A bank risk control model data sharing method and system

The application provides a bank risk control model data sharing method and system, the method comprises the following steps: when a bank risk control model needs to register a public data table, configuring the public data table according to the bank risk control model, and configuring parameters of the bank risk control model; the public data table comprises data shared among the bank risk control models; when the bank risk control model is executed, reading and writing data in the public data table according to the configuration of the public data table, and calling the data in the public data table according to a calling rule. The method guarantees the consistency of the data, reduces the design work of data reuse logic among the bank risk control models, and improves the data sharing capability.
Owner:SHENZHEN YLINK COMPUTING SYST

Multi-axis linkage track optimization system of numerical control planer type milling machine

The invention provides a numerical control planer type milling machine multi-axis linkage track optimization system, belongs to the technical field of numerical control planer type milling machine multi-axis linkage machining, and aims at solving the problems that an existing system is poor in track adaptability, insufficient in inter-axis synchronization, single in optimization dimension and complex in operation. The system comprises a data acquisition module, a trajectory planning module, a trajectory optimization module, an execution control module, a state monitoring module, a parameter storage module, a multi-axis cooperative compensation module and a man-machine interaction module, and all the modules are in closed-loop signal linkage. The data acquisition module acquires full-dimensional basic parameters, the trajectory planning module adapts to machining precision through NURBS interpolation, the trajectory optimization module gives consideration to precision, efficiency and inter-axis impact degree through multi-objective optimization, the multi-axis cooperative compensation module corrects errors in real time, and the parameter storage module achieves data reuse. The multi-axis linkage machining precision and stability can be improved, the operation threshold is lowered, and the multi-axis linkage machining method is suitable for high-precision complex workpiece machining.
Owner:JONAK CNC EQUIPMENT (JIANGSU) CO LTD

A non-functional interactive personalized random input generation method

This invention relates to the field of artificial intelligence processing, specifically to a method for generating non-functional interactive personalized random input, comprising the following steps: S1, basic data collection; S2, personalized random input generation; S3, output of generated content and reception of interactive feedback; S4, emotional feature fusion and extraction, obtaining surface emotional features and scene emotional features and performing emotional feature fusion; S5, lightweight emotional classification and multimodal response scheduling; S6, emotional response output, simultaneously outputting text, emoticons, and voice resources to the user; S7, quantitative analysis of user interaction willingness; S8, dynamic threshold iteration and adjustment of interaction intensity; S9, cross-stage data reuse and full-link reverse iteration, solving the problems of insufficient content generation adaptability, insufficient accuracy and real-time performance of emotional responses, and inefficient interaction links and adjustment mechanisms in traditional technologies in terms of underlying algorithms, model architecture, and technical mechanisms.
Owner:SICHUAN JISU POWER TECH CO LTD

Power amount control device

A power amount control device (10) has a storage unit (180) that stores situation information as training data (550). The power amount control device (10) also comprises a training data reuse unit (161) that, if new situation information including an unencountered outside air temperature range that is not stored as the training data (550) is generated in an operation stage, identifies situation information having the same in-room initial temperature level, load arrangement pattern, and air-conditioning control level in an outside air temperature range closest to the outside air temperature of the unencountered outside air temperature range, reuses an air-conditioning power function and a temperature function indicated by the identified situation information, and sets the air-conditioning control level indicated by the identified situation information to the new situation information after a prescribed adjustment process has been carried out.
Owner:NT T INC

A gpu scan path optimization method

The present application relates to the field of scientific computing and GPU technology, and particularly relates to a GPU scan path optimization method.The method comprises: constructing a three-dimensional Stencil calculation GPU framework based on an optimal scan traversal path; based on the three-dimensional Stencil calculation GPU framework, constructing a scan traversal path mathematical model, abstracting the scan traversal path into an integer programming problem, taking the total memory access amount as an optimization target, and optimizing to obtain an optimal two-dimensional window scan traversal path.Compared with the traditional direct calculation, the method of the present application can significantly reduce the data transfer in the global memory, improve the utilization rate of the accessed data, i.e., improve the data reuse rate, reduce the invalid waiting caused by the data blocking and high memory delay of the calculation unit, and thus improve the calculation efficiency.
Owner:COMP NETWORK INFORMATION CENT CHINESE ACADEMY OF SCI

ARCHITECTURE TOWARDS THE REUSE OF FRAME DATA VIA MULTI-DIMENSIONAL DATA PROCESSORS

Aspects of this technical solution can increase processing speed in low-latency applications while maintaining the integrity of image feature recognition at these higher speeds. For example, in image processing environments associated with autonomous navigation (e.g., driving), a large amount of image data must be processed quickly and accurately to maintain the reliability and timeliness of the physical environment. For example, embodiments according to this disclosure can provide powerful and accurate image feature recognition of input frame data that exceeds the capabilities of CPU processing or general-purpose GPU processing.
Owner:NVIDIA CORP

Implementation method, device and medium of a general fixed-point matrix multiplier based on FPGA high-performance computing architecture

This invention discloses a method, apparatus, and medium for implementing a general-purpose fixed-point matrix multiplier based on a high-performance computing architecture in an FPGA. The method includes: designing matrix partitioning strategies at different levels based on the parallelism of the AI ​​engine array resources on the Versal ACAP platform; data scheduling and multiplexing based on the data packet stream and data packet exchange of the AI ​​engine and AXI stream, accelerating the kernel through matrix partitioned multiplication, and designing data scheduling and multiplexing strategies; and implementing a high-throughput vectorized matrix multiplication pipeline on the AI ​​engine vector processor. This invention achieves multi-level partitioning of matrix multiplication based on the AXI stream transport protocol and AI engine array on the Versal ACAP platform, enabling efficient utilization of hardware resources, effectively improving data reuse rate, and achieving high parallelism, while achieving high computational speed under the high-speed clock of the AI ​​engine. This invention can be widely applied in the field of high-performance computing.
Owner:SOUTH CHINA UNIV OF TECH

Hardware accelerator facing sparse matrix vector multiplication, equipment and application method

The invention discloses a sparse matrix vector multiplication-oriented hardware accelerator, sparse matrix vector multiplication-oriented hardware accelerator equipment and an application method, and the hardware accelerator comprises an off-chip storage system, an on-chip network used for carrying out data exchange routing, and an on-chip processing system used for executing access and multiplication calculation and matrix in-row element merging, the off-chip storage system comprises HBM channels used for storing five types of data of a column index, a row index, a vector value, a matrix value and a result vector, each HBM channel comprises an HBM stack and a memory controller, and the HBM channels used for storing the vector values are connected with an on-chip network through second-level caches. And the other HBM channels are directly connected with the on-chip processing system. The method aims at improving on-chip data reuse of the hardware accelerator for sparse matrix vector multiplication, reducing off-chip memory access times and improving performance and energy efficiency performance of the hardware accelerator.
Owner:NAT UNIV OF DEFENSE TECH

A data distribution method, apparatus, device, storage medium, and product

The application provides a data allocation method, device, equipment, storage medium and product, which are applied to the technical field of computers. The method comprises the following steps: acquiring a plurality of time sequence characteristics related to a data access request; converting the plurality of time sequence characteristics from a time domain to a frequency domain, and performing machine learning analysis in the frequency domain; inversely converting the result of the machine learning analysis from the frequency domain to the time domain to obtain a data reuse probability prediction value; and based on the data reuse probability prediction value, allocating data corresponding to the data access request to an SLC cache area or a QLC main storage area of a solid state disk. The technical scheme solves the problems that the existing method cannot dynamically adapt to load changes and has low prediction accuracy.
Owner:SHANDONG YUNHAI GUOCHUANG CLOUD COMPUTING EQUIP IND INNOVATION CENT CO LTD

Direct current system insulation detection method and portable insulation detection and verification integrated device

The application discloses a DC system insulation detection method and a portable insulation detection and verification integrated device, and the method comprises the following steps: constructing a DC system bus insulation detection circuit, and establishing a mathematical model of bus voltage and bus insulation resistance; detection switches K4 and K12 are arranged in the bus insulation detection circuit; balance bridge switches K5, K13, K6 and K14 are arranged; K4, K12, K5, K13 and K6 are closed, and K14 is opened, and the positive and negative bus voltages to the ground are measured; after each measurement, the regularization data reuse recursive least square method is used to update the to-be-identified parameters; the states of the remaining switches remain unchanged, K6 is opened, K14 is closed, and the positive and negative bus voltages to the ground are measured; after each measurement, the regularization data reuse recursive least square method is used to update the to-be-identified parameters; and the to-be-identified parameters are combined with the mathematical model to calculate the positive and negative bus insulation resistance values, and the accuracy of bus insulation resistance measurement is improved.
Owner:STATE GRID ANHUI ELECTRIC POWER CO LTD ELECTRIC POWER SCI RES INST

A binary weighted convolutional neural network accelerator and RISC-V system-on-a-chip

This invention discloses a binary weighted convolutional neural network accelerator and a RISC-V system-on-a-chip. The binary weighted convolutional neural network accelerator includes an instruction parsing module, an address generation module, a finite state machine, a shift window module, a feature map storage module, a weight storage module, a multi-channel shared computing array, an accumulation module, and batch normalization and pooling modules. The RISC-V system-on-a-chip includes a FLASH module, a DDR module, an E203 RISC-V soft core module, an AXI Interconnect module, an AXI data transmission path, a cross-clock domain module, and the binary weighted convolutional neural network accelerator. This invention can significantly improve data reuse and storage retrieval efficiency, reduce resource consumption, and thus improve the overall operating efficiency of the accelerator.
Owner:UNIV OF ELECTRONICS SCI & TECH OF CHINA

Hardware-network collaborative optimization FPGA (Field Programmable Gate Array) digital identification system and method

According to the hardware-network collaborative optimization FPGA digital recognition system and method, the LeNet-5 network architecture is improved, the number of channels is adjusted to be a multiple of 4 so as to be matched with an FPGA parallel computing array, a ReLU activation function is adopted to replace traditional Sigmoid / Tanh, and 8-bit integer quantization processing is implemented; a four-level assembly line architecture (convolution-quantization-pooling-activation) is adopted in the aspect of hardware, a snakelike movement strategy and a double-Block dynamic storage scheme are combined, the energy efficiency ratio and the reasoning speed are remarkably improved while high recognition precision is kept, an end-to-end real-time processing function chain from image collection, neural network reasoning to result display is achieved, and the method has the advantages of being high in recognition precision and high in practicability. The data reuse rate is obviously improved, the storage access bandwidth is reduced, a referenceable low-power-consumption and high-performance solution is provided for edge computing applications such as intelligent monitoring and industrial quality inspection, and high real-time performance and practicability are achieved.
Owner:XI'AN PETROLEUM UNIVERSITY

Method, device, equipment, medium and program product for generating machine tool servo parameters

This application relates to the field of CNC machine tool servo control technology, and particularly to a method, apparatus, device, medium, and program product for generating machine tool servo parameters. The method includes: acquiring parameter data of the machine tool servo system to be optimized; calculating the actual distortion error value between the input trajectory data and the actual output trajectory data of the machine tool servo system to be optimized based on the parameter data; determining the speed loop proportional gain value and / or speed loop integral gain value of the machine tool servo system to be optimized based on the actual distortion error value, and controlling the machine tool servo system to be optimized to operate according to the speed loop proportional gain value and / or speed loop integral gain value. This solves the problems in related technologies, such as long trial cycles and high costs leading to low equipment utilization; reliance on experience leading to local optima; lack of trajectory spectrum adaptation analysis, requiring readjustment of new trajectories and difficulty in data reuse; and the lack of a quantitative relationship between parameters and contour errors, making it difficult to guide optimization and cross-device adaptation.
Owner:TSINGHUA UNIVERSITY

Fine-grained quantization matrix multiplication device and method based on systolic array

The invention relates to the technical field of computer systems, and provides a fine-grained quantization matrix multiplication device and method based on a systolic array, and the device comprises an input scheduling module which is used for receiving input matrix data and corresponding quantization factor data, and adding type identifiers to the input matrix data and the quantization factor data; the systolic array is formed by connecting a plurality of processing units in a two-dimensional net structure, and each processing unit is used for selectively executing multiply-accumulate operation or inverse quantization operation according to the data received at the upstream and the type identifier of the data; wherein the input scheduling module is used for scheduling data added with a type identifier to a systolic array according to rows and columns; and the output module is used for receiving and outputting a calculation result calculated by the systolic array. The problems that quantization factor insertion is low in efficiency, inverse quantization is separated from calculation flow and task switching overhead is large are solved, and the data reuse rate, the calculation efficiency and the overall energy efficiency ratio of the systolic array in a fine-grained quantization scene are remarkably improved.
Owner:NEW ZIGUANG GROUP CO LTD

Transform distributed computing parallel task scheduling-oriented accelerator

The invention discloses an accelerator architecture oriented to Transform distributed computing parallel task scheduling. The accelerator architecture comprises an input data controller, a systolic array operation block, a nonlinear SoftMax operation block, an intermediate data controller, a data flow controller, a multi-FPGA interconnection controller and an output controller. The accelerator has the following characteristics: the accelerator runs based on an independently developed instruction; systolic array operation can be flexibly adapted to matrix multiplication operation in the neural network; the intermediate data controller is compatible with BIAS calculation or BIAS calculation; the data flow controller supports operation data self-request and data flexible multiplexing; the multi-FPGA interconnection controller supports data communication between high-speed equipment and completes transmission of data and instructions; the output controller can configure an operation result storage position according to the instruction. According to the accelerator, instruction encapsulation and high-speed transmission are fully utilized, so that hardware is friendly and universal to a scheduling algorithm of software while reasoning acceleration is obtained, and load balance and efficient reasoning during distributed calculation are guaranteed.
Owner:BEIJING UNIV OF TECH

Memory access optimization device for convolutional neural network accelerator

The invention provides a memory access optimization device for a convolutional neural network accelerator, and belongs to the technical field of integrated circuit design, the memory access optimization device comprises a pulsation processing unit array which is used as a computing core, adopts a regular two-dimensional grid structure and comprises processing units, and each processing unit is provided with a basic multiplication and addition operation unit and a local register; the ifmaps second-level cache system adapts to locality of convolution operation through hierarchical scheduling, the reuse rate of an input feature map is improved, and redundant memory access is reduced; and a part of the sum register file is directly connected with the pulse processing unit array and is used for caching a part of the sum register file and an intermediate result generated in the calculation process. According to the method, a hierarchical data multiplexing system is established, so that the data multiplexing efficiency of the input feature map and the partial sum is remarkably improved on the premise of not excessively increasing the hardware complexity, and the storage access frequency and the system power consumption are effectively reduced.
Owner:YUANQIXIN (SHANDONG) SEMICONDUCTOR TECHNOLOGY CO LTD

A resource scheduling method for cloud rendering clusters based on AI distributed deployment

This invention provides a resource scheduling method for cloud rendering clusters based on AI distributed deployment. First, the original instruction stream is divided into instruction particles carrying energy consumption labels, latency labels, and complementarity weights. A cross-task instruction resonance graph is constructed to locate complementary candidate resonance particle groups. Placeholder particles are generated based on time-series prediction to fill upcoming instruction stream gaps. Particles are exchanged and spliced ​​across GPUs in a dynamic splicing network, solidifying into a continuous target instruction stream across GPUs. Simultaneously, real-time power consumption and performance indicators are collected, and a joint performance-energy consumption evaluation is performed on multiple candidate splicing schemes to select and apply the optimal scheme. Without compromising business correctness, this invention significantly improves GPU execution unit utilization, reduces end-to-end latency, and lowers energy consumption through continuous pipeline, zero-copy data reuse, and rollback anchor point guarantees. It is suitable for large-scale parallel scenarios such as film and television rendering, cloud gaming, and digital twins.
Owner:SHANGHAI ITHELP NETWORK TECH CO LTD

Cache aware kernel processing

PendingUS20260203220A1Data streamMulti processor
Retrieving data blocks from a memory cache on a processor unit to be used by more than one compute thread can be improved to increase data reuse and reduce power consumption. Conventionally, data blocks are retrieved using a standard data retrieval model regardless of how the data is reused by the multiple compute threads accessing that data. By dynamically using a combination of dataflow retrieval models, a more optimized process can be implemented increasing data reuse and lowering power consumption of the processor unit. The dataflow retrieval models can be adaptive swizzling, continuous rasterization, alternating k-order, periodic compute thread array synchronization, or explicit tile eviction. As the size and number of memory caches increase on a processing unit, as well as the number of logic units and streaming multiprocessors, these optimizations become more valuable to overall efficiency.
Owner:NVIDIA CORP