Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

1615 results about "Hardware acceleration" patented technology

In computing, hardware acceleration is the use of computer hardware specially made to perform some functions more efficiently than is possible in software running on a general-purpose CPU. Any transformation of data or routine that can be computed, can be calculated purely in software running on a generic CPU, purely in custom-made hardware, or in some mix of both. An operation can be computed faster in application-specific hardware designed or programmed to compute the operation than specified in software and performed on a general-purpose computer processor. Each approach has advantages and disadvantages. The implementation of computing tasks in hardware to decrease latency and increase throughput is known as hardware acceleration.

AI Serving Hardware and Software Frontier Enhancements

A computer system implements a unified framework integrating an adaptive elastic funnel (AEF) with a convergent intelligence fabric (CIF) for multi-agent AI collaboration. The system provides a universal multi-modal key-value subsystem for sharing partial computations, implements hybrid placement strategies for dynamic memory management, and incorporates quantum-resistant secure enclaves. The architecture integrates hardware acceleration through GPU-FPGA hybrid caching and neuromorphic processors, applies adaptive energy and thermal management across hardware generations, and implements autonomous flash resource orchestration with multi-dimensional wear management. The system orchestrates tensor workflows using hierarchical scheduling, enables cross-agent collaboration with privacy preservation, and supports continuous learning without catastrophic forgetting. This integration delivers unprecedented computational efficiency and security in high-dimensional decision-making environments while supporting incremental adoption through modular interfaces.
Owner:QOMPLX INC

Large language model reasoning acceleration method and system based on dynamic video memory compression and memory isomerism

The invention discloses a big language model reasoning optimization method and system based on dynamic video memory compression and memory isomerism, and intelligent management of video memory resources is realized by integrating a dynamic compression strategy of KV Cache and a memory parallel architecture. The method comprises the following steps: 1) analyzing the spatial-temporal characteristics of the KV Cache in real time, adaptively selecting a quantization compression algorithm, a rarefaction algorithm or a low-rank decomposition algorithm, performing hierarchical storage based on attention head importance scores, keeping high precision of a core head, and implementing low-bit quantization on a secondary head; (2) the compressed inactive data are divided into a plurality of data blocks to be stored in a system memory, a parallel data channel group is established according to the number of physical channels, the compressed blocks are concurrently read through multiple channels during loading, and parallel decompression of a sparse matrix is accelerated through a GPU tensor core; and 3) constructing a KV Cache multiplexing mechanism and a parallel channel, and parallelizing a compression / decompression process and model calculation by adopting a hardware acceleration compression and asynchronous pipeline mechanism.
Owner:HANGZHOU AMTD YINGANG DIGITAL TECH CO LTD

Platform for orchestrating a scalable, privacy-enabled network of collaborative and negotiating agents utilizing modular hybrid computing architecture

A scalable platform for orchestrating networks of collaborative AI agents utilizing modular hybrid computing architecture. The platform integrates classical, quantum, and neuromorphic computing paradigms through hardware-accelerated translation layers and cross-paradigm coordination mechanisms. A central orchestration engine manages interactions between domain-specific AI agents, dynamically distributing workloads across heterogeneous computing cores based on task complexity, computational requirements, and resource availability. The platform employs hardware-accelerated translation between paradigms, enabling efficient cross-paradigm information exchange while maintaining semantic consistency and computation integrity across different architectures. Specialized monitoring and optimization systems continuously adjust resource allocation and fine-tune performance across computing paradigms. Advanced cache management and fault tolerance mechanisms ensure reliable operation, while privacy-preservation techniques enable secure collaboration. The platform's modular architecture supports integration of different computational approaches, enabling complex multi-domain problem solving that leverages the unique advantages of each paradigm while maintaining system-wide efficiency, scalability, and coherence.
Owner:QOMPLX INC

Hierarchical collaborative management method for virtual power plant based on multi-modal deep learning

The invention discloses a hierarchical collaborative management method and system for a virtual power plant based on multi-modal deep learning, and the method comprises the steps: constructing a four-dimensional data collection system, and achieving privacy enhancement preprocessing through federated learning and a differential privacy technology; a Bi-LSTM and a heterogeneous graph neural network are adopted to construct a three-mode deep fusion model, the weight is dynamically adjusted in combination with an environment-user dual-drive attention mechanism, and the load prediction precision and the space resource utilization rate are improved; a multi-target scheduling strategy is generated based on a five-dimensional target function and an improved DDPG algorithm, and physical feasibility is ensured through digital twinborn pre-verification; efficient execution and excitation transparency are realized through edge layer FPGA + NPU hardware acceleration and block chain evidence storage; and constructing a user participation ecology by using a natural language interaction strategy engine and a stepped incentive mechanism. The power grid economy, the equipment reliability and the user participation degree are remarkably improved, and intelligent upgrading of the virtual power plant is promoted.
Owner:TIANSHENGQIAO FIRST-CLASS HYDROPOWER DEV CO LTD HYDROPOWER PLANT

Biding document multi-mode duplicate checking method and system based on large model

The invention belongs to the technical field of natural language processing and information retrieval. The invention provides a bidding document multi-modal duplicate checking method and system based on a large model, and the method comprises the steps: carrying out the structural analysis and multi-modal feature extraction of an input bidding document, and generating the feature representation of semantic blocks and non-text elements; performing deep semantic coding on text blocks by using a dynamic context-aware large language model, and retaining logic relevance of a long text in combination with hierarchical position coding; efficient matching of massive semantic vectors is achieved through a mixed retrieval framework, and calculation efficiency is optimized in combination with a distributed calculation framework and a hardware acceleration instruction; and finally, performing structured feature reconstruction on non-text contents such as tables, charts and the like to realize cross-modal semantic association analysis.
Owner:INSPUR GENERSOFT CO LTD

Unmanned aerial vehicle intelligent monitoring system for forest protection

The invention relates to the field of forest protection, and discloses an unmanned aerial vehicle intelligent monitoring system for forest protection. Comprising a multispectral data dynamic acquisition module, an environment field real-time modeling module, an image intelligent enhancement module, a multi-scale frequency domain feature extraction module, a Bayesian anomaly probability inference module, a space-time correlation verification module, an unmanned aerial vehicle cluster scheduling module and a heterogeneous hardware acceleration module. According to the system, real-time monitoring and accurate identification of the hidden ecological damage behavior are realized through a closed-loop process of multi-modal sensor collaborative acquisition, environmental parameter dynamic prediction, image illumination invariance transformation, cross-scale texture fingerprint extraction, particle filter dynamic threshold optimization, multi-constraint space-time clustering, resource allocation optimization and reconfigurable hardware acceleration. According to the invention, on the basis of an image preprocessing architecture based on illumination reflection separation and edge preserving enhancement, the feature distortion bottleneck of a traditional algorithm in a complex illumination environment is broken through, and a high-fidelity data base is provided for hidden ecological damage detection.
Owner:腾冲市曲石镇综合保障和技术服务中心 +1

Regional energy internet hierarchical control method considering flexible load

The invention discloses a regional energy internet hierarchical control method considering flexible loads, which comprises the following steps of: constructing a global layer, dividing a regional energy internet into a plurality of electrically strongly coupled sub-regions by adopting a dynamic partitioning algorithm, aggregating equivalent adjustable capacities of the flexible loads in the sub-regions, and performing rolling optimization in an hour-level time scale, an improved Benders decomposition algorithm is adopted to accelerate solving and generate a partition-level power regulation instruction; constructing a regional layer, designing a thermodynamic-economic hybrid constraint model according to flexible load physical characteristics and user behavior elasticity, introducing a user comfort recovery time constraint and elastic electricity price response function, and decomposing an adjustment instruction issued by the global layer to each flexible load cluster in a minute-level time scale by adopting a parallel ADMM optimization algorithm; and constructing a load layer, constructing a millisecond-level adjustable load priority switching mechanism based on FPGA hardware acceleration and finite-state machine logic, calculating a load adjustable power boundary in real time, and dynamically adjusting a response queue in combination with system frequency deviation.
Owner:STATE GRID CORPORATION OF CHINA +2

Multistage compression collaborative optimization neural network deployment method and device based on memristor and storage medium

The invention relates to the field of artificial intelligence hardware acceleration, and discloses a memristor-based multilevel compression collaborative optimization neural network deployment method and device, and a storage medium. The method comprises the following steps: carrying out progressive structured pruning on a pre-training model based on a dual-drive scoring mechanism of an L1 norm and gradient sensitivity and hardware feedback, and generating a hardware-friendly sparse weight structure; the characteristics of the memristor are simulated through a micro-nonlinear conductance modeling function, and network weight and conductance parameters are synchronously optimized to reduce errors by adopting mixed precision quantification of four bits of a convolutional layer and two bits of a full-connection layer; a conductance drift and read-write noise model is injected, and the anti-interference capability of the model is improved in combination with adaptive noise enhancement and KL divergence loss; and mapping the optimized model to a memristor memory architecture to complete weight coding and reasoning. Through collaborative optimization of pruning, quantification and distillation, the problems of insufficient storage density, non-ideal characteristic interference and algorithm and hardware mismatch in memristor deployment are solved.
Owner:YANGTZE DELTA REGION INST (QUZHOU) UNIV OF ELECTRONIC SCI & TECH OF CHINA

Sampling implementation method and system suitable for motor protection quick response

The invention discloses a sampling implementation method and system suitable for motor protection quick response, belongs to the technical field of motor protection, and can increase the sampling frequency and improve the quick response of motor fault protection. Comprising the steps that motor phase current and bus voltage are collected, after anti-aliasing filtering is conducted, parallel sampling is conducted through a main ADC module and a redundant ADC module which are independently configured, and the following operations are synchronously executed based on a double-buffer-area framework: in a first buffer area, a compensation coefficient updating period is dynamically adjusted, and convergence is restrained; in the second buffer area, generating a voltage gradient prediction sequence, and compensating data abrupt change; when the harmonic frequency spectrum amplitude or the voltage gradient prediction sequence exceeds the limit, an oversampling mode is triggered, a redundant ADC module is allocated to load a multi-order digital filter, and an independent hardware acceleration unit is started to process harmonic compensation and gradient prediction in parallel; and generating a dynamic safety envelope, and if the gradient extreme value in the preset continuous window breaks through the safety envelope, generating an overvoltage fault signal.
Owner:THE 704TH RES INST OF CHINA STATE SHIPBUILDING CORP

Distributed computing-based intelligent design method and system for electric power engineering cloud resources

The invention relates to the technical field of intelligent power grids, in particular to an intelligent design method and system for electric power engineering cloud resources based on distributed computing. Comprising the following steps: constructing an electric power engineering multi-dimensional resource modeling system, dividing calculation nodes into three types of heterogeneous resource units including a real-time processing core, a data analysis core and a disaster recovery backup core, and establishing a dynamic attribute matrix containing time delay sensitivity, an energy consumption coefficient and a risk assessment value; the real-time processing core is configured with a hardware acceleration instruction set and supports microsecond response; a self-adaptive dynamic clustering algorithm is adopted, elastic resource clusters facing task requirements are generated according to the spatial topological relation of resource units and the load change trend, a quantum genetic optimization mechanism is introduced in the clustering process to dynamically adjust the inter-cluster coupling degree, and a cross-cluster communication link based on credibility evaluation is established; according to the design, the problems of low resource utilization efficiency and service quality degradation caused by a static allocation mode can be fundamentally solved.
Owner:ZHONGKE WANYING POWER GRP CO LTD

Integrated AI-driven and compliance-aware multi-state encoding framework

The present invention relates to adaptive multi-state encoding and processing in virtualized computing environments. The system includes a virtualized state selection module that dynamically transitions between binary, ternary, quaternary, and higher-order encoding states based on workload, bandwidth, security posture, and compliance requirements. A virtual encoding engine utilizes hardware-accelerated components, such as vFPGAs, vGPUs, and cTPUs, to enhance encoding throughput. A compliance-driven feedback controller continuously monitors encoding efficiency, threat levels, and adherence to mandates such as GDPR, HIPAA, and FIPS 140-3. Additional features include AI-based anomaly detection, federated model refinement, distributed ledger-backed audit trails, and quantum-resistant encoding techniques. By integrating intelligent encoding decisions with scalable compliance enforcement, the system enables high-performance, secure, and regulation-ready data processing across distributed cloud and edge environments, delivering measurable improvements in system responsiveness, data integrity, and operational trust.
Owner:SGM INFOTECH LLC

Method for NPU firmware to support multi-model fast switching

The invention discloses a method for supporting fast switching of multiple models by NPU (Network Processing Unit) firmware. According to the invention, the system real-time performance and the resource utilization rate in a multi-task scene are obviously improved. Through a dynamic hierarchical caching strategy and a priority scheduling mechanism, the system can intelligently allocate cache resources and preferentially guarantee rapid loading of a high-frequency and high-urgency model, for example, in an automatic driving scene, switching delay of a path planning model can be reduced to a millisecond level, and non-perceptual switching of key tasks is guaranteed. The incremental parameter loading technology and firmware-level context management are deeply fused, only model difference data are transmitted, and hardware is utilized to accelerate and recover a calculation state, so that bandwidth waste caused by traditional full-amount loading is greatly reduced, the model switching efficiency of edge computing equipment is improved by more than 10 times when multiple tasks such as voice recognition and image processing are carried out in parallel, and the efficiency of the edge computing equipment is improved. And meanwhile, the calculation precision is kept lossless.
Owner:SUZHOU SUXIAN MICROELECTRONICS TECH CO LTD

Inference acceleration method and device, electronic equipment and storage medium

The invention relates to the technical field of computers, in particular to a reasoning acceleration method and device, electronic equipment and a storage medium, and is used for improving the reasoning speed when a model executes a task. The method comprises the steps that when a target model is adopted to execute a to-be-reasoned task, a matched reasoning mode is activated according to hardware state information of computing resources occupied by the target model; when the reasoning mode is an instruction acceleration mode, at least one round of reasoning operation is executed, and each round of reasoning operation comprises the steps that storage continuity check is conducted in computing resources for reasoning dependency data related to the input sequence of the round; the reasoning dependency data is generated and stored through the reasoning operation of the previous round; when the storage position of the reasoning dependency data is in a non-continuous state in the computing resources, the storage position is adjusted to be in a continuous state by rearranging the reasoning dependency data; and calling a specified hardware acceleration instruction to load the adjusted reasoning dependency data, and determining the reasoning result of the round based on the loading result.
Owner:TENCENT TECH (BEIJING) CO LTD

AI-powered anomaly detection system for high-volume managed file transfers

ActiveDE202025102388U1Platform integrity maintainanceTransmissionManaged file transferEdge node
A real-time anomaly detection system for high-volume managed file transfers (MFT), consisting of: a secure, hardware-accelerated monitoring unit configured to interface with a managed file transfer server and intercept file transfer session data at wire-speed; a metadata extraction engine embedded in the hardware-accelerated unit, the metadata extraction engine configured to analyze protocol-specific session attributes, including, but not limited to, file size, transfer duration, encryption status, source and destination endpoints, transfer frequency, and payload entropy; a contextual AI inference engine communicatively coupled to the metadata extraction engine, the AI inference engine comprising a deep learning model trained on labeled historical MFT activity logs to detect contextual deviations from normative behavior; a federated learning architecture with a plurality of edge nodes, each hosting a local anomaly detection model trained on localized transmission metadata and configured to synchronize with a central aggregator using differentially private gradient updates; an Explainable AI (XAI) subsystem integrated into and configured to generate human-readable anomaly justifications, feature importance maps, and threat categorization labels; and a policy orchestration module configured to dynamically execute pre-configured or AI-based security responses, where the security responses include selective session termination, quarantining of transferred files, generation of alerts, or redirection of MFT workflows.
Owner:CHELLU RAGHAVA ALPHARETTA

Method and system for optimizing mold opening and closing parameters of injection molding machine for energy conservation and efficiency improvement

The invention provides an injection molding machine mold opening and closing parameter optimization method and system oriented to energy conservation and efficiency improvement, and relates to the technical field of industrial process intelligent control. The method comprises the steps that mold temperature, hydraulic pressure and vibration signals are collected in real time through a multi-source sensor, and a composite optimization target of a dynamic energy consumption coefficient and a process stability index is constructed; an improved multi-target genetic algorithm is adopted to iteratively generate a candidate parameter set meeting energy consumption and stability constraints, and the algorithm fuses hydraulic system safety boundary dynamic constraints and an adaptive variation strategy; a parameter adjustment effect is rehearsed through a digital twin model, and motion track deviation is calibrated online in combination with a recursive least square method, so that parameter dynamic injection is realized; and triggering abnormal working condition recovery and material self-adaptive optimization based on a closed-loop feedback mechanism. The system comprises a multi-source data acquisition module, a composite optimization calculation module, a dynamic parameter injection module and a closed-loop feedback control module, and the real-time performance and generalization ability are improved through an FPGA hardware acceleration optimization engine and a transfer learning algorithm.
Owner:ZHEJIANG JIEZHONG SCI & TECH CO LTD

Agricultural spraying robot system with visual navigation function

The invention relates to the field of agricultural robots, and discloses an agricultural spraying robot system with visual navigation, and the system comprises a multi-mode sensing module, a dynamic environment modeling module, a time sequence prediction module, a multi-target path planning module, a spraying control module, and a hardware acceleration unit. The method comprises the steps that a dynamic semantic map is constructed through multi-modal perception, agricultural semantics and an attention mechanism are fused to predict an obstacle trajectory, a safe path is generated in combination with multi-objective optimization, the spraying amount is adjusted based on crop density, and real-time operation is ensured by means of hardware acceleration. According to the method, a dynamic semantic map is constructed through multi-modal perception, agricultural semantics is combined to enhance trajectory prediction, a safe path is generated by utilizing multi-objective optimization, the spraying amount is adaptively adjusted according to the crop density, and real-time control is realized by adopting heterogeneous hardware acceleration; the defects of a traditional agricultural robot in the aspects of navigation precision, obstacle avoidance safety, resource utilization and response speed are overcome, and the intelligent level of complex farmland operation is improved.
Owner:DONGYING SAMLEE ELECTRONIC INFORMATION TECH CO LTD

Internet of Things test platform adaptation method, system and device based on post quantum cryptography

The invention discloses an Internet of Things test platform adaptation method, system and device based on post quantum cryptography. The method comprises the following steps: constructing a quantum threat perception model; establishing a multi-dimensional adaptive decision matrix, and generating a candidate post-quantum cryptography algorithm set; dynamic injection of a cryptographic protocol is executed, and lossless replacement of a communication layer encryption suite is achieved; deploying a runtime performance monitoring engine, and collecting encryption and decryption delay, memory occupancy rate and electromagnetic radiation characteristic data in real time; and constructing an algorithm switching decision tree based on reinforcement learning, and when a quantum attack feature or an equipment performance threshold is detected to break through, triggering cryptographic algorithm dynamic migration. A quantum threat perception driven Internet of Things security adaptation system is constructed, and by introducing a dynamic password switching mechanism, a hardware acceleration optimization scheme and a trusted execution environment, the security upgrading problem of traditional Internet of Things equipment in the post-quantum era is solved. According to the invention, the execution efficiency of the PQC algorithm of the resource-constrained equipment can be improved, and the key updating delay is reduced.
Owner:HANGZHOU YUNXIANG NETWORK TECH

Federal knowledge retrieval and big language model enhancement system and method

The invention relates to a federal knowledge retrieval and large language model enhancement system and method. The system comprises a server and a hardware accelerator. The server preprocesses an input query and extracts a head word of the query; the server retrieves the head word in the caching process through a cache module of the server, and if the cache of the cache module is not hit, the server sends a retrieval instruction to the hardware accelerator so as to retrieve in a local cache of the hardware accelerator; if the local cache is still not hit, the hardware accelerator generates parallel subtasks associated with the head word in a mode of segmenting a local knowledge graph, so that deep search is carried out, and noise is added into knowledge items retrieved by the parallel subtasks to protect sensitive information; and the hardware accelerator integrates the knowledge items added with the noise, performs reasoning through a large language model and generates an enhanced answer. According to the method, a two-stage dynamic cache system is constructed, the cross-device communication frequency can be effectively reduced, the query request is responded preferentially through the local high-frequency cache, and the overall network load of the system is reduced.
Owner:HUAZHONG UNIV OF SCI & TECH

Hardware-accelerated flexible steering rules over service function chaining (SFC)

Technologies for configuring flexible hardware-accelerated rules in a Service Function Chaining (SFC) architecture are described. A DPU includes an acceleration hardware engine to provide a single accelerated data plane, and a processing device that generates a first virtual bridge and a second virtual bridge. The first virtual bridge is controlled by a first network service hosted on the DPU and has a first set of one or more network rules. The second virtual bridge has a second set of one or more user-defined network rules. The processing device generates a combined set of network rules based on the first set of one or more network rules and the second set of one or more user-defined network rules. The acceleration hardware engine processes network traffic data in the single accelerated data plane using the combined set of network rules.
Owner:MELLANOX TECHNOLOGIES LTD(IL)

Low-power-consumption control method and system of controller

The invention relates to the field of energy-saving computing, and discloses a low-power-consumption control method and system of a controller, which are used for realizing prediction-sleep-wake-up-response full-link hardware in the controller. Time sequence state data of a controlled object is collected in real time, and a predictive sleep window value is dynamically generated by utilizing a predictive cooperative computing unit, so that intelligent scheduling of sleep and awakening is realized, and the power consumption of a controller is effectively reduced. Meanwhile, in combination with the technologies of threshold comparison, DMA data transmission optimization and the like, the data processing efficiency is improved, and invalid data transmission is reduced. And on the awakening mechanism, an asynchronous awakening circuit and an interrupt request mechanism are adopted to ensure that power supply is quickly recovered and control response is executed when an effective instruction is received, so that the real-time performance of the system is guaranteed. According to the invention, a prediction algorithm, hardware acceleration and an intelligent scheduling technology are fused, the performance of the controller is ensured, the power consumption is obviously reduced, and the endurance time of equipment is prolonged.
Owner:WUXI DENVEL INTELLIGENT ELECTRONIC INC

Visual precise plastic mold surface milling system and milling method

The invention discloses a visual precise plastic mold surface milling system and method, and the method comprises the steps: obtaining and analyzing the surface topography data of a milling region and a tool displacement track in real time through an optical sensor, and extracting wear-related micro-deformation features through multi-band filtering processing; generating a deviation vector matrix through three-dimensional mapping of the deviation vector matrix and a reference curved surface, and realizing dynamic calculation and path correction of the cutter offset; the high efficiency and stability of the compensation cutting process are ensured by combining the self-adaptive adjustment of the rotating speed and the feeding rate of the main shaft; meanwhile, the compensation effect is projected to a visual interface in real time to form a comparison layer, hardware acceleration is started to maintain the system response speed when the frame rate fluctuates, and finally the material removal volume is evaluated according to the deviation vector matrix to intelligently control the machining cycle termination time. Therefore, the problem of microtopography deviation caused by tool abrasion in the milling process of the complex curved surface mold is effectively solved, and the machining precision and efficiency as well as the robustness and the automation level of the system are remarkably improved.
Owner:SHENZHEN XUDONG JINXIN TECH CO LTD

Hardware-accelerated policy-based routing (PBR) over service function chaining (SFC)

Technologies for creating an optimized and accelerated network pipeline using a network pipeline abstraction layer (NPAL) for policy-based routing (PBR) over Service Function Chaining (SFC) are described. A DPU includes acceleration hardware engine to provide a single accelerated data plane. A processing device can generate a first virtual bridge and a second virtual bridge, the first virtual bridge to be controlled by a first network service hosted on the DPU and having a set of one or more network rules, and the second virtual bridge having a policy-based routing policy (PBR policy). The processing device can add the virtual port between the first virtual bridge and the second virtual bridge. The acceleration hardware engine, in the single accelerated data plane, can route network traffic data using the PBR policy and process the network traffic data using the set of one or more network rules.
Owner:MELLANOX TECHNOLOGIES LTD(IL)

Rendering equipment based on three-dimensional Gaussian sputtering

The invention discloses rendering equipment based on three-dimensional Gaussian sputtering. The hardware architecture includes a memory interface, a pre-processing engine, a cardinality ordering engine, a rasterization engine, and the like. The preprocessing engine converts the three-dimensional Gaussian point into a two-dimensional representation and generates a key value pair containing a tile identifier and depth information; the cardinal number sorting engine performs staged efficient sorting through a parallel processing mechanism and a BRAM alternate working mode; and the rasterization engine performs rendering processing by adopting a full-pipeline design and an early-stage stopping mechanism. According to the method, parallel execution of depth sequencing and rasterization is achieved, memory access and calculation redundancy are reduced through optimized data flow design, an efficient and low-power-consumption hardware acceleration solution is provided for three-dimensional Gaussian sputtering rendering, and the method is particularly suitable for application scenes such as virtual reality, games and scientific visualization requiring real-time rendering.
Owner:TSINGHUA UNIVERSITY

Electric power business data analysis processing method and system

The invention provides a power business data analysis processing method and system. The method comprises the following steps: receiving a business expansion work order containing a structured text and an unstructured image through a business hall terminal; analyzing text key entity fields, extracting image anti-counterfeiting feature points, and mapping the image anti-counterfeiting feature points to a unified vector space through a multi-modal alignment layer to generate semantic image feature flow; loading to a hardware acceleration chip to execute parallel matrix operation, identifying a logic binding relation between entities, and outputting a nested relation topological graph; calling a historical work order decision path to train a reinforcement learning agent to generate a control strategy; comparing the similarity between the to-be-anti-fake feature point and a pre-stored template, and outputting a verification result containing a missing identifier and anti-fake abnormity; and if the field is complete and the anti-counterfeiting similarity is greater than or equal to a threshold value, releasing the work order to a service system, otherwise, freezing the work order and triggering an alarm. According to the invention, the processing efficiency is improved, the abnormal generation risk is blocked, and automatic anti-counterfeiting verification and flow control of the electric power work order are realized.
Owner:BEIJING SHUYANG SMART TECH CO LTD

A method of optimizing linear transformation

A method and system for optimizing compute runtime and memory footprint of a linear transformation process are provided. The method includes determining a set of optimal rotation parameters, wherein the optimal rotation parameters provide an optimal tradeoff between runtime compute resources and a memory footprint for a runtime execution of the linear transformation process; initializing the linear transformation process to run a boosting technique with the determined set of optimal rotation parameters, wherein the boosting technique, when executed at runtime as part of the linear transformation process, performs at least one iteration that yields rotated ciphertexts, and wherein the at least one iteration is based on the determined set optimal rotation parameters and at least one key switching key (KSK); and loading the initialized linear transformation process to an internal memory of a hardware accelerator.
Owner:CHAIN REACTION LTD

AI compiler and compiling method based on multistage intermediate representation framework

PendingCN120560627ABiological modelsIntelligent editorsActivation functionComposite operator
The invention relates to the technical field of artificial intelligence compilers, in particular to an AI compiler and compiling method based on a multi-level intermediate representation framework, and the compiler comprises a high-level semantic retention layer which converts models of different AI frameworks into Lalg-on-Tensor IR intermediate representations; the hardware perception optimization layer comprises a tensor packaging and propagation module which is used for performing block packaging, layout propagation and folding of redundant packaging / unpackaging operation on the input tensor; the dynamic partitioning module is used for automatically selecting the partitioning size based on the cache capacity and the core number of the target hardware; the microkernel fusion module is used for fusing matrix multiplication, bias addition and an activation function into a single composite operator; and the microkernel collaboration layer is in butt joint with the hardware acceleration library through the XSMM dialect to generate a target hardware code. The hardware perception optimization layer can perform optimization according to different hardware characteristics, so that codes generated by the compiler can better adapt to target hardware, and the hardware utilization rate is improved.
Owner:SHANDONG INSPUR SCI RES INST CO LTD

Miniature data transaction device and intelligent transaction system thereof

The invention relates to the field of edge computing and Internet of Things, and discloses a miniature data transaction device which comprises an edge computing unit used for executing a centralized or decentralized matchmaking algorithm and a transaction protocol; the multi-dimensional detection module comprises a compliance detection unit, a safety detection unit, a quality evaluation unit and a scene applicability evaluation unit; the hardware accelerator is used for accelerating a detection task and encryption operation; the communication module supports multi-protocol data transmission; the storage unit is used for caching transaction data and intelligent contract codes; the power management unit is used for controlling the total power consumption of the equipment, and the edge computing unit comprises a hierarchical trusted execution environment architecture consisting of a safe enclave and a general computing unit. According to the method, a hardware isolation architecture of the secure enclave and the general computing unit is adopted, and the dynamic key encryption shared memory technology is combined, so that the effects of privacy computing and efficient matching parallel execution are achieved.
Owner:CHENGDU PATZHILIHU DIGITAL TECHNOLOGY CO LTD

Configurable number-theory transformation parallel computing acceleration method and device for post quantum cryptography algorithm

The invention discloses a configurable number-theory transformation parallel computing acceleration method and device for a post quantum cryptographic algorithm, and relates to the technical field of cryptographic algorithms. The number theory transformation is a calculation bottleneck in lattice-based post-quantum cryptography and PQC; a hardware accelerator specially designed for number theory transformation (NTT) is an effective solution for improving the execution speed. At present, a mainstream design neglects the influence of the scale of a calculation array, so that the design of an NTT calculation circuit is poor in flexibility and low in memory utilization rate, and a two-dimensional reconfigurable number theory transformation acceleration circuit is provided; the method is applied to a key generation stage, an encryption stage and a decryption stage of the post-quantum cryptographic algorithm. The NTT reconfigurable computing circuit provided by the invention can keep the utilization rate of hardware resources, and is suitable for mobile terminal equipment with limited resources; the NTT circuit adopts a low-complexity memory mapping scheme, so that the address control logic is greatly simplified, and the hardware overhead is reduced.
Owner:BEIJING INST OF TECH +1

End side model reasoning method and device based on RWKV architecture, electronic equipment and storage medium

The invention provides an end side model reasoning method and device based on an RWKV architecture, electronic equipment and a storage medium, and the method comprises the steps: obtaining a target input request of a target object, and converting the target input request into target model input data; loading a historical reasoning state corresponding to the target input request in a preset state storage space; determining a corresponding RWKV core operator according to the hardware platform type of the terminal equipment; based on the RWKV core operator, performing reasoning calculation on the target model input data and the historical reasoning state to obtain an output token sequence; wherein in the reasoning calculation process, the real-time reasoning state of the large language model is stored in a preset state accelerator memory for multiplexing; converting the output token sequence into a text format and outputting the output token sequence; and updating the historical reasoning state according to the real-time reasoning state after reasoning calculation. According to the method, calculation optimization and hardware acceleration can be carried out on the large language model of the RWKV architecture, so that the reasoning performance of the RWKV architecture model is improved on the end side.
Owner:SHENZHEN YUANSHI INTELLIGENT CO LTD