Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

45 results about "Algorithm acceleration" patented technology

Algorithm Acceleration. Algorithm acceleration uses code generation technology to generate fast executable code. Accelerated algorithms must comply with MATLAB ® Coder™ code generation requirements and rules.

Micro-isolation and differential encryption method and system based on industrial protocol perception and medium

The invention provides a micro-isolation and differential encryption method and system based on industrial protocol perception and a storage medium, and the method comprises the steps: firstly carrying out the static filtering of an original flow through a preset port, and obtaining a to-be-analyzed first communication message; analyzing a message behavior mode by using a time sequence feature analysis model, and extracting a semantic tag containing a function code, a data point address and a data value; if the data value exceeds the safety threshold value, real-time blocking and alarming are carried out; otherwise, generating a dynamic communication strategy based on the function code and the data point address; when the micro-isolation rule is triggered, dynamically limiting the access authority of the equipment; when the difference encryption rule is triggered, an encryption algorithm is selected according to the data sensitivity, and an algorithm acceleration module is called for execution. According to the method, micro-isolation is realized through protocol semantic analysis, and illegal function code injection is prevented; a differential encryption mechanism is adopted to reduce encryption overhead, and meanwhile, the real-time requirement of an industrial scene is met; in addition, encryption strength is dynamically adjusted in combination with channel quality, and operation reliability under extreme working conditions is guaranteed.
Owner:中亿(深圳)信息科技有限公司

State cryptographic algorithm acceleration system supporting SM4-ZUC dual-core dynamic collaboration

The invention discloses a national cryptographic algorithm acceleration system supporting SM4-ZUC dual-core dynamic collaboration, and belongs to the technical field of information security and hardware acceleration. The system comprises an AXI read-write module, a configuration register module, an SM4 algorithm module and a ZUC algorithm module. The configuration register receives and distributes information such as a secret key, an initial vector and a mode through an APB interface; the AXI read-write module supports burst transmission and dynamic scheduling data read-write; the SM4 algorithm module supports three modes of GCM, CTR and ECB; and the ZUC algorithm module is compatible with ZUC-128 / 256. And the corresponding ZUC-128 / 256 and SM4 algorithms can be dynamically selected for operation by judging an encryption mode in the ZUC and SM4 algorithms. The system adopts a three-level assembly line architecture, space-time multiplexing of read-write data and encryption is realized, and throughput rate is improved.
Owner:BEIJING TECH & BUSINESS UNIV

Hardware acceleration calculation method and system based on post quantum cryptography algorithm

The invention relates to a hardware acceleration calculation method and system based on a post quantum cryptography algorithm, and is applied to a chip. The method comprises the steps that hardware engines in a hardware acceleration engine array are adopted to execute post-quantum cryptographic algorithms of corresponding types, and the post-quantum cryptographic algorithms of the corresponding types comprise at least one of polynomial multiplication, multi-branch hash calculation, a tree structure and sparse polynomial operation; in the execution process, resources needed by the quantum cryptography algorithm after the hardware engines execute are dynamically adjusted based on the prediction mechanism and the fuzzy control logic, and the resources comprise at least one of the core number, the cache bandwidth, the voltage and the frequency. Namely, on one hand, a hardware engine is adopted to perform accelerated calculation on the post quantum cryptography algorithm, and on the other hand, resources required by calculation are dynamically adjusted during accelerated calculation, so that the resource utilization rate is increased, and calculation and energy efficiency regulation are optimized. Cooperative improvement of algorithm acceleration efficiency and performance can be realized.
Owner:CHINA SOUTHERN POWER GRID NEW POWER SYSTEM (BEIJING) RESEARCH INSTITUTE CO LTD

Pinyin decoding input method, embedded system and storage medium

The invention provides a pinyin decoding input method, an embedded system and a storage medium. The method comprises the following steps: constructing a hierarchical code table index database; obtaining a pinyin character string input by a user, and determining a first index of the pinyin character string in the initial index layer according to the initial of the pinyin character string; calculating a Hash value of the pinyin character string by using a Hash algorithm, and obtaining a second index corresponding to the Hash value in a Hash table; searching in the pinyin data block according to the second index to obtain a pointer corresponding to the pinyin character string; and matching in the word data layer by using the pointer to obtain corresponding word data. According to the method, the memory pressure of the embedded system is reduced, the three-layer index structure of the hierarchical code table index database is utilized, and the optimized hash search algorithm is adopted to accelerate data access, so that high-precision matching and quick positioning are realized in an embedded environment.
Owner:NINGBO FOTILE KITCHEN WARE CO LTD

National cryptographic algorithm acceleration system based on Numba just-in-time compiling technology

The invention relates to the technical field of national cryptographic algorithm acceleration, and discloses a national cryptographic algorithm acceleration system based on a Numba just-in-time compiling technology. The numerical calculation mode reconstruction module is used for performing numerical calculation mode reconstruction on SM2 elliptic curve scalar multiplication, SM3 message extension round function, SM4 nonlinear transformation and ZUC flow generation logic contained in the national cryptographic algorithm core calculation module by utilizing an LLVM compiling chain of Numba, and converting an interpretively executed Python code into an optimized machine code adaptive to hardware; key function compiling acceleration is achieved through a (at) njit decorator, hot spot operation is dynamically recognized through compiling scheduling, and a differential instruction optimization strategy is loaded; the platform shows multiple technical advantages: on the development efficiency level, by presetting a standardized acceleration module library and an automatic compiling tool chain, the integration complexity of a national cryptographic algorithm is greatly reduced; in the aspect of operation performance, the compilation optimization depth is better than that of a general interpreter acceleration scheme; in the aspect of safety controllability, the mathematical theory basis of the national cryptographic algorithm is completely reserved, and hidden risks possibly introduced by black box type hardware acceleration are avoided.
Owner:SOUTHWEST PETROLEUM UNIV

Target scattering characteristic efficient solving method of multilayer fast multi-pole collaborative principal component analysis compression algorithm

The invention discloses a target scattering characteristic efficient solving method of a multilayer fast multi-pole collaborative principal component analysis compression algorithm. The method comprises the following steps: dividing a target model by using a hybrid tree, changing an impedance matrix into a laminated matrix by binary tree grouping, and dividing a target into a near field and a far field by octree grouping; a multi-layer fast multi-pole algorithm acceleration principal component analysis compression algorithm is adopted to compress the low-rank impedance matrix; performing nearest layer-by-layer inversion on the impedance matrix according to an SMW formula; and finally, current is obtained through matrix vector multiplication of a series of inverse matrixes and a right vector, and target scattering characteristics are solved. The method achieves the solving of the electromagnetic scattering characteristics of the complex target, and can improve the calculation efficiency of the analysis of the electromagnetic scattering characteristics of the complex target.
Owner:NANJING UNIV OF SCI & TECH

Decision fusion algorithm acceleration operation device

The utility model relates to the technical field of data processing equipment, in particular to decision fusion algorithm acceleration operation equipment. According to the decision fusion algorithm acceleration operation equipment provided by the utility model, the technical problem of poor heat dissipation performance of the decision fusion algorithm acceleration operation equipment in the prior art is solved. The utility model discloses decision fusion algorithm acceleration operation equipment, which comprises a box body, the data processing assembly is arranged in the box body; the heat dissipation assembly is connected into the box body in a sliding mode, and the data processing assembly is installed on the heat dissipation assembly; by arranging the air cooler, the air inlet hole and the air outlet hole, the data processing assembly is cooled in a forced heat dissipation mode, and therefore the heat dissipation efficiency of the equipment is effectively improved. And the air cooler blows air flow with lower temperature to the data processing assembly, and the air flow with lower temperature exchanges heat with the data processing assembly and then is discharged in time through the air outlet holes, so that the heat dissipation efficiency of the equipment is greatly improved.
Owner:ZHEJIANG DAXIN TECH CO LTD

Biped robot nonlinear model predictive controller FPGA heterogeneous acceleration method

The invention relates to a biped robot nonlinear model predictive controller FPGA heterogeneous acceleration method, which comprises the following steps: establishing an NMPC model according to a biped robot dynamic model, iteratively solving a quadratic programming form of the NMPC model, and completing a single iteration process: solving by a primitive-dual interior point method and solving a symmetric linear equation set by a minres algorithm; the process of solving the solution of the linear equation set constructed by using the KKT condition based on the quadratic programming form through the primitive-dual interior point method is configured to be executed by a CPU; the process of solving the symmetric linear equation set by the minres algorithm is configured to be executed by the FPGA. According to different calculation characteristics of a solution operator in a biped robot NMPC model, a logic reusable minres algorithm in the solution process is deployed to an FPGA calculation platform, and nonlinear operation is deployed to a CPU for implementation, so that the calculation power characteristic of a heterogeneous chip is fully utilized. According to the method, robot dynamics modeling, an efficient optimization solution algorithm and appropriate computing power resource allocation are considered, and the acceleration effect and verification of the algorithm can be ensured.
Owner:江淮前沿技术协同创新中心

Independent data isolation system and method based on hardware-level physical isolation

The application provides an independent data isolation system based on hardware-level physical isolation, comprising: a special security processing unit which is a multi-core heterogeneous system-level chip, and integrated with the following through a hardware security bus: a national encryption algorithm acceleration core for realizing encryption and decryption and operation processing of national encryption algorithms; a general control core for executing task scheduling, inter-core communication and external interface protocol processing; a security storage core for generating a unique physical key and supporting quantum key injection; an isolation control core for controlling physical isolation of data flow through a hardware mutex and a bus tag controller; and a fuse self-destruction module connected with the special security processing unit and controlled by the isolation control core, for quickly destroying key data and paths. The application has the beneficial effects of optimizing data security isolation, key protection and emergency response, effectively improving the overall security of the system and resisting more complex attacks.
Owner:BEIJING ZHONGHAI WATSON MEDICAL TECHNOLOGY CO LTD

Acceleration method and device of post quantum cryptography algorithm based on ASIP architecture

The invention belongs to the technical field of cryptographic algorithm hardware implementation, and particularly discloses a post quantum cryptographic algorithm acceleration method and device based on an ASIP architecture, and the method comprises the steps: obtaining a program code of a Kyber algorithm; continuously executing the microcode instruction of the Kyber algorithm until the Kyber algorithm is completely executed; executing the microcode instruction of the Kyber algorithm, wherein the microcode instruction to be executed is obtained; analyzing the microcode instruction to be executed to obtain instruction analysis information; acquiring data of a source operand address as source data; on the basis of the operation code, an operation object is determined in hardware resources, input of the operation object is configured on the basis of the source data, and the hardware resources comprise an ALU and a Hash and sampling module; and configuring the data of the target operand address based on the data output by the operation object. According to the method and the device, the operation efficiency of the Kyber algorithm can be improved through hardware acceleration.
Owner:WUHAN SHIP COMM RES INST (NO 722 RES INST OF CHINA STATE SHIPBUILDING CORP)

Device and method for accelerating high-precision ADC data operation

The invention relates to the technical field of integrated circuit design, and particularly discloses a device and method for accelerating high-precision ADC data operation, and the device comprises a bus which comprises an address bus and a data bus; the hardware algorithm accelerator comprises an address mapping logic module, an input and output mapping module, a register module and an algorithm acceleration core function module; the address mapping logic module is used for splitting an address access request from an address bus into a read-write access command of the register module and a parameter and operation command of the algorithm acceleration core function module; the register module includes a configuration register, a plurality of operand registers, and a configuration register. According to the method and the device, the address access request from the address bus is divided into two parts, meanwhile, calculation input and output are selected through address mapping, a calculation result can be directly imported into a newly-initiated calculation input end, read-write of a processor to an operand register can be reduced, and the calculation efficiency is improved.
Owner:苏州领慧立芯科技有限公司

Universal interface algorithm acceleration device and acceleration method

The invention provides a universal interface algorithm acceleration device and an acceleration method. The universal interface algorithm acceleration device comprises a configuration unit used for performing information configuration on a to-be-executed software task; the algorithm scheduling acceleration unit is used for reading a software task, performing data analysis on the software task to obtain configuration information and data required by an algorithm, and writing the configuration information and the data into the algorithm module; the SPIRNG unit is used for acquiring a random number from an external random source according to a task requirement and writing the random number into the algorithm module; after algorithm calculation is finished, the algorithm scheduling acceleration unit obtains an algorithm end mark and reads an algorithm result, the algorithm result is placed at a result output address specified by the configuration unit, an algorithm state is updated, and the updated algorithm state is transmitted to the configuration unit; the configuration unit schedules the updated task state and the related address of the acceleration unit according to the algorithm so as to execute the task. According to the method, extra time except algorithm calculation can be reduced, and the actual application efficiency of the algorithm is improved.
Owner:TIANJIN C CORE TECH CO LTD

RAFT consensus algorithm acceleration system

The invention discloses an RAFT consensus algorithm acceleration system, which belongs to the technical field of block chain algorithm acceleration and comprises a CPU (central processing unit), a bus bridging system, a network communication module, an RAFT algorithm acceleration module, a memory and a network interface. The CPU is used for completing control of an acceleration system, scheduling data to the RAFT algorithm acceleration module and converting message data conforming to a message frame format; the bus bridging system is formed by combining a plurality of buses and is used for carrying out protocol conversion on data according to different data frame formats; the network communication module is used for communicating with the RAFT algorithm acceleration module; the RAFT algorithm acceleration module is used for maintaining nodes, performing voting and election initiating operations through the network communication module, and completing RAFT consensus acceleration; the memory is composed of a high-speed memory array. The implementation of the RAFT consensus algorithm in the block chain is completed, the time delay is greatly reduced, and the effect of optimizing the performance is achieved.
Owner:BEIJING TECH & BUSINESS UNIV

A GPU multi-thread parallel-based hawk algorithm acceleration method

ActiveCN122044804BComputational scienceRotation factor
The application provides a Hawk algorithm acceleration method based on GPU multi-thread parallelism. The Hawk algorithm acceleration method based on GPU multi-thread parallelism comprises the following steps: (1) thread resource division and task mapping; (2) kernel fusion and throughput peak detection; (3) NTT / iNTT parallelization reconstruction and memory access optimization; (4) FFT / iFFT structured parallelization optimization and butterfly operator fusion. The application realizes a coarse-grained parallel butterfly operator execution mode without complex address calculation, without shared memory synchronization, and without competition between threads, greatly improving the throughput efficiency on the GPU; meanwhile, the application unifies the parallel execution framework in the forward and inverse transformations of FFT / NTT, wherein the inverse transformation only needs to use the corresponding inverse rotation factor and perform simple normalization processing at the end to complete the overall recovery.
Owner:NANJING UNIV OF POSTS & TELECOMM

Game and optimization-based robot online motion planning method

The invention discloses a game and optimization-based robot online motion planning method, which comprises the following steps of: establishing a game model of a robot and a threat target, and outputting a game action instruction: acquiring current parameters of the robot and the threat target; defining a discrete action set according to the current parameter, and setting a game grade rule; judging the game grade, determining the robot game grade, and predicting the future trajectory of the threat target; based on a reward function, combining action space reduction and MCTS algorithm accelerated search, and selecting the first action of the optimal action sequence as a game action instruction; and optimizing the motion path on line according to a game result. According to the embodiment, the game theory and the optimization method are introduced, the motion path of the robot is adjusted in real time in the dynamic environment, and the task execution capacity and safety of the robot in the complex and dangerous environment can be effectively improved. The method not only considers the moving target of the robot, but also considers the strategy and dynamic change of the threatening target in the environment, and has higher adaptability and robustness.
Owner:NORTHWESTERN POLYTECHNICAL UNIV

Micro-isolation and differential encryption method, system and medium based on industrial protocol perception

The application provides a micro-isolation and differential encryption method and system based on industrial protocol awareness, a storage medium, and first acquires a first communication message to be analyzed by performing static filtering on original traffic through a preset port; then analyzes the message behavior mode using a time sequence feature analysis model to extract semantic tags containing function codes, data point addresses and data values; if the data values exceed a safety threshold, real-time blocking and warning are performed; otherwise, a dynamic communication strategy is generated based on the function codes and data point addresses; when a micro-isolation rule is triggered, the access permission of the equipment is dynamically limited; when a differential encryption rule is triggered, an encryption algorithm is selected according to the data sensitivity, and an algorithm acceleration module is called to execute. The application realizes micro-isolation through protocol semantic analysis to prevent illegal function code injection; adopts a differential encryption mechanism to reduce encryption overhead while meeting the real-time requirements of industrial scenarios; in addition, the encryption strength is dynamically adjusted in combination with channel quality to ensure operation reliability under extreme working conditions.
Owner:中亿(深圳)信息科技有限公司

A winograd-based correlation algorithm accelerator storage method

The present disclosure belongs to the technical field of neural network storage, and relates to a Winograd-based correlation algorithm accelerator storage method, which comprises the following steps: S1, acquiring the size of a correlation result matrix block and a real-time graph matrix block, and acquiring the size of a correlation result matrix and a real-time graph tensor and the channel parallelism of an acceleration unit; S2, storing a reference graph tensor block from off-chip storage to a first area of a reference tensor; S3, storing a real-time graph tensor block from off-chip storage to a real-time tensor cache; S4, reading data from the first area in the reference graph tensor cache and writing the last two rows of the read data to the first two rows in the second area in the reference graph tensor cache; S5, reading a tensor block from the reference tensor cache and prewriting the tensor block to a reference tensor register group; S6, writing a tensor block from the real-time graph tensor cache to a real-time graph tensor graph register; S7, moving the front column data of the reference register group to the rear column, and reading data from the reference tensor cache to the front column of the register group; and S8, writing a tensor register group after processing and calculation between different register groups.
Owner:BEIJING AEROSPACE AUTOMATIC CONTROL RES INST

GPU multi-thread parallel Hawk algorithm acceleration method

The invention provides a Hawk algorithm acceleration method based on GPU (Graphics Processing Unit) multi-thread parallel. The GPU multi-thread parallel-based Hawk algorithm acceleration method comprises the following steps of (1) thread resource division and task mapping; (2) kernel fusion and throughput peak detection; (3) NTT / iNTT parallel reconstruction and memory access optimization are carried out; and (4) carrying out FFT / iFFT structured parallel optimization and butterfly operator fusion. According to the method, a coarse-grained parallel butterfly operator execution mode which does not need complex address calculation, does not need shared memory synchronization and does not compete among threads is realized, and the throughput efficiency on the GPU is greatly improved; meanwhile, a parallel execution framework is unified in forward and reverse transformation of the FFT / NTT, and overall recovery can be completed only by using a corresponding reverse twiddle factor and performing simple normalization processing at the tail end in the reverse transformation.
Owner:NANJING UNIV OF POSTS & TELECOMM

High-speed parallel Viterbi decoding method based on CUDA (Compute Unified Device Architecture)

The invention discloses a high-speed parallel Viterbi decoder based on a CUDA (Compute Unified Device Architecture) and a decoding method. The decoder comprises an input cache module, a parallel data de-puncturing module, a parallel measurement calculation module, a backtracking path processing module and an output control module, and multi-level parallel processing is realized by utilizing a CUDA (Compute Unified Device Architecture) of a GPU (Graphics Processing Unit). The method comprises the following steps: segmenting a receiving sequence and then distributing the segmented receiving sequence to GPU thread blocks; rapidly reading a state transition matrix and an output matrix by adopting a pre-calculation table look-up strategy; adopting a parallel prefix scanning algorithm to accelerate state measurement calculation; and reducing the global memory access delay by using the shared memory. According to the invention, the performance bottleneck of the traditional serial Viterbi algorithm is broken through, on the premise of keeping the decoding precision, the throughput rate is obviously improved, the real-time processing of multi-code-rate and multi-code-stream convolutional codes is supported, and a remarkable acceleration effect can be obtained compared with CPU (Central Processing Unit) implementation. The method is especially suitable for high-throughput ACM scenes, such as 5G communication and satellite navigation, requiring real-time processing of multiple code streams and long constraint length convolutional codes.
Owner:HUNAN INSTITUTE OF SCIENCE AND TECHNOLOGY +1

Lightweight twofish encryption algorithm accelerator and acceleration method thereof

The application provides a lightweight Twofish encryption algorithm accelerator and an acceleration method thereof, wherein main modules include a controller module, a sub-key generation module, a round operation module and an input / output whitening module. The application provides a high-efficiency hardware acceleration circuit for realizing the S-box unit permutation function, and introduces a linear feedback shift register for randomly selecting a permutation circuit in the S-box in each round operation, so as to improve the security of the encryption process. The round operation module and the extended sub-key generation unit, which are two core parts, are highly shared in hardware resources, and are alternately operated according to the switching function of the control signal, so that the resource utilization is less, the hardware implementation scale is small and lightweight, and the module can be well adapted to the module integration in the SoC.
Owner:NANJING UNIV

An apparatus and method for accelerating high-precision ADC data operation

The application relates to the technical field of integrated circuit design, and particularly discloses an apparatus and method for accelerating high-precision ADC data operation, which comprises a bus including an address bus and a data bus; a hardware algorithm accelerator including an address mapping logic module, an input-output mapping module, a register module and an algorithm acceleration core function module; the address mapping logic module is used for splitting an address access request from the address bus into two parts of a read-write access command of the register module and parameters and operation commands of the algorithm acceleration core function module; and the register module comprises a configuration register, a plurality of operand registers and the configuration register. According to the application, the address access request from the address bus is split into two parts, the calculation input and output are selected through address mapping, the calculation result can be directly introduced into a newly initiated calculation input end, the read-write of the operand register by the processor can be reduced, and the operation efficiency is improved.
Owner:苏州领慧立芯科技有限公司

Data stream-oriented efficient pattern mining heterogeneous parallel acceleration method

The invention discloses a data flow-oriented efficient pattern mining heterogeneous parallel acceleration method, relates to the technical field of data mining and heterogeneous algorithm acceleration, and fully considers storage resource limitation, data access efficiency and calculation task balance. Through a hybrid task allocation strategy, a storage optimization strategy, an efficient data interaction mechanism and a customized parallel hotspot code IP core of dynamic and static combined parallel calculation, the calculation performance and the iteration efficiency of a PFUM-Miner algorithm in an efficient item set mining algorithm on a heterogeneous architecture are effectively improved; and a solid foundation is laid for subsequent heterogeneous transplantation research of efficient item set mining.
Owner:XIDIAN UNIV

Quantum nerve fusion patent drawing automatic generation and credibility verification system

The invention provides an automatic generation and credibility verification system for patent drawings of quantum neural fusion, which belongs to the technical field of verification, and is characterized by comprising a quantum calculation module for performing parallel calculation and processing high-complexity data and operation tasks by using superposition and entanglement characteristics of quantum bits to obtain a quantum calculation result; the method is particularly suitable for carrying out rapid feature extraction and pattern recognition on mass data and providing an efficient data processing basis for a subsequent verification process, and a quantum-classical cooperative computing architecture is adopted as follows: a quantum computing module realizes parallel feature extraction (for example, a Shor algorithm accelerates large number decomposition) by utilizing quantum gate operation; and converting the processed quantum state data into classical vectors through a quantum-classical interface, and inputting the classical vectors into a neural network module to carry out multilayer convolution learning so as to form an assembly line work of quantum rapid preprocessing and neural depth analysis.
Owner:李建业

Bit-serial computing transmission structure for high speed data processing in chip interconnect systems

The application discloses a bit serial computing transmission structure for high-speed data processing in a chip interconnection system, and belongs to the technical field of semiconductor integrated circuits. The structure mainly integrates two bit serial multiply-accumulate modules, a bit serial adder and a high-speed voltage mode logic VML physical interface. The VML interface mainly comprises a driver, an equalizer, a comparator, a bias current source and an RS flip-flop. The structure directly connects the computation and data transmission without passing through parallel-serial conversion or serial-parallel conversion circuits, and has the potential to reduce the delay and energy consumption. The scalable chip interconnection component can be the basis of a future modular algorithm accelerator. Different accelerator hardware structures are formed by expanding or reconfiguring the basic component, so that different algorithms are adapted, and the reconfigurability of the system is improved.
Owner:UNIV OF ELECTRONICS SCI & TECH OF CHINA

Pulse-space division parallel BP algorithm acceleration system based on FPGA

The invention relates to the technical field of radar, and provides an FPGA-based pulse-space segmentation parallel BP algorithm acceleration system, which is realized based on a full-pipeline FPGA parallel BP imaging architecture, and comprises a communication module, a scene partitioning module, a pulse-space segmentation module and a BP algorithm full-pipeline processing module. The communication module is used for transmitting radar original data corresponding to a target scene to the FPGA onboard DDR3 memory from an upper computer through a PCIE interface in a zero-copy manner; the scene partitioning module is used for decomposing a target scene into a plurality of sub-blocks; the pulse-space segmentation module is used for decoupling the imaging pulse and sub-blocks decomposed from a target scene into a plurality of parallel calculation units; the BP algorithm full-pipeline processing module is used for realizing full-pipeline processing of a BP algorithm between pulses for each calculation unit in parallel; and the communication module is also used for carrying out data fusion on the BP imaging results of all the sub-blocks to obtain a BP imaging result of the target scene and uploading the BP imaging result to an upper computer.
Owner:XIDIAN UNIV

Long and short time memory network computing chip and method for MIMO time sequence signals

The invention relates to the technical field of integrated circuit chip design and artificial intelligence, aims to solve the problems of poor acceleration calculation timeliness and high cost performance of an LSTM (Long Short Term Memory) time sequence algorithm in the existing long short-term memory network calculation chip, and provides a long short-term memory network calculation chip and method for MIMO (Multiple Input Multiple Output) time sequence signals. The long and short time memory network computing chip for the MIMO time sequence signals comprises a Sigmoid function computing module, wherein the Sigmoid function computing module computes output data and transmits the output data to a multiply-accumulate computing module; the data flow control module is used for providing an input data signal sequence and parameter data for the Sigmoid function calculation module, outputting data obtained by Sigmoid function calculation to the multiply-accumulate calculation module, and transmitting an output result to the output data module; the assembly line control module is used for controlling assembly line calculation of the long-short term memory network. The LSTM calculation efficiency can be improved, and the calculation cost can be reduced.
Owner:SHANDONG LINGNENG ELECTRONIC TECH CO LTD

Method and system for protecting instructions of a large model based on four-layer cooperative closed loop

The application discloses a kind of big model instruction security protection method, system and storage medium based on four-layer cooperative closed-loop protection mechanism, belong to artificial intelligence security and trusted computing field.The application is deployed non-intrusive independent protection layer between user input layer and big model inference execution layer, through instruction weight dynamic control, context integrity check, multi-modal public power subject credible verification, four big links cooperative linkage of security rule priority rigid locking, combined with shared semantic label library, dynamic threshold value synchronization mechanism and cross-layer abnormal feedback channel constructs whole-link closed-loop protection system.The application realizes high semantic weight content attenuation by instruction weight system, resists split bypass attack by context governance, guarantees public power identity credible by multi-modal double verification, realizes permission control and algorithm transparent by rule priority rigid locking, built-in global multi-jurisdiction compliance rule library, adapt to global regulatory requirements and have anti-black-box auditable characteristics.The application does not need to modify model bottom code, supports hardware-level isolation and national encryption algorithm acceleration, can be widely applied to government, finance, military industry and other high-end security scene, provides non-intrusive, global, high-security compliance protection solution for big model.
Owner:高鹏

Multi-core collaborative board-level signal processing system integrated with storage and calculation integrated unit

The invention belongs to the technical field of electronic information engineering, and particularly discloses a multi-core collaborative board-level signal processing system integrated with a storage and calculation integrated unit. The system comprises an FPGA, an ARM, an AI co-processing chip and a storage and calculation integrated chip. The FPGA is connected with the ARM, and the AI co-processing chip and the storage and calculation integrated core are respectively connected with the AMR and the FPGA; the FPGA is a core calculation engine, executes parallel data processing and algorithm acceleration, and responds to a complex signal processing task in real time; the AI co-processing chip processes a multi-thread task, executes neural network reasoning with complex calculation, processes mathematical operation and data analysis tasks and provides a calculation basis; the storage and calculation integrated chip directly carries out data processing in the storage unit; the embedded ARM processor is used for overall system control, task scheduling and resource management to ensure cooperative work of each unit. According to the scheme, the technical problem of processing performance improvement limitation caused by a single-core processor architecture is solved.
Owner:BEIJING INST OF REMOTE SENSING EQUIP

Miller-rabin algorithm acceleration method and device based on hardware and software fusion

The application relates to a Miller-Rabin algorithm acceleration method and device based on software and hardware fusion. When the Miller-Rabin algorithm is executed, target data to be verified is loaded into a data splitting hardware unit, the target data is subjected to data splitting by using the data splitting hardware unit, and a decomposition array matched with the target data is generated after data splitting; the decomposition array is used to generate modulus power operation object data, the modulus power operation object data is subjected to modulus power operation by using a modulus power operation hardware unit, a logic judgment unit is used to perform logic judgment on a first modulus power operation result value and a second modulus power operation result value, and a data verification state corresponding to the target data is generated based on the logic judgment state of the first modulus power operation result value and the second modulus power operation result value. The application can effectively improve the efficiency of large-batch prime number judgment by using the Miller-Rabin algorithm, reduce the algorithm power requirement of running the Miller-Rabin algorithm, and improve the prime number judgment and generation efficiency.
Owner:WUXI INST OF INTERCONNECT TECH CO LTD