Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

26 results about "Critical path delay" patented technology

Critical path screening method and system based on chip time sequence report

The invention relates to the technical field of chip design, and provides a critical path screening method based on a chip time sequence report, which comprises the following steps: collecting chip time sequence report data, screening out first N critical paths from the data based on time sequence index data, and establishing a path ID-delay type-scene three-dimensional data model; and calculating the delay ratio of the unit delay to the interconnection line delay in each critical path. Performing single-type delay, type relevance and RC condition influence comparison on different scenes to obtain delay characteristic data. The invention discloses a system for executing the method, which can avoid full data redundancy and improve the analysis efficiency by screening the first N critical paths. And the delay change rule in different scenes is quantified. The delay characteristic data reflects the actual influence of the key path delay on the time sequence, so that a designer can quickly position the delay bottleneck of the key path, and a basis and a reference are provided for the targeted optimization of the total delay of the chip in the later period.
Owner:CHINA NAT INST OF STANDARDIZATION

Distributed dynamic time sequence monitoring circuit and self-adaptive adjusting system

The invention discloses a distributed dynamic time sequence monitoring circuit and a self-adaptive adjusting system, which comprise a plurality of monitoring units deployed in different IR voltage drop hot spot areas of a chip, each monitoring unit is configured as a delay chain, and the delay of the delay chain is determined by the local power supply voltage and temperature of the area where the monitoring unit is located. The delay variation is equal to the delay variation of the longest key path determined by simulation; the input end of the dynamic aggregation unit is connected with the monitoring unit, and the dynamic aggregation unit is used for comparing delays of all the input signals and outputting the input signal with the longest delay; a central monitor including a configurable delay pad circuit for providing a reference delay; the input end of the delay gasket circuit is connected with the output end of the dynamic aggregation unit, and the delay of the input signal with the longest delay and the reference delay are summed to obtain the total delay of the virtual key path for the real-time longest key path delay of the mirror image chip. According to the invention, non-intrusive monitoring can be realized, and all possible IR voltage drop hot spot areas can be fully covered.
Owner:SOUTHEAST UNIV

Method and device for merging and decomposing FPGA lookup table, equipment and medium

The invention provides an FPGA lookup table merging and decomposing method and device, equipment and a medium, and relates to the technical field of integrated circuit netlist optimization. The method comprises the following steps: acquiring a mapped LUT network of an FPGA; then, candidate pair identification is carried out on the LUT network, conflict-free LUT candidate pairs with the optimal weight are selected to be merged, an LUT connection relation graph is constructed after merging, and a key path is identified; on the basis of the LUT connection relation graph, decomposing the LUT on a non-critical path into smaller cascaded LUT units according to a LUT priority sequence; and dynamically updating the network topology structure and time sequence information of the LUT connection relation graph after each decomposition to complete optimization of the FPGA netlist. According to the method, the number of LUTs can be remarkably reduced and the resource utilization rate of the FPGA can be improved on the premise of ensuring that the delay of the critical path is not changed.
Owner:XIAMEN UNIV OF TECH

A three-stage positive feedback comparator circuit based on fast timing and electronic equipment

This invention relates to the field of integrated circuit technology and discloses a three-stage positive feedback comparator circuit and electronic device based on fast timing. The circuit includes a first-stage positive feedback module, a second-stage triggering and clocking module, a second-stage positive feedback amplification module, and a third-stage positive feedback latching module. By employing a three-stage positive feedback structure, this invention places the corresponding differential signal nodes in the same feedback loop at each stage, significantly enhancing the comparator's amplification rate and metastability suppression capability. Furthermore, by introducing a fast timing control mechanism, the second-stage amplification, clock reset generation, and third-stage latching operations can be executed in parallel, effectively shortening the critical path delay and significantly improving the overall comparison speed. The electronic device includes the aforementioned comparator circuit. This invention effectively solves the bottlenecks of traditional comparators in terms of speed and reliability and is suitable for high-speed, high-precision data conversion systems.
Owner:CHENGDU NACHUAN MICROELECTRONICS TECH CO LTD

Multi-agent reasoning optimization method and system based on Token calculation acceleration

PendingCN122021905AArtificial lifeInference methodsPathPingIncremental encoding
The invention provides a multi-agent reasoning optimization method and system based on Token calculation acceleration, and relates to the field of intelligent decision making, and the method comprises the steps: generating and maintaining a Token-level directed acyclic graph on line; constructing a reasoning path consistency sharing strategy, an approximate semantic sharing strategy and a forced sharing strategy to perform KV cache sharing; local change features are processed in an incremental coding mode, and context difference detection, incremental coding calculation and KV cache splicing are carried out; defining a contribution degree function for each Token so as to carry out Token contribution degree evaluation, and introducing a dynamic pruning mechanism based on the contribution degree; based on the Token-level directed acyclic graph, the parallel execution of the independent Token is realized through topology analysis and dynamic scheduling. According to the method, generation of redundant Token can be avoided, and the communication bandwidth and computing resources of the spacecraft are saved; the delay of the critical path is shortened, the decision-making efficiency of a multi-agent cooperation system can be improved, and the strict requirements of spacecraft tasks on the decision-making efficiency and the decision-making safety are met.
Owner:BEIHANG UNIV

A routing method and system based on timing criticality

The application relates to the technical field of chip design, and particularly discloses a routing method and system based on timing criticality, which abstracts physical layout into a node network, constructs a weighted cost function of total wiring length and critical path length, adopts a minimum spanning tree algorithm to iteratively expand a path and record a parasitic parameter, calculates a merging cost value based on the product of the length of a coincident path and the timing criticality of a node, generates a virtual node by merging a node pair with the maximum benefit, and cyclically optimizes until a convergence condition is met, so that the dual objectives of wiring resource optimization and timing critical path delay reduction are achieved. The method is based on a more accurate time delay model, defines a target function which comprehensively considers the overall routing cost and the routing cost of a timing critical load node, makes timing analysis more accurate, reduces the signal delay and the size of a clock cycle of a load node according to the criticality of the load node, and effectively increases the running speed of a chip.
Owner:SHAOXING XINNA TECHNOLOGY CO LTD +1

A node decision based binary adder array and method

This invention discloses a binary adder array and method based on node decision, relating to the field of integrated circuits. It aims to solve the problems of high dynamic power consumption, reliance on critical path delay, and complex asynchronous adder circuitry in traditional adders. The method includes: calculating the carry propagation signal Pi and the carry generation signal Gi in parallel using operands A and B; generating a decision node Di for each bit i based on Pi and Gi; setting Di low when Gi=1, setting Di high when Gi=0 and Pi=0, and assigning the state Di-1 to Di when Pi=1; generating the next bit sum signal Si+1 based on Di; outputting ~Pi+1 when Di is low, and outputting Pi+1 when Di is high. Signal transmission is achieved by controlling the on / off state of physical channels, rather than cascading logic gates. The device includes a logic operation module and an output control module. The output control module includes a node control network with level injection units and level transmission switches, all connected to the decision node Di. This invention also provides an asynchronous adder array composed of multiple cascaded adder units and a chip integrating this circuit. This invention achieves conditional signal propagation through a unique decision node network and physical channel control, which can significantly reduce dynamic power consumption when the input data is sparse, and naturally supports asynchronous operation, with the advantages of high energy efficiency and easy expansion.
Owner:汤斌

A digital filter structure for sigma-delta ADC down-sampling

The application discloses a kind of digital filter structures for Σ-ΔADC downsampling, adopt full digital method to construct three-order single loop feedforward system, series integrated integral-comb cascaded filter, compensation filter and half-band filter.Integral-comb cascaded filter is composed of multistage integrator series recursion structure, support downsampling multiple programmable, improve parallelism and hardware utilization by moving delay unit position in advance.Compensation filter adopts odd-order integrator multi-branch parallel, reduce resource occupation, significantly improve processing speed;Half-band filter uses retiming technology to adjust coefficient position, reduce critical path delay, improve the highest frequency of system, while reducing the amount of calculation through symmetric structure and multiplexing extraction path, front extraction reduces the demand of operation rate and data loss.The application realizes low power consumption, low hardware consumption and high design flexibility of downsampling filter in Σ-ΔADC digital module.
Owner:XI AN JIAOTONG UNIV

Cryptographic algorithm hardware implementation method based on Montgomery modular multiplication

The invention discloses a cryptographic algorithm hardware implementation method based on Montgomery modular multiplication. A hardware circuit comprises an APB interface module, a host register interface module, a control logic module, an IRQ interface module and an RSA core operation module. According to the hardware implementation method, loading of RSA parameters and result output are achieved through a host register interface. And full-process control from data loading, Montgomery domain conversion, modular exponentiation iterative operation to final result output is realized through a control logic module and an internal state machine. A high-base Montgomery algorithm is adopted in the RSA core operation module to be combined with a base 4 Booth carry-save multiplier and a CSA partial product compressor, parallel operation of a multi-stage depth assembly line is achieved, and a modular multiplier of 12-path parallel operation is adopted outside the RSA core operation module to achieve an external assembly line structure. Through cooperative processing of internal and external double-layer assembly lines, the delay of a critical path is remarkably shortened, and the throughput rate of modular exponentiation calculation and the utilization rate of hardware resources are improved.
Owner:HANGZHOU DIANZI UNIV

AVS3 advanced entropy coding hardware acceleration architecture and method based on bypass acceleration

The invention discloses an AVS3 advanced entropy coding hardware acceleration architecture and method based on bypass acceleration, the architecture comprises a register, a one-in-sixteen data selector, and a plurality of LPS units, MPS units and bypass units which form a bypass calculation path and 15 original path calculation paths, the bypass calculation path is formed by connecting the bypass units in series, and the original path calculation path is formed by connecting the bypass units in series. The original calculation path is formed by connecting LPS units or MPS units in series; the output end of the register is connected with the input end of each calculation path, the output end of each calculation path is connected to the input end of the one-out-of-sixteen data selector, and the output end of the one-out-of-sixteen data selector is connected to the input end of the register; according to the architecture and the method, the LPS processing path can be effectively shortened, the key path delay can be reduced, the hardware resource utilization rate is considered, and the architecture and the method are of great importance to high-requirement application scenes such as 8K and 120fps ultra-high-definition real-time coding.
Owner:SOUTHEAST UNIV

Circuit for reducing the number of iterations of a pipelined cordic and method of implementation thereof

ActiveCN122431637BPathPingTiming margin
The application relates to a circuit for reducing the number of CORDIC pipeline iterations and a method for implementing the same. The method combines multiple iterations in the same clock margin safety window at the initial stage of iteration by accurately evaluating the cumulative delay of each stage of iteration, realizes dynamic redivision of the physical boundary of the pipeline, breaks through the mapping limit of the single-stage pipeline, effectively reduces the total iteration number of the system, and significantly reduces the data processing delay of the system end to end. The dynamic division method of the pipeline boundary based on delay accumulation takes the critical path delay of the maximum bit width pipeline stage as a global constraint condition, dynamically evaluates and optimizes the insertion position of the register through a greedy strategy, fully transforms the early idle timing margin under the premise of not reducing the highest working frequency of the system, solves the problem of serious imbalance of the timing between the pipeline stages, and significantly reduces the data processing delay of the system end to end while further compressing the chip area and dynamic power consumption.
Owner:NAT UNIV OF DEFENSE TECH

Encoding apparatus, system, method for transmission identifier

The embodiment of the disclosure provides a kind of encoding device, system, method of transmission identifier, the embodiment of the disclosure is applied to the data transmission between the master port and slave port of AXI4 protocol, master port connects single external host, slave port connects single external slave machine, device includes: lookup table module, for storing the mapping information of M-bit transmission identifier of master port and N-bit transmission identifier of slave port, dynamically update the allocation and release of transmission identifier;Mapping module is used for according to mapping information, the mapping processing of M-bit transmission identifier and N-bit transmission identifier between master port and slave port, M and N are natural numbers.The embodiment of the disclosure solves the problems of low transmission identifier management efficiency, high critical path delay, large storage overhead and other problems existing when traditional AXI4 bus is accessed by multiple master devices concurrently, without changing the premise of AXI4 protocol standard, significantly improve the bus performance, applicable to the chip design of strict low delay, high bandwidth requirement.
Owner:ZHIHE COMPUTING TECHNOLOGY (HANGZHOU) CO LTD

A low-latency pseudo-random function implementation method supporting error detection

The application provides a low-delay pseudo-random function implementation method supporting error detection, comprising the following steps: step 1, obtaining and copying a 128-bit input message; step 2, generating a round key; step 3, performing a pseudo-random function operation to obtain a 128-bit initial output; and step 4, performing an error detection mechanism to obtain a 128-bit final output. Through the whole-process cooperative design from a nonlinear S-box to a complete pseudo-random function construction, the application realizes the minimization of the critical path delay, significantly reduces the hardware delay compared with the existing low-delay cryptography primitives, and is more suitable for the low-delay requirements of the Internet of Things platform. The application sets an error detection mechanism based on a linear block code at the algorithm level, detects the transmission errors and countermeasure fault injection induced by environmental interference with a small area cost, and fills the gap of the existing low-delay cryptography primitives in the reliability design.
Owner:XIANGTAN UNIV

A method, device and medium for merging and decomposing FPGA lookup tables

The application provides a FPGA lookup table merging and decomposition method, device, equipment and medium, and relates to the integrated circuit netlist optimization technical field.The application obtains a mapped LUT network of the FPGA;then, candidate pair identification is performed on the LUT network, a conflict-free and optimal-weight LUT candidate pair is selected for merging, and a LUT connection relationship graph is constructed after the merging, and a critical path is identified;based on the LUT connection relationship graph, LUTs on a non-critical path are decomposed into smaller cascade LUT units in the order of LUT priority;after each decomposition, the network topology structure and timing information of the LUT connection relationship graph are dynamically updated, and the optimization of the FPGA netlist is completed.The application can significantly reduce the number of LUTs and improve the FPGA resource utilization rate under the premise of ensuring that the delay of the critical path remains unchanged.
Owner:XIAMEN UNIV OF TECH

Reinforcement learning-based combinatorial logic clustering optimization method and system

The application discloses a kind of based on reinforcement learning's combination logic clustering optimization method and system, method includes: S1, the fan-out information of register and combination logic unit in circuit netlist is counted;S2, according to fan-out information, the register and combination logic unit that the fan-out quantity is greater than first preset threshold and the worst timing path timing violation value of output signal is less than second preset threshold are screened out, and target unit list is generated;S3, the fan-out combination logic information of each unit in target unit list is acquired;S4, based on fan-out combination logic information, based on reinforcement learning algorithm respectively to fan-out logic unit is clustered operation;S5, after completing logic clustering, the netlist of circuit is revised;After revision is completed, equivalence check is carried out;S6, according to the netlist of modification circuit is re-layouted and routed.The application can effectively reduce critical path delay, optimize circuit performance, especially suitable for large-scale integrated circuit design.
Owner:NAT UNIV OF DEFENSE TECH

Three-stage positive feedback comparator circuit based on rapid time sequence and electronic equipment

The invention relates to the technical field of integrated circuits, and discloses a three-stage positive feedback comparator circuit based on a fast time sequence and electronic equipment. The circuit comprises a first-stage positive feedback module, a second-stage trigger and clock module, a second-stage positive feedback amplification module and a third-stage positive feedback latch module. According to the invention, a three-stage positive feedback structure is adopted, and corresponding differential signal nodes are arranged in the same feedback loop at each stage, so that the amplification rate and the metastable state suppression capability of the comparator are obviously enhanced; and by introducing a rapid time sequence control mechanism, second-stage amplification, clock reset generation and third-stage latch operation can be executed in parallel, so that the key path delay is effectively shortened, and the overall comparison speed is greatly improved. The electronic equipment comprises the comparator circuit. The invention effectively solves the bottleneck of the traditional comparator in the aspects of speed and reliability, and is suitable for a high-speed and high-precision data conversion system.
Owner:CHENGDU NACHUAN MICROELECTRONICS TECH CO LTD

An approximate floating point adder based on operand truncation

The application provides an approximate floating point adder based on operand truncation, comprising a preprocessing module for performing size comparison of floating point data, data exchange, calculation of exponential difference value and shift and order operation; an approximate mantissa addition module for approximately adding floating point mantissas; an approximate leading 1 detection module for approximately detecting the position of the leading 1 of the mantissa addition result; a normalization module for performing sign bit judgment and mantissa normalization processing of the mantissa addition result; an exponential adjustment module for performing correction processing on the exponent of the normalized result; and a zero judgment correction circuit module for correcting zero judgment errors caused by approximation. The application effectively reduces circuit power consumption area and shortens critical path delay, compensates for the constant value 1 at the truncation of the mantissa addition output result, maximally realizes error compensation without consuming any additional resources, avoids the occurrence of large errors, and can realize a high-precision low-power approximate floating point adder.
Owner:NANJING UNIV OF AERONAUTICS & ASTRONAUTICS

Fpga layout method based on reinforcement learning

The application belongs to the field of integrated circuits, and particularly relates to a FPGA layout method based on reinforcement learning. First, according to an input netlist file, logical units contained in a FPGA design circuit are extracted, and then initialization layout operation of the logical units is completed; in view of the slow convergence problem of a traditional simulated annealing method layout, multiple search region construction methods are proposed, which can effectively improve the search efficiency of the layout solution space; on this basis, an optimal search region selection method based on reinforcement learning is proposed, which can adaptively select the optimal search region to perform the exchange operation of the logical units. The layout method can greatly reduce the time required for FPGA layout while maintaining the required wire length and critical path delay.
Owner:BEIJING MXTRONICS CORP +1

Modular multiplier hardware construction method and system based on constant decomposition and collaborative optimization

The application discloses a constant decomposition and cooperative optimization-based hardware construction method and system for a modulus multiplier. The method comprises the following steps: obtaining an original constant modulus multiplication operation, decomposing a large-bit-width constant to obtain an expression composed of small-bit-width parameters through multiplication and addition and subtraction operations, wherein each level of decomposition splits a current level value into a combination of a high-bit part and a low-bit part, and each parameter value is limited in a modulus range; based on a regular signed number coding method, converting each small-bit-width parameter obtained through decomposition into a binary representation form with the least non-zero bits to determine the least number of adders required for multiplication of each parameter; in combination with the characteristics of a modulus reduction algorithm, evaluating the resource consumption and critical path delay of different decomposition expressions during hardware construction; and through iterative optimization, selecting a decomposition expression with the least number of adders under the constraint of critical path delay to generate a hardware circuit structure of a constant modulus multiplier. The application can effectively reduce resource consumption and improve operation speed.
Owner:SUN YAT SEN UNIV

A configurable modular multiplier

The application provides a configurable modulo multiplier, comprising an n*n-bit binary multiplier, a bit extractor, an (n-1-k)*k-bit binary multiplier, a first configuration register stack, a compression tree adder, a second configuration register stack, an n+1-bit adder, a 2-bit adder, a carry-save adder, an n+2-bit adder and a modulo correction unit. The configurable modulo multiplier of the application proposes a new modulo correction scheme for high-bit partial product modulo correction, greatly reduces the configuration capacity of the configurable general-purpose modulo multiplier, reduces the critical path delay of the general-purpose modulo multiplier, and solves the variable remainder base requirement of the remainder system application compared with the prior art. By completely separating the multiplier and the modulo correction part, the hardware compatibility problem of the regular application and the remainder application is solved.
Owner:CHENGDU SANLINGJIA MICRO-ELECTRONICS CO LTD

Galois field index calculation method and device, decoder and storage medium

PendingCN122316548APathPingAlgorithm
This application relates to the field of digital communication technology and discloses a method, apparatus, decoder, and storage medium for calculating the Galois field index. The method includes: performing a group indexing operation on the target index to obtain a base address and an offset within the group; reading the corresponding base value from a pre-stored compressed index table based on the base address; wherein the index entries stored in the compressed index table are subsets selected at equal intervals with a fixed step size from the original complete index space, and the fixed step size is an integer power of 2 greater than 1; and performing a single-step recursive operation corresponding to the number of times the offset within the group is set based on the base value and the offset within the group to generate the Galois field element value corresponding to the target index. This application can reduce the storage and wiring overhead of Galois field index calculation and shorten the critical path delay.
Owner:SHENZHEN CITY TECHWIN SEMICONDUCTOR COMPANY LIMITED

Routing method and system based on time sequence criticality

The invention relates to the technical field of chip design, and particularly discloses a routing method and system based on time sequence criticality, and the method comprises the steps: abstracting a physical layout into a node network, and constructing a weighted cost function of a total wiring length and a critical path length; iteratively expanding the path by adopting a minimum spanning tree algorithm and recording parasitic parameters; and calculating a merging cost value based on the product of the coincidence path length and the node time sequence criticality, merging the maximum-income node pair to generate a virtual node, and performing loop optimization until a convergence condition is met, thereby realizing dual targets of wiring resource optimization and time sequence critical path delay reduction. According to the method, on the basis of a more accurate time delay model, a target function which comprehensively considers the total routing cost and the routing cost of a time sequence key load node is defined, so that time sequence analysis is more accurate. The overall routing cost can be reduced, meanwhile, the signal delay of the load node and the clock period are reduced according to the criticality degree, and the running speed of a chip is effectively increased.
Owner:SHAOXING XINNA TECHNOLOGY CO LTD +1

High-parallel rejection sampler for post-quantum cryptography and working method thereof

PendingCN122339692AData compressionPathPing
This invention discloses a highly parallel rejection sampler for post-quantum cryptography and its operating method. The 8-bit sampled data is split into symmetrical high-4 and low-4 bit parallel sampling units: the high-4 bit unit performs data compression towards the least significant bit, and the low-4 bit unit performs data compression towards the most significant bit. The two compressed outputs are combined through physical splicing to form a pre-aligned data stream with continuous data in the middle and possible invalid data bits at the beginning and end. Subsequently, the system uses the effective data count value of the high-bit unit as a control signal to drive a multiplexer network to directly select the output result according to five preset alignment paths, thereby achieving a compact arrangement of fully parallel data. This invention avoids the complex multi-level shift operations in traditional architectures, greatly optimizes critical path latency and hardware resource overhead, and is suitable for high-speed hardware accelerator design for post-quantum cryptography algorithms.
Owner:STATE GRID JIANGSU ELECTRIC POWER CO LTD RESEARCH INSTITUTE +1

Circular shift registers, data circular shift methods, and storage media

This application provides a circular shift register, a data circular shift method, and a storage medium. The circular shift register includes a data input port, a data output port, a control port, and a multi-level processing network. The multi-level processing network is constructed using a hybrid of basic units consisting of 2x2 switching units or 2x1 multiplexers. Each processing unit includes multiple basic units, and the multiple basic units of each processing unit are of the same type. At least two processing units in the entire network use different types of basic units. The multiple basic units of each processing unit are independently controlled by a common control signal generated by the shift control word generation stage, enabling rapid circular shift operations. This circular shift register, with its multi-level parallel processing architecture, significantly improves data processing efficiency. While maintaining ultra-low latency and high parallelism, it reduces the fan-out load of control signals and global interconnect pressure, significantly reducing critical path latency and dynamic power consumption.
Owner:合肥康芯威存储技术有限公司

Local-cloud collaborative RTL code optimization method for IP security

PendingCN121979528Asolve protection problemsResolve conflicts between optimization effectsIntelligent editorsCAD circuit designPathPingLinguistic model
The invention discloses a local-cloud collaborative RTL code optimization method for IP security. The method comprises the following steps: step 1, constructing K pairs of proprietary modules and first draft by using a local large language model; step 2, extracting an optimization principle set from a proprietary module to a first draft according to a local principle; 3, performing IP leakage risk detection verification on the optimization principle set by using a local large language model; step 4, feeding back the leakage feedback value to the local large language model, and repeating the step 2 to the step 4 until the leakage verification of the optimization principle set is passed; and 5, uploading the optimization principle set passing the leakage verification to the cloud big language model, and optimizing the target module to be optimized. According to the method, the optimization success rate of 50.85% is achieved in a key path delay optimization task; and the optimization success rate reaches 66.67% in a power consumption optimization task. According to the method, on the premise of ensuring the safety of intellectual property, the hardware description language code can be effectively optimized by utilizing the optimization capability of the cloud large model.
Owner:NANJING UNIV OF POSTS & TELECOMM +1

A memory-computing integrated large language model DRAM fault-tolerant method

The application relates to the technical field of computer architecture, and discloses a memory-computing integrated large language model DRAM fault-tolerant method, which comprises the following steps: constructing a space-time related multi-granularity DRAM fault model, defining a memory bank level space constraint, and a single-bit or multi-bit flip fault mode, and establishing a DRAM fault probability calculation method based on the Arrhenius formula; performing multi-precision and module level fault sensitivity quantitative analysis on the large language model; based on the fault sensitivity quantitative result, performing precision-aware bit-level differentiated fault injection, and using a fine-tuning framework of a low-rank self-adaptive adapter based on module heterogeneity perception to perform fault-tolerant enhancement on the target large language model; the application has small parameter redundancy, does not add a hardware error correction circuit, and does not increase the hardware critical path delay, thereby improving the tolerance and robustness of the large language model to single-bit and multi-bit flip faults under the DRAM memory-computing integrated architecture.
Owner:UNIV OF SCI & TECH OF CHINA