Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

61 results about "Approximate computing" patented technology

Approximate computing is a computation technique which returns a possibly inaccurate result rather than a guaranteed accurate result, and can be used for applications where an approximate result is sufficient for its purpose. One example of such situation is for a search engine where no exact answer may exist for a certain search query and hence, many answers may be acceptable. Similarly, occasional dropping of some frames in a video application can go undetected due to perceptual limitations of humans. Approximate computing is based on the observation that in many scenarios, although performing exact computation requires large amount of resources, allowing bounded approximation can provide disproportionate gains in performance and energy, while still achieving acceptable result accuracy. For example, in k-means clustering algorithm, allowing only 5% loss in classification accuracy can provide 50 times energy saving compared to the fully accurate classification.

Real-valued super-resolution direction-of-arrival estimation method for high-order acoustic field sensor array

The application relates to a real-valued super-resolution direction estimation method of a high-order acoustic field sensor array, which uses a received signal covariance matrix to approximately calculate an estimated value of noise power, avoids the iterative operation process of noise power in the prior art, improves the calculation efficiency, and reduces the floating-point operation amount. By constructing an augmented matrix, the array receiving data matrix and the array manifold matrix with a multi-dimensional structure of an element are made into Hermitian matrices, so that the unitary transformation processing of the related parameters of the high-order acoustic field sensor array is realized. On the basis of obtaining the array receiving data matrix and the array manifold matrix in the real number domain, the array receiving data matrix and the array manifold matrix are applied to a sparse approximate minimum variance method with a variable exponential factor, and the completely real-valued sparse approximate minimum variance direction estimation with a variable exponential factor is realized.
Owner:NORTHWESTERN POLYTECHNICAL UNIV

A heterogeneous platform approximate computing task optimization mapping method based on DVFS and DPM

The application discloses a heterogeneous platform approximate computing task optimization mapping method based on DVFS and DPM. Firstly, real-time tasks with correlation are modeled as an approximate computing task model, so that a task directed acyclic graph (DAG), a task correlation matrix and a six-tuple representing task characteristics can be obtained; then, a mechanism of DVFS and DPM combination is introduced based on a heterogeneous multi-core platform; a problem description of task mapping based on QoS and energy joint optimization is constructed; a variable substitution method and a Big-M reconstruction method are used to process nonlinear terms in the problem, the task mapping problem is linearized, and an optimal solution is obtained through a Gurobi solver; a task layering method and a greedy algorithm are used to design a low-complexity heuristic algorithm, and the scalability of the mapping method is improved. The method of the application adopts the mechanism of DVFS+DPM joint optimization under the premise of meeting the system real-time, energy efficiency and reliability constraints, and improves the QoS of the system.
Owner:SOUTHEAST UNIV

Layered multi-objective optimization method and system for approximate calculation parameter optimization

The invention relates to a hierarchical multi-objective optimization method and system for approximate calculation parameter optimization. The method comprises the following steps: constructing a hierarchical search model for approximate calculation commitment optimization; and in the category-type decision-making layer, taking each approximation strategy as an arm, maintaining the arms by adopting a multi-arm tiger machine algorithm, obtaining random sampling values in each round of iteration according to the obtained Beta distribution, and evaluating the target arm corresponding to the maximum value so as to obtain an arm to be optimized. And inputting the parameter point to be evaluated into a numerical decision-making layer to perform parameter iterative optimization, and generating a next parameter point to be evaluated by maximizing a qNEHVI acquisition function, so that a program, corresponding to the parameter point, of the arm to be optimized is executed through a multi-target Bayesian optimization engine, and an optimization evaluation result is obtained. And obtaining the most approximate calculation configuration according to the result and the preset evaluation budget. The method can face mixed and hierarchically dependent variable space, and can improve the tuning efficiency and reduce the cost.
Owner:NAT UNIV OF DEFENSE TECH

An aging-aware precision reconfigurable logic synthesis method

PendingCN122334124AHemt circuitsLogisim
The application discloses an aging-aware precision reconfigurable logic synthesis method and system, and belongs to the technical field of integrated circuit EDA. The application aims to solve the problems of limited optimization space, introduced error from the beginning and lack of aging modeling in the existing aging-aware synthesis. The application constructs an aging standard cell library, iteratively applies constant replacement to the original circuit to generate multiple approximate candidate circuits, and then constructs a precision reconfigurable circuit composed of reconfigurable modules, which dynamically switches between accurate mode and approximate mode. Through aging-aware timing modeling, a power function model of gate delay with running time is established, and the aging continuity during mode switching is processed. The accurate mode life and the approximate mode remaining life are predicted, and the circuit with the highest score is automatically selected according to the score function. The application can improve the circuit life by about 9.5 times with an additional area overhead of about 3.72% under the premise of meeting the error constraint, which significantly enhances the long-term reliability of digital circuits.
Owner:SHANGHAI JIAOTONG UNIV

A method for predicting the centering state of a permanent magnet coupling based on a time sequence signal

The application belongs to the technical field of permanent magnet magnetic drive, and provides a permanent magnet coupling centering state prediction method based on time sequence signals. First, two mutually orthogonal state vectors are introduced to represent the eccentricity and position angle of the inner and outer rotors of the permanent magnet coupling, and the axial section of the inner and outer rotors of the permanent magnet coupling is theoretically modeled; in combination with the approximate calculation model of the Monte Carlo algorithm, the measured results are referenced to the calculation to reevaluate the weight of the data set, and the fast prediction and evaluation of the centering state of the permanent magnet coupling are realized. The application can realize the prediction and evaluation of the centering state of different types of permanent magnet couplings, and has strong adaptability; in addition, the measured data is reevaluated by the statistical probability method, the calculation result is corrected, the accuracy of the prediction value is ensured, and important technical support is provided for the multi-parameter state evaluation and prediction of the permanent magnet coupling. In the aspect of actual engineering application, it has good actual application performance, simple operation and small calculation amount.
Owner:DALIAN UNIV OF TECH

Method and device for evaluating reliability of engineering production system

The invention relates to the technical field of reliability engineering, in particular to a reliability evaluation method and evaluation device for an engineering production system, and the method comprises the steps: obtaining a reliability block diagram of the engineering production system, determining the metadata of each function block, generating a reliability function network, marking a typical reliability structure unit, and obtaining a reliability function network; the method comprises the following steps: determining a function subnet, configuring a time domain strategy to generate a self-adaptive time grid, solving the equivalent reliability of the function subnet at a self-adaptive time point of the self-adaptive time grid to obtain a reliability sampling sequence of the engineering production system, and generating a reliability evaluation result. Therefore, the problems that in the related technology, due to the fact that semantic information corresponding to a reliability block diagram depends on manual disassembly, and a reliability evaluation result is calculated through fixed step size discrete approximation, missed judgment and misjudged are likely to occur due to complex topology, the calculation precision is poor in a high-risk time period, the calculation efficiency is low in a stable time period, and engineering decision is difficult to support are solved.
Owner:TSINGHUA UNIVERSITY

A mechanical arm obstacle avoidance path planning method based on an improved ant colony algorithm

PendingCN122323217ARobotic armMetric tensor
This invention proposes a robotic arm obstacle avoidance path planning method based on an improved ant colony algorithm, relating to the fields of robot path planning and intelligent optimization algorithms. This method endows the joint configuration space with a non-uniform metric structure induced by a metric tensor G(q), where G(q) is derived from the normalized joint inertia matrix W. I The algorithm consists of three parts: the singularity gradient outer product term and the obstacle spacing gradient outer product term. Local geodesic distances are approximated using the mean of the metric tensors at both ends of the node. All edge weights are pre-calculated and cached during the PRM graph construction phase. The ant colony uses the reciprocal of the geodesic distance as a heuristic function, drives non-uniform pheromone evaporation using the normalized value of the metric tensor trace increment, and uses the weighted sum of geodesic length cost and inertia-weighted velocity mutation penalty as the comprehensive path cost. The beetle whisker algorithm performs pre-search and completes non-uniform pheromone initialization under geodesic metrics. This method effectively improves path safety, continuity, and dynamic adaptability.
Owner:LUDONG UNIVERSITY

Approximate computing digital circuit for post-quantum cryptography applications

A digital circuit for including a scalar product between two vectors (a0, a1, . . . , ai, . . . , aN-1) and (s0, s1, . . . , si, . . . , sN-1). The digital circuit includes a multiplier, an accumulator including at least one adder and a register, as well as a control circuit of the accumulator. At a clock tick of index i, the multiplier is configured to compute the result ri of the multiplication ai×si, and the accumulator is configured to add ri with the current value of the register. Afterwards, the result of the addition is memorised in the register. The control circuit is configured to control the accumulator so as to perform the addition in an approximate manner for at least one addition amongst the N additions of the computation of the scalar product. In particular, the digital circuit is intended to be used in an electronic device implementing a cryptographic algorithm based on a “Learning With Errors” (LWE) technology.
Owner:COMMISSARIAT A LENERGIE ATOMIQUE ET AUX ENERGIES ALTERNATIVES

Improved model-free current prediction control method, device and system based on forgetting factor

The application discloses an improved model-free current prediction control method and device and system based on a forgetting factor. The method does not need to acquire system parameters in advance, only needs to collect a load current, and then approximately calculates a current increment from historical current data, and then outputs an optimal inverter switch state from a minimized current value function, and meanwhile, a forgetting factor is introduced to weaken the adverse effect of a sampling error on the system, so that the model-free current prediction control is realized. On the basis of guaranteeing the realization of the parameter-free control, the forgetting factor is introduced to weaken the adverse effect of the sampling error on the system, the current quality is improved while the response speed is faster, and the robustness of the control is improved. The application is suitable for different types of power electronic topological structures, and does not need to separately perform mathematical modeling on different types of power electronic topological structures, has a reference significance for the model-free prediction current control of the power electronic converter, and has a wide application prospect.
Owner:TONGJI UNIV

Nonlinear operator approximate calculation device and method, neural network processor and medium

The embodiment of the invention provides a nonlinear operator approximate calculation device and method, a neural network processor and a medium, and belongs to the technical field of neural networks. The device comprises: a floating point number evaluation unit receiving original floating point data of an original neural network operator, and performing compensation interval evaluation on the original floating point data according to a preset floating point value domain range to obtain compensation interval evaluation information; if the compensation interval evaluation information represents that the original floating point data is not in the preset floating point value domain range, the index splitting unit splits the original floating point data into a first floating point number and a second floating point number; the operation compensation unit performs fitting compensation on the first floating-point number to obtain first output data; and the splicing unit splices the first output data and the second floating-point number passing through the original neural network operator to obtain target output data. According to the invention, the computing resources and time of a computer system can be reduced, the full-value-domain compensation of the neural network operator is completed with few hardware resources, the computing precision is improved, and the reasonability of network reasoning is ensured.
Owner:SHENZHEN WEIXUN TECH CO LTD

A 2-group signed tensor computing circuit structure based on 6-bit approximate full adder

ActiveCN115840556BBinary multiplierNeural network hardware
The application discloses a 2-group signed tensor calculation circuit structure based on a 6-bit approximate full adder, relates to the field of neural network hardware acceleration, and comprises a 6-bit approximate full adder module, a signed 8*8 approximate multiplier circuit structure based on the 6-bit approximate full adder and the 2-group signed tensor calculation circuit structure based on the 6-bit approximate full adder. Since the neural network accelerator can sacrifice part of the accuracy of data in exchange for the optimization of delay, area and power consumption of a circuit structure, the 6-bit approximate full adder module ignores part of the carry, thereby reducing the circuit area and lowering the circuit power consumption, the signed 8*8 multiplier calculation process and the 2-group signed tensor calculation process are optimized by using the 6-bit approximate full adder module, approximate calculation is introduced at some positions, and in exchange for the improvement of area and power consumption of the circuit structure, part of the accuracy is lost.
Owner:SOUTHEAST UNIV

Fast approximate analysis method for mutual coupling-containing pattern of a large-scale ultra-wideband heterogeneous array

The present disclosure discloses a fast approximate analysis method for mutual coupling-containing pattern of a large-scale ultra-wideband heterogeneous array, comprises: dividing the entire large-scale ultra-wideband heterogeneous array into a plurality of sub-arrays according to a type of element used by a heterogeneous array, ensuring that each sub-array contains elements of the same type; classifying the elements in the sub-arrays; selecting representative elements for representing environmentally similar elements, removing rows and columns that do not have the representative elements, and retaining all key features of the heterogeneous array; performing full-wave simulation on a constructed compact representative array, extracting AEPs of all the representative elements, and storing same; and replacing AEPs of the environmentally similar elements with the AEPs of the representative elements, performing approximate computing to obtain patterns of all the sub-arrays, and superimposing the obtained results to obtain a mutual-coupling-containing pattern of the heterogeneous array.
Owner:UNIV OF ELECTRONICS SCI & TECH OF CHINA

Chaotic sequence prediction network training and use method, interference system and medium

The present application relates to a kind of chaos sequence prediction network training and use method, interference system and medium, method includes: obtaining the derivative of several chaos sequence data of pre-set chaos system;Approximate calculation of the derivative of several chaos sequence data is based on pre-set numerical difference equation, determine the trend data of chaos sequence;Obtain initial KANs, the network parameter of initial KANs is trained based on the trend data and the chaos sequence data, obtain initial prediction network;Obtain the prediction fixed point of initial prediction network, the training state of initial prediction network is evaluated based on the prediction fixed point, when the training state is training completion, the initial prediction network is output as chaos sequence prediction network.On the basis of high prediction accuracy, training speed is improved, and then the real-time, accurate prediction of chaos sequence in intelligent interference system is realized.
Owner:成都流体动力创新中心

Homomorphic encryption maximum value calculation method based on multivariate symmetric polynomial approximation

The invention discloses a homomorphic encryption maximum value calculation method based on multivariate symmetric polynomial approximation, and belongs to the field of information security. The method comprises the following steps: firstly, setting a polynomial maximum number of times for fitting a maximum value function according to precision and depth budget; secondly, mapping an input variable into a finite dimension moment variable, and training a fitting polynomial based on the original data set; through rearrangement and layer number analysis on monomials in a polynomial, identifying and reconstructing border-crossing items exceeding depth budget, and converting the border-crossing items into a tree multiplication structure conforming to depth constraint; and finally, generating a final low-depth polynomial expression by caching and multiplexing the intermediate variable. Through the moment variable dimension reduction and low-depth reconstruction technology, the problem of combinatorial explosion in a traditional method is effectively avoided, the multiplication depth and multiplication times of ciphertext calculation are remarkably reduced, efficient and low-delay maximum value approximate calculation is achieved in homomorphic encryption schemes such as CKKS, BGV and BFV, and the throughput rate and deployability of ciphertext reasoning are improved.
Owner:CHONGQING INST OF GREEN & INTELLIGENT TECH CHINESE ACAD OF SCI

High-precision random calculation method based on binary partial product and multiplier

The invention belongs to the field of heterogeneous approximate calculation, and relates to a high-precision random calculation method based on a binary partial product and a multiplier. The high-precision random calculation method based on the binary partial product directly uses the weight of a bit corresponding to a binary number to generate a random bit stream. Bit streams are combined in a uniform distribution mode, phase and sum are carried out bit by bit, and all phase and results are added to obtain the random calculation multiplier with the optimal precision. According to the method, binary weights are combined, a shifting and splicing mode is used for replacing addition, partial product addition is used for achieving random calculation, and an approximate parallel counter is adopted for optimization. The multiplier provided by the invention fully combines the advantages of random calculation and binary calculation, the hardware cost is less than half of that of a binary multiplier, and the delay is lower. When the provided multiplier is applied to neural network reasoning, the precision of model loss can be almost ignored.
Owner:BEIJING INFORMATION SCI & TECH UNIV

Method and device for enhancing NR channel estimation frequency offset measurement

The invention discloses a method and a device for enhancing NR channel estimation frequency offset measurement, which relate to the technical field of communication, convert nonlinear operation into linear operation, and comprise the following steps of: performing complex conjugate multiplication and average processing on channel estimation values of adjacent pilot symbols to obtain an average complex value; mapping the complex value phase angle to a first quadrant, and determining a Taylor expansion point in a preset equally-spaced division interval; a phase angle is approximately calculated by using a first-order Taylor expansion formula, and a final phase angle is obtained by combining quadrant judgment, so that a high-precision frequency offset value is calculated. According to the method, through theoretical innovation, the calculation complexity is reduced from a nonlinear high order to a linear low order, and the required parameter storage space is remarkably reduced; experimental data show that under the same resource condition, the method can realize about 17dB frequency offset estimation precision improvement; when similar precision is achieved, required resource consumption can be reduced to 1 / 50 or below of that of a traditional method, and operation efficiency and engineering applicability are greatly improved.
Owner:BEIJING CHANGKUN TECHNOLOGY LTD

A matrix multiplication approximate calculation method based on multi-hash voting mechanism

This application relates to a matrix multiplication approximation calculation method based on a multi-hash voting mechanism. The method includes: a source computing node in the inference system obtains the token activation value matrix of the data to be inferred, which is to be sent to the target computing node where the target expert resides; the source and target computing nodes perform hash compression and voting operations on the token activation value matrix and the expert weight matrix respectively based on multiple preset hash functions to generate corresponding sketch sets; during the All-to-All communication process at the MoE layer, the source computing node sends the sketch set of the token activation value matrix to the target computing node; the target computing node performs approximate calculations based on the rows selected from the intersection of the two sketch sets to obtain the inference result. This method improves the inference efficiency of the inference system by introducing multiple independent hash functions for collaborative sampling and voting, thereby reducing computation, memory usage, network load, and GPU idle waiting time.
Owner:GREEN IND INNOVATION RES INST OF ANHUI UNIV

Vision transformer neural network acceleration system and method based on cpu and fpga

This invention proposes a visual Transformer neural network acceleration system and method based on CPU and FPGA. The method's implementation steps are as follows: an embedding module constructs a feature matrix; a preprocessing module preprocesses the feature matrix and weight matrix; the preprocessed feature matrix and weight matrix are moved and cached; an acceleration unit constructs a self-attention matrix; the self-attention matrix is ​​cached and moved; a post-processing unit constructs a global feature matrix; and an MLP Head module obtains the classification result. The preprocessing module in the CPU reduces the additional data processing time of the FPGA acceleration unit by preprocessing the feature matrix and weight matrix, effectively improving the neural network's computation speed. Simultaneously, the normalization operation module in the acceleration unit utilizes exponential and logarithmic approximation results during approximation calculations, eliminating floating-point exponentiation and division operations, effectively reducing hardware resource consumption.
Owner:XIDIAN UNIV

An apparatus for approximating a softmax function

The application discloses a device for approximating calculation of a softmax function. The device comprises a maximum value and last input maximum value unit, a subtraction operation unit, an approximate exp solution unit, a tree-shaped solution unit, a local and accumulation unit and an approximate ln solution unit. The maximum value and last input maximum value unit is used to obtain the maximum value of input data and store the final maximum value after comparison in an input data storage unit. The subtraction operation unit is used to perform subtraction operation on the data output by the maximum value and last input maximum value unit and the input data storage unit. The approximate exp solution unit is used to obtain the result of the exponential function of any input. The tree-shaped solution unit is used to perform tree-shaped accumulation summation on the input data. The local and accumulation unit is used to accumulate the local and of multiple inputs, and finally obtain the accumulated value of multiple inputs. The approximate ln solution unit is used to obtain the result of the logarithmic function of any input, and store the result in the input data storage unit. The device can reduce the power consumption, area and delay cost of the hardware architecture while maintaining a certain precision.
Owner:NANJING UNIV

A power flow fast approximation calculation method adaptable to grid structure and parameter changes

The application discloses a power flow fast approximation calculation method which can adapt to the change of power grid structure and parameters, comprising the following steps: determining a set of topological structures of a target power grid; for each topological structure in the set, performing the following operations: randomly generating multiple groups of power grid operation parameters, performing accurate power flow calculation for the topological structure based on each group of power grid operation parameters; based on all the accurate power flow calculation results under each topological structure, constructing a sample set for each branch; solving the solutions of each undetermined parameter in the active power linear estimation model of the branch and the solutions of each undetermined parameter in the reactive power linear estimation model of the branch; constructing a branch linearization power flow calculation model; based on a given topological structure and power grid operation parameters, performing power flow calculation by using the constructed branch linearization power flow calculation model. The application can adapt to the change of topological structure and power grid parameters while ensuring the calculation accuracy and calculation efficiency.
Owner:HUNAN UNIV

Data processing device and data processing method

The data processing device (10) includes a learning unit (11) that acquires a projective transformation matrix and a singular value matrix that are generated for each of a plurality of pieces of learning data based on the plurality of pieces of learning data, the projective transformation matrix and the singular value matrix having a reduced number of dimensions compared to the corresponding learning data; a combining unit (13) that calculates a new projective transformation matrix and a singular value matrix by performing a weighted multiplication and combining process on the plurality of projective transformation matrices and the singular value matrices acquired by the learning unit (11); and a calculation unit (15) that approximately calculates a covariance matrix obtained by performing a weighted multiplication and combining process on the covariance matrices corresponding to each of the plurality of pieces of learning data using the new projective transformation matrix and the singular value matrix calculated by the combining unit (13).
Owner:MITSUBISHI ELECTRIC CORP

Safety helmet wearing detection and identification method and system based on construction site scene context information guidance and medium

The invention discloses a safety helmet wearing detection and identification method and system based on construction site scene context information guidance, and a medium, and belongs to the technical field of visual identification, and the method comprises the steps: firstly, extracting the feature information of a picture sample through a backbone network part; dynamic context information attention is utilized to guide refined feature information, and the focusing capability of the model on a safety helmet target is enhanced; multi-scale feature fusion is performed based on guidance of a multi-stage dynamic context information attention module, and feature information in a model learning process is refined. And identifying the gradient change of the attention distribution through a second-order central difference gradient approximate calculation method. And finally, detecting and identifying by a classifier and a detector respectively. According to the method, the feature information in the model learning process is refined, and the gradient information is added, so that the discrimination of fine-grained features can be adaptively improved, and the regional focusing capability is enhanced, and therefore, the detection precision and robustness of the safety helmet worn by the human body in a construction site scene can be remarkably enhanced.
Owner:SICHUAN JINHAI CHRISTIE BIG DATA TECHNOLOGY CO LTD

Methods, devices, and electronic equipment for processing feature data of neural network models

PendingCN122311484ANetwork modelFeature data
This application provides a method, apparatus, and electronic device for processing feature data of a neural network model. The method includes: acquiring feature data; performing fixed-point processing on the feature data to obtain first fixed-point data; scaling the first fixed-point data to obtain second fixed-point data; and performing approximate calculation on the second fixed-point data using a preset error function to obtain approximate result data; and integrating the approximate result data and the first fixed-point data to obtain nonlinear result data. This solution reduces hardware implementation complexity while improving the computational efficiency of model feature data, making it suitable for lightweight deployment in real-time scenarios.
Owner:GUANGZHOU ZHONO ELECTRONICS TECH CO LTD

Hierarchical adder tree structure module with fine-grained exact reconfigurable approximation computation

ActiveCN115220689BRealize approximate addition calculationSolve the problem of not being able to cope with data vector calculationsTheoretical computer scienceApproximate computing
The application relates to a hierarchical addition tree structure module with fine-grained accurate reconfigurable approximate calculation, comprising a multi-data input module, an adder tree structure generation module, a calculation and result output module; the multi-data input module receives addends and defines approximate adder calculation, and receives user required precision configuration; the adder tree structure generation module input end receives addition numbers; after initialization, the number of addition tree layers will be transmitted to the approximate adder tree structure module to complete generation of the fine-grained accurate reconfigurable approximate adder module; the calculation and result output module performs approximate addition operation on the approximate adder generated in the previous stage to complete the final approximate calculation task, and controls the precision required by each layer in the approximate calculation process. The application realizes approximate addition calculation on a data vector, higher energy efficiency and appropriate precision, and solves the problem that the existing approximate adder configuration scheme cannot cope with data vector calculation.
Owner:NANJING RES INST OF ELECTRONICS TECH

Approximate calculation of activation functions with taylor series

PendingCN121713191ANeural architecturesPhysical realisationAccumulator (computing)Activation function
An activation function unit is capable of calculating an activation function calculated by Taylor series approximation. The activation function unit may include a plurality of computing elements. Each computing element may include two multipliers and an accumulator. The first multiplier may calculate an intermediate product using an activation, such as an output activation of the DNN layer. The second multiplier may calculate terms of the Taylor series of the approximate calculation activation function based on the intermediate product from the first multiplier and coefficients of the Taylor series. The accumulator may compute a partial sum of terms as an output of an activation function. The number of terms may be determined based on a predetermined accuracy of the output of the activation function. The activation function unit may process a plurality of activations. Different activations may be input into different computing elements in different clock cycles. The activation function unit can calculate activation functions with different precisions.
Owner:INTEL CORP

A chiral origami honeycomb optimization design method

The application belongs to the field of honeycomb design, and relates to a chiral origami honeycomb optimization design method, which comprises the following steps: step one, determining design parameters of the chiral origami honeycomb and a value range of the design parameters; step two, selecting model calculation samples in the value range of the design parameters of the chiral origami honeycomb; step three, establishing an electromagnetic calculation model of the chiral origami honeycomb in electromagnetic analysis software, calculating the model calculation samples, and obtaining corresponding total effective anti-reflection bandwidth EAB; step four, establishing a characteristic calculation model of the chiral origami honeycomb in finite element analysis software, calculating the model calculation samples, and obtaining corresponding equivalent stiffness E and equivalent density Density; step five, establishing an EAB-E-Density approximate calculation model of the chiral origami honeycomb; and step six, taking the maximum total effective anti-reflection bandwidth EAB as a target, performing optimization in the value range of the design parameters of the chiral origami honeycomb under the constraints of the equivalent stiffness E and the equivalent density Density, and obtaining optimized design parameters.
Owner:CHINA AIRPLANT STRENGTH RES INST

Any-in first-out memory device

PendingUS20260186706A1Parallel computingLookup table
The disclosed technique is an any-in, first-out (“AIFO”) memory device that improves upon conventional first-in, first-out memory devices. The technique mimics a first-in, first-out queue to provide sequential storage and retrieval in scenarios where transactional ordering cannot be guaranteed. By associating an identifier and a write-pointer in a look-up table with read requests prior to transmission, responses may be received out of order and / or interleaved without delaying processing until the next response in the sequence is received. An internal validity flag may be included in the AIFO memory array to represent the validity of each line in the memory array. Once a read pointer reaches a row that is marked as invalid, the read command pauses and discontinues further reading until a write to the row that marks the row as valid. To further optimize writing, a polarity of the validity may be toggled each read cycle. The level of the AIFO queue may be approximated using a vector-based algorithm that analyzes the look-up table.
Owner:MOBILEYE VISION TECH LTD

Heterogeneous task adaptive scheduling system for edge computing nodes

PendingCN121858220AGuaranteed smooth executionImplement task-level adaptive executionProgram initiation/switchingResource allocationEdge computingParallel computing
The invention relates to the technical field of heterogeneous task scheduling, in particular to an edge computing node-oriented heterogeneous task adaptive scheduling system, which comprises the following modules: a task monitoring module, which is used for collecting the resource state and task execution information of each edge node in real time and generating corresponding task execution state data; and the task graph construction module is used for constructing a task execution graph according to the task execution state data, nodes represent calculation units, edges represent task dependency relationships, the task execution graph is dynamically adjusted according to the real-time state of each node, and the execution sequence and the dependency relationships of the tasks are reconstructed. In the invention, the task graph construction module and the mode adjustment module are used for realizing the self-adaptive reconstruction of the task execution mode, the task can be split into parallel subtasks or staged tasks under the condition of resource limitation or node load fluctuation, and an approximate model is established for an approximate calculation task for processing; therefore, the key task can be smoothly executed in a complex resource environment.
Owner:SHANXI UNIV

Large language model parameter optimization method and system based on random GMRES

The invention provides a large language model parameter optimization method and system based on random GMRES, and is applied to the technical field of large language models, and the method comprises the steps: determining at least one to-be-optimized problem in a large language model, and converting the to-be-optimized problem into a least square problem; inputting the least square problem into the constructed randomized iteration solver for processing, and outputting a solution vector; the randomization iteration solver comprises a random AB-GMRES solver and a random BA-GMRES solver; the random iteration solver comprises a random AB-GMRES solver and a random A random AB-GMRES solver and a random BA-GMRES solver project a high-dimensional orthogonalization process in a Krylov subspace into a low-dimensional space constructed by a random subspace embedding matrix by introducing the random subspace embedding matrix to carry out approximate calculation; parameters in the large language model are updated or generation of a next reasoning token is propelled based on the solution vector. According to the invention, on the premise of ensuring the model performance, the training speed is greatly improved and the reasoning delay is reduced.
Owner:TONGJI UNIV

Key value cache dynamic compression method based on multi-dimensional semantic entropy and task adaptive perception

The invention discloses a key value cache dynamic compression method based on multi-dimensional semantic entropy and task adaptive perception, which comprises the following steps: receiving an input text sequence, extracting task features and predicting task types, and loading a corresponding cache strategy configuration template according to a prediction result; in the autoregressive reasoning process, the normalized semantic entropy of the attention head of each network layer of the Transform is calculated in real time and is used for evaluating the global aggregation degree of information of each layer; according to the semantic entropy value, dynamically allocating a key value cache retention budget of each network layer; key value pairs exceeding the budget are subjected to grading processing according to semantic importance and are respectively stored in a high-precision storage area, a semantic compression area or a low-bit storage area; when new lexical elements are generated, data reconstruction or approximate calculation is triggered for semantic nodes with relevancy exceeding a threshold value according to an interaction result of a query vector and a compressed cache; the invention ensures that key information is not lost.
Owner:NANJING UNIV OF INFORMATION SCI & TECH