Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

160 results about "Precomputation" patented technology

In algorithms, precomputation is the act of performing an initial computation before run time to generate a lookup table that can be used by an algorithm to avoid repeated computation each time it is executed. Precomputation is often used in algorithms that depend on the results of expensive computations that don't depend on the input of the algorithm. A trivial example of precomputation is the use of hardcoded mathematical constants, such as π and e, rather than computing their approximations to the necessary precision at run time.

Graph neural network execution on neural processing unit

Workloads for executing a graph neural network (GNN) may be divided among various processing units, such as a central processing unit (CPU) and a neural processing unit (NPU). The NPU may include a data processing unit (DPU) and a digital signal processor (DSP). The CPU may perform precomputation, model optimization, hardware optimization, and compilation. For example, the CPU may precompute a parameter matrix and use the parameter matrix as internal parameters of a GNN. The CPU may also perform node padding, approximation computation, or transfer of DSP operations to DPU to optimize the GNN. The CPU may also perform sparsity data compute and storage, vertical fusion of DSP operations and DPU operations, or data quantization to optimize performance of the NPU. The compiled GNN may be provided to the NPU, and the DPU and DSP may perform the operations in the compiled GNN to produce a prediction of the GNN.
Owner:INTEL CORP

Low-bit-width high-energy-efficiency floating point storage and calculation integrated circuit based on partial pre-alignment architecture

The invention belongs to the technical field of storage and calculation integration, and particularly relates to a low-bit-width and high-energy-efficiency floating point storage and calculation integrated circuit based on a partial pre-alignment framework. The circuit comprises a memory array, a pre-calculation unit, an adder tree, a configurable arithmetic unit and a normalization unit, and supports mixed precision operation of FP8MACFP4 and FP8MACFP8. The method is characterized in that a partial pre-alignment strategy dominated by an activation value is adopted, the maximum index of the activation value is dynamically counted, the mantissa of the maximum index is aligned, and multiple partial pre-alignment intermediate results are pre-calculated and latched for reuse; in combination with a customized lookup table and a multiplexer, a pre-calculation result is directly selected to replace real-time multiplication and displacement; and through the reconfigurable hardware, the FP8MACFP8 high-precision operation is realized by utilizing the FP8MACFP4 unit combination. According to the method, complete online floating point multiplication and addition operation is realized, and excellent energy efficiency ratio and operation speed are obtained while high precision is kept.
Owner:FUDAN UNIVERSITY

Multi-modal KV cache retrieval method and system based on hybrid architecture

The invention discloses a multi-modal KV cache retrieval method and system based on a hybrid architecture, relates to the technical field of data processing, and aims to solve the technical problem that in the prior art, due to the adoption of a full-online computing architecture and lack of a unified cross-modal representation framework, the computing efficiency and the modal fusion effect are difficult to consider at the same time. By constructing a hybrid processing architecture of offline pre-calculation and online processing, multi-modal data is subjected to unified fragmentation coding and KV cache and semantic fingerprints of the multi-modal data are pre-calculated in an offline stage, and cache fragments are intelligently assembled and position codes are dynamically remapped based on a user query intention in an online stage; therefore, online repeated calculation of multi-modal content is fundamentally avoided, coherent alignment and efficient fusion of cross-modal semantics are achieved, the real-time response capability, the cache reuse rate and the multi-modal generation quality of the system are remarkably improved, and high-concurrency scenes can be dynamically adapted.
Owner:SHANGHAI COSUNET NETWORK TECH CO LTD

Lightweight deployment method and system of large model on edge computing device

The invention provides a lightweight deployment method and system of a large model on edge computing equipment. According to the method, the acceleration capability portrait is constructed by extracting the hardware instruction set architecture type of the target edge device and the number of parallel computing units. Based on the instruction type, the large model weight is grouped, divided and pre-calculated through a unified lookup table vectorization engine, and a pre-calculation vector matched with the target instruction set is generated; and according to the number of the parallel units and the instruction-level parallel capability, compiling the pre-calculation vector to generate an adaptive parallel table look-up instruction block, distributing execution threads with the same number as the parallel units, and eliminating data dependence conflicts. And finally, loading the instruction block to a shared memory area, configuring topological logic of the photoconductive switch matrix based on an instruction type, and dynamically switching a data transmission path in a hardware instruction period. According to the method, the efficient deployment of the large model in the edge equipment and the low-delay reasoning in the resource-constrained environment are realized.
Owner:LUSTER LIGHTWAVE CO LTD

Generative intelligent optimization method and device

The invention provides a generative intelligent optimization method and device, and relates to the technical field of generative artificial intelligence optimization computation.The method comprises the steps that a mathematical model containing decision variables, an objective function and constraint conditions is established for a target optimization problem, and scene parameters are pre-calculated through a traditional optimization algorithm to obtain a theoretical optimal solution; constructing a training data set containing scene parameters and a theoretical optimal solution; training the data set through a conditional generation type artificial intelligence model, minimizing the difference between a generation decision and a theoretical optimal solution, and establishing a mapping relation from scene parameters to an optimal decision, so that the mapping relation implicitly learns a constraint condition satisfaction mode; current scene parameters are collected in real time and input into the trained model, a near-optimal decision scheme is generated through reverse denoising or hidden variable decoding, and the near-optimal decision scheme is applied to real-time scenes such as industrial control. According to the method, a high-quality solution close to theoretical optimum is realized, the online solution time consumption is remarkably reduced, and the real-time requirement is met.
Owner:BEIHANG UNIV

Communication optimization method for topological table persistent storage in efficient parallel computing

The invention provides a communication optimization method for topological table persistent storage in efficient parallel computing, belongs to the technical field of storage communication, and aims to avoid pseudo sharing by aligning memory allocation through cache lines, expand local topological coverage through a topological entropy increment driven prefetching mechanism, and improve the reliability of the topological table persistent storage. Boundary processing is accelerated through a pre-calculation period mapping lookup table and a frequency domain transfer function vector, synchronization overhead is optimized through a concurrency control mechanism perceived by a read-write ratio, targeted cache preloading is achieved through stability and jitter degree two-dimensional evaluation, bandwidth consumption is reduced through an increment synchronization mechanism, and the stability of the cache is improved. The communication template is selected or the communication parameters are generated through adaptive conversion by matching the matching degree decision, and the technical problem that the parallel computing performance is reduced due to the fact that the communication overhead is too large in the topological table persistent storage process is solved.
Owner:青岛国实科技集团有限公司

Resource allocation method and device based on bit map precomputation

The invention discloses a resource allocation method and device based on bit map pre-calculation, and the method comprises the steps: generating a global resource state bitmap according to the CORESET configuration in a PDCCH, and carrying out the pre-calculation of each possible candidate PDCCH to form a candidate pre-allocation bit map set; extracting a corresponding pre-allocation bitmap from the candidate pre-allocation bitmap set as a candidate set bitmap; and comparing the candidate set bitmap with the global resource state bitmap, judging whether the resources are conflicted or not, and if the resources are not conflicted, determining that the resources are available and further occupying the resources. Complex function calculation at the scheduling moment is transferred to pre-calculation at the initialization stage, only simple bitmap retrieval and bit operation are needed during scheduling, the PDCCH resource allocation time consumption is greatly shortened, real-time intensive calculation is replaced with pre-stored bitmaps, the CPU load is remarkably reduced, a processor focuses on a core scheduling algorithm, resource redundancy consumption is reduced, and the scheduling efficiency is improved. And the overall capacity and the energy utilization efficiency of the communication system are effectively improved.
Owner:BEIJING BLUE TOWER OPTICAL TRANSMISSION INTELLIGENT TECHNOLOGY CO LTD

NoC low-delay data transmission method based on dynamic routing algorithm

The invention provides an NoC low-delay data transmission method based on a dynamic routing algorithm, relates to the technical field of data transmission, and aims to solve the technical problems that an existing NoC routing algorithm cannot adapt to a link dynamic load, is lack of congestion trend prejudgment, is slow in path search convergence and is insufficient in QoS differential scheduling. The method comprises the following steps: constructing a dynamic sensing module containing a double-branch LSTM time sequence prediction model, collecting states such as link bandwidth and queue length in real time, and predicting a congestion trend in 50-100ms in the future; constructing a multi-objective evaluation function taking delay as a core, and dynamically adjusting the weight according to the QoS level; solving an optimal path by adopting an improved ant colony algorithm which introduces a congestion penalty and pruning strategy; during transmission, dynamic reselection is triggered through a pre-calculation + fast matching mechanism; and iteratively optimizing the prediction model based on incremental learning. According to the method, congestion is avoided in advance, the path search efficiency and the transmission reliability are improved, the delay is remarkably reduced compared with a traditional algorithm, and the method is adaptive to delay sensitive scenes such as a high-performance processor and an AI chip.
Owner:兰奎龙

Privacy protection neural network reasoning method based on multi-point evaluation

The invention discloses a privacy protection neural network reasoning method based on multi-point evaluation. According to the method, in the reasoning process of the privacy protection neural network, a given nonlinear function is constructed into a K-segment k-degree polynomial, a safe multiplication tree and a safe remainder tree are established, and then the remainder polynomial is safely evaluated. Due to the fact that the polynomial power needing to be evaluated in the online stage is reduced, the number of self-multiplication terms is reduced, the number of communication rounds of overall polynomial evaluation is reduced, the number of calling times of a secure multiplication triple in secret sharing is reduced, and the overhead of offline pre-calculation is correspondingly reduced. In addition, the hierarchical structure of the secure multiplication tree and the remainder tree supports multi-point input parallel computing, and the parallel computing capability of a modern multi-core processor or a GPU can be fully utilized.
Owner:HARBIN INSTITUTE OF TECHNOLOGY (SHENZHEN) (INSTITUTE OF SCIENCE AND TECHNOLOGY INNOVATION HARBIN INSTITUTE OF TECHNOLOGY SHENZHEN)

System and method for adaptive ideation and innovation management using graph-based similarity computation and real-time social facilitation

A system and method for adaptive ideation and innovation management using graph-based similarity computation and real-time social facilitation are disclosed, and the system may include a plurality of client devices operated by a human user and one or more AI Agents. The system may further include a computing platform running a Precompute Similarity Engine that may be configured to embed and index idea objects, precompute similarity clusters, and generate candidate merge or purge operations together with diversity and novelty signals; and running a Social Physics Engine that may be configured to monitor human user and machine participant (AI agent) interactions, compute group metrics, and issue digital facilitation interventions. The Precompute Similarity Engine and the Social Physics Engine may operate in an interdependent feedback loop such that diversity and novelty signals guide the generation of digital facilitation interventions that are communicated via the network to the plurality of client devices and the one or more AI Agents.
Owner:MA MOSES T

Elliptic curve scalar multiplication optimization method based on multiple windows and multiple base chains

The invention discloses an elliptic curve scalar multiplication optimization method, an elliptic curve scalar multiplication optimization device and elliptic curve scalar multiplication optimization equipment based on multiple windows and multiple base chains, which can be used for a communication system core network. The method comprises the steps that parameter sets arranged according to the window width in a descending order are configured, each parameter set comprises a cardinal number set and the window width w, and a corresponding candidate odd number set is generated; sliding recoding is carried out on an input scalar k, 0 is added to even-numbered bits and the even-numbered bits are shifted rightwards by k, a congruent recoding number with the minimum absolute value is searched for odd-numbered bits according to the priority of a parameter group, if the recoding number is not found, the wNAF rule is returned, and a sparse digital sequence D is obtained; according to a fast / ct mode, generating a pre-calculation lookup table by combining a power chain with a binary combination; and scanning D from a high order to a low order, sequentially executing point doubling operation by the accumulator, executing conditional point adding operation when meeting a non-zero number, and finally obtaining a scalar multiplication result. According to the method, the non-zero bit density and the side channel risk are reduced, the operation efficiency is improved, and the method is suitable for various security scenes.
Owner:NORTHWESTERN POLYTECHNICAL UNIV

Neutral atom quantum compiling method and system for off-line pre-calculation

The invention discloses a neutral atom quantum compiling method and system for off-line pre-calculation, and belongs to the field of quantum computation.The method comprises the two stages of off-line template construction and on-line compiling, specifically, in the off-line stage, a neutral atom array is abstracted into a two-dimensional grid, an effective gate candidate set is constructed based on the Rydberg blocking radius and the parallel execution limiting radius, and an effective gate candidate set is constructed based on the effective gate candidate set; generating a space template configured by the maximum parallel gate by using a satisfiability model theory solver, and expanding a template library through rotational symmetry; in the online stage, quantum circuits are layered, matching templates are retrieved, slot distribution, direction selection and conflict evaluation are carried out, and finally an executable operation sequence is generated through timing sequence routing coloring. According to the method, the compiling efficiency, the expandability and the circuit execution fidelity are remarkably improved, and the method is suitable for a large-scale neutral atom quantum computing platform.
Owner:UNIV OF ELECTRONICS SCI & TECH OF CHINA +1

Multi-machine four-dimensional cooperative path planning method for pre-calculating deviation path and dynamically re-planning

The invention discloses a multi-machine four-dimensional cooperative path planning method for pre-calculating a deviation path and dynamically re-planning, and belongs to the field of cooperative control and path planning of unmanned aerial vehicle clusters. The method comprises the following steps: dividing a three-dimensional airspace into cubic empty blocks, constructing a directed connected graph, and removing the empty blocks and edges which do not meet a safe distance; generating an initial three-dimensional path of each unmanned aerial vehicle by using a dynamic priority fast expansion random tree algorithm; taking an end point empty block as a starting point, constructing a reverse reachable graph by means of breadth-first search, pre-calculating a plurality of deviation paths from a neighborhood empty block to an end point in an off-line manner, and storing the deviation paths into the empty block; when a dynamic threat is detected in the task, calling an improved heuristic artificial potential field algorithm to generate a local obstacle avoidance section; and retrieving a pre-stored deviation path after obstacle avoidance, and selecting an optimal path for splicing in combination with a cost function. According to the method, the calculation amount is moved forward to an offline stage, the dynamic environment re-planning time is shortened, simultaneous arrival of multiple machines and space safety are guaranteed, and the expansion capability and the real-time performance are improved.
Owner:SHENYANG AEROSPACE UNIVERSITY

Lens vignetting correction method and device, computer equipment and storage medium

The invention discloses a lens vignetting correction method and device, computer equipment and a storage medium, and the method comprises the steps: setting a reference data index according to a vignetting intensity parameter of a lens, and extracting a reference coefficient from pre-calibrated lens vignetting data; constructing a correction lookup table in combination with the vignetting intensity parameter, the reference data index and the reference coefficient; obtaining a target image collected by a lens, and calculating an optical center coordinate and a normalized radius reference value of the target image; obtaining a to-be-processed area of the target image, and performing area initialization on the to-be-processed area in combination with the optical center coordinate and the normalized radius reference value; and based on the initialized to-be-processed region, correcting the target image in combination with the correction lookup table. According to the vignetting correction method and device, the real-time complex mathematical operation is converted into the pre-calculated table query operation, efficient and accurate vignetting correction is achieved, and the problems that in the prior art, calculation complexity is high, and real-time performance is poor are solved.
Owner:AFIRSTSOFT CO LTD

A method, apparatus, device, and medium for processing multi-dimensional data

This invention relates to the field of data processing technology, and in particular to a method, apparatus, device, and medium for processing multi-dimensional data. The method instantiates a template based on input tensor parameters to obtain an offset calculation instance. The offset calculation instance is initialized to determine the step size and shape of each input array in each output dimension, and these are recorded in a designated storage space. The offset calculation instance is then started to perform index calculations on the step size and shape of each input array recorded in the storage space in each output dimension to obtain the offset corresponding to the output array. The data corresponding to the offset is then processed according to a set calculation rule to obtain the output result. By pre-calculating the step size and shape of each input array in each output dimension, unnecessary memory accesses can be reduced. Constructing offset calculation instances simplifies the index calculation process, enabling efficient and accurate data storage, retrieval, and calculation.
Owner:INSPUR SUZHOU INTELLIGENT TECH CO LTD

Accelerating Quantum Algorithms with Precomputation

The disclosure is directed to a method including executing, at a first time, a precompute algorithm that is a first portion of a quantum algorithm. Executing the precompute algorithm generates a precompute output that includes first quantum information encoded in a first set of qubits. A runtime input for the quantum algorithm is received at a second time that is subsequent to the first time. A runtime algorithm is executed, at a third time that is subsequent to the second time. The runtime algorithm is a second portion of the quantum algorithm. Executing the runtime algorithm is based on the first quantum information encoded in the first set of qubits. Executing the runtime algorithm generates a runtime output that includes second quantum information encoded in a set of qubits. An output of the quantum algorithm is provided. The output of the quantum algorithm is based on the second quantum information.
Owner:GOOGLE LLC

Idle time-based transaction data processing method and device, equipment and medium

The invention relates to the technical field of big data, and provides a transaction data processing method and device based on idle time, equipment and a medium, which can start a business thread to scan a market queue and a return queue in real time in a current transaction period so as to respond in time. In the continuous scanning process of preset times, when no new market information slice is scanned in the market information queue and no new return information is scanned in the return queue, calculating the current idle time according to an idle time determination strategy after the market information and the return are ensured to be processed; the operator type with high real-time performance is divided into the pre-calculation operator and the decision operator, the pre-calculation operator is processed in the current idle time, the decision operator is processed in the current transaction time, the idle time can be fully utilized, glitch time delay in the transaction process is reduced, and the performance and time delay penetration stability of a transaction system are improved.
Owner:SHANGHAI GUIYAN TECHNOLOGY CO LTD

Distributed optical path tracking method and system based on precomputation

The invention relates to a distributed light path tracking method and system based on precomputation, and the method comprises the following steps: S1, dividing scene data, and then carrying out distributed storage, dividing the scene data into a plurality of data blocks containing a BVH structure and geometric data according to the spatial distribution characteristics and storage overhead of a scene object, the static data are distributed to the distributed storage unit in a balanced mode to serve as initial static data of the distributed storage unit; s2, generating a pre-rendering camera according to the data access frequency of the KD tree; s3, predicting the data access frequency; s4, reordering the data items according to the data access frequency; and S5, performing distributed rendering.
Owner:SHANDONG UNIV OF FINANCE & ECONOMICS

Software and hardware collaborative design method for three-dimensional imaging internal structure reconstruction

The invention discloses a software and hardware collaborative design method for three-dimensional imaging internal structure reconstruction. The method comprises the steps that a three-nested loop general programming model is established, a hardware-friendly algorithm is designed, and a novel near memory computing architecture Waffle is designed; defining a unified programming model to cover different three-dimensional reconstruction tasks; a data stream of the algorithm is redesigned to improve data locality, and redundant calculation is reduced through pre-calculation and simplification of a synchronization scheme; a Waffle hardware accelerator with a near memory computing architecture is constructed and used for accelerating algorithm execution, and related operation is executed under customized computing logic; according to the method, the advantages of near-storage calculation are fully exerted through software and hardware collaborative design, and the speed and energy efficiency of three-dimensional imaging internal structure reconstruction can be improved.
Owner:PEKING UNIV

SIFT (Scale Invariant Feature Transform) algorithm hardware circuit implementation with low resource consumption

PendingCN121073746AProcessor architectures/configurationAlgorithmGaussian image
The invention discloses an SIFT (Scale Invariant Feature Transform) algorithm hardware circuit implementation, which optimizes the hardware circuit implementation of an original SIFT algorithm, and solves the problems of slow operation of the SIFT algorithm and large resource consumption when the SIFT algorithm is deployed on an FPGA (Field Programmable Gate Array). According to the specific implementation scheme, the method comprises the steps that a one-layer six-group Gaussian pyramid is built, and an independent precalculated 21 * 21 Gaussian filtering kernel is used in each group; extreme points of three adjacent groups of differential pyramids are detected by adopting a threshold method, and edge response points are eliminated; a CORDIC algorithm is adopted to calculate the gradient direction and amplitude of a Gaussian image where the feature points are located, and a 16-row 16-column calculation result is output for gradient histogram statistics and descriptor generation; and normalization of a 128-dimensional descriptor is realized by using a single divider. The method has the advantages of low resource consumption, high real-time performance and the like, and is suitable for scenes, such as an FPGA (Field Programmable Gate Array), an ASIC (Application Specific Integrated Circuit) and the like, needing hardware acceleration of the SIFT algorithm.
Owner:HARBIN INST OF TECH AT WEIHAI +1

Privacy-preserving neural network inference method based on multi-point evaluation

The application discloses a privacy protection neural network inference method based on multi-point evaluation. K Segment k Polynomial, and a safe multiplication tree and a safe remainder tree are established, and then a remainder polynomial is safely evaluated. Since the polynomial power to be evaluated in the online stage is reduced, the number of self-multiplication items is reduced, thereby reducing the communication round number of overall polynomial evaluation, reducing the number of calls of the safe multiplication triplets in the secret sharing, and correspondingly reducing the offline pre-computation cost. In addition, the layered structure of the safe multiplication tree and the remainder tree supports parallel computation of multi-point input, and can fully utilize the parallel computation capacity of a modern multi-core processor or a GPU.
Owner:HARBIN INSTITUTE OF TECHNOLOGY (SHENZHEN) (INSTITUTE OF SCIENCE AND TECHNOLOGY INNOVATION HARBIN INSTITUTE OF TECHNOLOGY SHENZHEN)

Zero knowledge proof-oriented multi-scalar multiplication hardware parallel acceleration method and device

The invention discloses a zero-knowledge proof-oriented multi-scalar multiplication hardware parallel acceleration method and device, and relates to the technical field of cryptographic acceleration and parallel computing. The method is based on a field programmable gate array, an improved multi-channel parallel Pippenger algorithm is utilized, and the parallel acceleration of the multi-scalar multiplication hardware is realized. The method comprises the following steps: firstly, performing scalar preordering, performing conflict-free data scheduling and designing a compact large integer multiplier to accelerate large-scale elliptic curve scalar multiplication, and selecting a window size and a precomputation scale by adopting an adaptive extensible 3D-Pippenger algorithm; in a data stream level, conflict between empty bucket reading and bucket writing is avoided through scalar slice merging sorting and pointer caching strategies. The acceleration circuit is suitable for various elliptic curves and different zero-knowledge proof protocols, and can be widely applied to scenes such as block chain private transactions, verifiable database outsourcing and verifiable machine learning.
Owner:BEIJING INST OF TECH +1

A rendering precomputation system for power grid data

The application discloses a rendering precalculation system for power grid data, and relates to the technical field of power grid data.The rendering precalculation system comprises a data acquisition module, a data integration processing module, a rendering processing module, a display module, a data selection module, a data preprocessing module, a system verification module and a display selection module.The data acquisition module is responsible for collecting a large amount of real-time and historical data from various sensors, monitoring devices and databases of a power grid.The data integration processing module is used for integrating and processing the data collected by the data acquisition module, so that the data of the data acquisition module is more convenient for the rendering processing module to use after being arranged.The application filters and pre-processes data in priority through the data selection module, calls historical rendering data through a data comparison module, and starts system verification at regular time intervals through a timing module in the system verification module, so that the system operation efficiency is improved, the waste of computing resource is reduced, and the system operation stability is enhanced.
Owner:GUIZHOU POWER GRID CO LTD

A method for optimizing shadow precomputation based on light source clustering technique

The application discloses a method for optimizing shadow precomputation based on light source clustering technology, which comprises the following steps: firstly, uploading a scene file; secondly, analyzing light source properties; thirdly, setting rendering parameters and submitting the scene file and the rendering parameters to a cloud rendering system; fourthly, clustering light sources through script processing of node machines of the cloud rendering farm, and then identifying key light sources; fifthly, allocating machine resources and precomputing allocation; sixthly, generating a map after power allocation; seventhly, synthesizing shadows; eighthly, dynamically adjusting; ninthly, performing rendering; and tenthly, generating and outputting a rendering output image. The application not only improves rendering efficiency and shadow quality, but also provides an efficient, simple and reliable offline rendering solution for users through intelligent management and optimization.
Owner:SHENZHEN RENDERBUS TECH

A mechanical arm obstacle avoidance path planning method based on an improved ant colony algorithm

PendingCN122323217ARobotic armMetric tensor
This invention proposes a robotic arm obstacle avoidance path planning method based on an improved ant colony algorithm, relating to the fields of robot path planning and intelligent optimization algorithms. This method endows the joint configuration space with a non-uniform metric structure induced by a metric tensor G(q), where G(q) is derived from the normalized joint inertia matrix W. I The algorithm consists of three parts: the singularity gradient outer product term and the obstacle spacing gradient outer product term. Local geodesic distances are approximated using the mean of the metric tensors at both ends of the node. All edge weights are pre-calculated and cached during the PRM graph construction phase. The ant colony uses the reciprocal of the geodesic distance as a heuristic function, drives non-uniform pheromone evaporation using the normalized value of the metric tensor trace increment, and uses the weighted sum of geodesic length cost and inertia-weighted velocity mutation penalty as the comprehensive path cost. The beetle whisker algorithm performs pre-search and completes non-uniform pheromone initialization under geodesic metrics. This method effectively improves path safety, continuity, and dynamic adaptability.
Owner:LUDONG UNIVERSITY

Fast number theoretic transform ntt acceleration chip

The embodiment of the present specification provides an NTT acceleration chip. The chip comprises: a rotation factor calculation module comprising a plurality of pre-calculation units and a plurality of real-time calculation units, the plurality of pre-calculation units are used for performing, in parallel, calculation of initial rows of rotation factors and inter-row stepping factors by using primitive roots; the plurality of real-time calculation units are used for performing, in parallel, repeated iteration calculation of rotation factors of other rows according to the inter-row stepping factors based on the initial rows of rotation factors obtained by the plurality of pre-calculation units; a transposition module comprising a plurality of transposition units, which are used for performing, in parallel, transposition processing on an input sequence to obtain a transposed sequence; and an NTT calculation module comprising a plurality of butterfly calculation units, which are used for performing, in parallel, butterfly operation of each order and point number based on the rotation factors obtained by the rotation factor calculation module and the transposed sequence obtained by the transposition module to obtain an output sequence. The efficiency of NTT calculation can be improved.
Owner:ALIPAY (HANGZHOU) INFORMATION TECH CO LTD

Optimising quantisation in attention-based neural networks through tightening of distributions of input tensors to offline transformations

The invention concerns a computer program and a method for optimising quantisation in an attention-based neural network (40). The dataflow architecture of the network includes quantisation stages (41Q-47Q) downstream of respective transformation stages, which involve transformations (representable through transformation matrices) and offline transformations. The method first comprises optimising a transformation stage (41, 423) for subsequent quantisation by a respective quantisation stage (41Q, 42Q), by optimising an offline transformation of this stage in an unsupervised manner, the offline transformation representable through a dot product of an offline input tensor and an offline transformation matrix, using an objective function designed to reduce a degree of tailedness of a distribution of tensor components of a tensor resulting from said dot product. Next, this stage is precomputed at least partly by performing the optimised offline transformation as an affine transformation based on said dot product, and stored to ready the network for inferences.
Owner:AXELERA AI BV

Non-local problem adaptive finite element method and system based on Cartesian grid

The invention discloses a non-local problem adaptive finite element method based on a Cartesian grid, and the method comprises the steps: calculating an element stiffness matrix, and carrying out the data storage in a linear table; calculating the number of non-zero elements in each row of the stiffness matrix according to the topology of the Cartesian grid, and pre-distributing the memory of the total stiffness matrix at one time; according to the result of memory pre-allocation, a correct numerical value is called from the linear table and added to the correct position of the stiffness matrix, and a total stiffness matrix used for solving is obtained; and assembling an algebraic system according to the total stiffness matrix, performing iterative solution in combination with a self-adaptive encryption index given based on posterior error estimation, and outputting a self-adaptive grid generated by iteration and a corresponding solution. According to the method, by pre-calculating possible numerical values of the stiffness matrix and providing efficient and reliable posterior error estimation, acceleration of the assembly process and improvement of the grid quality are achieved.
Owner:WUHAN UNIV

Automatic query and data retrieval optimization through procedural generation of data tables from query patterns

Latency, response times, and efficiency improvements for data querying are provided herein, particularly in the context of querying large database systems and data tables from disparate data sources. There are provided systems and methods for automatic query and data retrieval optimization through procedural generation of data tables from query patterns. A service provider may utilize different computing services for query processing and data retrieval for different applications and services used by internal and / or external users. Instead of querying large database systems and numerous data tables, pre-aggregated data tables may instead be used and searched by procedurally generating such tables based on precomputation rules and query patterns. Once patterns have been identified in queries, corresponding data may be aggregated from data sources in a pre-aggregated data table. Query optimization rules may then be used to have these data tables queried in place of their original sources.
Owner:PAYPAL INC

Path planning method, system and equipment based on mixed A satellites and medium

The invention provides a path planning method, system and equipment based on a mixed satellite A and a medium, and belongs to the technical field of automatic driving path planning. The method comprises the following steps: constructing a variable-resolution grid map according to environment obstacle distribution, adopting a high resolution in an obstacle area, and adopting a low resolution in an open area; performing pose validity verification on the starting point and the target point; dynamically updating map obstacle information based on real-time sensor data; pre-calculating the RS curve connectivity from all discretized vehicle kinematics reachable points to a target point by adopting a parallel calculation mode, and establishing a directly connectable node set; in the process of executing the mixed A star algorithm, when the expansion nodes are located in the directly connectable node set, the pre-calculated RS curve is directly called to complete path planning, and search is terminated. According to the method, through the variable-resolution map and the pre-calculation mechanism, the calculation complexity is remarkably reduced, and the real-time performance and efficiency of path planning are improved.
Owner:SINO TRUK JINAN POWER CO LTD