Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

619results about "Computation using non-contact making devices" patented technology

Calculation acceleration method based on cooperation of fast Fourier transform and neural network reasoning

The invention belongs to the field of edge computing acceleration, and relates to a fast Fourier transform and neural network reasoning collaborative computing acceleration method, which comprises the following steps of: deploying a computing method which is based on butterfly computing merging and a tensor mapping strategy and is mixed with DFT (Discrete Fourier Transform) and FFT (Fast Fourier Transform) in an operator deployment level; on the interface integration level, a bus interface of a register access path in a tensor accelerator control path is subjected to lightweight reconstruction, so that the bus interface is adaptive to an edge computing platform; in an operation level, a user-defined instruction is introduced into a tensor calculation unit, and an accelerator is enabled to independently complete a whole-process task of FFT signal processing and NN intelligent identification. According to the method, an operation instruction set is expanded on hardware, and delay overhead caused by returning a large amount of intermediate data to a processor is avoided, so that efficient execution of FFT and neural network tasks on unified hardware is guaranteed, and multiple requirements of high real-time performance, high precision and low power consumption are met.
Owner:ZHEJIANG UNIV

Fusion multiply-add operation circuit compatible with multiple floating-point number formats

The invention discloses a fusion multiply-add operation circuit compatible with multiple floating-point number formats, and the circuit comprises an input device which is used for obtaining three floating-point numbers and control information in a fusion multiply-add operation instruction; the input decoding device is used for analyzing the control information and extracting the symbol, the index and the mantissa of each floating-point number; the multiplication correlation device is used for calculating an initial index, a mantissa product and an initial symbol; the addend aligning and shifting device is used for completing addend mantissa aligning and shifting; the additive operation correlation device is used for calculating an addition result, a symbol, a normalized shift bit number and a direction flag bit according to the mantissa product sum; the original result processing device is used for generating an initial rounding bit, a normalized shift pasting bit and intermediate result data; and the final result output device is used for outputting a final operation result compatible with the IEEE-754 or Point floating-point number according to the rounding information. The circuit area and power consumption are remarkably reduced through hardware integration, and meanwhile the calculation efficiency is improved.
Owner:SHANGHAI HIGH-PERFORMANCE INTEGRATED CIRCUIT DESIGN CENT

Method and system for generative design based on deep learning and topology optimization

A generative machine learning model, such as a convolutional neural network (CNN), can be trained with solutions from a topology optimization solver for a solution for a topology of a set of structures so that the generative machine learning model can generate a plurality of alternative designs for a structure that are alternative topology optimizations (for the structure) for a set of initial setup parameters. The generative model when being trained includes a generative network and a discriminator network. The generative model can be trained using outputs from a CNN autoencoder for densities and a CNN autoencoder for strain energies.
Owner:ANSYS INC

Internal memory, chip and related electronic equipment

The embodiment of the invention discloses an internal memory, a chip and related electronic equipment. A target computing unit of the internal memory comprises a plurality of groups of primary multipliers, a plurality of adder trees and a plurality of secondary multipliers, one group of first-level multipliers is connected with one second-level multiplier through one adder tree; the group of first-level multipliers comprises a plurality of pairs of first multipliers, and each adder tree comprises a first adder, a second adder and a third adder; each pair of first multipliers respectively receive data to perform multiplication operation, and respectively input output results into the first adder to perform addition operation; each second adder receives the output results of the plurality of first adders or the second adder at the upper level to carry out addition operation, and respectively inputs the output results to the second adder at the lower level or the third adder; and each third adder receives the output results of the plurality of second adders to carry out addition operation, and inputs the output results to the secondary multiplier to carry out multiplication operation. By implementing the embodiment of the invention, the memory computing capability can be improved.
Owner:HUAWEI TECH CO LTD

Hardware accelerator with matrix block streaming

A hardware accelerator including tiles arranged in a systolic array. At each of the tiles, the systolic array receives a first input block that includes first input matrix elements of a first input matrix. In each of a plurality of multiplication iterations, at each of the tiles, the systolic array receives a respective second input block. The systolic array computes tile products of the first input matrix elements and second input matrix elements included in the second input blocks. The systolic array adds the tile products to column-wise partial sums and transmits the column-wise partial sums to subsequent tiles along accumulator rings included in array columns of the systolic array. In a subset of the multiplication iterations, the systolic array outputs product block rows of a product matrix. The product block rows each include product matrix blocks computed as rows of the column-wise partial sums.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Configurable module-lattice post-quantum cryptography processor for key-encapsulation mechanism

Disclosed is a reconfigurable module-lattice-based key encapsulation mechanism (ML-KEM) post-quantum cryptography system and method using memory-based numbers theoretic transform (NTT). A post-quantum cryptography method of a post-quantum cryptography system including a plurality of internal submodules includes reconfiguring the plurality of internal submodules by variably selecting one security level from among the plurality of security levels; reconfiguring execution of the plurality of internal submodules to be changed through a main controller; and variably processing data according to the selected security level to perform key generation, encapsulation, and decapsulation through the reconfigured plurality of internal submodules.
Owner:INHA UNIV RES & BUSINESS FOUNDATION

Validating autonomous vehicle simulation scenarios

Validating a simulation scenario for use in training a machine learning model for an autonomous vehicle includes determining a simulation scenario; executing a simulation based on a simulation scenario; monitoring execution of the simulation and receiving messages from the execution of the simulation; determining, by a first simulation monitor, whether the messages satisfy a first incident; and validating the simulation scenario responsive to the messages satisfying the first incident to produce validated simulation data. A system and method may also include determining, by a first simulation validator, whether the messages satisfy a condition; and validating the simulation scenario responsive to the messages satisfying the condition to produce validated simulation data.
Owner:AURORA OPERATIONS INC

Polynomial multiplication accelerator

The invention provides a polynomial multiplication accelerator, and relates to the technical field of multiplication accelerators, and the polynomial multiplication accelerator comprises a calculation module which comprises a number theory transformation unit, an iterative convolution unit and an inverse number theory transformation unit; the number-theory transformation unit is configured to execute number-theory transformation operation on the to-be-calculated polynomial in the global ring by using the first butterfly subunit of the execution layer number to obtain a first intermediate result; the iterative convolution unit is configured to execute iterative convolution operation on the first intermediate result in the quotient ring to obtain a second intermediate result; and the inverse number theory transformation unit is configured to execute inverse number theory transformation operation on the second intermediate result in the global ring by utilizing the second butterfly subunit of the execution layer number to obtain a target result, so as to solve the problem that the current polynomial multiplication accelerator can forcibly and completely execute butterfly operations of all layers, and the operation efficiency is high. And therefore, a part of computing resources are used for processing low significant bit operations with relatively low contribution degrees to results, and significant computing power redundancy is caused.
Owner:NANJING UNIV

Synthetic generation of simulation scenarios and probability-based simulation evaluation

Techniques are discussed herein for generating and evaluating driving simulations based on synthetic scenarios. Simulated objects may be controlled based on parameters determining the attributes and behaviors of the objects, and scenarios may be synthetically modified by changing the parameters for a simulated object. For a driving scenario with a synthetic simulated object, a simulation system may analyze driving log data to determine a probability or prevalence associated with the driving scenario. In various examples, the simulation system may determine marginal density estimates for individual attributes of the simulated object, as well as a cumulative distribution function modeling the dependence between the attributes. The joint probability distribution determined for the synthetic simulated object can be used for evaluating the efficacy of the simulation and the performance of the simulated vehicle controllers.
Owner:ZOOX INC

Hardware accelerator with scale factor applied at tensor processor

A hardware accelerator including input memory that receives first and second input matrices. The hardware accelerator further includes processing circuitry including one or more tiles that each include a respective tensor processor configured to receive a first and second input block of the first and second input matrices. Each tile receives a first block scale factor associated with rows of the first input block and a second block scale factor associated with columns of the second input block. Each tile multiplies the first input block by the second input block, applies the first block scale factor to rows of the result block, and applies the second block scale factor to columns of the result block to obtain a scaled result block. The processing circuitry further includes an accumulator that accumulates scaled result blocks to obtain a scaled result matrix, and output memory that receives and output the scaled result matrix.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Multi-precision floating point fusion multiply-add structure and microprocessor architecture

The invention provides a multi-precision floating point fusion multiply-add structure and a microprocessor architecture, and the structure is characterized in that in a first-stage assembly line, an input preprocessing module decomposes an input operand to obtain a sign bit, an index bit and a mantissa bit; the Booth encoder generates partial products through a radix-4 Booth encoding algorithm, and the tree array multiplier compresses the partial products into three groups; in the second-stage assembly line, the CSA4-2 compression adder compresses the mantissa bits after the three groups of partial products and shift alignment into two groups; the summator module sums the two groups of partial products to obtain a mantissa summation result; in the third-stage assembly line package, a leading zero detection module performs leading zero detection on the mantissa summation result, and performs index adjustment and mantissa adjustment to obtain a normalized result; and in the fourth-stage assembly line, rounding of floating point data is carried out. According to the application, the operation performance, the energy efficiency ratio and the adaptability of the floating point fusion multiply-add structure are improved, so that the high performance and the high energy efficiency of the microprocessor are balanced.
Owner:EHIWAY MICROELECTRONIC SCI & TECH (SUZHOU) CO LTD

Systems and methods for generating interior designs using artificial intelligence

Systems, methods and / or interfaces are provided for generating interior design options. In some implementations, the method includes obtaining an inspiration image of a portion of a room, a furniture, fixture and equipment, or a combination. The method also includes parsing the image, including segmenting the inspiration image into sub-images corresponding to each furniture, fixture or equipment. The method also include identifying alternatives based on the sub-images and a layout of the room, including coordinating the alternatives with respect to each other and coordinating the alternatives with respect to the room. The method also includes generating and / or displaying variations of the room with images corresponding to the alternatives. In some implementations, the method includes parsing an image to generate an empty room sketch or schematic, generating alternatives to be placed in the empty room according to the room, and generating and / or displaying a visualization by placing the alternatives in the room.
Owner:INHABITR INC

High-speed fixed-point multiplication circuit

The invention discloses a high-speed fixed-point multiplication circuit. The multiplication circuit is mainly composed of a multiplier coding module, a partial product generation module, a partial product compression module and a traveling wave carry adder module. The multiplier coding module is composed of a radix-4-Booth coding algorithm and an opposite number generation module, and is used for carrying out three-bit block coding on input multiplication data and generating an opposite number of a multiplicand in advance. The partial product generation module generates a plurality of groups of partial products with symbol extension according to the multiplicand coded signal. And the partial product compression module adopts an improved Wallace compression structure to perform layered compression on the partial product. And the traveling wave carry adder module sums the two groups of results output by compression and outputs a multiplication result. According to the invention, multiplication accumulation series can be reduced, the switching times of invalid signals in the circuit can be reduced, and the longest delay path in an operation link can be shortened, so that fixed-point multiplication which is high in speed, low in power consumption and more favorable for a comprehensive tool in structure is realized.
Owner:SHENYANG UNIVERSITY OF TECHNOLOGY

MAC operational circuit based on double-bit-line self-circulation switching, CIM circuit and chip

The invention relates to the field of integrated circuits, in particular to an MAC operational circuit, a CIM circuit and a chip based on double-bit-line self-circulation switching. The circuit comprises a calculation array, a transmission array, a quantization unit, a signal switching unit and an output module. The calculation array is composed of two twinborn calculation units connected to each SRAM unit, and the two twinborn calculation units have the same multiplication function; the signal switching unit is used for outputting a calculation control signal so as to alternately discharge the calculation bit lines connected with the two twin calculation units, and controlling the two calculation bit lines in the transmission array to be alternately connected to the ADC; the ADC of the quantization unit combines the total discharge capacity of the two calculation bit lines to obtain the value of the low-order MAC result, and the output unit obtains the value of the low-order MAC result according to the number of times of overturning of the output signal of the signal switching unit. The problems that an existing circuit is insufficient in precision and large in area and power consumption are solved.
Owner:ANHUI UNIV

Efficient data compression in processing systems

Certain aspects of the present disclosure provide techniques and apparatus for efficiently performing operations using data compression. An example method generally includes identifying, for a block of data samples, a number of leading bits to remove from each data sample in the block of data samples. A block of compressed data samples is generated based on the identified number of leading bits and truncation of a number of least significant bits from each data sample in the block of data samples. A bitstream including the block of compressed data samples and an indication of a type of compression applied to the block of data samples is generated and output for further processing.
Owner:QUALCOMM INC

Method for automatically generating design solutions for an engineering design project

One variation of a method includes: accessing a descriptor of a project; extracting a set of language signals from the descriptor; accessing a set of input parameters and a set of output characteristics; querying a language model for a range of values of input parameters exhibited within historical engineering solutions correlated with the set of language signals; accessing a virtual model for the project and a function representing a relationship between the set of input parameters and the set of output characteristics; for each analysis instance in a count of analysis instances, defining a combination of values of input parameters within corresponding ranges and based on the virtual model, the function, and the combination of values, executing the analysis instance to calculate a set of values of output characteristics; and rendering representations of combinations of values of input parameters and sets of values of output characteristics within a user interface.
Owner:GENERATIVE VISION LTD

System and method for generating a landscape design

A method for generating an online landscape design for a property includes first providing a computing system comprising at least a memory storing computer-executable instructions of a landscaping application, and a processor coupled to the memory. The landscaping application comprises a calculator engine, a landscape design engine, and a scoring engine. Next, calculating a landscape score of a property by retrieving landscape data from online databases and comparing an existing landscape design of the property to the retrieved landscape data via the scoring engine. Next, determining property landscape improvements by the scoring engine, and then generating an improved landscape design for the property based on the property landscape improvements of the scoring engine using the landscape design engine. Next, calculating an improved landscape score of the improved design via the scoring engine and displaying a 3D-image of the property with the improved landscape design and the improved score.
Owner:HOME OUTSIDE INC

RELU neuron chip circuit

A RELU neuron chip circuit belongs to the field of chip circuits, and is characterized by comprising a vector summation circuit, a shift circuit, a subtraction circuit, a logic judgment circuit, an input port and an output port, the input ports comprise a data input port, a weight input port, a bias data input port and a data counting port; the output port comprises a data output port and a logic output port; by directly realizing the function of the neurons on the chip circuit, when a neural network system calls a certain neuron to carry out corresponding function calculation, data input and output can be directly carried out at the bottommost circuit level, so that a large amount of data cross-layer conversion time is saved. Only after the whole neural network completes one time of complete training or judgment operation, the operation result of the bottom layer circuit can be transmitted to the high layer from the bottom layer circuit in a cross-layer mode and displayed in a human eye recognizable mode. Therefore, the overall operation performance of the neural network system is improved.
Owner:XIAN UNVERSITY OF ARTS & SCI

Large-scale logic optimization multiplier verification method

According to the technical scheme, the large-scale logic optimization multiplier verification method is characterized in that a four-stage pipeline processing mode is adopted, the first stage is a partial adder tree recovery stage, and the second stage is a partial adder tree recovery stage; the second stage is an architecture generation stage; the third stage is a function verification stage; and the fourth stage is an equivalence checking stage. The verification capability is improved in a breakthrough manner and is far better than that of an existing method; the innovative three-stage decomposition strategy significantly reduces the calculation complexity, the ILP algorithm optimizes resource allocation, the multi-stage scheduling balances the calculation load, and the QACO algorithm efficiently explores the design space, so that the RefSCAT-2.0 framework provided by the invention becomes the most effective solution for verifying a large-scale logic optimization multiplier at present. From the perspective of practice, the technical blank of large-scale logic optimization multiplier verification is filled up, and important tool support is provided for reliability guarantee of key computing systems such as artificial intelligence chips, high-performance CPUs and GPUs.
Owner:SHANGHAI TECH UNIV

FPGA (Field Programmable Gate Array) coprocessor for accelerated execution of intelligent contract

The invention discloses an FPGA (Field Programmable Gate Array) coprocessor for accelerated execution of an intelligent contract, which belongs to the field of hardware acceleration and comprises a controller module, an operation module, a memory module and a stack module, the controller module is used for receiving a byte code instruction of an external smart contract and generating a control signal to control the operation module to carry out corresponding operation and / or the stack module to carry out corresponding operation; the memory module is used for receiving and storing to-be-operated data output by the smart contract and outputting corresponding operation data to the stack module, and then the stack module inputs the data to the operation module; the operation module comprises a calculator module, a multiplier module, a divider module and a hash operation module; and the stack module is used for storing the data obtained by the calculation of the operation module, and can also execute the operations of in-stack, out-stack, copying and exchange according to the control of the control signal, and outputting the data stored in the stack module to the memory module. The method is suitable for block chains and other high-performance computing applications.
Owner:ZHEJIANG UNIV +2

Cryptographic System Pipelined Number Theoretic Transform Accelerator

A cryptographic accelerator utilizes a combination of parallel and pipelined butterfly operator circuit to perform number theoretic transform (NTT) or inverse NTT (INTT.) The accelerator includes a first set of pipelined pairs of parallel butterfly operator circuits configured to operate on pairs of polynomial coefficients to provide output coefficients. A first buffer is coupled to store the output coefficients. A second set of pipelined pairs of parallel butterfly operator circuits are configured to operation on pairs of coefficients obtained from the first buffer to provide coefficients of the polynomial in a number theoretic transform (NTT) domain or out of the NTT domain.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Approximate multiplier, operating method thereof, processor and chip

The invention discloses an approximate multiplier, an operation method thereof, a processor and a chip, and belongs to the field of integrated circuits. The approximate multiplier comprises a high-order processing circuit, a low-order processing circuit, a high-low-order fusion circuit and an error compensation circuit, and the high-order processing circuit adopts a plurality of negative deviation approximate addition units, performs approximate addition on a high-order area of a partial product step by step and outputs an error signal; the low-order processing circuit adopts a positive deviation addition unit to process the low-order area to generate an intermediate result of positive deviation; the fusion circuit merges the high and low position results and outputs an initial multiplication result; the error compensation circuit compensates the result high order according to the error signal to offset the deviation. According to the approximate multiplier provided by the invention, the comprehensive performance of the multiplication unit in the aspects of area, power consumption and speed can be improved while the basic calculation precision is ensured.
Owner:SHANGHAI XINCHE WUXIAN SEMICONDUCTOR TECHNOLOGY CO LTD

Low-cost masking for post-quantum cryptography

Devices, systems, and methods for secure modular addition and subtraction are provided. A modular adder and subtractor circuit with masking circuit includes an arithmetic to Boolean (A2B) conversion operator configured to convert (i) a second sum and (ii) a value determined based on a first sum, to Boolean resulting in first and second Boolean values, a shifter configured to (i) make a most significant bit of the first Boolean value a least significant bit resulting in a shifted first Boolean value and (ii) make the most significant bit of the second Boolean value a least significant bit resulting in a shifted second Boolean value, and a Boolean to arithmetic (B2A) conversion operator, configured to convert a representation of the shifted first Boolean value and a representation of the shifted second Boolean value to arithmetic representation resulting in first and second arithmetic values, respectively.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Predicting interior models of structures

Methods and systems for improved prediction and generation of structure interiors are provided. In one embodiment a method is provided that includes receiving exterior imagery of the structure and determining an exterior surface of the structure with a machine learning model. The exterior surface may enclose exterior portions of the structure. The machine learning model may further determine exterior features of the structure and may determine, based on the exterior surface of the exterior features, an interior model of the structure. A three-dimensional representation of interior and exterior portions of structure may be generated based on the exterior surface and the interior model.
Owner:UNEARTHED LAND TECHNOLOGIES LLC

Probability symbol generation method, probability symbol operation method and equipment

The invention discloses a probability symbol generation method, a probability symbol operation method and equipment, relates to the field of digital signal processing, and solves the problems of high hardware complexity, low nonlinear operation efficiency, insufficient probability calculation precision and large conversion loss of traditional fixed-point calculation. According to the technical scheme, by fusing the advantages of fixed-point and probability calculation, a numerical value is divided into a high-bit fixed-point integer and a low-bit probabilistic part, a multi-bit probability symbol which meets the requirement for expectation unbiasedness and is controllable in variance is constructed, a class fixed-point, class probability and iterative calculation triple system is designed, and the probability of the multi-bit probability symbol is calculated. The multiplexing standard arithmetic unit processes the probability symbol stream to guarantee the linear calculation precision; single-stage multiply-accumulate operation and nonlinear interpolation are realized through a multiplexer, and discrete value selection and weighting are driven by probability; iterative resources are optimized in combination with dynamic filtering conversion. According to the method, a new normal form is provided for a high-energy-efficiency communication chip and a storage and calculation integrated framework, and digital signal processing is promoted to evolve towards the high-precision and low-power-consumption direction.
Owner:UNIV OF ELECTRONICS SCI & TECH OF CHINA

High-speed high-precision operational circuit of arcsine and arccosine functions and electronic equipment

The invention relates to the technical field of digital circuits, in particular to a high-speed and high-precision operational circuit for arcsine and arccosine functions, which comprises a preprocessing circuit for inputting a variable value of an initial function and outputting an input end connected with a first-order sub-function circuit and a second-order sub-function circuit. The output end of the first-order sub-function circuit and the output end of the second-order sub-function circuit are connected with the two input ends of the connecting circuit respectively, the output end of the connecting circuit is connected with the post-processing circuit, and an operation result is output. The preprocessing circuit processes the initial function to obtain a target original function with a definition domain and a value domain normalized to a target interval. The first-order sub-function circuit and the second-order sub-function circuit respectively process a target original function to obtain a first result and a second result; the connection circuit multiplies the two to obtain a third result; and the post-processing circuit multiplies the third result to obtain a final result. Therefore, the problems of complicated operation, high delay, high storage occupation, interpolation error and the like in related technologies are solved.
Owner:WUHAN UNIV +1

In-memory processing device and in-memory processing package with asymmetric internal and external bandwidths

The invention relates to an in-memory processing device and an in-memory processing package with asymmetric internal and external bandwidths. In one embodiment, an in-memory processing (PIM) device includes a processing unit die configured to communicate with an external device over an external bandwidth, and a storage die configured to communicate with the processing unit die over an internal bandwidth. The external bandwidth and the internal bandwidth are asymmetrically configured to operate at different data transfer rates.
Owner:SK HYNIX INC

Incorporating a ternary matrix into a neural network

Artificial neural networks (ANNs) are computing systems inspired by the human brain by learning to perform tasks by considering examples. These ANNs are typically created by connecting several layers of artificial neurons using connections, where each artificial neuron is connected to every other artificial neuron either directly or indirectly to create fully connected layers within the ANN. By substituting ternary matrices for one or more fully connected layers within the ANN, a complexity and resource usage of the ANN may be reduced, while improving the performance of the ANN.
Owner:NVIDIA CORP