Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

889results about "Computation using non-contact making devices" patented technology

Electro-photonic network for machine learning

Various embodiments provide for electro-photonic networks, including a plurality of processing elements connected by bidirectional photonic channels, suited for implementing neural-network models. Weights of the model may be preloaded into memory of the processing elements based on assignments of neural nodes to processing elements implementing them, and routers of the processing elements can be configured to stream activations between the processing elements based on a predetermined flow of activations in the model.
Owner:SICILY MERGER SUB II INC

Dynamic data type adjustment during neural network training

One embodiment provides a graphics processor comprising a base die including a plurality of chiplet sockets and a plurality of chiplets coupled with the plurality of chiplet sockets. At least one of the plurality of chiplets including first circuitry configured to perform a multi-dimensional matrix multiply accumulate operation to facilitate training for a neural network and second circuitry configured to perform dynamic datatype adjustment during the training for the neural network. The dynamic datatype adjustment is performed based on statistics generated based on output of the multi-dimensional matrix multiply accumulate operation, neural network model metadata, and training metadata associated with training operations to be performed for the neural network.
Owner:INTEL CORP

Calculation acceleration method based on cooperation of fast Fourier transform and neural network reasoning

The invention belongs to the field of edge computing acceleration, and relates to a fast Fourier transform and neural network reasoning collaborative computing acceleration method, which comprises the following steps of: deploying a computing method which is based on butterfly computing merging and a tensor mapping strategy and is mixed with DFT (Discrete Fourier Transform) and FFT (Fast Fourier Transform) in an operator deployment level; on the interface integration level, a bus interface of a register access path in a tensor accelerator control path is subjected to lightweight reconstruction, so that the bus interface is adaptive to an edge computing platform; in an operation level, a user-defined instruction is introduced into a tensor calculation unit, and an accelerator is enabled to independently complete a whole-process task of FFT signal processing and NN intelligent identification. According to the method, an operation instruction set is expanded on hardware, and delay overhead caused by returning a large amount of intermediate data to a processor is avoided, so that efficient execution of FFT and neural network tasks on unified hardware is guaranteed, and multiple requirements of high real-time performance, high precision and low power consumption are met.
Owner:ZHEJIANG UNIV

Hardware compression of sparse matrix content

The disclosure describes hardware compression of sparse matrix content. One embodiment provides a graphics processor, the graphics processor comprising: a base die, the base die comprising a plurality of chiplet slots; and a plurality of chiplets, the plurality of chiplets being coupled with the plurality of chiplet slots. At least one chiplet of the plurality of chiplets comprises: a graphical core cluster comprising a plurality of processing elements; a shared local memory coupled with the plurality of processing elements; a plurality of matrix engines coupled to the shared local memory; and codec circuitry coupled with the shared local memory and the plurality of matrix engines. Codec circuitry is configured to decode matrix data stored in a first format in a shared local memory into a second format for consumption by a plurality of matrix engines.
Owner:INTEL CORP

Fusion multiply-add operation circuit compatible with multiple floating-point number formats

The invention discloses a fusion multiply-add operation circuit compatible with multiple floating-point number formats, and the circuit comprises an input device which is used for obtaining three floating-point numbers and control information in a fusion multiply-add operation instruction; the input decoding device is used for analyzing the control information and extracting the symbol, the index and the mantissa of each floating-point number; the multiplication correlation device is used for calculating an initial index, a mantissa product and an initial symbol; the addend aligning and shifting device is used for completing addend mantissa aligning and shifting; the additive operation correlation device is used for calculating an addition result, a symbol, a normalized shift bit number and a direction flag bit according to the mantissa product sum; the original result processing device is used for generating an initial rounding bit, a normalized shift pasting bit and intermediate result data; and the final result output device is used for outputting a final operation result compatible with the IEEE-754 or Point floating-point number according to the rounding information. The circuit area and power consumption are remarkably reduced through hardware integration, and meanwhile the calculation efficiency is improved.
Owner:SHANGHAI HIGH-PERFORMANCE INTEGRATED CIRCUIT DESIGN CENT

Method and system for generative design based on deep learning and topology optimization

A generative machine learning model, such as a convolutional neural network (CNN), can be trained with solutions from a topology optimization solver for a solution for a topology of a set of structures so that the generative machine learning model can generate a plurality of alternative designs for a structure that are alternative topology optimizations (for the structure) for a set of initial setup parameters. The generative model when being trained includes a generative network and a discriminator network. The generative model can be trained using outputs from a CNN autoencoder for densities and a CNN autoencoder for strain energies.
Owner:ANSYS INC

Multiplication hardware block with adaptive fidelity control system

Methods and systems relating to computational hardware are disclosed herein. One disclosed method for executing a multiplication computation using a computational hardware block includes storing a first operand and a second operand for the multiplication computation. The first operand includes a first set of bit strings. The second operand includes a second set of bit strings. The method also includes multiplying the first set of bit strings and the second set of bit strings in a set of temporal phases using the computational hardware block. Each temporal phase uses a different group of bit strings from the first set of bit strings and the second set of bit strings. A cardinality of the set of temporal phases is determined by a fidelity control value. The fidelity control value adaptively sets a fidelity of execution of the multiplication computation.
Owner:TENSTORRENT AI ULC

Complex operation lightweight architecture and method based on FPGA (Field Programmable Gate Array)

The invention provides a complex operation lightweight architecture and method based on an FPGA, the architecture comprises a main scheduling module, a calculation module and a floating point operator module, the calculation module comprises a storage unit and a calculation unit, the storage unit obtains and caches input data of a target algorithm, and the main scheduling module carries out the calculation of the target algorithm according to the operation sequence of the target algorithm. The target algorithm is divided into a plurality of calculation steps, and corresponding operation enable signals are generated and transmitted to the calculation unit; the calculation unit obtains target parameters required by calculation steps from the storage unit according to the operation enable signal, obtains floating point operator units required by the calculation steps from the floating point operator module, and executes the calculation steps to determine an operation intermediate result; and storing the operation intermediate result in a storage unit, generating an operation completion signal and transmitting the operation completion signal to a main scheduling module to control the next calculation step until a final operation result is obtained. Therefore, the FPGA architecture with low power consumption and miniaturization is realized.
Owner:EHIWAY MICROELECTRONIC SCI & TECH (SUZHOU) CO LTD

Internal memory, chip and related electronic equipment

The embodiment of the invention discloses an internal memory, a chip and related electronic equipment. A target computing unit of the internal memory comprises a plurality of groups of primary multipliers, a plurality of adder trees and a plurality of secondary multipliers, one group of first-level multipliers is connected with one second-level multiplier through one adder tree; the group of first-level multipliers comprises a plurality of pairs of first multipliers, and each adder tree comprises a first adder, a second adder and a third adder; each pair of first multipliers respectively receive data to perform multiplication operation, and respectively input output results into the first adder to perform addition operation; each second adder receives the output results of the plurality of first adders or the second adder at the upper level to carry out addition operation, and respectively inputs the output results to the second adder at the lower level or the third adder; and each third adder receives the output results of the plurality of second adders to carry out addition operation, and inputs the output results to the secondary multiplier to carry out multiplication operation. By implementing the embodiment of the invention, the memory computing capability can be improved.
Owner:HUAWEI TECH CO LTD

Hardware accelerator with matrix block streaming

A hardware accelerator including tiles arranged in a systolic array. At each of the tiles, the systolic array receives a first input block that includes first input matrix elements of a first input matrix. In each of a plurality of multiplication iterations, at each of the tiles, the systolic array receives a respective second input block. The systolic array computes tile products of the first input matrix elements and second input matrix elements included in the second input blocks. The systolic array adds the tile products to column-wise partial sums and transmits the column-wise partial sums to subsequent tiles along accumulator rings included in array columns of the systolic array. In a subset of the multiplication iterations, the systolic array outputs product block rows of a product matrix. The product block rows each include product matrix blocks computed as rows of the column-wise partial sums.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Annotation tools for reconstructing three-dimensional roof geometry

Methods and systems for generating geometric models of pitched roofs of buildings based on multiview imagery are provided. An example method involves establishing a three-dimensional coordinate space for the multiview imagery, displaying, through an annotation platform, multiview imagery depicting the pitched roof from different points of view, and providing, through the annotation platform, functionality for a user to define a perimeter of the pitched roof, define a ridge line of the pitched roof, and define at least part of a roof facet of the pitched roof by connecting the ridge line of the pitched roof to the perimeter of the pitched roof.
Owner:ECOPIA TECH CORP

Hardware compression for sparse matrix content

One embodiment provides a graphics processor comprising a base die including a plurality of chiplet sockets and a plurality of chiplets coupled with the plurality of chiplet sockets. At least one of the plurality of chiplets include a graphics core cluster including a plurality of processing elements, a shared local memory coupled with the plurality of processing elements, a plurality of matrix engines coupled with the shared local memory, and codec circuitry coupled with the shared local memory and the plurality of matrix engines. The codec circuitry is configured to decode matrix data stored in the shared local memory in a first format into a second format for consumption by the plurality of matrix engines.
Owner:INTEL CORP

Configurable module-lattice post-quantum cryptography processor for key-encapsulation mechanism

Disclosed is a reconfigurable module-lattice-based key encapsulation mechanism (ML-KEM) post-quantum cryptography system and method using memory-based numbers theoretic transform (NTT). A post-quantum cryptography method of a post-quantum cryptography system including a plurality of internal submodules includes reconfiguring the plurality of internal submodules by variably selecting one security level from among the plurality of security levels; reconfiguring execution of the plurality of internal submodules to be changed through a main controller; and variably processing data according to the selected security level to perform key generation, encapsulation, and decapsulation through the reconfigured plurality of internal submodules.
Owner:INHA UNIV RES & BUSINESS FOUNDATION

Quaternary data processing system based on multi-level signal and dual-storage architecture

The invention discloses a quaternary data processing system based on a multi-level signal and a dual-storage architecture, and relates to the technical field of data processing, in particular to a quaternary data processing system combining multi-level signal coding, a dual-storage architecture synchronization strategy and a quantum compatibility design. According to the system, through collaborative design of multi-level signals and a double-storage architecture, the density bottleneck and anti-interference shortages of a traditional binary system are solved, meanwhile, a classic data interface is provided for quantum calculation, and the system can be widely applied to the fields of power line communication, industrial automation, quantum calculation acceleration and the like.
Owner:王启枝

Validating autonomous vehicle simulation scenarios

Validating a simulation scenario for use in training a machine learning model for an autonomous vehicle includes determining a simulation scenario; executing a simulation based on a simulation scenario; monitoring execution of the simulation and receiving messages from the execution of the simulation; determining, by a first simulation monitor, whether the messages satisfy a first incident; and validating the simulation scenario responsive to the messages satisfying the first incident to produce validated simulation data. A system and method may also include determining, by a first simulation validator, whether the messages satisfy a condition; and validating the simulation scenario responsive to the messages satisfying the condition to produce validated simulation data.
Owner:AURORA OPERATIONS INC

Polynomial multiplication accelerator

The invention provides a polynomial multiplication accelerator, and relates to the technical field of multiplication accelerators, and the polynomial multiplication accelerator comprises a calculation module which comprises a number theory transformation unit, an iterative convolution unit and an inverse number theory transformation unit; the number-theory transformation unit is configured to execute number-theory transformation operation on the to-be-calculated polynomial in the global ring by using the first butterfly subunit of the execution layer number to obtain a first intermediate result; the iterative convolution unit is configured to execute iterative convolution operation on the first intermediate result in the quotient ring to obtain a second intermediate result; and the inverse number theory transformation unit is configured to execute inverse number theory transformation operation on the second intermediate result in the global ring by utilizing the second butterfly subunit of the execution layer number to obtain a target result, so as to solve the problem that the current polynomial multiplication accelerator can forcibly and completely execute butterfly operations of all layers, and the operation efficiency is high. And therefore, a part of computing resources are used for processing low significant bit operations with relatively low contribution degrees to results, and significant computing power redundancy is caused.
Owner:NANJING UNIV

Complex multiplication implementation method for parameter serial and parallel mixed input

The invention discloses a parameter serial and parallel mixed input complex multiplication implementation method, which is implemented by adopting a complex multiplier circuit, and the complex multiplier circuit comprises a serial input cache unit, a parallel input cache unit, a real number multiplication unit and a real number addition / subtraction unit. The method comprises the specific implementation steps that S1, serial input complex number parameters and parallel input complex number parameters are sequentially input from a serial input cache unit and a parallel input cache unit respectively; s2, the serial input complex parameters and the parallel input complex parameters are subjected to combined calculation through a real number multiplication unit and a real number addition / subtraction unit; and S3, outputting a real part and an imaginary part of a complex multiplication result in sequence in a serial mode. Compared with a traditional and optimized triple multiplier version complex multiplication operation circuit, only a small number of registers are additionally arranged, latch resources are conditionally selected, and the complex multiplication operation circuit has obvious advantages in hardware resource consumption.
Owner:CHENGDUSCEON TECH

Synthetic generation of simulation scenarios and probability-based simulation evaluation

Techniques are discussed herein for generating and evaluating driving simulations based on synthetic scenarios. Simulated objects may be controlled based on parameters determining the attributes and behaviors of the objects, and scenarios may be synthetically modified by changing the parameters for a simulated object. For a driving scenario with a synthetic simulated object, a simulation system may analyze driving log data to determine a probability or prevalence associated with the driving scenario. In various examples, the simulation system may determine marginal density estimates for individual attributes of the simulated object, as well as a cumulative distribution function modeling the dependence between the attributes. The joint probability distribution determined for the synthetic simulated object can be used for evaluating the efficacy of the simulation and the performance of the simulated vehicle controllers.
Owner:ZOOX INC

Method for performing matrix multiplication operation by processor comprising plurality of computing units

The invention relates to the technical field of computers, and provides a method for executing matrix multiplication by a processor comprising a plurality of computing units. In the method, a quantization parameter matrix is divided into a plurality of parameter blocks, and the parameter blocks are sequentially and respectively distributed to a plurality of calculation units in a non-repeated manner, so that at most one calculation unit is distributed with less than a first number of continuous parameter blocks in the plurality of parameter blocks, each computing unit is allocated a first number of consecutive parameter blocks in the plurality of parameter blocks; each calculation unit is used for carrying out dequantization on the distributed parameter blocks; each of the plurality of calculation units obtains a portion of the input matrix that should be multiplied by the allocated parameter block, and performs matrix multiplication on the dequantized allocated parameter block and the portion of the input matrix to complete matrix multiplication of the input matrix and the model parameter. Therefore, only one-time solution quantization needs to be carried out on the quantized parameter matrix globally, calculation resources are saved, and the reasoning efficiency is improved.
Owner:SHANGHAI INFINIGENCE AI INTELLIGENT TECHNOLOGY CO LTD

Hardware accelerator with scale factor applied at tensor processor

A hardware accelerator including input memory that receives first and second input matrices. The hardware accelerator further includes processing circuitry including one or more tiles that each include a respective tensor processor configured to receive a first and second input block of the first and second input matrices. Each tile receives a first block scale factor associated with rows of the first input block and a second block scale factor associated with columns of the second input block. Each tile multiplies the first input block by the second input block, applies the first block scale factor to rows of the result block, and applies the second block scale factor to columns of the result block to obtain a scaled result block. The processing circuitry further includes an accumulator that accumulates scaled result blocks to obtain a scaled result matrix, and output memory that receives and output the scaled result matrix.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

FPGA (Field Programmable Gate Array) comprehensive tool carry chain optimization method and device oriented to continuous addition of same numbers

The invention provides an FPGA (Field Programmable Gate Array) comprehensive tool carry chain optimization method and device oriented to continuous addition of the same number. The method comprises the following steps: acquiring the same signal and continuous addition times n in a first addition unit; acquiring a binary array b [0-log2n] of continuous addition times; if b [i] = 1, generating a left shift operation unit, mapping the left shift operation unit into a trigger unit chain table, and storing an output signal of the trigger unit chain table into an array N1; generating log2n second addition units according to the array N1, and storing output signals of the second addition units to an array C to obtain sets N2 and N3; acquiring a mapping set from the first-stage output signal which does not need to be processed to a second addition unit, and updating sets N2 and N3; and traversing each signal sig in the sets N2 and N3, and mapping a second addition unit corresponding to the signal sig into a carry chain unit chain table. The resource allocation of the carry chain can be optimized, and the number of logic resources after synthesis is reduced.
Owner:EHIWAY MICROELECTRONIC SCI & TECH (SUZHOU) CO LTD

Multi-precision floating point fusion multiply-add structure and microprocessor architecture

The invention provides a multi-precision floating point fusion multiply-add structure and a microprocessor architecture, and the structure is characterized in that in a first-stage assembly line, an input preprocessing module decomposes an input operand to obtain a sign bit, an index bit and a mantissa bit; the Booth encoder generates partial products through a radix-4 Booth encoding algorithm, and the tree array multiplier compresses the partial products into three groups; in the second-stage assembly line, the CSA4-2 compression adder compresses the mantissa bits after the three groups of partial products and shift alignment into two groups; the summator module sums the two groups of partial products to obtain a mantissa summation result; in the third-stage assembly line package, a leading zero detection module performs leading zero detection on the mantissa summation result, and performs index adjustment and mantissa adjustment to obtain a normalized result; and in the fourth-stage assembly line, rounding of floating point data is carried out. According to the application, the operation performance, the energy efficiency ratio and the adaptability of the floating point fusion multiply-add structure are improved, so that the high performance and the high energy efficiency of the microprocessor are balanced.
Owner:EHIWAY MICROELECTRONIC SCI & TECH (SUZHOU) CO LTD

Systems and methods for generating interior designs using artificial intelligence

Systems, methods and / or interfaces are provided for generating interior design options. In some implementations, the method includes obtaining an inspiration image of a portion of a room, a furniture, fixture and equipment, or a combination. The method also includes parsing the image, including segmenting the inspiration image into sub-images corresponding to each furniture, fixture or equipment. The method also include identifying alternatives based on the sub-images and a layout of the room, including coordinating the alternatives with respect to each other and coordinating the alternatives with respect to the room. The method also includes generating and / or displaying variations of the room with images corresponding to the alternatives. In some implementations, the method includes parsing an image to generate an empty room sketch or schematic, generating alternatives to be placed in the empty room according to the room, and generating and / or displaying a visualization by placing the alternatives in the room.
Owner:INHABITR INC

High-speed fixed-point multiplication circuit

The invention discloses a high-speed fixed-point multiplication circuit. The multiplication circuit is mainly composed of a multiplier coding module, a partial product generation module, a partial product compression module and a traveling wave carry adder module. The multiplier coding module is composed of a radix-4-Booth coding algorithm and an opposite number generation module, and is used for carrying out three-bit block coding on input multiplication data and generating an opposite number of a multiplicand in advance. The partial product generation module generates a plurality of groups of partial products with symbol extension according to the multiplicand coded signal. And the partial product compression module adopts an improved Wallace compression structure to perform layered compression on the partial product. And the traveling wave carry adder module sums the two groups of results output by compression and outputs a multiplication result. According to the invention, multiplication accumulation series can be reduced, the switching times of invalid signals in the circuit can be reduced, and the longest delay path in an operation link can be shortened, so that fixed-point multiplication which is high in speed, low in power consumption and more favorable for a comprehensive tool in structure is realized.
Owner:SHENYANG UNIVERSITY OF TECHNOLOGY

MAC operational circuit based on double-bit-line self-circulation switching, CIM circuit and chip

The invention relates to the field of integrated circuits, in particular to an MAC operational circuit, a CIM circuit and a chip based on double-bit-line self-circulation switching. The circuit comprises a calculation array, a transmission array, a quantization unit, a signal switching unit and an output module. The calculation array is composed of two twinborn calculation units connected to each SRAM unit, and the two twinborn calculation units have the same multiplication function; the signal switching unit is used for outputting a calculation control signal so as to alternately discharge the calculation bit lines connected with the two twin calculation units, and controlling the two calculation bit lines in the transmission array to be alternately connected to the ADC; the ADC of the quantization unit combines the total discharge capacity of the two calculation bit lines to obtain the value of the low-order MAC result, and the output unit obtains the value of the low-order MAC result according to the number of times of overturning of the output signal of the signal switching unit. The problems that an existing circuit is insufficient in precision and large in area and power consumption are solved.
Owner:ANHUI UNIV

Efficient data compression in processing systems

Certain aspects of the present disclosure provide techniques and apparatus for efficiently performing operations using data compression. An example method generally includes identifying, for a block of data samples, a number of leading bits to remove from each data sample in the block of data samples. A block of compressed data samples is generated based on the identified number of leading bits and truncation of a number of least significant bits from each data sample in the block of data samples. A bitstream including the block of compressed data samples and an indication of a type of compression applied to the block of data samples is generated and output for further processing.
Owner:QUALCOMM INC

Method for automatically generating design solutions for an engineering design project

One variation of a method includes: accessing a descriptor of a project; extracting a set of language signals from the descriptor; accessing a set of input parameters and a set of output characteristics; querying a language model for a range of values of input parameters exhibited within historical engineering solutions correlated with the set of language signals; accessing a virtual model for the project and a function representing a relationship between the set of input parameters and the set of output characteristics; for each analysis instance in a count of analysis instances, defining a combination of values of input parameters within corresponding ranges and based on the virtual model, the function, and the combination of values, executing the analysis instance to calculate a set of values of output characteristics; and rendering representations of combinations of values of input parameters and sets of values of output characteristics within a user interface.
Owner:GENERATIVE VISION LTD

Data processing method, device, storage medium, and electronic device

The present invention provides a method and apparatus for processing data, a storage medium, and an electronic device; wherein, the method includes: reading the feature map data of M*N of all input channels and the weights of a preset number of output channels, where the values of M*N and the preset number are respectively determined by a preset Y*Y weight; inputting the read feature map data and the weights of the output channels into a multiply-accumulate array of the preset number of output channels for convolution calculation; wherein, the convolution calculation method includes: not performing convolution calculation when the feature map data or the weights of the output channels are zero; when there are multiple identical values of the feature map data, selecting one from the multiple identical values for convolution calculation; outputting the result of the convolution calculation. Through the present invention, the problem in the related art of the lack of how to efficiently accelerate the convolution part in artificial intelligence is solved.
Owner:SANECHIPS TECH CO LTD

System and method for generating a landscape design

A method for generating an online landscape design for a property includes first providing a computing system comprising at least a memory storing computer-executable instructions of a landscaping application, and a processor coupled to the memory. The landscaping application comprises a calculator engine, a landscape design engine, and a scoring engine. Next, calculating a landscape score of a property by retrieving landscape data from online databases and comparing an existing landscape design of the property to the retrieved landscape data via the scoring engine. Next, determining property landscape improvements by the scoring engine, and then generating an improved landscape design for the property based on the property landscape improvements of the scoring engine using the landscape design engine. Next, calculating an improved landscape score of the improved design via the scoring engine and displaying a 3D-image of the property with the improved landscape design and the improved score.
Owner:HOME OUTSIDE INC