Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

172results about "Handling data according to predetermined rules" patented technology

Tensor transpose processor

The present invention relates to a processor designed to optimize memory bandwidth utilization for tensor transpositions in machine learning. An example processor includes an input tensor shift buffer, a staging buffer, and an output tensor shift buffer. The input tensor shift buffer reads an input tensor from input memory and performs multiple cycles of input tensor shifting. The shifted tensor data is then written into the staging buffer. The output tensor shift buffer reads the shifted tensor data from the staging buffer and performs multiple cycles of output tensor shifting. Finally, the result is written to the output memory. This configuration facilitates efficient handling and transformation of tensor data, optimizing the computational processes required in machine learning tasks.
Owner:MOFFETT TECH CO LTD

Method and system for processing operation data of brine discharging pipeline of salt cavern gas storage

The invention relates to the technical field of data processing, and discloses a salt cavern gas storage brine discharging pipeline operation data processing method and system, and the method comprises the steps: periodically receiving a link detection data packet which is transmitted by a collection device and is provided with a transmission timestamp, and recording the arrival time of the link detection data packet; according to the sending timestamp and the arrival time, transmission characteristics of the link detection data packet are analyzed, and the transmission characteristics comprise transmission delay, jitter dispersion and out-of-order conditions; adaptively predicting a time offset pattern based on the transmission characteristics; generating a real-time calibration parameter according to the time migration mode; and receiving an actual data packet, and adjusting the timestamp of the actual data packet by using the real-time calibration parameter. According to the method, the limitation of traditional fixed compensation parameters is overcome, the time alignment precision of pressure and flow data flow is remarkably improved, leakage misinformation caused by time deviation is avoided, and the reliability and the operation and maintenance efficiency of a salt cavern gas storage brine discharging pipeline operation monitoring system are improved.
Owner:WUHAN CENT CHINA GEOLOGICAL SURVEY CENT SOUTH CHINA INNOVATION CENT FOR GEOSCIENCES

Data processing method and device, electronic equipment and nonvolatile storage medium

The invention discloses a data processing method and device, electronic equipment and a nonvolatile storage medium. The method comprises the following steps: dividing a data matrix to be calculated into a plurality of data blocks with target sizes; data corresponding to the data blocks are allocated to a plurality of stream processors for data reading operation, the data reading operation is used for loading the data from a global memory of the edge computing device to a shared memory, and the data reading and writing efficiency of the shared memory is higher than that of the global memory; and loading the data in the memory block with the preset vector length in the shared memory into a register by taking the preset vector length as a unit, and operating the data loaded into the register by adopting a stream processor to obtain an operation result. The technical problem that the computing efficiency is not high due to the fact that the related technology is limited by the particularity of the memory access and computing rule of the correction operator and the hardware performance of the edge computing equipment cannot be fully exerted is solved.
Owner:北京大学长沙计算与数字经济研究院

Floating point data parallel computing method and device of vector processor and vector processor

The invention provides a floating point data parallel computing method and device based on a vector processor and the vector processor, and the method comprises the steps: carrying out vectorization and parallel absolute value calculation on a plurality of pieces of input floating point data, and dividing each piece of data into different computing intervals according to the calculated absolute value; for data located in different calculation intervals, calculating branches in different intervals are combined, branch differences are uniformly represented through symbol transformation variables, and fitting intermediate variables corresponding to all the data are obtained through uniform vector operation instruction parallel calculation; and performing parallel calculation on the basis of the intermediate variable for fitting to obtain a preliminary result corresponding to each data, and correcting the preliminary result according to a symbol of original input data to output a final calculation result in a vector form. According to the method, the problems of low hardware resource utilization rate and poor calculation efficiency of high-precision floating point function calculation realized by adopting a scalar serial processing mode in the prior art are solved.
Owner:SHANGHAI SMARTLOGIC TECHNOLOGY LTD

Accelerating table lookups using a system-on-chip, separate lookup table accelerator

To provide a VPU and associated components that are optimized to improve VPU performance and throughput.SOLUTION: A VPU includes a min / max collector, an automatic store predication function, a SIMD data path configuration allowing inter-lane sharing, a transposed load / store including a stride parameter function, a load including a permute and zero insertion function, a hardware, logic device, and memory layout function allowing two point and two by two point lookups, and per memory bank load caching ability. Decoupled accelerators are used to offload VPU processing tasks, and a hardware sequencer is included in a DMA system. The DMA and VPU execute a VPU configuration mode allowing the VPU and DMA to operate without a processing controller for executing dynamic region-based data movement operations.SELECTED DRAWING: Figure 9A
Owner:NVIDIA CORP

Continuous bit parallel counting device and method for preamble and mantissa of processor

ActiveCN121050687AHandling data according to predetermined rulesComputation using non-denominational number representationComputer architectureHemt circuits
The invention relates to the field of integrated circuit design, and discloses a processor preamble and mantissa continuous bit parallel counting device and method. The device comprises a 32-bit bit sequence flipping module, a 32-bit bit negation module, eight 4-bit leading zero detection modules, four 8-bit leading zero processing modules, two 16-bit leading zero processing modules, a 32-bit leading zero processing module and a 32-bit result output module. According to the invention, parallel counting can be carried out on leading zeros, leading 1, mantissa zeros or mantissa 1 of various data with different precisions in a single processor clock period, so that the required time sequence path length is greatly shortened, the higher processor clock frequency is realized, the multiplexing of a hardware circuit is realized, and the cost is reduced. The circuit greatly reduces the use of transistor resources, reduces the area of a chip, reduces the cost and power consumption of the chip, improves the operation efficiency, universality and adaptability of the processor, and is suitable for being applied to application scenes such as a high-performance digital signal processor and a CPU (Central Processing Unit).
Owner:青岛本原微电子有限公司

Data processing method and device in assembly line, medium, equipment and product

The invention discloses a data processing method and device in an assembly line, a medium, equipment and a product, and the method comprises the steps: obtaining a barrier identifier before the execution of a to-be-processed data stage in the current assembly line according to an assembly line identifier and a stage identifier of the current assembly line; comparing the completion signal count associated with the barrier identifier of the current assembly line with the expected signal number, and triggering the current assembly line to execute a to-be-processed data stage under the condition that the completion signal count is equal to the expected signal number; after the current assembly line finishes executing the to-be-processed data stage, obtaining a barrier identifier before executing the to-be-processed data stage in the next assembly line; and sending a completion signal to the synchronization barrier indicated by the barrier identifier of the next assembly line through the current assembly line, and updating the stage identifier of the current assembly line so as to enter a next data processing stage. According to the method, the execution time sequence of the assembly line can be accurately controlled, and more efficient and more accurate scheduling and utilization of hardware resources are realized.
Owner:SHANGHAI BIREN TECH CO LTD

Performance management in data orchestrated environments

This disclosure provides methods, devices, and systems for data management. The present implementations more specifically relate to a data orchestration system that can dynamically reconfigure a data processing pipeline based on telemetry received from various steps or data operations in the pipeline. For example, the telemetry may indicate a success, failure, time of entry, time of exit, or total duration of a given step or data flow in the processing pipeline. In some aspects, the data orchestration system may dynamically invoke new data flows based on the received telemetry. In some implementations, the new data flows may allocate additional memory and / or processing resources for the data processing pipeline. In some other implementations, the new data flows may deallocate memory and / or processing resources for the data processing pipeline. Still further, in some implementations, the new data flows may trigger an alert to a user or manager of the data processing pipeline.
Owner:VIEW SYSTEMS INC

Transposing at-speed in a vector-matrix accelerator

A system including one or more processors configured to receive a transpose instruction indicating to transpose a source matrix to a result matrix, provide data elements of the source matrix to input switching circuits, reorder the data elements using the input switching circuits, provide the data elements from the input switching circuits to one or more lanes of a datapath, provide the data elements from the datapath to output switching circuits, undo the reordering of the data elements using the output switching circuits, and provide the data elements from the output switching circuits to a result matrix. Each respective lane of the datapath receiving data elements receives multiple data elements directed to different respective non-overlapping portions of the lane.
Owner:GOOGLE LLC

Masked Shift-Add Operation

The computer-implemented method includes receiving, by a processing unit, an instruction to perform a masked shift-and-add operation with a set of operands. A logical AND operation is performed on a first pair of operands in the set of operands to obtain a first intermediate result. The first intermediate result is shifted by a first shift amount based on a first operand of the first pair of operands. A logical AND operation is performed on a second pair of operands in the set of operands to obtain a second intermediate result. The second intermediate result is shifted by a second shift amount based on the first operand of the second pair of operands. The shifted first intermediate result is added to the shifted second intermediate result. The method further includes outputting an output of the addition as a result of the masked shift-and-add operation.
Owner:INTERNATIONAL BUSINESS MACHINE CORPORATION

Processor operand management using fused buffers

Techniques relating to operand management using a fused buffer are disclosed. A processor includes an operand management circuit, wherein the operand management circuit includes a fusion buffer; and an execution circuit. In one embodiment, operand management circuitry is configured to: detect a first store instruction operation executable to store an operand value usable by one or more consumer instruction operations; and storing the first store instruction operation in the fusion buffer. In response to detecting a drop condition associated with the first store instruction operation, the operand management circuitry is configured to remove the first store instruction operation from the fusion buffer without forwarding the first store instruction operation for execution. The operand management circuitry is configured to forward the first stored instruction operation for execution by the execution circuitry in response to detecting a buffer emptied condition and not detecting a discard condition.
Owner:APPLE INC

An iris image matching method, device, equipment and medium

The application relates to the field of image processing, in particular to an iris image matching method and device, equipment and medium, which are used to solve the problem of large calculation error in the iris image matching process. The method determines a plurality of target Mahalanobis distances based on each feature point of a to-be-detected iris image. It should be noted that each target Mahalanobis distance is determined based on the gradient parameters corresponding to any two target feature points of the to-be-detected iris image. Based on the sum of the plurality of target Mahalanobis distances and the sum of a plurality of standard Mahalanobis distances, the matching degree between the to-be-detected iris image and a pre-stored standard iris image is determined, wherein the standard Mahalanobis distance is determined based on the standard iris image. The above-mentioned iris image determination method based on Mahalanobis distance considers the internal relationship between the feature points, improves the matching accuracy between the to-be-detected iris image and the pre-stored standard iris image, and improves the reliability of the safe box in use.
Owner:CHINA CONSTRUCTION BANK +1

Systems, methods, and apparatuses for tile transpose

Embodiments detailed herein relate to matrix operations. In particular, support for a matrix transpose instruction is detailed. In some embodiments, decode circuitry to decode an instruction having fields for an opcode, a source matrix operand identifier, and a destination matrix operand identifier; and execution circuitry to execute the decoded instruction to transpose each row of elements of the identified source matrix operand into a corresponding column of the identified destination matrix operand are detailed.
Owner:INTEL CORP

Interconnect mode for computational arrays

A processing engine array is provided with an interconnect mode of operation to use the array as an interconnect to move data elements to different locations in memory such as to perform a matrix transpose operation. In this interconnect mode of operation, although computations are still being performed in the array, the computations are not carried out to modify or change the values of the data elements, but are instead carried out to rearrange the data elements in memory. As such, the computations carried out in the interconnect mode of operation can deviate from the expected behavior of floating-point calculations. A mode selection signal can be used to provide the proper outputs of the processing elements of the array depending on the mode of operation.
Owner:AMAZON TECH INC

Memristor measurement matrix determination method and device, equipment, storage medium and product

The application provides a memristor measurement matrix determination method, device, equipment, storage medium and product. The method comprises the following steps: determining the maximum value of the candidate measurement matrix bit width according to the conductance deviation characteristics of the target memristor and the target compression ratio; obtaining the target bit width interval based on the least significant bit width corresponding to the target memristor and the maximum value; generating the candidate measurement matrix corresponding to each bit width in the target bit width interval based on the statistical distribution characteristics of the original measurement matrix; performing signal reconstruction simulation based on the preset noise simulation model corresponding to the target memristor and each candidate measurement matrix to obtain the signal reconstruction accuracy corresponding to each candidate measurement matrix; and taking the candidate measurement matrix with the highest signal reconstruction accuracy as the measurement matrix of the memristor array corresponding to the target memristor, so as to enable the memristor array to perform compressive sensing. The measurement matrix used can meet the actual situation of the operation of the target memristor, and overcome the negative influence of the device characteristics on the reconstruction accuracy of compressive sensing.
Owner:TSINGHUA UNIVERSITY

Matrix-based intra prediction using upsampling

Devices, systems, and methods for digital video coding, which includes matrix-based intra prediction methods for video coding, are described. In a representative aspect, a method for video processing includes performing a conversion between a current video block of a video and a bitstream representation of the current video block using a matrix based intra prediction (MIP) mode in which a prediction block of the current video block is determined by performing, on previously coded samples of the video, a boundary downsampling operation, followed by a matrix vector multiplication operation, and followed by an upsampling operation, where the upsampling operation is performed, in both a vertical direction and a horizontal direction in a fixed order, on samples obtained from the matrix vector multiplication operation.
Owner:BYTEDANCE INC +1

ALU, processor, chip, device and texel index calculation method

The invention discloses an ALU, a processor, a chip, equipment and a texel index calculation method, and belongs to the technical field of chips. The ALU comprises a multiplier used for calculating a product of a mantissa of an input texture coordinate in a floating-point number format and a texture size to obtain a first product; the first conversion unit is used for leftwards shifting the input offset by N bits to obtain the shifted input offset; the adder is used for calculating the sum of the first product and the shifted input offset to obtain a first calculation result; the second conversion unit is used for rightwards shifting the first calculation result by M bits to obtain a shifted first calculation result; based on the shifted first calculation result and the rounding mode, a second calculation result is obtained, and the second calculation result is used for determining a texel index corresponding to the input texture coordinates. According to the method, the second calculation result in the preset fixed-point format can be obtained only by executing the unified right-shift operation once, and the number of times of the right-shift operation is reduced, so that the determined texel index is more accurate.
Owner:MOORE THREADS TECH CO LTD

Matrix data protocol processing device and method

The invention discloses a matrix data protocol processing device and method. According to the embodiment of the invention, a cooperative effect is achieved under a closed loop of tile loading, hardware grouping rearrangement, parallel protocol and fixed prefix slot single-beat write-back; according to the method, higher throughput, lower end-to-end delay, less data migration / instruction quantity and better energy efficiency are obtained under the same example and resource constraint, and consistent and efficient hardware support is provided for a matrix protocol in the row / column direction through a unified instruction interface.
Owner:SUNMMIO SCIENCE & TECHNOLOGY (BEIJING) CO LTD

Methods and related devices for determining model output results

This application discloses a method and related apparatus for determining the output result of a model, belonging to the field of data processing technology. The network model includes an attention mechanism module; the method includes: determining a query matrix, a key matrix, and a value matrix through the attention mechanism module based on the current input information of the network model; processing the transpose of the key matrix to obtain an intermediate matrix, wherein the intermediate matrix has the same number of rows and columns as the transpose matrix, and the average value of each column element in the intermediate matrix is ​​less than the average value of the corresponding column elements in the transpose matrix; determining an output matrix through the attention mechanism module based on the query matrix, the intermediate matrix, and the value matrix; and determining the current output result of the network model based on the output matrix. Therefore, this application reduces the value of elements in the transpose matrix by processing it, thereby preventing numerical overflow and computational instability at the computational source.
Owner:HUAWEI TECH CO LTD

Security Device

According to various embodiments, a security device is provided comprising a modular reducer configured to perform a modulo reduction by a modulus of each binary number of a sequence of binary numbers forming a data word, wherein each binary number consists of n bits by one or more first iterations comprising, in reaction to a first detector of the security device detecting that the most significant bit (MSB) of the binary number is set, changing the binary number by deleting its MSB and adding the difference between 2n−1 and the modulus to the binary number, followed by one or more second iterations comprising, in reaction to a second detector of the security device detecting that the MSB of the sum of the binary number with the difference between 2n−1 and the modulus is set, setting the binary number to that sum, wherein the MSB of the sum is deleted.
Owner:INFINEON TECHNOLOGIES AG

Method and apparatus for unified codebook for orthogonal sequence transmission

The present disclosure relates to methods and devices (including apparatuses, such as UEs and / or base stations) for wireless communication. The apparatus can select one or more rows or one or more columns of a DFT matrix, the one or more rows being even rows in the DFT matrix and the one or more columns being even columns in the DFT matrix. The apparatus can also determine an orthogonal matrix based on the one or more rows or the one or more columns of the DFT matrix, the orthogonal matrix having a size of (MxN)x(MxN) with MxN rows and MxN columns. Additionally, the apparatus can determine a codebook based on the orthogonal matrix, the codebook including a plurality of codepoints. The apparatus can also transmit at least one signal including a first codepoint of the plurality of codepoints in the codebook in an uplink resource.
Owner:QUALCOMM INC

Data processing apparatus and methods for tensor transform operation

Data processing apparatus (DPA) for processing resource to perform transform operation on input tensor for processing resource. Input tensor is formed of blocks, each block being a portion of the inpu
Owner:ARM LTD

Multiple operation circuits, multiplication / accumulation operators having the multiple operation circuits, and processing-in-memory devices having the multiple operation circuits

A multiple operation circuit includes a multiplier, an adder, a latch circuit, and a plurality of selectors. The multiplier performs a multiplying calculation of first input data and second input data to generate and output multiplication result data. The adder performs an adding calculation of third input data and fourth input data to generate and output addition result data. The latch circuit latches fifth input data input to an input terminal of the latch circuit to generate and output feedback data. The plurality of selectors change transmission paths of first result data, the first input data, the second input data, the multiplication result data, and the addition result data according to a first operation mode, a second operation mode, or a third operation mode.
Owner:SK HYNIX INC

Vector processing circuit and vector processing method

The invention provides a vector processing circuit and a vector processing method. A vector processing circuit includes an instruction queue, a plurality of computing circuits, and a control circuit. The instruction queue includes a first reduction instruction and a second reduction instruction. The computing circuit has a plurality of pipeline stages. The control circuit is electrically connected to the instruction queue and the calculation circuit. The computing circuit alternately generates results of the first reduced instruction and the second reduced instruction in the plurality of clocks. Therefore, the overall throughput can be increased.
Owner:ANDES TECH

Variable bit width matrix multiplication

Systems and methods for performing variable bit width matrix multiplication are provided. For example, a processor device may include dot product hardware configured to perform a plurality of dot products at a first bit width to produce a plurality of first bit width dot product outputs. The processor device may include programmable adder hardware. The programmable adder hardware may be configured to obtain data indicative of one or more target bit widths. The programmable adder hardware may be configured to combine one or more subsets of the plurality of first bit width dot product outputs according to the one or more target bit widths based on data indicative of the one or more target bit widths.
Owner:GDM HOLDING LLC

Secure computer-implemented method for preventing a recovery of embedded data within a neural network model

The invention relates to a secure computer-implemented method (1) for preventing a recovery of embedded data (d) within an neural network model (NN), said neural network model (NN) comprising a plurality of layers (L), each layer (L) having a related matrix of parameters (M) and being configured to receive at least one input tensor (t1), wherein said secure computed implemented method (1) comprises: - for at least one layer (L), permuting sets (s) of parameters (P) within its related matrix of parameters (M) so as to change their initial positions (p) in said matrix of parameters (M), - applying said matrix of permuted parameters (M') to the at least one input tensor (t1) so as to generate an output tensor (t2').
Owner:THALES DIS FRANCE SA