Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

275results about "Handling data according to predetermined rules" patented technology

Matrix transposition method of DDR3 read-write controller based on AXI4 bus

The invention discloses a matrix transposition method of a DDR3 read-write controller based on an AXI4 bus. The matrix transposition method solves the problems that an existing transposition method is low in data access efficiency, serious in storage resource waste and poor in system portability. The method comprises the following steps: dividing an original echo matrix into a plurality of sub-matrix blocks, mapping and storing the sub-matrix blocks into a Bank of DDR3 to obtain a three-dimensional mapping result; reading data in each sub-matrix block in the three-dimensional mapping result by utilizing a DDR3 read control module, and writing the read data into an RAM (Random Access Memory) of an FPGA (Field Programmable Gate Array) to obtain a transposed matrix; writing the transposed matrix into the original Bank of the DDR3 by using a DDR3 write control module to obtain the DDR3 in which the data is written; according to the invention, low on-chip storage resource occupation independent of matrix scale is realized, and high-efficiency processing requirements of large-scale echo data in satellite-borne SAR real-time imaging processing are better met.
Owner:XIDIAN UNIV

Character recognition model training method and apparatus, character recognition method and apparatus, device and storage medium

The present disclosure provides a character recognition model training method and apparatus, a character recognition method and apparatus, a device and a medium, relating to the technical field of artificial intelligence, and specifically to the technical fields of deep learning, image processing and computer vision, which can be applied to scenarios such as character detection and recognition technology. The specific implementing solution is: partitioning an untagged training sample into at least two sub-sample images; dividing the at least two sub-sample images into a first training set and a second training set; where the first training set includes a first sub-sample image with a visible attribute, and the second training set includes a second sub-sample image with an invisible attribute; performing self-supervised training on a to-be-trained encoder by taking the second training set as a tag of the first training set, to obtain a target encoder.
Owner:BEIJING BAIDU NETCOM SCI & TECH CO LTD

Data sorting device and method, chip, electronic equipment and storage medium

The invention discloses a data sorting device and method, a chip, electronic equipment and a storage medium, and the device comprises a controller which is used for obtaining a plurality of first data groups; a parity sorting network for sorting the first data set into a second data set; the data selector is used for acquiring a target data group; the controller is used for comparing the target data set with the ith second data set at an interval of N bits to obtain an (i-1) th third data set; when the TopN is sorted, the double-adjustment merging network is used for sorting the (i-1) th third data group into the (i-1) th fourth data group; when i is equal to the number of the plurality of first data groups, taking the (i-1) th fourth data group as a TopN sorting result; and during full sorting, the double-adjustment merging network is used for carrying out iterative merging on the sorted fourth data group to obtain a full sorting result. The method can be compatible with TopN and a dual-tone full-sorting algorithm, and saves computing resources.
Owner:SHANGHAI ORIENTAL COMPUTER TECHNOLOGY CO LTD

Multi-scalar multiplication acceleration method based on resource pre-estimation and pre-calculation strategy

The invention discloses a multi-scalar multiplication acceleration method based on resource estimation and a pre-calculation strategy, and provides a systematic solution for the problems of insufficient resource estimation, pre-calculation factor stiffness and low sorting efficiency of a Pippenger algorithm in zero-knowledge proof. A multi-resolution point multiplication table is dynamically generated through a hierarchical displacement pre-calculation strategy, and global memory access delay is remarkably reduced by combining memory layout optimization and asynchronous flow task scheduling of window inner barrel continuous storage. A dynamic resource adaptation mechanism is designed, thread block topology, grid division and sorting algorithm selection are adjusted in real time based on GPU hardware features, and load balancing and cache utilization rate maximization are achieved. Iterative reduction and double temporary bucket strategies are introduced, and data scale is compressed and boundary processing is optimized through multi-round reduction. In the large-scale MSM operation in the block chain and privacy computing field, the computing efficiency and the resource utilization rate can be remarkably improved, and the method is suitable for an elliptic curve cryptography acceleration task.
Owner:BEIHANG UNIV

Tensor transpose processor

The present invention relates to a processor designed to optimize memory bandwidth utilization for tensor transpositions in machine learning. An example processor includes an input tensor shift buffer, a staging buffer, and an output tensor shift buffer. The input tensor shift buffer reads an input tensor from input memory and performs multiple cycles of input tensor shifting. The shifted tensor data is then written into the staging buffer. The output tensor shift buffer reads the shifted tensor data from the staging buffer and performs multiple cycles of output tensor shifting. Finally, the result is written to the output memory. This configuration facilitates efficient handling and transformation of tensor data, optimizing the computational processes required in machine learning tasks.
Owner:MOFFETT TECH CO LTD

Method and system for processing operation data of brine discharging pipeline of salt cavern gas storage

The invention relates to the technical field of data processing, and discloses a salt cavern gas storage brine discharging pipeline operation data processing method and system, and the method comprises the steps: periodically receiving a link detection data packet which is transmitted by a collection device and is provided with a transmission timestamp, and recording the arrival time of the link detection data packet; according to the sending timestamp and the arrival time, transmission characteristics of the link detection data packet are analyzed, and the transmission characteristics comprise transmission delay, jitter dispersion and out-of-order conditions; adaptively predicting a time offset pattern based on the transmission characteristics; generating a real-time calibration parameter according to the time migration mode; and receiving an actual data packet, and adjusting the timestamp of the actual data packet by using the real-time calibration parameter. According to the method, the limitation of traditional fixed compensation parameters is overcome, the time alignment precision of pressure and flow data flow is remarkably improved, leakage misinformation caused by time deviation is avoided, and the reliability and the operation and maintenance efficiency of a salt cavern gas storage brine discharging pipeline operation monitoring system are improved.
Owner:WUHAN CENT CHINA GEOLOGICAL SURVEY CENT SOUTH CHINA INNOVATION CENT FOR GEOSCIENCES

Processing system and method related to encryption

A processing system and method related to encryption are provided. The processing system includes multiple computing circuits. A memory stores data. The computing circuit performs a number theoretic transform (NTT) calculation on a polynomial. A matrix multiplication calculation is performed on the polynomial through the NTT calculation. Therefore, the computing efficiency of encryption or decryption can be improved.
Owner:WISTRON CORP

Data processing method and device, electronic equipment and nonvolatile storage medium

The invention discloses a data processing method and device, electronic equipment and a nonvolatile storage medium. The method comprises the following steps: dividing a data matrix to be calculated into a plurality of data blocks with target sizes; data corresponding to the data blocks are allocated to a plurality of stream processors for data reading operation, the data reading operation is used for loading the data from a global memory of the edge computing device to a shared memory, and the data reading and writing efficiency of the shared memory is higher than that of the global memory; and loading the data in the memory block with the preset vector length in the shared memory into a register by taking the preset vector length as a unit, and operating the data loaded into the register by adopting a stream processor to obtain an operation result. The technical problem that the computing efficiency is not high due to the fact that the related technology is limited by the particularity of the memory access and computing rule of the correction operator and the hardware performance of the edge computing equipment cannot be fully exerted is solved.
Owner:北京大学长沙计算与数字经济研究院

Calculation apparatus, calculation method, calculation system, chip, device, and medium

The invention provides a computing device, system and method, a chip, equipment and a medium. The computing device comprises a reading circuit, a computing circuit, a storage circuit, a control circuit and a plurality of first registers. The at least one first register is configured with task configuration information of the tasks, and the control circuit reads register information corresponding to the tasks one by one according to a preset mode; enabling a first register identified by a currently read register information to output first configuration information to a reading circuit, output second configuration information to a computing circuit, output third configuration information to a storage circuit, and controlling the reading circuit to perform time-sharing data reading according to the first configuration information of each task, and controlling the calculation circuit to perform data operation in a time-sharing manner according to the second configuration information of each task, and controlling the storage circuit to perform data storage in a time-sharing manner according to the third configuration information of each task, so that the resource utilization rate and the calculation efficiency of the calculation device can be improved.
Owner:BEIJING HORIZON INFORMATION TECH CO LTD

Method for permuting dimensions of a multi-dimensional tensor

A method performed by a processor for permuting dimensions of a multi-dimensional tensor is described. The multi-dimensional tensor contains an array of tensor values in three or more dimensions that are stored in a first storage unit. The array of tensor values is transferred from the first storage unit to a second storage unit by reading tensor values from the first storage that are arrayed along a first dimension of the multi-dimensional tensor and writing the corresponding tensor values to the second storage in locations corresponding to a second dimension of the multi-dimensional tensor. The dimensions of the multi-dimensional tensor may be further permuted by a programmable engine within the processor.
Owner:ARM LTD

Floating point data parallel computing method and device of vector processor and vector processor

The invention provides a floating point data parallel computing method and device based on a vector processor and the vector processor, and the method comprises the steps: carrying out vectorization and parallel absolute value calculation on a plurality of pieces of input floating point data, and dividing each piece of data into different computing intervals according to the calculated absolute value; for data located in different calculation intervals, calculating branches in different intervals are combined, branch differences are uniformly represented through symbol transformation variables, and fitting intermediate variables corresponding to all the data are obtained through uniform vector operation instruction parallel calculation; and performing parallel calculation on the basis of the intermediate variable for fitting to obtain a preliminary result corresponding to each data, and correcting the preliminary result according to a symbol of original input data to output a final calculation result in a vector form. According to the method, the problems of low hardware resource utilization rate and poor calculation efficiency of high-precision floating point function calculation realized by adopting a scalar serial processing mode in the prior art are solved.
Owner:SHANGHAI SMARTLOGIC TECHNOLOGY LTD

Accelerating table lookups using a system-on-chip, separate lookup table accelerator

To provide a VPU and associated components that are optimized to improve VPU performance and throughput.SOLUTION: A VPU includes a min / max collector, an automatic store predication function, a SIMD data path configuration allowing inter-lane sharing, a transposed load / store including a stride parameter function, a load including a permute and zero insertion function, a hardware, logic device, and memory layout function allowing two point and two by two point lookups, and per memory bank load caching ability. Decoupled accelerators are used to offload VPU processing tasks, and a hardware sequencer is included in a DMA system. The DMA and VPU execute a VPU configuration mode allowing the VPU and DMA to operate without a processing controller for executing dynamic region-based data movement operations.SELECTED DRAWING: Figure 9A
Owner:NVIDIA CORP

Continuous bit parallel counting device and method for preamble and mantissa of processor

ActiveCN121050687AHandling data according to predetermined rulesComputation using non-denominational number representationComputer architectureHemt circuits
The invention relates to the field of integrated circuit design, and discloses a processor preamble and mantissa continuous bit parallel counting device and method. The device comprises a 32-bit bit sequence flipping module, a 32-bit bit negation module, eight 4-bit leading zero detection modules, four 8-bit leading zero processing modules, two 16-bit leading zero processing modules, a 32-bit leading zero processing module and a 32-bit result output module. According to the invention, parallel counting can be carried out on leading zeros, leading 1, mantissa zeros or mantissa 1 of various data with different precisions in a single processor clock period, so that the required time sequence path length is greatly shortened, the higher processor clock frequency is realized, the multiplexing of a hardware circuit is realized, and the cost is reduced. The circuit greatly reduces the use of transistor resources, reduces the area of a chip, reduces the cost and power consumption of the chip, improves the operation efficiency, universality and adaptability of the processor, and is suitable for being applied to application scenes such as a high-performance digital signal processor and a CPU (Central Processing Unit).
Owner:青岛本原微电子有限公司

Data processing method and device in assembly line, medium, equipment and product

The invention discloses a data processing method and device in an assembly line, a medium, equipment and a product, and the method comprises the steps: obtaining a barrier identifier before the execution of a to-be-processed data stage in the current assembly line according to an assembly line identifier and a stage identifier of the current assembly line; comparing the completion signal count associated with the barrier identifier of the current assembly line with the expected signal number, and triggering the current assembly line to execute a to-be-processed data stage under the condition that the completion signal count is equal to the expected signal number; after the current assembly line finishes executing the to-be-processed data stage, obtaining a barrier identifier before executing the to-be-processed data stage in the next assembly line; and sending a completion signal to the synchronization barrier indicated by the barrier identifier of the next assembly line through the current assembly line, and updating the stage identifier of the current assembly line so as to enter a next data processing stage. According to the method, the execution time sequence of the assembly line can be accurately controlled, and more efficient and more accurate scheduling and utilization of hardware resources are realized.
Owner:SHANGHAI BIREN TECH CO LTD

Performance management in data orchestrated environments

This disclosure provides methods, devices, and systems for data management. The present implementations more specifically relate to a data orchestration system that can dynamically reconfigure a data processing pipeline based on telemetry received from various steps or data operations in the pipeline. For example, the telemetry may indicate a success, failure, time of entry, time of exit, or total duration of a given step or data flow in the processing pipeline. In some aspects, the data orchestration system may dynamically invoke new data flows based on the received telemetry. In some implementations, the new data flows may allocate additional memory and / or processing resources for the data processing pipeline. In some other implementations, the new data flows may deallocate memory and / or processing resources for the data processing pipeline. Still further, in some implementations, the new data flows may trigger an alert to a user or manager of the data processing pipeline.
Owner:VIEW SYSTEMS INC

Transposing at-speed in a vector-matrix accelerator

A system including one or more processors configured to receive a transpose instruction indicating to transpose a source matrix to a result matrix, provide data elements of the source matrix to input switching circuits, reorder the data elements using the input switching circuits, provide the data elements from the input switching circuits to one or more lanes of a datapath, provide the data elements from the datapath to output switching circuits, undo the reordering of the data elements using the output switching circuits, and provide the data elements from the output switching circuits to a result matrix. Each respective lane of the datapath receiving data elements receives multiple data elements directed to different respective non-overlapping portions of the lane.
Owner:GOOGLE LLC

Masked Shift-Add Operation

The computer-implemented method includes receiving, by a processing unit, an instruction to perform a masked shift-and-add operation with a set of operands. A logical AND operation is performed on a first pair of operands in the set of operands to obtain a first intermediate result. The first intermediate result is shifted by a first shift amount based on a first operand of the first pair of operands. A logical AND operation is performed on a second pair of operands in the set of operands to obtain a second intermediate result. The second intermediate result is shifted by a second shift amount based on the first operand of the second pair of operands. The shifted first intermediate result is added to the shifted second intermediate result. The method further includes outputting an output of the addition as a result of the masked shift-and-add operation.
Owner:INTERNATIONAL BUSINESS MACHINE CORPORATION

Streaming engine with separately selectable element and group duplication

A streaming engine employed in a digital data processor specifies a fixed read only data stream defined by plural nested loops. An address generator produces address of data elements. A steam head register stores data elements next to be supplied to functional units for use as operands. An element duplication unit optionally duplicates data element an instruction specified number of times. A vector masking unit limits data elements received from the element duplication unit to least significant bits within an instruction specified vector length. If the vector length is less than a stream head register size, the vector masking unit stores all 0's in excess lanes of the stream head register (group duplication disabled) or stores duplicate copies of the least significant bits in excess lanes of the stream head register.
Owner:TEXAS INSTRUMENTS INC

A method for denoising distributed optical sensor signals based on curvelet transform

The present invention discloses a method for denoising distributed optical sensing signals based on curvelet transform, comprising the following steps: continuously collecting original Rayleigh backscattered signals generated by multiple optical pulses; performing orthogonal demodulation on the original signals to obtain demodulated data; stacking the demodulated data in the order of the optical pulse time to form a two-dimensional matrix; performing fast discrete curvelet transform and scale decomposition on the two-dimensional matrix to obtain a curvelet coefficient matrix; performing threshold processing on the submatrix blocks corresponding to each direction of each scale to obtain a processed curvelet coefficient matrix; performing fast discrete inverse curvelet transform on the processed curvelet coefficient matrix to obtain a transformed matrix; performing moving difference on the transformed matrix to determine the disturbance position on the optical fiber using the amplitude difference between the backscattered signals generated by optical pulses at different times. The present invention has the beneficial effects of being simple and direct, having strong adaptability, and being able to perform fast and effective noise reduction.
Owner:NANCHANG NORMAL UNIV OF APPLIED TECH +3

Transposing matrices based on a multi-level crossbar

Embodiments of the present disclosure include systems and methods for transposing matrices based on a multi-level crossbar. A system may include a memory configured to store a matrix comprising a plurality of elements arranged in a set of rows and a set of columns. A system may include an input buffer configured to retrieve a subset of a plurality of elements from the memory. Each element in the subset of the plurality of elements is retrieved from a different column in the matrix. A system may include a multi-level crossbar configured to perform a transpose operation on the subset of the plurality of elements. A system may include an output buffer configured to receive the transposed subset of the plurality of elements and store, in the memory, each element in the transposed subset of the plurality of elements in a different column in the matrix.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Processor operand management using fused buffers

Techniques relating to operand management using a fused buffer are disclosed. A processor includes an operand management circuit, wherein the operand management circuit includes a fusion buffer; and an execution circuit. In one embodiment, operand management circuitry is configured to: detect a first store instruction operation executable to store an operand value usable by one or more consumer instruction operations; and storing the first store instruction operation in the fusion buffer. In response to detecting a drop condition associated with the first store instruction operation, the operand management circuitry is configured to remove the first store instruction operation from the fusion buffer without forwarding the first store instruction operation for execution. The operand management circuitry is configured to forward the first stored instruction operation for execution by the execution circuitry in response to detecting a buffer emptied condition and not detecting a discard condition.
Owner:APPLE INC

Processing non-power-of-two work unit in neural processor circuit

A neural processor includes one or more neural engine circuits for performing convolution operations on input data corresponding to one or more tasks to generate output data. The neural engine circuits process the input data having a power-of-two (P2) shape. The neural processor circuit also includes a data processor circuit. The data processor circuit fetches source data having a non-power-of-two (NP2) shape. The source data may correspond to data of a machine learning model. The data processor circuit also reshapes the source data to generate reshaped source data with the P2 shape. The data processor circuit further sends the reshaped source data to the one or more neural engine circuits as the input data for performing convolution operations. In some cases, the data processor circuit may also perform padding on the source data before the source data is reshaped to the P2 shape.
Owner:APPLE INC

An iris image matching method, device, equipment and medium

The application relates to the field of image processing, in particular to an iris image matching method and device, equipment and medium, which are used to solve the problem of large calculation error in the iris image matching process. The method determines a plurality of target Mahalanobis distances based on each feature point of a to-be-detected iris image. It should be noted that each target Mahalanobis distance is determined based on the gradient parameters corresponding to any two target feature points of the to-be-detected iris image. Based on the sum of the plurality of target Mahalanobis distances and the sum of a plurality of standard Mahalanobis distances, the matching degree between the to-be-detected iris image and a pre-stored standard iris image is determined, wherein the standard Mahalanobis distance is determined based on the standard iris image. The above-mentioned iris image determination method based on Mahalanobis distance considers the internal relationship between the feature points, improves the matching accuracy between the to-be-detected iris image and the pre-stored standard iris image, and improves the reliability of the safe box in use.
Owner:CHINA CONSTRUCTION BANK +1

Systems, methods, and apparatuses for tile transpose

Embodiments detailed herein relate to matrix operations. In particular, support for a matrix transpose instruction is detailed. In some embodiments, decode circuitry to decode an instruction having fields for an opcode, a source matrix operand identifier, and a destination matrix operand identifier; and execution circuitry to execute the decoded instruction to transpose each row of elements of the identified source matrix operand into a corresponding column of the identified destination matrix operand are detailed.
Owner:INTEL CORP

Fault-tolerant cubature Kalman filtering-based multi-machine power system elastic state estimation method

The invention discloses a fault-tolerant cubature Kalman filtering-based multi-machine power system elastic state estimation method. The method comprises the following steps of establishing a multi-machine power system dynamic state estimation model considering sensor abnormality; initializing parameter values of the state estimation method; calculating a state prediction value and a prediction error covariance matrix at the moment k; whether a fault occurs or not is judged by calculating filtering innovation of a state estimation method; and introducing a fault-tolerant factor to enforcedly execute correction filtering innovation, and updating a state estimation value and a state error covariance matrix. According to the method provided by the invention, the dynamic state estimation model of the multi-machine power system under the abnormal condition of the sensor is established, the method capable of improving the fault-tolerant capability and the stability of the system under the abnormal condition is provided, stable state estimation is kept, and the stability and the reliability of the system can be improved.
Owner:ZHENGZHOU UNIV

Transposed convolution using systolic arrays

In one example, a neural network accelerator can execute a set of instructions to: load a first weight data element from a memory into a systolic array, the first weight data element having a first coordinate; extract, from the instructions, information indicating a first subset of input data elements to be obtained from the memory, the first subset based on a stride of a transpose convolution operation and a second coordinate of the first weight data element in a rotated array of weight data elements; obtain the first subset of input data elements from the memory based on the information; load the first subset of input data elements into the systolic array; and control the systolic array to perform a first computation based on the first weight data element and the first subset of input data elements to produce an output data element of an array of output data elements.
Owner:AMAZON TECH INC

Shuffle accelerator for graphics processing unit

Shuffle accelerators for shuffling data between a plurality of instances executing a shader on a shader core of a graphics processing unit. The shuffle accelerators include routing logic, slave logic and master logic. The routing logic comprises a plurality of data input ports, a plurality of data output ports, and hardware to selectively connect one or more of the plurality of data input ports to one or more of the plurality of data output ports. The slave logic is configured to selectively provide data from a first set of instances to one or more of the plurality of data input ports and receive data from one or more of the plurality of data output ports for a second set of instances. The master logic is configured to, in response to receiving a shuffle instruction that identifies a shuffle of data between the plurality of instances, cause the routing logic and the slave logic to perform the identified shuffle of data in a plurality of phases, wherein in each phase of the plurality of phases a subset of the instances of the plurality of instances receive data from a subset of the instances of the plurality of instances.
Owner:IMAGINATION TECH LTD