Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

367results about "Register arrangements" patented technology

Heterogeneous processor-oriented reciprocal calculation instruction sequence generation method

The invention discloses a reciprocal calculation instruction sequence generation method oriented to a heterogeneous processor, and belongs to the field of compilation optimization and code generation. Aiming at the problems of instruction redundancy, weak precision control, poor hardware adaptation and high manual dependence of an existing method in a heterogeneous environment, characteristics of a reciprocal instruction and an operand are accurately identified by linearly scanning heterogeneous object codes (including vectorization, scalar and complex instruction sequences); in combination with hardware characteristics of RISC / SIMD / VLIW / DSP and the like, a multi-round iteration precision improvement and temporary register optimization allocation strategy is adopted, differential generation logic is formulated, and a high-precision low-redundancy instruction sequence is generated. The method comprises linear code scanning classification, reciprocal instruction and operand identification, cross-architecture generation logic rule formulation, instruction sequence generation and legality verification. Full-process automation is achieved, manual intervention is reduced, the execution efficiency and precision of reciprocal calculation of the heterogeneous processor are improved, and the method is suitable for embedded systems, high-performance calculation and other scenes.
Owner:HUNAN UNIV OF SCI & TECH

Padding and suppressing rows and columns of data

A method is described herein. The method generally includes receiving stream parameters that defines an array, wherein the stream parameters include a first null element count and a second null element count. The method generally includes forming a stream of vectors for the multidimensional array responsive to the stream parameters. The stream of vectors generally includes a vector of null elements at a beginning of the stream of vectors based on the first null element count. The stream of vectors generally includes a null element at a beginning of each vector of the stream of vectors based on the second null element count. The stream of vectors generally includes a set of data distributed across a subset of the stream of vectors. The method generally includes providing the stream of vectors.
Owner:TEXAS INSTRUMENTS INC

Artificial intelligence chip, parallel method for vector and scalar execution pipeline, computing device, medium and program product

The invention relates to an artificial intelligence chip, a method for parallel vector and scalar execution assembly lines, a computing device, a medium and a program product. The artificial intelligence chip comprises an execution unit, the execution unit is configured with a vector execution assembly line and a scalar execution assembly line, and the scalar execution assembly line at least comprises a scalar instruction decoding unit which is configured to at least obtain an operand type, address information and scalar operation control information of a scalar instruction; a scalar instruction operand acquisition unit configured to acquire an operand source of a scalar instruction; and a scalar instruction operation unit configured to execute scalar calculation at least based on an operand type, an operand source and scalar operation control information of the scalar instruction, and write a calculation result to the scalar register group included in the execution unit. According to the method, the utilization rate and the actual computing power of hardware resources of the execution unit of the artificial intelligence chip can be remarkably improved.
Owner:SHANGHAI BIREN TECH CO LTD

Automatic illegal character cleaning system

ActiveCN121807413AResource allocationRegister arrangementsTopology mappingSpeculative execution
The invention relates to the technical field of computer data processing and network security, in particular to an illegal character automatic cleaning system, which comprises a rule compiling module for monitoring rule change, performing semantic fusion and topological mapping on a rule set, constructing a deterministic finite automaton and mapping the deterministic finite automaton into a state transition table; the state switching module is used for constructing a double-buffer context and operating lock-free switching through an atomic pointer to realize hot updating; the speculation execution module is used for carrying out vectorization pre-scanning based on a state transition table by utilizing single-instruction multi-data stream parallel loading, identifying a walk path, falling into a safe state for releasing and falling into a trap state for triggering external verification; the self-adaptive feedback module is used for counting trap state triggering frequency, generating a rule allergy report and dynamically adjusting the size of a read fragment; according to the method, the contradiction between rule flexibility and execution efficiency is solved, and high-performance cleaning based on speculative execution is realized.
Owner:北京啄木鸟云健康科技有限公司

RISC-V data processing system and data processing method

The embodiment of the invention discloses a data processing system and a data processing method based on RISC-V. The data processing system comprises an instruction coding space and a reconfigurable execution unit, the instruction coding space comprises an operation code area, a reserved area and an expansion area; the instruction coding space is used for defining a basic operation to be executed by an instruction in the operation code area, defining an extended operation to be executed by the instruction in the reserved area and defining the length of the instruction in the extended area so as to configure different types of instructions; the reconfigurable execution unit is used for constructing a data path matrix of the instruction based on the configuration information of the instruction; transmitting data required when the instruction is executed based on the data path matrix; and executing an operation corresponding to the instruction based on the data of the instruction.
Owner:CHENGDU KAIYUAN COMPUTING ECOLOGICAL TECHNOLOGY CO LTD

SoC system control method and device, electronic equipment and readable storage medium

The invention provides an SoC system control method and device, electronic equipment and a readable storage medium, and the method comprises the steps: controlling a CPU large core to generate a control instruction, and writing the control instruction into a register; the control register analyzes the control instruction, determines at least one target function module associated with the control instruction, generates a corresponding enable signal and sends the enable signal to a control circuit corresponding to each target function module; and aiming at the control circuit of each target function module, controlling the control circuit to receive and analyze the corresponding enable signal, generating control information and sending the control information to the corresponding target function module, so that each target function module is configured according to the control information and executes the control instruction. Therefore, the configuration time of the CPU for configuring the functional modules one by one through the on-chip internet can be reduced, and the control efficiency of the SoC system is further improved.
Owner:CIX TECH (SHANGHAI) CO LTD

Update of seamless container in vehicles system based on container

A method for updating a container in a container-based vehicle system includes: executing an application included in the container, the application configured to access at least one item included in the container to operate a service related to a vehicle, receiving, from a server, a container image of an update target based on determining an update of the container being needed, the container image including a set of items, storing the received container image of the update target in a system registry, generating a new container by combining a user data item that is generated in the container and the stored container image of the update target, and updating the container by using a container management table on the new container while executing the application. The application reflects the updated container based on items included in the updated container.
Owner:LG ELECTRONICS INC

Resource allocation method, resource allocation device, electronic equipment and medium

The invention provides a resource allocation method, a resource allocation device, electronic equipment and a medium, and the method comprises the steps: determining whether the number of available memory logic sub-blocks in a shared memory meets the demand of a work group task or not according to the demand number of shared memory resources; under the condition that the number of available memory logic sub-blocks in the shared memory meets the requirements of the working group tasks, the corresponding memory logic sub-blocks are allocated to the working group tasks, and for each thread bundle task in the working group tasks, the corresponding thread bundle task is allocated to the corresponding thread bundle task based on the required number of universal register resources of the corresponding thread bundle task. Determining whether the number of available register logic sub-blocks in one execution unit in the plurality of execution units meets the requirement of a corresponding thread bundle task; and under the condition that the number of the available register logic sub-blocks in one execution unit in the plurality of execution units meets the requirement of the corresponding thread beam task, allocating the corresponding register logic sub-blocks in the execution unit to the corresponding thread beam task. The problem of storage space fragmentation is solved.
Owner:SUZHOU YIZHU INTELLIGENT TECH CO LTD

Supporting 8-bit floating point format operands in a computing architecture

An apparatus to facilitate supporting 8-bit floating point format operands in a computing architecture is disclosed. The apparatus includes a processor comprising: a decoder to decode an instruction fetched for execution into a decoded instruction, wherein the decoded instruction is a matrix instruction that operates on 8-bit floating point operands to cause the processor to perform a parallel dot product operation; a controller to schedule the decoded instruction and provide input data for the 8-bit floating point operands in accordance with an 8-bit floating data format indicated by the decoded instruction; and systolic dot product circuitry to execute the decoded instruction using systolic layers, each systolic layer comprises one or more sets of interconnected multipliers, shifters, and adder, each set of multipliers, shifters, and adders to generate a dot product of the 8-bit floating point operands.
Owner:INTEL CORP

Oral care device

An oral care device includes a head for caring for a user's oral cavity, the oral cavity including a plurality of oral regions, and an inertial measurement unit (IMU) operable to output signals that depend on the position and / or movement of the head. The oral care device includes a controller configured to receive from the IMU signals indicative of the position and / or movement of the head relative to the oral cavity, and to process the received signals using a trained non-linear classification algorithm to obtain classification data. The classification algorithm is trained to identify the oral region in which the head of the oral care device is located from the plurality of oral regions. The controller is configured to use the obtained classification data to control the oral care device to perform an action.
Owner:DYSON TECH LTD

Register reading method, device and equipment for cross-chip communication system

The invention provides a register reading method, device and equipment of a cross-chip communication system. The method comprises the steps that a reading instruction is sent to first transceiving transmitter equipment through a first peripheral bus; controlling the first transceiving transmitter device to package the read instruction, and sending the packaged read instruction to the second transceiving transmitter device, so that the second transceiving transmitter device analyzes the packaged read instruction and sends the read instruction to the second peripheral bus, so that the second peripheral bus sends the read instruction to the first transceiving transmitter device, and the second transceiving transmitter device sends the read instruction to the second peripheral bus. Reading the mounted register according to the reading instruction, and returning a plurality of read data to the second transceiving transmitter equipment, so that the second transceiving transmitter equipment packages the plurality of read data to obtain a read reply data packet; after controlling the first transceiving transmitter equipment to receive the read reply data packet, analyzing to obtain analyzed read data; and the analyzed read data is received through the first peripheral bus, so that cross-chip register reading is realized.
Owner:XIAMEN UNISOC TECH CO LTD

A circuit and method for long sequence data sorting

The application discloses a circuit and method for long sequence data sorting. The circuit is used for sorting N data in K base, and the maximum element in the data contains m bits in K base. The circuit comprises a radix counting unit, a first address generating unit, a data distribution unit and two sorting buffers. The two sorting buffers are used as a source buffer and a target buffer respectively, and can read data in a given address or write a data into a specified address. The radix counting unit, the first address generating unit and the data distribution unit are connected in sequence. The two sorting buffers are connected with the data distribution unit and the radix counting unit respectively. The data sorting circuit of the application has simple structure, can be flexibly adjusted according to specific requirements, and the data sorting method has linear order time complexity and short sorting time.
Owner:NANJING UNIV

A method, apparatus, device, and medium for processing multi-dimensional data

This invention relates to the field of data processing technology, and in particular to a method, apparatus, device, and medium for processing multi-dimensional data. The method instantiates a template based on input tensor parameters to obtain an offset calculation instance. The offset calculation instance is initialized to determine the step size and shape of each input array in each output dimension, and these are recorded in a designated storage space. The offset calculation instance is then started to perform index calculations on the step size and shape of each input array recorded in the storage space in each output dimension to obtain the offset corresponding to the output array. The data corresponding to the offset is then processed according to a set calculation rule to obtain the output result. By pre-calculating the step size and shape of each input array in each output dimension, unnecessary memory accesses can be reduced. Constructing offset calculation instances simplifies the index calculation process, enabling efficient and accurate data storage, retrieval, and calculation.
Owner:INSPUR SUZHOU INTELLIGENT TECH CO LTD

Data structure processing

An apparatus includes an instruction decoder and a processing circuitry system. In response to a data structure processing instruction specifying at least one input data structure identifier and an output data structure identifier, the instruction decoder controls the processing circuitry system to perform processing operations on at least one input data structure to produce an output data structure. Each input / output data structure includes a data arrangement corresponding to multiple memory addresses. The apparatus includes two or more sets of one or more data structure metadata buffers, each set associated with a corresponding data structure identifier and designated to store metadata indicating the memory address of the data structure identified by the corresponding data structure identifier.
Owner:ARM LTD

Compiler automatic debugging method and system for VLIW and SIMD architecture

The application discloses a kind of compiler automatic debugging method and system for VLIW and SIMD architecture, the method of the application includes the semantic correctness verification for the program to be checked to judge whether the program to be checked exists semantic error relative to source program, if semantic correctness verification finds that there is semantic error, then determine that debugging is not passed, otherwise the physical register verification for the program to be checked to judge whether the program to be checked exists physical register allocation error, if there is physical register allocation error, then determine that debugging is not passed, otherwise determine that debugging is passed;When determining that debugging is not passed, then generate feedback verification report.The application can automatically check the semantic correctness of program in the process of compiling, and provide accurate verification report (including error occurrence position and type etc.) for developer, improve the efficiency of compiler development, reduce the burden of developer.
Owner:NAT UNIV OF DEFENSE TECH

Method and apparatus for controlling input / output operation of vector processor in mixed precision environment

A The present invention relates to a technique for controlling input / output operation of a vector processor, which is designed to optimize vector operation in a mixed precision environment, and to a technique for maximizing data processing performance while minimizing waste of operation resources by improving the data conversion process between the memory and the vector processor.
Owner:SEOUL NATIONAL UNIVERSITY R&DB FOUNDATION

Methods and apparatus to generate graphics processing unit long instruction traces

Methods, apparatus, systems and articles of manufacture are disclosed to generate a graphics processing unit (GPU) long instruction trace (GLIT). An example apparatus includes at least one memory, and at least one processor to execute instructions to at least identify a first routine based on an identifier of a second routine executed by the GPU, the first routine based on an emulation of the second routine, execute the first routine to determine a first value of a GPU state of the GPU, the first routine having (i) a first argument associated with the second routine and (ii) a second argument corresponding to a second value of the GPU state prior to executing the first routine, and control a workload of the GPU based on the first value of the GPU state.
Owner:INTEL CORP

Memory device, operating method thereof, and in-memory processing device

The invention discloses a memory device, an operating method thereof, and an in-memory processing device. The memory device includes an in-memory processing (PIM) block configured to perform an operation between a weight value and an input value, the weight value being represented by a weight scaling factor and a weight element, the input value being represented by an input scaling factor and an input element, where the PIM block includes: a first scaling register file storing the input scaling factor; a second scaling register file storing a weight scaling factor; a scalar register file (SRF) storing input elements; a plurality of arithmetic logic units (ALU) configured to perform a first operation between the input scaling factor and the weight scaling factor and a second operation between the input element and the weight element in parallel in response to an operation command received from a host; and an accumulator configured to accumulate and store operation results of the first operation and the second operation.
Owner:SAMSUNG ELECTRONICS CO LTD

Programmable logic device dynamic register configuration device and control method thereof

The invention provides a programmable logic device dynamic register configuration device and a control method thereof. The configuration device for the dynamic register of the programmable logic device comprises an external controller used for providing a configuration instruction and configuration data; the main controller is used for receiving and analyzing the configuration data, generating a gating signal and recombining a data packet; at least one multiplexing module, each multiplexing module being configured to include a register block; and one or more stages of sub-controllers are arranged between the main controller and the multiplexing module and are used for further analyzing the data packet and controlling the target register group to complete read-write operation. In addition, the invention further provides a control method corresponding to the configuration device. Some embodiments of the invention aim to solve the problems of high configuration switching time consumption and uncertain configuration switching time in the prior art on the premise of not increasing too much complexity and area overhead.
Owner:SHANGHAI XINLU TECH CO LTD

Performance counting device and chip

The invention relates to a performance counting device and a chip, and the performance counting device comprises a performance counting module which is used for receiving at least one instruction memory address from a program counter associated with the performance counting module, and according to the at least one instruction memory address and a pre-configured instruction memory address interval, counting the at least one instruction memory address; obtaining a statistical value of a performance signal of at least one instruction memory address in the instruction memory address interval; the storage module is used for storing the statistical value of the performance signal of the at least one instruction memory address of the performance statistical module; the write-in control module is used for writing the statistical value, obtained by the performance statistical module, of the performance signal of the at least one instruction memory address into the storage module; and the configuration module is used for configuring an instruction memory address interval of the performance statistics module. According to the method and the device, the configuration flexibility of performance statistics can be improved, performance statistics of a plurality of specified program segments can be realized at the same time on the basis, and parallel processing of the performance statistics and accurate positioning of the program segments are further realized.
Owner:SHANGHAI BIREN TECH CO LTD

Hexadecimal floating-point multiply-add instruction

An instruction is executed to perform an operation selected from a plurality of operations configured for the instruction. The executing includes determining a value of a selected operand of the instruction. Determining the value includes reading the selected operand of the instruction from the selected operand location to obtain a value of the selected operand based on control of the instruction, the control including a first value, and using a default value as the value of the selected operand based on the control including a second value. The value and another selected operand of the instruction are multiplied to obtain a product. An arithmetic operation is performed using the product and the selected operand of the instruction to obtain an intermediate result. A result is obtained from the intermediate result and placed in the selected location.
Owner:INTERNATIONAL BUSINESS MACHINE CORPORATION

Tensor data transformation method, tensor processing unit, processor, system on chip, and computing device

Provided in the embodiments of the present disclosure are a tensor data transformation method, a tensor processing unit, a processor, a system on a chip, and a computing device. The tensor processing unit comprises: a register array unit, which comprises at least a source address register, a destination address register, and a source tensor parameter register; an instruction decoding unit, which parses an acquired transformation operation instruction, so as to obtain a transformation operation type indicated by an operation type field and obtain respective register serial numbers of the source address register, the destination address register, and the source tensor parameter register; and an instruction execution unit, which accesses the source address register, the destination address register, and the source tensor parameter register on the basis of their respective register serial numbers, reads a source memory address field, a destination memory address field, and a source tensor parameter, reads a source tensor on the basis of the source memory address field, performs, on the source tensor, a transformation operation indicated by the transformation operation type, and writes, on the basis of the source tensor parameter, a destination tensor obtained after the transformation operation into the destination memory address field.
Owner:ALIBABA DAMO (HANGZHOU) TECH CO LTD

Method and device for managing register

The invention relates to the field of computers, in particular to a register management method and device. In the prior art, an assembly line register is reused for temporarily storing data, then a data transmission mechanism of a processor is used for transmitting the temporarily stored data to subsequent instructions on an assembly line for use, and due to the fact that the number of the instructions entering the assembly line is limited, the method is only suitable for the instructions with the short life cycle, and the application scene is limited. The first indication information indicates the first life cycle of the register associated with the first instruction set, the first life cycle is shorter than the second life cycle determined by hardware, the processor can be prompted to release the register in advance, and therefore the use efficiency of register resources is improved. The first indication information indicates the life cycle of the non-pipeline register, so that the technology provided by the invention not only can optimize the register resource used by the instruction with the shorter life cycle, but also can optimize the register resource used by the instruction with the longer life cycle.
Owner:HUAWEI TECH CO LTD

Vector Scatter and Gather with Single Memory Access

PendingUS20260119176A1Register arrangementsMemory hierarchyComputer architecture
Disclosed embodiments provide techniques for improved performance in processing vector instructions. A processor core is accessed. The processor core is coupled to a memory hierarchy, and the processor core includes one or more vector execution units (VUs), and one or more load store units (LSUs). The processor core includes a vector register file (VRF). The VRF includes multiple vector registers, and each vector register includes multiple vector elements. Vector elements that have a source or destination in contiguous memory are identified. Load store units (LSUs) take advantage of the contiguous memory condition by executing a vector load or vector store operation as a single memory access, requiring a reduced number of clock cycles. The single memory access satisfies each memory operation for each vector element within the vector register file.
Owner:AKEANA INC

Scalable hardware event forwarding in an expandable system

PendingDE102025132662A1Register arrangementsTransmissionComputer hardwareScalable system
A test and measurement device is described, comprising one or more detection devices configured to receive an event from a corresponding event generator. The test and measurement device also includes a transmitter configured to receive the event from the one or more detection devices and transmit the event. The test and measurement device also includes one or more comparators configured to receive the event from the batch transmitter and forward the event to a corresponding event detector based on a comparison of an event identifier and a detector identifier.
Owner:KEITHLEY INSTRUMENTS LLC

Arithmetic processing unit and arithmetic processing method

To efficiently execute floating point arithmetic.SOLUTION: An arithmetic processing device 1 includes: an instruction storage section 161 that stores an arithmetic instruction; a data cache section 18 that caches an arithmetic result of the arithmetic instruction; a plurality of floating point registers 172 that is arranged on a side of the instruction storage section 161 and stores a register value for executing the arithmetic instruction transferred from the instruction storage section 161; and a plurality of floating point arithmetic units 171 that is arranged on a side of the data cache section 18 and performs floating point arithmetic on the basis of the arithmetic instruction. The number of cycles is one when the register value is transferred from the instruction storage section 161 to one or more floating point registers 172, among the plurality of floating point registers 172, arranged at positions closest in a distance to the instruction storage section 161.SELECTED DRAWING: Figure 8
Owner:FUJITSU LTD

Register configuration system, method, and electronic device

The present disclosure provides a register configuration system, method, and electronic device. The register configuration system includes a display processing unit, memory, and a hardware automatic control unit. The display processing unit includes multiple registers. The memory includes at least one pre-stored configuration instruction, wherein the configuration instruction is configured to indicate target registers that need to be configured in the display processing unit. The hardware automatic control unit is connected to a register configuration interface of the display processing unit. The hardware automatic control unit is configured to receive a startup command and, based on the startup command, retrieve and parse the configuration instructions from the memory. After parsing the configuration instructions, the hardware automatic control unit configures the target registers in the display processing unit. The startup command includes the storage address of the configuration instructions.
Owner:VERISILICON MICROELECTRONICS (SHANGHAI) CO LTD +1