A processor applied to an intelligent sensor

By employing mixed-length instruction parallel processing and acceleration unit optimization in the intelligent sensor processor, the shortcomings of existing processors in terms of real-time performance, power consumption, and hardware resource consumption are addressed, achieving low-power and high-efficiency intelligent sensor processing.

CN115904313BActive Publication Date: 2026-07-21NANJING INST OF INTELLIGENT TECH INST OF MICROELECTRONICS OF THE CHINESE ACAD OF
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
NANJING INST OF INTELLIGENT TECH INST OF MICROELECTRONICS OF THE CHINESE ACAD OF
Filing Date
2022-11-08
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

Existing smart sensor processors are inadequate in terms of real-time performance, versatility, power consumption, area, and cost, making it difficult to meet the needs of smart sensors.

Method used

A processor comprising a chip core (CORE) and pin pads (PADs) was designed. The core includes a data memory (DM), a program memory (PM), a register file (RF), an auxiliary register (AR), and a processing module. It employs mixed-length instructions and instruction-level parallel processing, combining SF-NEAS CORDIC arithmetic subunits and inverse square root arithmetic subunits based on successive approximation algorithms to optimize program control and accelerate computation.

Benefits of technology

It reduces the number of processor execution cycles and power consumption, improves computing speed and accuracy, meets the real-time and low-power requirements of smart sensors, and reduces hardware resource consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115904313B_ABST
    Figure CN115904313B_ABST
Patent Text Reader

Abstract

The application relates to a processor applied to an intelligent sensor. The chip kernel CORE in the processor comprises a DM, a PM, an RF, an AR and a processing module; the PM is used for storing application program codes and obtaining instructions; the DM is used for storing temporary operation numbers and operation results; the RF and the AR are used for storing operation numbers; the processing module is used for reading instructions and operation numbers from the PM, decoding and performing corresponding operations, and writing operation results back to the RF and the AR; a program control unit in the processing module is used for performing mixed-length instruction and instruction-level parallel processing, outputting control signals to a data path and an AGU, and decoding an immediate number; an acceleration unit in the processing module comprises an SF-NEAS CORDIC operation subunit and an inverse square root operation subunit based on a successive approximation algorithm; and the application has the characteristics of low power consumption and programmability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of smart sensors, and in particular to a processor for use in smart sensors. Background Technology

[0002] Smart sensors are fundamental to the Fourth Industrial Revolution. A smart sensor is a sensor with a microprocessor that combines information detection and processing capabilities; therefore, the processor is a crucial component. Sensor information detection and processing include data calibration and data fusion (attitude estimation). Consequently, as smart sensor algorithms continue to evolve and systems operate in real-time for extended periods, the demands on processors are increasing.

[0003] To address the aforementioned issues, there is an urgent need for a dedicated processor for smart sensors that features low power consumption and programmability, in order to meet the requirements of smart sensors in terms of real-time performance, versatility, power consumption, area, performance, and cost. Summary of the Invention

[0004] The purpose of this invention is to provide a processor for use in smart sensors, which has the characteristics of low power consumption and programmability, and can meet the requirements of smart sensors in terms of real-time performance, versatility, power consumption, area, performance and cost.

[0005] To achieve the above objectives, the present invention provides the following solution:

[0006] A processor for use in smart sensors includes: a chip core (CORE) and pins (PAD); the chip core (CORE) acquires external signals through the pins (PAD).

[0007] The chip core (CORE) includes: a data memory (DM), a program memory (PM), a register file (RF), auxiliary registers (AR), and a processing module.

[0008] The program memory (PM) is used to store application code and fetch instructions; the data memory (DM) stores temporary operands and calculation results; the register file (RF) and the auxiliary register (AR) are used to store operands; the processing module is used to read instructions and operands from the program memory (PM), decode them, execute the corresponding calculations, and write the calculation results back to the register file (RF) and the auxiliary register (AR);

[0009] The processing module includes: an arithmetic logic unit (ALU), a multiplier, a shifter, an acceleration unit, a control path, a data path, a data bus, an address generation unit (AGU), a program control unit, and peripheral devices;

[0010] The program control unit is used to perform mixed-length instruction and instruction-level parallel processing, output control signals to the data path and address generation unit AGU, and decode immediate values;

[0011] The acceleration unit includes an SF-NEAS CORDIC operation subunit and an inverse square root operation subunit based on a successive approximation algorithm; the SF-NEAS CORDIC operation subunit is used to configure different operation functions under different acceleration instructions.

[0012] Optionally, the program control unit employs a state transition mechanism.

[0013] Optionally, the SF-NEAS CORDIC operational subunit employs α i =2 -i Determine the iterative rotation angle; α is the iterative rotation angle, and i is the shift in hardware, in units of -bit.

[0014] Optionally, the acceleration commands include: SIN, COS, ARCSIN, ARCCOS, ARCTAN, and SQRT.

[0015] Optionally, the inverse square root operation subunit based on the successive approximation algorithm is used to calculate the inverse square root based on the input data.

[0016] According to specific embodiments provided by the present invention, the present invention discloses the following technical effects:

[0017] This invention provides a processor for smart sensors. The program control unit performs mixed-length instruction and instruction-level parallel processing, outputs control signals to the data path and address generation unit (AGU), and decodes immediate values, reducing the number of execution cycles of the assembler in the processor by 24.7% and reducing processor power consumption by 18%. The SF-NEAS CORDIC operation subunit consumes less hardware resources compared to CORDIC circuits with the same architecture and calculation error, consuming only 2.05uW at 1.5MHz. Simultaneously, using this operation subunit significantly reduces the number of clock cycles required for the processor to complete a single algorithm iteration, lowering the overall processor power consumption. The inverse square root operation subunit based on the successive approximation algorithm has lower hardware resource consumption compared to similar circuits under the same calculation error. Using the calculation result of this inverse square root operation subunit as an intermediate variable can improve the execution speed of vector normalization, division, and square root operations in the sensor algorithm. Furthermore, designing corresponding fixed-point normalization programs results in higher accuracy of the calculation results. Attached Figure Description

[0018] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0019] Figure 1 This is a schematic diagram of a processor structure for use in a smart sensor provided by the present invention;

[0020] Figure 2 This is a block diagram of the hardware structure of the program control unit;

[0021] Figure 3 A schematic diagram showing the different combinations of 16 / 32-bit instructions in the program memory (PM);

[0022] Figure 4 This is a block diagram of the hardware structure of the SF-NEAS CORDIC computing subunit;

[0023] Figure 5 This is a block diagram of the sub-unit structure for the inverse square root operation based on the successive approximation algorithm. Detailed Implementation

[0024] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0025] The purpose of this invention is to provide a processor for use in smart sensors, which has the characteristics of low power consumption and programmability, and can meet the requirements of smart sensors in terms of real-time performance, versatility, power consumption, area, performance and cost.

[0026] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0027] Figure 1 This is a schematic diagram of a processor structure for use in a smart sensor provided by the present invention, as shown below. Figure 1 As shown, the present invention provides a processor for use in smart sensors, comprising: a chip core (CORE) and pins (PAD); the chip core (CORE) acquires external signals through the pins (PAD).

[0028] The chip core (CORE) includes: data memory (DM), program memory (PM), register file (RF), auxiliary register (AR), and processing module.

[0029] The SRAM of the data memory DM is placed at one end of the chip core CORE, and the SRAM of the program memory PM is placed at the other end of the chip core CORE; the middle area of ​​the chip core CORE is used for the logic part of the circuit, where various units and wiring are placed.

[0030] The program memory (PM) is used to store application code and fetch instructions; the data memory (DM) stores temporary operands and calculation results; the register file (RF) and the auxiliary register (AR) are used to store operands; the processing module is used to read instructions and operands from the program memory (PM), decode them, perform corresponding calculations, and write the calculation results back to the register file (RF) and the auxiliary register (AR).

[0031] The processing module includes: an arithmetic logic unit (ALU), a multiplier, a shifter, an acceleration unit, a control path, a data path, a data bus, an address generation unit (AGU), a program control unit, and peripheral devices.

[0032] The program control unit is used to perform mixed-length instruction and instruction-level parallel processing, output control signals to the data path and address generation unit AGU, and decode immediate values.

[0033] The inputs to the program control unit include flags and instructions from the data path.

[0034] The program control unit employs a state transition mechanism.

[0035] When the program runs normally in sequence (i.e., without jumps), the new instruction address is obtained by adding the address offset step to the current instruction address. The instruction is then read from that address in the Program Manager Unit (PM) and decoded for execution. For conditional jumps, the value of the special register (CNR) determines whether a jump is necessary. Conditional jumps are all relative jumps, calculated by adding the offset IMM to the current instruction address. offset The address obtained later is used as the address of the new instruction. For the case of 16 / 32-bit mixed-length instructions in the instruction set, a corresponding state jump mechanism is designed in the program control unit to solve the problem of inconsistent instruction read and write rates for mixed-length instructions. Taking advantage of the characteristic that the assembly program compiled by the sensor algorithm consists largely of arithmetic and logical operation instructions and data transfer instructions, a program control unit is designed that can detect the aforementioned instruction-data conflicts and achieve instruction-level parallelism.

[0036] Program control unit such as Figure 2As shown, its workflow is roughly as follows: it determines the instruction combination for the current cycle based on the instruction combination pre-decoded in the previous cycle, the pre-decoding result, the current PC value, and whether a jump occurred in the previous cycle. In normal execution mode, the PC step size is 1 or 2 depending on the actual length of the instructions that can be issued. When a jump occurs, the PC is biased relative to the jump by an IMM. offset Function calls and returns jump to addresses in the absolute addresses IMM and LR, respectively. The instruction address is offset by 2 or 3 from the current PC value, and the first 11 bits are truncated as the PM address. However, if... Figure 3 In the case of instruction combinations 3 and 6, the instruction processing speed is slower than the instruction loading speed, and in this case, it is necessary to stop for one cycle to read the PM.

[0037] Figure 3 The instruction set contains combinations of 16-bit and 32-bit instructions at different locations within the instruction set (PM). Because the instruction lengths are inconsistent, and there are instances where two 16-bit instructions execute in parallel, instruction prefetching is necessary to prevent the loading speed from lagging behind the processor's processing speed. The processor has two 16-bit instruction buffers (IB) during the IF phase: IB0 and IB1. Specifically, as... Figure 3 There are 8 possible scenarios, and in each scenario, the current PC points to instruction 2. For combination 1, the previous state is that a 32-bit instruction has already been issued with an even PC value, and the instruction currently pointed to by the PC has just been fetched from the PM. Therefore, the two 16-bit instructions being tested are directly fetched from the PM. If the test result indicates that the instruction set is a 32-bit instruction or can be executed in parallel, the PC is incremented by 2. Simultaneously, IB0 and IB1 also cache the first and last 16 bits of the instruction read from the PM. The purpose of caching is to prevent situations where the tested instruction combination cannot be issued as a 32-bit group; in such cases, the instruction can be read from IB1 and the subsequent instructions can continue to be executed. This method enables parallel execution of mixed-length instructions, reducing the number of assembler cycles in the processor by 24.7% and reducing processor power consumption by 18%.

[0038] The acceleration unit includes an SF-NEAS CORDIC operation subunit and an inverse square root operation subunit based on a successive approximation algorithm; the SF-NEAS CORDIC operation subunit is used to configure different operation functions under different acceleration instructions.

[0039] The instruction set design of this invention mainly includes the design of processor data types, arithmetic and logical operation instructions, control flow instructions, data transfer instructions, and dedicated acceleration instructions. For example, the data types supported by the processor mainly include unsigned int, int, unsigned short, short, void, pointer type *, and array type []. By analyzing the source code of typical sensor data fusion (attitude estimation) algorithms and calibration algorithm program sets, the operation types, hotspot operators, and control flow statement types of the target application are determined, and a dedicated instruction set for the above algorithms is designed based on this.

[0040] Acceleration units, combined with instruction sets, can reduce program execution time and system power consumption. Like the ALU, acceleration units are invoked by instructions. Acceleration instructions first fetch instructions from the PM (Processing Unit) based on their address. After decoding, the instructions are executed in the execution unit. The focus of these instructions is the execution phase; after calling various computational units, the computation task is completed. Since the operands of the processor's arithmetic units must come from the RF (Processing Logic Unit), the data stored in the DM (Processing Descriptor) must be loaded into the RF before the processor performs calculations, and the result is then stored in the DM after the calculations are completed. Therefore, the processor reads the source operands from the RF, performs calculations through the ALU or acceleration unit, and writes the result back to the RF, thus completing one instruction operation.

[0041] The traditional CORDIC algorithm iteratively rotates by decomposing the rotation angle into a set of predefined small angles, α i =arctan(2 -i This algorithm is implemented by shifting i (-bits) in hardware. However, it requires a constant multiplier when calculating the square root, which is unsuitable for area and power-sensitive applications, and it cannot be directly used to calculate arctan and arccos. To address these issues, the SF-NEAS CORDIC arithmetic subunit employs an α... i =2 -i Determine the iterative rotation angle; α is the iterative rotation angle, and i is the shift in hardware, in units of -bit.

[0042] 2 -i The unit is radians (rad). In practice, the angle data uses a 16-bit word length. Since the angle coverage should be [-π, π], the angle uses the Q3.13 fixed-point format. Because this algorithm ensures that the length of the vector remains unchanged during rotation, it is very convenient in calculations such as trigonometric functions. For sin and cos calculations, the unit vector (1, 0) rotated to the input angle θ results in a vector (x... n y n Therefore, the results of cosine and sinine operations can be directly obtained as x.n y n In this process, x n and y n No scaling factor is needed. For the square root operation, the two inputs, Input1 and Input2, are assigned to the initial vector (x0, y0), and then y0 is continuously increased during the iteration process. i The square root tends towards 0, and after iteration, the result of the square root calculation is directly x. n For the arccos, arcsin, and arctan operations, the iteration process ensures that x and y approach the target value, respectively. The final accumulated rotation angle is then the corresponding inverse trigonometric function. During this process, the length of the rotation vector remains constant, thus guaranteeing the correct determination of the rotation direction when comparing it with the target value. This CORDIC uses an iterative structure, and as shown... Figure 4 As shown.

[0043] The operation type indicates the type of function calculation performed by the arithmetic unit. There are six function types (sin, cos, arcsin, arccos, arctan, sqrt), therefore this signal uses a 3-bit width. The enable signal indicates when the arithmetic unit starts working. Only input 1 is used when calculating sin and cos, representing the input angle. Similarly, only input 1 is used when calculating arcsin and arccos, representing the input value of the inverse trigonometric function. For the calculation of arctan and sqrt, since the ratio of two numbers and the coordinate system combination are used respectively, there are two inputs (input 1 and input 2). (input 2 / input 1) represents the input for arctan, and the coordinates (input 1, input 2) represent the input for calculating sqrt. Coefficients 1, 2, and 3 represent the coefficients of the three-term expansion of this arithmetic unit.

[0044] Table 1 compares the resource consumption of different implementations of the proposed SF-NEAS CORDIC with existing methods. Since CORDIC resource consumption is related to the architecture, FPGA version, function computation capabilities, and computational accuracy, for fair comparison, the CORDIC in Method 1 also uses an iterative structure and supports both rotation and vector modes simultaneously as much as possible. The computational errors of the CORDICs within the table are kept as consistent as possible. For Table 1, since the CORDIC in Method 1 uses 20 bits, the proposed SF-NEAS CORDIC uses a 20 / 16-bit data width and a 4-term expansion implementation. It can be seen that the proposed CORDIC achieves better Slice-Delay-Product (SDP) and Area-Delay-Product (ADP). When implemented on an SMIC 55nm CMOS process at a frequency of 1.5MHz, the power consumption after DC synthesis is 2.05uW. Although the programmable CORDIC arithmetic unit increases the chip area, it significantly reduces the execution time of the entire program, resulting in a marked reduction in the clock cycle of the entire processor system and lowering the overall power consumption of the processor system.

[0045] Table 1

[0046]

[0047] The acceleration commands include: SIN, COS, ARCSIN, ARCCOS, ARCTAN, and SQRT.

[0048] The inverse square root operation subunit based on the successive approximation algorithm is used to calculate the inverse square root based on the input data. This algorithm eliminates the multipliers required by Newton's iteration method, unified parabolic synthesis method, etc., thus reducing resource consumption.

[0049] Basic operators such as addition, subtraction, and multiplication are relatively easy to design and can achieve high performance and throughput. However, advanced operators such as division, square root, and inverse square root require more hardware and have lower performance and throughput. Implementing these operations directly in hardware involves significant resources and high latency. Division can be easily implemented by obtaining the inverse square root. square root Vector normalization operation Calculating the inverse square root, as an intermediate data point, can both accelerate the processor's computation and ensure its versatility. For input x, its inverse square root... Squaring both sides of the equation and moving x to the left side, we get: x*y 2=1 (4.29). Since x can be any non-negative number, a range restriction needs to be imposed on x during the specific iteration implementation so that the iterative algorithm can select an initial value within this range and iterate within this range. At the same time, any non-negative number can be shifted into this range. For the input x range [0.25, 1), the output y range is (1, 2]. For any input binary fixed-point number, it can always be shifted to be within the interval [0.25, 1). Given an input value x in the range [0.25, 1), and an initial value y of 1, y needs to be selected as the inverse square root within the range (1, 2). First, the midpoint of the range, 1.5, is selected for comparison. The input value x and the test value y, 1.5, are substituted into formula (4.29). If the result is greater than or equal to 1, it means that the test value 1.5 is too large, and the search should continue within the range (1, 1.5). If the result is less than 1, it means that the test value 1.5 is too small, and the search should continue within the range (1, 2) to approximate it. This process is iterated until a suitable value is found.

[0050] Figure 5 The initial values ​​of each parameter are x = Input, y = 1.0, xys = x, and dxy = 2*x.

[0051] The constant here refers to 0.5 j 0.5 (j+1) Truncation refers to expressions that contain only x (such as 2x * 0.5). 2j The input x ranges from [0.25, 1), and a 32-bit fixed-point representation is used. After the 13th iteration, the average error of the inverse square root is basically stable at 10. -9 Magnitude.

[0052] Table 2 summarizes the parameters of the hardware implementation of the inverse square root calculation unit and compares it with other methods. Since the calculation errors of the inverse square root unit differ among methods, to ensure fairness in the comparison, the proposed successive approximation inverse square root algorithm uses different implementation versions. Each version achieves different calculation errors by changing the iteration order, ultimately achieving calculation errors that are essentially consistent with other methods. The three implementation versions, successive approximation 1, 2, and 3, in Table 1 correspond to methods 1, 2, and 3, respectively. Since method 2 does not provide a relative error, implementation version 2 uses a 16-bit maximum iteration count. As can be seen from Table 2, the proposed successive approximation inverse square root algorithm achieves lower resource consumption while maintaining a smaller relative error. The processor used in this application employs a 32-bit iterative structure, and as shown in Table 2, the iterative structure maintains low resource consumption while achieving the highest accuracy.

[0053] Table 2

[0054]

[0055] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.

[0056] This document uses specific examples to illustrate the principles and implementation methods of the present invention. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of the present invention. Furthermore, those skilled in the art will recognize that, based on the ideas of the present invention, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of the present invention.

Claims

1. A processor for use in smart sensors, characterized in that, include: Chip core (CORE) and pin pads; The chip core CORE obtains external signals through the PAD pin; The chip core (CORE) includes: a data memory (DM), a program memory (PM), a register file (RF), auxiliary registers (AR), and a processing module. The program memory (PM) is used to store application code and fetch instructions; the data memory (DM) stores temporary operands and calculation results; the register file (RF) and the auxiliary register (AR) are used to store operands; the processing module is used to read instructions and operands from the program memory (PM), decode them, execute the corresponding calculations, and write the calculation results back to the register file (RF) and the auxiliary register (AR); The processing module includes: an arithmetic logic unit (ALU), a multiplier, a shifter, an acceleration unit, a control path, a data path, a data bus, an address generation unit (AGU), a program control unit, and peripheral devices; The program control unit is used to perform mixed-length instruction and instruction-level parallel processing, output control signals to the data path and address generation unit AGU, and decode immediate values; The acceleration unit includes: an SF-NEAS CORDIC operation subunit and an inverse square root operation subunit based on a successive approximation algorithm; the SF-NEAS CORDIC operation subunit is used to configure different operation functions under different acceleration instructions; The SF-NEAS CORDIC operational subunit adopts α i =2 -i Determine the iterative rotation angle; α is the iterative rotation angle, and i is the shift in hardware, in units of -bit; The acceleration commands include: SIN, COS, ARCSIN, ARCCOS, ARCTAN, and SQRT.

2. The processor for use in a smart sensor according to claim 1, characterized in that, The program control unit employs a state transition mechanism.

3. A processor for use in smart sensors according to claim 1, characterized in that, The inverse square root operation subunit based on the successive approximation algorithm is used to calculate the inverse square root based on the input data.