Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

24 results about "4-bit" patented technology

In computer architecture, 4-bit integers, memory addresses, or other data units are those that are 4 bits wide. Also, 4-bit CPU and ALU architectures are those that are based on registers, address buses, or data buses of that size. A group of four bits is also called a nibble and has 2⁴ = 16 possible values.

Quantization method of large language model, related equipment and computer program product

The invention provides a large language model quantification method, related equipment and a computer program product, and the method comprises the steps: carrying out the reasoning of a to-be-quantized large language model through calibration data, obtaining the activation of the large language model, and carrying out the statistics of the activation distribution of a target layer according to channels; calculating a smoothing factor of each channel according to the activation distribution; compensating the weight of the target layer channel by channel according to the smoothing factor to obtain a compensated weight; performing 4-bit quantization on the compensated weight; in the model reasoning process, activation of a target layer is smoothed, and 16-bit quantization is carried out on the smoothed activation. According to the method, a W4A16 quantification scheme is adopted, compared with W8A8, the weight storage amount is compressed by half, meanwhile, the precision loss is controlled to be smaller than 3%, and deployment of a large language model on edge equipment is facilitated.
Owner:SHANGHAI ZHICHEN MICRO TECHNOLOGY CO LTD

Simd-based four-bit fixed-point BANKROUND implementation method

The invention provides a four-bit fixed-point BANKROUND implementation method based on simd. The four-bit fixed-point BANKROUND implementation method based on simd comprises the steps that S1, floating point conversion shaping and four-bit fixed-point operation are carried out; s2, performing right shift by 4 bits, i.e., after multiplying by 16, dividing by 16 for reduction, retaining an integer part of the data, and reducing the fixed-point data into original data; s3, retaining a decimal part of the data, wherein the decimal part is expressed as result3 = result1amp; 0 * 00001111, 0 * 00001111; s4, judging whether a decimal part is 0.5 or not: judging whether the decimal part is equal to 0.5 or not in banker rounding; s5, according to the BANKROUND standard, the condition of entering 1 needs to be met at the same time: case 1) decimal part gt; the condition is completed in the process of S1, and all the parts larger than or equal to 0.5 directly enter 1; 2) the decimal part is equal to 0.5, and the unit of the integer part needs to be ensured to be an odd number; s6, setting each 16-bit number meeting the condition in each result5 as 1, and setting each 16-bit number not meeting the condition as 0; and S7, subtracting the calculation result in the step S6 from the result after shifting calculated in the step S2.
Owner:HEFEI JUNZHENG TECH CO LTD

R-2r digital-to-analog converter with auxiliary calibration structure

The application discloses an R-2R digital-analog converter with an auxiliary calibration structure and relates to the technical field of high-precision digital-analog conversion. The converter comprises a main DAC, a calibration quantity calculation module, a compensation code generation module, an auxiliary DAC and an operational amplifier circuit. The main DAC converts a digital input into an analog current. The calibration quantity calculation module generates a 15-bit compensation code containing a 1-bit sign bit and a 14-bit data bit through a 24-bit ADC, a comparator and successive approximation logic, and only calibrates the high 12 bits of the main DAC bit by bit, and the low 4 bits do not need to be calibrated. The compensation code generation module generates a total compensation code through gating by a multiplexer and superposition by an adder. The auxiliary DAC is a 14-bit binary current type structure, and realizes bidirectional compensation in cooperation with a positive and negative reference voltage source. The application has the advantages of short calibration period, small storage consumption, no need for additional bias calibration, excellent linearity, and suitability for application in the fields of wireless communication, biological medicine and the like, and meets the application requirements of high resolution and high precision.
Owner:TONGJI UNIV

High-performance block cipher construction method and device, electronic equipment and storage medium

The application discloses a high-performance block cipher construction method and device, electronic equipment and a storage medium, wherein the method comprises: an encryption process: using a preset round function generation algorithm to encrypt plaintext to obtain ciphertext, wherein a round key in the preset round function generation algorithm is obtained by performing position permutation on 32 positions in units of 4 bits through a 128-bit key scheduling algorithm; and a decryption process: decrypting the ciphertext through a decryption algorithm to obtain the plaintext. Thus, the problem that the existing block cipher algorithm has high resource occupation when hardware is efficiently implemented, and has low software implementation efficiency is solved.
Owner:TSINGHUA UNIVERSITY

Electronic device performing calculation using artificial intelligence model, and method for operating electronic device

An electronic device for performing a multiply and accumulate (MAC) operation includes at least one MAC unit, a memory, at least one 8b×8b operator, at least one 8b×4b operator, at least one 4b×4b operator, a bit shifter, an adder, at least one accumulator, and a processor. The processor receives first bit string data and second bit string data which include 12 bits, divides the first bit string data into 4 bits and 8 bits, divides the second bit string data into 4 bits and 8 bits, outputs two 8-bit results corresponding to activation 8-bit, weight 8-bit (A8W8) by using the bit shifter and a first accumulator on the basis of determining to output an A8W8 result, and outputs one 12-bit result corresponding to activation 12-bit, weight 12-bit (A12W12) on the basis of determining to output an A12W12 result.
Owner:SAMSUNG ELECTRONICS CO LTD

A high-precision wind turbine blade pitch angle detection system and method

This invention belongs to the field of wind turbine generator condition monitoring technology, and in particular to a high-precision wind turbine blade pitch angle detection system and method. The high-precision wind turbine blade pitch angle detection system includes a video acquisition module, high-reflectivity markers, an integrated processing unit, and a digital output module. The video acquisition module is installed inside the turbine hub and includes a global shutter camera and a ring infrared light source. The high-reflectivity markers are fixed on the surface of the pitch bearing and are ceramic-based reflective patches. The integrated processing unit adopts a RISC-V main control chip and FPGA collaborative architecture. This invention accelerates image preprocessing and optical flow tracking through FPGA, and combines Kalman filtering to achieve 0.1 accuracy angle detection. The result is encoded into a 7+4 bit binary signal and transmitted to the PLC via an isolated CAN bus. It is particularly suitable for real-time dynamic deviation detection in high-vibration and wide-temperature-varying environments, and especially suitable for real-time pitch control of large-megawatt wind turbine generators.
Owner:BEIJING YADESHI ENGINEERING TECHNOLOGY CONSULTING SERVICE CO LTD

Hybrid parallel time sequence Pipelined-SAR ADC (Synthetic Aperture Radar Analog to Digital Converter)

The invention discloses a mixed parallel time sequence Pipelined-SAR ADC (Synthetic Aperture Radar Analog to Digital Converter), and relates to the technical field of integrated circuits, the Pipelined-SAR ADC comprises a first-stage SAR, a ring amplifier, a second-stage SAR, an inverter-based open-loop amplifier and a third-stage SAR; wherein the resolutions of the first-stage SAR, the second-stage SAR and the third-stage SAR are respectively set to be 5bit, 4bit and 5bit, and 2-bit redundancy is included in the first-stage SAR, the second-stage SAR and the third-stage SAR; the first-stage SAR adopts bottom plate sampling, and the second-stage SAR and the third-stage SAR adopt top plate sampling. According to the invention, a hybrid parallel time sequence scheme is provided, the problem that the time sequence of the second-stage SAR is tense is solved, the speed of the whole Pipelined-SAR ADC is improved, and the robustness of the ring amplifier is improved.
Owner:SUN YAT SEN UNIV

A data interleaving method and apparatus

The application discloses a data interleaving method and device, and relates to the field of optical fiber communication systems.The method comprises the following steps: dividing an input bit data stream into a plurality of first data blocks by taking 4 bits as a unit, exchanging 2 bits of data in the middle of each first data block; dividing the exchanged bit data stream into second data blocks by taking 512 bits as a unit, and performing intra-block interleaving; and performing inter-block interleaving on the bit data stream after the intra-block interleaving through two random access memories.The application guarantees the original probability distribution characteristics of the PCS technology, and improves the ability of correcting burst errors in the joint application of the FEC technology and the PCS technology.
Owner:FIBERHOME TELECOMMUNICATION TECHNOLOGIES CO LTD +1

Design method and use method of variable precision calculation unit applied to quantized neural network convolution layer

The application provides a design method and use method of a variable precision calculation unit applied to a quantized neural network convolution layer, shift operation is performed on a vector with a precision of 1, and the shift bit number, dimension and number of an operation basic block are determined. After the operation basic blocks are arranged, an addition tree is connected to form an array, and a variable precision fusion calculation unit is obtained. According to the actual precision of the current convolution layer activation value and the weight value, the variable precision fusion calculation unit array is configured, and each column of the fusion calculation unit is calculated in parallel. After the current convolution layer operation is completed, the two-dimensional array is reconfigured according to the data precision of the next convolution layer, and the operation of all convolution layers is completed in this way. The problems that the operation unit in the prior art is prone to causing waste of calculation resources in convolution operation less than 4 bits, and cannot dynamically adopt different quantization bit widths to process data according to the actual situation of the convolution layer are solved, and the resource utilization rate and the calculation efficiency are improved.
Owner:NAT INNOVATION INST OF DEFENSE TECH PLA ACAD OF MILITARY SCI

Bus signal for vehicle-gauge-level MCU chip adaptation evaluation, and generation method and device of bus signal for vehicle-gauge-level MCU chip adaptation evaluation

The invention relates to the technical field of data transmission, and discloses a bus signal for vehicle gauge level MCU chip adaptation evaluation, a generation method and a device, the bus signal is composed of a start frame, a 6-bit racing frame, a variable length control frame, a data frame, a 16-bit CRC check frame, a response frame and an end frame in sequence; 3-bit priorities and 3-bit data type identifiers are built in the race frame, the length of the control frame is dynamically expanded by multiplying the number of data type identifiers in the race frame as 0 by 4 bits, and the byte length of integer and single / double-precision floating point type data is sequentially defined 4 bits by 4 bits, so that the self-description of a protocol layer is realized. The generating device automatically completes frame coding, data filling, periodic sending and response backward reading after parameters are configured on the upper computer, a test vector can be switched in a zero script mode within the range of 1-200 ms, and a single frame bears 192-byte mixed data to the maximum extent, so that real-time blind decoding can be achieved without preset analysis of an MCU chip. According to the invention, the adaptation verification period of the vehicle-gauge-level MCU chip is obviously shortened, the test coverage rate is improved, and the manual coding error risk is reduced.
Owner:BEIJING NEW ENERGY VEHICLE TECH INNOVATION CENT CO LTD

Design method and use method of variable precision calculation unit applied to quantization neural network convolutional layer

The invention provides a design method and a use method of a variable precision calculation unit applied to a quantization neural network convolutional layer, and the method comprises the steps: carrying out the shift operation of a vector with the precision of 1, and determining the shift digit, dimension and number of an operation basic block; and arranging operation basic blocks and then connecting the operation basic blocks by using an adder tree to form an array to obtain a fusion calculation unit with variable precision. According to the actual precision of the activation value and the weight value of the current convolutional layer, a variable-precision fusion calculation unit array is configured, and all columns of fusion calculation units carry out parallel calculation; and after the operation of the current convolutional layer is completed, according to the data precision of the next convolutional layer, the two-dimensional array is reconfigured, and the operation of all the convolutional layers is completed by parity of reasoning. The problems that in the prior art, an operation unit is prone to causing waste of calculation resources in convolution operation smaller than 4 bits, different quantization bit widths cannot be flexibly and dynamically adopted to process data according to the actual situation of a convolution layer, and the resource utilization rate and the calculation efficiency cannot be improved are solved.
Owner:NAT INNOVATION INST OF DEFENSE TECH PLA ACAD OF MILITARY SCI

Matmul transpose weight splitting method based on simd

The invention provides a matmul transpose weight splitting method based on simd, in the method, weight splitting is formed by combining a series of 2bit transpose logics, an input column is transposed to an output row, the method is suitable for matmul calculation in any shape, and the method comprises the following steps: S1, reading data in a first row, namely data before transpose, from an input address by using an mx512lad (vr0, 0, input tradr, 0) instruction; s2, aggregating the 4 bit data together by using an mx512gt4bi instruction; s3, all the 2bits are aggregated together by using an mx512gt2bi instruction; s4, all the 16 bits are aggregated together by using an mx512th instruction; and S5, transposing the generated rows of data by using an mx512ilve instruction and an mx512ilvo instruction to obtain a weight splitting target output format. According to the method, additional input and output space development is reduced; a large number of la and sa instructions are reduced; and a large number of redundant instructions are removed.
Owner:HEFEI JUNZHENG TECH CO LTD

Method for transmitting video signal containing Alpha channel based on high-definition multimedia interface, receiving method and related electronic equipment

The invention discloses a method for transmitting a video signal containing an Alpha channel based on a high-definition multimedia interface, a receiving method and related electronic equipment, and relates to the technical field of digital video transmission. The method applied to the electronic equipment at the transmitting end comprises the following steps: acquiring to-be-transmitted video data with an RGB color signal and an Alpha channel signal; the RGB signal is converted into a YCbCr signal containing 8 bit Y, 4 bit Cb and 4 bit Cr; respectively filling YCbCr and Alpha channel signals into exclusive bearing intervals of the RGB three channels according to a preset mapping rule and a 24-bit width requirement, and generating to-be-sent data of which the single-frame total amount, the resolution and the time sequence are consistent with those of RGB4: 4: 4; and synchronous transmission of the video data to be transmitted is realized through a single interface. According to the scheme, the RGB color signals and the Alpha channel signals are synchronously transmitted through a single interface, and data integrity and transmission compliance are guaranteed.
Owner:BEIJING TIMES OSEE TECH CO LTD

A gamma / vcom voltage value storage method

The application discloses a Gamma / Vcom voltage value storage method, which comprises the following steps: receiving a plurality of voltage values generated by a first Gamma IC, converting each bit of the voltage value into a corresponding four-bit binary number, and sequentially storing the binary number into a preset position; reading the binary number, converting the binary number of every 4 bits into a decimal number, and outputting the voltage value composed of the converted decimal number by a second Gamma IC. For the Gamma / Vcom voltage value required by a customer, the application does not need to convert between two Gamma ICs, but only needs to output the Gamma / Vcom voltage value stored in the XB board EEP or FLASH, so that the application has universality and versatility, and can be used in SoCs designed by various different Gamma IC models, and has strong adaptability.
Owner:XIANYANG CAIHONG OPTOELECTRONICS TECH CO LTD

Optoelectronic adder

The present invention relates to a 4-bit optoelectronic adder using IR LED's, IR detectors and supporting circuit. The adder performs binary operations on two numbers each of 4 bits. The adder uses only Infrared light intensity and its combination in various forms to give the result of addition which is detected by IR detector and identified by supporting electronic circuitry to give output of addition.
Owner:KUMAR TEJASVI

Intra-mode JVET coding

This provides a suitable method for dividing video coding blocks for JVET. [Solution] The MPM set includes a set of intra predictive coding modes other than the six, and is encoded using truncated unary binarization. The 16 selected intra predictive coding modes are encoded using 4 bits of fixed-length code, and the remaining unselected coding modes are encoded using truncated binary coding. The JVET coding tree unit is coded as the root node of a QTBT structure. The QTBT structure has a quadtree branching from the root node and binary trees branching from each of the multiple leaf nodes of the quadtree, and using asymmetric binary partitioning, the coding unit represented by the leaf nodes of the quadtree is divided into multiple child nodes, and the multiple child nodes are represented as multiple leaf nodes of binary trees branching from the leaf nodes of the quadtree.
Owner:ARRIS ENTERPRISES LLC

Lightweight dynamic s-box construction method for circuit implementation

The application discloses a lightweight dynamic S-box construction method for circuit implementation, and by using an automatic tool STP solver, a dynamic S-box structure satisfying specific cryptographic security properties can be automatically searched under given hardware resource constraints, the cryptographic security properties of the S-box under 1-4 bit dynamic control are considered, the logic gate circuit cost required can be effectively constrained, and the balance between the security performance of the dynamic S-box and the hardware implementation overhead can be explored, so that the result of considering both security and lightweight is obtained.
Owner:HANGZHOU DIANZI UNIV

An activation function-based double-table 128-entry optimization method

This invention provides a dual-table 128-entry optimization method based on activation functions. It accelerates the table lookup logic using SIMD instructions and obtains the correct table entry calculation by setting a corresponding mask through index merging. The method includes: dual-table lookup: dividing table entries into positive and negative entries; selecting either positive or negative numbers; S1, clipping the input data to a threshold range, making it a number between 0 and 127; S2, determining the range of the input data for selecting positive or negative entries; selecting Neg_dy and Neg_y0 for -127 to 0, and Pos_dy and N for 0 to 127. eg_y0; S3, Original index calculation: Original table entry index -127 to 127, each entry has 4 bits, totaling 128 numbers; (a) Word index; (b) Byte index; (c) Entry merging; S4, Determine the table entry index, use the instruction to look up the table: for positive numbers, update the table entry value corresponding to the positive number; for negative numbers, update the table entry value corresponding to the negative number; S5, Separate the high and low 4 bits to obtain the valid bit width data; when the data is negative, the corresponding data of the table entry needs to be updated; S6, Perform a multiplication operation on the data obtained from the table entry to obtain the final result.
Owner:HEFEI JUNZHENG TECH CO LTD

Data processing method, device and system

The embodiment of the invention discloses a data processing method, device and system. Specifically, PCS processing is performed on a first bit set in a plurality of bits to obtain a second bit set. And performing first interleaving on the second bit set and a third bit set except the first bit set in the plurality of bits to obtain two fourth bit sets. And respectively carrying out forward error correction (FEC) coding on the two fourth bit sets to obtain two fifth bit sets. And respectively carrying out second interleaving on the two fifth bit sets to obtain two sixth bit sets. And performing third interleaving on the two sixth bit sets to obtain a seventh bit set. It needs to be explained that continuous 4 * Z + 4 bits in the seventh bit set are used for mapping to obtain one dual-polarization symbol, the dual-polarization symbol comprises a first polarization symbol and a second polarization symbol, and Z is an integer greater than or equal to 3.
Owner:HUAWEI TECH CO LTD

Finite-precision quantization hierarchical nonsurjective finite character set decoding method applicable to 5G LDPC codes

The application provides a limited-precision quantization layered non-surjective finite alphabet (NS-FAID) decoding method suitable for 5G LDPC codes.Under the condition of limited-precision quantization, especially limited-precision quantization less than or equal to 4 bits, the bit width of soft quantity information in the general quantization decoding algorithm is generally consistent with the bit width of input channel information, while the layered NS-FAID decoding method designed for 5G LDPC codes provided by the application can further reduce the bit width of soft quantity information used in the check node update process in the iterative decoding process by using the lookup table mapping mode, so that the complexity of the check node internal calculation is further reduced, thereby saving hardware resources, reducing power consumption, and the performance loss is smaller compared with the standard decoding algorithm.
Owner:SOUTHEAST UNIV

High-parallel rejection sampler for post-quantum cryptography and working method thereof

PendingCN122339692AData compressionPathPing
This invention discloses a highly parallel rejection sampler for post-quantum cryptography and its operating method. The 8-bit sampled data is split into symmetrical high-4 and low-4 bit parallel sampling units: the high-4 bit unit performs data compression towards the least significant bit, and the low-4 bit unit performs data compression towards the most significant bit. The two compressed outputs are combined through physical splicing to form a pre-aligned data stream with continuous data in the middle and possible invalid data bits at the beginning and end. Subsequently, the system uses the effective data count value of the high-bit unit as a control signal to drive a multiplexer network to directly select the output result according to five preset alignment paths, thereby achieving a compact arrangement of fully parallel data. This invention avoids the complex multi-level shift operations in traditional architectures, greatly optimizes critical path latency and hardware resource overhead, and is suitable for high-speed hardware accelerator design for post-quantum cryptography algorithms.
Owner:STATE GRID JIANGSU ELECTRIC POWER CO LTD RESEARCH INSTITUTE +1

An activation function-based single-table 128 entry optimization method

This invention provides an optimization method for a single 128-entry table based on activation functions. The method accelerates the table lookup logic using SIMD instructions, merges indexes, sets corresponding masks, obtains the correct table entries, and then performs calculations. The method includes: Single-table lookup: Input data is determined to be positive, using only positive entries; Pos_dy: Positive entry 0, a total of 128 entries, each entry is 8 bits, totaling 512 bits, which can be stored in a register; Pos_y0: Positive entry 1, a total of 128 entries, each entry is 8 bits, totaling 512 bits, which can be stored in a register; S1: Input data is cilpped to a threshold range, making it a number between 0 and 127; S2: Original index calculation; S3: Determine the table entry index and use instructions to look up the table; S4: Separate the high and low 4 bits to obtain effective bit width data; S5: Multiply the data obtained from the table entries to obtain the final result. By using SIMD instructions and designing corresponding masks, the table lookup function based on a single 128-entry table based on activation functions is completed.
Owner:HEFEI JUNZHENG TECH CO LTD

Satellite-borne image storage system based on row and column combination extended Hamming code error correction

The invention discloses a satellite-borne image storage system based on row and column combination extended Hamming code error correction, and the system comprises a power module which is used for providing a stable required voltage for the satellite-borne image storage system; the main control module transmits the data to the storage array module in real time through a DMA controller, and is internally provided with a row-column combined extended Hamming code to perform verification and error correction on the data to be stored so as to realize real-time processing and monitoring of the data; the storage array module is used for storing data; the interface communication module is used for realizing data transmission between the main control module and the central computer; and the temperature monitoring module is used for monitoring the temperature of the circuit board in real time, and regulating and controlling the internal temperature through the main control module and the external thermal insulation structure. The limitation of single-bit error correction and double-bit detection of traditional Hamming codes in satellite-borne storage can be solved, error correction within 4 bits is achieved, and the size of the satellite-borne storage is reduced.
Owner:NANJING UNIV OF SCI & TECH