Optimization processing method for image fusion in time epitome

By using 8bit data and simd instructions for image fusion processing, the problem of slow processing speed in the prior art is solved, real-time processing is realized, and processing speed is significantly improved.

CN120339083APending Publication Date: 2025-07-18INGENIC SEMICON CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410066786.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-17
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The prior art has slow processing speed in time microscopic image fusion, which cannot meet the real-time processing requirements, and the processing speed is slow when implementing C language programs.

Method used

Using 8bit data and 8bit simd instructions, the calculation amount is reduced through coefficient selection and processing, and using load data, save data, 8bit addition, subtraction and shift instructions to achieve image fusion and avoid data conversion and saturation processing.

Benefits of technology

It significantly improves processing speed and meets real-time processing needs. It is about 10 times higher than ordinary C programs, and is especially suitable for chips of Junzheng Integrated Circuit Co., Ltd.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339083A_ABST
    Figure CN120339083A_ABST
Patent Text Reader

Abstract

The invention provides an optimization processing method for image fusion in time epitome, and the method comprises the steps: S1, coefficient selection and processing: employing 8-bit data and a 8-bit simd instruction, and carrying out the decomposition of parameters; the processing of each coefficient is # imgabs0 #, wherein n is equal to 1, 2, 3, 4, 5, 6 and 7; the coefficient of the first picture in the nth fused image is # imgabs1 #, the sum of the coefficients of the two pictures is 1, the coefficient of the other picture is # imgabs2 #, and the coefficient of the other picture is # imgabs3 # S2, and the design of the fused image is as follows: the data numbers of the two pictures of the fused image are the same, and only each piece of data needs to be correspondingly processed and then added; a data loading instruction, a data storage instruction, an 8-bit addition instruction, an 8-bit subtraction instruction and an 8-bit shift instruction are used; 64 pieces of 8-bit data are stored in each register; all data are unsigned data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image processing, and particularly relates to an optimized processing method for image fusion in a time-lapse summary. Background Art

[0002] In the prior art, especially in the field of image processing, the process of fusing pictures in a time-lapse summary is a gradual fusion process, that is, a process in which the proportion coefficient of the pre-added pictures gradually increases. The commonly used coefficients are 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1.0. In addition, the image fusion technology converts 8-bit data into 32-bit integers, then converts them into floating points, multiplies by the specified floating-point coefficient in the fusion, adds the data, converts the obtained data back into integers, and finally performs saturation processing to 0-255 to generate a fused picture.

[0003] However, the defects in the prior art are as follows:

[0004] 1. In the prior art, this fusion technology is processed in a video stream, which has a high requirement for processing speed, and the traditional processing method cannot meet real-time processing.

[0005] 2. In addition, the prior art is implemented using a C language program, and the processing speed is very slow.

[0006] In addition, the commonly used technical terms in the prior art include:

[0007] 1. Time-lapse summary, where several historical video photos are fused into a single video, and the fused pictures gradually enter the video and finally ghost with the images in the video.

[0008] 2. Image fusion, the process of multiplying the first and second pictures by their corresponding coefficients and then adding them to obtain an image is called image fusion.

[0009] 3. SIMD, an instruction for vector operations. Summary of the Invention

[0010] To solve the above problems, the purpose of the present application is: through the selection and processing of coefficients and the derivation of calculation formulas, ultimately reduce the amount of calculation, reduce the number of instructions used, complete all data within 8 bits, no longer require data conversion, no saturation processing, improve the processing speed, and meet the actual situation. It can be implemented through SIMD instructions and is applicable to chips of the SIMD instruction type, especially applicable to the chip products of Beijing Junzheng Integrated Circuit Co., Ltd. (abbreviation: Junzheng).

[0011] Specifically, the present invention provides an optimized processing method for image fusion in a time-lapse summary, and the method includes the following steps:

[0012] S1. Selection and processing of coefficients:

[0013] Decompose the parameters using 8-bit data and 8-bit SIMD instructions;

[0014] The basic coefficients adopted are:

[0015] 0.125, 0.25, 0.375, 0.5, 0.625, 0.75, 0.875;

[0016] Therefore, the corresponding data is:

[0017]

[0018] Processing of each coefficient: where n = 1, 2, 3, 4, 5, 6, 7;

[0019] The coefficient of the first picture in the nth fused image is Since the sum of the coefficients of the two pictures is 1,

[0020] Therefore, the coefficient of the other picture is The coefficient is

[0021] S2. Design of the fused image:

[0022] The number of data of the two pictures in the fused image is the same. Just process each corresponding data and then add them up;

[0023] Let the data of the first picture in the ith fused image be x, and the corresponding processing coefficient be xm i , let the data of the first picture of the fused image be y, and the corresponding processing coefficient be ym i ; the data of the fused image is z; where ym i = 1 - xm i , i ∈ [1, 7]; the calculation formula is:

[0024] z = x × xm i + y × ym i ;

[0025] z = x × xm i + y × ym i ;

[0026] S2.1. Deduction of the calculation formula:

[0027] According to the calculation formula, substitute the coefficients into the formula to calculate each fused image; all calculations use addition, subtraction, and shift processing, and the rounding method for shift is the downward rounding method;

[0028] Assume that each picture has a data size of size;

[0029] Calculation method for the i-th fused image:

[0030]

[0031] S2.2, Design of SIMD:

[0032] Five SIMD instructions are used: data loading instruction, data saving instruction, 8-bit addition instruction, 8-bit subtraction instruction, 8-bit shift instruction; 64 8-bit data are stored in each register; all data here are unsigned;

[0033] SIMD implementation process:

[0034] Design SIMD according to the calculation formula;

[0035] Assume that the number of data in the picture is size, the first picture is src1, load the data into vr1, the second picture is src2, load the data into vr2, and the fused image is dst;

[0036] Calculation method for the first fused image: z = x >> 3 + y - (y >> 3);

[0037] Calculation method for the second fused image: z = x >> 2 + y - (y >> 2);

[0038] Calculation method for the third fused image: z = x >> 1 + x >> 3 + y >> 1 - (y >> 3);

[0039] Calculation method for the fourth fused image: z = x >> 1 + y >> 1;

[0040] Calculation method for the fifth fused image: z = y >> 1 + y >> 3 + x >> 1 - (x >> 3);

[0041] Calculation method for the sixth fused image: z = y >> 2 + x - (x >> 2);

[0042] Calculation method for the seventh fused image: z = y >> 3 + x - (x >> 3).

[0043] In the step S1, it further includes:

[0044] Processing of each coefficient:

[0045]

[0046]

[0047]

[0048]

[0049]

[0050]

[0051]

[0052] The first picture coefficient in the first fused image is Since the sum of the coefficients of the two pictures is 1, the coefficient of the other picture is That is The coefficient is

[0053] The first picture coefficient in the second fused image is Since the sum of the coefficients of the two pictures is 1, the coefficient of the other picture is That is The coefficient is

[0054] The first picture coefficient in the third fused image is That is Since the sum of the coefficients of the two pictures is 1, the coefficient of the other picture is That is The coefficient is

[0055] The first picture coefficient in the fourth fused image is Since the sum of the coefficients of the two pictures is 1, the coefficient of the other picture is The coefficient is

[0056] The first picture coefficient in the fifth fused image is That is Since the sum of the coefficients of the two pictures is 1, the coefficient of the other picture is That is The coefficient is

[0057] The first picture coefficient in the sixth fused image is That is Since the sum of the coefficients of the two pictures is 1, the coefficient of the other picture is The coefficient is

[0058] The first picture coefficient in the seventh fused image is That is Since the sum of the coefficients of the two pictures is 1, the coefficient of the other picture is Coefficient is

[0059] In the said step S2.1,

[0060] Calculation method of the first fused image:

[0061]

[0062]

[0063] z = x >> 3 + y - (y >> 3)

[0064] Calculation method of the second fused image:

[0065]

[0066]

[0067] z = x >> 2 + y - (y >> 2)

[0068] Calculation method of the third fused image:

[0069]

[0070]

[0071] z = x >> 1 + x >> 3 + y >> 1 - (y >> 3)

[0072] Calculation method of the fourth fused image:

[0073]

[0074] z = x >> 1 + y >> 1

[0075] Calculation method of the fifth fused image:

[0076]

[0077]

[0078] z = y >> 1 + y >> 3 + x >> 1 - (x >> 3)

[0079] Calculation method of the sixth fused image:

[0080]

[0081]

[0082] z = y >> 2 + x - (x >> 2)

[0083] Calculation method of the seventh fused image:

[0084]

[0085]

[0086] z = y >> 3 + x - (x >> 3);

[0087] Among them, the calculation method of the seventh fused image is to swap x and y in the calculation method of the first fused image; the calculation method of the sixth fused image is to swap x and y in the calculation method of the second fused image; the calculation method of the fifth fused image is to swap x and y in the calculation method of the third fused image.

[0088] In the step S2.2, the description of the simd instruction:

[0089] Data loading instruction. Let the 512-bit register be vr1 and the input picture data pointer be src; this instruction is to load 512-bit data in src into vr1:

[0090] ingenic_la(vr1,src);

[0091] Data saving instruction. Let the 512-bit register be vr1 and the saved picture data pointer be dst; this instruction is to save 512-bit data into dst:

[0092] ingenic_sa(vr1,dst);

[0093] 8-bit addition instruction. Let the input registers be vr1 and vr2, and the addition calculation result output register be vr3:

[0094] ingenic_add_u8bit(vr3,vr2,vr1);

[0095] 8-bit subtraction instruction. Let the input registers be vr1 and vr2, and the subtraction calculation result output register be vr3. The minuend is vr2, the subtrahend is vr1, and the difference is vr3:

[0096] ingenic_sub_u8bit(vr3,vr2,vr1);

[0097] 8-bit shift instruction. Let the input register be vr1, and the shift calculation result output register be vr3. Each data is shifted to the right by a constant li bits, and the calculation is performed by rounding down:

[0098] ingenic_srli_u8bit(vr3,vr2,li).

[0099] The calculation method of the first fused image:

[0100] Initialize the integer i = 0, expressed as int i = 0;

[0101] When it is judged that i < size, execute the following loop body, otherwise jump out of the loop body. After the loop body is executed, perform i = i + 64, expressed as for(; i < size; i += 64);

[0102] Start of the loop body:

[0103] Load data from src1 to vr1, expressed as ingenic_la(vr1, src1);

[0104] Jump the src1 pointer by 64 8-bit units, expressed as src1 += 64;

[0105] Load data from src2 to vr2, expressed as ingenic_la(vr2, src2);

[0106] Jump the src2 pointer by 64 8-bit units, expressed as src2 += 64;

[0107] Perform a right shift by 3 bits calculation, expressed as ingenic_srli_u8bit(vr1, vr1, 3);

[0108] Perform a right shift by 3 bits calculation, expressed as ingenic_srli_u8bit(vr3, vr2, 3);

[0109] Perform a subtraction calculation, expressed as ingenic_sub_u8bit(vr2, vr2, vr3);

[0110] Perform an addition calculation, expressed as ingenic_add_u8bit(vr1, vr2, vr1);

[0111] Save the data in the register to dst, expressed as ingenic_sa(vr1, dst);

[0112] Jump the dst pointer by 64 8-bit units;

[0113] End of the loop body;

[0114] For the part that cannot be divided evenly by 64, it is expressed as:

[0115] Initialize the integer left, expressed as int left = size - i;

[0116] When it is judged that i < left, the following loop body is executed; otherwise, the loop body is jumped out. After the loop body is executed, it is expressed as for(i = 0; i < left; i++);

[0117] Start of the loop body:

[0118] An integer x, expressed as int x = src1[i];

[0119] An integer y, expressed as int y = src2[i];

[0120] An integer z, expressed as int z = x >> 3 + y - (y >> 3);

[0121] Save the picture data pointer, expressed as dst[i] = z;

[0122] End of the loop body;

[0123] This part of the loop is for the part that cannot be divided by 64 evenly, and the remaining part is implemented using ordinary C programs, processing one data each time.

[0124] The calculation method of the second fused image:

[0125] Initialize the integer i = 0, expressed as int i = 0;

[0126] When it is judged that i < size, the following loop body is executed; otherwise, the loop body is jumped out. After the loop body is executed, i = i + 64, expressed as

[0127] for(; i < size; i += 64)

[0128] Start of the loop body:

[0129] Load the data of src1 into vr1, expressed as ingenic_la(vr1, src1);

[0130] Jump the pointer of src1 by 64 8-bit units, expressed as src1 += 64;

[0131] Load the data of src2 into vr2, expressed as ingenic_la(vr2, src2);

[0132] Jump the pointer of src2 by 64 8-bit units, expressed as src2 += 64;

[0133] Right shift by 2 calculation, expressed as ingenic_srli_u8bit(vr1, vr1, 2);

[0134] Right shift by 2 calculation, expressed as ingenic_srli_u8bit(vr3, vr2, 2);

[0135] Subtraction calculation, expressed as ingenic_sub_u8bit(vr2, vr2, vr3);

[0136] Addition calculation, expressed as ingenic_add_u8bit(vr1, vr2, vr1);

[0137] Save the data in the register to dst, expressed as ingenic_sa(vr1, dst);

[0138] Jump the dst pointer by 64 8-bit units;

[0139] End of the loop body;

[0140] The part that cannot be divided evenly by 64 is expressed as

[0141] Initialize the integer left, expressed as int left = size - i;

[0142] When judging i < left, execute the following loop body, otherwise jump out of the loop body. After the loop body is executed, it is expressed as for(i = 0; i < left; i++);

[0143] Start of the loop body:

[0144] The integer x, expressed as int x = src1[i];

[0145] The integer y, expressed as int y = src2[i];

[0146] The integer z, expressed as int z = x >> 2 + y - (y >> 2);

[0147] Save the image data pointer, expressed as dst[i] = z;

[0148] End of the loop body;

[0149] This part of the loop is for the part that cannot be divided evenly by 64. The remaining part is implemented using ordinary C programs, processing one data at a time.

[0150] The calculation method of the third fused image:

[0151] Initialize the integer i = 0;

[0152] When judging i < size, execute the following loop body, otherwise jump out of the loop body. After the loop body is executed, perform i = i + 64, expressed as for(; i < size; i += 64);

[0153] Start of the loop body:

[0154] Load the data from src1 into vr1, represented as ingenic_la(vr1,src1);

[0155] Jump the src1 pointer by 64 8-bit units, represented as src1 += 64;

[0156] Load the data from src2 into vr2, represented as ingenic_la(vr2,src2);

[0157] Jump the src2 pointer by 64 8-bit units, represented as src2 += 64;

[0158] Perform a right shift by 1 calculation, represented as ingenic_srli_u8bit(vr4,vr1,1);

[0159] Perform a right shift by 3 calculations, represented as ingenic_srli_u8bit(vr5,vr1,3);

[0160] Perform a right shift by 3 calculations, represented as ingenic_srli_u8bit(vr3,vr2,3);

[0161] Perform a subtraction calculation, represented as ingenic_sub_u8bit(vr2,vr2,vr3);

[0162] Perform an addition calculation, represented as ingenic_add_u8bit(vr1,vr5,vr4);

[0163] Perform an addition calculation, represented as ingenic_add_u8bit(vr1,vr2,vr1);

[0164] Save the data in the register to dst, represented as ingenic_sa(vr1,dst);

[0165] Jump the dst pointer by 64 8-bit units;

[0166] The loop body ends;

[0167] The part that cannot be divided evenly by 64 is represented as

[0168] Initialize the integer left, represented as int left = size - i;

[0169] When it is judged that i < left, execute the following loop body, otherwise jump out of the loop body. After the loop body is executed, it is represented as for(i = 0; i < left; i++);

[0170] The loop body starts:

[0171] An integer x, represented as int x = srcl[i];

[0172] An integer y, represented as int y = src2[i];

[0173] An integer z, represented as int z = x >> 1 + x >> 3 + y >> 1 - (y >> 3);

[0174] Save the pointer of the image data, represented as dst[i] = z;

[0175] The loop body ends;

[0176] The calculation method of the fourth fused image:

[0177] Initialize the integer i = 0;

[0178] When it is judged that i < size, execute the following loop body, otherwise jump out of the loop body. After the loop body is executed, perform i = i + 64, represented as for(; i < size; i += 64);

[0179] The loop body starts:

[0180] Load the data of src1 into vr1, represented as ingenic_la(vr1,src1);

[0181] Jump the pointer of src1 by 64 8-bit units, represented as src1 += 64;

[0182] Load the data of src2 into vr2, represented as ingenic_la(vr2,src2);

[0183] Jump the pointer of src2 by 64 8-bit units, represented as src2 += 64;

[0184] Perform a right shift by 1 calculation, represented as ingenic_srli_u8bit(vr1,vr1,1);

[0185] Perform a right shift by 1 calculation, represented as ingenic_srli_u8bit(vr2,vr2,1);

[0186] Perform an addition calculation, represented as ingenic_add_u8bit(vr1,vr2,vr1);

[0187] Save the data in the register to dst, represented as ingenic_sa(vr1,dst);

[0188] Jump the dst pointer by 64 8-bit units;

[0189] The loop body ends;

[0190] The part that cannot be divided evenly by 64 is expressed as:

[0191] Initialize the integer left, expressed as int left = size - i;

[0192] When it is judged that i < left, execute the following loop body, otherwise jump out of the loop body. After the loop body is executed, it is expressed as for(i = 0; i < left; i++);

[0193] The loop body starts:

[0194] The integer x is expressed as int x = src1[i];

[0195] The integer y is expressed as int y = src2[i];

[0196] The integer z is expressed as int z = x >> 1 + y >> 1;

[0197] Save the picture data pointer, expressed as dst[i] = z;

[0198] The loop body ends;

[0199] The calculation method of the fifth fused image:

[0200] Initialize the integer i = 0;

[0201] When it is judged that i < size, execute the following loop body, otherwise jump out of the loop body. After the loop body is executed, perform i = i + 64, expressed as for(; i < size; i += 64);

[0202] The loop body starts:

[0203] Load the data of src2 into vr1, expressed as ingenic_la(vr1, src2);

[0204] Jump the pointer of src2 by 64 8-bit units, expressed as src2 += 64;

[0205] Load the data of src1 into vr2, expressed as ingenic_la(vr2, src1);

[0206] Jump the pointer of src1 by 64 8-bit units, expressed as src1 += 64;

[0207] Calculate the right shift by 1, expressed as ingenic_srli_u8bit(vr4, vr1, 1);

[0208] Calculate the right shift by 3, expressed as ingenic_srli_u8bit(vr5, vr1, 3);

[0209] Right shift by 3 is calculated and expressed as ingenic_srli_u8bit(vr3, vr2, 3);

[0210] Subtraction calculation is expressed as ingenic_sub_u8bit(vr2, vr2, vr3);

[0211] Addition calculation is expressed as ingenic_add_u8bit(vr1, vr5, vr4);

[0212] Addition calculation is expressed as ingenic_add_u8bit(vr1, vr2, vr1);

[0213] Save the data in the register to dst, which is expressed as ingenic_sa(vr1, dst);

[0214] Jump the dst pointer by 64 8 - bits;

[0215] The loop body ends;

[0216] The part that cannot be divided evenly by 64 is expressed as:

[0217] Initialize the integer left, which is expressed as int left = size - i;

[0218] When it is judged that i < left, execute the following loop body, otherwise jump out of the loop body. After the loop body is executed, it is expressed as for(i = 0; i < left; i++);

[0219] The loop body starts:

[0220] The integer x is expressed as int x = src2[i];

[0221] The integer y is expressed as int y = src1[i];

[0222] The integer z is expressed as int z = x >> 1 + x >> 3 + y >> 1 - (y >> 3);

[0223] Save the picture data pointer, which is expressed as dst[t] = z;

[0224] The loop body ends;

[0225] The calculation method of the sixth fused image:

[0226] Initialize the integer i = 0;

[0227] When it is judged that i < size, the following loop body is executed; otherwise, the loop body is exited. After the loop body is executed, i = i + 64, which is expressed as for(; i < size; i += 64);

[0228] Start of the loop body:

[0229] Load data from src2 to vr1, which is expressed as ingenic_la(vr1, src2);

[0230] Jump the src2 pointer by 64 8 - bits, which is expressed as src2 += 64;

[0231] Load data from src1 to vr2, which is expressed as ingenic_la(vr2, src1);

[0232] Jump the src1 pointer by 64 8 - bits, which is expressed as src2 += 64;

[0233] Perform a right - shift by 2 calculation, which is expressed as ingenic_srli_u8bit(vr1, vr1, 2);

[0234] Perform a right - shift by 2 calculation, which is expressed as ingenic_srli_u8bit(vr3, vr2, 2);

[0235] Perform a subtraction calculation, which is expressed as ingenic_sub_u8bit(vr2, vr2, vr3);

[0236] Perform an addition calculation, which is expressed as ingenic_add_u8bit(vr1, vr2, vr1);

[0237] Save the data in the register to dst, which is expressed as ingenic_sa(vr1, dst);

[0238] Jump the dst pointer by 64 8 - bits;

[0239] End of the loop body;

[0240] For the part that cannot be divided evenly by 64, it is expressed as:

[0241] Initialize the integer left, which is expressed as int left = size - i;

[0242] When it is judged that i < left, the following loop body is executed; otherwise, the loop body is exited. After the loop body is executed, it is expressed as for(i = 0; i < left; i++);

[0243] Start of the loop body:

[0244] Integer x, which is expressed as int y = src2[i];

[0245] An integer y, represented as int y = src1[i];

[0246] An integer z, represented as int z = x >> 2 + y - (y >> 2);

[0247] Save the picture data pointer, represented as dst[i] = z;

[0248] End of the loop body;

[0249] The calculation method of the seventh fused image:

[0250] Initialize the integer i = 0;

[0251] When it is judged that i < size, execute the following loop body, otherwise jump out of the loop body. After the loop body is executed, perform i = i + 64, represented as for(; i < size; i += 64);

[0252] Start of the loop body:

[0253] Load the data of src2 into vr1, represented as ingenic_la(vr1, src2);

[0254] Jump the pointer of src2 by 64 8-bit units, represented as src2 += 64;

[0255] Load the data of src1 into vr2, represented as ingenic_la(vr2, src1);

[0256] Jump the pointer of src1 by 64 8-bit units, represented as src1 += 64;

[0257] Right shift by 3-bit calculation, represented as ingenic_srli_u8bit(vr1, vr1, 3);

[0258] Right shift by 3-bit calculation, represented as ingenic_srli_u8bit(vr3, vr2, 3);

[0259] Subtraction calculation, represented as ingenic_sub_u8bit(vr2, vr2, vr3);

[0260] Addition calculation, represented as ingenic_add_u8bit(vr1, vr2, vr1);

[0261] Save the data in the register to dst, represented as ingenic_sa(vr1, dst);

[0262] Jump the pointer of dst by 64 8-bit units;

[0263] The loop body ends;

[0264] The part that cannot be divided evenly by 64 is expressed as:

[0265] Initialize the integer left, expressed as int left = size - i;

[0266] When it is determined that i < left, execute the following loop body; otherwise, jump out of the loop body. After the loop body is executed, it is expressed as for(i = 0; i < left; i++);

[0267] The loop body starts:

[0268] The integer x, expressed as int x = src2[i];

[0269] The integer y, expressed as int y = srcl[i];

[0270] The integer z, expressed as int z = x >> 3 + y - (y >> 3);

[0271] Save the picture data pointer, expressed as dst[i] = z;

[0272] The loop body ends;

[0273] Among them, the calculation method of the seventh fused image is to swap src1 and src2 in the calculation method of the first fused image. The calculation method of the sixth fused image is to swap src1 and src2 in the calculation method of the second fused image. The calculation method of the fifth fused image is to swap src1 and src2 in the calculation method of the third fused image.

[0274] In the described method, all data is processed using 8-bit data, and no data conversion processing is performed. The simd instructions are also 8-bit related instructions.

[0275] In the described time microcosm, dividing it into 8 segments can meet the actual requirements, and the effect is the same as dividing it into 10 segments; if the coefficient of 10 segments is used, the nearest neighbor data can be processed using 8 segments.

[0276] Therefore, the advantages of this application are as follows: For this method, the amount of calculation is less, the number of simd instructions used is less, the speed is faster. Implementing other methods using simd instructions on the same chip, compared with implementing this method, the speed of this method is more than doubled. Compared with using ordinary C programs, it is increased by about 10 times. Description of the Drawings

[0277] The drawings described here are used to provide a further understanding of the present invention, constitute a part of this application, and do not constitute a limitation to the present invention.

[0278] Figure 1 It is a flow schematic diagram of this method. Specific implementation mode

[0279] In order to more clearly understand the technical content and advantages of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings.

[0280] This application provides an optimized processing method for image fusion in a time microcosm. All data in the present invention is processed using 8-bit data, without data conversion processing, and related SIMD instructions are also 8-bit related instructions. The basic coefficients used in the time microcosm are 0.125, 0.25, 0.375, 0.5, 0.625, 0.75, and 0.875. Dividing it into 8 segments can meet the actual requirements. The effect is the same as dividing it into 10 segments. If the coefficients of 10 segments are used, the nearest neighbor data can be processed using 8 segments. As Figure 1 shown, it includes the following steps:

[0281] S1. Selection and processing of coefficients.

[0282] In order to only use 8-bit data and 8-bit SIMD instructions later, the parameters are decomposed.

[0283] The basic coefficients used

[0284] 0.125, 0.25, 0.375, 0.5, 0.625, 0.75, 0.875,

[0285] The corresponding data is

[0286]

[0287] Processing of each coefficient:

[0288]

[0289]

[0290]

[0291]

[0292]

[0293]

[0294]

[0295] The coefficient of the first picture in the first fused image is Since the coefficients of the two images add up to 1, the coefficient of the other image is (i.e., ), and the coefficient is

[0296] The coefficient of the first image in the second fused image is Since the coefficients of the two images add up to 1, the coefficient of the other image is (i.e., ), and the coefficient is

[0297] The coefficient of the first image in the third fused image is (i.e., ), and since the coefficients of the two images add up to 1, the coefficient of the other image is (i.e., ), and the coefficient is

[0298] The coefficient of the first image in the fourth fused image is Since the coefficients of the two images add up to 1, the coefficient of the other image is The coefficient is

[0299] The coefficient of the first image in the fifth fused image is (i.e., ), and since the coefficients of the two images add up to 1, the coefficient of the other image is (i.e., ), and the coefficient is

[0300] The coefficient of the first image in the sixth fused image is (i.e., ), and since the coefficients of the two images add up to 1, the coefficient of the other image is The coefficient is

[0301] The coefficient of the first image in the seventh fused image is (i.e., ), and since the coefficients of the two images add up to 1, the coefficient of the other image is The coefficient is

[0302] S2. Design of the fused image

[0303] The number of data of the two images in the fused image is the same. Only the corresponding processing for each data and then adding them up is needed. Let the data of the first image in the i-th fused image be x, and the corresponding processing coefficient be xm i Let the data of the first image in the fused image be y, and the corresponding processing coefficient be ymi The fused image data is z. Among them, ym i = 1 - xm i , i ∈ [1, 7]. The calculation formula is:

[0304] z = x × xm i + y × ym i

[0305] z = x × xm i + y × ym i

[0306] S2.1, Derivation of the calculation formula

[0307] According to the calculation formula, substitute the coefficients into the formula to calculate each fused image. All calculations use addition, subtraction, and shift operations. The rounding method for shift is the downward rounding method. Since using the rounding method or the banker's rounding method may cause the result of the subsequent addition operation to be 256, for 8-bit operations, an overflow occurs, resulting in an error. Therefore, only the downward rounding method can be used.

[0308] Suppose each picture has size data;

[0309] Calculation method for the first fused image:

[0310]

[0311]

[0312] z = x >> 3 + y - (y >> 3)

[0313] Calculation method for the second fused image:

[0314]

[0315]

[0316] z = x >> 2 + y - (y >> 2)

[0317] Calculation method for the third fused image:

[0318]

[0319]

[0320] z = x >> 1 + x >> 3 + y >> 1 - (y >> 3)

[0321] Calculation method for the fourth fused image:

[0322]

[0323] z = x >> 1 + y >> 1

[0324] Calculation method of the fifth fused image:

[0325]

[0326]

[0327] z = y >> 1 + y >> 3 + x >> 1 - (x >> 3)

[0328] Calculation method of the sixth fused image:

[0329]

[0330]

[0331] z = y >> 2 + x - (x >> 2)

[0332] Calculation method of the seventh fused image:

[0333]

[0334]

[0335] z = y >> 3 + x - (x >> 3)

[0336] Among them, the calculation method of the seventh fused image is to swap x and y in the calculation method of the first fused image. The calculation method of the sixth fused image is to swap x and y in the calculation method of the second fused image. The calculation method of the fifth fused image is to swap x and y in the calculation method of the third fused image.

[0337] S2.2, Design of SIMD

[0338] Five SIMD instructions are used. Data loading instruction, data saving instruction, 8-bit addition instruction, 8-bit subtraction instruction, 8-bit shift instruction. 64 8-bit data are stored in each register. Here all the data are unsigned.

[0339] a) Instruction description, that is, the description of SIMD instructions:

[0340] Data loading instruction. Assume that the 512-bit register is vr1 and the input picture data pointer is src. This instruction is to load 512-bit data in src into vr1, ingenic_la(vr1,src);

[0341] Instruction to save data. Assume that the 512-bit register is vr1 and the pointer to save the picture data is dst. This instruction saves 512-bit data to dst, ingenic_sa(vr1,dst);

[0342] 8-bit addition instruction. Assume that the input registers are vr1 and vr2, and the register to output the addition result is vr3, ingenic_add_u8bit(vr3,vr2,vr1);

[0343] 8-bit subtraction instruction. Assume that the input registers are vr1 and vr2, and the register to output the subtraction result is vr3. The minuend is vr2, the subtrahend is vr1, and the difference is vr3, ingenic_sub_u8bit(vr3,vr2,vr1) 8-bit shift instruction. Assume that the input register is vr1, and the register to output the shift result is vr3. Each data is shifted to the right by a constant li bits, calculated using the floor method, ingenic_srli_u8bit(vr3,vr2,li);

[0344] b) simd implementation process:

[0345] Design simd according to the calculation formula. Assume that the number of data in the picture is size, the first picture is src1, load the data into vr1, the second picture is src2, load the data into vr2, and the fused image is dst;

[0346] Calculation method for the first fused image:

[0347] Initialize the integer i = 0, expressed as int i = 0;

[0348] When it is judged that i < size, execute the following loop body, otherwise jump out of the loop body. After the loop body is executed, perform i = i + 64, expressed as for(;i<size;i += 64);

[0349] Start of the loop body:

[0350] Load data from src1 into vr1, expressed as ingenic_la(vr1,src1);

[0351] Jump the src1 pointer by 64 8-bit units, expressed as src1 += 64;

[0352] Load data from src2 into vr2, expressed as ingenic_la(vr2,src2);

[0353] Jump the src2 pointer by 64 8-bit units, expressed as src2 += 64;

[0354] Right shift by 3 bits for calculation, expressed as ingenic_srli_u8bit(vr1,vr1,3);

[0355] Right shift by 3 bits for calculation, expressed as ingenic_srli_u8bit(vr3,vr2,3);

[0356] Subtraction calculation, expressed as ingenic_sub_u8bit(vr2,vr2,vr3);

[0357] Addition calculation, expressed as ingenic_add_u8bit(vr1,vr2,vr1);

[0358] Save the data in the register to dst, expressed as ingenic_sa(vr1,dst);

[0359] Jump the dst pointer by 64 8-bit units;

[0360] End of the loop body;

[0361] For the part that cannot be divided evenly by 64, it is implemented using c, expressed as:

[0362] Initialize the integer left, expressed as int left = size - i;

[0363] When judging i < left, execute the following loop body, otherwise jump out of the loop body. After the loop body is executed, it is expressed as for(i = 0; i < left; i++);

[0364] Start of the loop body:

[0365] Integer x, expressed as int x = src1[i];

[0366] Integer y, expressed as int y = src5[i];

[0367] Integer z, expressed as int z = x >> 3 + y - (y >> 3);

[0368] Save the picture data pointer, expressed as dst[i] = z;

[0369] End of the loop body;

[0370] Calculation method for the second fused image:

[0371] Initialize the integer i = 0, expressed as int i = 0;

[0372] When judging i < size, execute the following loop body, otherwise jump out of the loop body. After the loop body is executed, perform i = i + 64, expressed as for(; i < size; i += 64);

[0373] Start of the loop body:

[0374] Load data from src1 to vr1, denoted as ingenic_la(vr1,src1);

[0375] Jump the src1 pointer by 64 8-bit units, denoted as src1 += 64;

[0376] Load data from src2 to vr2, denoted as ingenic_la(vr2,src2);

[0377] Jump the src2 pointer by 64 8-bit units, denoted as src2 += 64;

[0378] Right shift by 2 bits, denoted as ingenic_srli_u8bit(vr1,vr1,2);

[0379] Right shift by 2 bits, denoted as ingenic_srli_u8bit(vr3,vr2,2);

[0380] Subtraction operation, denoted as ingenic_sub_u8bit(vr2,vr2,vr3);

[0381] Addition operation, denoted as ingenic_add_u8bit(vr1,vr2,vr1);

[0382] Save the data in the register to dst, denoted as ingenic_sa(vr1,dst);

[0383] Jump the dst pointer by 64 8-bit units;

[0384] End of the loop body;

[0385] For the part that cannot be divided evenly by 64, it can be implemented using C, denoted as:

[0386] Initialize the integer left, denoted as int left = size - i;

[0387] When i < left, execute the following loop body, otherwise jump out of the loop body. After the loop body is executed, denoted as for(i = 0; i < left; i++);

[0388] Start of the loop body:

[0389] Integer x, denoted as int x = src1[i];

[0390] Integer y, denoted as int y = src2[i];

[0391] An integer z, expressed as int z = x >> 2 + y - (y >> 2);

[0392] Save the pointer of the image data, expressed as dst[i] = z;

[0393] The loop body ends;

[0394] The calculation method of the third fused image:

[0395] Initialize the integer i = 0, expressed as int i = 0;

[0396] When it is judged that i < size, execute the following loop body, otherwise jump out of the loop body. After the loop body is executed, perform i = i + 64, expressed as for(; i < size; i += 64);

[0397] The loop body starts:

[0398] Load the data of src1 into vr1, expressed as ingenic_la(vr1, src1);

[0399] Jump the pointer of src1 by 64 8 - bits, expressed as src1 += 64;

[0400] Load the data of src2 into vr2, expressed as ingenic_la(vr2, src2);

[0401] Jump the pointer of src2 by 64 8 - bits, expressed as src2 += 64;

[0402] Right - shift 1 calculation, expressed as ingenic_srli_u8bit(vr4, vr1, 1);

[0403] Right - shift 3 calculation, expressed as ingenic_srli_u8bit(vr5, vr1, 3);

[0404] Right - shift 3 calculation, expressed as ingenic_srli_u8bit(vr3, vr2, 3);

[0405] Subtraction calculation, expressed as ingenic_sub_u8bit(vr2, vr2, vr3);

[0406] Addition calculation, expressed as ingenic_add_u8bit(vr1, vr5, vr4);

[0407] Addition calculation, expressed as ingenic_add_u8bit(vr1, vr2, vr1);

[0408] Save the data in the register to dst, denoted as ingenic_sa(vr1, dst);

[0409] Jump the dst pointer by 64 8-bit units;

[0410] End of the loop body;

[0411] For the part that cannot be divided evenly by 64, it is implemented using c, denoted as:

[0412] Initialize the integer left, denoted as int left = size - i;

[0413] When i < left, execute the following loop body, otherwise jump out of the loop body. After the loop body is executed, it is denoted as for(i = 0; i < left; i++);

[0414] Start of the loop body:

[0415] Define the integer x, denoted as int x = src1[i];

[0416] Define the integer y, denoted as int y = src2[i];

[0417] Define the integer z, denoted as int z = x >> 1 + y >> 3 + y >> 1 - (y >> 3);

[0418] Save the image data pointer, denoted as dst[t] = z;

[0419] End of the loop body;

[0420] Calculation method of the fourth fused image:

[0421] Initialize the integer i = 0;

[0422] When i < size, execute the following loop body, otherwise jump out of the loop body. After the loop body is executed, increment i by 64, denoted as for(; i < size; i += 64);

[0423] Start of the loop body:

[0424] Load data from src1 to vr1, denoted as ingenic_la(vr1, src1);

[0425] Jump the src1 pointer by 64 8-bit units, denoted as src1 += 64;

[0426] Load data from src2 to vr2, denoted as ingenic_la(vr2, src2);

[0427] Jump the src2 pointer by 64 8-bit units, denoted as src2 += 64;

[0428] Shift right by 1 calculation, expressed as ingenic_srli_u8bit(vr1, vr1, 1);

[0429] Shift right by 1 calculation, expressed as ingenic_srli_u8bit(vr2, vr2, 1);

[0430] Addition calculation, expressed as ingenic_add_u8bit(vr1, vr2, vr1);

[0431] Save the data in the register to dst, expressed as ingenic_sa(vr1, dst);

[0432] Jump the dst pointer by 64 8-bit units;

[0433] End of the loop body;

[0434] The part that cannot be divided evenly by 64 is expressed as:

[0435] Initialize the integer left, expressed as int left = size - i;

[0436] When i < left, execute the following loop body, otherwise jump out of the loop body. After the loop body is executed, it is expressed as for(i = 0; i < left; i++);

[0437] Start of the loop body:

[0438] Integer x, expressed as int x = src1[i];

[0439] Integer y, expressed as int y = src2[i];

[0440] Integer z, expressed as int z = x >> 1 + y >> 1;

[0441] Save the image data pointer, expressed as dst[i] = z;

[0442] End of the loop body;

[0443] This part of the loop is for the remaining part that cannot be divided evenly by 64, and is implemented using ordinary C programs, processing one data at a time.

[0444] Calculation method of the fifth fused image:

[0445] Initialize the integer i = 0;

[0446] When it is judged that i < size, the following loop body is executed; otherwise, the loop body is jumped out. After the loop body is executed, i = i + 64, which is expressed as for(; i < size; i += 64);

[0447] Start of the loop body:

[0448] Load data from src2 to vr1, which is expressed as ingenic_la(vr1,src2);

[0449] Jump the src2 pointer by 64 8-bit units, which is expressed as src2 += 64;

[0450] Load data from src1 to vr2, which is expressed as ingenic_la(vr2,src1);

[0451] Jump the src1 pointer by 64 8-bit units, which is expressed as src1 += 64;

[0452] Perform a right shift by 1 calculation, which is expressed as ingenic_srli_u8bit(vr4,vr1,1);

[0453] Perform a right shift by 3 calculation, which is expressed as ingenic_srli_u8bit(vr5,vr1,3);

[0454] Perform a right shift by 3 calculation, which is expressed as ingenic_srli_u8bit(vr3,vr2,3);

[0455] Perform a subtraction calculation, which is expressed as ingenic_sub_u8bit(vr2,vr2,vr3);

[0456] Perform an addition calculation, which is expressed as ingenic_add_u8bit(vr1,vr5,vr4);

[0457] Perform an addition calculation, which is expressed as ingenic_add_u8bit(vr1,vr2,vr1);

[0458] Save the data in the register to dst, which is expressed as ingenic_sa(vr1,dst);

[0459] Jump the dst pointer by 64 8-bit units;

[0460] End of the loop body;

[0461] The part that cannot be divided evenly by 64 is expressed as:

[0462] Initialize the integer left, which is expressed as int left = size - i;

[0463] When it is judged that i < left, the following loop body is executed; otherwise, the loop body is exited. After the loop body is executed, it is expressed as for(i = 0; i < left; i++);

[0464] Start of the loop body:

[0465] An integer x, expressed as int x = src2[i];

[0466] An integer y, expressed as int y = src1[i];

[0467] An integer z, expressed as int z = x >> 1 + x >> 3 + y >> 1 - (y >> 3);

[0468] Save the picture data pointer, expressed as dst[i] = z;

[0469] End of the loop body;

[0470] Calculation method of the sixth fused image:

[0471] Initialize the integer i = 0;

[0472] When it is judged that i < size, the following loop body is executed; otherwise, the loop body is exited. After the loop body is executed, i = i + 64 is performed, expressed as for(; i < size; i += 64);

[0473] Start of the loop body:

[0474] Load data from src2 to vr1, expressed as ingenic_la(vr1,src2);

[0475] Jump the src2 pointer by 64 8-bit units, expressed as src2 += 64;

[0476] Load data from src1 to vr2, expressed as ingenic_la(vr2,src1);

[0477] Jump the src1 pointer by 64 8-bit units, expressed as src2 += 64;

[0478] Right shift by 2 calculation, expressed as ingenic_srli_u8bit(vr1,vr1,2);

[0479] Right shift by 2 calculation, expressed as ingenic_srli_u8bit(vr3,vr2,2);

[0480] Subtraction calculation, expressed as ingenic_sub_u8bit(vr2,vr2,vr3);

[0481] Addition calculation, represented as ingenic_add_u8bit(vr1,vr2,vr1);

[0482] Save the data in the register to dst, represented as ingenic_sa(vr1,dst);

[0483] Jump the dst pointer by 64 8-bit;

[0484] End of the loop body;

[0485] The part that cannot be divided evenly by 64 is expressed as:

[0486] Initialize the integer left, represented as int left = size - i;

[0487] When it is judged that i < left, execute the following loop body, otherwise jump out of the loop body. After the loop body is executed, it is represented as for(i = 0; i < left; i++);

[0488] Start of the loop body:

[0489] The integer x, represented as int x = src2[i];

[0490] The integer y, represented as int y = src1[i];

[0491] The integer z, represented as int z = x >> 2 + y - (y >> 2);

[0492] Save the image data pointer, represented as dst[i] = z;

[0493] End of the loop body;

[0494] Calculation method of the seventh fused image:

[0495] Initialize the integer i = 0;

[0496] When it is judged that i < size, execute the following loop body, otherwise jump out of the loop body. After the loop body is executed, perform i = i + 64, represented as for(; i < size; i += 64);

[0497] Start of the loop body:

[0498] Load the data of src2 into vr1, represented as ingenic_la(vr1,src2);

[0499] Jump the src2 pointer by 64 8-bit, represented as src2 += 64;

[0500] Load the data of src1 into vr2, represented as ingenic_la(vr2,src1);

[0501] Jump the src1 pointer by 64 8-bit units, expressed as src1 += 64;

[0502] Right shift by 3 bits for calculation, expressed as ingenic_srli_u8bit(vr1,vr1,3);

[0503] Right shift by 3 bits for calculation, expressed as ingenic_srli_u8bit(vr3,vr2,3);

[0504] Subtraction calculation, expressed as ingenic_sub_u8bit(vr2,vr2,vr3);

[0505] Addition calculation, expressed as ingenic_add_u8bit(vr1,vr2,vr1);

[0506] Save the data in the register to dst, expressed as ingenic_sa(vr1,dst);

[0507] Jump the dst pointer by 64 8-bit units;

[0508] End of the loop body;

[0509] The part that cannot be divided evenly by 64 is expressed as:

[0510] Initialize the integer left, expressed as int left = size - i;

[0511] When i < left, execute the following loop body, otherwise jump out of the loop body. After the loop body is executed, it is expressed as for(i = 0; i < left; i++);

[0512] Start of the loop body:

[0513] Integer x, expressed as int x = src2[i];

[0514] Integer y, expressed as int y = src1[i];

[0515] Integer z, expressed as int z = x >> 3 + y - (y >> 3);

[0516] Save the picture data pointer, expressed as dst[i] = z;

[0517] End of the loop body;

[0518] The calculation method of the seventh fused image is to swap src1 and src2 in the calculation method of the first fused image. The calculation method of the sixth fused image is to swap src1 and src2 in the calculation method of the second fused image. The calculation method of the fifth fused image is to swap src1 and src2 in the calculation method of the third fused image.

[0519] It is used in the Junzheng t41 model chip for fusion processing of 720p NV12 image data, and each time it takes 1.4 ms, which fully meets the requirements of the actual scenario. This method is suitable for any chip acceleration processing. In particular, using simd-class chips, the acceleration is more obvious.

[0520] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, various changes and modifications can be made to the embodiments of the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. An optimized processing method for image fusion in the time microcosm, characterized in that, The method includes the following steps: S1. Coefficient selection and processing: Using 8-bit data and 8-bit SIMD instructions, decompose the parameters; The basic coefficients adopted are: 0.125,0.25,0.375,0.5,0.625,0.75,0.875; Therefore, the corresponding data is: Processing of each coefficient: where n = 1, 2, 3, 4, 5, 6, 7; The first picture coefficient in the nth fused image is Since the sum of the coefficients of the two pictures is 1, the coefficient of the other picture is The coefficient is S2. Design of the fused image: The number of data in the two pictures of the fused image is the same. Just process each corresponding data and then add them up; Let the first image data in the $i$-th fused image be $x$, and the corresponding processing coefficient be $x_m$. i Let the first image data of the fused image be $y$, and the corresponding processing coefficient be $y_m$. i ; The fused image data is $z$; where $y_m$ i $= 1 - x_m$ i , $i\in[1, 7]$; The calculation formula is: z = x × xm i + y × ym i ; z = x × xm i + y × ym i ; S2.

1. Deduction of the calculation formula: According to the calculation formula, substitute the coefficients into the formula to calculate each fused image; all calculations use addition, subtraction, and shift operations. The rounding method used for shifting is the floor method; Assume that each picture has size data; Calculation method for the i-th fused image: S2.

2. Design of SIMD: Five SIMD instructions are used: data loading instruction, data saving instruction, 8-bit addition instruction, 8-bit subtraction instruction, 8-bit shift instruction; 64 8-bit data are stored in each register; all here are unsigned data; SIMD implementation process: Design SIMD according to the calculation formula; Assume that the number of data in the picture is size, the first picture is src1, load the data into vr1, the second picture is src2, load the data into vr2, and the fused image is dst; Calculation method for the first fused image: z = x >> 3 + y - (y >> 3); Calculation method for the second fused image: z = x >> 2 + y - (y >> 2); Calculation method for the third fused image: z = x >> 1 + x >> 3 + y >> 1 - (y >> 3); Calculation method for the fourth fused image: z = x >> 1 + y >> 1; Calculation method for the fifth fused image: z = y >> 1 + y >> 3 + x >> 1 - (x >> 3); Calculation method for the sixth fused image: z = y >> 2 + x - (x >> 2); Calculation method for the seventh fused image: z = y >> 3 + x - (x >> 3).

2. The optimized processing method for image fusion in a time-lapse microcosm according to claim 1, wherein In step S1, it further includes: Processing of each coefficient: The first picture coefficient in the first fused image is Since the sum of the coefficients of the two pictures is 1, the coefficient of the other picture is That is The coefficient is The first picture coefficient in the second fused image is Since the sum of the two picture coefficients is 1, the other picture coefficient is That is The coefficient is The first picture coefficient in the third fused image is That is Since the sum of the coefficients of the two pictures is 1, the coefficient of the other picture is That is The coefficient is The first picture coefficient in the fourth fused image is Since the sum of the coefficients of the two pictures is 1, the coefficient of the other picture is The coefficient is The first picture coefficient in the fifth fused image is That is Since the sum of the coefficients of the two pictures is 1, the coefficient of the other picture is That is The coefficient is The first picture coefficient in the sixth fused image is That is Since the sum of the coefficients of the two pictures is 1, the coefficient of the other picture is The coefficient is The first picture coefficient in the seventh fused image is That is Since the sum of the coefficients of the two pictures is 1, the coefficient of the other picture is The coefficient is 3. The optimized processing method for image fusion in a time miniature according to claim 1, characterized in that In step S2.1, Calculation method for the first fused image: z = x >> 3 + y - (y >> 3) Calculation method for the second fused image: z = x >> 2 + y - (y >> 2) Calculation method for the third fused image: z = x >> 1 + x >> 3 + y >> 1 - (y >> 3) Calculation method for the fourth fused image: z = x >> 1 + y >> 1 Calculation method for the fifth fused image: z = y >> 1 + y >> 3 + x >> 1 - (x >> 3) Calculation method for the sixth fused image: z = y >> 2 + x - (x >> 2) Calculation method for the seventh fused image: z = y >> 3 + x - (x >> 3); Among them, the calculation method for the seventh fused image is to swap x and y in the calculation method for the first fused image; the calculation method for the sixth fused image is to swap x and y in the calculation method for the second fused image; the calculation method for the fifth fused image is to swap x and y in the calculation method for the third fused image.

4. An optimized processing method for image fusion in a time-lapse microcosm according to claim 1, characterized in that, In step S2.2, the description of the SIMD instructions: Load data instruction. Set the 512-bit register as vr1 and the input image data pointer as src. This instruction loads 512-bit data from src into vr1: ingenic_la(vr1,src); Save data instruction. Set the 512-bit register as vr1 and the save image data pointer as dst. This instruction saves 512-bit data into dst: ingenic_sa(vr1,dst); 8-bit addition instruction. Set the input registers as vr1 and vr2, and the addition calculation result output register as vr3: ingenic_add_u8bit(vr3,vr2,vr1); 8-bit subtraction instruction. Set the input registers as vr1 and vr2, and the subtraction calculation result output register as vr3. The minuend is vr2, the subtrahend is vr1, and the difference is vr3: ingenic_sub_u8bit(vr3,vr2,vr1); 8-bit shift instruction. Set the input register as vr1, and the shift calculation result output register as vr3. Each data is shifted to the right by a constant li bits, and the floor function is used for calculation: ingenic_srli_u8bit(vr3,vr2,li).

5. The optimized processing method for image fusion in a time miniature according to claim 4, wherein The Calculation method of the first fused image: Initialize the integer i = 0, expressed as int i = 0; When it is judged that i < size, execute the following loop body, otherwise jump out of the loop body. After the loop body is executed, perform i = i + 64, expressed as for(; i < size; i += 64); Start of the loop body: Load data from src1 into vr1, expressed as ingenic_la(vr1,src1); Jump the src1 pointer by 64 8-bits, expressed as src1 += 64; Load data from src2 into vr2, expressed as ingenic_la(vr2,src2); Jump the src2 pointer by 64 8-bits, expressed as src2 += 64; Perform a right shift calculation by 3 bits, expressed as ingenic_srli_u8bit(vr1,vr1,3); Perform a right shift calculation by 3 bits, expressed as ingenic_srli_u8bit(vr3,vr2,3); Perform a subtraction calculation, expressed as ingenic_sub_u8bit(vr2,vr2,vr3); Perform an addition calculation, expressed as ingenic_add_u8bit(vr1,vr2,vr1); Save the data in the register into dst, expressed as ingenic_sa(vr1,dst); Jump the dst pointer by 64 8-bits; End of the loop body; For the part that cannot be divided evenly by 64, it is expressed as: Initialize the integer left, expressed as int left = size - i; When it is judged that i < left, execute the following loop body, otherwise jump out of the loop body. After the loop body is executed, it is expressed as for(i = 0; i < left; i++); Start of the loop body: An integer x, expressed as int x = src1[i]; An integer y, expressed as int y = src2[i]; An integer z, expressed as int z = x >> 3 + y - (y >> 3); Save the pointer of the image data, expressed as dst[i] = z; The loop body ends; This part of the loop is for the data that cannot be divided evenly by 64, and the remaining part is implemented using ordinary C programs, processing one data each time; Calculation method for the second fused image: Initialize the integer i = 0, expressed as int i = 0; When it is judged that i < size, execute the following loop body, otherwise jump out of the loop body. After the loop body is executed, perform i = i + 64, expressed as: for(; i < size; i += 64) The loop body starts: Load the data of src1 into vr1, expressed as ingenic_la(vr1, src1); Jump the pointer of src1 by 64 8-bit units, expressed as src1 += 64; Load the data of src2 into vr2, expressed as ingenic_la(vr2, src2); Jump the pointer of src2 by 64 8-bit units, expressed as src2 += 64; Right shift by 2 calculation, expressed as ingenic_srli_u8bit(vr1, vr1, 2); Right shift by 2 calculation, expressed as ingenic_srli_u8bit(vr3, vr2, 2); Subtraction calculation, expressed as ingenic_sub_u8bit(vr2, vr2, vr3); Addition calculation, expressed as ingenic_add_u8bit(vr1, vr2, vr1); Save the data in the register to dst, expressed as ingenic_sa(vr1, dst); Jump the pointer of dst by 64 8-bit units; The loop body ends; The part that cannot be divided evenly by 64 is expressed as: Initialize the integer left, expressed as int left = size - i; When it is judged that i < left, execute the following loop body, otherwise jump out of the loop body. After the loop body is executed, it is expressed as for(i = 0; i < left; i++); The loop body starts: An integer x, expressed as int x = src1[i]; An integer y, expressed as int y = src2[i]; An integer z, expressed as int z = x >> 2 + y - (y >> 2); Save the pointer of the image data, expressed as dxt[i] = z; The loop body ends; This part of the loop is for the data that cannot be divided evenly by 64, and the remaining part is implemented using ordinary C programs, processing one data each time; Calculation method for the third fused image: Initialize the integer i = 0; When it is judged that i < size, execute the following loop body, otherwise jump out of the loop body. After the loop body is executed, perform i = i + 64, expressed as for(; i < size; i += 64); The loop body starts: Load the data of src1 into vr1, expressed as ingenic_la(vr1, src1); Jump the pointer of src1 by 64 8-bit units, expressed as src1 += 64; Load data from src2 to vr2, denoted as ingenic_la(vr2,src2); Jump the src2 pointer by 64 8-bit units, denoted as src2 += 64; Right shift by 1 calculation, denoted as ingenic_srli_u8bit(vr4,vr1,1); Right shift by 3 calculation, denoted as ingenic_srli_u8bit(vr5,vr1,3); Right shift by 3 calculation, denoted as ingenic_srli_u8bit(vr3,vr2,3); Subtraction calculation, denoted as ingenic_sub_u8bit(vr2,vr2,vr3); Addition calculation, denoted as ingenic_add_u8bit(vr1,vr5,vr4); Addition calculation, denoted as ingenic_add_u8bit(vr1,vr2,vr1); Save the data in the register to dst, denoted as ingenic_sa(vr1,dst); Jump the dst pointer by 64 8-bit units; The loop body ends; The part that cannot be divided evenly by 64 is denoted as: Initialize the integer left, denoted as int left = size - i; When i < left, execute the following loop body, otherwise jump out of the loop body. After the loop body is executed, it is denoted as for(i = 0; i < left; i++); The loop body starts: The integer x, denoted as int x = src1[i]; The integer y, denoted as int y = src2[i]; The integer z, denoted as int z = x >> 1 + x >> 3 + y >> 1 - (y >> 3); Save the image data pointer, denoted as dxt[i] = z; The loop body ends; The calculation method of the fourth fused image: Initialize the integer i = 0; When i < size, execute the following loop body, otherwise jump out of the loop body. After the loop body is executed, perform i = i + 64, denoted as for(; i < size; i += 64); The loop body starts: Load data from src1 to vr1, denoted as ingenic_la(vr1,src1); Jump the src1 pointer by 64 8-bit units, denoted as src1 += 64; Load data from src2 to vr2, denoted as ingenic_la(vr2,src2); Jump the src2 pointer by 64 8-bit units, denoted as src2 += 64; Right shift by 1 calculation, denoted as ingenic_srli_u8bit(vr1,vr1,1); Right shift by 1 calculation, denoted as ingenic_srli_u8bit(vr2,vr2,1); Addition calculation, denoted as ingenic_add_u8bit(vr1,vr2,vr1); Save the data in the register to dst, denoted as ingenic_sa(vr1,dst); Jump the dst pointer by 64 8-bit units; The loop body ends; The part that cannot be divided evenly by 64 is expressed as: Initialize the integer left, expressed as int left = size - i; When i < left is judged, execute the following loop body, otherwise jump out of the loop body. After the loop body is executed, it is expressed as for(i = 0; i < left; i++); Start of the loop body: The integer x is expressed as int x = src1[i]; The integer y is expressed as int y = src2[i]; The integer z is expressed as int z = x >> 1 + y >> 1; Save the picture data pointer, expressed as dxt[i] = z; End of the loop body; This part of the loop is for the part that cannot be divided evenly by 64. The remaining part is implemented using ordinary C programs, and one data is processed each time; Calculation method of the fifth fused image: Initialize the integer i = 0; When i < size is judged, execute the following loop body, otherwise jump out of the loop body. After the loop body is executed, perform i = i + 64, expressed as for(; i < size; i += 64); Start of the loop body: Load the data of src2 into vr1, expressed as ingenic_la(vr1, src2); Jump the src2 pointer by 64 8-bit units, expressed as src2 += 64; Load the data of src1 into vr2, expressed as ingenic_la(vr2, src1); Jump the src1 pointer by 64 8-bit units, expressed as src1 += 64; Right shift by 1 calculation, expressed as ingenic_srli_u8bit(vr4, vr1, 1); Right shift by 3 calculation, expressed as ingenic_srli_u8bit(vr5, vr1, 3); Right shift by 3 calculation, expressed as ingenic_srli_u8bit(vr3, vr2, 3); Subtraction calculation, expressed as ingenic_sub_u8bit(vr2, vr2, vr3); Addition calculation, expressed as ingenic_add_u8bit(vr1, vr5, vr4); Addition calculation, expressed as ingenic_add_u8bit(vr1, vr2, vr1); Save the data in the register to dst, expressed as ingenic_sa(vr1, dst); Jump the dst pointer by 64 8-bit units; End of the loop body; The part that cannot be divided evenly by 64 is expressed as: Initialize the integer left, expressed as int left = size - i; When i < left is judged, execute the following loop body, otherwise jump out of the loop body. After the loop body is executed, it is expressed as for(i = 0; i < left; i++); Start of the loop body: The integer x is expressed as int x = src2[i]; The integer y is expressed as int y = src1[i]; The integer z is expressed as int z = x >> 1 + x >> 3 + y >> 1 - (y >> 3); Save the picture data pointer, expressed as dt[i] = z; End of the loop body; Calculation method of the sixth fused image: Initialize the integer i = 0; When it is judged that i < size, the following loop body is executed; otherwise, the loop body is jumped out. After the loop body is executed, i = i + 64, which is expressed as for(; i < size; i += 64); The loop body starts: Load data from src2 to vr1, which is expressed as ingenic_la(vr1, src2); Jump the src2 pointer by 64 8-bit, which is expressed as src2 += 64; Load data from src1 to vr2, which is expressed as ingenic_la(vr2, src1); Jump the src1 pointer by 64 8-bit, which is expressed as src2 += 64; Perform a right shift by 2 calculation, which is expressed as ingenic_srli_u8bit(vr1, vr1, 2); Perform a right shift by 2 calculation, which is expressed as ingenic_srli_u8bit(vr3, vr2, 2); Perform a subtraction calculation, which is expressed as ingenic_sub_u8bit(vr2, vr2, vr3); Perform an addition calculation, which is expressed as ingenic_add_u8bit(vr1, vr2, vr1); Save the data in the register to dst, which is expressed as ingenic_sa(vr1, dst); Jump the dst pointer by 64 8-bit; The loop body ends; For the part that cannot be divided evenly by 64, it is expressed as: Initialize the integer left, which is expressed as int left = size - i; When it is judged that i < left, the following loop body is executed; otherwise, the loop body is jumped out. After the loop body is executed, it is expressed as for(i = 0; i < left; i++); The loop body starts: The integer x, which is expressed as int x = src2[i]; The integer y, which is expressed as int y = src1[i]; The integer z, which is expressed as int z = x >> 2 + y - (y >> 2); Save the picture data pointer, which is expressed as dst[i] = z; The loop body ends; The calculation method of the seventh fused image: Initialize the integer i = 0; When it is judged that i < size, the following loop body is executed; otherwise, the loop body is jumped out. After the loop body is executed, i = i + 64, which is expressed as for(; i < size; i += 64); The loop body starts: Load data from src2 to vr1, which is expressed as ingenic_la(vr1, src2); Jump the src2 pointer by 64 8-bit, which is expressed as src2 += 64; Load data from src1 to vr2, which is expressed as ingenic_la(vr2, src1); Jump the src1 pointer by 64 8-bit, which is expressed as src1 += 64; Perform a right shift by 3 calculation, which is expressed as ingenic_srli_u8bit(vr1, vr1, 3); Perform a right shift by 3 calculation, which is expressed as ingenic_srli_u8bit(vr3, vr2, 3); Perform a subtraction calculation, which is expressed as ingenic_sub_u8bit(vr2, vr2, vr3); Addition calculation, expressed as ingenic_add_u8bit(vr1,vr2,vr1); Save the data in the register to dst, expressed as ingenic_sa(vr1,dst); Jump the dst pointer by 64 8-bit units; End of the loop body; The part that cannot be divided evenly by 64 is expressed as: Initialize the integer left, expressed as int left = size - i; When it is judged that i < left, execute the following loop body, otherwise jump out of the loop body. After the loop body is executed, it is expressed as for(i = 0; i < left; i++); Start of the loop body: Integer x, expressed as int x = src2[i]; Integer y, expressed as int y = src1[i]; Integer z, expressed as int z = x >> 3 + y - (y >> 3); Save the picture data pointer, expressed as dxt[i] = z; End of the loop body; Among them, the calculation method of the seventh fused image is to swap src1 and src2 in the calculation method of the first fused image; the calculation method of the sixth fused image is to swap src1 and src2 in the calculation method of the second fused image; the calculation method of the fifth fused image is to swap src1 and src2 in the calculation method of the third fused image.

6. The optimized processing method for image fusion in a time microcosm according to claim 1, characterized in that All data in the method are processed using 8-bit data, and no data conversion processing is performed. The simd instructions are also 8-bit related instructions.

7. An optimized processing method for image fusion in a time-lapse microcosm according to claim 1, characterized in that, In the time microcosm described, dividing it into 8 segments can meet the actual requirements, and the effect is the same as dividing it into 10 segments; if the coefficient of 10 segments is used, the nearest neighbor data can be processed with 8 segments.