Local window based linear dynamic compensation solving method, device and electronic equipment

CN122820522APending Publication Date: 2026-09-25ZHEJIANG YUANSUAN TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611309193.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-27
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

在待处理数据和参考数据具有较高分辨率时,内积调用次数极多,造成巨大的计算量

Benefits of technology

本发明提供了一种基于局部窗口的线性动态补偿求解方法、装置和电子设备,方法包括:获取待处理数据和多个参考数据,并基于预设的局部窗口模板提取局部窗口;其中,局部窗口包括对应于待处理数据的标准局部窗口和多个对应于参考数据的参考局部窗口;基于标准局部窗口与参考局部窗口的内积结果构建第一矩阵;基于参考局部窗口两两之间的内积结果构建第二矩阵;其中,构建第二矩阵的过程包括:按照预设的遍历顺序依次选取参考局部窗口对,遍历顺序被配置为使得在选取序列中至少有一组相邻的参考局部窗口对在计算内积时具有重叠区域,并且,对于后一参考局部窗口对,基于前一参考局部窗口对的内积结果,通过增量计算方法得到当前参考局部窗口对的内积结果;基于第一矩阵和第二矩阵求解线性方程组,得到用于对待处理数据进行局部线性动态补偿的补偿参数;通过特定的遍历顺序和增量计算手段,显著减少冗余内积运算,提高求解效率。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122820522A_ABST
    Figure CN122820522A_ABST
Patent Text Reader

Abstract

The application provides a local window-based linear dynamic compensation solving method and device and electronic equipment, and relates to the technical field of image processing. In constructing a second matrix, a traversal order is configured so that there is an overlapping area between adjacent reference local window pairs in a selected sequence. On this basis, instead of independently calculating all inner products of the current reference local window pair, the current result is obtained through incremental calculation based on the inner product result of the previous reference local window pair. Since most of the pixel areas between adjacent window pairs are overlapping, the incremental calculation only needs to process a small amount of pixel contribution that moves in and out, without repeated multiplication and addition of the overlapping area, thereby greatly reducing the number of multiplication and addition operations and the amount of data reading required for constructing the second matrix, effectively reducing the overall computing cost, and achieving acceleration of the local window-based linear dynamic compensation solving process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image processing technology, and in particular to a linear dynamic compensation solution method, apparatus, and electronic device based on local windows. Background Technology

[0002] In scenarios such as industrial inspection, image stitching, brightness compensation, local field value correction, and CAE result field consistency processing, it is often necessary to solve a set of compensation parameters based on the local correlation between the data to be processed and multiple reference data to reduce the numerical differences between adjacent regions or adjacent data blocks.

[0003] Taking a typical linear dynamic compensation solution process as an example, existing technologies usually include the following steps: First, read the data to be processed and extract a target data region of a specified size; then read multiple reference data to form a reference data set; next, define a local window template and extract local windows from the standard region and the reference region; then construct a first matrix and a second matrix for linear solution, where the first matrix is ​​composed of the inner product results between the standard local window and each reference local window, and the second matrix is ​​composed of the inner product results between all pairs of reference local windows; finally, obtain the compensation parameters by solving the linear equation system, and output or save the compensation parameters for subsequent brightness correction, defect identification assistance, or field value consistency processing, etc.

[0004] The main time-consuming aspect of existing technologies lies in the calculation of inner products for a large number of local windows. When constructing the second matrix, a large-scale pairwise inner product is required for multiple local windows. For example, for R reference local windows, R×R inner products need to be executed, with each inner product independently calculating the multiplication and summation of all elements within the window. When the data to be processed and the reference data have high resolution, the number of inner product calls is extremely high, resulting in a huge computational load. Furthermore, existing technologies typically employ an independent calculation method for each element when performing inner products, failing to consider the overlapping areas between the reference local windows. This leads to a large amount of pixel data being repeatedly read and multiplied / accumulated, wasting computational resources and increasing memory access overhead, making it difficult to meet real-time requirements. Summary of the Invention

[0005] The purpose of this invention is to provide a linear dynamic compensation solution method, apparatus, and electronic device based on a local window, which significantly reduces redundant inner product operations and improves solution efficiency through a specific traversal order and incremental calculation method.

[0006] In a first aspect, the present invention provides a linear dynamic compensation solution method based on a local window, comprising: The system acquires the data to be processed and multiple reference data, and extracts local windows based on a preset local window template; wherein, the local window includes a standard local window corresponding to the data to be processed and multiple reference local windows corresponding to the reference data; The first matrix is ​​constructed based on the inner product of the standard local window and the reference local window; A second matrix is ​​constructed based on the inner product results between each pair of reference local windows. The process of constructing the second matrix includes: selecting reference local window pairs in a preset traversal order, the traversal order being configured such that at least one pair of adjacent reference local window pairs in the selection sequence has an overlapping region when calculating the inner product; and, for the next reference local window pair, the inner product result of the current reference local window pair is obtained by incremental calculation based on the inner product result of the previous reference local window pair. Solving the system of linear equations based on the first and second matrices yields compensation parameters used for local linear dynamic compensation of the data to be processed.

[0007] In some preferred embodiments of the present invention, a window relationship matrix is ​​pre-constructed; wherein, the window relationship matrix is ​​determined based on the order of extracting local windows; the traversal order is to start from the upper right corner of the window relationship matrix and traverse the window relationship matrix along the diagonal from the upper left to the lower right.

[0008] In some preferred embodiments of the present invention, extracting a local window based on a preset local window template includes: First, fix the current data, then traverse all local windows within the current data. After extracting all local windows, switch to the next data and perform the same processing until all data has been processed.

[0009] In some preferred embodiments of the present invention, after extracting the local window, the method further includes: Pre-calculate and cache the local region index, rectangular range, and / or sub-block description information of the local window in the reference data; During the construction of the first and second matrices, the pre-calculated and cached local region indices, rectangular ranges, and / or sub-block description information are directly used.

[0010] In some preferred embodiments of the present invention, for a subsequent reference local window pair, the inner product result of the current reference local window pair is obtained by incremental calculation based on the inner product result of the previous reference local window pair, including: When the next reference local window pair is translated relative to the previous reference local window pair, the inner product result of the previous reference local window pair is subtracted from the cumulative contribution of the pixel values ​​that moved out of the overlapping area due to the translation, and then the cumulative contribution of the pixel values ​​that newly entered the overlapping area due to the translation is added.

[0011] In some preferred embodiments of the present invention, constructing the second matrix based on the inner product results between each pair of reference local windows further includes: Calculate the upper triangular part or the lower left triangular part of the second matrix, and complete the second matrix according to the symmetry relationship.

[0012] In some preferred embodiments of the present invention, the method further includes: The inner product result is calculated using Single Instruction Multiple Data (SIMD) vectorized instructions. Both the data to be processed and the reference data are in 8-bit unsigned integer format. The vectorized instructions perform parallel data processing for the 8-bit unsigned integer format.

[0013] In some preferred embodiments of the present invention, the method further includes: The data to be processed, the reference data, the first matrix, and the second matrix are divided and calculated into sub-blocks of a preset size.

[0014] Secondly, the present invention provides a linear dynamic compensation solution device based on a local window, comprising: The data processing module is used to acquire the data to be processed and multiple reference data, and extract local windows based on a preset local window template; wherein, the local window includes a standard local window corresponding to the data to be processed and multiple reference local windows corresponding to the reference data; The first matrix construction module is used to construct the first matrix based on the inner product of the standard local window and the reference local window; The second matrix construction module is used to construct a second matrix based on the inner product results between each pair of reference local windows. The process of constructing the second matrix includes: selecting reference local window pairs in a preset traversal order, the traversal order being configured such that at least one pair of adjacent reference local window pairs in the selection sequence has an overlapping region when calculating the inner product; and, for the next reference local window pair, obtaining the inner product result of the current reference local window pair by an incremental calculation method based on the inner product result of the previous reference local window pair. The compensation parameter determination module is used to solve the linear equation system based on the first matrix and the second matrix to obtain the compensation parameters used for local linear dynamic compensation of the data to be processed.

[0015] Thirdly, the present invention provides an electronic device including a processor and a memory, the memory storing computer-executable instructions executable by the processor, the processor executing the computer-executable instructions to implement the method provided in the first aspect above.

[0016] This invention brings the following beneficial effects: This invention provides a method, apparatus, and electronic device for solving linear dynamic compensation based on local windows. The method includes: acquiring data to be processed and multiple reference data, and extracting local windows based on a preset local window template; wherein the local window includes a standard local window corresponding to the data to be processed and multiple reference local windows corresponding to the reference data; constructing a first matrix based on the inner product results of the standard local window and the reference local windows; constructing a second matrix based on the inner product results between each pair of reference local windows; wherein the process of constructing the second matrix includes: sequentially selecting pairs of reference local windows according to a preset traversal order, the traversal order being configured such that at least one pair of adjacent reference local window pairs in the selection sequence has an overlapping region when calculating the inner product, and, for a subsequent pair of reference local windows, obtaining the inner product result of the current pair of reference local windows based on the inner product result of the previous pair of reference local windows through an incremental calculation method; solving a system of linear equations based on the first and second matrices to obtain compensation parameters for performing local linear dynamic compensation on the data to be processed; through a specific traversal order and incremental calculation method, redundant inner product operations are significantly reduced, and the solution efficiency is improved. Attached Figure Description

[0017] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0018] Figure 1 This is a flowchart of the linear dynamic compensation solution for industrial inspection images in existing technologies; Figure 2 A flowchart illustrating a linear dynamic compensation solution method based on a local window, provided in an embodiment of the present invention; Figure 3 This is a schematic diagram illustrating an optimized traversal order according to an embodiment of the present invention. Figure 4 This is a schematic diagram illustrating the principle of local window overlap and incremental calculation provided in an embodiment of the present invention; Figure 5 A schematic diagram illustrating a method for reducing computational load by utilizing the symmetry of a second matrix, as provided in an embodiment of the present invention; Figure 6 A schematic diagram illustrating a block-based computation method provided in an embodiment of the present invention; Figure 7 A schematic diagram of a linear dynamic compensation solution device based on a local window provided in an embodiment of the present invention; Figure 8 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention.

[0019] Icons: 310 - Data processing module; 320 - First matrix construction module; 330 - Second matrix construction module; 340 - Compensation parameter determination module; 400 - Memory; 401 - Processor; 402 - Bus; 403 - Communication interface. Detailed Implementation

[0020] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.

[0021] Therefore, the following detailed description of the embodiments of the invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the invention without inventive effort are within the scope of protection of the invention.

[0022] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.

[0023] In the description of this invention, it should be noted that the terms "center," "upper," "lower," "left," "right," "vertical," "horizontal," "inner," and "outer," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings, or the orientation or positional relationship commonly used when the product of this invention is in use. They are only for the convenience of describing this invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of this invention. In addition, the terms "first," "second," "third," etc., are only used to distinguish descriptions and should not be construed as indicating or implying relative importance.

[0024] Furthermore, terms such as "horizontal," "vertical," and "sag" do not imply that components must be absolutely horizontal or suspended, but rather that they can be slightly tilted. For example, "horizontal" simply means that its direction is more horizontal relative to "vertical," and does not mean that the structure must be completely horizontal, but can be slightly tilted.

[0025] In the description of this invention, it should also be noted that, unless otherwise explicitly specified and limited, the terms "set," "install," "connect," and "link" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances.

[0026] See Figure 1 The flowchart shown in the prior art for solving the linear dynamic compensation of industrial inspection images is as follows: Input the industrial inspection image to be processed (2k×2k, U8) → Read multiple reference inspection images and crop the 2k×2k target region → Reduce the boundary by 2 pixels and construct a mask matrix → Determine the 5×5 local window template → Use the input image and mask matrix to construct a standard window → Extract reference windows iteratively based on the 5×5 local window → Construct matrix B through a large number of inner products → Construct matrix A through the pairwise inner products between reference windows → Solve the linear equation system → Output the compensation parameters (for image correction).

[0027] From an implementation perspective, the main time consumption of existing solutions is concentrated in the large-scale local window inner product calculation. For industrial inspection image scenarios (2k×2k image, 4 reference images, 5×5 window, 10 iterations), according to existing tests, the overall time consumption is approximately 4617 ms to 5066 ms, of which more than 98% of the time consumption is concentrated in the cv::Mat::dot() function. Although the existing technology has a clear computational logic, there is still significant room for optimization in terms of the number of inner product calls, data reuse methods, access order, and vectorization implementation. Furthermore, it does not have optimization strategies designed for the high resolution, U8 pixels, and uniform local texture of industrial inspection images, and cannot meet the latency requirements of real-time industrial inspection. If there are 4 reference images, constructing matrix B requires 4×(5×5)=100 inner product operations. Constructing matrix A requires (4×(5×5))×(4×(5×5))=100,000 inner product operations. When constructing matrix A, since r1 and r2 are both taken from the reference graph, they are completely equivalent. Therefore, the resulting matrix A is symmetric.

[0028] Based on this, embodiments of the present invention provide a linear dynamic compensation solution method, apparatus, and electronic device based on a local window, which significantly reduces redundant inner product operations and improves solution efficiency through a specific traversal order and incremental calculation method.

[0029] The following detailed description of some embodiments of the present invention is provided in conjunction with the accompanying drawings. Unless otherwise specified, the following embodiments and features can be combined with each other.

[0030] This invention provides a linear dynamic compensation solution method based on a local window, see [link to relevant documentation]. Figure 2 The flowchart shown in this embodiment of the invention provides a linear dynamic compensation solution method based on a local window. The method includes: Step S102: Obtain the data to be processed and multiple reference data, and extract local windows based on a preset local window template; wherein, the local window includes a standard local window corresponding to the data to be processed and multiple reference local windows corresponding to the reference data.

[0031] Specifically, the data to be processed refers to the target data object that needs to undergo local linear dynamic compensation correction, and the reference data refers to the auxiliary data object used to establish a local correlation with the data to be processed and to calculate compensation parameters. In this embodiment of the invention, the data to be processed and the reference data can be in various forms.

[0032] In the first implementation, the data to be processed is a single frame of industrial inspection image, and the reference data consists of multiple frames of reference industrial inspection images. These image data were all captured by industrial cameras under identical or similar conditions, possessing high resolution and an 8-bit unsigned integer (U8) pixel format, primarily reflecting the grayscale, texture, and defect information of the industrial parts' surfaces. Due to variations in lighting conditions, imaging angles, and other factors, local brightness differences or grayscale deviations may exist between the image to be processed and the reference images, requiring the calculation of a set of local compensation parameters for correction. This scenario places strict requirements on processing speed, needing to meet the latency requirements of real-time industrial inspection.

[0033] In the second implementation, the data to be processed is the result field data under a specific working condition obtained from finite element simulation, and the reference data is the reference result field data under multiple different working conditions. The result field can be a two-dimensional numerical matrix formed by expanding scalar or vector fields such as temperature field, stress field, strain field, and displacement field. The data format can be single-precision floating-point number (FP32) or double-precision floating-point number (FP64). The purpose of this embodiment is to establish a local mapping relationship between the target working condition and multiple sets of reference working conditions by using the result field under the target working condition as a benchmark and through local window linear dynamic compensation. The obtained compensation parameters can be used to predict the field distribution under the new working condition, perform consistency fusion of multiple sets of simulation results, or detect abnormal local areas in the simulation results.

[0034] After acquiring the data to be processed and multiple reference data, the target processing region in each data set needs to be determined. In image processing scenarios, the boundaries of the original image are typically reduced appropriately to eliminate interference from edge artifacts. A preset local window template refers to predefined shape and size parameters used to extract local sub-regions from the data matrix. The preset local window template is preferably a rectangular sliding window, and the window size can be determined based on the spatial resolution of the data to be processed, the scale of the target features, and the constraints of computational resources.

[0035] Local window extraction is performed based on this local window template. Standard local windows are extracted from the data to be processed and used as a benchmark for subsequent compensation and correction; there can be one or more of these. Reference local windows are extracted from each reference data point, typically by sliding the window across the entire target processing region. Starting from the top left corner of the target processing region, the window template is slid row by row and column by column with a preset step size, extracting one reference local window per slide. The total number of reference local windows depends on the size of the target processing region, the window template size, and the sliding step size.

[0036] Step S104: Construct the first matrix based on the inner product of the standard local window and the reference local window.

[0037] Specifically, the first matrix (which can be denoted as matrix B) is used to characterize the correlation between the standard local window and each reference local window. The construction process is as follows: Let the standard local window be a vector s containing M elements, where M is the product of the width and height of the window template. For each reference local window, let the i-th reference local window be a vector ri containing M elements, and calculate the inner product between this reference local window and the standard local window. This yields a scalar value. The inner product scalar values ​​corresponding to all reference local windows are then arranged sequentially according to the extraction order of the reference local windows, forming the first matrix B. The first matrix B is a column vector with the number of rows equal to the total number of reference local windows.

[0038] In the inner product calculation process, the numerical types of the data to be processed and the reference data determine the data type of the multiplication-accumulation operation. In industrial inspection image scenarios, pixel values ​​are in U8 format, and the inner product involves integer multiplication-accumulation operations; in CAE result field scenarios, the data is in FP32 or FP64 format, and the inner product involves floating-point multiplication-accumulation operations. Inner product calculation can be implemented using a general function interface or accelerated using vectorized instructions.

[0039] Step S106: Construct a second matrix based on the inner product results between each pair of reference local windows; wherein, the process of constructing the second matrix includes: selecting reference local window pairs in sequence according to a preset traversal order, the traversal order being configured such that at least one pair of adjacent reference local window pairs in the selection sequence has an overlapping region when calculating the inner product, and, for the next reference local window pair, obtaining the inner product result of the current reference local window pair by incremental calculation based on the inner product result of the previous reference local window pair.

[0040] Specifically, the second matrix (denoted as matrix A) is used to characterize the pairwise correlations between all reference local windows. Let the total number of reference local windows be R, then the second matrix A is an R×R square matrix, where the (i, j)th element A[i][j] is the inner product of the i-th reference local window ri and the j-th reference local window rj. The construction of the second matrix is ​​the most computationally intensive part of the entire solution process. In existing techniques, a complete inner product operation is performed independently for each pair (i, j), resulting in a total computational complexity of O(R²). 2 (×M). When the number of reference local windows is huge, the computational load is extremely large. This invention significantly reduces the computational load required to construct the second matrix through two core methods: traversal order control and incremental computation.

[0041] The preset traversal order refers to selecting reference local window pairs sequentially according to a specific spatial sequence, rather than simply selecting them according to matrix row or column numbers. This traversal order is configured such that at least one pair of adjacent reference local windows in the selection sequence overlaps. Spatially, each subsequent window pair is slightly shifted relative to the preceding one, so most of the pixel regions between the two window pairs overlap. This provides the foundation for subsequent incremental calculations.

[0042] During the selection of reference local window pairs according to the above traversal order, for each subsequent reference local window pair, if it overlaps with the previous one, the inner product result of the current reference local window pair is obtained through incremental calculation based on the inner product result of the previous reference local window pair. The core principle of incremental calculation is that when the current reference local window pair is translated relative to the previous one, most of the pixel correspondence between the two window pairs remains unchanged, with only the boundary changing. Specifically, some pixels move out of the window overlap range due to translation and no longer contribute to the inner product; other pixels enter the window overlap range due to translation and need to contribute. Therefore, the new inner product value can be obtained by subtracting the multiplicative cumulative contribution of the outgoing pixels from the old inner product value and adding the multiplicative cumulative contribution of the incoming pixels, without performing any multiplicative addition operations on the majority of overlapping pixels that remain unchanged in the middle. In scenarios with a large target processing area and a high window overlap ratio, the computational savings of incremental calculation are considerable.

[0043] Step S108: Solve the system of linear equations based on the first and second matrices to obtain compensation parameters for local linear dynamic compensation of the data to be processed.

[0044] Specifically, after constructing the first matrix B and the complete second matrix A, the linear equation system A×kernel=B is solved to obtain the compensation parameter kernel. This linear equation system can be solved using standard numerical linear algebra methods, such as Cholesky decomposition, LU decomposition, QR decomposition, or singular value decomposition (SVD). Since the second matrix A is constructed using inner product operations, it is itself symmetric positive semi-definite. To further improve numerical stability, damping values ​​can be added to the diagonal of the second matrix A to improve the condition number of the matrix and enhance the numerical stability of the solution.

[0045] The compensation parameter kernel is a vector of size R×1 (where R is the total number of reference local windows), where each element represents the weight coefficient of the corresponding reference local window in local linear compensation. After obtaining the compensation parameter, it can be used to perform local linear dynamic compensation on the data to be processed. In industrial inspection image scenarios, the compensation parameter is used for brightness correction, local grayscale compensation, or defect region enhancement of the industrial inspection image. In Computer-Aided Engineering (CAE) result field scenarios, the compensation parameter is used for result field consistency correction, multi-condition fusion, or local suppression of abnormal regions.

[0046] Furthermore, in some preferred embodiments of the present invention, a window relationship matrix is ​​pre-constructed; wherein, the window relationship matrix is ​​determined based on the order of extracting local windows; the traversal order is to start from the upper right corner of the window relationship matrix and traverse the window relationship matrix along the diagonal from the upper left to the lower right.

[0047] For details, see Figure 3 The diagram shown illustrates a traversal order optimization method provided by an embodiment of the present invention. This diagram uses a simplified 3×3 window template as an example. Figure 3 In (b), the index of the top-left corner of the left matrix (the sliding window template corresponding to one of the reference data in the local window) is... Figure 3 The x-coordinate in (a) is... Figure 3 In (b) of the diagram, the index of the top-left corner of the right matrix (the sliding window template corresponding to another reference data in the local window pair) is... Figure 3 The ordinate in (a) is Figure 3 The element in (a) represents the inner product when the top-left corners of the two reference data are moved to their corresponding positions in the corresponding window templates. For example, 08 indicates that the top-left corner of the left matrix is ​​at position 0 and the top-left corner of the right matrix is ​​at position 8, and the inner product is calculated between them. Figure 3In the window relation matrix (a) shown in the diagram, elements of the same color indicate data reuse when calculating the inner product. Therefore, the traversal order is along the diagonal of the matrix. There exists at least one calculation of the inner product and result of two adjacent elements. Data can be reused, so the traversal order is from the top left to the bottom right. The traversal order can be 08, 07, 18, 06, 17, 28, 05, 16, 27, 38, 04, 15, 26, 37, 48, ..., 00, 11, 22, 33, 44, 55, 66, 77.

[0048] Each reference local window is considered as a node, and arranged spatially into a concept matrix. The element (i, j) of the second matrix A represents the inner product of window i and window j.

[0049] For example, see [link to example]. Figure 3 , Figure 3 This illustrates the specific execution method of traversing along the diagonal direction used in embodiments of the present invention. For example... Figure 3 In (c), the red, yellow, and green arrows indicate different diagonal traversal directions. The diagonal traversal order used in this embodiment of the invention is as follows: starting from the upper right element (0, 8) of the second matrix, then proceeding along the secondary diagonal direction, then 07, 18; 06, 17, 28; 16, 27; 26; 05, 38; 04, 15, 37, 48; 03, 14, 25, 36, 47, 58; 13, 24, 46, 57; 23, 56; 02, 35, 68; 01, 12, 34, 45, 67, 78; 00, 11, 22, 33, 44, 55, 66, 77, 88.

[0050] In practical applications, when the number of reference local windows is large, the dimension of the second matrix A is R×R. When calculating the upper triangular part of the second matrix, starting from the upper right element (0, R-1), the calculation proceeds sequentially along the second diagonal, advancing diagonally downwards and to the left. On each diagonal, the calculated window pairs have a stable translation relationship in image space. Taking window pairs (0, 8) and (1, 7) as an example, window 0 and window 1 are spatially adjacent (translation relationship), and window 8 and window 7 are also spatially adjacent; therefore, the pixel correspondence between the two window pairs highly overlaps. After traversing one diagonal, the calculation proceeds one unit to the left, starting the calculation for the next diagonal. This traversal method ensures that incremental computation can be applied continuously throughout the entire computation process, rather than only sporadically between individual window pairs. It also allows the reference local window pairs of adjacent computations to be translated diagonally within the target processing region, resulting in the largest window overlap area. This maximizes the applicability of incremental computation, significantly improves data reuse efficiency, and reduces memory access costs and computational load.

[0051] Furthermore, in some preferred embodiments of the present invention, extracting local windows based on a preset local window template includes: first fixing the current data, traversing all local windows within the current data, and after completing the extraction of all local windows, switching to the next data and performing the same processing until all data has been processed.

[0052] Specifically, when multiple reference data exist, traditional methods may traverse the data in the order of window position as the outer loop and reference data as the inner loop, that is, sequentially reading the corresponding windows of all reference data at each window position. This approach requires frequent switching between memory regions of different reference data during processing, resulting in frequent cache invalidation and poor data locality. This invention adopts a method of first fixing the current reference data object and then traversing all local windows within that reference data object. For example, in an industrial inspection image scenario, there are 4 reference images, and N reference local windows can be extracted from each reference image. First, fix the first reference image, extract all N reference local windows within that image, and cache the relevant description information; then, fix the second, third, and fourth reference images sequentially and perform the same operation. The same order is used when constructing the first and second matrices: first process all window pairs involving the first reference image, then process all window pairs involving the second reference image, and so on. Since local windows within the same reference data are stored contiguously in memory or have a regular step size relationship, processing all local windows continuously within a single reference data set can fully utilize CPU cache and prefetching mechanisms, avoiding cache thrashing caused by repeatedly traversing different data buffers. This embodiment reduces frequent switching between different reference data sets, improves the locality of data access, reduces data switching overhead, and increases cache hit rate, thereby accelerating the entire extraction and computation process. The benefits of this optimization are particularly significant in scenarios with a large amount of reference data.

[0053] Furthermore, in some preferred embodiments of the present invention, after extracting the local window, the method further includes: pre-calculating and caching the local region index, rectangular range, and / or sub-block description information of the local window in the reference data; and directly using the pre-calculated and cached local region index, rectangular range, and / or sub-block description information during the construction of the first matrix and the second matrix.

[0054] Specifically, the construction of the first and second matrices requires repeated access to the data of each local window, and each access requires determining the specific position of the window in the reference data. If the start and end row and column numbers, rectangular regions, or data pointers of the window are recalculated each time within the main loop, it will introduce a large amount of repetitive region resolution overhead. This embodiment of the invention advances this work to after the local window extraction is completed and before matrix construction begins, calculating and caching it all at once.

[0055] The local region index refers to the starting and ending row and column numbers of each reference local window within the target processing area of ​​the reference data. For example, for the i-th reference local window, the row index range is [row_start_i, row_end_i], and the column index range is [col_start_i, col_end_i]. The rectangular range refers to the window position and size described by a rectangular data structure, such as using a rectangular object containing the coordinates of the top-left corner (x, y), width, and height to represent the data area covered by the window. The sub-block description information refers to the access description of the local window in the reference data memory, such as the window data start pointer (pointing to the address of the top-left element of the window in the reference data memory) and the row step size (the number of bytes per row of the reference data matrix, i.e., the offset from the start of one row to the start of the next row). With the start pointer and row step size, any element within the window can be read directly through pointer offset operations, without needing to address pixel-by-pixel or grid-by-grid points through function calls, significantly reducing access overhead.

[0056] The aforementioned information is calculated once and cached in an array or vector container before the main loop begins. In subsequent loops constructing the first and second matrices, when data for a specific reference local window is needed, the corresponding local region index, rectangular range, or sub-block description information can be retrieved directly from the cache using the index, allowing for quick location and access to all data within that window. By moving repetitive region parsing work outside the loop and caching it all at once, the repetitive calculation of region location information during large-scale window computations is avoided, effectively reducing the overhead of repetitive region parsing. This is particularly suitable for scenarios involving large-scale window computations with high-resolution data, improving overall computational efficiency.

[0057] Furthermore, in some preferred embodiments of the present invention, for the subsequent reference local window pair, the inner product result of the current reference local window pair is obtained by incremental calculation based on the inner product result of the previous reference local window pair, including: when the subsequent reference local window pair is translated relative to the previous reference local window pair, the inner product result of the previous reference local window pair is subtracted from the cumulative contribution of the pixel values ​​that moved out of the overlapping area due to the translation, and then the cumulative contribution of the pixel values ​​that newly entered the overlapping area due to the translation is added.

[0058] For details, see Figure 4 The diagram shown illustrates the principle of local window overlap and incremental calculation provided by an embodiment of the present invention. Taking a 5×5 pixel reference local window as an example, it demonstrates several typical overlap situations between two reference local windows. r1 and r2 represent the two reference local windows respectively, the pink area in the middle is the overlapping part of the two, that is, the pixel area jointly covered by the two windows; the other colored areas are the unique parts of each. Figure 4The top left image shows the general situation: r1 has a unique upper brown area and a left yellow area, r2 has a unique right blue area and a lower cyan area, and the two windows overlap in the middle pink area. Figure 4 The top right image shows the situation when translating horizontally: r1 has the yellow area on the left, and r2 has the blue area on the right. Figure 4 The bottom left image shows the situation when translating vertically: r1 has a unique upper brown area, and r2 has a unique lower cyan area. Figure 4 The bottom right image shows the situation when the diagonal is translated, which is different from... Figure 4 The general case shown in the top left figure is equivalent.

[0059] Let the window template size be W×W. For a reference local window pair (r1, r2) whose inner product has been fully calculated, its inner product value is... Where (u, v) are the row and column coordinates within the window. When the window pair is shifted horizontally to the right by d pixels, a new window pair (r1', r2') is obtained. r1' shifts out of a region of one or more columns of pixels (denoted as region L) of width d from the left relative to r1. out The region with width d (denoted as region R) was moved into the right side. in Similarly, r2' also undergoes a corresponding regional change relative to r2. The inner product value P of the new window pair... new It can be calculated using the incremental formula: Typically, the sliding step size d=1, meaning that each shift is one pixel, with both the outgoing and incoming shifts involving one column, totaling W pixels. For a 5×5 window, independent calculations require 25 multiplications and 24 additions, while incremental calculations only require processing the multiplication and addition adjustments of 2×5=10 pixels, reducing the computational load by approximately 60%.

[0060] Similarly, when a window is moved downwards in the vertical direction, the area moved out is one or more rows above (denoted as area U). out The area to be moved into is one or more rows below (denoted as area D). in The incremental formula is: .

[0061] When a window is translated diagonally, the outgoing region consists of the left and top boundaries, while the incoming region consists of the right and bottom boundaries. Incremental calculations need to handle boundary changes in both directions simultaneously. Figure 4 In the case shown in the lower right figure, r1 has only the brown area at the top and the yellow area on the left, while r2 has only the blue area on the right and the cyan area at the bottom. The incremental calculation needs to subtract the contribution of the moved-out area (the part moved out of the brown and yellow areas) and add the contribution of the moved-in area (the part added to the blue and cyan areas).

[0062] In real-world high-resolution scenarios, such as when the target processing area is 2044×2044 and the window template is 5×5, the proportion of the overlapping area of ​​the window in a single calculation is typically over 99% (for example, when shifting only one pixel to the left or down, the overlapping pixels are (2044×2043) / (2044×2044) = 99.951%). Incremental calculations do not require repeated multiplication and addition for the vast majority of pixels, resulting in substantial computational savings. In CAE result fields, the data is in floating-point format, and the formula for incremental calculations remains identical, only the data type changes from integer to floating-point. Since floating-point operations are inherently more time-consuming than integer operations, the absolute computational time reduction achieved by incremental calculations in floating-point scenarios is even more significant.

[0063] This embodiment uses incremental calculation, performing multiplication and accumulation adjustments only on a small number of pixels moved out and in due to translation, avoiding large-scale repetitive multiplication and accumulation operations in overlapping window areas, thus significantly reducing the computational load. This reduction in computational load directly translates into shorter execution time, enabling the local window dynamic compensation solution for high-resolution data to meet real-time requirements.

[0064] Furthermore, in some preferred embodiments of the present invention, constructing the second matrix based on the inner product results between pairs of reference local windows further includes: calculating the upper triangular part or the lower left triangular part of the second matrix, and completing it according to the symmetry relationship to obtain the complete second matrix.

[0065] For details, see Figure 5 The illustrated embodiment of the present invention provides a schematic diagram of reducing computational load by utilizing the symmetry of a second matrix. Since the inner product operation is commutative, i.e., dot(ri, rj) = dot(rj, ri) holds for any i and j, the second matrix A is a symmetric matrix satisfying A[i][j] = A[j][i]. This means that only about half of the elements in matrix A need to be actually calculated through the inner product; the other half can be obtained directly by copying based on the symmetry.

[0066] Figure 5Taking a 25×25 second matrix A as an example, this symmetry is visually demonstrated. The elements of the lower left triangle (white) and upper right triangle (colored) are perfectly symmetrical about the main diagonal. In this embodiment, it is possible to choose to calculate only the upper triangular part (i.e., elements i ≤ j) or only the lower triangular part (i.e., elements i ≥ j). Taking the calculation of the upper triangular part as an example, when selecting window pairs according to the aforementioned diagonal traversal order, only window pairs satisfying i ≤ j are selected and calculated, i.e., elements located on or above the main diagonal in the second matrix. After calculating each upper triangular element A[i][j], the value is immediately assigned to its symmetrical position A[j][i], thus obtaining the complete second matrix A while completing the calculation of all upper triangular elements. Since the upper triangular (or lower triangular) part contains R(R+1) / 2 elements, and the entire matrix has R... 2 There are 4 elements, therefore the computational cost of the inner product is reduced from R. 2 This reduction is approximately R. 2 / 2 times, reducing by about half.

[0067] This can be used in conjunction with the aforementioned incremental calculation. Incremental calculation applies to adjacent window pairs, while symmetry utilization applies to any two windows. During the traversal of the upper triangular portion, for adjacent window pairs, incremental calculation is still applied to reduce the computational cost of a single inner product; for non-adjacent window pairs (e.g., when switching to a new diagonal), a full inner product calculation can be used as the starting point for incremental calculation, and the result is also copied to the symmetrical position through symmetry. The two methods complement each other, reducing the overall computational cost from different dimensions.

[0068] This embodiment utilizes the symmetry of the second matrix to reduce the computational complexity of the inner product from R... 2 This reduction is approximately R. 2 This reduces the computational load by about half, significantly lowering the computational burden in high-resolution, multi-reference-frame, and multi-window scenarios, and further improving the overall solution speed.

[0069] Furthermore, in some preferred embodiments of the present invention, the method further includes: calculating the inner product result using a single instruction multiple data (SIMD) vectorized instruction, wherein the data to be processed and the reference data are both in 8-bit unsigned integer format, and the vectorized instruction performs parallel data processing for the 8-bit unsigned integer format.

[0070] Specifically, Advanced Vector Extensions (AVX) is a Single Instruction Multiple Data (SIMD) instruction set extension for x86 architecture CPUs. It allows the same operation to be performed on multiple data elements within a single clock cycle, thereby improving data parallel processing capabilities. In industrial image inspection scenarios, the pixel format of the data to be processed and the reference data is 8-bit unsigned integers (U8), with each pixel occupying 1 byte. AVX2 instructions can use a 256-bit wide vector register (such as the YMM register) to load 32 U8 data elements at a time; AVX512 instructions can use a 512-bit wide vector register (such as the ZMM register) to load 64 U8 data elements at a time. By loading consecutive pixel data into vector registers in batches and using vector multiplication and vector addition instructions for parallel processing, the throughput of inner product calculations can be significantly improved.

[0071] In practical implementation, for the U8 format, a single AVX2 vector multiplication instruction can simultaneously perform 32 multiplications of U8 data, producing 32 intermediate results. These intermediate results are then accumulated using vector addition and horizontal addition instructions to finally obtain the scalar result of the inner product. For window sizes that are not divisible by the width of the vector register (e.g., 5×5=25 pixels, not divisible by 32 or 64), the implementation needs to handle data alignment and remaining elements, which is typically achieved efficiently through loop unrolling and masking operations.

[0072] Through practical testing and comparison of various inner product kernel implementations for industrial image inspection in U8 format, the results are as follows: OpenCVdot, as the baseline implementation, takes approximately 616~663 microseconds for a single inner product; SSE takes approximately 493~557 microseconds; AVX2 takes approximately 437~538 microseconds, showing superior overall performance; AVX512 takes approximately 442~492 microseconds, with performance close to AVX2 but higher CPU hardware requirements. These data demonstrate that the vectorized AVX implementation has significant advantages over general-purpose interfaces.

[0073] Furthermore, in the specific implementation of vectorized computation, this embodiment preferably adopts a hybrid vectorization approach: maintaining the original data in memory in U8 format, after loading the data into the CPU vector register, it is then converted to a data format suitable for vector multiplication and addition operations (e.g., 16-bit integer) using vector conversion instructions, and then the vector multiplication and addition operation is performed. This approach avoids the additional time and memory overhead caused by performing full format pre-conversion on all data to be processed and reference data. In contrast, if all U8 data is pre-converted to 32-bit floating-point numbers or 32-bit integers before vectorized computation, not only is an additional traversal for data conversion required, but the data volume also increases by 2 to 4 times, putting enormous pressure on cache capacity and memory bandwidth. The hybrid vectorization implementation performs on-the-fly conversion only at the register level, taking into account both the parallel advantages of vectorization and the advantages of compact data storage.

[0074] Vectorized parallel processing of U8 format data allows more data elements to be processed simultaneously in a single instruction, fully leveraging the CPU's vector computing capabilities and significantly improving the throughput of inner product calculations. At the same time, the hybrid vectorization implementation avoids the additional overhead of full pre-conversion, further meeting the high throughput requirements of industrial real-time detection scenarios.

[0075] Furthermore, in some preferred embodiments of the present invention, the method further includes: dividing the data to be processed, the reference data, the first matrix, and the second matrix into sub-blocks of a preset size (e.g., 128×128) for calculation to adapt to the L1 cache.

[0076] For details, see Figure 6 The illustrated embodiment of the present invention provides a schematic diagram of block-based computation. During the construction of the second matrix, a large amount of reference local window data needs to be accessed, and this data volume is proportional to the total number of reference local windows, R. In industrial image inspection scenarios, if the target processing area is 2044×2044, and each reference data provides N=2040×2040≈4.16×10⁻⁶... 6 If we consider a local reference window, then the dimension of the second matrix A is approximately 4N×4N≈1.66×10⁻⁶. 7 ×1.66×10 7 Because the complete second matrix A is extremely large, it is impossible to load it into the cache or memory at once and calculate the entire matrix directly. Similarly, the data volume of all reference local windows is also enormous and cannot all reside in the CPU's L1 cache.

[0077] This embodiment solves the above problem by dividing the construction process of the second matrix into multiple sub-blocks for computation. Specifically, the second matrix A is divided into sub-blocks of a certain size, and the inner product is calculated on a sub-block basis. The size of the sub-block is determined so that the computational working set of a single sub-block is adapted to the capacity of the processor's L1 cache. In this embodiment, the block size used is 128×128. This means that the second matrix A is calculated in batches in 128×128 element sub-blocks. When calculating a 128×128 sub-block, it is necessary to access the data of two matrices. Each element of the matrix is ​​U8 (1B), occupying a total of 128×128×2×1=32768B=32KB, which is just suitable for the 32KB level L1 data cache of mainstream CPUs. Figure 6 Taking a 2048×2048 overall image as an example, this illustrates the layout of dividing the image into 16×16 128×128 sub-blocks.

[0078] Block-based computation allows the entire computation process of a single sub-block (including data loading, inner product operation, incremental calculation, symmetric completion, etc.) to be completed in the L1 cache, significantly improving cache hit rate and avoiding high memory access latency caused by frequent access to main memory. In the specific implementation, the outer loop operates on a sub-block basis, while the inner loop performs the aforementioned diagonal traversal, incremental calculation, and symmetry utilization operations within the sub-block. The block-based strategy has good compatibility and complementarity with the aforementioned optimization techniques.

[0079] Furthermore, when performing block calculations on the second matrix, the inner loop can be expanded, preferably by expanding it to 4 rows. Loop expansion is a compiler and manual code optimization technique that improves the utilization efficiency of the CPU instruction pipeline and achieves instruction-level parallelism by reducing the number of loop iterations and the proportion of loop checks and branch jump instructions. Expanding to 4 rows means that the calculation work of 4 rows is processed simultaneously in each loop iteration, further reducing loop control overhead and allowing the processor to focus more on performing actual data operations.

[0080] By partitioning the computational workset to fit the processor's L1 cache capacity, cache hit rate is significantly improved and memory access latency is reduced. At the same time, loop unrolling reduces loop control overhead and improves instruction-level parallelism, further improving overall computational efficiency, especially suitable for high-resolution and large-data-volume processing scenarios.

[0081] The embodiments of this invention can significantly reduce the execution latency of the linear dynamic compensation solution process based on local windows, especially for industrial inspection image processing scenarios. Considering the characteristics of high-resolution images, U8 pixels, uniform local texture, multiple reference frames, and high real-time requirements in this scenario, the effect is not a local improvement brought about by a single optimization point, but is jointly produced by reducing the number of independent inner products, improving the reusability rate, improving cache locality, and enhancing vectorized execution efficiency. The specific effects are as follows: First, due to the symmetry of the second matrix A, and considering the computational scenario of multiple reference frames and multiple windows in industrial inspection images, only half of the region needs to be calculated to complete the overall construction, thus directly reducing the number of related inner products by 50%. Furthermore, leveraging the characteristics of uniform local texture and high overlap between adjacent windows in industrial inspection images, the original single-step reconstruction of the inner product can be transformed into an incremental update of 'subtracting the contribution from the moved-out region and adding the contribution from the moved-in region'. Therefore, it is unnecessary to repeatedly multiply and add to a large amount of overlapping data, further reducing the computational load.

[0082] Secondly, by changing the window traversal order to proceed along the diagonal and adopting a method of fixing the reference object first and then switching the local window at the implementation level, combined with the characteristics of multiple reference frames in industrial inspection images, the data correlation between continuous calculations can be significantly improved, making it easier for loaded data and obtained results to be reused in subsequent steps, thereby reducing memory access costs and improving computational efficiency.

[0083] Furthermore, by employing AVX vectorized instructions to implement the inner product kernel, optimized for the U8 pixel format of industrial inspection images, more data elements can be processed in parallel within a single instruction. Test results show that AVX2 and AVX512 implementations have significant advantages over general-purpose interfaces; at the same time, the hybrid implementation that maintains the original data as U8 and converts it after reading is more suitable for industrial inspection image scenarios than the full pre-conversion scheme, avoiding additional overhead.

[0084] Furthermore, by dividing the computation process into 128×128 blocks matching the L1 cache and combining it with 4 rows of loop expansion, this embodiment of the invention can reduce loop control overhead, improve instruction-level parallelism, and hide some memory access latency, thus further improving overall performance, especially considering the high resolution of 2k×2k industrial inspection images.

[0085] Based on the synergistic effect of the aforementioned technical means, for industrial image inspection scenarios (single process, 4 sets of 2k×2k data, 10 iterations), the original processing time was approximately 4428 ms, which was optimized to approximately 211 ms, achieving a speedup of 20.99, meeting the requirement of ≤500ms latency for single-frame processing in real-time industrial inspection. In a scenario with 19 concurrent processes, 10 sets of 2k×2k data, and 10 iterations, the original processing time was approximately 139923 ms, which was optimized to approximately 1589 ms, achieving a speedup of 88.06 and an improvement of 98.86%, meeting the needs of multi-task concurrent processing in industrial inspection. These results demonstrate that this invention is suitable not only for low-latency single-task requirements in industrial inspection scenarios but also for high-throughput multi-task concurrent scenarios.

[0086] Taking linear dynamic compensation acceleration in industrial inspection image scenarios as an example, the input is an industrial inspection image to be processed (U8 format, 2k×2k resolution, reflecting the grayscale information of the surface of industrial parts) and 4 reference industrial inspection images (same specifications, same scene). The region of interest (ROI) is reduced by 2 pixels in each of the top, bottom, left and right directions (to avoid invalid edge data); the local window template size is 5×5 (to adapt to the local compensation of minor defects in parts); iterates 10 times (simulating the continuous processing scenario of industrial inspection).

[0087] The process involves reading the industrial inspection image to be processed and four reference inspection images; establishing a 5×5 local window template; calculating two layers of compensation kernels; reducing computation by employing symmetry during matrix A construction (calculating only the upper triangular region and completing the lower triangular region); traversing the window combination in diagonal order (maximizing window overlap reuse); incrementally reusing the inner product results of adjacent windows (utilizing the characteristics of uniform local texture and high window overlap in industrial inspection images); implementing core calculations using AVX hybrid vectorization (adapting to U8 pixel format); dividing the image into 128 blocks (adapting to L1 cache and reducing memory access latency for high-resolution images); expanding the inner loop with four rows (reducing loop control overhead); and finally solving the linear equation system to obtain the kernel list, which is used for brightness correction of the image to be processed.

[0088] In a single-process scenario with 4 reference data points, 2k×2k resolution, and 10 iterations, the original processing time was 4428 ms, while the current processing time is 211 ms, representing a speedup of 20.99 and an improvement of 95.23%. The single-frame processing latency has been reduced to 211 ms, meeting the requirements for real-time industrial inspection. The corrected industrial inspection images have uniform brightness and clear edges, which can effectively assist in subsequent defect identification and improve the detection accuracy by more than 10%.

[0089] Taking the solution of local window compensation parameters in the CAE result field as an example, the input is a result field of a certain working condition obtained from finite element simulation and multiple reference working condition result fields. The result field can be a two-dimensional matrix after expansion of temperature field, stress field, strain field, displacement field or other scalar / vector field.

[0090] The target analysis region is extracted from the result field; a local window template is established; the local window of the target working condition is used as the standard window, and the windows corresponding to multiple reference working conditions are used as reference windows to construct the first matrix B; then the second matrix A is constructed through the correlation between the reference windows; the symmetry utilization, repeated reuse, diagonal traversal, vectorized kernel, block calculation and loop expansion strategy of the present invention are adopted to accelerate the solution; finally, the local linear compensation parameters are obtained, which are used for result field consistency correction, abnormal region suppression or local fitting analysis.

[0091] For data matrices formed after mapping large-scale two-dimensional or three-dimensional result fields, the method provided in this embodiment of the invention can significantly reduce the solution latency of local window modeling, and is especially suitable for CAE applications that require batch analysis of multiple working conditions or interactive post-processing.

[0092] Furthermore, taking an industrial inspection image scenario as an example, the time consumption for different optimization stages in an industrial inspection image scenario (2k×2k, 4-frame reference, 5×5 window, 10 iterations) is as follows: 1. Original implementation stage: The overall time is about 4800 ms to 5000 ms, with more than 98% of the time spent in the dot function, which cannot meet the requirements of real-time detection.

[0093] 2. Function splitting only: This is easier to measure, but the overall time improvement is limited, approximately 4750 ms.

[0094] 3. Pre-compute rectangular region (adapting to fixed boundary features): Move the parsing of repetitive regions outside the loop, reducing the total time to approximately 4510 ms.

[0095] 4. Using pointer offsets to replace some general access interfaces (adapting to U8 format): the total time is reduced to about 2350 ms.

[0096] 5. Utilizing only symmetry (adapting to multiple reference frames and multiple window features): can reduce some computation, but the individual benefit is limited, with a total time consumption of approximately 1820 ms.

[0097] 6. Introducing reuse and diagonal traversal (adapting to window overlap features): The total time was reduced to approximately 461 ms, initially meeting the real-time requirements.

[0098] 7. Combined with AVX hybrid implementation (adapted to U8 format): The total time for the two functions is approximately 420 ms, and the total time for the single function is approximately 408 ms.

[0099] 8. Add matrix partitioning (to adapt to high-resolution features): the total time for a single process is about 254 ms, and the full process takes about 330 ms.

[0100] 9. Finally, loop unrolling is added: the total time for a single process is about 216 ms, and for a full process it is about 279 ms.

[0101] Therefore, it can be seen that the significant effect of the method provided by the embodiments of the present invention comes from the synergistic cooperation of multiple optimization methods, and each optimization point is designed for specific features of industrial inspection images, rather than simply replacing a basic library function. The optimization effect is even more prominent in industrial inspection scenarios with high resolution, U8 format, and multiple reference frames.

[0102] Furthermore, considering the characteristics of U8 format industrial inspection images, a comparison of different inner product kernel implementations is shown in the table below:

[0103] DPBUS refers to the _mm_dpbusd_epi32 class of instructions in AVX. Comparative results show that simply rewriting using general-purpose C++ does not outperform mature library implementations; rather, a vectorized implementation tailored to U8 format and high-resolution features for industrial image inspection, especially when combined with reuse and computational link reconstruction, demonstrates significant advantages and meets the requirements of real-time industrial inspection.

[0104] Furthermore, for inner product calculation in industrial inspection image scenarios, a comparison of common calculation libraries is shown in the table below:

[0105] Test results show that common solutions such as OpenCV, Eigen, MKL, and NumPy have similar performance in general inner product scenarios, indicating that it is difficult to achieve orders-of-magnitude optimization simply by replacing common libraries. However, the method provided in this embodiment of the invention achieves better results in the overall business scenario by reconstructing the computation path and data reuse method, and designing optimization strategies in combination with the specific features of industrial inspection images, thus meeting the requirements of real-time detection.

[0106] Furthermore, for industrial image inspection scenarios, the performance test results under different numbers of processes, data volume, and data scale are shown in the table below:

[0107] A process count of 1 indicates single-device, single-task detection, meeting real-time requirements. Multiple reference frames allow for adaptation to complex scenarios. A process count of 19 indicates concurrent detection by multiple devices, improving throughput. Suitable for large-scale industrial detection.

[0108] This invention provides a linear dynamic compensation solution method based on local windows. When constructing the second matrix, the traversal order is configured to ensure overlapping regions between adjacent reference local window pairs in the selected sequence. Based on this, the entire inner product of the current reference local window pair is no longer calculated independently; instead, the current result is obtained through incremental calculation based on the inner product result of the previous reference local window pair. Since most pixel regions between adjacent window pairs overlap, this incremental calculation only needs to handle the small contributions of pixels that have moved out and in, eliminating the need for repeated multiplication and addition in the overlapping regions. This significantly reduces the number of multiplication and addition operations and data reading required to construct the second matrix, effectively reducing the overall computational overhead and accelerating the linear dynamic compensation solution process based on local windows.

[0109] Based on the above embodiments, this invention provides a linear dynamic compensation solution device based on a local window, see [link to relevant documentation]. Figure 7 The diagram shown is a structural schematic of a linear dynamic compensation solution device based on a local window, according to an embodiment of the present invention. The device includes: The data processing module 310 is used to acquire the data to be processed and multiple reference data, and extract local windows based on a preset local window template; wherein, the local window includes a standard local window corresponding to the data to be processed and multiple reference local windows corresponding to the reference data; The first matrix construction module 320 is used to construct the first matrix based on the inner product result of the standard local window and the reference local window; The second matrix construction module 330 is used to construct a second matrix based on the inner product results between pairs of reference local windows. The process of constructing the second matrix includes: selecting reference local window pairs in a preset traversal order, the traversal order being configured such that at least one pair of adjacent reference local window pairs in the selection sequence has an overlapping region when calculating the inner product, and for the next reference local window pair, obtaining the inner product result of the current reference local window pair by an incremental calculation method based on the inner product result of the previous reference local window pair. The compensation parameter determination module 340 is used to solve the linear equation system based on the first matrix and the second matrix to obtain the compensation parameters used for local linear dynamic compensation of the data to be processed.

[0110] Furthermore, in some preferred embodiments of the present invention, a window relationship matrix is ​​pre-constructed; wherein, the window relationship matrix is ​​determined based on the order of extracting local windows; the traversal order is to start from the upper right corner of the window relationship matrix and traverse the window relationship matrix along the diagonal from the upper left to the lower right.

[0111] Furthermore, in some preferred embodiments of the present invention, the data processing module 310 is used to first fix the current data, traverse all local windows within the current data, extract all local windows, and then switch to the next data for the same processing until all data has been processed.

[0112] Furthermore, in some preferred embodiments of the present invention, the apparatus further includes a data and calculation module, used to pre-calculate and cache the local region index, rectangular range and / or sub-block description information of the local window in the reference data; and to directly retrieve the pre-calculated and cached local region index, rectangular range and / or sub-block description information during the construction of the first matrix and the second matrix.

[0113] Furthermore, in some preferred embodiments of the present invention, the second matrix construction module 330 is used to subtract the cumulative contribution of pixel values ​​that have moved out of the overlapping region due to the translation from the inner product result of the previous reference local window pair when the subsequent reference local window pair is translated relative to the previous reference local window pair, and add the cumulative contribution of pixel values ​​that have newly entered the overlapping region due to the translation.

[0114] Furthermore, in some preferred embodiments of the present invention, the second matrix construction module 330 is used to calculate the upper triangular part or the lower left triangular part of the second matrix and complete it according to the symmetry relationship to obtain the complete second matrix.

[0115] Furthermore, in some preferred embodiments of the present invention, the apparatus further includes: an inner product calculation module, used to calculate the inner product result using a single instruction multiple data stream vectorized instruction, wherein the data to be processed and the reference data are both in 8-bit unsigned integer format, and the vectorized instruction performs parallel data processing on the 8-bit unsigned integer format.

[0116] Furthermore, in some preferred embodiments of the present invention, the apparatus further includes: a data segmentation module, used to segment and calculate the data to be processed, the reference data, the first matrix, and the second matrix into sub-blocks of a preset size.

[0117] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the linear dynamic compensation solution device based on local windows described above can be referred to the corresponding process in the aforementioned embodiments of the linear dynamic compensation solution method based on local windows, and will not be repeated here.

[0118] This invention also provides an electronic device for running a linear dynamic compensation solution method based on a local window; see [link to related documentation]. Figure 8The schematic diagram of an electronic device provided by the embodiment of the present invention shown includes a memory 400 and a processor 401. The memory 400 is used to store one or more computer instructions, which are executed by the processor 401 to implement the above-mentioned linear dynamic compensation solution method based on local windows.

[0119] Furthermore, Figure 8 The electronic device shown also includes a bus 402 and a communication interface 403. The processor 401, the communication interface 403 and the memory 400 are connected via the bus 402.

[0120] The memory 400 may include high-speed random access memory (RAM) and may also include non-volatile memory, such as at least one disk storage device. Communication between this system network element and at least one other network element is achieved through at least one communication interface 403 (which can be wired or wireless), such as the Internet, wide area network, local area network, metropolitan area network, etc. The bus 402 can be an ISA bus, PCI bus, or EISA bus, etc. The bus can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 8 The symbol is represented by a single double-headed arrow, but this does not mean that there is only one bus or one type of bus.

[0121] Processor 401 may be an integrated circuit chip with signal processing capabilities. In implementation, each step of the above method can be completed by the integrated logic circuitry in the hardware of processor 401 or by instructions in software form. Processor 401 can be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it can also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this invention. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this invention can be directly manifested as execution by a hardware decoding processor, or execution by a combination of hardware and software modules in the decoding processor. The software module can reside in a readily available storage medium in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, or registers. This storage medium is located in memory 400, and processor 401 reads information from memory 400 and, in conjunction with its hardware, completes the steps of the method described in the foregoing embodiments.

[0122] This invention also provides a computer-readable storage medium storing computer-executable instructions. When these computer-executable instructions are called and executed by a processor, they cause the processor to implement the aforementioned linear dynamic compensation solution method based on local windows. For specific implementation details, please refer to the method embodiments, which will not be repeated here.

[0123] The computer program product of the linear dynamic compensation solution method, apparatus and electronic device based on local window provided in the embodiments of the present invention includes a computer-readable storage medium storing program code. The instructions included in the program code can be used to execute the methods in the preceding method embodiments. For specific implementation, please refer to the method embodiments, which will not be repeated here.

[0124] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the system and / or device described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0125] Furthermore, in the description of the embodiments of the present invention, unless otherwise explicitly specified and limited, the terms "installation," "connection," and "linking" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in the present invention based on the specific circumstances.

[0126] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, essentially, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0127] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A linear dynamic compensation solution method based on a local window, characterized in that, include: The system acquires data to be processed and multiple reference data, and extracts local windows based on a preset local window template; wherein, the local window includes a standard local window corresponding to the data to be processed and multiple reference local windows corresponding to the reference data; A first matrix is ​​constructed based on the inner product of the standard local window and the reference local window; A second matrix is ​​constructed based on the inner product results between each pair of the reference local windows; wherein, the process of constructing the second matrix includes: selecting reference local window pairs in a preset traversal order, the traversal order being configured such that at least one pair of adjacent reference local window pairs in the selection sequence has an overlapping region when calculating the inner product, and, for the latter reference local window pair, obtaining the inner product result of the current reference local window pair by an incremental calculation method based on the inner product result of the former reference local window pair; Solving the linear equations based on the first and second matrices yields compensation parameters used for local linear dynamic compensation of the data to be processed.

2. The method according to claim 1, characterized in that, A window relationship matrix is ​​pre-constructed; wherein the window relationship matrix is ​​determined based on the order in which the local windows are extracted; the traversal order is to start from the upper right corner of the window relationship matrix and traverse the diagonal of the window relationship matrix from the upper left to the lower right.

3. The method according to claim 1, characterized in that, Extracting a local window based on a preset local window template includes: First, fix the current data, then traverse all the local windows within the current data. After extracting all the local windows, switch to the next data and perform the same processing until all data has been processed.

4. The method according to claim 1, characterized in that, After extracting the local window, the method further includes: Pre-calculate and cache the local region index, rectangular range, and / or sub-block description information of the local window in the reference data; During the construction of the first matrix and the second matrix, the pre-calculated and cached local region index, the rectangular range and / or the sub-block description information are directly used.

5. The method according to claim 1, characterized in that, For the latter reference local window pair, based on the inner product result of the former reference local window pair, the inner product result of the current reference local window pair is obtained through an incremental calculation method, including: When the latter reference local window pair is translated relative to the former reference local window pair, the inner product result of the former reference local window pair is subtracted from the cumulative contribution of the pixel values ​​that moved out of the overlapping area due to the translation, and then the cumulative contribution of the pixel values ​​that newly entered the overlapping area due to the translation is added.

6. The method according to claim 1, characterized in that, Constructing the second matrix based on the pairwise inner product results of the aforementioned reference local windows further includes: Calculate the upper triangular part or the lower left triangular part of the second matrix, and complete the second matrix according to the symmetry relationship.

7. The method according to claim 1, characterized in that, The method further includes: The inner product result is calculated using a single instruction multiple data stream vectorized instruction. Both the data to be processed and the reference data are in 8-bit unsigned integer format. The vectorized instruction performs parallel data processing on the 8-bit unsigned integer format.

8. The method according to claim 1, characterized in that, The method further includes: The data to be processed, the reference data, the first matrix, and the second matrix are divided into sub-blocks of preset size for calculation.

9. A linear dynamic compensation solution device based on a local window, characterized in that, include: The data processing module is used to acquire data to be processed and multiple reference data, and extract local windows based on a preset local window template; wherein, the local window includes a standard local window corresponding to the data to be processed and multiple reference local windows corresponding to the reference data; The first matrix construction module is used to construct a first matrix based on the inner product result of the standard local window and the reference local window; The second matrix construction module is used to construct a second matrix based on the inner product results between each pair of the reference local windows. The process of constructing the second matrix includes: selecting reference local window pairs in a preset traversal order, wherein the traversal order is configured such that at least one pair of adjacent reference local window pairs in the selection sequence has an overlapping region when calculating the inner product; and, for the latter reference local window pair, the inner product result of the current reference local window pair is obtained by incremental calculation based on the inner product result of the former reference local window pair. The compensation parameter determination module is used to solve a system of linear equations based on the first matrix and the second matrix to obtain compensation parameters for local linear dynamic compensation of the data to be processed.

10. An electronic device, characterized in that, The method includes a processor and a memory, the memory storing computer-executable instructions executable by the processor, the processor executing the computer-executable instructions to implement the method according to any one of claims 1 to 8.