Computer vision point-line feature extraction acceleration method and device based on Shiwu chip
By optimizing the computer vision feature extraction algorithm on Cambrian chips, using parallel computing and asynchronous flow optimization technology, the shortcomings of traditional algorithms in terms of computing efficiency and real-time performance are solved, and efficient feature extraction is achieved, meeting the needs of real-time application scenarios.
Patent Information
- Application Number
- CN202510235705.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-06-17
AI Technical Summary
Traditional computer vision feature extraction algorithms have shortcomings in computing efficiency and real-time performance, and it is difficult to meet the needs of high-resolution image processing and real-time application scenarios.
By implementing memory alignment optimization, vector parallel computing and asynchronous flow optimization technologies on Cambrian chips, the feature extraction process of SIFT and CANNY operators is optimized, and the parallel computing capabilities of Cambrian chips are fully utilized.
It significantly improves the efficiency and performance of computer vision point-line feature extraction, and meets application scenarios with high real-time requirements, such as autonomous driving and robot navigation.
Smart Images

Figure CN120163845A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the fields of computer vision and artificial intelligence technologies, and in particular, to a method and apparatus for accelerating computer vision point-line feature extraction based on Cambrian chips. Background Art
[0002] This section aims to provide background or context for the embodiments of the present disclosure recited in the claims. The description herein is not admitted to be prior art merely by including it in this section.
[0003] In the field of computer vision, feature extraction is a core technology for implementing numerous applications, and its performance directly determines the performance of the entire system. Traditional feature extraction algorithms, such as SIFT (Scale-Invariant Feature Transform) and CANNY edge detection algorithm, are widely used due to their accuracy and robustness in feature extraction. The SIFT algorithm detects key points in an image and generates feature descriptors by constructing Gaussian pyramids and difference pyramids, and has good invariance to scale, rotation, and illumination changes. The CANNY algorithm accurately detects edge information in an image through steps such as Gaussian filtering, gradient calculation, non-maximum suppression, and double-threshold detection. However, the computational complexity of these algorithms is relatively high, especially when processing high-resolution images, which often requires a long calculation time and is difficult to meet the application scenarios with high real-time requirements.
[0004] With the continuous development of computer vision technology, higher requirements are put forward for the real-time performance and efficiency of feature extraction algorithms. In practical applications, such as in scenarios like autonomous driving and robot navigation, it is necessary to quickly and accurately extract image features in order to make decisions and responses in a timely manner. Therefore, how to improve the computational efficiency of feature extraction algorithms has become an important research direction.
[0005] In recent years, with the development of artificial intelligence chip technology, Cambrian chips, as a kind of high-performance AI chips, have powerful parallel computing capabilities and low-power consumption characteristics, providing new possibilities for accelerating computer vision feature extraction algorithms. By utilizing the parallel computing capabilities of Cambrian chips, the operation efficiency of feature extraction algorithms can be significantly improved, thereby meeting the application scenarios with high real-time requirements. However, how to effectively apply the parallel computing capabilities of Cambrian chips to the optimization of feature extraction algorithms remains a challenging problem. Summary of the Invention
[0006] In view of this, the purpose of the present disclosure is to propose a method and device for accelerating the extraction of point-line features in computer vision based on Cambrian chips, so as to solve the deficiencies of traditional feature extraction algorithms in terms of computational efficiency and meet application scenarios with high real-time requirements. The present disclosure realizes the efficient acceleration of point-line feature extraction in computer vision through a combination of algorithm optimization and hardware acceleration, providing strong technical support for application scenarios with high real-time requirements.
[0007] Based on the above purpose, the first aspect of the exemplary embodiment of the present disclosure provides a method for accelerating point feature extraction based on Cambrian chips, and the method includes:
[0008] Perform memory alignment optimization to ensure that the storage location of data in memory matches the access pattern, reduce memory access latency, ensure that the Cambrian BANG C high-performance vector computing operator can be applied, and improve data processing efficiency;
[0009] Utilize the vector acceleration unit of the Cambrian chip to perform vector parallel computing, allocate image processing tasks to multiple computing units for parallel processing, and significantly improve the computing speed;
[0010] Through asynchronous pipelining optimization technology, parallelize data transmission and computing tasks, significantly reduce the impact of data transmission time on computing efficiency, and improve the computing performance of the operator.
[0011] Based on the same inventive concept, the second aspect of the exemplary embodiment of the present disclosure provides a method for accelerating line feature extraction based on Cambrian chips, including:
[0012] Deeply optimize the Gaussian filtering operator, adopt a separable convolution strategy, and utilize the parallel computing unit of the chip to perform convolution operations on the image in the horizontal and vertical directions respectively. By reasonably allocating computing tasks, each computing core is responsible for performing convolution on a specific row or column of the image, reducing the computing complexity and improving the processing speed. At the same time, since the Gaussian kernel coefficients are fixed, these coefficients are calculated and cached in advance to avoid repeated calculations during convolution, further improving the computing efficiency.
[0013] Use the SOBEL operator with better effects for gradient calculation, and implement computing parallelization using the Cambrian chip. Through vector parallel computing, accelerate the gradient calculation process;
[0014] Perform non-maximum suppression parallelization. Since it involves conditional branches, it causes certain difficulties for parallel computing. Adopt an optimization strategy based on masked branch processing, calculate the calculation results of each branch in turn, and generate a mask vector according to the branch conditions. Utilize the vector operation ability of the Cambrian chip to perform logical operations on the mask vector and the calculation results to achieve the branch effect, ensuring both the high efficiency of parallel computing and the accuracy of the calculation results.
[0015] Based on the same inventive concept, a third aspect of the exemplary embodiments of the present disclosure provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the method described in the first aspect is implemented.
[0016] Based on the same inventive concept, a fourth aspect of the exemplary embodiments of the present disclosure provides a non-transitory computer-readable storage medium storing computer instructions for causing a computer to execute the method described in the first aspect.
[0017] Based on the same inventive concept, a fifth aspect of the exemplary embodiments of the present disclosure provides a computer program product including computer program instructions, which when running on a computer, cause the computer to execute the method described in the first aspect.
[0018] As can be seen from the above, the computer vision point-line feature extraction acceleration method and device based on Cambrian chips proposed in the embodiments of the present disclosure, through the in-depth optimization of various point-line feature extraction algorithms, give full play to the powerful advantages of Cambrian chips and significantly improve the efficiency and performance of computer vision point-line feature extraction. This invention can better meet the requirements of computer vision application scenarios with high real-time requirements, such as autonomous driving, robot navigation, real-time monitoring, etc., and provides strong support for the development of related technologies. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In order to more clearly illustrate the technical solutions in the present disclosure or related technologies, the following will briefly introduce the drawings required for use in the embodiments or related technology descriptions. Obviously, the drawings in the following description are only the embodiments of the present disclosure. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0020] Figure 1 Schematic diagram of the overall steps of the SIFT operator provided by the embodiments of the present disclosure;
[0021] Figure 2 Schematic diagram of the construction process of the SIFT image pyramid provided by the embodiments of the present disclosure;
[0022] Figure 3 Schematic diagram of the SIFT Gaussian pyramid provided by the embodiments of the present disclosure;
[0023] Figure 4 Schematic diagram of the construction process of the SIFT difference pyramid provided by the embodiments of the present disclosure;
[0024] Figure 5Schematic diagram of the MLU instruction stream provided by the embodiments of the present disclosure;
[0025] Figure 6 Schematic diagram of the development steps of the CANNY operator provided by the embodiments of the present disclosure;
[0026] Figure 7 Schematic diagram of the splitting of the two-dimensional Gaussian kernel provided by the embodiments of the present disclosure;
[0027] Figure 8 Schematic diagram of the memory data movement provided by the embodiments of the present disclosure;
[0028] Figure 9 Schematic diagram of the non-maximum suppression mask vector provided by the embodiments of the present disclosure;
[0029] Figure 10 Schematic diagram of the hardware structure of the electronic device provided by the embodiments of the present disclosure; Detailed implementation manners
[0030] It can be understood that before using the technical solutions disclosed in the embodiments of the present application, the types, usage scopes, usage scenarios, etc. of the personal information involved in the present application should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.
[0031] For example, when responding to receiving an active request from the user, a prompt message is sent to the user to clearly prompt the user that the operation requested by the user will require obtaining and using the user's personal information. Thus, the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, an application program, a server, or a storage medium that executes the operations of the technical solutions of the present application according to the prompt message.
[0032] As an optional but non-limiting implementation manner, the manner of sending a prompt message to the user in response to receiving an active request from the user can be, for example, in the form of a pop-up window, and the prompt message can be presented in text in the pop-up window. In addition, the pop-up window can also carry a selection control for the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0033] It can be understood that the above process of notifying and obtaining the user's authorization is only illustrative and does not limit the implementation manner of the present application, and other manners that meet relevant laws and regulations can also be applied to the implementation manner of the present application.
[0034] It can be understood that the data involved in the technical solution of the present application (including but not limited to the data itself, the acquisition or use of the data) should comply with the requirements of the corresponding laws, regulations and related provisions.
[0035] To make the objectives, technical solutions, and advantages of the present disclosure more clearly understood, the principles and spirit of the present disclosure will be described below with reference to several exemplary embodiments. It should be understood that these embodiments are provided only to enable those skilled in the art to better understand and thereby implement the present disclosure, and do not limit the scope of the present disclosure in any way. On the contrary, these embodiments are provided to make the present disclosure more thorough and complete, and to fully convey the scope of the present disclosure to those skilled in the art.
[0036] In this document, it should be understood that any number of elements in the drawings is for illustration rather than limitation, and any naming is only for distinction and does not have any limiting meaning.
[0037] It should be noted that unless otherwise defined, the technical terms or scientific terms used in the embodiments of the present disclosure should be the ordinary meanings understood by those of ordinary skill in the art to which the present disclosure pertains. The "first", "second", and similar terms used in the embodiments of the present disclosure do not denote any order, quantity, or importance, but are only used to distinguish different components. The terms such as "comprising" or "including" mean that the elements or items appearing before this term cover the elements or items listed after this term and their equivalents, without excluding other elements or items. The terms such as "connected" or "coupled" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The terms such as "upper", "lower", "left", "right", etc. are only used to represent relative positional relationships, and when the absolute position of the object being described changes, the relative positional relationship may also change accordingly. The article "a" or "an" before an element does not exclude the existence of multiple such elements.
[0038] The principles and spirit of the present disclosure will be elaborated in detail below with reference to several representative embodiments of the present disclosure.
[0039] Existing visual SLAM systems often face many challenges under conditions such as low texture and weak illumination. Under these conditions, since traditional feature extraction algorithms only consider local information in the image and do not consider the structure and semantic information of the image, existing SLAM systems may get stuck because it is difficult to extract and match accurate and stable features, resulting in unstable or even ineffective tracking. For example, in a low-light environment, the extraction and tracking stability of feature points in the image are poor, which easily leads to tracking loss and a decrease in positioning accuracy. In addition, in weakly textured areas, traditional feature extraction algorithms have difficulty finding enough feature points, further exacerbating the performance problems of the SLAM system.
[0040] To solve these problems, feature extraction operators with better performance are needed, such as the SIFT and CANNY operators. Although they perform well in terms of the accuracy and robustness of feature extraction, their computational complexity is relatively high. Especially when dealing with large-scale image data, it is difficult to meet the requirements of application scenarios with high real-time requirements. In addition, with the wide application of computer vision technology in fields such as autonomous driving and robot navigation, higher requirements are put forward for the real-time performance and efficiency of feature extraction algorithms.
[0041] The inventors of the present disclosure found that as a kind of hardware specifically designed for artificial intelligence computing, Cambrian chips have powerful parallel computing capabilities, efficient memory management mechanisms, and customized instruction sets, and can provide hardware-level acceleration support for computer vision point-line feature extraction. By making full use of these characteristics of Cambrian chips, optimizing and improving traditional feature extraction algorithms can significantly improve the efficiency and accuracy of feature extraction. That is, through the optimization of the SIFT and CANNY operators, the hardware advantages of Cambrian chips can be fully utilized to achieve efficient acceleration of the feature extraction process.
[0042] To overcome the limitations in the prior art, the present disclosure provides a method and device for accelerating computer vision point-line feature extraction based on Cambrian chips. By adopting optimized SIFT and CANNY operators and combining the parallel computing capabilities of Cambrian chips, efficient acceleration of the feature extraction process is achieved. Specifically, the present disclosure accelerates the construction of the Gaussian pyramid and the feature point extraction process of the SIFT operator through technologies such as memory alignment optimization, vector parallel computing, and asynchronous pipelining optimization; and accelerates the edge detection process of the CANNY operator through technologies such as Gaussian filtering parallelization, gradient calculation parallelization, and non-maximum suppression parallelization. These optimization measures significantly improve the computational efficiency of feature extraction while maintaining the accuracy and robustness of feature extraction.
[0043] The principle and spirit of the present disclosure lie in making full use of the parallel computing capabilities of Cambrian chips to optimize traditional feature extraction algorithms, thereby improving computational efficiency and meeting the real-time and efficient requirements in practical applications.
[0044] The following specifically elaborates on the two core acceleration operators in this embodiment respectively.
[0045] SIFT Operator Based on Cambrian Parallel Acceleration
[0046] For the SIFT operator on the Cambrian MLU chip, an algorithm design based on parallel computing is proposed to optimize the feature point extraction process to make full use of the MLU hardware characteristics.
[0047] The feature point extraction part of the SIFT operator is specifically split into five steps: image downsampling, Gaussian filtering, image differencing, calculating the gradient magnitude and direction, and key point detection, as Figure 1 shown. The core construction process of the Gaussian image pyramid is as Figure 2 shown.
[0048] According to different scaling factors, multiple downsampled images are generated from the original image to form an image pyramid. At this time, there is only one downsampled original image data in each layer of the pyramid. Image downsampling is the first step in building the image pyramid, that is, the original image is first downsampled multiple times with different parameters to generate an image pyramid as Figure 3 shown. The purpose of the downsampling operation is to generate an image hierarchical structure with different resolutions by repeatedly shrinking the original image, so as to provide multi-scale image information for subsequent feature point detection and description.
[0049] For each image in each layer of the image pyramid, Gaussian filtering is performed multiple times with different parameters to generate a Gaussian image pyramid, as Figure 2 shown. Each core of the MLU processes 1 / 16 of the image. If the Gaussian kernel radius is gR and the number of data rows in a single operation is row r , then theoretically the number of data rows row c to be copied in each batch is as shown in Equation (1).
[0050] row c = 2gR + 1 + row r (1)
[0051] Since vector operations are located in the NRAM space with the highest cost in the Cambrian MLU chip and the space resources are limited, in order to make full use of the space resources and optimize the performance, this study makes full use of the symmetry of the Gaussian kernel. Specifically, in the current algorithm, when calculating pixels at symmetric positions, the corresponding Gaussian kernel coefficients are the same, resulting in the same multiplication operation being calculated twice. By reasonably caching data, not only can the number of data rows copied each time be reduced, thus reducing the occupancy of the NRAM space, but also the number of operations of Gaussian blur can be significantly reduced, improving the calculation efficiency. After optimization, the number of data rows row c to be copied in each batch is as shown in Equation (1).
[0052] row c = gR + 1 + row r (2)
[0053] After that, a Difference of Gaussian pyramid (DoG pyramid) is constructed. First, it is necessary to perform differencing calculations on each layer of images in the Gaussian pyramid to generate corresponding difference images and generate a Difference of Gaussian image pyramid. The generation process is as Figure 4As shown. Specifically, for two adjacent images in the same layer of the Gaussian pyramid, their difference is calculated to obtain a difference image. This process is used in the SIFT algorithm to detect key feature points of the image. Each core of the MLU processes 1 / 16 of the image.
[0054] Based on the Gaussian difference image pyramid, calculating the gradient magnitude and direction is a key step in the SIFT algorithm. This process aims to assist in subsequent feature point extraction and descriptor generation to achieve rotational invariance. The gradient magnitude represents the intensity of gray-scale change at each pixel point in the image. The formula for calculating the magnitude adopted by this operator is shown in Equation (3).
[0055]
[0056] Among them, L(x, y) represents the gray-scale value of the image at position (x, y). This formula calculates the gradient magnitude of each pixel point by computing the gray-scale differences in the horizontal and vertical directions. At the same time, to make the feature point descriptor rotationally invariant, the SIFT algorithm will consider the main direction of the area where the key point is located during feature extraction. Specifically, the algorithm calculates the histogram of gradient directions in the neighborhood of the key point, and the main direction corresponds to the peak of the histogram. Then, the neighborhood coordinate system of the key point is rotated to the main direction, making the descriptor rotationally invariant. Therefore, calculating the gradient direction is necessary. The formula for calculating the gradient direction is shown in Equation (4). This formula calculates the gradient direction of each pixel point by computing the gray-scale differences in the horizontal and vertical directions.
[0057]
[0058] Finally, after obtaining the gradient pyramid, key points are detected. Each pixel point is compared with all its adjacent points to see if it is larger or smaller than its adjacent points in the image domain and scale domain. The pixel is compared with a total of 26 points, which are the 9 points in the corresponding position neighborhoods of the two upper and lower images in its vector scale space. If the point is the largest or smallest point, then the point is temporarily listed as a feature point.
[0059] During the operator optimization process, experimental results show that the time taken for copying image data between memory and the on-chip storage (GDRAM) of Cambrian MLU is relatively high, almost equivalent to the time taken for the operator to actually compute on the MLU chip. This bottleneck restricts the improvement of the overall performance of the operator. Therefore, it is necessary to effectively reduce the impact of data transfer time on computing efficiency through optimization strategies.
[0060] When designing the Cambrian MLU hardware, the basic characteristics of artificial intelligence applications were fully considered, and multiple levels of parallelism were set. The MLU hardware supports both data-level parallelism and instruction-level parallelism. Data-level parallelism means that multiple data are processed simultaneously in one instruction. For example, vector instructions are typical data-level parallelism. The MLU hardware also provides multiple pipelines that can be executed in parallel, each corresponding to a different function, such as Figure 5 as shown. Instructions in different pipelines can be executed in parallel, thus achieving instruction-level parallelism between different pipelines.
[0061] The MLU hardware also provides synchronization instructions to achieve synchronization at positions where dependencies need to be maintained. The system reduces the interruption of different streams by synchronization instructions and reduces data dependencies between streams by reasonably arranging the execution order of vector / tensor operations and memory access instructions, and tries to reduce them through instruction scheduling or rearrangement, so as to achieve latency hiding of vector / tensor calculations and I / O within the MLU Core.
[0062] In order to make full use of the hardware characteristics of the Cambrian MLU, an asynchronous pipeline optimization strategy was adopted in the process of optimizing the construction of the difference pyramid of the SIFT operator in this study. Specifically, by decoupling the image calculation task from the data transfer process between the host memory and the GDRAM, an asynchronous processing framework was constructed. Under this framework, the IO stream (data transfer) and the Compute stream (calculation task) can run in parallel, thus hiding the latency of data transfer and reducing the overall running time of the operator.
[0063] CANNY operator based on Cambrian parallel acceleration
[0064] For the CANNY operator on the Cambrian MLU chip, an algorithm design based on parallel computing was proposed to optimize the feature point extraction process to make full use of the MLU hardware characteristics.
[0065] The CANNY operator was specifically split into five steps: Gaussian filtering, gradient calculation, non-maximum suppression, threshold screening, and edge connection, as Figure 6 shown.
[0066] First, the original image is preprocessed such as memory alignment and copied to the MLU memory. Then, Gaussian filtering is applied to smooth the image to reduce the influence of noise. On this basis, the gradient of the image is calculated. This step involves the calculation of image difference, and the Sobel operator is used to improve the accuracy of the result. Subsequently, the gradient magnitude and direction of the image are calculated. Since this process involves square root and division operations, considering the low efficiency of the Cambrian hardware in performing these operations, we reduce the use of division and square root operations by optimizing the calculation order and reusing the calculated results, so as to minimize the impact of the calculation bottleneck as much as possible.
[0067] Next, the focus will be on the optimization of Gaussian filtering, gradient calculation, and non-maximum suppression.
[0068] Gaussian filtering is to perform Gaussian blur on the original image to reduce the impact of noise.
[0069] Gaussian blur is an important operation in image processing. Its main function is to smooth the original image, thereby reducing the impact of noise and providing more stable input data for subsequent feature extraction and calculation. In the SLAM system, Gaussian blur is usually used as a pre-step for constructing the scale space. By performing multi-scale Gaussian filtering on the image, feature information at different resolutions is generated. However, Gaussian blur is essentially a two-dimensional convolution operation, that is, a two-dimensional Gaussian kernel slides on the image, and convolution calculations are performed on each pixel point. The computational complexity of this two-dimensional convolution is relatively high, especially when processing large-resolution images, which may bring a large computational burden. Therefore, in this study, algorithm optimization was carried out for the Gaussian blur operation, and combined with the parallel computing characteristics of the Cambrian MLU hardware, a significant acceleration effect was achieved.
[0070] The optimization of Gaussian blur is based on its mathematical properties. The one-dimensional definition of the Gaussian kernel is shown in Equation (5).
[0071]
[0072] And the two-dimensional definition of the Gaussian kernel is shown in Equation (6).
[0073]
[0074] From Equation (4-1) and Equation (4-2), the method of decomposing the two-dimensional Gaussian into one-dimensional can be deduced as shown in Equation (7).
[0075]
[0076] Therefore, the two-dimensional Gaussian kernel can be decomposed into the combination of two one-dimensional Gaussian kernels, which means that the two-dimensional convolution operation can be decomposed into two independent one-dimensional convolution operations: first perform one-dimensional convolution in the horizontal direction, and then perform one-dimensional convolution in the vertical direction. As Figure 7 shown, assuming the size of the two-dimensional Gaussian kernel is 5×5, it can be decomposed into a 1×5 horizontal kernel and a 5×1 vertical kernel. This decomposition reduces the computational complexity of the convolution from O(n 2 ) to O(2n), significantly reducing the amount of calculation.
[0077] Meanwhile, when performing Gaussian blur, the one-dimensional convolution in the horizontal direction is executed first, followed by the one-dimensional convolution in the vertical direction. This way of separating the convolution kernel is suitable for parallelization, and the Cambrian MLU parallel computing unit can be utilized to further improve the processing speed and achieve efficient parallel processing. This optimized parallelization scheme significantly enhances the processing efficiency of Gaussian blur, enabling it to meet the real-time requirements of the SLAM system.
[0078] This study proposed an optimization scheme. Considering the characteristic that image data is stored row by row in memory, and the data of adjacent rows is adjacent in memory. By shifting the entire image data left (right) by n units, in the resulting new image data, each pixel position corresponds to the nth pixel on the right (left) side of the original image, as Figure 8 shown. This method avoids the transpose operation and directly performs convolution calculation on the image, thereby reducing the time overhead and improving the efficiency of parallel computing.
[0079] The following explains the optimization for gradient calculation.
[0080] Image gradient calculation is one of the important steps in image processing and is widely used in tasks such as edge detection, feature extraction, and texture analysis. Gradient calculation is used to describe the rate and direction of pixel intensity changes in an image, which can help detect edge information in the image. Common gradient calculation methods include the difference method and the convolution method, among which the Sobel operator is a classic and efficient image gradient calculation operator. By optimizing the Sobel operator, the gradient calculation process can be accelerated, thereby improving the overall image processing efficiency.
[53] 。
[0081] In image processing, the gradient is generally achieved through differences, as shown in Equation (8).
[0082]
[0083] However, in the image, Δx = 1, so generally f(x + 1, y) - f(x, y) is used to approximate the gradient magnitude in the x direction, and the same applies to the y direction.
[0084] In the CANNY algorithm, the Sobel operator is often used to calculate the gradient, and the operator is as shown in Equations (9) and (10).
[0085]
[0086] From this, the gradients G x , G y in the x and y directions can be obtained, as shown in Equations (11) and (12).
[0087]
[0088] Use the Sobel operator to calculate the gradients G in the x and y directions respectively x , G y After that, the modulus value of the sum of the gradients in the two directions can be obtained as: The gradient direction is:
[0089] Although the Sobel operator can effectively calculate the gradient of an image, its calculation process faces performance bottlenecks, especially in real-time applications such as feature extraction and edge detection in visual SLAM systems. In this case, the following strategies are adopted in this project to optimize the calculation process of the Sobel operator and improve the performance of the operator.
[0090] (1) Parallel computing optimization
[0091] With the acceleration support of Cambrian chips, the calculation process of the Sobel operator can be parallelized. Specifically, the image is divided into multiple data blocks, and the gradient calculation tasks within each block can be performed simultaneously, thus making full use of the multi-core parallel processing ability of Cambrian chips. Each computing unit is responsible for processing the gradient calculation of a part of the image, and finally the results are summarized. In this way, the calculation speed of the Sobel operator is significantly improved, meeting the requirements of real-time image processing.
[0092] (2) Memory access optimization
[0093] Memory access in image processing is an important factor affecting the calculation speed. The Sobel operator needs to frequently access the neighboring pixels of the image, and the traditional memory access mode may lead to cache misses, thereby affecting the calculation efficiency. By optimizing the data storage and access mode, such as adopting data prefetching technology and memory alignment optimization, the memory access latency can be reduced. The on-chip storage of Cambrian chips can accelerate data access. By reasonably planning the storage mode of image data in memory, the data transfer time can be reduced, thereby improving the calculation efficiency.
[0094] (3) Operator fusion and optimization
[0095] In practical applications, the gradient calculation of the Sobel operator is usually part of multiple image processing steps, such as edge detection, feature extraction, etc. Through operator fusion technology, the gradient calculation can be combined with subsequent operations (such as non-maximum suppression, edge tracking, etc.) to reduce duplicate calculations and memory accesses in the calculation steps. For example, in Canny edge detection, the gradient calculation of the Sobel operator and non-maximum suppression in the edge detection process can be jointly optimized to further improve the performance.
[0096] The optimization for non-maximum suppression is explained below.
[0097] Due to the multi-quadrant characteristic of the gradient direction, multiple conditional branches are involved in the NMS calculation, which significantly increases the difficulty of implementing parallel computing on Cambrian MLU chips. The traditional branch processing method will cause pipeline stalls and reduce the computing efficiency. To solve this problem, this project proposes an optimization strategy based on masked branch processing.
[0098] This optimization strategy based on masked branch processing is to calculate each branch in sequence and multiply by the relevant mask according to the branch conditions to achieve the effect of the branch. This not only ensures the efficiency of parallel computing but also ensures correct calculation according to the branch conditions. Specifically, as Figure 9 shown, after calculating the results of each branch through parallel computing, multiply by the corresponding mask vector to obtain the results of all data corresponding to this branch. Finally, add up the results of the four branches to obtain the result of non-maximum suppression.
[0099] Figure 10 FIG. shows a more specific schematic diagram of the hardware structure of an electronic device provided in this embodiment. The device may include: an input / output interface 101, a CPU processor 102, a memory 103, an MLU chip 104, and a bus 105. Among them, the CPU processor 102, the memory 103, the input / output interface 101, and the MLU chip 104 are communicatively connected to each other inside the device through the bus 105.
[0100] The processor 102 uses a general-purpose CPU (Central Processing Unit) to execute relevant programs to implement the technical solutions provided in the embodiments of this specification.
[0101] The memory 103 may be implemented in the form of a ROM (Read Only Memory), a RAM (Random Access Memory), a static storage device, a dynamic storage device, etc. The memory 103 may store an operating system and other application programs. When implementing the technical solutions provided in the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 103 and called and executed by the processor 102.
[0102] The input / output interface 101 is used to connect to an input / output module to implement information input and output. The input / output module may be configured as a component in the device (not shown in the figure) or externally connected to the device to provide corresponding functions. Among them, the input device may include a monocular camera, a binocular camera, a hard disk file, etc., and the output device may include a display, etc.
[0103] The MLU chip 104 is used for the computing tasks related to the acceleration of computer vision point-line feature extraction based on Cambrian chips in the present invention. Specifically, the MLU220 or MLU270 chip can be adopted to implement the technical solutions provided in the embodiments of this specification.
[0104] The bus 105 includes a path for transmitting information among various components of the device (such as the processor 102, the memory 103, the input / output interface 101, and the MLU chip 104).
[0105] It should be noted that although the above device only shows the processor 102, the memory 103, the input / output interface 101, the MLU chip 104, and the bus 105, in the specific implementation process, the device may also include other components necessary for normal operation. In addition, those skilled in the art can understand that the above device may also only include the components necessary to implement the solution of the embodiments of this specification, and do not necessarily include all the components shown in the figure.
[0106] The electronic device in the above embodiment is used to implement the corresponding computer vision point-line feature extraction acceleration method based on Cambrian chips in any of the foregoing embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be elaborated here.
[0107] Based on the same inventive concept, corresponding to the method in any of the above embodiments, the present disclosure also provides a non-transitory computer-readable storage medium. The non-transitory computer-readable storage medium stores computer instructions, and the computer instructions are used to make the computer execute the computer vision point-line feature extraction acceleration method based on Cambrian chips as described in any of the foregoing embodiments.
[0108] The computer-readable medium in this embodiment includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible by a computing device.
[0109] The above non-transitory computer-readable storage medium can be any available medium or data storage device accessible by a computer, including but not limited to magnetic memories (such as floppy disks, hard disks, magnetic tapes, magneto-optical disks (MO), etc.), optical memories (such as CDs, DVDs, BDs, HVDs, etc.), and semiconductor memories (such as ROMs, EPROMs, EEPROMs, non-volatile memories (NAND FLASH), solid-state drives (SSD)), etc.
[0110] The computer instructions stored in the storage medium of the above embodiments are used to cause the computer to execute the computer vision point-line feature extraction acceleration method based on Cambrian chips as described in any of the embodiments in the above exemplary method section, and have the beneficial effects of the corresponding method embodiments, which will not be elaborated here.
[0111] Based on the same inventive concept, corresponding to the computer vision point-line feature extraction acceleration method based on Cambrian chips described in any of the above embodiments, the present disclosure also provides a computer program product, which includes computer program instructions. In some embodiments, the computer program instructions can be executed by one or more processors of a computer to cause the computer and / or the processor to execute the computer vision point-line feature extraction acceleration method based on Cambrian chips. Corresponding to the execution subjects of the respective steps in the respective embodiments of the computer vision point-line feature extraction acceleration method based on Cambrian chips, the processors executing the corresponding steps can belong to the corresponding execution subjects.
[0112] The computer program product of the above embodiments is used to cause the computer and / or the processor to execute the computer vision point-line feature extraction acceleration method based on Cambrian chips as described in any of the above embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be elaborated here.
[0113] Those skilled in the art know that the embodiments of the present disclosure can be implemented as a system, method, or computer program product. Therefore, the present disclosure can be specifically implemented in the following forms: complete hardware, complete software (including firmware, resident software, microcode, etc.), or a combination of hardware and software, generally referred to as "circuit", "module", or "system" in this article. In addition, in some embodiments, the present disclosure can also be implemented in the form of a computer program product in one or more computer-readable media, which contains computer-readable program code.
[0114] Any combination of one or more computer-readable media may be employed. The computer-readable media may be a computer-readable signal medium or a computer-readable storage medium. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the foregoing. More specific examples (a non-exhaustive list) of the computer-readable storage medium may include: an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. In the present document, a computer-readable storage medium may be any tangible medium that contains or stores a program which can be used by or in connection with an instruction execution system, apparatus, or device.
[0115] A computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which computer-readable program code is carried. Such a propagated data signal may take many forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the foregoing. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device.
[0116] The program code contained on the computer-readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wireline, optical fiber cable, RF, etc., or any suitable combination of the foregoing.
[0117] The computer program code for performing the operations of the present disclosure may be written in one or more programming languages or combinations thereof, including object-oriented programming languages such as Python, Smalltalk, C++, and also including conventional procedural programming languages such as the "C" language or similar programming languages. The program code may execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer, or entirely on the remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0118] It should be understood that each block of the flowchart and / or block diagram, and combinations of blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the computer program instructions, when executed by the computer or other programmable data processing apparatus, create means for implementing the functions / operations specified in the blocks of the flowchart and / or block diagram.
[0119] These computer program instructions can also be stored in a computer-readable medium that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable medium produce an article of manufacture including instruction means for implementing the functions / operations specified in the blocks of the flowchart and / or block diagram.
[0120] The computer program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process such that the instructions executed on the computer or other programmable apparatus provide processes for implementing the functions / operations specified in the blocks of the flowchart and / or block diagram.
[0121] In addition, although the operations of the methods of the present disclosure are depicted in the figures in a particular order, this is not required or implied to perform the operations in that particular order, or to perform all of the illustrated operations to achieve the desired result. On the contrary, the steps depicted in the flowchart may be performed in a different order. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step and executed, and / or one step may be decomposed into multiple steps and executed.
[0122] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present application. Each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of code, and the above-mentioned module, segment of a program, or portion of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions noted in the blocks may occur in a different order than noted in the figures. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and combinations of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0123] It should be noted that although several modules or units of a device for action execution are mentioned in the above detailed description, such a division is not mandatory. In fact, according to the embodiments of the present application, the features and functions of two or more of the above-described modules or units can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.
[0124] Those of ordinary skill in the art should understand that: The discussion of any of the above embodiments is only exemplary and is not intended to imply that the scope of the present application (including the claims) is limited to these examples; Under the concept of the present application, the technical features between the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations in different aspects of the embodiments of the present application as described above, and they are not provided in detail for the sake of brevity.
[0125] In addition, for the sake of simplicity of description and discussion, and in order not to make the embodiments of the present application difficult to understand, the well-known power / ground connections to integrated circuit (IC) chips and other components may or may not be shown in the provided drawings. In addition, the device may be shown in block diagram form to avoid making the embodiments of the present application difficult to understand, and this also takes into account the fact that the details of the implementation of such block diagram devices are highly dependent on the platform on which the embodiments of the present application are to be implemented (i.e., these details should be completely within the understanding of those skilled in the art). In the case where specific details (such as circuits) are set forth to describe the exemplary embodiments of the present application, it will be apparent to those skilled in the art that the embodiments of the present application can be implemented without these specific details or with variations of these specific details. Therefore, these descriptions should be considered illustrative rather than restrictive.
[0126] Although the present application has been described in conjunction with specific embodiments of the present application, many alternatives, modifications, and variations of these embodiments will be apparent to those of ordinary skill in the art based on the foregoing description. For example, other memory architectures (such as dynamic RAM (DRAM)) can be used with the embodiments discussed.
[0127] The embodiments of the present application are intended to cover all such alternatives, modifications, and variations that fall within the broad scope of the appended claims. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the embodiments of the present application shall be included within the protection scope of the present application.
[0128] Although the spirit and principles of the present disclosure have been described with reference to several specific embodiments, it should be understood that the present disclosure is not limited to the specific embodiments disclosed, and the division of each aspect does not mean that the features in these aspects cannot be combined for benefits. Such division is only for the convenience of expression. The present disclosure is intended to cover various modifications and equivalent arrangements included within the spirit and scope of the appended claims. The scope of the appended claims is to be accorded the broadest interpretation so as to encompass all such modifications and equivalent structures and functions.
Claims
1. A computer vision point and line feature extraction acceleration method based on Cambrian chip, characterized in that: include: Obtaining a reference image and a query image; Optimize memory alignment of input images to ensure that data storage and access patterns match to call Cambrian BANG C high-performance vector computing operators; Using the vector acceleration unit of the Cambrian chip, image processing tasks are assigned to multiple computing units for parallel computing; Through asynchronous pipeline optimization technology, data transmission and computing tasks are decoupled and executed in parallel, hiding data transmission delays; Gaussian pyramid and difference pyramid are constructed based on the optimized SIFT operator to extract multi-scale feature points; Perform edge detection based on the optimized CANNY operator, including parallelized Gaussian filtering, gradient calculation, and non-maximum suppression; Output point and line feature extraction results for real-time computer vision applications.
2. The method according to claim 1, characterized in that The Gaussian pyramid and the differential pyramid are constructed based on the optimized SIFT operator, including: Downsample the input image to generate a multi-resolution image pyramid; Use the separation convolution strategy to perform Gaussian filtering in the horizontal and vertical directions on each layer of the image to generate a Gaussian pyramid; By differentially calculating adjacent Gaussian images, a differential pyramid is constructed; Detect extreme points in the difference pyramid as candidate feature points; Calculate the gradient magnitude and direction of the feature points to generate a feature descriptor with rotation invariance.
3. The method according to claim 2, characterized in that The method of optimizing Gaussian filtering by using a separation convolution strategy includes: Decompose the two-dimensional Gaussian kernel into one-dimensional kernels in the horizontal and vertical directions; Horizontal and vertical convolutions are performed separately through the parallel computing units of the Cambrian chip; Cache Gaussian kernel coefficients to reduce repeated calculations; According to the storage characteristics of the Cambrian MLU chip, the number of data copy rows is optimized to reduce the NRAM space occupied.
4. The method according to claim 1, characterized in that The edge detection is performed based on the optimized CANNY operator, including: Parallelize Gaussian filtering of input images and accelerate calculations through separate convolution strategies; The Sobel operator is used to parallelly calculate the image gradient magnitude and direction; Optimize non-maximum suppression through masked branch processing to eliminate the impact of conditional branches on parallel computing; Double threshold screening and edge connection algorithm are applied to generate the final edge detection results.
5. The method according to claim 4, characterized in that The masked branch processing optimizes non-maximum suppression, including: Calculate the gradient suppression results under all possible branch conditions; Generate a mask vector corresponding to the branch condition; The mask vector and the calculation result are logically operated through the vector operation of the Cambrian chip to merge the final results.
6. The method according to claim 1, characterized in that The asynchronous pipeline optimization technology includes: Dividing the image data into a plurality of data blocks; Execute data transmission and computing tasks in parallel through IO streams and Compute streams; The synchronization instructions of the Cambrian chip are used to coordinate dependencies between streams and reduce computing interruptions.
7. The method according to claim 1, characterized in that The memory alignment optimization includes: Adjust the image data storage format according to the vector calculation bit width of the Cambrian chip; Reduce memory access latency through data prefetching technology; Align image row data to adapt to the vectorized memory access instructions of the Cambrian chip.
8. A computer vision point and line feature extraction acceleration device based on Cambrian chip, characterized in that: include: A memory optimization module, configured to perform memory alignment optimization on the input image; SIFT optimization module, configured to use Cambrian chip to build Gaussian pyramid and difference pyramid and extract feature points; The CANNY optimization module is configured to perform Cambrian vector optimization operators such as Gaussian filtering, gradient calculation, and non-maximum suppression; The output module is configured to output the visual feature extraction result.
9. An electronic device, characterized in that: The method comprises a memory, a processor, an MLU chip and a computer program stored in the memory and executable on the processor, wherein the processor implements the method according to any one of claims 1 to 8 when executing the program.
10. A non-transitory computer-readable storage medium, characterized in that: The non-transitory computer-readable storage medium stores computer instructions, and the computer instructions are used to cause a computer to execute the method according to any one of claims 1 to 8.
11. A computer program product, characterized in that The method comprises computer program instructions, which, when executed on a computer, cause the computer to execute the method according to any one of claims 1 to 8.