Optimization method of VSIPL library for ARM platform

CN118916035BActive Publication Date: 2026-09-25XIDIAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410958060.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-17
Publication Date
2026-09-25
Estimated Expiration
2044-07-17

AI Technical Summary

Technical Problem

[0004]然而,目前已提出的VSIPL库优化方法,学习曲线越陡峭,开发难度较大

Benefits of technology

[0007]通过上述方式,选择在ARM平台上对VSIPL库中的存储的数据进行数据处理,可以提升VSIPL库的稳定性,在进行数据处理时,先在ARM平台上添加目标函数库的头文件,以链接目标函数库,并对目标函数库进行初始化,随后从VSIPL库中的目标视图对象中读取出待处理数据,确定待处理数据对应的处理需求,并基于待处理数据配置所需参数,再之后基于处理需求从目标函数库中调用对应的处理函数,并基于处理函数和所需参数,对待处理数据进行数据处理,得到数据处理结果,最后将数据处理结果按照VSIPL库规定的标准进行输出,并拷贝至VSIPL库的结果视图对象中,针对性地选择函数库并对函数库进行了优化,有效提升了VSIPL库处理大规模数据和实现复杂算法的能力,满足信号处理的实时性要求,同时增强了VSIPL库的普遍适用性,具有广泛的应用前景。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118916035B_ABST
    Figure CN118916035B_ABST
Patent Text Reader

Abstract

Embodiments of the present application relate to the technical field of data processing, in particular to a VSIPL library optimization method for ARM platform, which comprises the following steps: adding a header file of a target function library on the ARM platform to link the target function library and initialize the target function library; reading out to-be-processed data from a target view object in a VSIPL library, determining a processing requirement corresponding to the to-be-processed data, and configuring required parameters; calling a corresponding processing function from the target function library based on the processing requirement, and performing data processing on the to-be-processed data based on the processing function and the required parameters to obtain a data processing result; and outputting the data processing result according to a standard of the VSIPL library and copying the data processing result into a result view object of the VSIPL library. The method selects a function library and optimizes it in a targeted manner, stably improves the data processing speed and system performance, and has a wide application prospect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of this application relate to the field of data processing technology, and in particular to a method for optimizing the VSIPL library for the ARM platform. Background Technology

[0002] VSIPL (Vector Signal Image Processing Library) is a vector, signal, and image processing library primarily used for advanced embedded signal, information, and image processing. By providing a set of open standard application programming interfaces (APIs) and a common computational library for different platforms, VSIPL decouples software applications from hardware resources, enhancing application portability. Developers only need to use the computational library provided by VSIPL to design and develop algorithm components without worrying about hardware details. Furthermore, through code reuse, it effectively reduces system development time and extends the software lifecycle.

[0003] While the VSIPL library offers rich signal and image processing capabilities, it still has several shortcomings. For example, when handling complex algorithms in fields such as radar, wireless communication, and multimedia processing, its performance may be affected by memory access latency or computational bottlenecks, preventing it from delivering consistently high performance. The VSIPL library performs well with small datasets, but as the data size increases, the required memory and computational resources also increase, leading to memory bandwidth bottlenecks or processor overload, thus impacting its performance. Furthermore, the performance of the VSIPL library is highly dependent on hardware characteristics; its performance can vary significantly across different hardware platforms. Therefore, selecting a suitable hardware platform, using targeted optimization techniques, and employing appropriate algorithms are crucial for improving the performance and stability of the VSIPL library.

[0004] However, the steeper the learning curve of the proposed VSIPL library optimization methods, the more difficult they are to develop. Summary of the Invention

[0005] In view of this, embodiments of this application propose a VSIPL library optimization method for the ARM platform. By selecting the ARM platform to process the data in the VSIPL library, and selectively optimizing function libraries, the data processing speed and system performance are steadily improved, which has broad application prospects.

[0006] In a first aspect, embodiments of this application propose a VSIPL library optimization method for the ARM platform, applicable to data processing of data stored in the VSIPL library on the ARM platform. The method includes: adding a header file of a target function library to the ARM platform to link the target function library and initializing the target function library; wherein the target function library is an NE10 function library or an OpenBLAS function library; reading the data to be processed from the target view object in the VSIPL library, determining the processing requirements corresponding to the data to be processed, and configuring the required parameters based on the data to be processed; calling the corresponding processing function from the target function library based on the processing requirements, and processing the data to be processed based on the processing function and the required parameters to obtain the data processing result; outputting the data processing result according to the VSIPL library standard, and copying the output to the result view object of the VSIPL library.

[0007] By employing the above method, processing data stored in the VSIPL library on the ARM platform can improve the stability of the VSIPL library. During data processing, the header file of the target function library is first added to the ARM platform to link the target function library and initialize it. Then, the data to be processed is read from the target view object in the VSIPL library, the processing requirements corresponding to the data are determined, and the necessary parameters are configured based on the data. Next, the corresponding processing function is called from the target function library based on the processing requirements, and the data to be processed is processed based on the processing function and the necessary parameters to obtain the data processing result. Finally, the data processing result is output according to the standards specified by the VSIPL library and copied to the result view object of the VSIPL library. This targeted selection and optimization of the function library effectively improves the VSIPL library's ability to process large-scale data and implement complex algorithms, meeting the real-time requirements of signal processing. It also enhances the universal applicability of the VSIPL library and has broad application prospects.

[0008] Optionally, the VSIPL library encapsulates memory management through object-based design. The smallest storage object in the VSIPL library is a data array, and several data arrays form a block object. The block object is a contiguous memory region used to store data. Several block objects can be constructed as a view object in the form of vectors, matrices, or heights. Each view object has a specific offset, length, and stride. Accessing the data arrays stored in the VSIPL library is indirectly achieved through view objects and block objects. If the stride of the data array to be accessed is 1 and the complex data storage method is interleaved, the data pointer of the data array to be accessed is directly provided to the library function interface. If the stride of the data array to be accessed is not 1 and the complex data storage method is block storage, the data array to be accessed needs to be rearranged to achieve memory alignment, and the memory-aligned data pointer is provided to the library function interface.

[0009] Optionally, configuring the required parameters based on the data to be processed includes: obtaining the required parameters of the data to be processed from the information carried by the target view object and configuring them; or, obtaining the required parameters of the data to be processed and configuring them through a data parameter structure specific to the VSIPL library.

[0010] Optionally, the ARM platform supports the NEON instruction set, and the processor deployed on the ARM platform supports NEON. If the target function library does not contain a processing function corresponding to the processing requirement, the method further includes: adding a header file containing arm_neon.h to the source file, wherein the header file containing arm_neon.h defines all NEON intrinsics; aligning the data to be processed to 128 bits to obtain aligned data; writing a processing function corresponding to the processing requirement using NEON intrinsics, and processing the aligned data based on the written processing function and the required parameters to obtain a data processing result; outputting the data processing result according to the VSIPL library standard, and copying the output to the result view object of the VSIPL library. For processing functions not found in the target function library, this application supports writing functions based on NEON, thereby better optimizing the VSIPL library and further improving the performance, stability, and universal applicability of the VSIPL library.

[0011] Optionally, the NEON instruction set can be enabled at compile time using either the -mfpu=neon compilation option or the -mfloat-abi=softfp compilation option. Choosing the appropriate compilation option can fully utilize the performance of the NEON instructions.

[0012] Optionally, if data processing related to signal processing is required, the NE10 function library is selected as the target function library; if data processing related to matrix operations is required, the OpenBLAS function library is selected as the target function library.

[0013] Optionally, after outputting the data processing results according to the VSIPL library standard and copying the output to the result view object of the VSIPL library, the method further includes: releasing the target function library. Releasing the target function library immediately after completing one data processing operation can save resources and quickly prepare for the next data processing operation.

[0014] Secondly, embodiments of this application propose a VSIPL library optimization system for the ARM platform. The system includes a function library linking module for adding header files of a target function library to the ARM platform to link the target function library and initialize it. The target function library is either an NE10 function library or an OpenBLAS function library. A data reading module is used to read data to be processed from a target view object in the VSIPL library, determine the processing requirements corresponding to the data to be processed, and configure the required parameters based on the data to be processed. A data processing module is used to call the corresponding processing function from the target function library based on the processing requirements, and perform data processing on the data to be processed based on the processing function and the required parameters to obtain the data processing result. A data output module is used to output the data processing result according to the VSIPL library standard and copy the output to the result view object of the VSIPL library.

[0015] Thirdly, embodiments of this application provide an electronic device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the VSIPL library optimization method for the ARM platform as described in the first aspect above.

[0016] Fourthly, embodiments of this application propose a computer-readable storage medium storing a computer program that, when executed by a processor, can implement the VSIPL library optimization method for the ARM platform as described in the first aspect above.

[0017] It is understood that the beneficial effects of the second to fourth aspects mentioned above can be found in the relevant descriptions in the first aspect mentioned above, and will not be repeated here. Attached Figure Description

[0018] Figure 1This is a flowchart illustrating a VSIPL library optimization method for the ARM platform, provided in one embodiment of this application.

[0019] Figure 2 This is a code diagram of a data array provided in one embodiment of this application;

[0020] Figure 3 This is a flowchart provided by one embodiment of the present application, illustrating the process of writing processing functions corresponding to processing requirements and using the written processing functions to perform data processing;

[0021] Figure 4 This is a schematic diagram of the component structure of a VSIPL library optimization system for the ARM platform provided in another embodiment of this application;

[0022] Figure 5 This is a schematic diagram of the structure of an electronic device provided in another embodiment of this application. Detailed Implementation

[0023] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the various embodiments of this application will be described in detail below with reference to the accompanying drawings. However, those skilled in the art will understand that many technical details have been provided in the various embodiments of this application to help readers better understand this application. However, the technical solutions claimed in this application can be implemented even without these technical details and various changes and modifications based on the following embodiments. The division of the various embodiments below is for the convenience of description and should not constitute any limitation on the specific implementation of this application. The various embodiments can be combined with and referenced by each other without contradiction.

[0024] Optimization of the VSIPL library on the ARM platform can be achieved using NEON technology. NEON technology provides a dedicated extension to the ARM instruction architecture, namely additional instructions that can execute mathematical operations in parallel on multiple data streams. NEON accelerates data operations by providing 32 128-bit vector registers, each register corresponding to multiple data channels, and supports simultaneous operations on these multi-channel data using SIMD instructions. The use of NEON technology can be referenced as follows:

[0025] 1) Compiler-automatic vector optimization

[0026] Automatic vector optimization by the compiler refers to the compiler automatically identifying and utilizing the SIMD architecture for instruction-level parallel processing during program compilation. Common optimization aspects include vectorization, loop unrolling, instruction merging, data layout adjustments, and control flow optimization. In practice, these optimizations can usually be achieved through specific compiler options.

[0027] 2) Manually assemble NEON programs

[0028] For experienced programmers, hand-coded NEON assemblers can achieve very high performance.

[0029] 3) Supports NEON's open-source libraries:

[0030] Common libraries include NE10, OpenBLAS, Eigen, etc.

[0031] 4) NEON intrinsics

[0032] NEON intrinsics are a set of C and C++ functions defined in arm_neon.h, supported by both the ARM compiler and GCC. The compiler replaces these function calls with an appropriate NEON instruction or a sequence of NEON instructions. Programmers can focus on the high-level behavior of the algorithm without worrying about the low-level implementation.

[0033] These four methods have different levels of optimization. The more fundamental the optimization method, the better the effect, but the steeper the learning curve and the greater the development difficulty. There are still problems that need to be solved.

[0034] The first method offers higher development efficiency. Developers don't need to explicitly use NEON intrinsics or assembly language; by enabling the O3 optimization option during compilation, the compiler automatically identifies and optimizes the code. The optimized code remains a high-level language, making it easy to understand and maintain. However, its performance is inferior to manual optimization. The code cannot be automatically optimized into NEON instructions, and the compiler needs to analyze the code and apply vector optimizations, often resulting in additional time consumption and increased compilation time.

[0035] The second method, through manually writing assembly code, allows for precise control over the execution of NEON instructions, thus achieving optimal performance. However, this method is more difficult and costly to develop, requiring developers to have in-depth knowledge of programming languages ​​and architectures. Assembly code is also difficult to read and maintain, and prone to errors. Furthermore, assembly code is highly dependent on specific hardware platforms, resulting in poor portability.

[0036] The third approach typically utilizes the NEON instruction set for high optimization and features simplified, easy-to-use API interfaces. This aims to make it easier for developers to call the functions of these libraries, while also ensuring high portability across different platforms. Specifically, the NE10 library has been specifically optimized for signal processing functions on the ARM platform, and the OpenBLAS library optimizes matrix-related operations using basic linear algebra library interfaces and linear algebra function library interfaces.

[0037] The fourth method provides direct control over the NEON instruction set, allowing for performance tuning for specific algorithms and precise control over the execution of NEON instructions, thus offering high flexibility. Programmers using this method only need to focus on the high-level behavior of the algorithm, without needing to worry about the underlying code implementation, making it simpler than manually writing assembly language.

[0038] However, the third and fourth methods are not yet mature and are currently only in the theoretical stage.

[0039] In conclusion, considering factors such as performance improvement, portability across different platforms, and code operability, choosing the NEON open-source library and NEON intrinsics to optimize the VSIPL library is more appropriate.

[0040] Based on this, one embodiment of this application proposes a VSIPL library optimization method for the ARM platform, which is applied to electronic devices, wherein the electronic devices can be terminals or servers. This embodiment and the following embodiments all use servers as examples for illustration. The implementation details of the VSIPL library optimization method for the ARM platform proposed in this embodiment are described in detail below. The following content is only for the convenience of understanding and is not necessary for implementing this solution.

[0041] The specific process of the VSIPL library optimization method for the ARM platform proposed in this embodiment is as follows: Figure 1 As shown, it includes:

[0042] Step 101: Add the header file of the target function library to the ARM platform to link the target function library and initialize the target function library.

[0043] In the specific implementation, the server first needs to add the header file of the target function library to the ARM platform to link the target function library and initialize it. The target function library can be either the NE10 library or the OpenBLAS library. If signal processing-related data processing is required, the NE10 library is selected as the target function library; if matrix operation-related data processing is required, the OpenBLAS library is selected.

[0044] In one example, if the NE10 function library is chosen as the target function library, the header files that need to be added on the ARM platform include ne10_types.h and ne10_fft.h.

[0045] In one example, if the OpenBLAS function library is chosen as the target function library, the header files that need to be added on the ARM platform include cblas.h and lapack.h.

[0046] Step 102: Read the data to be processed from the target view object in the VSIPL library, determine the processing requirements corresponding to the data to be processed, and configure the required parameters based on the data to be processed.

[0047] In the actual implementation, after the server links to the target function library, it can read the data to be processed from the target view object in the VSIPL library, determine the processing requirements corresponding to the data to be processed, and configure the required parameters based on the data to be processed.

[0048] It's worth noting that the VSIPL library encapsulates memory management through object-based design. The smallest storage object in VSIPL is a data array. Several data arrays can form a block object, which is a contiguous memory region used to store data. Several block objects can be constructed into a view object as vectors, matrices, or heights. Each view object has a specific offset, length, and stride. Accessing data arrays stored in the VSIPL library requires indirect access through view objects and block objects. If the stride of the data array to be accessed is 1 and the complex data storage method is interleaved, the data pointer of the data array to be accessed is directly provided to the library function interface. If the stride of the data array to be accessed is not 1 and the complex data storage method is block storage, the data array to be accessed needs to be rearranged to achieve memory alignment. The memory-aligned data pointer is then provided to the library function interface to read the data to be processed.

[0049] In one example, the code structure for a data array can be as follows: Figure 2 As shown.

[0050] In one example, the required parameters for the data to be processed can be obtained and configured from the information carried by the target view object, or obtained and configured through a specific data parameter structure of the VSIPL library.

[0051] Step 103: Based on the processing requirements, call the corresponding processing function from the target function library, and based on the processing function and the required parameters, perform data processing on the data to be processed to obtain the data processing result.

[0052] In the specific implementation, after the server reads the data to be processed from the target view object in the VSIPL library, determines the processing requirements corresponding to the data to be processed, and configures the required parameters based on the data to be processed, it can call the corresponding processing function from the target function library based on the processing requirements corresponding to the data to be processed, and perform data processing on the data to be processed based on the processing function and the required parameters to obtain the data processing result.

[0053] The NE10 library provides a large number of signal processing-related functions, such as the ne10_fft_c2c_1d_float32_neon function for complex Fourier transform. The OpenBLAS library is an open-source, optimized BLAS library that provides a large number of high-performance linear algebra functions, including matrix multiplication, matrix inversion, and vector dot product, such as the cblas_sgemm function for matrix multiplication.

[0054] Step 104: Output the data processing results according to the VSIPL library standard, and copy the output to the result view object of the VSIPL library.

[0055] In practice, neither the Ne10 function library nor the OpenBLAS function library outputs the same standard as the VSIPL library. Such output cannot be stored in the VSIPL library. Therefore, the server needs to output the data processing results according to the VSIPL library standard and then copy the output to the result view object (data array) of the VSIPL library.

[0056] In one example, after the server outputs the data processing results according to the VSIPL library standard and copies the output to the VSIPL library's result view object, it also needs to release the target function library. Releasing the target function library immediately after completing a data processing operation conserves resources and quickly prepares for the next data processing iteration.

[0057] In this embodiment, data processing of the data stored in the VSIPL library is performed on the ARM platform, which improves the stability of the VSIPL library. During data processing, the header file of the target function library is first added to the ARM platform to link the target function library and initialize it. Then, the data to be processed is read from the target view object in the VSIPL library, the processing requirements corresponding to the data are determined, and the required parameters are configured based on the data. Next, the corresponding processing function is called from the target function library based on the processing requirements, and the data to be processed is processed based on the processing function and the required parameters to obtain the data processing result. Finally, the data processing result is output according to the VSIPL library standard and copied to the result view object of the VSIPL library. By specifically selecting and optimizing the function library, the ability of the VSIPL library to process large-scale data and implement complex algorithms is effectively improved, meeting the real-time requirements of signal processing. At the same time, the versatility of the VSIPL library is enhanced, giving it broad application prospects.

[0058] In one embodiment, the ARM platform supports the NEON instruction set, and the processor deployed on the ARM platform supports NEON. If the target function library does not contain a processing function corresponding to the processing requirements, the server can use methods such as... Figure 3 The steps shown involve writing and processing functions corresponding to the processing requirements, and then using these functions to process the data. Specifically, this includes:

[0059] Step 201: Add a header file containing arm_neon.h to the source file. The header file containing arm_neon.h defines all NEON intrinsics.

[0060] In a specific implementation, the server adds a header file containing arm_neon.h to the C or C++ source files. This header file, which includes arm_neon.h, defines all NEON intrinsics.

[0061] Step 202: Align the data to be processed by 128 bits to obtain the aligned data.

[0062] In its implementation, the NEON instruction requires that data access must be 128-bit aligned. In other words, the data used must be 128-bit aligned. Therefore, the server needs to align the data to be processed to 128-bit to obtain the aligned data, thereby improving the performance of the VSIPL library.

[0063] Step 203: Use NEON intrinsics to write processing functions corresponding to the processing requirements, and based on the written processing functions and required parameters, perform data processing on the aligned data to obtain the data processing results.

[0064] In the specific implementation, the server uses NEON intrinsics to write processing functions corresponding to the processing requirements. The written processing functions are usually used to perform vector operations. Based on the written processing functions and the required parameters, the server processes the aligned data and obtains the data processing results.

[0065] In one example, the NEON instruction set can be enabled at compile time via the -mfpu=neon compilation option or the -mfloat-abi=softfp compilation option, ensuring that the -O2 or higher optimization option is included in the compiler command line to fully utilize the performance of NEON instructions.

[0066] In one example, if the data processing results do not meet the requirements, meaning the performance improvement of the VSIPL library is insufficient, then the written processing functions need to be readjusted and optimized.

[0067] Step 204: Output the data processing results according to the VSIPL library standard, and copy the output to the result view object of the VSIPL library.

[0068] In practice, the output of NEON inline functions does not conform to the VSIPL library standard. Such output cannot be stored in the VSIPL library. Therefore, the server needs to output the data processing results according to the VSIPL library standard and then copy the output to the data array of the result view object in the VSIPL library.

[0069] In this embodiment, for processing functions not found in the target function library, it is supported to write functions based on NEON, thereby better optimizing the VSIPL library and further improving the performance, stability and universal applicability of the VSIPL library.

[0070] Another embodiment of this application proposes a VSIPL library optimization system for the ARM platform. The details of this ARM platform-oriented VSIPL library optimization system are described below. The following implementation details are provided for ease of understanding and are not essential for implementing this example. Figure 4 This is a schematic diagram of the VSIPL library optimization system for the ARM platform proposed in this embodiment, including: function library linking module 301, data reading module 302, data processing module 303, and data output module 304.

[0071] The function library linking module 301 is used to add the header file of the target function library on the ARM platform to link the target function library and initialize the target function library, which is either the NE10 function library or the OpenBLAS function library.

[0072] The data reading module 302 is used to read the data to be processed from the target view object in the VSIPL library, determine the processing requirements corresponding to the data to be processed, and configure the required parameters based on the data to be processed.

[0073] The data processing module 303 is used to call the corresponding processing function from the target function library based on the processing requirements, and to process the data to be processed based on the processing function and the required parameters to obtain the data processing result.

[0074] The data output module 304 is used to output the data processing results according to the VSIPL library standard and copy the output to the result view object of the VSIPL library.

[0075] It is not difficult to see that this embodiment is a system embodiment corresponding to the above method embodiments, and this embodiment can be implemented in conjunction with the above method embodiments. The relevant technical details and technical effects mentioned in the above embodiments are still valid in this embodiment, and will not be repeated here to reduce repetition. Accordingly, the relevant technical details mentioned in this embodiment can also be applied to the above embodiments.

[0076] It is worth mentioning that all modules involved in this embodiment are logical modules. In practical applications, a logical unit can be a physical unit, a part of a physical unit, or a combination of multiple physical units. Furthermore, to highlight the innovative aspects of this application, this embodiment does not introduce units that are not closely related to solving the technical problems proposed in this application; however, this does not mean that other units are absent in this embodiment.

[0077] To illustrate the superiority of the VSIPL library optimization method for the ARM platform proposed in this application, the performance comparisons before and after optimizing VSIPL library functions in different ways are listed below.

[0078] Table 1: Comparison of the vsip_ccfftop_f function before and after optimization using the Ne10 function library.

[0079]

[0080]

[0081] As can be seen from Table 1, the signal processing function vsip_ccfftop_f can be 3 to 5 times faster after optimization by the Ne10 function library.

[0082] Table 2: Comparison of the vsip_ccholsol_f function before and after optimization using the OpenBLAS function library.

[0083]

[0084] As can be seen from Table 2, the performance of the matrix operation function vsip_ccholsol_f drops significantly with large numbers of points. After optimization by the OpenBLAS function library, the speed can be improved by 21 times.

[0085] Table 3: Comparison of compiler optimization and NEON intrinsic optimization results for the vsip_cvdot_f function.

[0086] 1024 2.6 0.9 2.8 2048 5.3 2.0 2.7 4096 10.9 3.9 2.8 8192 21.0 7.6 2.8

[0087] As can be seen from Table 3, the implementation using NEON inline functions is more effective than open compiler optimization, achieving a speedup of 2.8 times, and there is still room for further optimization.

[0088] The VSIPL library optimization method for the ARM platform proposed in this application has broad application prospects. Besides playing a crucial role in software-defined radar, it can also provide efficient, real-time signal processing capabilities in other fields, especially in applications requiring extensive mathematical operations, significantly improving processing speed and system performance. For example, in wireless communication, the NEON-optimized VSIPL library can be used to quickly implement signal processing algorithms such as Fast Fourier Transform, digital filters, and channel encoding / decoding. These algorithms are crucial in baseband processing, improving the data processing speed and efficiency of wireless communication systems. In multimedia processing, such as audio and video processing, the NEON-optimized VSIPL library can be used to implement audio effects processing, video encoding and decoding, and image processing. These applications require extensive mathematical operations, and NEON's vector instructions can significantly improve computational speed. Simultaneously, the VSIPL library can also be used in medical imaging, such as in computed tomography (CT) and magnetic resonance imaging (MRI), which require substantial image processing and analysis. The NEON-optimized VSIPL library can accelerate these calculations and shorten image processing time.

[0089] As ARM processors become increasingly widely used in various devices, the importance of NEON-optimized VSIPL library will become increasingly prominent.

[0090] Another embodiment of this application provides an electronic device, such as Figure 5 As shown, it includes: at least one processor 401; and a memory 402 communicatively connected to the at least one processor 401; wherein the memory 402 stores instructions executable by the at least one processor 401, the instructions being executed by the at least one processor 401 to enable the at least one processor 401 to execute the VSIPL library optimization method for the ARM platform as described in the above method embodiments.

[0091] The memory and processor are connected via a bus, which can include any number of interconnecting buses and bridges, connecting various circuits of one or more processors and the memory. The bus can also connect other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art and will not be described further herein. The bus interface provides an interface between the bus and the transceiver. The transceiver can be a single element or multiple elements, such as multiple receivers and transmitters, providing a unit for communicating with various other devices over a transmission medium. Data processed by the processor is transmitted over the wireless medium via an antenna, which further receives data and transmits it to the processor.

[0092] The processor manages the bus and general processing, and also provides various functions, including timing, peripheral interfaces, voltage regulation, power management, and other control functions. Memory is used to store data used by the processor during operation.

[0093] Another embodiment of this application proposes a computer-readable storage medium storing a computer program. When executed by a processor, the computer program implements the VSIPL library optimization method for the ARM platform as described in the above method embodiments.

[0094] That is, those skilled in the art will understand that all or part of the steps in the methods of the above embodiments can be implemented by a program instructing related hardware. This program is stored in a storage medium and includes several instructions to cause a device (which may be a microcontroller, chip, etc.) or processor to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, a portable hard drive, ROM (Read-Only Memory), RAM (Random Access Memory), a magnetic disk, or an optical disk.

[0095] Those skilled in the art will understand that the above embodiments are specific embodiments for implementing this application, and in practical applications, various changes can be made to them in form and detail without departing from the spirit and scope of this application.

Claims

1. A VSIPL library optimization method for the ARM platform, applicable to data processing of data stored in the VSIPL library on the ARM platform, characterized in that, The method includes: Add the header file of the target function library to the ARM platform to link the target function library and initialize the target function library; wherein the target function library is the NE10 function library or the OpenBLAS function library; Read the data to be processed from the target view object in the VSIPL library, determine the processing requirements corresponding to the data to be processed, and configure the required parameters based on the data to be processed; Based on the processing requirements, the corresponding processing function is called from the target function library, and based on the processing function and the required parameters, the data to be processed is processed to obtain the data processing result. The data processing results are output according to the standards of the VSIPL library, and the output is copied to the result view object of the VSIPL library; The VSIPL library encapsulates memory management through object-based design. The smallest storage object in the VSIPL library is a data array. Several data arrays form a block object, which is a contiguous memory area used to store data. Several block objects can be constructed as a view object in the form of vectors, matrices, or heights. Each view object has a specific offset, length, and stride. Accessing the data array stored in the VSIPL library is achieved indirectly through view objects and block objects. If the step size of the data array to be accessed is 1 and the complex data storage method is interleaved storage, the data pointer of the data array to be accessed is directly provided to the library function interface. If the step size of the data array to be accessed is not 1 and the complex data storage method is block storage, the data array to be accessed needs to be rearranged to achieve memory alignment, and the memory-aligned data pointer is provided to the library function interface. The ARM platform supports the NEON instruction set, and the processor deployed on the ARM platform supports NEON. If the target function library does not contain a processing function corresponding to the processing requirement, the method further includes: Add a header file containing arm_neon.h to the source file. This header file defines all NEON intrinsics. The data to be processed is aligned to 128 bits to obtain the aligned data; Using NEON intrinsics, write processing functions corresponding to the processing requirements, and based on the written processing functions and the required parameters, perform data processing on the aligned data to obtain the data processing results; The data processing results are output according to the standards of the VSIPL library, and the output is copied to the result view object of the VSIPL library.

2. The VSIPL library optimization method for the ARM platform as described in claim 1, characterized in that, The parameters required for configuring the data to be processed include: Obtain the required parameters for the data to be processed from the information carried by the target view object and configure them; Alternatively, the required parameters of the data to be processed can be obtained and configured through the specific data parameter structure of the VSIPL library.

3. The VSIPL library optimization method for the ARM platform as described in claim 1, characterized in that, During compilation, the NEON instruction set is enabled via the -mfpu=neon compilation option or the -mfloat-abi=softfp compilation option.

4. The VSIPL library optimization method for the ARM platform as described in any one of claims 1 to 3, characterized in that, If data processing related to signal processing is required, the NE10 function library should be selected as the target function library; if data processing related to matrix operations is required, the OpenBLAS function library should be selected as the target function library.

5. The VSIPL library optimization method for the ARM platform as described in any one of claims 1 to 3, characterized in that, After outputting the data processing results according to the VSIPL library standard and copying the output to the result view object of the VSIPL library, the method further includes: Release the target function library.

6. A VSIPL library optimization system for the ARM platform, used to implement the VSIPL library optimization method for the ARM platform as described in any one of claims 1 to 5, characterized in that, The system includes: a function library linking module, a data reading module, a data processing module, and a data output module; A function library linking module is used to add the header file of the target function library on the ARM platform to link the target function library and initialize the target function library, wherein the target function library is the NE10 function library or the OpenBLAS function library; The data reading module is used to read the data to be processed from the target view object in the VSIPL library, determine the processing requirements corresponding to the data to be processed, and configure the required parameters based on the data to be processed. A data processing module is used to call the corresponding processing function from the target function library based on the processing requirements, and to process the data to be processed based on the processing function and the required parameters to obtain the data processing result. The data output module is used to output the data processing results according to the VSIPL library standard and copy the output to the result view object of the VSIPL library.

7. An electronic device, characterized in that, include: At least one processor; And, a memory communicatively connected to the at least one processor; The memory stores instructions that can be executed by the at least one processor, which are executed by the at least one processor to enable the at least one processor to perform the VSIPL library optimization method for the ARM platform as described in any one of claims 1 to 5.

8. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the VSIPL library optimization method for the ARM platform as described in any one of claims 1 to 5.