Image rotation display method and device

By using NEON instructions to perform image rotation operations in parallel in ARM architecture devices, the problem of low image rotation display efficiency in embedded devices is solved, and efficient image rotation processing and smooth user experience is achieved.

CN119941519APending Publication Date: 2025-05-06QIAOHONG NUMERICAL CONTROL TECH SHANGHAI
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510019495.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-06
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

In embedded devices, image rotation display processing efficiency is low, especially in large-resolution screens and ARM architecture devices with no graphics processors configured, where frame rate drops are severe.

Method used

By obtaining the original image and the target rotation angle, a sliding window based on a preset size selects pixel blocks from the original image, and uses the NEON instruction to perform transpose operations on pixel points in the same pixel block in parallel, and store and display the rotated image.

Benefits of technology

It achieves the efficiency of image rotation processing, can provide smooth user interaction and real-time refresh experience in GPU-free devices, solving the problem of frame rate drop.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119941519A_ABST
    Figure CN119941519A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses an image rotation display method and device which are used for improving the efficiency of image rotation display. The scheme provided by the invention is applied to a processor with an ARM architecture, and comprises the steps of obtaining an original image and a target rotation angle, the original image comprising a plurality of pixel points arranged in an array; pixel blocks are sequentially selected from the original image based on a sliding window of a preset size, and transposition operation is executed on all pixel points belonging to the same pixel block in parallel based on an NEON instruction; storing each pixel point in the transposed pixel block to a buffer area according to the matching position of the target rotation angle; and displaying the rotated image based on the pixel points stored in the buffer area.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing, and in particular to a method and device for image rotation display. Background Art

[0002] Screen rotation is a common requirement in industrial control, portable terminals, vehicle-mounted systems and other application scenarios. Due to limited resource systems such as embedded devices, the processing efficiency of screen rotation display is often low.

[0003] When rotating an image on a high-resolution screen, it is usually necessary to rearrange and calculate a large number of pixels. The processing time of such algorithms, such as copying and storing line by line, is too long, which may cause the frame rate to drop. In ARM (Advanced Reduced Instruction Set Computer Machines) architecture devices without a graphics processor, it is particularly difficult to efficiently implement the image rotation display function.

[0004] How to improve the efficiency of image rotation display is a technical problem to be solved by this application. Summary of the invention

[0005] The purpose of the embodiments of the present application is to provide a method and device for image rotation display, so as to improve the efficiency of image rotation display.

[0006] In a first aspect, a method for image rotation display is provided, which is applied to a processor having an ARM architecture, comprising: Acquire an original image and a target rotation angle, wherein the original image includes a plurality of pixel points arranged in an array; Selecting pixel blocks from the original image in sequence based on a sliding window of a preset size, and performing a transposition operation on each pixel point in the same pixel block in parallel based on a NEON instruction; storing each pixel point in the transposed pixel block in a buffer according to the matching position of the target rotation angle; The rotated image is displayed based on the pixels stored in the buffer.

[0007] In a second aspect, a device for image rotation display is provided, which is applied to a processor having an ARM architecture, and includes: An acquisition module is used to acquire an original image and a target rotation angle, wherein the original image includes a plurality of pixel points arranged in an array; A transposition module, which sequentially selects pixel blocks from the original image based on a sliding window of a preset size, and performs a transposition operation on each pixel point in the same pixel block in parallel based on a NEON instruction; A storage module stores each pixel point in the transposed pixel block in a buffer according to a matching position of the target rotation angle; A display module displays the rotated image based on the pixels stored in the buffer.

[0008] According to a third aspect, an electronic device is provided. The electronic device comprises a processor, a memory, and a computer program stored in the memory and executable on the processor. When the computer program is executed by the processor, the steps of the method according to the first aspect are implemented.

[0009] According to a fourth aspect, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps of the method according to the first aspect are implemented.

[0010] In a fifth aspect, a computer program product is provided, the computer program product comprising a non-transitory computer-readable storage medium storing a computer program, the computer program being operable to cause a computer to execute part or all of the steps of the method of the first aspect.

[0011] The embodiment of the present application is applied to a processor with an ARM architecture. First, an original image and a target rotation angle are obtained, and the original image includes a plurality of pixel points arranged in an array; then, pixel blocks are sequentially selected from the original image based on a sliding window of a preset size, and a transposition operation is performed in parallel on each pixel point belonging to the same pixel block based on a NEON instruction; then, each pixel point in the transposed pixel block is stored in a buffer according to the matching position of the target rotation angle; then, the rotated image is displayed based on the pixel points stored in the buffer. This scheme applies the NEON instruction of an ARM architecture processor, and performs a transposition operation on the pixel points belonging to the same pixel block in parallel through a single instruction, which can realize batch transposition of pixel points and improve the efficiency of image rotation processing. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings: Figure 1a This is one of the flowcharts of a method for image rotation display in one embodiment of the present application; Figure 1b This is a second flow chart of a method for image rotation display according to an embodiment of the present application; Figure 2 This is a flowchart of a method for image rotation display according to an embodiment of the present application; Figure 3This is a fourth flow chart of a method for image rotation display according to an embodiment of the present application; Figure 4a This is a fifth flow chart of a method for image rotation display according to an embodiment of the present application; Figure 4b This is a sixth flow chart of a method for image rotation display according to an embodiment of the present application; Figure 5 This is a seventh flow chart of a method for image rotation display according to an embodiment of the present application; Figure 6 It is a structural schematic diagram of a device for image rotation display according to an embodiment of the present application. DETAILED DESCRIPTION

[0013] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application. The numbering of the drawings in this application is only used to distinguish the various steps in the scheme, and is not used to limit the execution order of the various steps. The specific execution order is subject to the description in the specification.

[0014] In embedded development scenarios, screen rotation often relies on hardware functions, which increases the complexity of software implementation and hardware costs. In addition, the implementation of rotation algorithms varies greatly between different devices and systems, and image rotation methods have the defects of poor versatility and low efficiency.

[0015] In some application scenarios, screen rotation can be accelerated based on GPU (Graphics Processing Unit) hardware. For example, in modern embedded devices such as smartphones and tablets, screen rotation often relies on GPU acceleration. Among them, GPU can accelerate the conversion of pixel position calculation through parallel processing capabilities, thereby achieving efficient screen rotation. However, in actual applications, the performance of GPU is limited, and it is difficult to efficiently implement real-time screen rotation function. In addition, some devices do not include GPU, and these devices without GPU cannot efficiently implement real-time screen rotation function.

[0016] In terms of hardware acceleration, the hardware acceleration function often requires specific hardware and corresponding hardware driver support, which also increases the difficulty of implementing the screen rotation function.

[0017] In addition, image rotation can also be achieved based on the CPU (Central Processing Unit). Specifically, the mapping position of each pixel can be calculated and rotated pixel by pixel. However, this method is difficult to ensure efficient real-time pixel rotation on resource-constrained devices, and the user experience is poor. Moreover, this pixel-by-pixel rotation method requires a large amount of calculation. In application scenarios with high image resolution, limited CPU computing resources cannot efficiently process large amounts of image data.

[0018] In order to solve the problems existing in the related art, the embodiment of the present application provides a method for image rotation display, which is applied to a processor with an ARM architecture. The ARM architecture is applied to a processor of an embedded system and has the advantages of high efficiency computing power and low power consumption. Figure 1a As shown, the method provided in the embodiment of the present application includes: S11: Acquire an original image and a target rotation angle, wherein the original image includes a plurality of pixel points arranged in an array.

[0019] Among them, the original image refers to the image to be rotated, and the original image includes pixels arranged in an array. In practical applications, the resolution of the original image can be, for example, 800×600, 1024×768, 1366×768, 1920×1080, etc. The resolution can represent the structure of the pixel array arrangement in the original image, that is, the number of pixels arranged in the horizontal and vertical directions of the original image.

[0020] The target rotation angle refers to the angle at which the original image needs to be rotated. The target rotation angle can be specified by the user or automatically generated by the device according to actual display requirements. For example, the rotation angle includes a clockwise or counterclockwise rotation direction and a desired angle, such as 90°, 180°, 270°, etc.

[0021] S12: sequentially selecting pixel blocks from the original image based on a sliding window of a preset size, and performing a transposition operation on each pixel point in the same pixel block in parallel based on a NEON instruction.

[0022] In this step, pixel blocks are sequentially selected from the original image based on the sliding window, and the selected pixel blocks have a preset size. For example, assuming that the preset size is 8×8, that is, 8 pixels in the horizontal direction and 8 pixels in the vertical direction. Then the pixel block selected from the original image based on the sliding window of the preset size includes 8×8 pixels.

[0023] The order of selecting pixel blocks based on the sliding window can be preset, for example, pixel blocks are selected in sequence along the horizontal direction from the initial corner of the original image until all pixel points in the current direction are exhausted or the remaining pixel points are insufficient to fill a sliding window, and then the sliding window is moved in the vertical direction based on the initial corner to the sliding window containing 8×8 unselected pixel points, and then pixel blocks are selected in sequence along the horizontal direction. Pixel blocks are selected in sequence from the original image through a small cycle selected in sequence along the horizontal direction and a large cycle selected in sequence along the vertical direction.

[0024] It should be understood that in actual applications, the order in which the sliding window selects pixel blocks can also be freely set, for example, pixel blocks can be selected in an S-shaped route, a circular route, etc.

[0025] For the pixel block selected by the sliding window, the transposition operation is performed in parallel on each pixel point in the pixel block based on the NEON instruction. Among them, the NEON instruction is an instruction in the NEON instruction set. The NEON instruction set is an advanced SIMD (Single Instruction Multiple Data) extension in the ARM architecture, which can efficiently realize data parallel processing and can process multiple data operations in a single cycle. The above-mentioned SIMD is a parallel computing model that allows a single instruction to operate multiple data units simultaneously.

[0026] In this example, the transposition operation is performed in parallel on each pixel point in the pixel block based on the NEON instruction, so that the pixel point transposition is performed in batches in parallel, and the efficient transposition of the pixel points in the pixel block is realized, the efficiency of the image transposition processing is improved, and then it is beneficial to improve the overall image rotation efficiency.

[0027] Optionally, in this step, after selecting a pixel block through the sliding window, a transposition operation is performed on each pixel point in the selected pixel block in parallel through a NEON instruction. After the transposition operation is completed on the pixel block, the next pixel block is selected through the sliding window, and the transposition operation is performed on the newly selected pixel block through a NEON instruction.

[0028] Optionally, in one scenario, multiple pixel blocks can be selected through a sliding window, and then the transposition operation can be performed on the multiple pixel blocks in parallel through NEON instructions. The multiple pixel points in any pixel block can be transposed in parallel based on the NEON instructions, thereby further improving the efficiency of the overall image transposition processing through parallel execution.

[0029] S13: storing each pixel point in the transposed pixel block in a buffer according to a matching position of the target rotation angle.

[0030] In this step, the pixel points in the transposed pixel block are stored based on the target rotation angle. The storage position of the pixel points matches the target rotation angle. Specifically, the transposed pixel block is stored according to the position of the pixel block in the image after the rotation according to the target rotation angle. For example, assuming that the target rotation angle is 90° clockwise, the transposed pixel block is horizontally flipped and stored. There is a preset matching relationship between 90° clockwise and horizontal flipping. Based on the preset matching relationship, the transposed pixel block is stored in the buffer, which can be used to efficiently display the rotated image.

[0031] S14: Displaying the rotated image based on the pixels stored in the buffer.

[0032] In practical applications, if the buffer stores a partial rotated image, the stored partial rotated image can be directly displayed, and the displayed image can be gradually updated with the progress of transposition and storage of the remaining pixel blocks until the complete rotated image is displayed. Alternatively, if the buffer stores a complete rotated image, the rotated image can be displayed as a whole.

[0033] The solution provided by the embodiment of the present application can be widely used in devices without GPU, and can also be applied to devices with insufficient processing resources such as CPU and GPU. This solution uses the NEON instruction of the ARM architecture processor to perform transposition operations on pixels belonging to the same pixel block in parallel through a single instruction, which can realize batch transposition of pixels and improve the efficiency of image rotation processing. Software acceleration based on the NEON instruction set can efficiently perform image rotation display, which is conducive to realizing real-time image rotation display, and can still provide smooth user interaction and real-time refresh experience in devices without GPU.

[0034] Optionally, in actual application scenarios, further display processing may be performed on the rotated image, such as horizontal or vertical mirror flipping, adding blur or clear display effects, color cast processing, etc.

[0035] The following provides a specific implementation of the above-mentioned method for rotating and displaying an image in combination with an example. Figure 1b .

[0036] The solution provided in the embodiment of the present application can be implemented based on a custom cross compiler. A cross compiler is a compiler that runs on one architecture (such as x86 or x86_64) but generates executable code for another architecture (such as ARM, MIPS, RISC-V, etc.). It can be used to implement embedded development and supports compiling programs for target devices in a resource-rich development environment.

[0037] Among them, the custom cross-compiler enables the optimized code of the NEON instruction set to run on a wide range of devices on the ARM platform. It is not limited to the hardware platform and can be widely applied to various ARM architectures (such as ARMv7, ARMv8, etc.). At the same time, it reduces the complexity of the solution implementation and enhances versatility.

[0038] In this solution, a cross-compiler can be built to generate a compiler tool chain that supports the NEON instruction set and hard floating point operations (Vector Floating Point, VFP). The code of high-level languages ​​(such as C, C++) is compiled into a binary file suitable for the target device to run, and the NEON instruction set and hard floating point acceleration are used to optimize the computing performance.

[0039] Specifically, a cross-compiler production tool such as crosstool-ng can be used to produce the above-mentioned cross-compiler that supports hard floating-point operations and NEON instruction sets, ensuring that the compilation environment can generate a cross-tool chain optimized for the target platform.

[0040] Then, execute the Buildroot configuration. Specifically, set the target options in Buildroot to support NEON SIMD extension and VFP extension, select EABIhf embedded binary application interface support, set the floating point operation strategy to NEON, and ensure that the compiled program can make full use of NEON instructions. Among them, Buildroot is a tool chain and file system building tool for embedded systems, which can support the compilation and generation of libraries and applications for specific architectures.

[0041] Among them, Buildroot can be used to configure and build the file system and related library support of the target device, enable NEON SIMD and VFP extensions, configure hard floating-point ABI (such as EABIhf), ensure the use of hardware acceleration in floating-point operations, and generate a runtime environment and library compatible with the NEON instruction set.

[0042] Then, kernel configuration is performed. Specifically, the Linux kernel is configured to support NEON and VFP hardware acceleration functions, ensuring that the operating system on the device can call the hardware NEON acceleration module and floating point computing unit, ensuring the correctness of NEON hardware acceleration, and ensuring that the operating system can perform image rotation calculations in NEON mode. Among them, Linux kernel support refers to enabling the function of NEON instruction set support in the Linux kernel, so that applications can call related instructions.

[0043] Afterwards, the screen rotation algorithm is executed, that is, the image rotation display method described in the above embodiment of the present application is implemented through the NEON SIMD instruction set library, image optimization processing is achieved, and the use of GPU hardware acceleration is avoided, so as to achieve energy saving and high efficiency, and ensure user experience and real-time refresh in the absence of GPU. Through vectorized calculation, pixels are calculated in parallel, the calculation time is reduced, memory access is optimized, latency is reduced, and a smooth rotation process is ensured.

[0044] Based on the solution provided in the above embodiment, optionally, Figure 2 As shown, in the above step S12, pixel blocks are sequentially selected from the original image based on a sliding window of a preset size, including: S21: sequentially selecting pixel blocks from an initial position of the original image along a first direction with a first preset step length based on the sliding window.

[0045] The above-mentioned initial position, first preset step length and first direction can be preset based on actual needs.

[0046] For example, for a rectangular original image, the position of the sliding window located at a corner of the original image is set as the initial position, and the sliding window at the initial position includes the pixel points of the first row and the first column of the original image.

[0047] The first preset step length can be set as the length of the sliding window in the first direction, and the first preset step length can specifically be a preset number of pixels, that is, each time the sliding window slides along the first direction, the preset number of pixels is moved to select pixels that have not been selected.

[0048] The first direction may be a row direction or a column direction, and may be a direction from a corner of the original image corresponding to the initial position to an adjacent corner. For example, from the upper left corner to the upper right corner of the original image is a row direction. From the upper left corner to the lower left corner of the original image is a column direction.

[0049] In this step, pixel blocks are selected sequentially from the original image based on the sliding window, and pixel blocks are selected step by step starting from the initial position, which can ensure that the pixel blocks selected along the first direction are orderly and not missed.

[0050] S22: If the number of unselected pixel points in the first direction is less than the first preset step size, the sliding window is moved along the second direction with a second preset step size based on the initial position, and pixel blocks are selected sequentially along the first direction until the number of unselected pixel points in the second direction is less than the second preset step size, wherein the first preset step size is the length of the sliding window of the preset size in the first direction, the second preset step size is the length of the sliding window of the preset size in the second direction, the first direction is the row direction or the column direction, and the second direction is perpendicular to the first direction in the original image.

[0051] In this step, if the unselected pixel points in the first direction are not enough to fill the sliding window, the sliding window is moved along the second direction based on the initial position. In the case where the first direction is the row direction, in this step, moving the sliding window along the second direction with a second preset step size based on the initial position can achieve "row wrapping", so that the sliding window performs pixel block selection from a new row position.

[0052] Through the solution provided by the embodiment of the present application, small cycle pixel selection can be achieved along the first direction, and large cycle pixel selection can be achieved along the second direction, so as to realize orderly pixel block division of the pixels in the original image. Among them, the first direction can be determined according to the row and column priority in the memory. In the case of row priority, the first direction is determined as the row direction and the second direction is determined as the column direction, which can better utilize the cache and memory bandwidth and effectively improve performance.

[0053] Optionally, in the process of sliding and selecting pixel blocks along the first direction, there may be unselected pixel points in the first direction whose number is less than the first preset step length. In addition, in the process of sliding and selecting pixel blocks along the second direction, there may also be unselected pixel points in the second direction. For these pixel points that have not been transposed, image rotation and storage can be performed one by one after the transposition processing is performed on the pixel points of each pixel block, so as to ensure the integrity of the image display after rotation.

[0054] Based on the solution provided in the above embodiment, optionally, Figure 3 As shown, in the above step S12, the transposition operation is performed in parallel on each pixel point belonging to the same pixel block based on the NEON instruction, including: S31: Load the pixel block to be transposed into the NEON register.

[0055] In the embodiment of the present application, NEON registers can be initialized in advance, and these registers can be used to implement image loading, image transposition and image storage after transposition. Among them, multiple NEON registers can be defined, each register is used to store corresponding channel data in a pixel block.

[0056] In this step, the pixel block selected by the sliding window is used as the pixel block to be transposed and loaded into the NEON register. For example, a pixel point can be stored as 4 bytes, which correspond to 4 channels, representing the data of the 4 channels of R, G, B, and A respectively.

[0057] S32: If the first data length of the pixel data in the target pixel block stored in the NEON register is less than the preset data length, a transposition operation is performed on the pixel data in the target pixel block in parallel based on a transposition instruction of the first data length, and the pixel data in the transposed target pixel block is reinterpreted as pixel data of a second data length based on a reinterpretation instruction until the data length of the reinterpreted pixel data is greater than or equal to the preset data length, wherein the second data length is greater than the first data length.

[0058] In the embodiment of the present application, the preset data length can be determined according to the preset size of the sliding window. For example, if the length of the sliding window in the horizontal direction is 8 pixels, the corresponding preset data length can be 32 bits.

[0059] In this step, it is determined whether the first data length of the pixel data in the target pixel block stored in the NEON register is less than the preset data length, that is, whether it is less than 32 bits. If less than, a transposition operation is performed on the pixel data based on the transposition instruction of the first data length. For example, when the first data length is 8 bits, a transposition operation is performed on the pixel data in the target pixel block in parallel based on the 8-bit transposition instruction.

[0060] Subsequently, the pixel data in the transposed target pixel block is reinterpreted as longer pixel data having a second data length based on the reinterpretation instruction. For example, the 8-bit pixel data in the transposed target pixel block is reinterpreted as 16-bit pixel data.

[0061] The above transposition and re-interpretation operations are repeatedly performed until the data length of the pixel data in the target pixel block reaches the above preset data length.

[0062] S33: If the data length of the pixel data in the target pixel block stored in the NEON register is greater than or equal to the preset data length, a transposition operation is performed on the pixel data in the target pixel block in parallel based on a transposition instruction of the preset data length, and the pixel data in the transposed target pixel block is reinterpreted as pixel data of the initial pixel length of the pixel block to be transposed based on a reinterpretation instruction.

[0063] In this step, when the data length of the pixel data in the target pixel block stored in the NEON register reaches a preset data length, a transposition operation is performed based on a transposition instruction of the preset data length to complete the transposition of each data block in the target pixel block.

[0064] In this example, assuming that the preset data length is 32 bits, after the data length of the pixel data in the target pixel block stored in the NEON register reaches 32 bits, a transposition operation is performed on the pixel data in the target pixel block based on a 32-bit transposition instruction.

[0065] In the embodiment of the present application, the above-mentioned transposition instruction may specifically refer to vtrn, and the reinterpret instruction may specifically refer to vreinterpret.

[0066] Based on the solution provided in the above embodiment, optionally, the row direction size of the sliding window of the preset size is 8 pixels, and the column direction size of the sliding window of the preset size is 8 pixels.

[0067] In practical applications, the image resolution is usually a multiple of 8. Setting an 8×8 sliding window can achieve both processing efficiency and universality. The 8×8 sliding window can effectively reduce the overhead of processing pixels one by one and improve the overall performance.

[0068] Based on the solution provided in the above embodiment, optionally, Figure 4a As shown, in the above step S12, the transposition operation is performed in parallel on each pixel point belonging to the same pixel block based on the NEON instruction, including: S41: Load the pixel block to be transposed into the NEON register; S42: performing a transposition operation in parallel on the pixel data in the target pixel block stored in the NEON register based on the 8-bit transposition instruction; S43: reinterpreting the 8-bit pixel data in the target pixel block into 16-bit pixel data based on the 16-bit reinterpretation instruction; S44: performing a transposition operation on the pixel data in the target pixel block in parallel based on the 16-bit transposition instruction; S45: reinterpreting the 16-bit pixel data in the target pixel block into 32-bit pixel data based on the 32-bit reinterpretation instruction; S46: performing a transposition operation on the pixel data in the target pixel block in parallel based on the 32-bit transposition instruction; S47: Reinterpret the 32-bit pixel data in the target pixel block into 8-bit pixel data based on the reinterpretation instruction.

[0069] The following is a further explanation of this solution with reference to an example. Figure 4b, this solution can be used to perform full-screen rotation on an image, which can be achieved by the following steps: The first step is to calculate the divisible height and width based on the resolution of the original image. The height of the original image refers to the length of the resolution in the column direction, and the width of the original image refers to the length of the resolution in the row direction. In this step, the resolution divisors in the row and column directions are calculated respectively to determine the preset size of the sliding window.

[0070] Among them, the image resolution is mostly designed to be a multiple of 8, and the preset size of the sliding window can be set by default to 8 × 8. In some scenarios, the resolution of the original image in a certain direction cannot be divided by 8, so the divisor can be determined based on the actual value of the resolution in the row and column directions, and one of the divisors can be selected as the preset size of the sliding window in that direction.

[0071] If there are multiple divisors, a divisor of appropriate size can be selected as the preset size of the sliding window in this direction according to the actual processing performance. If the processing performance is strong and the storage resources are sufficient, a divisor with a larger value can be selected to expand the size of the sliding window, reduce the number of pixel blocks, and improve the overall processing efficiency. If the processing performance is weak and the storage resources are tight, a divisor with a smaller value can be selected to ensure that the hardware performance can realize the transposition processing of the original image.

[0072] The second step is to initialize the NEON registers. Based on the actual needs of loading, transposing and storing image data, this step can flexibly define multiple NEON registers. Each register can be used to save the corresponding channel data in the image block.

[0073] The third step is to loop and select pixel blocks and perform transposition processing. The loop can be divided into an outer loop and an inner loop. The outer loop traverses the original image in the height direction, and the inner loop traverses the original image in the width direction. Assuming that the preset size of the sliding window is 8×8, the preset step size of the sliding window is 8 pixels. In any inner loop, the pixel block is moved along the row direction, and 8 columns of pixels are processed each time. In any outer loop, the pixel block is moved along the column direction, and 8 rows of pixels are processed each time. In this way, sliding selection is performed sequentially along the row and column directions to obtain an 8×8 pixel block.

[0074] The fourth step is to load the selected pixel block data into the NEON register. In actual applications, the vld4_u8() instruction can be used to load 8×8 pixel data from the original image into the NEON register. Among them, one pixel is 4 bytes, and the 4 bytes correspond to 4 channels, representing R, G, B, and A channel data respectively. The data is loaded into the register by channel.

[0075] The fifth step is to perform data transposition on the pixel block in the NEON register. The purpose of this data transposition is to achieve row and column swapping. In this example, 8-bit transposition is first performed, and the vtrn_u8() instruction is used to swap the adjacent bytes of each row, that is, the row and column parts of the 8-bit data are swapped. Then a 16-bit transposition is performed, in which the vreinterpret_u16_u8() instruction is first used to reinterpret every two 8-bit data as 16-bit data. Then vtrn_u16() is used to further transpose the 16-bit data, swap the adjacent 16-bit row and column data, and achieve 16-bit transposition. Next, the vreinterpret_u32_u16() instruction is used to reinterpret every two 16-bit data as 32-bit data, and then vtrn_u32() is used to complete the final row and column swap, and the transposition of the data block is completed. Among them, data type conversion can gradually expand 8-bit data to 16 bits and 32 bits, thereby increasing the amount of data for each transposition operation. In this step, SIMD instructions (such as vtrn_u8 and vtrn_u16) are used to optimize specific data types, and the operation is made efficient by step-by-step conversion. By gradually expanding to higher bit widths, batch data parallel transposition can be completed in stages. Compared with pixel-by-pixel processing, this solution can effectively improve the efficiency of data transposition processing.

[0076] Step 6: Store the transposed data. Specifically, the vst4_u8() instruction can be used to write the transposed data from the NEON register to the target buffer.

[0077] The selection and transposition of the above pixel blocks are executed in a loop until the loop ends.

[0078] In practical applications, if the new position is directly calculated for each pixel in the original image, the code needs to be processed pixel by pixel, resulting in a large number of loops and memory accesses, which seriously affects performance. Through the solution provided by the embodiment of the present application, the 8-bit data is first transposed, and the adjacent bytes are rearranged into 16-bit pairs, and then the 16-bit data is further transposed into 32-bit data pairs to complete the final transposition operation. The conversion of data types is to use the NEON instruction set to perform efficient parallel data processing and transposition operations. This conversion enables data to be effectively operated in registers of different bit widths, reducing the number of cycles and the number of instruction executions.

[0079] In terms of data types, this solution performs data transposition through multiple levels (8-bit, 16-bit, and 32-bit). The NEON library contains multiple data types such as 8x8x4, 16x4x2, and 32x2x2. This solution performs transposition on 8x8x4 and interprets it as 16x4x2 for storage. Then, it is further transposed and interpreted as 32x2x2 for storage, and finally, it is transposed and interpreted as 8x8x4, so that the data is rearranged from the row direction to the column direction, completing the overall data transposition.

[0080] Based on the solution provided in the above embodiment, optionally, Figure 5 As shown, after the above step S13, that is, after each pixel point in the transposed pixel block is stored in the buffer according to the matching position of the target rotation angle, it also includes: S51: If the original image includes boundary pixel points that are not selected by the sliding window, each of the boundary pixel points is stored in the buffer according to a matching position of the target rotation angle.

[0081] Based on the solution provided by the above embodiment, after executing the sixth step, the seventh step may be executed to perform separate processing on the boundary pixels to ensure the integrity of the rotated display image.

[0082] Step 7: Boundary processing. In some application scenarios, the original image also includes pixels that are not selected into the pixel block. These pixels are processed separately as boundary pixel data. Specifically, the remaining boundary pixels can be loaded, transposed and stored separately.

[0083] The solution provided by the embodiment of the present application does not require GPU acceleration, but uses the NEON instruction set of the ARM architecture to accelerate screen rotation, and no longer relies on external GPU hardware acceleration, and can achieve efficient screen rotation on embedded devices without GPU or with limited GPU performance. Through parallel processing and the NEON instruction set, it can achieve an image processing speed faster than the CPU rotation algorithm, solving the performance bottleneck and high resource occupancy caused by limited CPU and GPU resources.

[0084] The solution provided in the embodiment of the present application realizes parallel computing by using SIMD of the NEON instruction set, processing multiple pixel data each time, thereby realizing efficient screen rotation in an embedded environment, achieving real-time performance without hardware support, and being suitable for scenes requiring high response speed, such as industrial equipment and mobile devices.

[0085] In order to solve the problems existing in the related art, the embodiment of the present application further provides a device 60 for rotating and displaying an image, such as Figure 6 As shown, it is applied to processors with ARM architecture, including: An acquisition module 61 is used to acquire an original image and a target rotation angle, wherein the original image includes a plurality of pixel points arranged in an array; A transposition module 62, which sequentially selects pixel blocks from the original image based on a sliding window of a preset size, and performs a transposition operation on each pixel point in the same pixel block in parallel based on a NEON instruction; The storage module 63 stores each pixel point in the transposed pixel block in a buffer according to the matching position of the target rotation angle; The display module 64 displays the rotated image based on the pixels stored in the buffer.

[0086] The device provided in the embodiment of the present application is applied to a processor with an ARM architecture. First, an original image and a target rotation angle are obtained, and the original image includes a plurality of pixel points arranged in an array; then, pixel blocks are sequentially selected from the original image based on a sliding window of a preset size, and a transposition operation is performed in parallel on each pixel point belonging to the same pixel block based on a NEON instruction; then, each pixel point in the transposed pixel block is stored in a buffer according to the matching position of the target rotation angle; then, the rotated image is displayed based on the pixel points stored in the buffer. This scheme uses the NEON instruction of an ARM architecture processor, and performs a transposition operation on the pixel points belonging to the same pixel block in parallel through a single instruction, which can realize batch transposition of pixels and improve the efficiency of image rotation processing.

[0087] Among them, the above modules in the device provided by the embodiment of the present application can also implement the method steps provided by the above method embodiment. Alternatively, the device provided by the embodiment of the present application can also include other modules in addition to the above modules to implement the method steps provided by the above method embodiment. And the device provided by the embodiment of the present application can achieve the technical effects that can be achieved by the above method embodiment.

[0088] Preferably, an embodiment of the present application also provides an electronic device, including a processor, a memory, and a computer program stored in the memory and executable on the processor. When the computer program is executed by the processor, each process of the above-mentioned method embodiment for image rotation display is implemented, and the same technical effect can be achieved. To avoid repetition, it will not be described here.

[0089] The embodiment of the present application also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, each process of the above-mentioned method for rotating and displaying an image is implemented, and the same technical effect can be achieved. To avoid repetition, it is not repeated here. The computer-readable storage medium is, for example, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

[0090] An embodiment of the present application also provides a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program. The computer program can be operated to enable a computer to execute part or all of the steps of the above-mentioned image rotation display method embodiment, and can achieve the same technical effect. To avoid repetition, it will not be repeated here.

[0091] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present application may adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.

[0092] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate a device for implementing the functions specified in one process or multiple processes in the flowchart and / or one box or multiple boxes in the block diagram.

[0093] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory produce a manufactured product including an instruction device that implements the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.

[0094] These computer program instructions may also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.

[0095] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0096] Memory may include non-permanent storage in a computer-readable medium, in the form of random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.

[0097] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.

[0098] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.

[0099] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems or computer program products. Therefore, the present application may take the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware. Moreover, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.

[0100] The above is only an embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included in the scope of the claims of the present application.

Claims

1. A method for image rotation display, characterized in that: Applicable to processors with ARM architecture, including: Acquire an original image and a target rotation angle, wherein the original image includes a plurality of pixel points arranged in an array; Selecting pixel blocks from the original image in sequence based on a sliding window of a preset size, and performing a transposition operation on each pixel point in the same pixel block in parallel based on a NEON instruction; storing each pixel point in the transposed pixel block in a buffer according to the matching position of the target rotation angle; The rotated image is displayed based on the pixels stored in the buffer.

2. The method according to claim 1, characterized in that Sequentially selecting pixel blocks from the original image based on a sliding window of a preset size, including: Sequentially selecting pixel blocks along a first direction from an initial position of the original image with a first preset step length based on the sliding window; If the number of unselected pixel points in the first direction is less than the first preset step size, the sliding window is moved along the second direction with a second preset step size based on the initial position, and pixel blocks are selected sequentially along the first direction until the number of unselected pixel points in the second direction is less than the second preset step size, wherein the first preset step size is the length of the sliding window of the preset size in the first direction, the second preset step size is the length of the sliding window of the preset size in the second direction, the first direction is the row direction or the column direction, and the second direction is perpendicular to the first direction in the original image.

3. The method according to claim 2, characterized in that Based on the NEON instruction, the transposition operation is performed in parallel on each pixel point in the same pixel block, including: Load the pixel block to be transposed into the NEON register; If a first data length of the pixel data in the target pixel block stored in the NEON register is less than a preset data length, a transposition operation is performed in parallel on the pixel data in the target pixel block based on a transposition instruction of the first data length, and the pixel data in the transposed target pixel block is reinterpreted as pixel data of a second data length based on a reinterpretation instruction until the data length of the reinterpreted pixel data is greater than or equal to the preset data length, wherein the second data length is greater than the first data length; If the data length of the pixel data in the target pixel block stored in the NEON register is greater than or equal to the preset data length, a transposition operation is performed on the pixel data in the target pixel block in parallel based on a transposition instruction of the preset data length, and the pixel data in the transposed target pixel block is reinterpreted as pixel data of the initial pixel length of the pixel block to be transposed based on a reinterpretation instruction.

4. The method according to claim 1, characterized in that The row direction size of the sliding window of the preset size is 8 pixels, and the column direction size of the sliding window of the preset size is 8 pixels.

5. The method according to claim 4, characterized in that Based on the NEON instruction, the transposition operation is performed in parallel on each pixel point in the same pixel block, including: Load the pixel block to be transposed into the NEON register; Performing a transposition operation in parallel on the pixel data in the target pixel block stored in the NEON register based on an 8-bit transposition instruction; reinterpreting 8-bit pixel data in the target pixel block into 16-bit pixel data based on the 16-bit reinterpretation instruction; Performing a transposition operation on the pixel data in the target pixel block in parallel based on a 16-bit transposition instruction; reinterpreting the 16-bit pixel data in the target pixel block into 32-bit pixel data based on the 32-bit reinterpretation instruction; Performing a transposition operation on the pixel data in the target pixel block in parallel based on a 32-bit transposition instruction; The 32-bit pixel data in the target pixel block is reinterpreted into 8-bit pixel data based on the reinterpretation instruction.

6. The method according to any one of claims 1 to 5, characterized in that After storing each pixel point in the transposed pixel block in the buffer according to the matching position of the target rotation angle, the method further includes: If the original image includes boundary pixel points that are not selected by the sliding window, each of the boundary pixel points is stored in the buffer according to the matching position of the target rotation angle.

7. A device for rotating and displaying an image, characterized in that: Applicable to processors with ARM architecture, including: An acquisition module is used to acquire an original image and a target rotation angle, wherein the original image includes a plurality of pixel points arranged in an array; A transposition module, which sequentially selects pixel blocks from the original image based on a sliding window of a preset size, and performs a transposition operation on each pixel point in the same pixel block in parallel based on a NEON instruction; A storage module stores each pixel point in the transposed pixel block in a buffer according to a matching position of the target rotation angle; A display module displays the rotated image based on the pixels stored in the buffer.

8. An electronic device, characterized in that: include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program implements the steps of the method according to any one of claims 1 to 6 when executed by the processor.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

10. A computer program product, characterized in that The computer program product comprises a non-transitory computer-readable storage medium storing a computer program, the computer program being operable to cause a computer to perform the steps of the method according to any one of claims 1 to 6.

Citation Information

Cited By

  • Batch image data processing method and system for intelligent manufacturing production line

    CN120198858A