Frame insertion processing method and device, electronic equipment and storage medium

Through the interpolation model of mixed programming in GPU and NPU, the problem of time-consuming interpolation processing is solved, efficient interpolation processing is achieved, and the efficiency and real-timeness of game interpolation are improved.

CN120378691APending Publication Date: 2025-07-25BEIJING X RING TECHNOLOGY CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202411206310.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-08-29
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

In the prior art, the interpolation processing takes a long time and is inefficient, making it difficult to meet the needs of high-frame-rate games.

Method used

Using hybrid programming of GPU and NPU, the feature extraction network is deployed in the NPU and the interpolated frame network is deployed in the GPU to achieve efficient processing of the interpolated frame model.

Benefits of technology

It improves the processing efficiency of the interpolation model, reduces the overall model inference time, and meets the real-time requirements of game interpolation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120378691A_ABST
    Figure CN120378691A_ABST
Patent Text Reader

Abstract

The invention provides a frame insertion processing method and device, electronic equipment and a storage medium, and the method comprises the steps: obtaining a first texture and a second texture obtained through rendering before the first texture, the first texture and the second texture are UI-free textures, splicing the first texture and the second texture to obtain a spliced texture, and storing the spliced texture in a storage medium; a feature extraction network in a frame insertion model is adopted to perform multi-dimensional feature extraction on the spliced texture to obtain a target optical flow feature map and a target mask feature map, the feature extraction network is deployed in an NPU, and a frame insertion network in the frame insertion model is adopted to obtain frame insertion texture according to the target optical flow feature map and the target mask feature map, the frame insertion network is deployed in the GPU, and the frame insertion model is deployed in the GPU and the NPU, so that the hybrid programming of the GPU and the NPU is realized, the algorithm advantages of two kinds of hardware are fully played, and the processing efficiency of the frame insertion model is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and in particular, to an interpolation processing method, apparatus, electronic device, and storage medium. Background Art

[0002] The screen refresh rates of current high-performance mobile devices have reached 120HZ, 144HZ, or even higher. However, a large number of mobile games still remain at 60HZ. By using interpolation technology, users can experience high-frame-rate games and enhance the gaming experience.

[0003] In related technologies, during the interpolation processing, there are problems such as long interpolation time and low efficiency. Summary of the Invention

[0004] This application aims to solve at least one of the technical problems in the related technologies to some extent.

[0005] To this end, this application proposes an interpolation processing method, apparatus, electronic device, and storage medium, so as to achieve the hybrid programming of the GPU and NPU by deploying the interpolation model in the GPU and NPU, giving full play to the algorithm advantages of the two kinds of hardware, and improving the processing efficiency of the interpolation model.

[0006] An embodiment of one aspect of this application proposes an interpolation processing method, including:

[0007] Obtain a first texture and a second texture rendered before the first texture;

[0008] Stitch the first texture and the second texture to obtain a stitched texture;

[0009] Use the feature extraction network in the interpolation model to perform multi-dimensional feature extraction on the stitched texture to obtain a target optical flow feature map and a target mask feature map; wherein, the feature extraction network is deployed in a neural network processor NPU;

[0010] Use the interpolation network in the interpolation model to obtain an interpolated texture according to the target optical flow feature map and the target mask feature map; wherein, the interpolation network is deployed in a graphics processing unit GPU.

[0011] An embodiment of another aspect of this application proposes an interpolation processing method apparatus, including:

[0012] An obtaining module, configured to obtain a first texture and a second texture rendered before the first texture;

[0013] A stitching module, configured to stitch the first texture and the second texture to obtain a stitched texture;

[0014] A feature extraction module, configured to perform multi-dimensional feature extraction on the spliced texture by using a feature extraction network in an interpolation model, so as to obtain a target optical flow feature map and a target mask feature map; wherein, the feature extraction network is deployed in an NPU;

[0015] An interpolation module, configured to obtain an interpolated texture by using an interpolation network in the interpolation model according to the target optical flow feature map and the target mask feature map; wherein, the interpolation network is deployed in a GPU.

[0016] Another embodiment of this application provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the method described in the foregoing one aspect or the method described in the foregoing other aspect is implemented.

[0017] Another embodiment of this application provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the method described in the foregoing one aspect or the method described in the foregoing other aspect is implemented.

[0018] Another embodiment of this application provides a computer program product, on which a computer program is stored. When the program is executed by a processor, the method described in the foregoing one aspect or the method described in the foregoing other aspect is implemented.

[0019] The interpolation processing method, device, electronic device, and storage medium provided in this application obtain a first texture and a second texture rendered before the first texture. Both the first texture and the second texture are textures without UI. The first texture and the second texture are spliced to obtain a spliced texture. A feature extraction network in an interpolation model is used to perform multi-dimensional feature extraction on the spliced texture to obtain a target optical flow feature map and a target mask feature map. Among them, the feature extraction network is deployed in an NPU. An interpolation network in the interpolation model is used to obtain an interpolated texture according to the target optical flow feature map and the target mask feature map. Among them, the interpolation network is deployed in a GPU. By deploying the interpolation model in a GPU and an NPU, GPU and NPU hybrid programming is realized, giving full play to the algorithm advantages of the two kinds of hardware and improving the processing efficiency of the interpolation model.

[0020] Additional aspects and advantages of this application will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of this application. Description of the Drawings

[0021] The above and / or additional aspects and advantages of this application will become apparent and be readily understood from the following description of the embodiments in conjunction with the drawings, where:

[0022] Figure 1Schematic diagram of a frame interpolation processing method provided by an embodiment of the present application;

[0023] Figure 2 Schematic diagram of another frame interpolation processing method provided by an embodiment of the present application;

[0024] Figure 3A Schematic diagram of accessing a shared memory provided by an embodiment of the present application;

[0025] Figure 3B Schematic diagram of establishing the relationship between a shared memory and a texture provided by an embodiment of the present application;

[0026] Figure 4 Schematic diagram of another frame interpolation processing method provided by an embodiment of the present application;

[0027] Figure 5A Schematic diagram of the layout mode of a texture in a GPU provided by an embodiment of the present application;

[0028] Figure 5B Schematic diagram of the layout mode of a texture in an NPU provided by an embodiment of the present application;

[0029] Figure 6 Schematic diagram of texture layout matching provided by an embodiment of the present application;

[0030] Figure 7 Schematic diagram of another frame interpolation processing method provided by an embodiment of the present application;

[0031] Figure 8 Schematic diagram of the structure of a feature extraction network provided by an embodiment of the present application;

[0032] Figure 9 Schematic diagram of the structure of a feature extraction sub-network provided by an embodiment of the present application;

[0033] Figure 10 Schematic diagram of the structure of a frame interpolation network provided by an embodiment of the present application;

[0034] Figure 11 Schematic diagram of the structure of a frame interpolation device provided by an embodiment of the present application;

[0035] Figure 12 Schematic diagram of the structure of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0036] Embodiments of the present application will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below by referring to the accompanying drawings are exemplary and are intended to explain the present application, and should not be construed as limiting the present application.

[0037] The following describes an interpolation processing method, apparatus, electronic device, and storage medium according to embodiments of the present application with reference to the accompanying drawings.

[0038] Figure 1 It is a schematic flowchart of an interpolation processing method provided by an embodiment of the present application.

[0039] The execution subject of the interpolation processing method according to an embodiment of the present application is an interpolation device, which can be set in an electronic device. The electronic device can be a mobile terminal, such as a smart phone, a personal digital assistant, a smart wearable device, an IPAD, etc., which is not limited in this embodiment.

[0040] As Figure 1 shown, the method may include the following steps:

[0041] Step 101, obtain a first texture and a second texture rendered before the first texture.

[0042] As an implementation manner, the first texture may be rendered by the GPU at a target time, and the second texture may be rendered by the GPU at a historical time before the target time. The second texture and the first texture are two adjacent frames of textures. Among them, both the first texture and the second texture are textures without UI. Interpolation processing through textures without UI can reduce the influence of UI textures on the interpolation model algorithm and improve the interpolation effect of the subsequent interpolation model.

[0043] Step 102, splice the first texture and the second texture to obtain a spliced texture.

[0044] In an embodiment of the present application, the first texture and the second texture are respectively subjected to feature extraction to obtain the features of the first texture and the features of the second texture, and the features of the first texture and the features of the second texture are spliced to obtain a spliced texture. As an implementation manner, the features of the first texture and the features of the second texture can be spliced in the channel dimension through the feature connection concat method to obtain a spliced texture.

[0045] Step 103, use the feature extraction network in the interpolation model to perform multi-dimensional feature extraction on the spliced texture to obtain a target optical flow feature map and a target mask feature map.

[0046] Among them, the feature extraction network is deployed in the NPU.

[0047] In the embodiments of the present application, in the field of video frame interpolation (VFI), such as the game frame interpolation field on mobile devices, frame interpolation processing is performed on two adjacent frame textures through a frame interpolation model to obtain the interpolated texture between the adjacent front and rear frame textures. To improve the efficiency of the frame interpolation model processing, the algorithm of the frame interpolation model is jointly executed by the GPU and the NPU, that is, the feature extraction network of the frame interpolation model is deployed in the NPU, and the frame interpolation network of the frame interpolation model is deployed in the GPU. That is to say, GPU and NPU hybrid programming is adopted to make full use of the sampling acceleration hardware unit of the GPU and the convolution acceleration hardware unit of the NPU, reduce the time consumption of the entire path, and improve the efficiency of model processing.

[0048] In the embodiments of the present application, the feature extraction network is deployed in the NPU, and multi-dimensional feature extraction is performed on the spliced texture through the feature extraction network. Among them, the multi-dimensional feature extraction includes feature granularity from coarse to fine, from the whole to the local, and the extraction of optical flow features and mask features is performed based on the multi-dimensional feature information, which improves the accuracy of the extraction of optical flow features and mask features and can reduce the artifacts and picture distortion of the texture obtained by subsequent frame interpolation.

[0049] Step 104, use the frame interpolation network in the frame interpolation model to obtain the interpolated texture according to the target optical flow feature map and the target mask feature map.

[0050] In the embodiments of the present application, the frame interpolation network is deployed in the GPU. The optical flow information is mainly used to predict and simulate the pixel movement between adjacent frames during the frame interpolation process, and a new intermediate frame is generated by combining with the mask feature. The interpolated texture between the first texture and the second texture is predicted according to the intermediate feature. By adopting GPU and NPU hybrid programming, the sampling acceleration hardware unit of the GPU and the convolution acceleration hardware unit of the NPU are fully utilized, the time consumption of the entire path is reduced, and the time of overall model inference is reduced to meet the real-time requirements of game frame interpolation.

[0051] In the frame interpolation processing method of the embodiments of the present application, the first texture and the second texture rendered before the first texture are obtained. Both the first texture and the second texture are textures without UI. The first texture and the second texture are spliced to obtain the spliced texture. The feature extraction network in the frame interpolation model is used to perform multi-dimensional feature extraction on the spliced texture to obtain the target optical flow feature map and the target mask feature map. Among them, the feature extraction network is deployed in the NPU. The frame interpolation network in the frame interpolation model is used to obtain the interpolated texture according to the target optical flow feature map and the target mask feature map. Among them, the frame interpolation network is deployed in the GPU. By deploying the frame interpolation model in the GPU and the NPU, GPU and NPU hybrid programming is realized, the algorithm advantages of the two kinds of hardware are fully exerted, and the efficiency of the frame interpolation model processing is improved.

[0052] Based on the above embodiments,Figure 2 The following is a schematic flowchart of an interpolation processing method provided by an embodiment of the present application. As Figure 2 shown, the method includes the following steps:

[0053] Step 201: Obtain the file descriptor of the shared memory buffer.

[0054] Among them, the texture data is rendered in the GPU. The NPU needs to obtain the first texture from the GPU, but the NPU and the GPU cannot directly communicate with each other. That is to say, the NPU and the GPU cannot directly interact with data. Usually, data copying is implemented based on the CPU. This data copying method is inefficient and will cause a high CPU load, increasing the burden on the CPU and reducing the performance of the mobile device. To improve the efficiency of information interaction between the NPU and the GPU, a shared memory buffer is set up. For example, the DMA buffer, and the DMA buffer can act as shared memory. Figure 3A The following is a schematic diagram of accessing a shared memory provided by an embodiment of the present application. As Figure 3A shown, the GPU and the NPU can jointly access the DMA buffer. Specifically, the data written by the GPU to the DMA buffer can be read by the NPU, and the data written by the NPU to the DMA buffer can be read by the GPU. This process does not require the participation of the CPU, significantly reducing the CPU load and enabling the CPU to have more resources to process other tasks, improving the device performance. In the embodiment of the present application, the shared memory buffer is initialized and generated by the CPU. After the CPU generates the shared memory buffer, it will send the file descriptor of the shared memory buffer to the GPU and the NPU. The file descriptor is used to identify the shared memory, and the NPU and the GPU can find the corresponding shared memory buffer according to the file descriptor and perform data reading and writing.

[0055] As an implementation method, obtain the parameter information of the memory to be created. The parameter information is used to indicate that the memory to be created is a shared memory. The parameter information includes the parameters set by the Android system, that is, "GraphicBuffer:USAGE_HW_TEXTURE|GraphicBuffer:USAGE_HW_RENDER". According to the parameter information, create a shared memory buffer and generate the file descriptor of the shared memory. Specifically, first create a GraphicBuffer, and convert it into an android_native_buffer through the getNativeBuffer function. The android_native_buffer is a bridge, which is used to further create an EGLImageKHR and also used to obtain the file descriptor (File Descriptor, FD) of the DMABuffer.

[0056] Step 202: Access the shared memory buffer according to the file descriptor.

[0057] In the embodiment of the present application, the GPU can find the corresponding shared memory buffer according to the file descriptor, so as to implement writing the rendered texture data into the shared memory buffer.

[0058] Step 203: Perform texture rendering in the shared memory buffer. In response to determining that texture rendering is performed on the user interface UI, obtain the texture without UI rendered in the shared memory buffer, and use the texture without UI as the first texture.

[0059] In the embodiment of the present application, texture rendering in the shared memory buffer can be implemented through an external texture. As an implementation method, Figure 3B FIG. is a schematic diagram showing the relationship between the shared memory and the texture provided by the embodiment of the present application. As Figure 3B shown, first create a memory GraphicBuffer, and convert it into an android_native_buffer_t through the getNativeBuffer function. Among them, android_native_buffer_t is a bridge, which is used to further create an EGLImageKHR through the function eglCreateImage, and is also used to obtain the file descriptor FD of the shared memory buffer DMABuffer through the function get_plane_dma_buf. Finally, create an external texture using EGLImageKHR. So far, the association between the texture and the DMA buffer is completed. Access the shared memory buffer according to the interface of the external texture, and render the first texture in the shared memory buffer. That is to say, the rendering result of the GPU on the texture will be directly stored in the DMA buffer, and the NPU can directly access this DMA buffer through the file descriptor FD.

[0060] In the embodiments of the present application, frame interpolation processing is performed through the feature extraction sub-network in the frame interpolation model set in the NPU. In the related art, the texture used includes the texture of the UI. During the frame interpolation processing, the UI texture will interfere with the optical flow calculation. When applied to mobile game frame interpolation, the image quality effect is not good. Therefore, in the embodiments of the present application, the UI texture is separated to obtain the first texture that does not include the UI. Among them, the layer of the UI texture is not related to the movement of other objects. When performing frame interpolation, the UI image will cause great interference to the optical flow calculation or the calculation of the motion vector in the frame interpolation calculation process, resulting in poor accuracy of the frame interpolation result. During the process of drawing or rendering the display page, the painter's algorithm is usually used for texture rendering, that is, the scene farther away is rendered first, and then the scene closer is rendered to cover the farther part. In the game scene, the UI texture is the closest to the scene and is the last texture to be rendered. This makes it feasible to separate the UI texture. Only when other scenes have been rendered and the UI texture has not been rendered, the frame buffer framebufer is switched, that is, switched from the shared memory buffer to the frame buffer, and the UI texture is continuously rendered in the switched frame buffer, so as to achieve the separation of the UI texture and the texture that does not include the UI.

[0061] In the embodiments of the present application, during the process of texture rendering in the shared memory buffer, the process of texture rendering is monitored to determine whether it is the timing of texture separation. When it is determined that it is the timing of texture separation, that is, when it is determined that the GPU starts to perform UI texture rendering, and the texture that does not include the UI in the shared memory buffer has been rendered, the rendered texture that does not include the UI can be obtained from the shared memory buffer, and the texture that does not include the UI is used as the first texture. Subsequently, frame interpolation processing is performed according to the first texture that does not include the UI, and the influence of the UI texture on the frame interpolation result is reduced during the frame interpolation processing, and the accuracy of the frame interpolation result is improved.

[0062] Among them, for the method of determining the timing of texture separation, as an implementation method, the CPU monitors the rendering instructions to be sent to the GPU, and can identify whether the rendering instruction is the target rendering instruction for rendering the UI texture according to the identification information of the rendering instruction. If it is determined that the rendering instruction is the target rendering instruction, it is further identified whether the state of the state machine of the Graphics Application Programming Interface (OpenGL for Embedded Systems, GLES) is the target state. If it is determined that the rendering instruction is the target rendering instruction and the state of the state machine of GLES is the target state, it is determined that the GPU starts to perform texture rendering of the user interface UI based on the target rendering instruction, that is, it is determined that the timing of switching the frame buffer arrives, and the UI texture is separately rendered by switching the frame buffer, and the separation of the UI texture is achieved.

[0063] It should be noted that the target rendering instruction performs rendering by calling a rendering function. However, during the execution of the target rendering instruction, the objects to be rendered are not necessarily the same. Therefore, the rendering behavior can be identified through the rendering instruction. The rendering behavior indicates which rendering pass the rendering has reached. If it is the target render pass, it is considered that the UI texture needs to be rendered. Reaching the target render pass can also be understood as the number of times the gl*() function is called reaching the target number. For example, if the number of times the gl*() function is called reaches 12 times, the state machine of GLES will change to the target state. If the CPU queries that the state of the state machine is the target state, it confirms that the frame buffer needs to be switched, that is, switched from the first buffer to the second buffer. Among them, the gl*() function is, for example, the glDrawElements function. That is to say, only when it is determined that the rendering instruction is the target rendering instruction and the state of the state machine of GLES is the target state, can the timing of texture separation be accurately determined.

[0064] The target rendering instruction refers to the rendering instruction used for the first texture rendering of the UI. For example, it is the rendering instruction of the graphics application programming interface GLES or VulKan generated by the CPU based on the rendering request.

[0065] Step 204: Obtain the first texture and the second texture rendered before the first texture.

[0066] In the embodiment of the present application, the first texture is the texture stored in the shared memory buffer. The NPU obtains the first texture, that is, reads the first texture from the shared memory buffer. The second texture is a texture of the same type as the first texture. The second texture is a texture without UI rendered by the GPU in the shared memory buffer at a historical moment. The second texture can be stored in the shared memory buffer, or the NPU reads and stores it in the memory of the NPU from the shared memory buffer at a historical moment.

[0067] Step 205: Stitch the first texture and the second texture to obtain a stitched texture.

[0068] Step 206: Use the feature extraction network in the frame interpolation model to perform multi-dimensional feature extraction on the stitched texture to obtain a target optical flow feature map and a target mask feature map.

[0069] Among them, the feature extraction network is deployed in the NPU. In the embodiment of the present application, the NPU can store the generated target optical flow feature map and target mask feature map in the shared memory buffer, so that the GPU can directly access the shared memory buffer to obtain the target optical flow feature map and target mask feature map, thereby improving the data transmission efficiency.

[0070] Step 207: Use the interpolation network in the interpolation model to obtain the interpolation texture based on the target optical flow feature map and the target mask feature map.

[0071] Among them, the interpolation network is deployed in the GPU.

[0072] Among them, steps 204-207 can refer to the relevant explanations in the foregoing embodiments. The principles are the same and will not be elaborated here.

[0073] Furthermore, obtain the third texture including the UI, fuse the interpolation texture and the third texture to obtain the target interpolation texture. The target interpolation texture is a complete target interpolation texture to ensure the integrity of the finally displayed interpolation picture. Among them, the target interpolation texture is inserted into the image buffer management component bufferQueue for display on the mobile device.

[0074] Among them, for the rendering method of the third texture, as an implementation, in response to determining to perform texture rendering on the user interface UI, switch the rendering buffer from the shared memory buffer to the target buffer, and render the UI in the target buffer according to the target rendering instruction to obtain the third texture, realizing the separation of the UI texture and avoiding the influence of the UI texture on the interpolation scene, thereby improving the accuracy of the interpolation result.

[0075] In the interpolation processing method of the embodiment of the present application, the first texture without UI and the second texture without UI are spliced to obtain the spliced texture. The feature extraction network of the interpolation model deployed in the neural network processor NPU is used to perform multi-dimensional feature extraction on the spliced texture to obtain the target optical flow feature map and the target mask feature map, reducing the influence of the UI texture on the target optical flow feature. Store the target optical flow feature map and the target mask feature map in the shared memory buffer. The GPU obtains the target optical flow feature map and the target mask feature map from the shared memory buffer, and uses the interpolation network deployed in the GPU in the interpolation model to obtain the interpolation texture without UI according to the target optical flow feature map and the target mask feature map. By setting the shared memory buffer, data copying is reduced and processing efficiency is improved. By jointly arranging the interpolation network by the NPU and the GPU, the processing efficiency of the interpolation network is improved.

[0076] Based on the above embodiments, Figure 4 is a schematic flowchart of another interpolation processing method provided by the embodiment of the present application. As Figure 4 shown, the method includes the following steps:

[0077] Step 401: Obtain the file descriptor of the shared memory buffer.

[0078] Step 402: Access the shared memory buffer according to the file descriptor.

[0079] Among them, steps 401 and 402 can refer to the explanations in the foregoing embodiments. The principles are the same and will not be elaborated here.

[0080] Step 403, obtain a first rendering instruction.

[0081] Among them, the first rendering parameters carried by the first rendering instruction include the size information and rendering type of the texture to be rendered.

[0082] Step 404, according to the first rendering parameters, perform grayscale texture rendering in the shared memory buffer. In response to determining that texture rendering is performed on the user interface UI, obtain the grayscale texture without UI rendered in the shared memory buffer, and use the grayscale texture without UI as the first texture.

[0083] Among them, the first texture and the second texture are generated in the same way, the difference is the generation time. The first texture and the second texture are two adjacent front and back frame textures. The layout of texture data in the memory of the GPU and the NPU is different. That is to say, the texture data in the NPU and the GPU is not compatible in the memory layout, that is, the size information and the rendering type are not the same. In order to achieve the compatibility of textures in the GPU and the NUP, the following method is used to generate the first texture and the second texture.

[0084] In an implementation manner of the embodiment of the present application, as an example, Figure 5A is a schematic diagram of the layout of textures in the GPU provided by the embodiment of the present application. As shown in Figure 5A shown, Figure 5A shows the arrangement of RGB textures in the GPU. As shown in Figure 5A a shows an RGB three-channel texture with a width of 5 and a height of 3. In the memory, the RGB values of the first pixel are arranged first, that is, the three pixel values marked as the number 1 are the RGB values of the first pixel, then the RGB values of the second pixel are arranged, that is, the three pixel values marked as the number 2 are the RGB values of the second pixel, and then the RGB values of the third pixel are arranged, that is, the three pixel values marked as the number 3 are the RGB values of the third pixel, and so on. Details will not be described one by one here. Figure 5B is a schematic diagram of the layout of textures in the NPU provided by the embodiment of the present application. Figure 5B shows the arrangement of RGB textures in the NPU. In the NPU, the supported data format is channel first. As shown in Figure 5BAs shown, this is an RGB three-channel texture with a width of 5 and a height of 3. In memory, the R channels of all pixels are arranged first, then the G channels of all pixels are arranged, and finally the B channels of all pixels are arranged. It can be seen that the arrangement methods of textures in the GPU and NPU are not compatible. In related technologies, functions such as permute can be used to complete the conversion between NCHW and NHWC to achieve data format matching. However, these memory rearrangement functions require a large number of reads and writes to the DDR, severely restricting the running speed of the interpolation algorithm, significantly increasing the data transmission bandwidth and power consumption. In this application, the RGB texture is converted into a grayscale texture, merging the three RGB channels into one channel. When the value of the channel is 1, the layout methods in the GPU and NPU are the same. Therefore, setting the type of the texture rendered by the GPU to a grayscale texture can achieve data layout matching between the GPU and NPU, improving the data processing efficiency in the NPU.

[0085] In addition, the width and height values of the input texture required by the frame interpolation model are set values, such as integer multiples of 64. To match the size requirements of the input texture in the NPU, the texture output by the GPU needs to match the input required by the NPU. That is, during the rendering process, data padding also needs to be completed. Padding specifically refers to adding extra pixels or values at the boundaries of an image or feature map, so that the texture data output by the GPU meets the input size of the NPU.

[0086] As an implementation method, obtain the first rendering instruction, perform rendering according to the first rendering parameters carried in the first rendering instruction, and complete data rendering according to the first rendering parameters to deliver a grayscale image without UI. The width and height values of this grayscale image without UI are both integer multiples of 64. As an implementation method, the GPU completes layout matching by adding a render pass. The newly added RenderPass is as Figure 6 shown. By using the functions glViewPort and glScissor, the state machine of OpenGL ES (GLES) is modified to achieve size matching such as width and height of the data, that is, padding is achieved. A fragment shader (Fragment Shader, FS) is created. The fragment shader is used for texture sampling and blending to achieve the conversion of the grayscale texture. Furthermore, this fragment shader is compiled and linked into newProgram, and the state machine of GLES is modified through the glUseProgram function. Subsequently, rendering is performed through the first rendering instruction drawCall, and the rendered result is a grayscale image.

[0087] Among them, for each pixel point, the conversion formula for the grayscale information of this pixel point is as follows:

[0088] grepColor = blueColor * 0.5 + redColor * 0.5 + greenColor;

[0089] Among them, grepColor is the gray value of the pixel, and blueColor, redColor, and greenColor are the color values of the red, blue, and green of the pixel.

[0090] Step 405: Obtain the first texture and the second texture rendered before the first texture.

[0091] Step 406: Stitch the first texture and the second texture to obtain a stitched texture.

[0092] Step 407: Use the feature extraction network in the interpolation model to perform multi-dimensional feature extraction on the stitched texture to obtain a target optical flow feature map and a target mask feature map.

[0093] Step 408: Use the interpolation network in the interpolation model to obtain an interpolated texture according to the target optical flow feature map and the target mask feature map.

[0094] Among them, the interpolation network is deployed in the GPU.

[0095] Among them, Steps 405 to 408 can refer to the relevant explanations in the foregoing embodiments. The principles are the same and will not be elaborated here.

[0096] In the interpolation processing method of the embodiment of the present application, the feature extraction network deployed in the neural network processor NPU of the interpolation model is used to perform multi-dimensional feature extraction on the stitched texture to obtain a target optical flow feature map and a target mask feature map, and the target optical flow feature map and the target mask feature map are stored in the shared memory buffer. The GPU obtains the target optical flow feature map and the target mask feature map from the shared memory buffer, and uses the interpolation network deployed in the GPU of the interpolation model to obtain an interpolation texture without UI according to the target optical flow feature map and the target mask feature map, and renders through the GPU to obtain a gray texture with the same size as the texture required by the NPU, realizing the unification of the texture layout between the GPU and the NPU. By constructing the shared memory buffer, the data copy between the GPU and the NPU is reduced, and the processing efficiency is improved. And by jointly deploying the interpolation network by the NPU and the GPU, the processing efficiency of the interpolation network is improved.

[0097] Based on the above embodiments, Figure 7 is a schematic flowchart of another interpolation processing method provided by the embodiment of the present application. As Figure 7 shown, the method includes the following steps:

[0098] Step 701: Obtain the first texture and the second texture rendered before the first texture.

[0099] Among them, step 701 can refer to the relevant explanations in the foregoing embodiments. Since the principle is the same, it will not be elaborated here.

[0100] Step 702: Stitch the first texture and the second texture to obtain a stitched texture.

[0101] In the embodiment of the present application, the first texture and the second texture are stitched in the channel dimension to obtain a two-channel texture, and the two-channel texture is used as the stitched texture.

[0102] Step 703: For any feature extraction sub-network, determine the feature map to be processed corresponding to the feature extraction sub-network.

[0103] Among them, the feature to be processed is determined according to the first feature map obtained by performing feature extraction on the stitched texture, or is obtained by fusing the optical flow feature map and the mask feature map output by the previous feature extraction sub-network of any feature extraction sub-network with the first feature map corresponding to the stitched texture.

[0104] In the embodiment of the present application, the feature extraction network in the interpolation model is deployed in the NPU, and the feature extraction network includes at least one feature extraction sub-network connected in series. Figure 8 It is a schematic structural diagram of a feature extraction network provided by an embodiment of the present application. As Figure 8 shown, there are 3 feature extraction sub-networks, that is, 3 blocks. The output of the previous feature extraction sub-network will be used as a part of the input of the next feature extraction sub-network.

[0105] As an implementation manner, if the feature extraction sub-network is the first feature extraction sub-network and is determined according to the first feature map obtained by performing feature extraction on the stitched texture, the first feature map corresponding to the stitched texture can be used as the feature map to be processed.

[0106] As another implementation manner, if the feature extraction sub-network is not the first feature extraction sub-network, the first feature map corresponding to the stitched texture and the optical flow feature map and the mask feature map output by the previous feature extraction sub-network are used as the feature map to be processed by this feature extraction sub-network. By fusing the optical flow feature map and the mask feature map output by each feature extraction sub-network with the first feature map, the network can combine feature information at different levels, thicken the feature map, and increase the included information volume.

[0107] In the embodiment of the present application, the feature extraction network is composed of multiple feature extraction sub-networks, and gradually extracts the features of the stitched texture from coarse to fine. The multiple feature extraction sub-networks form a series structure, which has a simple structure and less computational complexity, and is suitable for mobile devices. As Figure 8As shown in the figure, the output of each of the 3 blocks is superimposed on the first feature map of the spliced texture of the original input to thicken the feature map, which is used as the input for the next block. By superimposing the output of the block (including high-level features) on the first feature map of the spliced texture of the original input (including low-level features), the fusion of multi-scale feature maps at different levels is achieved. The fusion of multi-scale feature maps helps to enhance the network's understanding of the texture content, making the feature representation more rich and comprehensive. The superimposition operation introduces more context information, making the model more stable when facing challenges such as noise, occlusion, or deformation in the spliced texture. Even if some local features are damaged, the features at other levels can still provide useful information to improve the accuracy of the finally obtained features.

[0108] Step 704: Use any one of the feature extraction sub-networks to process the feature to be processed, and obtain the optical flow feature map and the mask feature map output by any one of the feature extraction sub-networks.

[0109] Among them, any one of the feature extraction sub-networks includes a downsampling module, a first feature extraction module, a first upsampling module optical flow feature extraction network, and a mask feature extraction module.

[0110] As an implementation manner, use the downsampling module to downsample the feature map to be processed to obtain a downsampled feature map, and then use the first feature extraction module to extract features from the downsampled feature map to obtain an intermediate feature map. The intermediate feature map includes high-dimensional features. Use the upsampling module to upsample the intermediate feature map to obtain an upsampled feature map. Use the optical flow feature extraction module to extract optical flow features from the upsampled feature map to obtain an optical flow feature map. Use the mask feature extraction module to extract mask features from the upsampled feature map to obtain a mask feature map.

[0111] As an implementation manner, if the downsampling module includes two layers of downsampling sub-modules, the optical flow feature extraction module includes an optical flow feature extraction sub-module and an upsampling sub-module, and the mask feature extraction module includes a mask feature extraction sub-module and an upsampling sub-module.

[0112] In the embodiments of the present application, the structures of each feature extraction sub-network block are the same and the parameters are similar. As an example, Figure 9 is a schematic structural diagram of a feature extraction sub-network provided by an embodiment of the present application. As shown in Figure 9As shown, the downsampling module performs downsampling to obtain a downsampled feature map. The downsampling module includes two cascaded first convolutional layers. Each first convolutional layer is a convolutional layer with a stride of 2, which can achieve 4-fold downsampling. Each first convolutional layer halves the width and height of the input feature map. Through two first convolutional layers for 4-fold downsampling, the neural network parameters can be effectively reduced and the load can be decreased. Furthermore, multiple convolutional layers included in the first feature extraction module are used to perform convolutional operations for feature extraction to obtain an intermediate feature map. The intermediate feature map includes the extracted high-dimensional features. The extracted high-dimensional features usually contain information in multiple dimensions, and each dimension represents a specific attribute of the texture, such as the correlation between gray levels in the texture, including the correlation between features such as contrast, energy, and homogeneity. Since the feature map is downsampled by the downsampling module, the computational amount processed by the first feature extraction module is reduced. The first feature extraction module includes 3 convolutional layers conv3*3, and these 3 convolutional layers do not change the width and height of the feature map output by the previous layer, and high-dimensional features, that is, the intermediate feature map, are extracted.

[0113] Furthermore, the intermediate feature map is input into the upsampling module for upsampling, that is, ConvTranspose2d is used to perform upsampling to expand the width and height of the feature map to obtain an upsampled feature map. Furthermore, two parallel convolutional layers conv3*3 included in the second feature extraction module are used to respectively implement optical flow feature extraction and fusion mask feature extraction. Furthermore, the result of feature extraction is upsampled through PixelShuffle, so that the optical flow feature map and mask feature map output by this feature extraction sub-network match the size of the first feature map input to the next feature extraction sub-network, so as to facilitate the fusion of feature maps and increase the information volume.

[0114] Step 705, use the optical flow feature map and mask feature map output by the last feature extraction sub-network as the target optical flow feature map and target mask feature map.

[0115] Among them, the optical flow feature map is a two-dimensional vector field, which is used to characterize the motion offset of each pixel point in adjacent frames, that is, the horizontal and vertical components of the motion of each pixel, and is used to predict and estimate the pixel motion between adjacent frames, so as to generate high-quality intermediate frames. The mask feature map is used to characterize which regions need to perform interpolation operations, which regions should be avoided from being modified, or which regions have higher reliability, so as to perform frame synthesis more accurately.

[0116] In the embodiments of the present application, by setting multiple feature extraction sub-networks, the output of each feature extraction sub-network is superimposed on the first feature map of the spliced texture of the original input to thicken the feature map, which is used as the input of the next feature extraction sub-network. By superimposing the output of the feature extraction sub-network (including high-level features) and the first feature map of the spliced texture of the original input (including low-level features), the fusion of multi-scale feature maps at different levels is achieved. The fusion of multi-scale feature maps helps to enhance the understanding of the texture content by the feature extraction network, making the feature representation richer and more comprehensive. The superimposition operation introduces more context information, making the feature extraction network more stable when facing challenges such as noise, occlusion, or deformation in the spliced texture. Even if some local features are damaged, the features at other levels can still provide useful information. Therefore, the output of the last feature extraction sub-network is used as the target output to improve the accuracy of the finally obtained features.

[0117] Step 706: Divide the target optical flow feature map to obtain a first target optical flow feature map and a second target optical flow feature map.

[0118] In the embodiments of the present application, it is necessary to sample the adjacent first texture and second texture based on the obtained target optical flow feature map. The target optical flow feature map is a multi-channel feature map. The target optical flow feature map can be divided according to channels to obtain a first target optical flow feature map and a second target optical flow feature map. Among them, the first target optical flow feature map includes the feature maps of the horizontal component and the vertical component. The feature map of the horizontal component includes the displacement information of each pixel point in the horizontal direction, and the feature map of the vertical component includes the displacement information of each pixel point in the vertical direction.

[0119] As an example, Figure 10 is a schematic structural diagram of an interpolation network provided by the embodiments of the present application. As Figure 10 shown, the target optical flow feature map is usually an N-channel feature map, where N is greater than or equal to 1. For example, N is 4. The feature maps of the first two channels include the feature maps of the horizontal component and the vertical component, and the feature maps of the last two channels also include the feature maps of the horizontal component and the vertical component. Then, the feature maps of the first 2 channels are used as the first target optical flow feature map, and the feature maps of the last 2 channels are used as the second target optical flow feature map. Among them, the feature map of the horizontal component includes the displacement information of each pixel point in the horizontal direction, and the feature map of the vertical component includes the displacement information of each pixel point in the vertical direction.

[0120] In the embodiments of the present application, in order to avoid sampling out-of-bounds during the sampling process using the grid_sample sampling algorithm, it is necessary to limit the range of the optical flow feature map according to the set boundary range information. As an implementation, the target optical flow feature map is divided based on channel information to obtain a first candidate optical flow feature map and a second candidate optical flow feature map. The set boundary information is used to set the boundaries of the first candidate optical flow feature map and the second candidate optical flow feature map respectively, obtaining a first target optical flow feature map and a second target optical flow feature map. For example, the range of the feature map indicated by the boundary information is (-5, 5), which limits the value range and restricts the sampling of the sampling function grid_sample, preventing sampling out-of-bounds and enhancing the stability of the model.

[0121] Step 707: Sample the first texture using the first target optical flow feature map to obtain a first sampling result.

[0122] Among them, the first sampling result indicates the starting position information of each pixel point in the first texture and the ending position information in the second texture. The starting position and ending position of each pixel indicate the optical flow vector of the pixel, which can be understood as the starting point and ending point of the optical flow corresponding to each pixel point.

[0123] It should be noted that the first texture and the second texture have the same size and are two adjacent frames of textures. The positions and quantities of pixel points in two adjacent frames of textures are corresponding.

[0124] Step 708: Sample the second texture using the second target optical flow feature map to obtain a second sampling result.

[0125] Among them, the second sampling result indicates the starting position information of each pixel point in the second texture and the ending position information in the first texture. The starting position and ending position of each pixel indicate the optical flow vector of the pixel, which can be understood as the starting point and ending point of the optical flow corresponding to each pixel point.

[0126] It should be noted that the first target optical flow feature map is used to sample the first texture, and the second target optical flow feature map is used to sample the second texture, which are set during the model training process. If they are exchanged, only the interpolation model needs to be retrained.

[0127] Step 709: Perform weighted processing on the target mask feature according to the first sampling result and the second sampling result to obtain an interpolated texture.

[0128] As an implementation, the interpolated texture can be determined by the following formula:

[0129] image1 = warp_image0 * mask + warp_image2 * (1 - mask);

[0130] Among them, image1 is the frame interpolation texture, warp_image0 is the first sampling result, warp_image2 is the second sampling result, and mask is the target mask feature map. The target mask feature map indicates the offset of each pixel point relative to the starting position of the optical flow, and 1 - mask represents the offset of each pixel point relative to the ending position of the optical flow. In the embodiments of the present application, the target mask feature is weighted and averaged using the first sampling result and the second sampling result to generate the frame interpolation texture, which can make the generated frame interpolation texture more accurately reflect the motion information and visual consistency between adjacent frames.

[0131] The embodiments of the present application make full use of the characteristics of mobile games, consider the differences in the data layouts of the GPU and NPU, can reduce the size of the input, use DMABuffer to achieve zero - copy of data, and effectively reduce the CPU load. Moreover, this solution extracts features from coarse to fine and from the whole to the local, and fuses multi - dimensional characteristic information, which can reduce artifacts and image distortion. Only use gridsample in the post - processing of the model, that is, use gridsample in the GPU, reduce the use of the gridsample operator, reduce the read - write operations on the DDR memory, save the DDR bandwidth and power consumption. In addition, the present application can directly output the generated frame interpolation texture frame, there are no hole points, and there is no need to process the occlusion relationship again.

[0132] In the frame interpolation processing method of the embodiments of the present application, obtain the first texture without UI rendered by the GPU stored in the shared memory buffer at the target moment, and obtain the second texture without UI at the historical moment before the target moment; wherein, the values of each pixel point in the first texture and the second texture are grayscale values. Concatenate the first texture and the second texture to obtain a concatenated texture. Use the feature extraction network of the frame interpolation model deployed in the neural network processor NPU to perform multi - dimensional feature extraction on the concatenated texture to obtain the target optical flow feature map and the target mask feature map, and store the target optical flow feature map and the target mask feature map in the shared memory buffer. Among them, the target optical flow feature map and the target mask feature map are used for the GPU to obtain the target optical flow feature and the target mask feature from the shared memory buffer, and use the frame interpolation network of the frame interpolation model deployed in the GPU to obtain the frame interpolation texture without UI according to the target optical flow feature map and the target mask feature map, and obtain the grayscale texture through GPU rendering, which realizes the unification of the texture layout between the GPU and the NPU, stores the grayscale texture in the shared memory buffer, reduces data copying, improves the processing efficiency, and improves the processing efficiency of the frame interpolation network by jointly arranging the frame interpolation network by the NPU and the GPU.

[0133] To implement the above - mentioned embodiments, the embodiments of the present application also propose a frame interpolation device.

[0134] Figure 11 This is a schematic structural diagram of an interpolation frame device provided by an embodiment of the present application.

[0135] As Figure 11 shown, the device may include:

[0136] An acquisition module 111, configured to acquire a first texture and a second texture rendered before the first texture.

[0137] A splicing module 112, configured to splice the first texture and the second texture to obtain a spliced texture.

[0138] A feature extraction module 113, configured to perform multi-dimensional feature extraction on the spliced texture by using a feature extraction network in the interpolation frame model to obtain a target optical flow feature map and a target mask feature map; wherein, the feature extraction network is deployed in an NPU.

[0139] An interpolation frame module 114, configured to obtain an interpolation frame texture according to the target optical flow feature map and the target mask feature map by using an interpolation frame network in the interpolation frame model; wherein, the interpolation frame network is deployed in a GPU.

[0140] Furthermore, in an implementation manner of an embodiment of the present application, the device further includes a determination module, configured to:

[0141] Obtain a file descriptor of a shared memory buffer;

[0142] Access the shared memory buffer according to the file descriptor;

[0143] Perform texture rendering in the shared memory buffer, and in response to determining to perform texture rendering on a user interface UI, obtain a texture without UI rendered in the shared memory buffer;

[0144] Use the texture without UI as the first texture.

[0145] In an implementation manner of an embodiment of the present application, the determination module is further configured to:

[0146] Obtain a first rendering instruction; wherein, the first rendering parameters carried by the first rendering instruction include size information and rendering type of a texture to be rendered;

[0147] Perform grayscale type texture rendering in the shared memory buffer according to the first rendering parameters. In an implementation manner of an embodiment of the present application, the determination module is further configured to:

[0148] In response to determining to perform texture rendering on a user interface UI, switch the rendering buffer from the shared memory buffer to a target buffer;

[0149] Render the UI on the target buffer according to the target rendering instruction to obtain a third texture;

[0150] Fuse the interpolation texture and the third texture to obtain a target interpolation texture.

[0151] In an implementation manner of the embodiment of the present application, the feature extraction network includes at least one feature extraction sub-network. The feature extraction module 113 is further configured to:

[0152] For any one of the feature extraction sub-networks, determine the feature map to be processed corresponding to the feature extraction sub-network; wherein, the feature map to be processed is determined according to the first feature map obtained by performing feature extraction on the spliced texture, or is obtained by fusing the optical flow feature map and the mask feature map output by the previous feature extraction sub-network of any one of the feature extraction sub-networks with the first feature map;

[0153] Use any one of the feature extraction sub-networks to process the feature map to be processed to obtain the optical flow feature map and the mask feature map output by any one of the feature extraction sub-networks;

[0154] Use the optical flow feature map and the mask feature map output by the last feature extraction sub-network as the target optical flow feature map and the target mask feature map.

[0155] In an implementation manner of the embodiment of the present application, any one of the feature extraction sub-networks includes a downsampling module, a first feature extraction module, an upsampling module, and a second feature extraction module. The feature extraction module 113 is further configured to:

[0156] Use the convolutional layer in the downsampling module to downsample the feature map to be processed to obtain a downsampled feature map;

[0157] Use the first feature extraction module to perform feature extraction on the downsampled feature map to obtain an intermediate feature map;

[0158] Use the upsampling module to upsample the intermediate feature map to obtain an upsampled feature map;

[0159] Use the second feature extraction module to perform feature extraction on the upsampled feature map to obtain the optical flow feature map and the mask feature map output by any one of the feature extraction sub-networks.

[0160] In an implementation manner of the embodiment of the present application, the splicing module 112 is further configured to:

[0161] Splice the first texture and the second texture in the channel dimension to obtain the spliced texture.

[0162] In an implementation manner of the embodiment of the present application, the frame interpolation module 114 is further configured to:

[0163] Divide the target optical flow feature map to obtain a first target optical flow feature map and a second target optical flow feature map; wherein, both the first target optical flow feature map and the second target optical flow feature map include displacement information of pixel points in the horizontal direction and displacement information in the vertical direction;

[0164] Sample the first texture using the first target optical flow feature map to obtain a first sampling result; wherein, the first sampling result indicates the starting position information of each pixel point in the first texture and the ending position information in the second texture; sample the second texture using the second target optical flow feature map to obtain a second sampling result; wherein, the second sampling result indicates the starting position information of each pixel point in the second texture and the ending position information in the first texture;

[0165] Perform weighted processing on the target mask feature according to the first sampling result and the second sampling result to obtain the frame interpolation texture.

[0166] In an implementation manner of the embodiment of the present application, the frame interpolation module 114 is further configured to:

[0167] Divide the target optical flow feature map based on channel information to obtain a first candidate optical flow feature map and a second candidate optical flow feature map; wherein, the channel information includes horizontal channel information and vertical channel information;

[0168] Use the set boundary information to set the boundaries of the first candidate optical flow feature map and the second candidate optical flow feature map respectively to obtain a first target optical flow feature map and a second target optical flow feature map.

[0169] It should be noted that the foregoing explanation of the method embodiment also applies to the device of this embodiment, and will not be repeated here.

[0170] In the frame interpolation device of the embodiment of the present application, the first texture and the second texture rendered before the first texture are obtained. Both the first texture and the second texture are textures without UI. The first texture and the second texture are spliced to obtain a spliced texture. The feature extraction network in the frame interpolation model is used to perform multi-dimensional feature extraction on the spliced texture to obtain a target optical flow feature map and a target mask feature map. Among them, the feature extraction network is deployed in the NPU. The frame interpolation network in the frame interpolation model is used to obtain the frame interpolation texture according to the target optical flow feature map and the target mask feature map. Among them, the frame interpolation network is deployed in the GPU. By deploying the frame interpolation model in the GPU and the NPU, GPU and NPU hybrid programming is realized, giving full play to the algorithm advantages of the two kinds of hardware and improving the processing efficiency of the frame interpolation model.

[0171] To implement the above embodiments, the present application also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the method described in the foregoing method embodiments is implemented.

[0172] To implement the above embodiments, the present application also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the method described in the foregoing method embodiments is implemented.

[0173] To implement the above embodiments, the present application also provides a computer program product, on which a computer program is stored. When the computer program is executed by a processor, the method described in the foregoing method embodiments is implemented.

[0174] Figure 12 FIG. is a block diagram of an electronic device provided by an embodiment of the present application. For example, the electronic device 800 may be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.

[0175] Referring to Figure 12 , the electronic device 800 may include one or more of the following components: a processing component 802, a memory 804, a power component 806, a multimedia component 808, an audio component 810, an input / output (I / O) interface 812, a sensor component 814, and a communication component 816.

[0176] The processing component 802 generally controls the overall operation of the electronic device 800, such as operations associated with display, telephone calls, data communications, camera operations, and recording operations. The processing component 802 may include one or more processors 820 to execute instructions to complete all or part of the steps of the above method. In addition, the processing component 802 may include one or more modules to facilitate the interaction between the processing component 802 and other components. For example, the processing component 802 may include a multimedia module to facilitate the interaction between the multimedia component 808 and the processing component 802.

[0177] The memory 804 is configured to store various types of data to support the operation of the electronic device 800. Examples of such data include instructions for any application or method operating on the electronic device 800, contact data, phone book data, messages, pictures, videos, and the like. The memory 804 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk.

[0178] The power component 806 provides power to various components of the electronic device 800. The power component 806 may include a power management system, one or more power sources, and other components associated with generating, managing, and distributing power for the electronic device 800.

[0179] The multimedia component 808 includes a screen that provides an output interface between the electronic device 800 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can not only sense the boundaries of touch or swipe actions, but also detect the duration and pressure associated with the touch or swipe operation. In some embodiments, the multimedia component 808 includes a front camera and / or a rear camera. When the electronic device 800 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each of the front camera and the rear camera can be a fixed optical lens system or have focal length and optical zoom capabilities.

[0180] The audio component 810 is configured to output and / or input audio signals. For example, the audio component 810 includes a microphone (MIC) that is configured to receive external audio signals when the electronic device 800 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signals can be further stored in the memory 804 or transmitted via the communication component 816. In some embodiments, the audio component 810 further includes a speaker for outputting audio signals.

[0181] The I / O interface 812 provides an interface between the processing component 802 and a peripheral interface module, which can be a keyboard, a click wheel, buttons, etc. These buttons can include, but are not limited to: a home button, a volume button, a power button, and a lock button.

[0182] The sensor assembly 814 includes one or more sensors for providing an assessment of various aspects of the status of the electronic device 800. For example, the sensor assembly 814 can detect the on / off state of the electronic device 800, the relative positioning of components, such as the display and keypad of the electronic device 800. The sensor assembly 814 can also detect a change in the position of the electronic device 800 or a component of the electronic device 800, the presence or absence of user contact with the electronic device 800, the orientation or acceleration / deceleration of the electronic device 800, and a change in the temperature of the electronic device 800. The sensor assembly 814 can include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor assembly 814 can also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor assembly 814 can also include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.

[0183] The communication component 816 is configured to facilitate communication between the electronic device 800 and other devices in a wired or wireless manner. The electronic device 800 can access a wireless network based on communication standards, such as WiFi, 4G, or 5G, or a combination thereof. In an exemplary embodiment, the communication component 816 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 816 further includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.

[0184] In an exemplary embodiment, the electronic device 800 can be implemented by one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components for performing the above method.

[0185] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as the memory 804 including instructions, and the above instructions can be executed by the processor 820 of the electronic device 800 to complete the above method. For example, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc.

[0186] In the description of this specification, the descriptions referring to terms such as "one embodiment", "some embodiments", "examples", "specific examples", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of this application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, without contradiction, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0187] In addition, the terms "first" and "second" are used only for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In the description of this application, "a plurality of" means at least two, such as two, three, etc., unless otherwise specifically defined.

[0188] Any process or method description shown in a flowchart or described in other ways herein may be understood to represent a module, segment, or portion of code including one or more executable instructions for implementing a customized logic function or process, and the scope of the preferred embodiments of this application includes additional implementations, where the functions may be executed in a manner that is not shown or discussed in the order, including in a substantially simultaneous manner according to the functions involved or in the reverse order, which should be understood by those skilled in the art to which the embodiments of this application pertain.

[0189] The logic and / or steps represented in the flowchart or otherwise described herein can, for example, be considered as a definable sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or in conjunction with these instruction execution systems, apparatuses, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include the following: an electrical connection portion having one or more wirings (electronic device), a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or otherwise processing as appropriate, and then storing it in a computer memory.

[0190] It should be understood that various parts of the present application can be implemented using hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented using hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application-specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.

[0191] Those of ordinary skill in the art of this technology can understand that all or part of the steps carried by the method of the above-described embodiments can be completed by instructing relevant hardware through a program, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiments.

[0192] In addition, each functional unit in various embodiments of the present application may be integrated into one processing module, or each unit may exist physically alone, or two or more units may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. When the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.

[0193] The above-mentioned storage medium may be a read-only memory, a magnetic disk or an optical disc, etc. Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present application. Those of ordinary skill in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present application.

Claims

1. A frame interpolation processing method, characterized in that Including: Obtain a first texture and a second texture rendered before the first texture; Stitch the first texture and the second texture to obtain a stitched texture; Use the feature extraction network in the interpolation model to perform multi-dimensional feature extraction on the stitched texture to obtain a target optical flow feature map and a target mask feature map; wherein, the feature extraction network is deployed in a neural network processor NPU; Use the interpolation network in the interpolation model to obtain an interpolated texture according to the target optical flow feature map and the target mask feature map; wherein, the interpolation network is deployed in a graphics processing unit GPU.

2. The method according to claim 1, wherein The method further includes: Obtain a file descriptor of a shared memory buffer; Access the shared memory buffer according to the file descriptor; Perform texture rendering in the shared memory buffer, and in response to determining to perform texture rendering on a user interface UI, obtain a texture without UI rendered in the shared memory buffer; Use the texture without UI as the first texture.

3. The method according to claim 2, wherein The performing texture rendering in the shared memory buffer includes: Obtain a first rendering instruction; wherein, the first rendering parameters carried by the first rendering instruction include size information and rendering type of the texture to be rendered; Perform grayscale-type texture rendering in the shared memory buffer according to the first rendering parameters.

4. The method according to claim 2, wherein The method further includes: In response to determining to perform texture rendering on a user interface UI, switch the rendering buffer from the shared memory buffer to a target buffer; Render the UI in the target buffer according to a target rendering instruction to obtain a third texture; Fuse the interpolated texture and the third texture to obtain a target interpolated texture.

5. The method according to any one of claims 1 to 4, characterized in that, The feature extraction network includes at least one feature extraction sub-network, and the using the feature extraction network in the interpolation model to perform feature extraction on the stitched texture to obtain a target optical flow feature map and a target mask feature map includes: For any one feature extraction sub-network, determine a to-be-processed feature map corresponding to the feature extraction sub-network; wherein, the to-be-processed feature map is determined according to a first feature map obtained by performing feature extraction on the stitched texture, or is obtained by fusing an optical flow feature map and a mask feature map output by the previous feature extraction sub-network of the any one feature extraction sub-network with the first feature map; Use the any one feature extraction sub-network to process the to-be-processed feature map to obtain an optical flow feature map and a mask feature map output by the any one feature extraction sub-network; Use the optical flow feature map and the mask feature map output by the last feature extraction sub-network as the target optical flow feature map and the target mask feature map.

6. The method according to claim 5, wherein Any one of the feature extraction sub-networks includes a downsampling module, a first feature extraction module, an upsampling module, and a second feature extraction module, and the using the any one feature extraction sub-network to process the to-be-processed feature map to obtain an optical flow feature map and a mask feature map output by the any one feature extraction sub-network includes: Use a convolutional layer in the downsampling module to perform downsampling on the to-be-processed feature map to obtain a downsampled feature map; Use the first feature extraction module to extract features from the downsampled feature map to obtain an intermediate feature map; Use the upsampling module to upsample the intermediate feature map to obtain an upsampled feature map; Use the second feature extraction module to extract features from the upsampled feature map to obtain the optical flow feature map and the mask feature map output by any one of the feature extraction sub-networks.

7. The method according to any one of claims 1-4, characterized in that, The step of splicing the first texture and the second texture to obtain a spliced texture includes: Splice the first texture and the second texture in the channel dimension to obtain the spliced texture.

8. The method according to claim 5, characterized in that, The step of using the interpolation network in the interpolation model to obtain an interpolated texture according to the target optical flow feature map and the target mask feature map includes: Divide the target optical flow feature map to obtain a first target optical flow feature map and a second target optical flow feature map; wherein, both the first target optical flow feature map and the second target optical flow feature map include displacement information of pixel points in the horizontal direction and displacement information in the vertical direction; Sample the first texture with the first target optical flow feature map to obtain a first sampling result; wherein, the first sampling result indicates the starting position information of each pixel point in the first texture and the ending position information in the second texture; Sample the second texture with the second target optical flow feature map to obtain a second sampling result; wherein, the second sampling result indicates the starting position information of each pixel point in the second texture and the ending position information in the first texture; Perform weighted processing on the target mask feature according to the first sampling result and the second sampling result to obtain the interpolated texture.

9. The method according to claim 8, wherein The step of dividing the target optical flow feature map to obtain a first target optical flow feature map and a second target optical flow feature map includes: Divide the target optical flow feature map based on channel information to obtain a first candidate optical flow feature map and a second candidate optical flow feature map; wherein, the channel information includes horizontal channel information and vertical channel information; Use the set boundary information to set the boundaries of the first candidate optical flow feature map and the second candidate optical flow feature map respectively to obtain the first target optical flow feature map and the second target optical flow feature map.

10. An interpolation processing device, characterized in that, Comprises: An acquisition module, configured to acquire a first texture and a second texture rendered before the first texture; A splicing module, configured to splice the first texture and the second texture to obtain a spliced texture; A feature extraction module, configured to perform multi-dimensional feature extraction on the spliced texture by using a feature extraction network in an interpolation model to obtain a target optical flow feature map and a target mask feature map; wherein, the feature extraction network is deployed in a neural network processor NPU; An interpolation module, configured to use the interpolation network in the interpolation model to obtain an interpolated texture according to the target optical flow feature map and the target mask feature map; wherein, the interpolation network is deployed in a graphics processing unit GPU.

11. An electronic device, characterized in that, Comprises a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, the method described in any one of claims 1-9 is implemented.

12. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method according to any one of claims 1-9.

Citation Information

Cited By

  • Image processing method and device and medium

    CN121353491A

  • Method, medium, product and computing device for dynamically adjusting frame insertion computation

    CN121486636A