Method and device for optimizing rendering program

By building a resource consumption matrix and optimizing GPU rendering programs using machine learning algorithms, the problem of low accuracy of traditional manual experience optimization is solved, and more efficient GPU rendering performance is achieved.

CN120411274APending Publication Date: 2025-08-01JINAN INSPUR DATA TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510409384.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-01
Publication Date
2025-08-01

AI Technical Summary

Technical Problem

In the prior art, the accuracy of optimizing GPU rendering programs is low by relying on traditional manual experience, and it is difficult to significantly improve performance.

Method used

By obtaining the resource consumption data when different types of rendering programs run in the graphics processing unit, building a resource consumption matrix, determining key resource types and other resource types, using machine learning algorithms such as random forests and extreme random tree algorithms to determine the correlation between resource types, and performing sorting optimization.

Benefits of technology

It realizes more accurate evaluation of the performance characteristics of the GPU rendering program, improves optimization accuracy, and improves the performance of the rendering program and has a more ideal picture quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120411274A_ABST
    Figure CN120411274A_ABST
Patent Text Reader

Abstract

The invention discloses a method for optimizing a rendering program, which relates to the technical field of computers, and comprises the following steps: acquiring the quantity of resources actually consumed under M resource types when different types of rendering programs run in a graphic processing unit, determining a key resource type and other resource types in the M resource types, according to the incidence relation between the key resource type and other resource types, sorting the other resource types to obtain a target sorting result of the influence degree of each other resource type on the rendering performance of the rendering program, so that the main characteristics influencing the performance of the GPU rendering program can be assessed observably; and the rendering program is optimized based on the determined resource type with the top ranking in the target ranking result, so that the problem of relatively low accuracy rate of optimizing the GPU rendering program by depending on traditional artificial experience in related technologies is solved, and the effect of improving the optimization accuracy rate of the GPU rendering program is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular, to a method and device for optimizing a rendering program. Background Art

[0002] The GPU (Graphics Processing Unit) rendering program is a crucial link in graphics processing, and its performance directly affects the overall application response speed. With the wide application of computer graphics rendering in various fields of the digital age, the performance optimization of the GPU rendering program has become crucial. However, the performance optimization of the GPU rendering program is a complex process, and its complexity stems from the huge scale of the GPU rendering program code and data, as well as the complexity of the GPU working principle.

[0003] Currently, the GPU rendering program is usually optimized relying on traditional manual experience. However, manual experience has a certain degree of subjectivity and it is difficult to objectively evaluate the main features that affect the performance of the GPU rendering program. The optimization result of the GPU rendering program is not ideal and cannot significantly improve the performance of the GPU rendering program. Therefore, deeply understanding the main features of the GPU rendering program and optimizing it is an important task. Summary of the Invention

[0004] This application provides a method and device for optimizing a rendering program to at least solve the technical problem of low accuracy in optimizing the GPU rendering program relying on traditional manual experience in related technologies.

[0005] This application provides a method for optimizing a rendering program, including: obtaining resource consumption data of N rendering programs under M resource types when running in a graphics processing unit to obtain a resource consumption matrix, where N and M are integers greater than 0, and the N rendering programs are rendering programs belonging to different types of graphics processing units; determining a key resource type and other resource types among the M resource types, and determining the resource consumption data of the key resource type as key resource consumption data and the resource consumption data of other resource types as other resource consumption data; determining the relevance between each other resource type and the key resource type through a target sorting model according to the resource consumption matrix, and sorting the other resource types based on the relevance to obtain a target sorting result, where the target sorting result is used to identify the influence degree of each other resource type on the rendering performance of the N rendering programs; optimizing the rendering program through the target sorting result to obtain a target rendering program.

[0006] The present application also provides a device for optimizing a rendering program, including: an acquisition module, configured to acquire resource consumption data of N rendering programs under M resource types when running in a graphics processing unit, so as to obtain a resource consumption matrix, where N and M are integers greater than 0, and the N rendering programs are rendering programs belonging to different types of graphics processing units; a determination module, configured to determine a key resource type and other resource types among the M resource types, and determine the resource consumption data of the key resource type as key resource consumption data and the resource consumption data of other resource types as other resource consumption data; a sorting module, configured to determine the relevance between each of the other resource types and the key resource type through a target sorting model according to the resource consumption matrix, and sort the other resource types based on the relevance to obtain a target sorting result, where the target sorting result is used to indicate the influence degree of each of the other resource types on the rendering performance of the N rendering programs; an optimization module, configured to optimize the rendering program through the target sorting result to obtain a target rendering program.

[0007] The present application also provides an electronic device, including: a memory, configured to store a computer program; a processor, configured to implement the steps of any one of the above methods for optimizing a rendering program when executing the computer program.

[0008] The present application also provides a computer-readable storage medium, in which a computer program is stored, where the computer program implements the steps of any one of the above methods for optimizing a rendering program when being executed by a processor.

[0009] The present application also provides a computer program product, including a computer program, where the computer program implements the steps of any one of the above methods for optimizing a rendering program when being executed by a processor.

[0010] Through this application, by obtaining the actual amount of resources consumed by different types of rendering programs under M resource types when running in a graphics processing unit, key resource types and other resource types are determined among the M resource types, and the resource types that are most important to the overall rendering performance of the rendering program are obtained, that is, the resource types that best represent the rendering performance of the rendering program. Then, according to the correlation between the key resource types and other resource types, the other resource types are sorted, and a ranking result of the degree of influence of each other resource type on the rendering performance of the rendering program is obtained, that is, the target ranking result. In this way, the main features that affect the performance of the GPU rendering program can be objectively evaluated, and the rendering program can be optimized based on the resource types that are ranked high in the target ranking result, thereby avoiding the problem of low accuracy caused by the subjectivity of manual experience in optimizing the GPU rendering program in the prior art, and thus solving the problem of low accuracy in optimizing the GPU rendering program by relying on traditional manual experience in the related art, achieving the effect of improving the optimization accuracy of the GPU rendering program, making the picture rendered by the GPU rendering program more ideal. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] In order to more clearly illustrate the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0012] Figure 1 A hardware structure block diagram of a mobile terminal according to a method for optimizing a rendering program provided in an embodiment of the present application;

[0013] Figure 2 A flowchart of a method for optimizing a rendering program according to an embodiment of the present application;

[0014] Figure 3 1 is a diagram showing a resource consumption matrix according to an embodiment of the present application;

[0015] Figure 4 Schematic diagram of a method for optimizing a rendering program according to an embodiment of the present application;

[0016] Figure 5 This is a comparison chart of the importance ranking of four GPU random forest features according to an embodiment of the present application;

[0017] Figure 6 This is a comparison chart of the ExtraTrees algorithm feature importance ranking of four GPUs according to an embodiment of the present application;

[0018] Figure 7Comparison and analysis of the true and predicted values of the sm_efficiency feature on four GPUs according to an embodiment of the present application;

[0019] Figure 8 Structural block diagram of an apparatus for optimizing a rendering program according to an embodiment of the present application. Detailed implementation manners

[0020] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.

[0021] It should be noted that in the description of the present application, the terms "include", "comprise" or any other variation thereof are intended to cover a non-exclusive inclusion, such that a process, method, article or device including a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. The terms "first", "second", etc. in the present application are used to distinguish similar objects and are not used to describe a specific order or sequence.

[0022] In order to enable those skilled in the art of the present technology to better understand the solutions of the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific implementation manners.

[0023] Combined with the specific application environment architecture or specific hardware architecture on which the execution of the method for optimizing the rendering program depends, the specific application environment architecture or specific hardware architecture is described herein.

[0024] The method embodiments provided in the embodiments of the present application can be executed on a mobile terminal, a computer terminal or a similar computing device. Taking the execution on a mobile terminal as an example, Figure 1 is a hardware structural block diagram of a mobile terminal for a method of optimizing a rendering program according to an embodiment of the present application. As Figure 1 shown, the mobile terminal may include one or more ( Figure 1 only one is shown in Figure 1 processors 102 (the processors 102 may include, but are not limited to, processing devices such as a microprocessor MCU or a programmable logic device FPGA) and a memory 104 for storing data. Among them, the above mobile terminal may further include a transmission device 106 for communication functions and an input / output device 108. Those of ordinary skill in the art can understand that Figure 1more or fewer components shown, or having a configuration different from that shown in Figure 1 that shown.

[0025] The memory 104 can be used to store computer programs, for example, software programs and modules of application software, such as the computer program corresponding to the method of optimizing the rendering program in the embodiments of the present application. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, implements the above-mentioned method. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some instances, the memory 104 may further include a memory remotely disposed relative to the processor 102, and these remote memories can be connected to the mobile terminal through a network. Examples of the above-mentioned network include but are not limited to the Internet, enterprise intranets, local area networks, mobile communication networks, and combinations thereof.

[0026] The transmission device 106 is used to receive or send data via a network. Specific examples of the above-mentioned network may include a wireless network provided by a communication provider of the mobile terminal. In one instance, the transmission device 106 includes a network adapter (Network Interface Controller, abbreviated as NIC), which can be connected to other network devices through a base station and thus can communicate with the Internet. In one instance, the transmission device 106 may be a radio frequency (Radio Frequency, abbreviated as RF) module, which is used to communicate with the Internet wirelessly.

[0027] Embodiments of the present application provide a method for optimizing a rendering program. The method is described in detail in combination with the execution flow of the method for optimizing the rendering program.

[0028] For a better understanding of the following embodiments, the following explains the professional terms in the embodiments of the present application:

[0029] GPU: Graphics Processing Unit, a graphics processing unit, is a processor specifically designed to process graphic information such as images and videos. The GPU was initially designed to improve the performance of computer graphics processing and has now been widely used in various scenarios that require a large amount of parallel computing, such as gaming, deep learning, scientific computing, and other fields.

[0030] PCA: Principal Component Analysis, a method of principal component analysis, is a commonly used data dimensionality reduction technique. It reduces the dimensionality of data by projecting the original data onto fewer principal component axes while retaining as much of the main variation information as possible. This method uses orthogonal transformation to convert a set of variables that may be correlated into a set of linearly uncorrelated variables, i.e., principal components, thus helping researchers better understand the internal structure and characteristics of the data and simplify subsequent data analysis and processing.

[0031] NVProf: A CUDA application performance analysis tool that can detailedly record and analyze the execution process of CUDA programs, including key performance metrics such as kernel run time and memory transfer, helping developers locate performance bottlenecks and optimize code. Through NVProf, developers can obtain comprehensive performance data, thus effectively improving the running efficiency of CUDA applications.

[0032] Rodinia: The Rodinia benchmark suite is a performance test suite designed specifically for parallel and multi-core processor architectures. It contains a series of scientific computing benchmark programs for evaluating and comparing the performance of different high-performance computing (HPC) platforms.

[0033] Parboil: The Parboil benchmark suite is a set of benchmark programs focused on evaluating parallel computing performance, designed specifically for evaluating the performance of modern parallel and multi-core processors. It contains multiple parallel computing-intensive applications from different fields, such as image processing and scientific computing.

[0034] Octane: The Octane renderer is a high-performance GPU-accelerated rendering engine designed specifically for previewing 3D graphics and visual effects. It makes full use of the powerful computing capabilities of modern GPUs to provide real-time rendering and high-quality preview functions, enabling designers and artists to create and iterate 3D works more quickly and efficiently.

[0035] Warp: In the GPU architecture, a computing unit warp refers to a set of threads that execute simultaneously. It is one of the basic units of GPU parallel computing. Each warp contains 32 threads, and these threads execute in SIMD (Single Instruction Multiple Data) mode, that is, they execute the same instruction simultaneously but operate on different data. The existence of warps enables GPUs to efficiently process a large number of parallel computing tasks and is the key to achieving high-performance computing.

[0036] DRAM: Dynamic Random-Access Memory, a type of computer memory that is widely used. It can store and read data in a short time and supports random access. DRAM maintains data by refreshing the charge in the capacitor regularly. Due to its high speed, low cost, and moderate storage capacity, it is often used as the main memory in computer systems, directly exchanging data with the CPU to meet the requirements of fast data reading and writing.

[0037] inst_integer: The number of integer instructions executed by non-predicted threads, which is a performance metric used to measure the number of integer instructions executed by non-predicted threads in a processor. This metric reflects the efficiency of the processor in handling integer arithmetic tasks. Especially in a multi-threaded environment, the number of integer instructions executed by non-predicted threads directly affects the overall processing power and performance. By analyzing and optimizing inst_integer, the processor resources can be utilized more effectively, improving the system's computing efficiency.

[0038] l2_utilization: L2 cache utilization rate, which is a performance metric measuring the usage level of the L2 cache relative to its peak utilization rate, with a value range from 0 to 10. This metric reflects the degree to which the L2 cache is effectively utilized within a given time.

[0039] achieved_occupancy: GPU occupancy rate, which is a metric measuring the ratio of the number of active warps in each active cycle in a GPU multiprocessor to the maximum number of warps supported by that multiprocessor. It reflects the actual utilization degree of the warp resources by the multiprocessor during the GPU execution process.

[0040] inst_fp_32: Refers to the number of single-precision floating-point instructions executed by non-predicted threads. These instructions cover various types such as arithmetic operations and comparison operations. This metric is a key parameter for measuring the workload of the processor in handling single-precision floating-point arithmetic tasks, reflecting the processor's response ability and execution efficiency for floating-point calculation requirements.

[0041] global_store_requests: Refers to the total number of global store requests issued by the multiprocessor. This metric covers all requests for the processor to write data to the global memory, excluding atomic operation requests. It is an important metric for measuring the frequency of interaction between the multiprocessor and the global memory. By analyzing global_store_requests, the performance of the processor in data storage and access can be understood, providing a basis for optimizing system performance and improving data processing efficiency.

[0042] NVCC: CUDA Compiler, that is, the CUDA compiler. It is a command-line tool for compiling and linking CUDA programs, which can compile C / C++ code containing CUDA extensions into binary files that can be executed on the GPU.

[0043] In this embodiment, a method for running an optimized rendering program on the above-mentioned mobile terminal is provided. Figure 2 It is a flowchart of the method for the optimized rendering program according to the embodiment of the present application, as Figure 2 shown, and this process includes the following steps:

[0044] Step S202, obtain the resource consumption data of N rendering programs under M resource types when running in the graphics processing unit, and obtain a resource consumption matrix, where N and M are integers greater than 0, and the N rendering programs are rendering programs belonging to different types of the graphics processing unit;

[0045] Optionally, in the embodiment of the present application, the rendering program may be a specific instance of a graphics processing task running in a graphics processing unit (GPU), and is usually used to generate visual effects of images or videos. Among them, the rendering program may include but is not limited to the following types: real-time rendering program, offline rendering program, physical rendering program, ray tracing program, rasterization rendering program, shader program, compute shader, hybrid rendering program, Octane renderer program, etc.

[0046] Optionally, in the embodiment of the present application, the graphics processing unit may be a microprocessor specifically designed to process image and video data, and it may be various types of devices, including but not limited to discrete graphics cards, integrated graphics cards, workstation-level GPUs, data center GPUs, mobile GPUs, embedded GPUs, cloud GPUs, etc. The architecture and performance characteristics of the GPU play an important role in different application scenarios. From these types of GPUs, a suitable GPU structure can be selected according to actual needs.

[0047] Optionally, in the embodiments of the present application, the resource type may be the resource type that may be consumed when a rendering program runs on a GPU, including but not limited to computing units, memory bandwidth, video memory, texture units, rendering output units, power consumption, shader clock frequency, cache usage, computing power utilization, thread concurrency, instruction throughput, PCIe, etc. By knowing the resource types that may be consumed when a rendering program runs on a GPU, the consumption of each rendering program on different resources can be understood, so as to make optimization or resource allocation decisions. For example, for a specific GPU architecture, the resource consumption matrix may include the following resource types: stream processor efficiency (sm_efficiency), global store requests (global_store_requests), number of integer instructions for non-predicted thread execution (inst_integer), number of single-precision floating-point instruction executions (inst_fp_32), L2 cache utilization (l2_utilization), GPU occupancy (achieved_occupancy), texture unit utilization (tex_sm_utilization), system memory write bandwidth (sysmem_write_throughput), shared memory read bandwidth (shared_load_throughput), L2 texture write hit rate (l2_tex_write_hit_rate), etc.

[0048] Optionally, in the embodiments of the present application, the resource consumption matrix may be a matrix composed of the actual consumed resource amounts when a rendering program runs on a GPU. Each row in this resource consumption matrix represents a rendering program, each column represents a resource type, and the element value in the matrix represents the consumption amount of the corresponding program under a specific resource type. The specific matrix form may be as Figure 3 shown. By analyzing this resource consumption matrix, it can be identified which resource types are bottlenecks and which are efficiently utilized in the rendering program, so as to guide the design of GPU hardware, the optimization of rendering algorithms, and the formulation of resource scheduling strategies to improve the overall performance and energy efficiency of the GPU.

[0049] Optionally, the above resource consumption matrix can be obtained by using performance analysis tools such as NVProf, Nsight Compute (a powerful CUDA kernel analysis tool that can provide detailed performance metrics and guidance to help optimize GPU programs), or other analysis tools. When running various rendering programs on the GPU, record the performance metric data of each program during execution. Specifically, by using a variety of performance counters provided by the GPU hardware, accurately collect data on GPU utilization, memory bandwidth, utilization of shader processing units, texture processing rate, frame rate, power consumption, etc., and access these counters in CUDA. By collecting detailed usage data using the performance counters provided by the GPU hardware and accessing these counters in CUDA, it can help developers and system administrators monitor and analyze the performance of the GPU, optimize the code accordingly to improve the overall execution efficiency. At the same time, it helps evaluate the performance of different versions of programs or different hardware platforms to make better hardware or software choices, significantly enhancing the performance, efficiency, and reliability of GPU-accelerated applications.

[0050] Through the above, obtain the resource consumption data of various types of rendering programs under M resource types and form a resource consumption matrix. Through the resource consumption matrix, it is possible to understand the consumption of various resources by different programs, better manage and schedule resources, avoid resource waste or contention, and improve the utilization rate of the graphics processing unit. At the same time, analyzing the usage of different rendering programs on different resource types helps identify the performance bottlenecks of the programs and provides a basis for optimizing resource allocation, thereby improving the overall system efficiency. In addition, through the resource consumption matrix, it is also possible to compare the resource consumption differences of different types of rendering programs on the same or different graphics processing units, providing data support for selecting the appropriate program and achieving comprehensive optimization and management of the graphics processing unit and rendering programs.

[0051] Step S204, determine the critical resource type and other resource types among the M resource types, and determine the resource consumption data of the critical resource type as critical resource consumption data, and the resource consumption data of the other resource types as other resource consumption data;

[0052] Optionally, in the embodiments of the present application, the critical resource type may be the resource type that has the most important impact on the overall performance, efficiency, etc. of the rendering program among the resource types that the rendering program may consume when running on the GPU. The critical resource type may be one or more; the other resource types may be the resource types other than the critical resource type among the M resource types.

[0053] Optionally, algorithms such as principal component analysis (PCA), successive projection algorithm (SPA), competitive adaptive reweighted sampling algorithm (CARS), etc. can be used to effectively select key resource types from M resource types.

[0054] Through the above content, classifying resources into key resources and other resources can accurately identify and manage resource types, significantly improve resource utilization efficiency and management level, so as to better carry out resource management and optimization.

[0055] Step S206, determine the relevance between each of the other resource types and the key resource types through the target sorting model based on the resource consumption matrix, and sort the other resource types based on the relevance to obtain a target sorting result, where the target sorting result is used to identify the influence degree of each of the other resource types on the rendering performance of the N rendering programs;

[0056] Optionally, in the embodiments of the present application, the target sorting result can be a list or sequence, which contains the sorting results of the relevance between each other resource type and the key resource type, and this sorting reflects the relative degree of influence of each other resource type on the rendering program performance (or key resource type). Specifically, the target sorting result can be arranged in descending or ascending order according to the influence degree of the other resource type on the key resource type, so as to help identify which resource types are the most critical to the rendering performance, and which resource type optimizations may be the most significant for overall performance improvement. For example, if the target sorting result is a list arranged in descending order, then the resource types ranked in the front have the greatest influence on the rendering performance and may need to be optimized or concerned about first; while the resource types ranked in the back have relatively less influence and can be temporarily ignored or placed in a secondary position for processing when resources are limited. Through the target sorting result, it helps with resource allocation, performance optimization decision-making, and related analysis to improve the overall efficiency and performance of the rendering program.

[0057] Optionally, an ensemble learning algorithm or other machine learning algorithms, such as a deep learning model (such as a neural network), etc. can be used to determine the target sorting result of the other resource types, so as to sort the feature importance of the other resources. It is also possible to perform a fusion strategy for different algorithms, such as combining ensemble learning and deep learning, to utilize the advantages of both and build a more complex and efficient model, thereby further improving the accuracy and efficiency of the feature importance sorting.

[0058] Through the above, the relevance between other resource types and the key resource type is determined by means of the resource consumption matrix and the target sorting model, and the other resource types are sorted to obtain the target sorting result, so as to understand the influence degree of different resource types on the rendering performance, identify which resource types have the greatest impact on the rendering performance, which helps to improve the performance tuning of the rendering program, effectively allocate and manage resources, ensure the priority supply of key resources, reduce bottlenecks, improve the rendering speed and quality, and thus improve the overall rendering efficiency. At the same time, the target sorting result can be used as the basis for performance monitoring and analysis, which helps to continuously track the resource usage and performance metrics, conduct long-term optimization, and improve the efficiency of resource management and the overall performance of the rendering program.

[0059] Step S208, optimize the rendering program according to the target sorting result to obtain a target rendering program.

[0060] Optionally, the execution subject of the above steps may be a background processor, or other devices with similar processing capabilities, or may also be a machine integrating at least an image acquisition device and a data processing device. Among them, the image acquisition device may include a graphics acquisition module such as a camera, and the data processing device may include terminals such as a computer and a mobile phone, but is not limited thereto.

[0061] Through the above steps, since the actual resource consumption amounts of different types of rendering programs when running in the graphics processing unit are obtained, the key resource type and other resource types are determined among the M resource types, and the resource type that is most important for the overall rendering performance of the rendering program, that is, the resource type that can best represent the rendering performance of the rendering program, is obtained. Furthermore, according to the association relationship between the key resource type and other resource types, the other resource types are sorted to obtain the sorting result of the influence degree of each other resource type on the rendering performance of the rendering program, that is, the target sorting result, and the rendering program is optimized according to the target sorting result. Thus, targeted optimization strategies can be formulated according to the characteristics of different rendering programs, the rendering quality and performance in different application scenarios can be improved, the overall optimization of the rendering program can be realized, and the overall performance of the graphics processing unit and the user experience can be improved. This solves the problem of low accuracy in optimizing the GPU rendering program relying on traditional manual experience in the related art, and improves the accuracy of optimizing the GPU rendering program.

[0062] In addition, through the innovative application of machine learning algorithms in feature importance sorting, a comprehensive and systematic solution is provided for the performance optimization of GPU rendering programs. It not only simplifies the performance analysis process, improves the model accuracy, but also intuitively shows the performance bottlenecks, provides a clear direction for optimization, and effectively improves the resource utilization efficiency and the accuracy of load characteristics.

[0063] As an alternative implementation, a target ranking model is constructed based on the critical resource consumption data and other resource consumption data. The relevance between each other resource type and the critical resource type is determined through the target ranking model, and the other resource types are ranked based on the relevance to obtain a target ranking result, including: determining the critical resource consumption data as the dependent variable and the other resource consumption data as the independent variable, and obtaining a first ranking result through the random forest algorithm; determining the critical resource consumption data as the dependent variable and the other resource consumption data as the independent variable, and obtaining a second ranking result through the extremely randomized trees algorithm; and obtaining the target ranking result through the first ranking result and / or the second ranking result.

[0064] Through the above content, by introducing machine learning algorithms, it is possible to more comprehensively evaluate the feature importance of GPU rendering programs, thereby enabling more effective optimization and improvement. At the same time, by integrating two weak classifiers, namely the random forest algorithm and the ExtraTrees algorithm, a more powerful and accurate model can be constructed, reducing the bias introduced by a single algorithm, effectively reducing the risk of overfitting, and significantly enhancing the generalization performance of the model, providing a solid algorithmic foundation for rendering performance optimization.

[0065] As an alternative implementation, the first ranking model determines the critical resource consumption data as the dependent variable and the other resource consumption data as the independent variable, and obtains a first ranking result through the random forest algorithm, including: obtaining the Gini importance feature values of each other resource type through the random forest algorithm; ranking each other resource type based on the Gini importance feature values of each other resource type to obtain a first ranking result.

[0066] Optionally, in the embodiments of the present application, the Gini importance feature value can be measured by calculating the contribution of each feature to the splitting of the classification node in the random forest model. Among them, the higher the Gini importance feature value, the higher the importance of the feature, that is, the greater the impact of the feature on the model decision-making.

[0067] Through the above content, the Gini importance feature values of each resource type are obtained through the random forest algorithm, so as to effectively evaluate the influence degree of each resource on predicting the target variable through the Gini importance feature values, and further identify which resource types contribute the most to the rendering performance of the model, thereby helping to make more informed decisions in the process of resource type selection and helping to improve the performance and efficiency of the model.

[0068] As an alternative implementation, obtaining the Gini importance feature values of each other resource type through the random forest algorithm includes: obtaining the Gini importance feature value of the i-th other resource type through the following formula, where i is an integer, and the i-th other resource type is any one of the other resource types:

[0069] GI(i) = (Gini_orig - Gini_split) × Split / Total

[0070] Among them, GI(i) represents the Gini importance eigenvalue of the i-th other resource type, Gini_orig is the Gini index of the resource consumption matrix, Gini_split represents the Gini index of the i-th other resource type, Split represents the minimum number of rendering programs included in each node when splitting N rendering programs based on the i-th other resource type, and Total represents N.

[0071] As an optional implementation manner, the key resource consumption data is determined as the dependent variable, the other resource consumption data is determined as the independent variable, and a second sorting result is obtained through the extreme random tree algorithm, including: obtaining the average impurity reduction value of the key resource type when each other resource type splits N rendering programs through the extreme random tree algorithm; sorting each other resource type according to the average impurity reduction value of the key resource type when each other resource type splits N rendering programs to obtain the second sorting result.

[0072] Optionally, in the embodiments of the present application, the average impurity reduction value can be a common metric for measuring the importance of a feature to the target variable, and is usually calculated based on a certain impurity metric (such as Gini impurity or information gain). Among them, the higher the average impurity reduction value, the higher the importance of the feature, that is, the greater the impact of the feature on the model decision.

[0073] Through the above content, in the process of sorting resource types through the extreme random tree algorithm, the impact of each resource type on the target variable (i.e., rendering program performance or key resource type) during node splitting is evaluated. Thus, by analyzing the contribution of each resource type to the average impurity reduction value, the resource that has the greatest impact on the rendering program performance is identified, and then resource allocation can be better performed to ensure that key resources are fully utilized, thereby improving the overall system performance stability.

[0074] As an optional implementation manner, sorting each other resource type according to the average impurity reduction value of the key resource type when each other resource type splits N rendering programs includes: sorting each other resource type through the following formula:

[0075] ImportanceScore(j) = Rank(MeanDecreaseImpurity(j))

[0076] Among them, ImportanceScore(j) represents the sorting result of the average decrease in impurity value of the key resource type when the j-th other resource type splits nodes for the N rendering programs. Here, j is an integer, and the j-th other resource type is any one of the various other resource types. MeanDecreaseImpurity(j) represents the average decrease in impurity value of the key resource type when the j-th other resource type splits nodes for the N rendering programs, and Rank represents sorting the average decrease in impurity value of the j-th other resource type.

[0077] As an optional implementation manner, obtaining the target sorting result through the first sorting result and / or the second sorting result includes: determining the first sorting result as the target sorting result; or, determining the second sorting result as the target sorting result; or, determining the weighted average result of the first sorting result and the second sorting result as the target sorting result.

[0078] Optionally, in the embodiments of the present application, the first sorting result or the second sorting result can be directly used as the final target sorting result, or the target sorting result can be jointly determined according to the first sorting result and the second sorting result. For example, the average value of the sorting values of the other resource type in the first sorting result and in the second sorting result is used as the final sorting value of the other resource type, and the target sorting result is determined according to the final sorting values of each other resource result. Here, the sorting value can be the number of sorting times of the other resource type in the sorting result. For example, the sorting times of the most forward sorting is 1, etc.; or the numerical values used for sorting in the first sorting result or the second sorting result are weighted and calculated according to the weight value to obtain the final sorted numerical value, and the target sorting result is obtained by re-sorting based on the final sorted numerical value. Here, the weight value can be preset according to the actual situation or determined according to the first sorting result and the second sorting result.

[0079] Through the above content, according to different requirements and scenarios, directly adopting a certain sorting result or obtaining a result that better meets specific requirements through a weighted method enhances the flexibility and adaptability of the sorting method. At the same time, by combining the first and second sorting results to determine the corresponding weights, the influence of multiple factors on the sorting result can be fully considered, the priority can be better determined, the interpretability of the sorting result can be enhanced, and thus the comprehensiveness and accuracy of the decision-making can be improved.

[0080] As an alternative implementation, the rendering program is optimized based on the target sorting result to obtain a target rendering program, including: determining the top S other resource types in the target sorting result as target resource types, where S is an integer greater than 0; determining the correlation between the target resource types and the key resource types; and optimizing the rendering program according to the correlation, the resource consumption data of the target resource types, and the resource consumption data of the key resource types to obtain the target rendering program.

[0081] Through the above content, by analyzing the association relationship between other resource types and key resource types, system resources can be allocated and used more effectively, thereby reducing unnecessary consumption. As a result, the optimized rendering program may reduce resource contention and bottleneck problems, improve overall efficiency, enhance the stability and responsiveness of the system, and provide a smoother and higher-quality user experience, especially in graphics-intensive applications or games.

[0082] As an alternative implementation, the main design method of this application mainly includes the following aspects:

[0083] (1) Algorithm selection and feature analysis framework:

[0084] Focus on the application of ensemble learning algorithms in feature importance ranking. The core algorithms include Random Forest and ExtraTrees algorithm (i.e., Extreme Random Trees algorithm). These two algorithms construct a more robust and accurate model by integrating multiple weak classifiers, effectively reducing the risk of overfitting and improving the generalization ability of the model.

[0085] (2) Feature dataset construction and preprocessing:

[0086] For the diverse rendering technology algorithms of the rendering program, such as ray tracing, rasterization rendering, path tracing, volume rendering, etc., each algorithm exhibits unique performance characteristics and resource consumption patterns when dealing with different scenarios and effects. Therefore, in order to comprehensively understand and optimize the performance of the rendering program, GPU load feature data for these diverse rendering technology algorithms are widely collected and systematically sorted. These data cover multiple aspects from basic pixel processing to advanced lighting models, shadow calculations, global illumination simulations, etc. At the same time, 160 features of 100 rendering programs on four mainstream GPU models, namely 2080Ti, A100, P100, and T4, are covered. Through the principal component analysis (PCA) method, the most critical feature - stream processor efficiency (sm_efficiency) on each GPU is identified as the core dependent variable for subsequent analysis.

[0087] (3) Feature importance ranking model construction:

[0088] Using two machine learning algorithms, namely the Random Forest algorithm and the ExtraTrees algorithm, with the streaming multiprocessor efficiency (sm_efficiency) as the dependent variable and the remaining 159 GPU features as independent variables, the feature importance ranking is carried out. Through training and optimization, this model can accurately evaluate the influence degree of each feature on the rendering performance and provide data support for performance optimization.

[0089] (4) Performance evaluation and result visualization:

[0090] Summarize and visualize the ranking results of the two machine learning algorithms, namely the Random Forest algorithm and the ExtraTrees algorithm, on four GPUs: 2080Ti, A100, P100, and T4, to intuitively show the contribution degree of different features in the rendering performance. The features ranked at the top are identified as the key factors with the most significant impact on performance, while the features ranked at the bottom are regarded as secondary factors, which helps to identify the key points for optimization.

[0091] (5) Performance bottleneck identification and optimization guidance:

[0092] Based on the feature importance ranking results, deeply analyze the performance bottlenecks in the rendering program and their causes. By comparing key indicators such as the computational amount and memory bandwidth requirements of different tasks, clarify the location of the bottlenecks and provide a scientific basis for the targeted optimization of the rendering algorithm and hardware configuration.

[0093] (6) Resource utilization analysis and optimization strategy:

[0094] Further analyze the utilization rate of resources such as bandwidth, memory, and storage by the rendering program, and evaluate the rationality of the existing resource configuration. Combining the feature ranking results, propose corresponding resource management and optimization strategies to improve the overall resource utilization efficiency and reduce the operation cost.

[0095] (7) Verification of load modeling accuracy and load feature analysis:

[0096] Using the feature importance ranking of two machine learning algorithms, namely the random forest algorithm and the ExtraTrees algorithm, the modeling accuracy of the streaming processor efficiency is analyzed on four GPUs respectively. During the verification process for the load characteristic analysis, the present application separately selected 40 GPU rendering applications from the open source library for experimental evaluation. Using the NVProf feature extraction tool, in-depth feature analysis was performed on these 40 GPU applications on the P100 device. Subsequently, the research ranked the importance of these features through the random forest algorithm and the ExtraTrees machine learning algorithm, aiming to compare the differences in GPU load characteristics shown in this ranking result with the previously selected characteristics. This step aims to verify whether the accuracy and effectiveness of the finally selected 10 GPU rendering application load characteristics in the conclusion of the present application are verified in widely representing GPU rendering applications.

[0097] As an optional implementation Figure 4 For the architecture diagram of the method for optimizing the rendering program according to the embodiment of the present application, as Figure 4 shown, the specific implementation process is as follows:

[0098] (1) Algorithm selection and construction of the feature analysis framework:

[0099] Prepare the framework of the ensemble learning algorithm, including the random forest algorithm and the ExtraTrees algorithm, analyze the characteristics of the two algorithms, and determine their application potential in the feature importance ranking.

[0100] The present application regards the 160 GPU features included in each GPU rendering program as a vector. The calculation formula is as follows:

[0101] Vector = {a1,…,a k ,…,a M}

[0102] Among them, the variables a1 to a M represent 160 GPU features (i.e., resource types), Vector represents a GPU rendering program of the present application, each GPU rendering program of the present application is represented by 160 GPU features, and the present application can regard a GPU program as a row vector. 100 GPU program vectors are a 100-row and 160-column matrix.

[0103] As Figure 3As shown, at this time N is equal to 100 and M is equal to 160. This application conducts importance ranking modeling analysis through machine learning methods. This application uses the principal component analysis method to find the GPU features that have the greatest impact on the GPU rendering program as the dependent variable (i.e., the key resource type), and the other 159 GPU features as independent variables (i.e., other resource types), and finally obtains the GPU features that have a greater impact on the GPU rendering program and sorts them according to the importance of the impact. The calculation formula is as follows:

[0104] IMPv = {a1, a2…, a p}

[0105] Among them, in the independent variables a M and a p , p is less than M. This application uses the principal component analysis method to find the GPU features that have the greatest impact on the GPU rendering program as the dependent variable, and the other 159 GPU features as independent variables, and conducts random forest algorithm and ExtraTrees machine learning importance ranking.

[0106] The random forest algorithm is an ensemble model composed of multiple decision trees. During the training process of each decision tree, the random forest constructs each tree by randomly selecting a subset of features and a subset of samples. The calculation of the feature importance score is based on the aggregation of the feature importance of each decision tree in the random forest. In this application, the calculation of the random forest feature importance score is based on the Gini importance feature index (i.e., the Gini importance feature value). The calculation formula is as follows:

[0107] GI(i) = (Gini_orig - Gini_split) × Split / Total

[0108] Among them, GI(i) represents the Gini importance feature (i.e., the Gini importance feature value of the i-th other resource type), Gini_orig represents the Gini index of the original node (i.e., the Gini index of the resource consumption matrix), Gini_split represents the Gini index of the split node (i.e., the Gini index of the i-th other resource type), Split represents the number of samples in the split node (i.e., the minimum number of rendering programs included in each node when splitting the N rendering programs based on the i-th other resource type), and Total represents the total number of samples (i.e., N). The Gini index is an indicator to measure the purity of a node, which measures the probability that two randomly selected samples from a dataset belong to different classes. For each node, calculate the change in the Gini index, that is, the degree to which the Gini index of the entire model decreases after splitting the node. The feature importance score can be measured by calculating the difference between the Gini index of the split node and the original node.

[0109] The ExtraTrees algorithm, as an ensemble learning algorithm based on decision trees, introduces randomness elements during the construction of decision trees. Specific practices include randomly selecting subsets of features and randomly determining split points to enhance the diversity of the model. The importance of features in ExtraTrees is calculated by measuring the reduction in average impurity. The reduction in impurity refers to the amount by which the impurity of the model decreases after node splitting compared to before splitting. For each feature j, its average impurity reduction can be calculated. The feature importance score is based on the calculation result of the average impurity reduction. By sorting the average impurity reduction of features, the feature importance score can be obtained. The calculation formula is as follows:

[0110] ImportanceScore(j) = Rank(MeanDecreaseImpurity(j))

[0111] Among them, ImportanceScore(j) represents the importance score of feature j (i.e., the sorting result of the average impurity reduction value of the key resource type when splitting nodes for the N rendering programs), Rank represents sorting the average impurity reduction (i.e., sorting the average impurity reduction value of the j-th other resource type), and MeanDecreaseImpurity(j) represents the average impurity reduction (i.e., the average impurity reduction value of the key resource type when splitting nodes for the N rendering programs by the j-th other resource type).

[0112] (2) Construction of the feature importance ranking model:

[0113] Using the random forest algorithm and the ExtraTrees machine learning algorithm, with the stream processor efficiency as the dependent variable and the remaining 159 features as independent variables, a feature importance ranking model is constructed. Through training and optimization, the impact of each feature on the rendering performance is accurately evaluated.

[0114] Table 1 Comparative analysis of ensemble learning algorithms

[0115]

[0116]

[0117] As shown in Table 1, the Random Forest algorithm and the ExtraTree algorithm belong to the ensemble learning method. Ensemble learning improves the performance and robustness of the overall model by integrating the prediction results of multiple models. Although these algorithms all adopt the idea of ensemble learning, there are some differences in their specific implementations and strategies to adapt to different problems and scenarios. For example, the Random Forest algorithm integrates multiple decision trees and improves diversity through bootstrap sampling and introducing feature randomness. The ExtraTrees algorithm introduces additional randomness and constructs decision trees through random features and random thresholds. There are some differences and connections between these two machine learning methods in terms of feature importance ranking. First, regarding the differences in parallelism and sequentiality during the training process, the Random Forest algorithm and the ExtraTrees algorithm build multiple decision trees in parallel, which improves the training speed and is especially suitable for large-scale datasets. This parallelism can fully utilize multi-core processors and accelerate the model construction process. Second, the differences in model interpretability are also worth delving into. The Random Forest improves the model performance by averaging or voting multiple decision trees. Although this integration method is effective, it reduces the interpretability of the overall model.

[0118] In addition, overfitting handling is also an important consideration in these methods. During the training process, the Random Forest and the ExtraTrees introduce more randomness, making the model more generalizable to the training data and relatively more resistant to overfitting. Finally, the differences in scalability and computational resource utilization are also worth in-depth study. The Random Forest has good scalability due to its characteristic of building multiple decision trees in parallel and is suitable for processing large-scale datasets. Generally speaking, there are rich differences between these two machine learning methods in terms of parallelism, interpretability, overfitting handling, and scalability. This makes it crucial to select the appropriate method according to the task requirements and data situation in practical applications. In-depth analysis of these differences in details helps to understand these machine learning models more comprehensively and deeply and provides more targeted suggestions and optimization strategies for solving practical problems.

[0119] The requirement of this application for the modeling accuracy of the GPU rendering program load characteristics is above 85%. The ratio of the absolute value of the difference between the predicted value of the dependent variable and the actual value of the dependent variable to the actual value of the dependent variable represents the prediction error. The load modeling accuracy refers to the prediction accuracy, that is, 100% minus the percentage of the prediction error. The calculation of the load modeling accuracy is as follows:

[0120]

[0121] where, sm_eff real is the actual value of the stream processor efficiency sm_efficiency feature, sm_eff predicateis the predicted value of the stream processor efficiency sm_efficiency feature, and Model_Precision represents the modeling accuracy of the GPU rendering program load characteristics. This application uses machine learning feature importance ranking to analyze GPU rendering programs and non-rendering programs. With the stream processor efficiency feature as the dependent variable, a total of 159 GPU features are used as independent variables. The analysis methods adopted include two machine learning methods, the random forest algorithm and the ExtraTrees algorithm, to rank the importance of the features of the GPU rendering program.

[0122] (3) Performance evaluation and result visualization:

[0123] Summarize and visualize the ranking results of the random forest algorithm and the ExtraTrees algorithm on four GPUs (i.e., the first ranking result and the second ranking result) to intuitively display the contribution degree of different features in the rendering performance, which is convenient for identifying the optimization focus. The feature importance ranking of the ensemble learning algorithm is of great significance for the interpretation and optimization of the model performance. First of all, both the random forest algorithm and the ExtraTrees algorithm play a key role in the feature importance ranking. By combining multiple weak classifiers or models, these two algorithms form a more robust and accurate overall model, reducing the risk of overfitting and improving the generalization performance. They perform well in dealing with classification, regression, and ranking problems and are widely used in various industries. The rendering program usually consists of multiple tasks or threads, such as rendering threads, physical simulation threads, light calculation threads, etc. By ranking the load characteristics, the relative importance and priority of each task can be understood, so as to better perform task scheduling and parallel processing and make the most of the computing resources. At the same time, the load feature ranking of the rendering program can help identify potential performance bottlenecks and the reasons for the bottlenecks. By understanding the characteristics such as the computational volume and memory bandwidth requirements of each task, the location of the bottleneck can be determined. For example, if a certain task has a very large computational volume but a low demand for memory bandwidth, then the bottleneck is the computing resource rather than the memory bandwidth. By understanding the load characteristics, the rendering algorithm and hardware configuration can be optimized and adjusted targeted to improve the performance.

[0124] In addition, the requirements for resources such as the width of load characteristics, memory, and storage can be evaluated. The sorting of resource usage can also help analyze the utilization rate of various resources by the rendering program. By understanding the situation of the rendering program and performing corresponding resource management and optimization, the overall resource utilization rate can be improved. In this application, principal component analysis is used for 160 GPU load characteristics of 100 rendering programs on four GPUs: 2080Ti, A100, P100, and T4. It is found that the feature with the largest comprehensive principal component score on the four GPUs is the stream processor efficiency sm_efficiency. Using it as the dependent variable and the remaining 159 GPU characteristics as independent variables, a random forest importance ranking is performed. The random forest importance ranking results (i.e., the first ranking results) of 100 GPU rendering programs on the four GPUs: NVIDIA 2080Ti, A100, P100, and T4 are as Figure 5 shown. Features with higher rankings indicate greater importance in different rendering programs and have a more significant impact on performance. In contrast, features with lower rankings have a smaller impact on performance and can even be ignored. By deeply analyzing the underlying principles, this application can further understand the relationship between rendering performance and various metrics. These metrics include thread execution efficiency, the number of executed instruction types, GPU utilization rate, memory occupancy, etc. By sorting the feature importance, it can be determined which metrics contribute more to performance.

[0125] For the random forest algorithm, through the integration of multiple weak classifiers, it can effectively improve the accuracy and robustness of the overall model. It performs well in reducing overfitting. Therefore, the feature importance ranking can reveal the contribution degree of each feature to the performance of the overall model. The ExtraTrees algorithm further randomly selects features on the basis of random partitioning, improves the diversity of the model by increasing randomness, and reduces variance. This makes the ExtraTrees algorithm effective in dealing with high-dimensional data sets and has the advantages of fast training and small memory occupancy.

[0126] Figure 5The random forest feature importance ranking results of 100 GPU rendering programs on four GPUs: 2080Ti, A100, P100, and T4. The dependent variable, stream processor efficiency, measures the proportion of time that at least one warp is active on a particular multiprocessor. This metric reflects the utilization of the multiprocessor, i.e., how much time the multiprocessor spends on executing valid computational tasks within a given time. GPU occupancy measures the ratio between the average number of active warps per active cycle and the maximum number of warps supported by the multiprocessor. This metric reflects the parallelism of the multiprocessor, i.e., the number of active warps that the multiprocessor can execute simultaneously in each cycle. There is a mutual relationship between these two metrics. A high stream processor efficiency means that at least one warp is active on the multiprocessor for most of the time, indicating a high utilization of the multiprocessor. A high GPU occupancy means that the multiprocessor can execute more active warps simultaneously in each cycle, i.e., has a higher parallelism. The effects of these two metrics are interrelated. A higher stream processor efficiency leads to a higher GPU occupancy. When the utilization of the multiprocessor is high, it has more opportunities to execute more active warps simultaneously, thus increasing the parallelism of the multiprocessor. Vice versa, a lower sm_efficiency leads to a lower GPU occupancy. When the utilization of the multiprocessor is low, the number of active warps it executes simultaneously in each cycle is less, thus reducing the parallelism of the multiprocessor. Therefore, the relationship between these two metrics can be summarized as: high utilization is usually accompanied by high parallelism, while low utilization leads to low parallelism. This is because when the utilization of the multiprocessor is high, it can make more full use of the available computing resources and execute more active warps simultaneously, thus increasing the parallelism.

[0127] Table 2 Comprehensive Ranking of Random Forest Feature Importance

[0128]

[0129] Table 2 shows the comprehensive ranking of feature importance for 100 GPU rendering programs, Rodinia and Parboil programs, and the Octane renderer on four GPUs: 2080Ti, A100, P100, and T4. In this application, the stream processor utilization sm_efficiecny feature is used as the dependent variable, and the remaining 159 features are used as independent variables using the random forest algorithm. Finally, 10 features with relatively high comprehensive ranking scores are selected. For calculating the scores of the finally selected 10 features in the comprehensive ranking (i.e., the target ranking result), this application first considers the feature frequency. That is, if a feature appears in the rankings of all four GPUs, it indicates that it can be used as a performance metric for the rendering program. Secondly, the relative feature ranking score is considered. That is, if two features appear in four GPUs or three GPUs, the one with the higher relative feature ranking score ranks higher in the comprehensive ranking. Finally, the average feature score is considered. That is, if two features appear in two GPUs or three GPUs and their relative feature rankings are close, the one with the higher average feature score ranks higher in the comprehensive ranking. The random forest importance ranking can help optimize the rendering process. Researchers can focus on those metrics that have an important impact on performance and optimize the corresponding algorithms or resource allocation strategies to improve the overall rendering performance.

[0130] The global store requests global_store_requests is a metric that measures the total number of global store requests on a multiprocessor, excluding atomic requests. First of all, rendering programs usually involve a large amount of computation and data transfer, which involves rendering image data and writing it to global memory. During the rendering process, multiple thread blocks read and write data in global memory simultaneously. This results in a large number of global store requests. When multiple thread blocks access global memory simultaneously, conflicts and race conditions occur. This leads to resource contention, latency, and reduced efficiency. Global store conflicts occur when multiple thread blocks attempt to write to the same global memory location simultaneously. To avoid such conflicts, GPUs usually use some mechanisms to resolve race conditions, such as access control, synchronization, and scheduling policies. However, these mechanisms also introduce a certain amount of overhead, reducing the performance of the GPU. At this time, the value of the stream processor efficiency has a certain impact on the global store request feature. A higher stream processor efficiency indicates that the computing resources on the multiprocessor are more fully utilized, and more warps are in the active state, which will increase the number of global store requests. Because active warps perform read and write operations simultaneously, increasing the frequency of global store requests.

[0131] However, it should be noted that an increase in global_store_requests does not necessarily mean a deterioration in performance. In some cases, the increased global store requests coincide with parallel computing and actually improve the performance of the rendering program. For example, if multiple thread blocks can read and write to different global memory locations in parallel without causing conflicts and race conditions, then the increased global store requests can improve the efficiency of the rendering program. Therefore, there is a complex relationship between the stream processor efficiency characteristic and the global store request characteristic in the rendering program. A higher active warp ratio increases the number of global store requests, but the specific impact also depends on the characteristics of the rendering program, the mode of parallel computing, and the optimization strategy of the GPU architecture. To maximize the performance of the GPU, developers need to consider these metrics comprehensively and use appropriate optimization techniques to improve the performance and efficiency of the rendering program.

[0132] Figure 6 The feature importance ranking of 100 GPU rendering programs on four GPUs, namely 2080Ti, A100, P100, and T4, using the ExtraTrees algorithm is presented. After the comparative evaluation and analysis of the load characteristics, this application will focus on two GPU load characteristics: GPU occupancy and the inst_replay_overhead characteristic. These two characteristics play an important role in the ExtraTrees ranking, indicating their key role in the performance of GPU rendering programs. In GPU rendering, a high GPU occupancy is usually associated with higher stream processor efficiency and computing core utilization. This shows that the GPU can effectively utilize computing resources when executing rendering tasks, improving the overall performance. In a specific GPU software and hardware environment, if the GPU has a large number of computing cores and high parallelism, then achieving a high GPU occupancy will become an optimization goal. For example, when rendering complex scenes or performing real-time lighting calculations, a high GPU occupancy can ensure that more computing cores execute tasks simultaneously, thereby improving parallelism and overall performance.

[0133] The inst_replay_overhead feature is related to the GPU architecture and instruction scheduling. In GPU rendering programs, the efficient execution of instructions is crucial for performance. Instruction replay causes performance loss due to data dependencies or other factors, which leads to the re-execution of instructions. In some cases, optimizing the instruction replay overhead involves adjusting the code structure, optimizing the data access pattern, etc. Specifically, if the GPU architecture adopts a pipelined instruction execution model, reducing instruction replay involves adjusting the instruction scheduling strategy to enable instructions to better utilize the pipeline and reduce the latency caused by instruction dependencies. For example, in the case of a smaller GPU memory, reducing the instruction replay overhead is more important because it can reduce the competition for memory bandwidth and improve the overall execution efficiency. For the rendering of large-scale scenes, such as complex scenes in virtual reality applications, optimizing the instruction replay overhead helps to improve the stability and responsiveness of the GPU rendering program. For example, in real-time lighting calculation scenarios, the GPU occupancy rate becomes a key factor in performance because lighting calculation usually involves a large number of parallel floating-point operations. On the other hand, for rendering tasks that require frequent modification of shader programs or processing of a large amount of dynamic data, the overhead of instruction replay is more sensitive because these operations lead to an increase in instruction replay, thereby affecting performance.

[0134] Table 3 Comprehensive Sorting of ExtraTrees Feature Importance

[0135]

[0136]

[0137] Table 3 is the comprehensive sorting of ExtraTrees feature importance (i.e., the second sorting result) with the sm_efficiecny feature as the dependent variable and the remaining 159 features as independent variables for 100 GPU rendering programs, Rodinia and Parboil programs, and Octane renderer on four GPUs: 2080Ti, A100, P100, and T4. The 10 features with relatively large comprehensive sorting scores are finally selected. This application will conduct a comparative evaluation and analysis of the results of 100 GPU rendering programs, Rodinia and Parboil programs, and Octane renderer one by one. Among all rendering programs, the machine learning comprehensive sorting score of the GPU occupancy rate is relatively large. The achieved occupancy rate reflects the ratio of active warps to the maximum number of warps that can be accommodated on the GPU and is a key indicator for evaluating the GPU resource utilization efficiency. In GPU rendering, a high achieved occupancy rate usually means higher parallel processing capabilities and less resource idling, thus bringing higher rendering performance. Therefore, optimizing strategies such as thread scheduling, memory access patterns, and computational load distribution to increase the achieved occupancy rate is an important way to improve GPU rendering performance.

[0138] Secondly, the instruction replay overhead inst_replay_overhead feature has a relatively high comprehensive ranking score in GPU rendering programs. Instruction replay overhead reflects the proportion of instruction replay caused by dependencies, resource contention, or improper instruction scheduling. Reducing instruction replay overhead can be achieved by optimizing the instruction stream, reducing dependencies, and improving resource utilization, which is of great significance for enhancing rendering performance. In the Rodinia and Parboil programs, the constant memory dependency stall stall_constant_memory_dependency and texture unit utilization tex_sm_utilization are identified as key features. Texture memory utilization reflects the access efficiency of the GPU to texture data during rendering, while constant memory dependency stall reveals the performance bottleneck caused by constant memory access. Optimizing these features, such as improving the texture mapping algorithm, optimizing the constant memory access pattern, and reducing memory dependency stalls, can significantly enhance the rendering performance of the Rodinia and Parboil programs.

[0139] For the Octane renderer, the L2 cache utilization l2_utilization and texture unit utilization tex_sm_utilization are also identified as important features. This indicates that optimizing the access efficiency of the L2 cache and the utilization of texture memory is crucial for enhancing performance during Octane rendering. By strategies such as improving data layout, increasing cache hit rate, and optimizing texture access patterns, the performance of the Octane renderer can be further improved. Take the volume rendering program as an example. This program involves a large number of operations such as data reading, texture mapping, and lighting calculation. When optimizing the volume rendering program, developers can focus on the key features mentioned above and make corresponding optimizations in combination with the specific characteristics of the program. For example, reducing texture access overhead by optimizing the texture mapping algorithm and increasing texture memory utilization; to reduce memory access latency, methods such as optimizing data layout and increasing L2 cache utilization can be adopted; reducing instruction replay overhead by optimizing the instruction stream and reducing dependencies. These optimization measures can significantly enhance the performance and image quality of the volume rendering program.

[0140] The significance of feature importance ranking lies in deeply understanding the contributions of different algorithms to model performance, helping with feature selection and engineering, and improving the efficiency and generalization performance of the model. By identifying key features, it can be better applied to specific business scenarios and optimize model design. Feature importance ranking also helps detect abnormal and redundant features, improving model interpretability and user trust. Generally speaking, deeply analyzing the feature importance ranking of ensemble learning algorithms is an important part of optimizing the model and improving prediction results.

[0141] (4) Performance bottleneck identification and optimization guidance:

[0142] By synthesizing the feature importance rankings of the above-mentioned Random Forest algorithm and ExtraTrees algorithm, a final feature importance ranking (i.e., the target ranking result) can be obtained. This ranking takes into account the evaluation results of different algorithms, thus more comprehensively evaluating the importance of features. Table 4 shows the final ranking of the feature importance of the rendering program. For each feature, the feature frequency, the relative ranking result, and the average ranking of the feature in different rankings can be considered. Next, a comprehensive consideration will be made in combination with the above selection criteria, and a ranking will be made accordingly to obtain the final comprehensive ranking result of feature importance. The advantage of this comprehensive ranking is that it can reduce the bias of a single algorithm and more objectively evaluate the importance of features.

[0143] Table 4 Final ranking of the feature importance of the rendering program

[0144]

[0145] As the core component of rendering, the performance and capabilities of the GPU have a significant impact on the execution speed and quality of the rendering program. The achieved_occupancy feature of GPU occupancy and the global_store_requests feature rank first and second respectively in the comprehensive proportion of feature importance rankings of the random forest algorithm and the ExtraTrees machine learning algorithm, which are 26.77% and 7.80% respectively. GPU occupancy represents the utilization degree of warps by the GPU core when executing computational tasks. A higher occupancy means that the GPU core can make more full use of warps and maintain a higher number of active warps in each computational cycle. This can improve the parallel computing ability of the GPU, thereby accelerating the execution of the rendering program. During the rendering process, a higher warp utilization rate can better handle rendering tasks such as geometric calculations and lighting calculations, improving the rendering speed and effect. The global store request is an indicator that measures the total number of global store requests on a multiprocessor, excluding atomic requests. In the GPU rendering program, a large amount of data needs to be loaded from the global memory to the processor for calculation and then the results are written back to the global memory. By monitoring the global store request indicator, the access pattern of the program to the global memory can be evaluated, and corresponding memory access optimizations can be carried out. For example, if it is found that there are too many global store requests, technologies such as shared memory or local memory can be considered to reduce the number of accesses to the global memory, thereby improving the memory access efficiency. In addition, the present application has conducted a comparative analysis of the Octane renderer, Rodinia, and Parboil programs in the above-mentioned machine learning feature importance rankings. To sum up, the GPU occupancy also has a relatively large comprehensive importance proportion in the random forest algorithm and the ExtraTrees machine learning algorithm in the Octane renderer, Rodinia, and Parboil programs because it is directly related to the hardware utilization rate of the GPU, latency hiding ability, instruction-level parallelism, and memory access efficiency. These aspects are all key factors for evaluating and optimizing GPU performance.

[0146] The inst_integer feature of integer instructions measures the number of integer instructions executed by non-conditional execution threads. This metric can provide information about the integer computation workload in a program. For ray tracing programs, it reflects the importance and execution of integer computations in the ray tracing algorithm. In ray tracing programs, integer instructions are typically used to perform some basic mathematical operations, index calculations, distance calculations, and logical operations for auxiliary algorithms. For example, ray-object intersection tests, pixel shading calculations, ray reflection and refraction calculations, etc. all involve the execution of integer instructions. Since ray tracing is a computationally intensive task that involves a large number of geometric calculations and vector operations, the importance of integer instructions in ray tracing programs cannot be ignored. By analyzing the inst_integer metric, this application can understand the quantity and workload of integer computations in the program, and then evaluate the complexity and execution efficiency of the program. In ray tracing programs, different types of integer instructions can be considered, such as addition, subtraction, multiplication, bitwise operations, etc. Analyzing the execution counts of different types of integer instructions can help understand the importance and load of various computational operations in the program. Analyzing the distribution of integer instructions in the program can help identify computational bottlenecks and hot spots. If integer instructions are overly concentrated on certain specific computations or operations, these parts need to be optimized more deeply to improve the overall performance of the program. By optimizing the selection of algorithms and data structures, the number of integer instructions can be reduced and the efficiency of the program can be improved. For example, using bitwise operations instead of multiplication and division operations, and using lookup tables to reduce redundant calculations and other optimization techniques can reduce the execution count of integer instructions. Ray tracing programs can usually perform parallel computations, using multiple processing units of the GPU to execute computational tasks simultaneously. Analyzing the distribution of integer instructions among different threads and thread blocks can help optimize load balancing, concurrency control, and data distribution strategies to improve the parallelism and overall performance of the program.

[0147] When it comes to ray tracing programs, it is very important to understand and analyze the relationship between the two metrics of stream processor efficiency and inst_integer. In-depth analysis of these metrics can help this application understand the parallelism of the program, the utilization of computing resources, and the execution of integer instructions, thereby optimizing the performance of the ray tracing program. First of all, stream processor efficiency is a metric that measures the utilization of multi-processors. It represents the proportion of time during which at least one warp (a group of parallel-executing threads) is active on a specific multi-processor. For ray tracing programs, a higher proportion of active warps means that the computing resources of the multi-processor are fully utilized, enabling better parallel computing. Ray tracing programs involve a large number of ray calculations and intersection tests, including vector and matrix operations, distance calculations, index calculations, etc. These calculations usually require the use of integer instructions. Therefore, an increase in active warps can lead to more integer instructions being executed by unconditional execution threads. This is because active warps can execute their respective computing tasks simultaneously, including integer operations.

[0148] However, this application needs to note that the number of integer instructions is not the only factor determining the performance of ray tracing programs. Other factors such as data access patterns, memory bandwidth, algorithm selection, etc. also have an impact on performance. Therefore, when optimizing ray tracing programs, this application needs to consider multiple factors comprehensively and adopt appropriate optimization strategies. Increasing the proportion of active warps can be achieved in various ways. One method is to ensure that there are sufficient active warps on the multi-processor by optimizing the thread block size and grid configuration. Selecting an appropriate thread block size and grid configuration can ensure the full utilization of the computing resources of the multi-processor, thereby increasing the proportion of active warps. In addition, optimizing the algorithm and data structure selection can also affect the number of integer instructions. By choosing more efficient algorithms and data structures, the number of integer operations involved in the calculation can be reduced, thereby reducing the number of executed integer instructions and improving the performance of the program. Another point to note is that in order to improve the performance of ray tracing programs, memory access patterns and data locality also need to be considered. Reasonable use of cache technologies such as caches and shared memory can reduce memory access latency and improve the execution efficiency of the program. In addition, use appropriate data layouts and access patterns to minimize the number of global memory accesses. At the same time, reducing the execution frequency of integer instructions can also effectively improve the performance of the program.

[0149] The number of single-precision floating-point instructions executed, inst_fp_32, measures the number of single-precision floating-point instructions executed by non-conditional execution threads, including arithmetic operations, comparison operations, etc. This metric reflects the importance and workload of single-precision floating-point calculations in a program. Take the Game of Life CUDA game as an example. The Game of Life is a classic cellular automaton game that shows interesting life evolution and pattern generation by iteratively updating the cell states. In the CUDA game, GPU parallel computing is used to accelerate the calculation process of the Game of Life. The Game of Life is a highly parallel algorithm, and the update of each cell's state is independent of each other. When at least one warp is active on a specific multiprocessor, it means that the computing resources of that multiprocessor are fully utilized, which helps to execute more cell state update operations. Therefore, a higher percentage of active warps can improve the parallelism and computing efficiency of the Game of Life CUDA game. In the Game of Life, the update of cell states usually involves the calculation and comparison of neighboring cells. These calculation processes contain a certain number of single-precision floating-point instructions, such as addition, subtraction, multiplication, etc. Therefore, the number of floating-point instructions reflects the workload and computational complexity of cell state updates in the game. There is a certain correlation between the percentage of active warps and the number of floating-point instructions. When the percentage of active warps is relatively high, it means that the computing resources on the multiprocessor are better utilized, and more floating-point calculation operations can be executed. Therefore, a higher percentage of active warps can indirectly lead to a higher number of floating-point instructions. When optimizing the Game of Life CUDA game, it can be considered to improve the percentage of active warps and make full use of the computing resources of the multiprocessor by means of reasonable task parallelization and scheduling strategies, as well as optimizing the division of thread blocks and grids. By using appropriate floating-point calculation techniques, reducing redundant calculations, using low-precision calculations, etc., the floating-point calculation process can be optimized, the number of floating-point instructions can be reduced, and the performance can be improved. The Game of Life involves a large amount of reading and writing of cell state data. By optimizing the memory access pattern and reducing data transfer and access latency, the overall performance of the program can be improved.

[0150] The warp non - conditional execution efficiency warp_nonpred_execution_efficiency feature, that is, the ratio of "the average number of active threads executing non - conditional instructions in each warp to the maximum number of threads supported by each warp" has a significant impact on the performance and efficiency of the algorithm. A high active thread ratio indicates that more threads are performing computational activities simultaneously, thereby improving the efficiency of parallel computing. In the progressive photon mapping algorithm, each thread is responsible for calculating a ray path and tracing and calculating the interaction between the ray and the scene. If the number of active threads in each warp is small, it means that more computing resources are idle and the parallel computing power of the GPU cannot be fully utilized. On the contrary, a higher active thread ratio can make the number of threads in each warp approach the maximum value, maximizing the use of the GPU's computing resources and improving the computing efficiency. In addition, a higher thread activity also helps to improve thread utilization. In the progressive photon mapping algorithm, the main tasks of the threads are to trace and calculate the ray path and intersect with the objects in the scene. If only a few threads in each warp are active, the other threads will be idle, wasting computing resources. By increasing the number of active threads, this application can more fully exploit the computing potential of the GPU, thereby enhancing thread utilization and reducing the consumption of computing resources.

[0151] For the progressive photon mapping algorithm, higher thread activity can also improve memory access efficiency. During the photon mapping process, a large amount of photon data needs to be read from memory for calculation, and memory access is a relatively slow operation. When more threads are active simultaneously, the memory bandwidth can be better utilized, the latency of memory access can be reduced, and the data reading speed can be increased. This is crucial for the performance of the progressive photon mapping algorithm, which can reduce waiting time and improve the calculation speed. To increase thread activity, some optimization strategies can be adopted. First, reasonably organize the sizes of thread blocks and thread grids so that the number of threads in each warp is close to the maximum value, maximizing the use of the GPU's computing resources. Second, optimize the loop structure and data access patterns in the algorithm to reduce the idle time of threads and the latency of memory access, improving thread activity and memory access efficiency. In addition, reasonably set the parameters and configuration options of the algorithm to maximize thread activity and improve the overall performance of the algorithm. For example, assume there is a rendering program for the progressive photon mapping algorithm, where each warp supports 32 threads. If during runtime, only 16 threads are active in each warp, then the ratio of thread activity is 16 / 32 = 0.5. This means that only half of the threads are executing non-conditional instructions, and the other half of the threads are idle. To increase thread activity, the implementation method of the algorithm can be adjusted, the data structure and algorithm logic can be optimized, and the sizes of thread blocks and thread grids can be reasonably set so that the number of threads in each warp is close to 32, thereby increasing thread activity and improving the performance and efficiency of the algorithm. Through in-depth analysis and optimization of the thread activity metrics, the progressive photon mapping algorithm can more efficiently utilize the parallel computing power of the GPU, thereby enhancing the computing efficiency, thread utilization rate, and memory access efficiency, thus accelerating the rendering process and obtaining a more rapid rendering result. <http: / / www.example.com /

[0152] In summary, the impacts of these GPU features on the rendering program are multi-faceted. High GPU occupancy and warp utilization can improve the parallel computing power and rendering speed. Higher single-precision floating-point operation and integer instruction counts improve the accuracy and realism of the rendering program. However, a high number of global memory requests can lead to memory access bottlenecks and latency, which need to be optimized. Therefore, during the optimization process of the rendering program, these features need to be comprehensively considered and appropriate adjustments and trade-offs need to be made to obtain the best rendering performance and quality.

[0153] (5) Resource Utilization Analysis and Optimization Strategies:

[0154] It should be noted that the content in seems to be an incorrect or incomplete link in the original. I have translated it as it is, but it might need to be corrected in the actual context.Analyze the utilization rate of resources such as bandwidth, memory, and storage by the rendering program. Combining the feature sorting results, propose resource management and optimization strategies aimed at improving the overall resource utilization efficiency. Table 5 shows the relationship between the number of GPU rendering programs and the average modeling accuracy. This application's research found that as the number of GPU rendering programs increases, the average modeling accuracy of the machine learning feature importance sorting algorithm also increases. When the number of GPU rendering programs is 80, the average modeling accuracy is 96.78%. When the number of GPU rendering programs is 100, the average modeling accuracy is 97.07%. When the number of GPU rendering programs is 140, the average modeling accuracy is 97.38%. Increasing the number of training samples can help the model better learn the true distribution and features of the data, thereby improving the modeling accuracy. Each additional rendering program provides more information, enabling the model to better capture the features of the data. When the number of GPU rendering programs reaches a certain level, increasing the sample number further will not significantly improve the modeling accuracy. This is because the model has already learned most of the features of the data, and further increasing the sample number will only introduce noise or duplicate information without bringing significant gains. Therefore, when the number of GPU rendering programs is 100, the average modeling accuracy has reached 97.07%. At this order of magnitude, the model can already capture the features of the data well. Further increasing the sample number will not significantly improve the model accuracy and will increase the computational cost and time. Therefore, choosing 100 GPU rendering programs as the number of training samples in this application is a reasonable choice.

[0155] Table 5 Relationship between the number of GPU rendering programs and the average modeling accuracy

[0156]

[0157] This application uses the random forest feature importance sorting on four GPUs to analyze the modeling accuracy of the streaming processor efficiency. Modeling accuracy is an important indicator for evaluating the performance of machine learning models. It measures the consistency between the predicted values and the true values of the model. The modeling accuracy is equal to 100% minus the relative error, and the relative error is equal to the absolute value of the difference between the predicted value and the true value divided by the true value. A high-precision model can more accurately capture the internal laws and relationships in the dataset, providing reliable predictions and decision-making support for practical applications. Such as Figure 7As shown, there are six groups of true and predicted values of the efficiency of stream processors on each GPU. The average modeling accuracies of 2080Ti, A100, P100, and T4 are 98.52%, 94.19%, 96.97%, and 96.79% respectively. Therefore, the average modeling accuracy of the random forest feature importance ranking is 96.62%. As shown in Table 6, the average modeling accuracies of the random forest algorithm and the ExtraTrees machine learning algorithm on the four GPUs are both higher than 96.62%. This indicates that these models have high accuracy and reliability when modeling the GPU rendering program load. In the GPU rendering program load modeling, the modeling accuracy reflects the model's ability to understand different GPU rendering program loads and prediction accuracy. A high modeling accuracy means that the model can accurately capture and understand the load characteristics of the GPU rendering program and can accurately predict the predicted value of the dependent variable based on the true value of the dependent variable.

[0158] Table 6 Average Modeling Accuracy of GPU Rendering Program Load

[0159]

[0160] The number of single-precision floating-point instruction executions represents the number of single-precision floating-point multiply-accumulate operations executed by the GPU core. In rendering, single-precision floating-point operations are commonly used in computationally intensive operations such as lighting calculations, texture sampling, and geometric transformations. A higher number of operations means that the GPU core has executed more floating-point calculations, which can improve the accuracy and realism of the rendering program. However, it should be noted that excessive floating-point calculations can lead to overconsumption of computing resources, so a proper balance needs to be struck between computational accuracy and performance. The integer instruction execution number feature represents the number of integer instructions executed by the GPU core. In the rendering program, integer instructions are commonly used in operations such as geometric calculations and texture indexing. A higher number of integer instructions indicates complex geometric calculations and texture indexing operations, which are crucial for handling complex scenes and large-scale textures. Efficient execution of integer instructions can improve the performance and quality of the rendering program. The global memory request execution number feature represents the total number of global memory requests issued from the multiprocessor, excluding atomic requests. During the rendering process, global memory requests are used to write the calculation results back to global memory, including the render target buffer and texture cache, etc. A higher number of global memory requests indicates that the rendering program needs to frequently read and write global memory, which results in a memory access bottleneck and latency. Therefore, when optimizing the rendering program, it is necessary to consider reducing the number of global memory requests and optimizing the memory access pattern to improve the rendering performance.

[0161] This application mainly studies the analysis results of two feature importance ranking algorithms, namely the random forest algorithm and the ExtraTrees algorithm, for GPU rendering programs, and provides a comprehensive and feasible analysis process by combining the load performance analysis of GPU rendering programs and the essence of rendering algorithms. By introducing machine learning algorithms, this application can more comprehensively evaluate the feature importance of GPU rendering programs and reduce the bias introduced by a single algorithm. The research methods and results of this application have practical significance for the development and performance optimization of GPU rendering programs. Developers can use the feature importance ranking method proposed in this application to identify the performance bottlenecks of rendering programs and optimize them targeted. And the designers of GPUs can also better meet the needs of rendering applications and provide more efficient hardware support by analyzing the characteristics of GPU rendering programs.

[0162] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware, but in many cases, the former is a better implementation method.

[0163] In this embodiment, a device for optimizing a rendering program is also provided. This device is used to implement the above embodiments and preferred implementation manners, and those that have been described will not be repeated. As used below, the term "module" can be a combination of software and / or hardware that can achieve a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.

[0164] Figure 8 is a structural block diagram of a device for optimizing a rendering program according to an embodiment of the present application. As Figure 8 shown, the device includes: an acquisition module 802, configured to acquire resource consumption data of N rendering programs under M resource types when running in a graphics processing unit, so as to obtain a resource consumption matrix, where N and M are integers greater than 0, and the N rendering programs are rendering programs belonging to different types of graphics processing units; a determination module 804, configured to determine a key resource type and other resource types among the M resource types, and determine the resource consumption data of the key resource type as key resource consumption data and the resource consumption data of other resource types as other resource consumption data; a sorting module 806, configured to determine the relevance between each other resource type and the key resource type through a target sorting model according to the resource consumption matrix, and sort the other resource types based on the relevance to obtain a target sorting result, where the target sorting result is used to identify the influence degree of each other resource type on the rendering performance of the N rendering programs; an optimization module 808, configured to optimize the rendering program through the target sorting result to obtain a target rendering program.

[0165] Optionally, the sorting module includes: a first sorting unit configured to determine the critical resource consumption data as the dependent variable and other resource consumption data as the independent variable, and obtain a first sorting result through a random forest algorithm; a second sorting unit configured to determine the critical resource consumption data as the dependent variable and other resource consumption data as the independent variable, and obtain a second sorting result through an extreme random tree algorithm; and a first determination unit configured to obtain a target sorting result based on the first sorting result and / or the second sorting result.

[0166] Optionally, the first sorting unit is further configured to obtain the Gini importance eigenvalue of each other resource type through a random forest algorithm; and sort each other resource type according to the Gini importance eigenvalue of each other resource type to obtain a first sorting result.

[0167] Optionally, the first sorting unit obtains the Gini importance eigenvalue of the i-th other resource type through the following formula, where i is an integer, and the i-th other resource type is any one of the other resource types:

[0168] GI(i) = (Gini_orig - Gini_split) × Split / Total

[0169] where GI(i) represents the Gini importance eigenvalue of the i-th other resource type, Gini_orig is the Gini index of the resource consumption matrix, Gini_split represents the Gini index of the i-th other resource type, Split represents the minimum number of rendering programs included in each node when splitting the N rendering programs based on the i-th other resource type, and Total represents N.

[0170] Optionally, the second sorting unit is further configured to obtain the average impurity reduction value of the critical resource type when splitting the N rendering programs by each other resource type through an extreme random tree algorithm; and sort each other resource type according to the average impurity reduction value of the critical resource type when splitting the N rendering programs by each other resource type to obtain a second sorting result.

[0171] Optionally, the second sorting unit is further configured to sort each other resource type through the following formula:

[0172] ImportanceScore(j) = Rank(MeanDecreaseImpurity(j))

[0173] Among them, ImportanceScore(j) represents the sorting result of the average impurity reduction value of the key resource type when the j-th other resource type performs node splitting on the N rendering programs, where j is an integer, and the j-th other resource type is any one of the various other resource types. MeanDecreaseImpurity(j) represents the average impurity reduction value of the key resource type when the j-th other resource type performs node splitting on the N rendering programs, and Rank represents the sorting of the average impurity reduction value of the j-th other resource type.

[0174] Optionally, the first determination unit is further configured to determine the first sorting result as the target sorting result; or, determine the second sorting result as the target sorting result; or, determine the weighted average result of the first sorting result and the second sorting result as the target sorting result.

[0175] Optionally, the optimization module includes: a second determination unit, configured to determine the top S other resource types with higher sorting in the target sorting result as the target resource types, where S is an integer greater than 0; a third determination unit, configured to determine the correlation between the target resource types and the key resource types; and an optimization unit, configured to optimize the rendering program according to the correlation, the resource consumption data of the target resource types, and the resource consumption data of the key resource types to obtain a target rendering program.

[0176] For the description of the features in the embodiments corresponding to the apparatus for optimizing a rendering program, reference may be made to the relevant descriptions in the embodiments corresponding to the method for optimizing a rendering program, which will not be elaborated here one by one.

[0177] An embodiment of the present application further provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the method embodiments for optimizing a rendering program described above.

[0178] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps in any one of the method embodiments for optimizing a rendering program when running.

[0179] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: USB flash drives, read-only memories (ROMs), random access memories (RAMs), mobile hard disks, magnetic disks, or optical discs and other various media that can store computer programs.

[0180] An embodiment of the present application also provides a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, the steps in any of the method embodiments of the above optimization rendering program are implemented.

[0181] An embodiment of the present application also provides another computer program product, including a non-volatile computer-readable storage medium. The non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in any of the method embodiments of the above optimization rendering program are implemented.

[0182] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in this article can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Skilled professionals can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0183] The above has introduced in detail a method and device for an optimization rendering program provided by this application. Specific examples are used in this article to elaborate on the principle and implementation manner of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application. It should be noted that for those of ordinary skill in the art in this technical field, without departing from the principle of this application, several improvements and modifications can be made to this application, and these improvements and modifications also fall within the protection scope of the claims of this application.

Claims

1. A method for optimizing a rendering program, characterized in that, Including: Obtain the resource consumption data of N rendering programs under M resource types when running on a graphics processing unit, and obtain a resource consumption matrix, where N and M are integers greater than 0, and the N rendering programs belong to rendering programs of different types of the graphics processing unit; Determine the critical resource type and other resource types among the M resource types, and determine the resource consumption data of the critical resource type as the critical resource consumption data, and determine the resource consumption data of the other resource types as the other resource consumption data; Determine the relevance between each of the other resource types and the critical resource type through a target sorting model according to the resource consumption matrix, and sort the other resource types based on the relevance to obtain a target sorting result, where the target sorting result is used to identify the influence degree of each of the other resource types on the rendering performance of the N rendering programs; Optimize the rendering program through the target sorting result to obtain a target rendering program.

2. The method for optimizing a rendering program according to claim 1, wherein Construct a target sorting model based on the critical resource consumption data and the other resource consumption data, Determine the relevance between each of the other resource types and the critical resource type through the target sorting model, and sort the other resource types based on the relevance to obtain a target sorting result, including: Determine the critical resource consumption data as the dependent variable, and determine the other resource consumption data as Independent variables, and obtain a first sorting result through a random forest algorithm; Determine the critical resource consumption data as the dependent variable, and determine the other resource consumption data as Independent variables, and obtain a second sorting result through an extremely randomized trees algorithm; Obtain the target sorting result through the first sorting result and / or the second sorting result.

3. The method for optimizing a rendering program according to claim 2, wherein The first sorting model determines the critical resource consumption data as the dependent variable and the other resource consumption data as independent variables, and obtains a first sorting result through a random forest algorithm, including: Obtain the Gini importance feature values of each of the other resource types through the random forest algorithm; Sort each of the other resource types through the Gini importance feature values of each of the other resource types to obtain the first sorting result.

4. The method for optimizing a rendering program according to claim 3, wherein Obtain the Gini importance feature values of each of the other resource types through the random forest algorithm, including: Obtain the Gini importance feature value of the i-th other resource type through the following formula, where i is an integer, and the i-th other resource type is any one of the other resource types: GI(i) = (Gini_orig - Gini_split) × Split / Total Among them, GI(i) represents the Gini importance eigenvalue of the i-th other resource type, Gini_orig is the Gini index of the resource consumption matrix, Gini_split represents the Gini index of the i-th other resource type, Split represents the minimum number of rendering programs included in each node when splitting the N rendering programs based on the i-th other resource type, and Total represents N.

5. The method for optimizing a rendering program according to claim 2, wherein: Determining the critical resource consumption data as the dependent variable and the other resource consumption data as the independent variable, and obtaining a second sorting result through the extreme random tree algorithm, including: Obtaining the average decrease in impurity of the critical resource type when each of the other resource types splits the N rendering programs through the extreme random tree algorithm; Sorting each of the other resource types according to the average decrease in impurity of the critical resource type when each of the other resource types splits the N rendering programs, to obtain the second sorting result.

6. The method for optimizing a rendering program according to claim 5, wherein Sorting each of the other resource types according to the average decrease in impurity of the critical resource type when each of the other resource types splits the N rendering programs, including: Sorting each of the other resource types through the following formula: ImportanceScore(j) = Rank(MeanDecreaseImpurity(j)) Among them, ImportanceScore(j) represents the sorting result of the average decrease in impurity of the critical resource type when the j-th other resource type splits the N rendering programs. Here, j is an integer, the j-th other resource type is any one of the other resource types, MeanDecreaseImpurity(j) represents the average decrease in impurity of the critical resource type when the j-th other resource type splits the N rendering programs, and Rank represents sorting the average decrease in impurity of the j-th other resource type.

7. The method for optimizing a rendering program according to claim 2, wherein: Obtaining the target sorting result through the first sorting result and / or the second sorting result, including: Determining the first sorting result as the target sorting result; or, Determining the second sorting result as the target sorting result; or, Determining the weighted average result of the first sorting result and the second sorting result as the target sorting result.

8. The method for optimizing a rendering program according to claim 1, wherein: Optimizing the rendering program through the target sorting result to obtain a target rendering program, including: Determining the top S other resource types with higher sorting in the target sorting result as target resource types, where S is an integer greater than 0; Determining the correlation between the target resource types and the critical resource types; Optimize the rendering program according to the relevance, the resource consumption data of the target resource type, and the resource consumption data of the key resource type, to obtain the target rendering program.

9. An electronic device, characterized in that, It includes: A memory for storing computer programs; A processor for implementing the steps of the method for optimizing the rendering program as described in any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, wherein the computer program implements the steps of the method for optimizing the rendering program as described in any one of claims 1 to 7 when executed by a processor.