Rendering program adjusting method and device, electronic equipment and storage medium

By analyzing the resource usage of the rendering program in the target processor and identifying and optimizing the features that have the greatest impact on its performance, the problem of low performance on different processors is solved, and more efficient rendering performance is achieved.

CN120371687APending Publication Date: 2025-07-25JINAN INSPUR DATA TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510412329.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-02
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

In the prior art, the operational performance optimization efficiency of rendering programs is low, and it is difficult to maintain stability and versatility on different processors.

Method used

By obtaining the running information of the target processor, analyzing the resource usage of the rendering program, identifying the features that have the greatest impact on rendering performance, and adjusting the rendering program based on these features to optimize its operation on the target processor.

Benefits of technology

It improves the operating performance of the rendering program on the target processor, enhances the adaptability of the program and processor resources, and solves the problem of low performance optimization efficiency of the rendering program on different processors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120371687A_ABST
    Figure CN120371687A_ABST
Patent Text Reader

Abstract

The invention discloses a rendering program adjusting method and device, electronic equipment and a storage medium, and relates to the field of computers.The method comprises the steps that target running information of a target processor on an initial rendering program is obtained, the target running information is used for indicating the use condition of program running resources of the target processor in the running process of the initial rendering program; target program features of the initial rendering program are converted according to the target operation information and reference operation information, the reference operation information is used for indicating the response condition of program operation resources when the test program is operated, and the target program features are program operation features with the influence degree on rendering performance larger than the target influence degree in the initial rendering program; and adjusting the initial rendering program according to the target program features to obtain a rendering program used for running on the processor. The technical problem that in the prior art, the optimization efficiency of the operation performance of the rendering program is low is solved, and the effect of improving the optimization efficiency of the operation performance of the rendering program is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computers, and in particular, to a method and device for adjusting a rendering program, an electronic device, and a storage medium. Background Art

[0002] With the rapid development of information technology, the performance requirements of computer systems are increasing day by day. Especially in the field of graphics rendering, unprecedented challenges have been posed to the performance of processors and rendering efficiency. As an important part of computer graphics, the rendering process involves a large number of computationally intensive tasks, such as ray tracing, rasterization, etc. These tasks have extremely high dependence on the micro-architecture characteristics of the processor.

[0003] In order to improve the rendering efficiency of graphics, algorithm optimization techniques are used in related technologies to optimize the rendering program. These techniques improve the performance of the rendering program by improving the rendering algorithm or parallel computing. However, there are many challenges in the actual application of these techniques. Algorithm optimization techniques often require in-depth understanding and improvement of the rendering algorithm, which requires a high level of professional knowledge and skills. In addition, the effect of algorithm optimization techniques is often limited by specific processor architectures and test environments. That is to say, although the rendering program is optimized, when the optimized rendering program is deployed on different processors, the rendering performance may be reduced instead, and it is difficult to ensure the generality and stability in different scenarios.

[0004] Aiming at the technical problem of low optimization efficiency of the running performance of the rendering program in related technologies, no effective solution has been proposed yet. Summary of the Invention

[0005] This application provides a method for adjusting a rendering program to at least solve the technical problem of low optimization efficiency of the running performance of the rendering program in related technologies.

[0006] This application provides a method for adjusting a rendering program, including: obtaining target running information of an initial rendering program on a target processor, where the target running information is used to indicate the usage of program running resources of the target processor during the running process of the initial rendering program;

[0007] Converting target program characteristics of the initial rendering program according to the target running information and reference running information, where the reference running information is used to indicate the response of program running resources when running a test program, the test program is used to test the resource performance of the program running resources, and the target program characteristics are program running characteristics in the initial rendering program that have a greater impact on the rendering performance than the target impact degree; adjusting the initial rendering program according to the target program characteristics to obtain a target rendering program for running on the target processor.

[0008] The present application also provides an adjustment device for a rendering program, including: an acquisition module, configured to acquire target operation information of an initial rendering program by a target processor, where the target operation information is used to indicate the usage of program operation resources of the target processor during the operation of the initial rendering program; a conversion module, configured to convert target program features of the initial rendering program according to the target operation information and reference operation information, where the reference operation information is used to indicate the response of program operation resources when running a test program, the test program is used to test the resource performance of program operation resources, and the target program features are program operation features in the initial rendering program whose influence on the rendering performance is greater than a target influence degree; an adjustment module, configured to adjust the initial rendering program according to the target program features to obtain a target rendering program for running on the target processor.

[0009] The present application also provides an electronic device, including: a memory, configured to store a computer program; a processor, configured to implement the steps of any one of the above-mentioned rendering program adjustment methods when executing the computer program.

[0010] The present application also provides a computer-readable storage medium, in which a computer program is stored, where the computer program implements the steps of any one of the above-mentioned rendering program adjustment methods when executed by a processor.

[0011] The present application also provides a computer program product, including a computer program, where the computer program implements the steps of any one of the above-mentioned rendering program adjustment methods when executed by a processor.

[0012] Through the present application, by acquiring the target operation information of the initial rendering program running on the target processor, it is possible to characterize the running performance of the initial rendering program on the target processor through the usage of program operation resources of the target processor during the operation of the initial rendering program. Furthermore, according to the target operation information and reference operation information, the target program features of the rendering program are converted to obtain the program operation features in the initial rendering program whose influence on the rendering performance is greater than the target influence degree by the program operation resources of the current target processor. Then, by adjusting the initial rendering program according to the target program features, it is possible to optimize the initial rendering program according to the program operation resources of the target processor, so that the adaptability between the adjusted target rendering program and the program operation resources of the target processor is higher, which can solve the technical problem of low optimization efficiency of the running performance of the rendering program in the related art and achieve the effect of improving the optimization efficiency of the running performance of the rendering program. Description of the Drawings

[0013] To more clearly illustrate the embodiments of the present application, the following will briefly introduce the accompanying drawings required in the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.

[0014] Figure 1 is a hardware block diagram of a method for adjusting a rendering program according to an embodiment of the present application;

[0015] Figure 2 is a flowchart of a method for adjusting a rendering program according to an embodiment of the present application;

[0016] Figure 3 is an optional architecture diagram of a method for adjusting a rendering program provided by the present application;

[0017] Figure 4 is an optional schematic diagram of instruction mix classification according to an embodiment of the present application Figure 1 ;

[0018] Figure 5 is an optional schematic diagram of instruction mix classification according to an embodiment of the present application Figure 2 ;

[0019] Figure 6 is an optional schematic diagram of instruction mix comparison according to an embodiment of the present application;

[0020] Figure 7 is an optional schematic diagram of instruction-level parallelism according to an embodiment of the present application Figure 1 ;

[0021] Figure 8 is an optional schematic diagram of instruction-level parallelism according to an embodiment of the present application Figure 2 ;

[0022] Figure 9 is an optional schematic diagram of program running characteristics according to the present application;

[0023] Figure 10 is a structural block diagram of a device for adjusting a rendering program according to an embodiment of the present application. Detailed implementation manners

[0024] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only some, rather than all, of the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the protection scope of the present application.

[0025] It should be noted that in the description of this application, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or device. The terms "first", "second", etc. in this application are used to distinguish similar objects, rather than to describe a specific order or sequence.

[0026] To enable those skilled in the art of this technology to better understand the solution of this application, the following further detailed description of this application will be given in conjunction with the accompanying drawings and specific embodiments.

[0027] Combined with the specific application environment architecture or specific hardware architecture on which the execution of the adjustment method of the rendering program depends, the specific application environment architecture or specific hardware architecture will be described here.

[0028] The method embodiments provided in the embodiments of this application can be executed on a server device or a similar computing device. Taking the operation on a server device as an example, Figure 1 is a hardware block diagram of an adjustment method of a rendering program according to an embodiment of this application. As Figure 1 shown, the server device may include one or more ( Figure 1 only one is shown in Figure 1 the processor 102 (the processor 102 may include, but is not limited to, a processing device such as a microprocessor MCU or a programmable logic device FPGA) and a memory 104 for storing data. Among them, the above-mentioned server device may further include a transmission device 106 for communication functions and an input / output device 108. Those of ordinary skill in the art can understand that Figure 1 the structure shown in Figure 1 is only illustrative and does not limit the structure of the above-mentioned server device. For example, the server device may also include more or fewer components than

[0029] The memory 104 can be used to store computer programs, for example, software programs and modules of application software, such as the computer program corresponding to the startup method of the operating system in the embodiments of the present application. The processor 102 executes various functional applications and data processing by running the computer programs stored in the memory 104, that is, implements the above-mentioned method. The memory 104 may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memories. In some instances, the memory 104 may further include a memory remotely disposed relative to the processor 102, and these remote memories can be connected to the server device through a network. Examples of the above-mentioned network include but are not limited to the Internet, intranet, local area network, mobile communication network, and combinations thereof.

[0030] The transmission device 106 is used to receive or send data via a network. Specific examples of the above-mentioned network may include a wireless network provided by a communication provider of the server device. In one instance, the transmission device 106 includes a network adapter (Network Interface Controller, abbreviated as NIC), which can be connected to other network devices through a base station and thus communicate with the Internet. In one instance, the transmission device 106 may be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0031] Embodiments of the present application provide a method for adjusting a rendering program. The method is described in detail in combination with the execution flow of the method for adjusting the rendering program.

[0032] The following explains the professional terms that appear in the present application:

[0033] CPU (Central Processing Unit): Central Processing Unit, which is the computing and control core of a computer system and is responsible for executing information processing and program running;

[0034] Rendering: Rendering, the process of converting a 3D model into a 2D image, which is widely used in computer graphics and game development;

[0035] MiBench: An embedded benchmark suite that contains multiple sub-category programs and is used to evaluate the performance of embedded processors;

[0036] Rasterization: Rasterization, which is the process in computer graphics of converting geometric shapes in a 3D scene into a 2D image;

[0037] Ray Tracing: Ray tracing is a computer graphics rendering technique that calculates the color of each pixel in a scene by simulating the path of light from the light source to the camera;

[0038] Renderer: A renderer is a software program that converts a 3D scene or model into a 2D image, enhancing the image's realism by applying effects such as lighting, shadows, and textures;

[0039] Instruction Mix: Instruction mix refers to the proportional distribution of different types of instructions in a program;

[0040] Instruction-level Parallelism (ILP): Instruction-level parallelism refers to the ability of a computer program to execute multiple instructions simultaneously;

[0041] Register Traffic: Register traffic refers to the usage and dependencies of registers in a program;

[0042] Working Set Size: Working set size refers to the number of memory pages or blocks that are frequently accessed during the execution of a program;

[0043] Data Stream Strides: Data stream strides refer to the stride size in the data access pattern in memory;

[0044] Branch Predictability: Branch predictability refers to the ability of a processor to predict the execution path of program branches;

[0045] Memory Reuse Distance: Memory reuse distance refers to the memory distance between two accesses when data in a program is re-accessed;

[0046] Kiviat Diagram: A Kiviat diagram is a graphical representation method used to display the relative sizes and trends of multidimensional data;

[0047] PMP (Partial Match Predictor): A partial match predictor, an algorithm used for branch prediction;

[0048] PAg PPM (Prediction by Partial Match): Prediction by partial match, an evaluation metric for the PMP predictor;

[0049] Global Load Stride: Global load stride is the address difference between adjacent memory read accesses;

[0050] Memory Reuse Distance: The memory reuse distance is the interval at which data is reused in the cache;

[0051] Control Flow: The control flow is the order and conditions of instruction execution in a program;

[0052] Floating-Point: Floating-point operations are arithmetic operations for handling decimals and fractions;

[0053] Micro-architecture: The micro-architecture is the internal logical structure and circuit design of a processor;

[0054] Benchmark: A benchmark is a standard test program used to evaluate the performance of a computer system or component.

[0055] In this embodiment, a method for adjusting a rendering program is provided. Figure 2 It is a flowchart of a method for adjusting a rendering program according to an embodiment of the present application, as Figure 2 shown. The method includes the following steps:

[0056] Step S202: Obtain the target running information of the target processor for the initial rendering program, where the target running information is used to indicate the usage of the program running resources of the target processor during the running of the initial rendering program;

[0057] Step S204: Convert the target program characteristics of the initial rendering program according to the target running information and the reference running information, where the reference running information is used to indicate the response of the program running resources when running the test program, the test program is used to test the resource performance of the program running resources, and the target program characteristics are the program running characteristics in the initial rendering program that have a greater impact on the rendering performance than the target impact degree;

[0058] Step S206: Adjust the initial rendering program according to the target program characteristics to obtain a target rendering program for running on the target processor.

[0059] Through the above steps, by obtaining the target running information of the target processor running the initial rendering program, the usage of the program running resources of the target processor during the running process of the initial rendering program is used to characterize the running performance of the initial rendering program on the target processor. Furthermore, based on the target running information and the reference running information, the target program characteristics of the rendering program are converted, and the program running characteristics in the initial rendering program that have a greater impact on the rendering performance than the target impact degree by the program running resources of the current target processor are obtained. Then, by adjusting the initial rendering program according to the target program characteristics, the program optimization of the initial rendering program is realized according to the program running resources of the target processor, so that the adjusted target rendering program has a higher adaptability to the program running resources of the target processor, which can solve the technical problem of low optimization efficiency of the running performance of the rendering program in the related art and achieve the effect of improving the optimization efficiency of the running performance of the rendering program.

[0060] In the embodiment provided in step S202, the target processor is a processor with functions related to executing graphics processing services for running the rendering program. The target processor may include, but is not limited to, a GPU (Graphic Processing Unit) and a CPU (Central Processing Unit); the initial rendering program is the rendering program to be optimized, and the initial rendering program is mainly used for two-dimensional and three-dimensional graphic rendering work; when the initial rendering program runs on the target processor, it needs to use the program running resources of the target processor, and the program running resources include, but are not limited to, computing instruction resources, such as memory read, memory write, control flow, arithmetic operation, floating-point operation, stack operation, shift operation, string operation, SSE instruction, NOP instruction, register transfer, and other types of computing instructions; memory resources, such as register dependence distance, cache, memory reuse distance, data flow stride, register traffic characteristics, and branch predictability, etc.

[0061] Optionally, in the embodiments of the present application, the program running resources are the running resources deployed on the target processor to implement the running of the rendering program. The program running resources may include, but are not limited to, the microarchitecture characteristics of the processor (referring to the specific implementation details inside the processor, which determine how the processor executes instructions, processes data, and how to organize its internal components). The microarchitecture characteristics may include, but are not limited to: instruction execution mode, pipeline design, memory hierarchy, organization of functional units, implementation of the instruction set architecture (ISA), multi-core and multi-thread support, register dependence distance, memory reuse distance, data flow stride, register traffic characteristics, etc. The present application does not make any limitations in this regard.

[0062] Optionally, in the embodiments of the present application, the initial rendering program can be, but is not limited to, constructed based on the resource characteristics of the program running resources of the target processor. For example, for the running resources corresponding to the instruction-level characteristics of the processor, the instruction set of the processor (such as the x86 instruction set) is divided and classified, so as to obtain multiple types (which can be, but are not limited to, including memory read, memory write, control flow, arithmetic operation, floating-point operation, stack operation, shift operation, string operation, SSE instruction, NOP instruction, register transfer, and other types, etc., a total of 12 types). Furthermore, through the analysis of the instruction set, it is found that arithmetic operation instructions account for the highest proportion among all instruction types, reaching 39%, while memory read, memory write, and control flow instructions also occupy an important position, and the sum of the three reaches 70%. Therefore, the proportion of each program type in the initial rendering program can be configured according to the proportion of various instructions in the instruction set. Or, the number of instruction parallelisms in the initial rendering program can also be configured according to the operation resource characteristics in the processor program running resources. Thus, the initial rendering program is configured according to the program running resources of the target processor, so as to improve the adaptation relationship between the initial rendering program and the program running resources of the target processor, reduce the workload in the subsequent optimization process of the initial rendering program, and improve the optimization efficiency of the initial rendering program.

[0063] Optionally, in the embodiments of the present application, the target running information is the rendering program execution data collected during the process of the target processor running the initial rendering program. The present application reflects the behavioral characteristics of the workload of the rendering program on the processor by the characteristics that are independent of the microarchitecture shown by the processor during the running and rendering process of the processor (that is, the performance of the program running resources when the processor runs the rendering program). Therefore, the target running information can be, but is not limited to, including instruction-level data and memory-level data. Among them, the instruction-level data includes instruction-level characteristic data such as the instruction mix and ILP of the rendering program; the memory-level data includes memory-level characteristic data such as the register dependence distance, memory reuse distance, and data flow stride of the rendering program.

[0064] In the embodiment provided in step S204, the test program is a program for testing the general performance of the target processor in program operation. The test program is a non-rendering graphics function program and has a performance test function for the general requirements of the processor when running different types of programs. The test program may but is not limited to include, for example, basicmath, bitcount, qsort-large, and susan programs in the automotive field, the jpeg program in the consumer electronics field, the patricia program in the network field, the blowfish and sha programs in the security field, and the CRC32, FFT, and gsm programs in the communication field; the reference running information is the program running resources required when the test program runs on the target processor. Comparing and analyzing the target running information and the reference running information obtains the target program features, that is, the key features that have the greatest impact on the performance of the initial rendering program.

[0065] Optionally, in the embodiment of the present application, the target program features are the program running features to be optimized. By optimizing the target program features, the running performance of the rendering program is improved. The target program features may but are not limited to include algorithm optimization features, parallel computing features, cache optimization features, and branch prediction optimization features. The algorithm optimization features are for compute-intensive tasks in the rendering program. The parallel computing features involve utilizing the parallel computing capabilities of multi-core processors. The cache optimization features are for frequent memory access operations in the rendering program. The branch prediction optimization features are for frequent control flow operations in the rendering program.

[0066] In the embodiment provided in step S206, the method of adjusting the initial rendering program according to the target program features may include: obtaining the program adjustment information corresponding to the target program features, and then adjusting the initial rendering program according to the adjustment method indicated by the program adjustment information to obtain the target rendering program. Or it can also analyze the program information corresponding to the target program features in the initial rendering program to obtain the program adjustment information of the initial rendering program, and then adjust the initial rendering program according to the program adjustment method indicated by the program adjustment information to obtain the target rendering program. The present solution does not limit this.

[0067] Optionally, in the embodiment of the present application, the target rendering program can be used to compare and analyze with the initial rendering program and the test program, and then iteratively optimize the target rendering program to obtain the iteratively optimized rendering program.

[0068] As an alternative implementation, obtain the target running information of the target processor for the initial rendering program, including: obtaining the microarchitecture features of the target processor, where the microarchitecture features are used to indicate the resource features of the program running resources of the target processor; constructing the initial rendering program for the target processor according to the microarchitecture features; calling the target processor to run the initial rendering program, and collecting the target running information during the running process of the initial rendering program.

[0069] Optionally, in the embodiments of the present application, the microarchitecture characteristics of the processor may, but are not limited to, characterize the characteristics of the program running resources of the processor from multiple dimensions, and the microarchitecture features may, but are not limited to, include the instruction-level features of the processor (which may include instruction mix features and instruction-level parallelism features), and memory-level features.

[0070] Optionally, in the embodiments of the present application, when constructing the initial rendering program, the proportional relationship of different types of rendering programs in the initial rendering program may be configured according to the instruction mix features of the target processor; the number of rendering instructions running in parallel in the initial rendering program may be configured according to the instruction-level parallelism features of the target processor; and the parameters in aspects such as instruction mix, register dependence distance, working set, data flow step size, branch predictability, and memory reuse distance in the initial rendering program may be configured according to the memory-level features of the target processor.

[0071] Through the above content, by analyzing the microarchitecture characteristics of the target processor, the initial rendering program is constructed according to the microarchitecture characteristics of the processor, thereby improving the matching relationship between the initial rendering program and the program running resources of the processor.

[0072] As an alternative implementation, call the target processor to run the initial rendering program, and collect the target running information during the running process of the initial rendering program, including: configuring multiple instruction windows for the initial rendering program, where each instruction window is used to indicate that the initial rendering program executes multiple rendering instructions in parallel, and the number of rendering instructions indicated to be executed by different instruction windows is different, controlling the initial rendering program to execute the rendering instructions of multiple instruction windows respectively; collecting the instruction execution information and memory access information after the initial rendering program responds to the instruction window, where the instruction execution information is used to indicate the execution situation of the initial rendering program for the parallel rendering instructions, and the memory access information is used to indicate the memory access situation of the rendering program to the target processor during the running process; determining the instruction execution information as the first running information of the initial rendering program, and determining the second running information of the initial rendering program according to the memory access information, where the target running information includes the first running information and the second running information.

[0073] Optionally, in the embodiments of the present application, the instruction window is a buffer with a limited size, which is used to store the instructions fetched from the Instruction Cache. These instructions wait for scheduling in the window so as to be sent to the execution unit at an appropriate time.

[0074] Optionally, in the embodiments of the present application, the number of rendering instructions to be executed configured in different instruction windows is different. By setting different instruction windows, the Instruction-level Parallelism (ILP) of the processor when running the rendering program is tested. Instruction-level Parallelism refers to the parallel or simultaneous execution of a series of instructions in a computer program. More specifically, ILP refers to the average number of instructions executed in each step during this parallel execution process.

[0075] Optionally, in the embodiments of the present application, the instruction execution information is the number of rendering instructions actually executed in parallel by the rendering program.

[0076] Optionally, in the embodiments of the present application, the second running information can be, but is not limited to, obtained by quantitatively analyzing the memory access information. In the present application, by collecting the running characteristics of the processor when running the rendering program, the running performance of the rendering program is reflected according to the running characteristics, and then through information conversion of the memory access information, the second running information characterizing the running performance of the rendering program at the processor memory level is obtained.

[0077] Optionally, in the embodiments of the present application, by setting dynamic instruction windows of different sizes, the present application deeply reveals the potential deficiencies of the rendering program in terms of instruction-level parallelism, and accordingly proposes a series of targeted performance improvement strategies. These strategies involve multiple aspects such as the optimization of instruction scheduling and the improvement of branch prediction accuracy, aiming to significantly improve the ILP level of the rendering program by finely regulating the instruction execution process inside the CPU, and thus greatly accelerate the rendering speed.

[0078] Through the above content, by deeply exploring the instruction-level and memory-level characteristics of the CPU microarchitecture, an unprecedented comprehensive analysis method is proposed, which can accurately depict the complex behavior patterns of the rendering program in terms of instruction mix, Instruction-level Parallelism (ILP), register dependence distance, memory reuse distance, data flow stride, and branch predictability. This in-depth analysis not only reveals the inherent workload characteristics of the rendering program but also lays a solid theoretical foundation for subsequent performance optimization, which is a major breakthrough in the analysis of the behavior characteristics of the rendering program in the prior art.

[0079] As an alternative implementation, determining the second running information of the initial rendering program according to the memory access information includes: using the principal component analysis algorithm to determine the data scores of each of the first access data of the first quantity, where the first access data of the first quantity is used to characterize the memory access situation of the target processor from different dimensions, and the data score is used to indicate the importance of the first access data to the rendering performance, and the memory access information includes the first access data of the first quantity; sorting the first access data in descending order of the data scores to obtain a target sequence; screening out the second access data of the second quantity with data scores greater than the target score from the target sequence, where the second quantity is less than the first quantity; and determining the second access data as the second running information.

[0080] Optionally, in the embodiments of the present application, the first access data of the first quantity is obtained by dimension expansion of the memory access information (processor microarchitecture-independent characteristics) of the processor. For example, on the basis of 47 microarchitecture-independent characteristics, we expand it to 99 microarchitecture-independent characteristics, so that the inherent behavior of the rendering program in aspects such as instruction mix, register dependence distance, working set, data flow stride, branch predictability, and memory reuse distance can be described more comprehensively.

[0081] Optionally, in the embodiments of the present application, the principal component analysis algorithm screens out relatively important access data from the first access data by calculating the scores of each first access data and sorting the access data according to the scores, so as to realize the dimensionality reduction processing of the access data according to the importance of the data.

[0082] Through the above content, by using the principal component analysis algorithm to perform dimensionality reduction processing on the data, the reliability of the second running information is ensured on the premise of reducing the amount of access data processing.

[0083] As an alternative implementation, adjusting the initial rendering program according to the target program characteristics includes: obtaining the program information corresponding to the target program characteristics in the initial rendering program; generating program adjustment information for the initial rendering program according to the program information, and using the program adjustment information to adjust the initial rendering program to obtain a target rendering program for running on the target processor.

[0084] Optionally, in the embodiments of the present application, for the compute-intensive tasks in the rendering program, the present application proposes a performance improvement method based on algorithm optimization. By deeply analyzing the rendering algorithm, the computing bottleneck and redundant calculations are identified, and the rendering performance is improved by means of optimizing the algorithm structure and reducing the computing complexity. For example, in the ray tracing algorithm, the algorithm efficiency can be improved by optimizing the ray intersection calculation and reducing unnecessary ray sampling.

[0085] Optionally, in the embodiments of the present application, leveraging the parallel computing capabilities of a multi-core processor, the present application proposes a method for improving rendering performance based on parallel computing. By dividing the rendering task into multiple subtasks and performing the calculations in parallel on multiple processor cores, the rendering speed can be significantly improved.

[0086] Optionally, in the embodiments of the present application, for the frequent memory access operations in the rendering program, the present application proposes a method for improving performance based on cache optimization. By deeply analyzing the memory access pattern of the rendering program, optimizing the cache usage strategy, increasing the cache hit rate, and reducing the memory access latency. For example, during the rendering process, the principle of locality can be utilized to store frequently accessed data blocks in the cache to reduce the number of memory accesses.

[0087] Optionally, in the embodiments of the present application, for the frequent control flow operations in the rendering program, the present application proposes a method for improving performance based on branch prediction optimization. By optimizing the branch prediction algorithm, improving the accuracy of branch prediction, and reducing the performance loss caused by branch mispredictions. For example, in the rendering program, historical branch information can be used to predict the future branch directions, thereby reducing the occurrence of branch mispredictions.

[0088] Through the above content, by analyzing the program information corresponding to the program characteristics, the program adjustment information for the rendering program is obtained, ensuring the accuracy and reliability of the program adjustment information and improving the adjustment efficiency of the rendering program.

[0089] As an alternative implementation, program adjustment information for the initial rendering program is generated based on program information, and the initial rendering program is adjusted using the program adjustment information to obtain a target rendering program for running on a target processor, including: when the target program feature is used to indicate that the running performance of the target rendering algorithm included in the initial rendering program is lower than a first threshold, identifying the computing nodes in the initial rendering algorithm, and adjusting the algorithm structure of the target rendering algorithm according to the node attributes of the computing nodes to obtain the target rendering algorithm; when the target program feature is used to indicate that the parallel computing performance of the initial rendering program is lower than a second threshold, dividing the rendering tasks to be executed by the initial rendering program into multiple subtasks; configuring the binding relationship between the multiple subtasks and the multiple processor cores on the target processor, where different subtasks are bound to different processor cores; when the target program feature is used to indicate that the memory access performance of the initial rendering program for the target processor is lower than a third threshold, obtaining the target data accessed by the initial rendering program, where the target data is the data in the memory whose access frequency by the initial rendering program is greater than the target frequency; configuring a target cache location for the target data in the cache space of the target processor; when the target program feature is used to indicate that the frequency of the control flow operations of the initial rendering program is greater than a fourth threshold, predicting the target score trend of the initial rendering program after the current moment according to the reference score information between the current moments, and adjusting the initial rendering program using the target score trend.

[0090] As an alternative implementation, target program features of the initial rendering program are converted based on target running information and reference running information, including: extracting candidate running information corresponding to the information type of the target running information from the reference running information; matching the candidate running information with the target running information; when the running performance indicated by the candidate running information is greater than or equal to the running performance indicated by the target running information, determining the program running feature corresponding to the target running information as the target program feature.

[0091] Optionally, in the embodiments of the present application, by matching the candidate running information corresponding to the information type of the target running information in the reference running information with the target running information, the running performance of the current rendering program on the processor is reflected according to the running performance of the processor on a general benchmark tool.

[0092] As an alternative implementation, the present application provides a method for analyzing the behavior characteristics and optimizing the performance of a rendering program based on CPU microarchitecture characteristics. This method deeply analyzes the independent characteristics of the CPU microarchitecture, comprehensively understands the workload behavior characteristics of the rendering program, and proposes targeted performance optimization strategies on this basis. The technical solution of the present invention aims to improve the execution efficiency of the rendering program and meet the requirements of the current graphics rendering field for high-performance processors.

[0093] The main design concept of this application is as follows:

[0094] Analysis of CPU microarchitecture characteristics:

[0095] (1) Instruction-level characteristic analysis

[0096] The present invention deeply analyzes the rendering program at the CPU instruction level, mainly including Instruction Mix and Instruction-Level Parallelism (ILP). First, the x86 instruction set is divided into 12 types, including memory read, memory write, control flow, arithmetic operation, floating-point operation, stack operation, shift operation, string operation, SSE instruction, NOP instruction, register transfer, and other types. By analyzing the Instruction Mix of the rendering program, it is found that arithmetic operation instructions account for the highest proportion among all instruction types, reaching 39%, while memory read, memory write, and control flow instructions also play an important role, and the sum of the three reaches 70%. This discovery provides an important design basis for processor architects, that is, when designing a rendering processor, the proportion of arithmetic operation instructions can be appropriately increased to meet the requirements of modern rendering applications.

[0097] In addition, the present invention also uses the MICA performance analysis tool to quantitatively analyze the ILP of the rendering program. By setting 32, 64, 128, and 256 dynamic instruction windows, and based on an ideal out-of-order processor and perfect cache and branch predictor, the ILP of the rendering program is quantitatively evaluated. The analysis results show that the ILP data of the rendering program are generally lower than those of the MiBench benchmark suite, indicating that there is still a large room for improvement in the instruction-level parallelism of the rendering program.

[0098] (2) Memory-level characteristic analysis

[0099] At the CPU memory level, the present invention deeply analyzes the register dependence distance, memory reuse distance, data flow stride, register traffic characteristics, and branch predictability of the rendering program. First, by extending to 99 microarchitecture-independent characteristics, the internal behavior of the rendering program is more comprehensively characterized. These characteristics include Instruction Mix, register traffic, working set size, data flow stride, branch predictability, and memory reuse distance, etc.

[0100] Through the quantitative analysis of these characteristics, the present invention discovers that the rendering program exhibits unique behavioral patterns in aspects such as register dependence distance, memory reuse distance, and data flow stride. For example, the average PAg PPM (prediction based on partial match) of the rendering program is 8 times that of the MiBench benchmark suite, indicating that the rendering program has strong data dependencies. In addition, the global load stride of the rendering program is relatively concentrated, while the MiBench benchmark suite contains fewer global memory read accesses. These findings provide an important basis for further optimizing the memory access pattern of the rendering program.

[0101] Rendering Program Performance Optimization Strategies:

[0102] Based on the above analysis of CPU microarchitecture characteristics, the present invention proposes performance optimization strategies for the rendering program, mainly including the following aspects:

[0103] (1) Algorithm Optimization

[0104] For the compute-intensive tasks in the rendering program, the present invention proposes a performance improvement method based on algorithm optimization. By deeply analyzing the rendering algorithm, identifying the computing bottlenecks and redundant computations, and improving the rendering performance by means of optimizing the algorithm structure and reducing the computing complexity. For example, in the ray tracing algorithm, the algorithm efficiency can be improved by optimizing the ray intersection calculation and reducing unnecessary ray sampling.

[0105] (2) Parallel Computing

[0106] Utilizing the parallel computing ability of multi-core processors, the present invention proposes a rendering performance improvement method based on parallel computing. By dividing the rendering task into multiple subtasks and performing the calculations in parallel on multiple processor cores, the rendering speed can be significantly improved.

[0107] (3) Cache Optimization

[0108] For the frequent memory access operations in the rendering program, the present invention proposes a performance improvement method based on cache optimization. By deeply analyzing the memory access pattern of the rendering program, optimizing the cache usage strategy, increasing the cache hit rate, and reducing the memory access latency. For example, during the rendering process, the principle of locality can be utilized to store frequently accessed data blocks in the cache to reduce the number of memory accesses.

[0109] (4) Branch Prediction Optimization

[0110] For the frequent control flow operations in the rendering program, the present invention proposes a performance improvement method based on branch prediction optimization. By optimizing the branch prediction algorithm, the accuracy of branch prediction is improved, and the performance loss caused by branch misprediction is reduced. For example, in the rendering program, historical branch information can be used to predict future branch directions, thereby reducing the occurrence of branch mispredictions.

[0111] To achieve the above, this application involves a data collection module, a data analysis module, and a performance optimization module, specifically as follows:

[0112] (1) Data collection and analysis module

[0113] The present invention designs a data collection and analysis module for collecting the execution data of the rendering program and performing quantitative analysis on it. This module includes an instruction-level data collection sub-module and a memory-level data collection sub-module. The instruction-level data collection sub-module is responsible for collecting instruction-level characteristic data such as the instruction mix and ILP of the rendering program; the memory-level data collection sub-module is responsible for collecting memory-level characteristic data such as the register dependence distance, memory reuse distance, and data flow stride of the rendering program.

[0114] The collected data will be transmitted to the data analysis sub-module for quantitative analysis. The data analysis sub-module will use statistical methods and principal component analysis methods to process and analyze the data to extract the key characteristics that have the greatest impact on the performance of the rendering program.

[0115] (2) Performance optimization module

[0116] Based on the results of the data analysis module, the present invention designs a performance optimization module for optimizing the performance of the rendering program. This module includes an algorithm optimization sub-module, a parallel computing sub-module, a cache optimization sub-module, and a branch prediction optimization sub-module. The algorithm optimization sub-module is responsible for optimizing the rendering algorithm according to the algorithm analysis results; the parallel computing sub-module is responsible for dividing the rendering task into multiple sub-tasks and performing parallel computing on a multi-core processor; the cache optimization sub-module is responsible for optimizing the cache usage strategy to improve the cache hit rate; the branch prediction optimization sub-module is responsible for optimizing the branch prediction algorithm to improve the accuracy of branch prediction.

[0117] (3) Visualization and debugging module

[0118] To facilitate users to understand and debug the optimized rendering program, the present invention also designs a visualization and debugging module. This module will provide a visualization display function of the rendering program performance, such as Kiviat diagrams, etc., to help users intuitively understand the performance bottlenecks and optimization effects of the rendering program. At the same time, this module will also provide debugging tools to help users quickly locate and solve problems existing in the rendering program.

[0119] In summary, the present invention proposes a method for analyzing the behavior characteristics and optimizing the performance of a rendering program based on the characteristics of the CPU microarchitecture. By deeply analyzing the independent characteristics of the CPU microarchitecture, this method comprehensively understands the workload behavior characteristics of the rendering program, and on this basis, proposes targeted performance optimization strategies. The technical solution of the present invention aims to improve the execution efficiency of the rendering program and meet the requirements for high-performance processors in the current graphics rendering field.

[0120] Figure 3 is an optional architecture diagram of a rendering program adjustment method provided by the present application. As Figure 3 shown, first, the rendering program is run and tested using profiler tools to obtain row data series, and then this data is analyzed to obtain integrated data independent of the microarchitecture. Then, this data is used to perform memory level analysis and instruction level analysis on the rendering program.

[0121] Rendering is the process of converting the models in a three-dimensional scene into two-dimensional images according to rendering parameters. We selected 60 rendering programs, as shown in Table I (where Abr. represents the benchmark test abbreviation), and divided them into six rendering categories, including 13 rasterization programs, 12 ray tracing programs, 13 renderer programs, 8 WebGL programs, 10 game rendering programs, and 3 photon mapping programs. Since rasterization, ray tracing, rendering, and game rendering account for the vast majority of the rendering scene, we gave them relatively high proportions. The ray tracing programs also include path tracing algorithms, rasterization includes point-based and voxel-based global illumination, and game renderers include algorithms such as ambient occlusion. Therefore, our rendering program benchmark tests are diverse and representative. Table 1 is an optional rendering program table according to an embodiment of the present application, as shown in Table 1:

[0122]

[0123]

[0124] Meanwhile, this paper selects the MiBench benchmark suite as the experimental application for CPU non-rendering programs. MiBench contains a total of 35 embedded programs, divided into six subcategories: automotive and industrial manufacturing, consumer electronics, office automation, networking, security, and communications. This paper selects 11 typical applications as experimental applications, including basicmath, bitcount, qsort-large, and susan in the automotive field, jpeg in the consumer electronics field, patricia in the networking field, blowfish and sha in the security field, and CRC32, FFT, and gsm in the communications field.

[0125] As the computing and control core of a computer system, the central processing unit (CPU) is the final execution unit for information processing and program operation. Since its inception, the CPU has made great progress in terms of logical structure, operating efficiency, and functional expansion. With the widespread popularity of graphics applications, processor architects have begun to pay more attention to the CPU's support for graphics. In this chapter, we will introduce the microarchitecture-independent features of CPU-level rendering programs to explore the behavioral characteristics of rendering program workloads. At the same time, we will also compare the behavioral characteristics of existing rendering test programs with well-known benchmark suites to reveal the differences in microarchitecture-independent features between rendering test programs and other benchmark suites.

[0126] First, we analyze the characteristics at the CPU instruction level, including instruction mix, instruction-level parallelism, and some important instruction types, such as memory reads, memory writes, control flow, and floating-point operations.

[0127] Instruction mix:

[0128] The instruction mix is evaluated by classifying the executed instructions. Considering that the x86 architecture is not a load-store architecture, memory reads / writes are calculated separately. In the instruction mix section, we divide the x86 instruction set into 12 categories, namely memory reads, memory writes, control flow, arithmetic operations, floating-point operations, stack operations, shift operations, string operations, SSE (Streaming SIMD Extensions) instructions, NOP (No Operation) instructions, register transfer instructions, and other instructions.

[0129] Figure 4 is a schematic diagram of an optional instruction mix classification according to an embodiment of the present application Figure 1 as Figure 4As shown below, our observations are as follows: (1) Among all 60 CPU rendering test programs, the average proportion of arithmetic operations is the highest, reaching 39%. (2) The three instruction types with the highest proportions are arithmetic operations, memory reads, and memory writes, and their total reaches 70%. (3) Among all instruction types, SSE (Streaming SIMD Extensions), register transfers, and NOP (No Operation) types are hardly used. The above observations on the instruction mix of rendering programs are very important because they indicate that when designing a CPU rendering processor, some appropriate changes can be made based on the existing CPU instruction set to improve the performance of the rendering processor. At the same time, this also provides an opportunity for processor architects to optimize the CPU processor. For example, the architect can appropriately increase the proportion of arithmetic operation instructions to adapt to modern rendering applications. In addition, we also selected MiBench (an embedded benchmark suite) to compare the instruction mix with the rendering programs.

[0130] Figure 5 is an optional schematic diagram of instruction mix classification according to an embodiment of the present application Figure 2 , as Figure 5 shown, we also selected 12 instruction types, and the observations are as follows: (1) In the MiBench test suite, the proportion of arithmetic operations is the highest, reaching 54%. Compared with the rendering programs, the proportion of arithmetic operations in MiBench exceeds 50%, which means that the program uses half of its instructions for computing instructions, requiring the processor to have a large number of powerful ALU (Arithmetic Logic Unit) computing units. (2) The top three with the highest proportions are arithmetic operations, memory reads, and control flow, and their total reaches 83%. Compared with the rendering programs, the total proportion of the top three types of instruction mix in the rendering programs is only 70%. At the same time, we noticed that the average percentage of floating-point operations in the rendering programs is 10%, indicating that the equally important floating-point operations in the rendering programs cannot be ignored. (3) Among all instruction types, the MiBench test suite also hardly uses SSE, register transfers, and NOP types. This observation provides good advice for processor architects, that is, when designing a new rendering processor, there is no need to design more SSE, register transfer, and NOP type instructions on the existing processor architecture.

[0131] Figure 6 is an optional schematic diagram of instruction mix comparison according to an embodiment of the present application, as Figure 6Shows the percentage of each instruction type in certain rendering programs and MiBench programs. Our observations are as follows: (1) Memory reads, memory writes, control flow, and arithmetic operations account for a relatively high proportion in all rendering programs and MiBench programs. (2) The proportion of stack instruction types in rendering programs (i.e., programs 18 to 87) is higher than that in MiBench programs (i.e., programs bm to FFT). We know that stack instructions are used to store intermediate results of program operations. Therefore, appropriately increasing stack instructions helps improve the performance of the computing unit. (3) The proportion of floating-point operations in MiBench programs is higher than that in rendering programs, indicating that there are relatively fewer floating-point operations in rendering programs running on the CPU, and architects can appropriately design fewer floating-point computing units.

[0132] (II) Instruction-Level Parallelism:

[0133] Instruction-level parallelism (ILP) refers to the parallel or simultaneous execution of a series of instructions in a computer program. More specifically, ILP refers to the number of instructions executed on average at each step during this parallel execution process. To quantify the instruction-level parallelism of rendering programs, we used the MICA performance analysis tool to characterize windows of 32, 64, 128, and 256 dynamic instructions based on a processor with ideal out-of-order execution capabilities, perfect caches, and branch predictor sizes.

[0134] Figure 7 Is an optional instruction-level parallelism schematic according to an embodiment of the present application Figure 1 As Figure 7 shown, shows the comparison of instruction-level parallel data between rendering programs and MiBench. From program 18 to program 87, the instruction-level parallelism data of rendering programs are generally lower than those of MiBench. Through statistical analysis, we found that the average values of ILP32, ILP64, ILP128, and ILP256 for 60 rendering programs are 4.29, 4.94, 5.42, and 5.70 respectively, while the average values of ILP32, ILP64, ILP128, and ILP256 for MiBench test programs are 5.42, 6.56, 7.38, and 7.97 respectively. This indicates that in each instruction-level parallel range, the parallelism of rendering programs is inferior to that of MiBench.

[0135] Figure 8 Is an optional instruction-level parallelism schematic according to an embodiment of the present application Figure 2 As Figure 8As shown, we divide the values of instruction-level parallelism into 7 value intervals. For ILP32, the values of the renderer mainly concentrate between 3 and 6, while the values of MiBench mainly concentrate between 4 and 7. For ILP32, ILP64, ILP128, and ILP256, the instruction-level parallelism values of MiBench are all one interval higher than those of the renderer, which also has a certain impact on the parallel performance of the CPU processor. Therefore, there is great room for improvement in CPU performance by increasing ILP.

[0136] (III) CPU Memory Hierarchy Characteristics

[0137] The memory hierarchy characteristics of a program have a significant impact on program performance. Based on 47 microarchitecture-independent characteristics, we expand it to 99 microarchitecture-independent characteristics, so that we can more comprehensively describe the inherent behaviors of the renderer in terms of instruction mix, register dependence distance, working set, data flow stride, branch predictability, and memory reuse distance. As above, the instruction mix and instruction-level parallelism have been explained. Now we will introduce the following characteristics: The register traffic has two values, namely the average number of input operands of an instruction and the average usage degree, and 7 probability values, representing the seven numbers of dynamic instructions between writing to a register and reading it. The working set focuses on the sizes of instructions and data flows for 32-byte blocks and 4KB page sizes. The data flow stride focuses on the local and global data strides between loads and stores. Branch predictability considers global and local history predictors, as well as per-address and global predictors using different history lengths of 4 bits, 8 bits, and 12 bits. In particular, we focus on the memory reuse distance, which describes the cache behavior of the application of interest. The reuse distances of all memory reads are reported in 20 buckets between different numbers of 64-byte cache blocks.

[0138]

[0139] Among them, Ai is the component matrix, λ1 is the eigenvalue, and Ui is the principal component loading matrix.

[0140] After calculation, the principal component loading matrix becomes 99 rows and 9 columns. Multiply it with the 60×99 standardized data matrix to obtain the principal component Yi. Subsequently, we use the ratio of the eigenvalues to calculate the comprehensive principal component score Y.

[0141]

[0142] We sorted the 99 features according to the comprehensive principal component scores. The higher the score, the stronger the explanatory power of the rendering program. Our dataset is a 60×99 data metric set, including 60 rendering benchmark tests and 99 microarchitecture-independent features. Obviously, this is a huge dataset. First, we standardized the data using statistical methods, which helps to compare and weigh metrics of different units or magnitudes. After that, we selected the principal component analysis method to reduce the 99 features to 9 features. In this example, the eigenvalues of the first 9 principal components are all greater than 1, and the cumulative contribution rate reaches 98.542%, indicating that these 9 factors have a high impact on the overall explanation rate. Therefore, the first 9 factors can be extracted. Now, we calculate the principal component scores of the 99 features. If a certain feature has a high comprehensive score, it means that the feature has a strong explanatory power for the overall, reflecting more overall features of the dataset. First, we use SPSS to calculate the 99×9 component matrix of the 60×99 data matrix, and then divide the component matrix by the square root of the eigenvalue to obtain the principal component loading matrix. Table 2 is a schematic table of CPU microarchitecture-independent features according to the embodiments of the present application, as shown in Table 2:

[0143] Table 2

[0144]

[0145]

[0146] We have quantitatively described program behavior according to a series of feature categories, as shown in Table 2: The instruction mix feature is measured by 12 percentages, representing the proportion of control flow operations among all dynamic instructions; the register traffic feature consists of 2 values and 7 probabilities, with the dependence distance selected as a power of 2 (i.e., 1, 2, 4, 8, 16, 32, 64), and the proportion of register dependence distances less than or equal to 16 gives the proportion from 1 to 16; the working set size feature is described by 4 numbers, and the data working set size (32-byte blocks) refers to the proportion of the number of blocks (64 bytes) and pages (4KB) in the instruction and data memory; the data flow stride feature contains 28 probabilities, with the stride characterized by a power of 8 (i.e., less than or equal to 0, 8, 64, 512, 4096, 32768, 262144), where the proportion of local store strides less than or equal to 4096 refers to the proportion of 0, 8, 64, 512, and 4096 in each static instruction write access, the proportion of global store strides less than or equal to 4096 refers to the proportion of these values in all instruction write accesses, and the proportion of global load strides less than or equal to 4096 is the proportion of these values in all instruction read accesses; the branch predictability feature is measured by 12 percentages, and PAg PPM represents the proportion of the global branch history at 4, 8, and 12-bit history lengths; the memory reuse distance feature contains 20 probabilities, representing the proportion of memory reuse distances less than or equal to 512, i.e., [0; 29] represents the proportion within the reuse distance range of [2n; 2n+1], where n ranges from 1 to 18; the instruction-level parallelism feature is described by 4 values, namely four different instruction window sizes (32, 64, 128, 256) measured under conditions such as assuming perfect caches and perfect branch prediction.

[0147] Figure 9 is a schematic diagram of an optional program running characteristic according to the present application, as Figure 9 shown, with the inherent behavior pattern of part of the CPU rendering program represented in blue and the inherent behavior pattern of the MiBench program represented in green. We have identified the top-ranked features and used a radar chart to visualize the inherent behavior of the benchmark tests. Figure 9Shows eight groups of features we selected based on the features with top-ranked principal component scores. We describe the meanings of these features in the table. Our observations are as follows: (1) Prediction by Partial Match predicts the next result based on historical lengths of 4 bits, 8 bits, and 12 bits and starts executing instructions in advance. The average PAg PPM (average page accesses of partial match prediction or a similar metric, the specific meaning needs to be determined according to the context) of the renderer is 8 times that of MiBench, indicating that the renderer has strong data dependencies. (2) The average data working set size (32-byte blocks) of MiBench is 4 times that of the renderer. (3) The average value of the global load stride less than or equal to 4096 in the renderer is 75.22%, while the average value of MiBench is only 55.89%. The global load stride is defined as the difference in data memory addresses between temporally adjacent memory read accesses. MiBench contains only nearly half of the global memory read accesses, while the renderer contains three-quarters of the global memory read accesses, indicating that the global load stride of the renderer is relatively concentrated.

[0148] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation.

[0149] The embodiment of the present application also provides an adjustment device for a renderer, Figure 10 is a structural block diagram of an adjustment device for a renderer according to an embodiment of the present application, as Figure 10 shown. The device includes:

[0150] An acquisition module 1002, configured to acquire target running information of an initial renderer by a target processor, where the target running information is used to indicate the usage of program running resources of the target processor during the running of the initial renderer;

[0151] A conversion module 1004, configured to convert the target program features of the initial renderer according to the target running information and reference running information, where the reference running information is used to indicate the response of program running resources when running a test program, the test program is used to test the resource performance of program running resources, and the target program features are program running features in the initial renderer that have an impact on rendering performance greater than the target impact level;

[0152] An adjustment module 1006, configured to adjust the initial renderer according to the target program features to obtain a target renderer for running on the target processor.

[0153] Through the above device, by obtaining the target running information of the target processor running the initial rendering program, the running performance of the initial rendering program on the target processor is characterized by the usage of the program running resources of the target processor during the running process of the initial rendering program. Furthermore, according to the target running information and the reference running information, the target program characteristics of the rendering program are converted, and the program running characteristics in the initial rendering program that have a greater impact on the rendering performance than the target impact degree by the program running resources of the current target processor are obtained. Then, by adjusting the initial rendering program according to the target program characteristics, the program optimization of the initial rendering program is realized according to the program running resources of the target processor, so that the adjusted target rendering program has a higher degree of adaptation to the program running resources of the target processor, which can solve the technical problem of low optimization efficiency of the running performance of the rendering program in the related technology and achieve the effect of improving the optimization efficiency of the running performance of the rendering program.

[0154] Optionally, the obtaining module includes:

[0155] A first obtaining unit, configured to obtain the microarchitecture characteristics of the target processor, where the microarchitecture characteristics are used to indicate the resource characteristics of the program running resources of the target processor;

[0156] A building unit, configured to build an initial rendering program for the target processor according to the microarchitecture characteristics;

[0157] A calling unit, configured to call the target processor to run the initial rendering program and collect the target running information during the running process of the initial rendering program.

[0158] Optionally, the calling unit is further configured to: configure multiple instruction windows for the initial rendering program, where each instruction window is used to indicate that the initial rendering program executes multiple rendering instructions in parallel, and the number of rendering instructions indicated by different instruction windows is different, and control the initial rendering program to execute the rendering instructions of multiple instruction windows respectively; collect the instruction execution information and memory access information after the initial rendering program responds to the instruction window, where the instruction execution information is used to indicate the execution situation of the initial rendering program for the parallel rendering instructions, and the memory access information is used to indicate the memory access situation of the rendering program to the target processor during the running process; determine the instruction execution information as the first running information of the initial rendering program, and determine the second running information of the initial rendering program according to the memory access information, where the target running information includes the first running information and the second running information;

[0159] Optionally, the calling unit is further configured to: determine a data score for each of the first quantity of first access data using a principal component analysis algorithm, where the first quantity of first access data is used to characterize the memory access situation of the target processor from different dimensions, the data score is used to indicate the importance degree of the first access data to the rendering performance, and the memory access information includes the first quantity of first access data; sort the first access data in descending order of the data scores to obtain a target sequence; screen out a second quantity of second access data with data scores greater than a target score from the target sequence, where the second quantity is less than the first quantity; and determine the second access data as second running information, where the target running information includes the first running information and the second running information.

[0160] Optionally, the adjustment module includes:

[0161] A second obtaining unit, configured to obtain program information corresponding to the target program feature in the initial rendering program;

[0162] A generating unit, configured to generate program adjustment information for the initial rendering program according to the program information, and use the program adjustment information to adjust the initial rendering program to obtain a target rendering program for running on the target processor.

[0163] Optionally, the generating unit is further configured to, when the target program feature is used to indicate that the running performance of the target rendering algorithm included in the initial rendering program is lower than a first threshold, identify calculation nodes in the initial rendering algorithm, and adjust the algorithm structure of the target rendering algorithm according to the node attributes of the calculation nodes to obtain a target rendering algorithm; when the target program feature is used to indicate that the parallel computing performance of the initial rendering program is lower than a second threshold, divide the rendering tasks to be executed by the initial rendering program into multiple subtasks; configure the binding relationship between the multiple subtasks and multiple processor cores on the target processor, where different subtasks are bound to different processor cores; when the target program feature is used to indicate that the memory access performance of the initial rendering program for the target processor is lower than a third threshold, obtain target data accessed by the initial rendering program, where the target data is data in the memory whose access frequency by the initial rendering program is greater than a target frequency; configure a target cache location for the target data in the cache space of the target processor; when the target program feature is used to indicate that the frequency of the control flow operations of the initial rendering program is greater than a fourth threshold, predict the target score trend of the initial rendering program after the current moment according to the reference score information between the current moments, and use the target score trend to adjust the initial rendering program.

[0164] Optionally, the conversion module includes:

[0165] An extraction unit, configured to extract candidate running information corresponding to the information type of the target running information from the reference running information;

[0166] A matching unit, configured to match candidate running information with target running information;

[0167] A determining unit, configured to determine the program running feature corresponding to the target running information as the target program feature when the running performance indicated by the candidate running information is greater than or equal to the running performance indicated by the target running information.

[0168] For the description of the features in the corresponding embodiment of the adjustment device of the rendering program, reference may be made to the relevant description in the corresponding embodiment of the adjustment method of the rendering program, which will not be elaborated here one by one.

[0169] An embodiment of the present application further provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above embodiments of the adjustment method of the rendering program.

[0170] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps in any one of the above embodiments of the adjustment method of the rendering program when running.

[0171] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: various media such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disc that can store a computer program.

[0172] An embodiment of the present application further provides a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, the steps in any one of the above embodiments of the adjustment method of the rendering program are implemented.

[0173] An embodiment of the present application further provides another computer program product, including a non-volatile computer-readable storage medium. The non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in any one of the above embodiments of the adjustment method of the rendering program are implemented.

[0174] Those skilled in the art may further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0175] The above has introduced in detail a method and device for adjusting a rendering program, an electronic device, a storage medium, and a computer program product provided by this application. Specific examples are used herein to elaborate on the principle and implementation manner of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application. It should be noted that for those of ordinary skill in the art of this technology, without departing from the principle of this application, several improvements and modifications can still be made to this application, and these improvements and modifications also fall within the protection scope of the claims of this application.

Claims

1. A method for adjusting a rendering program, characterized in that Including: Obtaining target running information of an initial rendering program by a target processor, where the target running information is used to indicate the usage of program running resources of the target processor during the running of the initial rendering program; Converting target program features of the initial rendering program according to the target running information and reference running information, where the reference running information is used to indicate the response of the program running resources when running a test program, the test program is used to test the resource performance of the program running resources, and the target program features are program running features in the initial rendering program whose influence on the rendering performance is greater than a target influence degree; Adjusting the initial rendering program according to the target program features to obtain a target rendering program for running on the target processor.

2. The method according to claim 1, wherein The obtaining of the target running information of the initial rendering program by the target processor includes: Obtaining microarchitecture features of the target processor, where the microarchitecture features are used to indicate resource features of the program running resources of the target processor; Constructing the initial rendering program for the target processor according to the microarchitecture features; Invoking the target processor to run the initial rendering program and collecting the target running information during the running of the initial rendering program.

3. The method according to claim 2, wherein The invoking of the target processor to run the initial rendering program and collecting the target running information during the running of the initial rendering program includes: Configuring a plurality of instruction windows for the initial rendering program, where each instruction window is used to indicate that the initial rendering program executes a plurality of rendering instructions in parallel, and the number of rendering instructions indicated to be executed by different instruction windows is different; Controlling the initial rendering program to execute the rendering instructions of the plurality of instruction windows respectively; Collecting instruction execution information and memory access information after the initial rendering program responds to the instruction windows, where the instruction execution information is used to indicate the execution situation of the initial rendering program for the parallel rendering instructions, and the memory access information is used to indicate the memory access situation of the target processor during the running of the rendering program; Determining the instruction execution information as the first running information of the initial rendering program, and determining the second running information of the initial rendering program according to the memory access information, where the target running information includes the first running information and the second running information.

4. The method according to claim 3, wherein The determining of the second running information of the initial rendering program according to the memory access information includes: Using a principal component analysis algorithm to determine data scores of each of the first access data of a first number of first access data, where the first number of first access data is used to characterize the memory access situation of the target processor from different dimensions, the data scores are used to indicate the importance degree of the first access data to the rendering performance, and the memory access information includes the first number of first access data; Sorting the first access data according to the order of the data scores from large to small to obtain a target sequence; Screen out a second quantity of second access data from the target sequence, where the data score of the second access data is greater than the target score, and the second quantity is less than the first quantity; Determine the second access data as the second running information.

5. The method according to claim 1, characterized in that The adjusting the initial rendering program according to the target program characteristics includes: Obtain the program information corresponding to the target program characteristics in the initial rendering program; Generate program adjustment information for the initial rendering program according to the program information, and use the program adjustment information to adjust the initial rendering program to obtain a target rendering program for running on the target processor.

6. The method according to claim 5, characterized in that, The generating program adjustment information for the initial rendering program according to the program information, and using the program adjustment information to adjust the initial rendering program to obtain a target rendering program for running on the target processor includes: When the target program characteristics are used to indicate that the running performance of the target rendering algorithm included in the initial rendering program is lower than a first threshold, identify the computing nodes in the initial rendering algorithm, and adjust the algorithm structure of the target rendering algorithm according to the node attributes of the computing nodes to obtain a target rendering algorithm; When the target program characteristics are used to indicate that the parallel computing performance of the initial rendering program is lower than a second threshold, divide the rendering tasks to be executed by the initial rendering program into multiple subtasks; configure the binding relationship between the multiple subtasks and multiple processor cores on the target processor, where different subtasks are bound to different processor cores; When the target program characteristics are used to indicate that the memory access performance of the initial rendering program for the target processor is lower than a third threshold, obtain the target data accessed by the initial rendering program, where the target data is data in the memory whose access frequency accessed by the initial rendering program is greater than the target frequency; configure a target cache location for the target data in the cache space of the target processor; When the target program characteristics are used to indicate that the frequency of the control flow operations of the initial rendering program is greater than a fourth threshold, predict the target score trend of the initial rendering program after the current moment according to the reference score information between the current moments, and use the target score trend to adjust the initial rendering program.

7. The method according to claim 1, wherein The converting the target program characteristics of the initial rendering program according to the target running information and the reference running information includes: Extract candidate running information corresponding to the information type of the target running information from the reference running information; Match the candidate running information with the target running information; When the running performance indicated by the candidate running information is greater than or equal to the running performance indicated by the target running information, determine the program running characteristics corresponding to the target running information as the target program characteristics.

8. An adjustment device for a rendering program, characterized in that, including: An obtaining module, configured to obtain target running information of a target processor for an initial rendering program, where the target running information is used to indicate the usage of program running resources of the target processor during the running of the initial rendering program; A conversion module, configured to convert target program features of the initial rendering program according to the target running information and the reference running information, where the reference running information is used to indicate the response of the program running resources when running a test program, the test program is used to test the resource performance of the program running resources, and the target program features are program running features in the initial rendering program whose influence on the rendering performance is greater than a target influence degree; An adjustment module, configured to adjust the initial rendering program according to the target program features to obtain a target rendering program for running on the target processor.

9. An electronic device, characterized in that, Comprising: A memory, configured to store a computer program; A processor, configured to implement the steps of the method for adjusting a rendering program according to any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, where the computer program, when executed by a processor, implements the steps of the method for adjusting a rendering program according to any one of claims 1 to 7.

Citation Information

Cited By

  • Method for determining influence factors of performance problem and electronic equipment

    CN122470483A